Methods and compositions for reducing GC bias
By circularizing single-stranded DNA templates using a helicase and ligase in the presence of a splint oligonucleotide, the method addresses GC bias and secondary structure issues, resulting in improved genome coverage and variant detection accuracy.
Patent Information
- Application Number
- PCT/US2024/061018
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
The presence of sequence-specific bias (GC bias) and secondary structures during library preparation and amplification leads to uneven genome coverage, resulting in data analysis issues such as gaps in sequences, shorter contigs, and inaccurate variant detection.
A method of circularizing a single-stranded DNA template by maintaining its linearity through helicase activity and ligating it on a splint oligonucleotide in a circularization mixture containing a ligase, thereby reducing GC bias.
The method improves sequencing coverage uniformity, reduces GC bias, and enhances the accuracy of variant detection by maintaining the linearity of single-stranded DNA templates during circularization.
Smart Images

Figure US2024061018_26062025_PF_FP_ABST
Abstract
Description
METHODS AND COMPOSITIONS FOR REDUCING GC BIASCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 613,028, filed December 20, 2023; the content of which is incorporated herein by reference in its entirety for all purposes.BACKGROUNDField
[0002] The present application generally relates to molecular biology and more specifically to amplification and sequencing of nucleic acids.Description of the Related Art
[0003] The presence of sequence specific bias (e.g., GC bias) and secondary structures during library preparation and amplification gives rise to difficulties determining the sequence of certain regions of a genome and non-uniformity in genome coverage. The resulting overrepresentation of certain regions and underrepresentation of other regions can translate into data analysis problems such as increases in the number of gaps in the analyzed sequence, a yield of shorter contigs, less accurate variant detection, and a need for increased coverage to sequence a genome. There is a need in the field to develop methods and compositions to reduce sequence coverage biases.SUMMARY
[0004] Disclosed herein include methods of circularizing a single-stranded DNA template. In some embodiments, a method of circularizing a single-stranded DNA template comprises maintaining linearity of a single-stranded DNA template through the activity of a helicase. The method can comprise circularizing the single-stranded DNA template by ligation on a splint oligonucleotide in a circularization mixture comprising a ligase for a time duration. The single-stranded DNA template can comprise a 5’ adapter and a 3’ adapter. The splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter. The complementarity can enable or allow the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template (or the complementarity can enable or allow the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide to enable ligation by the ligase to form a circularized DNA template). As a result of the complementarity,the 5’ adapter and the 3’ adapter can hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template (or as a result of the complementarity, the 5’ adapter and the 3’ adapter can hybridize to the splint oligonucleotide and enable ligation by the ligase to form a circularized DNA template).
[0005] In some embodiments, maintaining linearity of the single-stranded DNA template and circularizing the single-stranded DNA template are performed in a single step. In some embodiments, circularizing the single-stranded DNA template occurs in the presence of a single-stranded DNA binding protein bound to the single-stranded DNA template. Maintaining linearity of the single-stranded DNA template can be performed in the presence of the singlestranded DNA binding protein. The single-stranded DNA binding protein can bind to the singlestranded DNA template and facilitate association of the helicase with the single-stranded DNA template. In some embodiments, circularizing the single-stranded DNA template is performed in the presence of the single-stranded DNA binding protein. The single-stranded DNA binding protein can bind to the single-stranded DNA template. In some embodiments, the circularization mixture comprises the helicase. In some embodiments, maintaining linearity of the single-stranded DNA template comprises preventing and / or disrupting at least one of secondary structure, coiling, reannealing, and hybridization of adaptors across DNA templates. In some embodiments, the method comprises, prior to maintaining linearity of the single-stranded DNA template, denaturing a DNA template to form the single-stranded DNA template. The DNA template can be at least partially double-stranded. In some embodiments, the method is performed for a plurality of DNA templates each comprising a 5’ adapter and a 3’ adapter. At least a portion of each DNA template (e.g., the 5’ adapter or a portion thereof and / or the 3’ adapter or a portion thereof) is complementary to the splint oligonucleotide.
[0006] Disclosed herein include methods of circularizing a single-stranded DNA template. In some embodiments, a method of circularizing a single-stranded DNA template comprises: circularizing a single-stranded DNA template on a splint oligonucleotide in a circularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration. The single-stranded DNA template can comprise a 5 ’ adapter and a 3 ’ adapter. The splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter. The splint oligonucleotide can comprise a second portion complementary to at least a portion of the 3’ adapter. The complementarity can enable or allow the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase (or enable ligation by the ligase) to form a circularized DNA template. As a result of the complementarity, the 5’ adapter and the 3’ adapter can hybridize to the splint oligonucleotide and be ligated by the ligase (or to enable ligation by the ligase) to form a circularized DNA template.
[0007] In some embodiments, the method comprises: performing rolling circle amplification (RCA). RCA can be performed by contacting the circularized DNA template with a RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template.
[0008] Disclosed herein include methods of rolling circle amplification (RCA) for nucleic acids. In some embodiments, a method of RCA for nucleic acids comprises: (a) circularizing a single-stranded DNA template on a splint oligonucleotide in a circularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration. The single-stranded DNA template can comprise a 5 ’ adapter and a 3 ’ adapter. The splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter. The complementarity can enable or allow the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase (or to enable ligation by the ligase) to form a circularized DNA template. As a result of the complementarity, the 5’ adapter and the 3’ adapter can hybridize to the splint oligonucleotide and be ligated by the ligase (or to enable ligation by the ligase) to form a circularized DNA template. The method can comprise: (b) performing RCA by contacting the circularized DNA template with an RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template.
[0009] In some embodiments, the single-stranded DNA template has a high GC content or a high AT content. In some embodiments, the single-stranded DNA forms a secondary structure prior to being in contact with the helicase and / or the single-stranded DNA binding protein. The secondary structure can comprise a stem, a hairpin structure, a pseudoknot, a bulge, an internal loop, a multiple loop, or a combination thereof. In some embodiments, the singlestranded DNA template is a library constituent. In some embodiments, the single-stranded DNA template comprises a target nucleic acid positioned between the 5’ adapter and the 3’ adapter.
[0010] In some embodiments, the method comprises ligating the 5’ adapter and the 3’ adapter to a double-stranded DNA template. The method can comprise dehybridizing the doublestranded DNA template to form the single-stranded DNA template. In some embodiments, dehybridizing the double-stranded DNA template is performed in the presence of the helicase or prior to being in contact with the helicase. In some embodiments, dehybridizing the doublestranded DNA template comprises heat treatment, chemical treatment, enzymatic treatment by the helicase and the single-stranded DNA binding protein, or a combination thereof.
[0011] In some embodiments, the ligation is performed in solution or on a solid support. In some embodiments, one or more steps of the method is performed in solution or on a solid support. In some embodiments, the method is performed in solution or on a solid support. In some embodiments, circularizing the single-stranded DNA template is performed in solution oron a solid support. In some embodiments, the RCA is performed in solution or on a solid support. In some embodiments, the splint oligonucleotide is provided in solution. Maintaining linearity of the single-stranded DNA template and / or circularizing the single-stranded DNA template can be performed in solution. In some embodiments, the circularization and RCA are performed in solution. The splint oligonucleotide can function as a primer for the RCA. The method can comprise depositing the circularized DNA template on a surface. The method can comprise depositing the circularized DNA template on binding sites of a structured surface. In some embodiments, the circularization is performed in solution and the RCA is performed on a solid support. Contacting the circularized DNA template with the RCA mixture can be in the presence of a capture primer immobilized on the solid support. The method can comprise providing the solid support having the capture primer. In some embodiments, the method comprises hybridizing the circularized DNA template to the solid support or the capture primer. In some embodiments, the splint oligonucleotide is a capture primer immobilized on a solid support. Linearizing the DNA template can be performed on the solid support. Alternatively or additionally, circularizing the linear single-stranded DNA template can be performed on the solid support. Alternatively or additionally, performing RCA can be performed on the solid support. In some embodiments, performing RCA comprises extending the primer or the capture primer along the circularized DNA template.
[0012] In some embodiments, the time duration in the circularization can be (about) 5 minutes to (about) 2 hours, such as 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 minutes, or a number or a range between any two of these values.
[0013] In some embodiments, the helicase is at a concentration of (about) 0.005 nM to (about) 5 nM, such as (about) 0.005 nM to (about) 2 nM, such as (about) 1 nM to (about) 2 nM, or such as (about) 0.005 nM to (about) 0.01 nM. The helicase can be at a concentration of (about) 0.005 pM to (about) 5 pM, such as (about) 0.005 pM to (about) 2 pM, such as (about) 1 pM to (about) 2 pM, or such as (about) 0.005 pM to (about) 0.01 pM. In some embodiments, the helicase is Tte UvrD or T4 helicase (gp41).
[0014] In some embodiments, the single-stranded DNA binding protein is at a concentration of less than 1 pM, such as (about) 0.01 pM to (about) 1 pM, or such as (about) 100 nM, 125 nM, 150 nM, or 200 nM. In some embodiments, the single-stranded DNA binding protein is gp32.
[0015] In some embodiments, the ligase is at a concentration of less than 1 pM, such as (about) 0.01 pM to (about) 1 pM, optionally 100 nM, 125 nM, 150 nM, or 200 nM. In some embodiments, wherein the ligase is T4 DNA Ligase.
[0016] In some embodiments, maintaining linearity of the single-stranded DNA template is performed at a temperature from (about) 30 °C to (about) 50 °C, such as (about) 40 °C. In some embodiments, the circularizing is performed at a temperature from (about) 37 °C to (about) 60 °C.
[0017] In some embodiments, the single-stranded DNA template is about 100 bps to 500 bps in length, such as 100 bps, 150 bps, 200 bps, 250 bps, 300 bps, 350 bps, 400 bps, 450 bps, 500 bps, or a number of a range between any two of these values.
[0018] In some embodiments, the RCA is performed at about 37 °C. In some embodiments, the RCA mixture comprises a divalent metal cation, a DNA polymerase, and a dNTP mix. In some embodiments, the divalent metal cation is Ca2+, Mg2+, or Sr2+. In some embodiments, the divalent metal cation is Ca2+or Sr2+. In some embodiments, the RCA mixture does not comprise Mg2+. In some embodiments, the divalent metal cation is at a concentration from (about) 0.001 mM to (about) 1 M, such as from (about) 0.5 mM to (about) 200 mM. In some embodiments, the RCA mixture further comprises a polymer. In some embodiments, the polymer is linear or branched. In some embodiments, the polymer is a polyelectrolyte species, optionally the polyelectrolyte species is branched. In some embodiments, the branched polyelectrolyte species is a dendrimer species. In some embodiments, the dendrimer species is poly(amidoamine) (PAMAM) dendrimer. In some embodiments, the PAMAM dendrimer is a G1 PAMAM, G2 PAMAM, G3 PAMAM, G4 PAMAM, G5 PAMAM, or a combination thereof. In some embodiments the polymer is at a concentration from (about) 0.001 pM to (about) 1 M. In some embodiments, the divalent metal cation is Ca2+or Sr2+. The polymer can be at a concentration from about 1 mM to about 500 mM.
[0019] In some embodiments, performing RCA comprises introducing an amplification buffer into the vessel following contacting the circularized DNA template with the RCA mixture for, for example, (about) 10 minutes to (about) 60 minutes. The amplification buffer may not comprise any DNA polymerase. The amplification buffer can comprise a divalent metal cation and a polymer. The polymer can be branched polyelectrolyte species. In some embodiments, the divalent metal cation in the amplification buffer is Mg2+. In some embodiments, the divalent metal cation in the RCA mixture is Ca2+or Sr2+and the divalent metal cation in the amplification buffer is Mg2+. The RCA can be performed on the solid support. In some embodiments, the divalent metal cation in the amplification buffer is in a concentration of at least 10 mM, such as (about) 10 mM to (about) 10 M. In some embodiments, the branched polyelectrolyte in the amplification buffer is in a concentration of at least (about) 5 pM. In some embodiments, the method comprises removing the helicase, the single-stranded DNA binding protein, and / or the ligase from the vessel prior to introducing the amplification buffer.
[0020] In some embodiments, the method comprises removing the helicase, the singlestranded DNA binding protein, and / or the ligase prior to performing RCA. The helicase, the single-stranded DNA binding protein, and / or the ligase can be removed by washing excess helicase, single-stranded DNA binding protein, and / or ligase.
[0021] In some embodiments, the RCA is performed in the absence of the helicase, the single-stranded DNA binding protein, and / or the ligase, In some embodiments, wherein the RCA is performed in the presence of the helicase, the single-stranded DNA binding protein, and / or the ligase. In some embodiments, the method comprises adding helicase and / or the single-stranded DNA binding protein during the RCA. In some embodiments, the method comprises preparing a DNA library comprising a plurality of library constituents in the presence of the helicase and / or the single-stranded DNA binding protein, prior to maintaining linearity of the single-stranded DNA template and circularizing the single-stranded DNA template.
[0022] Disclosed herein include methods of rolling circle amplification (RCA) for nucleic acids. In some embodiments, a method of RCA for nucleic acids comprises contacting a circular DNA template and a capture primer with a RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template. The RCA mixture can comprise a DNA polymerase, a dNTP mix, a divalent metal cation, and a branched polyelectrolyte species. The divalent metal cation can be Ca2+or Sr2+. In some embodiments, the RCA mixture does not comprise Mg2+. In some embodiments, the divalent metal cation is at a concentration from (about) 0.001 mM to (about) 1 M, such as from (about) 0.5 mM to (about) 200 mM.
[0023] The branded polyelectrolyte species can be at a concentration from (about) 1 pM to (about) 1 M. In some embodiments, the branded polyelectrolyte species is at a concentration from (about) 1 mM to (about) 500 mM. In some embodiments, the branched polyelectrolyte species is a dendrimer species, optionally the dendrimer species is poly(amidoamine) (PAMAM) dendrimer. The PAMAM dendrimer can be a G1 PAMAM, G2 PAMAM, G3 PAMAM, G4 PAMAM, G5 PAMAM, or a combination thereof.
[0024] In some embodiments, the method comprises introducing an amplification buffer into the vessel following contacting the circular DNA template with the RCA mixture for about 10 minutes to about 60 minutes. The amplification buffer may not comprise any DNA polymerase. The amplification buffer can comprise a divalent metal cation and the branched polyelectrolyte species. In some embodiments, the divalent metal cation in the amplification buffer is Mg2+. In some embodiments, the divalent metal cation in the amplification buffer is in a concentration of from (about) 1 mM to (about) 10 M, such as from (about) 10 mM to (about) 5 M. In some embodiments, the branched polyelectrolyte in the amplification buffer is in aconcentration from (about) 10 mM to (about) 1 M, such as (about) 10 mM to (about) 500 mM. In some embodiments, the capture primer is immobilized on a solid surface.
[0025] In some embodiments, the vessel is a flow cell. In some embodiments, the method comprises sequencing the circularized DNA template by single molecule real-time sequencing. In some embodiments, the method comprises sequencing the amplified concatemers. The sequencing can comprise paired end sequencing. The sequencing can comprise SBB, SBS, single molecule real-time sequencing, or nanopore sequencing. The sequencing can comprise SBB or SBS and is performed on a flow cell. In some embodiments, the GC bias of reads obtained by the sequencing is less than 0.05. The GC bias of reads obtained by the sequencing can be less than 0.03. In some embodiments, the action of helicase and / or SSB reduces the GC bias to less than 75% of GC bias in the absence of helicase and SSB. The action of helicase and / or SSB can reduce the GC bias to less than 50% of GC bias in the absence of helicase and SSB. In some embodiments, the Fl score of the sequencing is greater than 0.99 for SNP and / or greater than for 0.98 indels. The Fl score of the sequencing can be greater than 0.993 for SNP and / or greater than for 0.987 indels. In some embodiments, the action of helicase and / or SSB maintains or increases at least one of average spot signal, uniformity of signal, and throughput.
[0026] Disclosed herein include kits of for rolling circle amplification. In some embodiments, a kit for rolling circle amplification comprises a first buffer comprising a divalent metal cation of Ca2+or Sr2+, a branched polyelectrolyte species, and a DNA polymerase. The kit can comprise a second buffer that does not comprise any DNA polymerase and comprises Mg2+and the branched polyelectrolyte species.
[0027] In some embodiments, the branched polyelectrolyte species in the first buffer is at a concentration from (about) 1 pM to (about) 1 M, such as at a concentration from (about) 1 mM to (about) 500 mM. In some embodiments, the branched polyelectrolyte species is a dendrimer species.
[0028] Disclosed herein include kits of for circularizing a single-stranded DNA template. In some embodiments, a kit for circularizing a single-stranded DNA template comprises: a helicase; a splint oligonucleotide; and a ligase. The kit can comprise: a single-stranded DNA binding protein.
[0029] Disclosed herein include methods of circularizing a single-stranded DNA template. In some embodiments, a method of circularizing a single-stranded DNA template comprises: maintaining linearity of a single-stranded DNA template in a circularization mixture through the activity of a helicase in the presence of a single-stranded DNA binding protein. The single-stranded DNA binding protein can bind to the single-stranded DNA template and facilitate association of the helicase with the single-stranded DNA template. The method can comprisecircularizing the linear single-stranded DNA template by ligation on a splint oligonucleotide in a ligase of the circularization mixture for a time duration to form a circularized DNA template. The single-stranded DNA template can comprise a 5’ adapter complimentary to a first portion of the splint oligonucleotide. The single-stranded DNA template can comprise a 3’ adapter complimentary to a second portion of the splint oligonucleotide. The complementarity can enable the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide to enable ligation by the ligase (or and be ligated by the ligase) to form the circularized DNA template. As a result of the complementarity, the 5’ adapter and the 3’ adapter can hybridize to the splint oligonucleotide to enable ligation by the ligase (or be ligated by the ligase) to form the circularized DNA template. Maintaining linearity of the single-stranded DNA template and circularizing the linear singlestranded DNA template can be performed in a single step.
[0030] Disclosed herein includes methods, systems and compositions of circularizing single-stranded nucleic acids and rolling circle amplifications for nucleic acids.BRIEF DESCRIPTION OF THE DRAWINGS
[0031] FIG. 1 is a non-limiting schematic illustration showing the circularization of a single-stranded nucleic acid. The diagram outlines keys steps involved converting ssDNA libraries with problematic secondary structures into unfolded molecules with improved capture and ring making efficiency.
[0032] FIG. 2A shows a non-limiting schematic illustration of producing a circular template nucleic acid hybridized to an immobilized nucleic acid primer (e.g. a capture primer). FIG. 2B shows a non-limiting schematic illustration of extending an immobilized capture primer along a circular template nucleic acid via rolling circle amplification to produce amplified concatemers.
[0033] FIG. 3 illustrates Integrative Genomics Viewer (IGV) windows showing two regions of the Pseudomonas aeruginosa genome with especially high %GC (A: 75% and B : 72%). Gray, red and blue bars indicate sequenced reads that are mapped to these genome regions. The upper panels reflect sequencing data collected from clustering conditions that include a ligation mixture containing Helicase and single-stranded DNA binding protein (SSB), while the lower panels reflect sequenced reads collected under the same conditions, but without the Helicase and SSB. The presence of Helicase and SSB during ssDNA library ligation significantly improved sequencing coverage for both high %GC loci.
[0034] FIG. 4 provides plots demonstrating that helicase and single-stranded DNA binding protein can reduce background signal. Upper panels - L01 / HS: circularization mixture with helicase (UvrD) and SSB (gp32) present and L02 / SOP: circularization mixture withouthelicase and SSB. Cycle 1 shows the raw sequencing images (Exam G) at the beginning of Read 1. Cycle 153 shows the raw sequencing images (Exam G) at the beginning of Read 2. The space between clusters (background) in L01 / HS panels on the left is darker than the L02 / SOP panels on the right for both Reads (Cycle 1 and 153). Lower panels - A histogram representation of the intensity along the yellow lines from the images above, confirming the background intensity is lower in the left L01 / HS panels than the L02 / SOP panels on the right.
[0035] FIG. 5A-B are non-limiting plots showing Picard GC bias plots for clustering conditions without (FIG. 5A) and with (FIG. 5B) helicase and SSB. FIG. 5C is a non-limiting table summarizing the sum difference between observed and expected base quality scores in both conditions. Picard GC bias plots show normalized sequencing coverage per %GC. The black distribution reflects the expected coverage distribution, while the red reflects the observed distribution. Blue circles show the normalized frequency of coverage, where a value =1 indicates a perfect correlation between observed and expected. Green line plots the observed base quality score for each GC bias bin. Figures were generated from HG002 PEI 50 sequencing data that used clustering conditions without (A) and with (B) Helicase and SSB.
[0036] FIGS. 6A-B are non-limiting Picard GC bias plots showing normalized sequencing coverage per GC percentage. The black distribution reflects the expected coverage distribution, while the red reflects the observed distribution. Blue circles show the normalized frequency of coverage, where a value =1 indicates a perfect correlation between observed and expected. Green line plots the observed base quality score for each GC bias bin. Picard figures (FIGS. 6 A and 6B) demonstrate that multiple reagent and condition optimizations based on the Helicase / SSB addition (Ml 1 condition shown in FIG. 6B) provides a significant improvement GC bias compared to no Helicase / SSB (M10 condition shown in FIG. 6A). Optimizations on the Helicase / SSB clustering are described in Table 1 and related description further herein. The Ml 1 condition (the optimized Helicase / SSB chemistry) significantly improves GC bias for HG002 (FIG. 6B). FIG. 6C provides Fl scores derived from Hap.py analysis of sequencing coverage SNP and Indel sites from the human genome. Fl scores are reported from M10 (FIG. 6A), Mi l (FIG. 6B) and publicly available data from third party sequencing platforms.DETAILED DESCRIPTION
[0037] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of thesubject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.
[0038] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.
[0039] Disclosed herein include methods, compositions, systems and kits for circularizing single-stranded nucleic acid templates and for performing rolling circle amplification (RCA) with the circularized single-stranded nucleic acid templates. The circularization and RCA methods disclosed herein can be performed in solution or on a surface (e.g., a structured surface).
[0040] Disclosed herein includes a method of circularizing a single-stranded DNA template. The method can comprise maintaining linearity of a single-stranded DNA template through the activity of a helicase and circularizing the single-stranded DNA template by ligation on a splint oligonucleotide in a circularization mixture comprising a ligase (e.g., T4 DNA ligase) for a time duration. The single-stranded DNA template can comprise a 5’ adapter and a 3’ adapter and the splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter, thereby allowing the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template. In some embodiments, maintaining linearity of the single-stranded DNA template and circularizing the single-stranded DNA template are performed in a single step. The ssDNA may be provided with a 5’ phosphate for ligation (e.g., through ligation to adaptors with a 5’ phosphate and / or amplification with primers having a 5’ phosphate). Alternatively, a 5’ phosphate may be added to the ssDNA by a kinase, such as a T4 Polynucleotide Kinase (T4 PNK) in solution. 5’ phosphate can be added by the kinase in the presence of the ligase in some embodiments.
[0041] Disclosed herein also includes a method of circularizing a single-stranded DNA template. The method can comprise circularizing a single-stranded DNA template on a splint oligonucleotide in a circularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration. The single-stranded DNA template can comprise a 5’ adapter and a 3’ adapter and the splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter, thereby allowing the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template.
[0042] Disclosed herein also includes a method rolling circle amplification (RCA) fornucleic acids. The method comprise (a) circularizing a single-stranded DNA template on a splint oligonucleotide in a circularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration, wherein the single-stranded DNA template comprises a 5’ adapter and a 3’ adapter and the splint oligonucleotide comprises a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter, thereby allowing the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template; and (b) performing RCA by contacting the circularized DNA template with a RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template.
[0043] Disclosed herein also includes rolling circle amplification (RCA) for nucleic acids. The method can comprise contacting a circular DNA template and a capture primer with a RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template, wherein the RCA mixture comprises a DNA polymerase, a dNTP mix, a divalent metal cation, and a branched polyelectrolyte species, wherein the branded polyelectrolyte species is at a concentration from about 1 pM to about 1 M and the divalent metal cation is Ca2+or Sr2+.
[0044] In some embodiments, the concatemer is a product of RCA with a strand displacing polymerase from a primer hybridized to a circular nucleic acid template. The circular nucleic acid template can comprise the target sequence and at least one adapter sequence. For example, the strand displacing polymerase may be Phi29 or a variant thereof. In further embodiments, the circular nucleic acid template is a product of circularization of a linear nucleic acid template comprising the target sequence and at least one adapter sequence. Circularization of a linear nucleic acid template can be performed using any method known in the art. In some embodiments, circularization of a linear nucleic acid template is performed using a ligase. Circularization of a linear nucleic acid template can be performed using any suitable ligase, such as, for example, T4 DNA ligase.
[0045] The methods, compositions, and systems can resolve secondary structures and sequence specific biases within individual library molecules, especially molecules comprising high or low GC percentage regions, maintain linearity of single-stranded nucleic acids by preventing and / or disrupting the formation of secondary structures, and ultimately reduce bias during ring template generation. In some embodiments, the methods, compositions and systems described herein can solve the problem of over-representing AT rich and underrepresenting GC rich regions in Whole Genome Sequencing (WGS), resulting in more uniform sequencing coverage across the board regardless of genome composition. That translates to less “dark”, “hard- to-sequence” regions and less gaps in the final assemblies and longer contigs.Definitions
[0046] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See e.g., Singleton et al., Dictionary of Microbiology and Molecular Biology 2nded., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.
[0047] As used herein, the term “nucleotide” can be used to refer to a native nucleotide or analog thereof. Examples include, but are not limited to, nucleotide triphosphates (NTPs) such as ribonucleotide triphosphates (rNTPs), deoxyribonucleotide triphosphates (dNTPs), or nonnatural analogs thereof such as dideoxyribonucleotide triphosphates (ddNTPs) or reversibly terminated nucleotide triphosphates (rtNTPs).
[0048] As used herein, the term “hybridizing” or “hybridize” refers to the pairing of substantially complementary or complementary nucleic acid sequences within two different molecules. Pairing can be achieved by any process in which a nucleic acid sequence joins with a substantially or fully complementary sequence through base pairing to form a hybridization complex. “Hybridizing” or “hybridize” can comprise denaturing the molecules to disrupt the intramolecular structure(s) (e.g., secondary structure(s)) in the molecule. In some embodiments, denaturing the molecules comprises heating a solution comprising the molecules to a temperature sufficient to disrupt the intramolecular structures of the molecules. In some instances, denaturing the molecules comprises adjusting the pH of a solution comprising the molecules to a pH sufficient to disrupt the intramolecular structures of the molecules. For purposes of hybridization, two nucleic acid sequences or segments of sequences are “substantially complementary” if at least 80% of their individual bases are complementary to one another. In some embodiments, a splint oligonucleotide sequence is not more than about 50% identical to one of the two polynucleotides (e.g., RNA fragments) to which it is designed to be complementary. The complementary portion of each sequence can be referred to herein as a ‘segment’, and the segments are substantially complementary if they have 80% or greater identity.
[0049] As used herein, the terms “complementarity” or “complementary” mean that a nucleic acid can form hydrogen bond(s) with another nucleic acid based on traditional Watson- Crick base paring rule. Complementarity can be complete or partial. Complete complementarity indicates that each and every nucleic acid base of one strand is capable of forming hydrogen bonds according to Watson-Crick canonical base pairing with a corresponding base in another, antiparallel nucleic acid sequence. Partial complementarity indicates that only a percentage of thecontiguous residues of a nucleic acid sequence can form Watson-Crick base pairing with the same number of contiguous residues in another, antiparallel nucleic acid sequence. “Substantially complementary” refers to a percentage of complementary of about, at least, or at least about 70%, 80%, 90%, 100% or a number or a range between any two of these values.
[0050] As used herein, the term “primer” refers to a nucleic acid having a sequence that binds to a nucleic acid at or near a template sequence. Generally, the primer binds in a configuration that allows replication of the template, for example, via polymerase extension of the primer. The primer can be a first portion of a nucleic acid molecule that binds to a second portion of the nucleic acid molecule, the first portion being a primer sequence and the second portion being a primer binding sequence (e.g., a hairpin primer). In some embodiments, the primer is a first nucleic acid molecule that binds to a second nucleic acid molecule having the template sequence. A primer can consist of DNA, RNA or analogs thereof. A primer can have an extendible 3’ end or a 3’ end that is blocked from primer extension.
[0051] As used herein, the term “polymerase” can be used to refer to a nucleic acid synthesizing enzyme, including but not limited to, DNA polymerase, RNA polymerase, reverse transcriptase, primase and transferase. Typically, the polymerase has one or more active sites at which nucleotide binding and / or catalysis of nucleotide polymerization may occur. The polymerase can catalyze the polymerization of nucleotides to the 3’ end of the first strand of the double stranded nucleic acid molecule. For example, a polymerase catalyzes the addition of a next correct nucleotide to the 3’ oxygen moiety of the first strand of the double stranded nucleic acid molecule via a phosphodiester bond, thereby covalently incorporating the nucleotide to the first strand of the double stranded nucleic acid molecule. In some embodiments, a polymerase need not be capable of nucleotide incorporation under one or more conditions used in a method set forth herein. For example, a mutant polymerase can be capable of forming a ternary complex but incapable of catalyzing nucleotide incorporation.
[0052] As used herein, a “vessel” is a container that functions to isolate one chemical process (e.g., a binding event; an incorporation reaction; etc.) from another, or to provide a space in which a chemical process can take place. Examples of vessels useful in connection with the disclosed technique include, but are not limited to, flow cells, wells of a multi-well plate; microscope slides; tubes (e.g., capillary tubes); droplets, vesicles, test tubes, trays, centrifuge tubes, features in an array, tubing, channels in a substrate etc.
[0053] As used herein, the term “circular,” when used in reference to a nucleic acid strand, means that the strand has no terminus (that is, the strand lacks a 3’ end and a 5’ end). Accordingly, the 3’ oxygen and the 5’ phosphate moi eties of every nucleotide monomer in a circular strand is covalently attached to an adjacent nucleotide monomer in the strand. A circularDNA strand can serve as a template for producing a concatemeric amplicon via rolling circle amplification (RCA), wherein each sequence unit of the concatemeric amplicon is the reverse complement of the circular nucleic acid strand. A circular nucleic acid can be double stranded. One or both strands in a double stranded nucleic acid can lack a 3’ end and a 5’ end. One strand in a double stranded nucleic acid can have a gap (absence of at least one nucleotide monomer relative to the other strand) or nick (absence of a phosphodiester bond between two nucleotide monomers), so long as the other strand is circular.
[0054] As used herein, the term “concatemer,” when used in reference to a nucleic acid molecule, means a continuous nucleic acid molecule that contains multiple copies of a common sequence linked in series. Similarly, the term “concatemer,” when used in reference to a nucleotide sequence, means a continuous nucleotide sequence that contains multiple copies of a common sequence in series. Each copy of the sequence can be referred to as a “sequence unit” of the concatemer. A sequence unit can have a length of at least 10 bases, 50 bases, 100 bases, 250 bases, 500 bases or more. A concatemer can include at least 2, 5, 10, 50, 100 or more sequence units. A sequence unit can include subregions having any of a variety of functions such as a primer binding region, target sequence region, tag region, unique molecular identifier (UMI), or the like.
[0055] As used herein, a “flow cell” is a reaction chamber that includes one or more channels that direct fluid in a predetermined manner to conduct a desired reaction. The flow cell can be coupled to a detector such that a reaction occurring in the reaction chamber can be observed. For example, a flow cell can contain primed template nucleic acid molecules, for example, tethered to a solid support, to which nucleotides and ancillary reagents are iteratively applied and washed away. The flow cell can include a transparent material that permits the sample to be imaged after a desired reaction occurs. For example, a flow cell can include a glass slide containing small fluidic channels, through which polymerases, dNTPs and buffers can be pumped. The glass inside the channels is decorated with one or more primed template nucleic acid molecules to be sequenced. An external imaging system can be positioned to detect the molecules on the surface of the glass. Reagent exchange in a flow cell is accomplished by pumping, drawing, or otherwise “flowing” different liquid reagents through the flow cell. Exemplary flow cells, methods for their manufacture and methods for their use are described in U.S. Patent Application Publications US2010 / 0111768 or US2012 / 0270305; or WO 05 / 065814, each of which is incorporated by reference herein.
[0056] As used herein, the term “immobilized,” when used in reference to a molecule, refers to direct or indirect, covalent or non-covalent attachment of the molecule to a surface such as a surface of a solid support. In some configurations, covalent attachment is preferred, butgenerally all that is required is that the molecules (e.g., nucleic acids) remain immobilized or attached to the surface under the conditions in which surface retention is intended.
[0057] As used herein, the term “cluster,” when used in reference to nucleic acids, refers to a population of nucleic acids. The population of nucleic acids can be attached to a solid support, for example, at a binding area in an array of binding areas on the solid support. Alternatively, the population of nucleic acids can be in solution and not attached to any solid support.Circularization of Nucleic Acids
[0058] Provided herein include a method of circularizing a single-stranded DNA template. In some embodiments, the method can comprise maintaining linearity of a singlestranded nucleic acid template (e.g., DNA template) through the activity of a helicase. Maintaining linearity of a single-stranded DNA template can comprise linearizing a single-stranded nucleic acid template (e.g., DNA template) in the presence of a helicase by removing or disrupting at least one secondary structure in the single-stranded nucleic acid template to produce a linear singlestranded nucleic acid template. In some embodiments, maintaining linearity of a single-stranded nucleic acid template can comprise preventing the formation of at least one secondary structure, coiling, reannealing, and hybridization of adaptors across nucleic acid templates. The method can further comprise circularizing the linear single-stranded nucleic acid template (e.g., DNA template) on a splint oligonucleotide in a circularization mixture comprising a ligase for a time duration to form a circularized DNA template. The single-stranded DNA template can comprise a 5’ adapter and a 3’ adapter and the splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter. The sequence complementarity between portions of the adapters and the splint oligonucleotide allows the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide, and subsequently be ligated by the ligase to form a circularized DNA template. The steps described herein, including maintaining linearity of the single-stranded nucleic acid template and circularizing the linear single-stranded nucleic acid template, can be carried out sequentially or concurrently in a single step. The method can take place in solution or on a solid surface.
[0059] FIG. 1 is a non-limiting schematic illustration showing the circularization of a single-stranded nucleic acid on a solid surface. Nucleic acid template can form secondary structures due to inter-molecular complementarity, such as a stem, a hairpin structure, a pseudoknot, a bulge, an internal loop, a multiple loop, or other secondary structures identifiable to a person skilled in the art. Helicase can resolve the secondary structures by separating strandsof a self-annealed nucleic acid molecule. The diagram shown in FIG. 1 outlines key steps involved converting single-stranded DNA (ssDNA) libraries with problematic secondary structures into unfolded molecules with improved capture and ring making efficiency. Surface bound capture / splint probes are used to hybridize one or both adapter ends of a ssDNA library molecule. In the event that the DNA strand contains a significant amount of secondary structure, the efficiency of ligating both DNA ends on a splint is reduced. The addition of helicase and optionally single-stranded DNA binding protein provides a mechanism for unfolding the strand to enable more efficient capture and ligation of library ends.
[0060] In some embodiments, maintaining linearity of the single-stranded nucleic acid template and / or circularizing the linear single-stranded nucleic acid template can be carried out in the presence of a single-stranded DNA binding protein. The single-stranded DNA binding protein can be provided at any time during the process. For example, the single-stranded DNA binding protein can be added during the linearization step. Alternatively or additionally, the singlestranded DNA binding protein can be added in the circularization mixture. During the linearization process, the single-stranded DNA binding protein can bind to the single-stranded DNA template and facilitates the association of the helicase with the single-stranded DNA template. During the circularization process, the single-stranded DNA binding protein can bind to the linear singlestranded DNA template and aid in maintaining the linearity of the single-stranded DNA template by stabilizing and protecting the separated single strand of the nucleic acid template. The singlestranded DNA binding protein can reduce the speed of re-annealing / re-coiling, allowing more efficient circularization. In some embodiments, the circularization mixture can further comprise a helicase and a single-stranded DNA binding protein. In some embodiments, the circularization is carried out in the absence of a helicase and a single-stranded DNA binding protein. The method can further comprise removing the helicase and the single-stranded DNA binding protein following the linearization step and prior to the circularization step. In some embodiments, the method can further comprise denaturing a DNA template to from the single-stranded DNA template, wherein the DNA template is double-stranded, partially double-stranded or contains one or more secondary structures, using for example heat treatment, chemical treatment, enzymatic treatment, or a combination thereof. The circularization method described herein can increase the representation of problematic library molecule such as library constituents with secondary structures such as hairpin or G-quad structures and / or with high GC or AT content.
[0061] In some embodiments, maintaining linearity of the single-stranded DNA template and circularizing the linear single-stranded DNA template are performed in a single step. Accordingly, in some embodiments, the method of circularizing a single-stranded DNA template can comprise circularizing a single-stranded DNA template on a splint oligonucleotide in acircularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration. The single-stranded DNA template can comprise a 5’ adapter and a 3’ adapter and the splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter. The sequence complementarity between portions of the adapters and the splint oligonucleotide allows the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide, and subsequently be ligated by the ligase to form a circularized DNA template.
[0062] As used herein, the term “helicase” refers to a class of conserved enzymes that hydrolyze adenosine triphosphate (ATP) in the presence of nucleic acids. These enzymes can translocate along nucleic acid strand unwind and separate the helical structure of double-stranded nucleic acid or secondary structure formed internally within a single-stranded nucleic acid by dissociating the hydrogen bonds between the bases. A helicase can be a DNA helicase or an RNA helicase. A helicase typically possesses sequence motifs located in the interior of their primary structure, involved in ATP binding, ATP hydrolysis and translocation along the nucleic acid substrate. In some embodiments, a putative helicase can be designated based on the presence of some or all of seven conserved amino acid motifs (see, for example Figure 1 of Marians KJ “Volume 5, Issue 9, 15 September 1997, Pages 1129-1134”, the content of which is incorporated herein by reference). The helicase genes are generally grouped into six super families based on their shared sequence motifs: SF1, SF2, SF3, SF4, SF5 and SF6. The helicase used herein can be a member of any one of SF1-SF6 families. In some embodiments, the helicase used herein belongs to SF1 family. Examples of SF1 helicase include, for example, PcrA helicase from gram-positive bacteria, Rep and UvrD from gram-negative bacteria, RecD and Dda helicases, bacteriophage T4 helicase (gp 1) and others identifiable by a person skilled in the art.
[0063] As used herein, the term “single-stranded DNA binding protein” or “SSB” refers to a class of proteins that can bind to ssDNA with high affinity in a sequence-independent manner to maintain the transient unwinding of duplex DNA in a single-stranded state. In general, there are four main structural folds in ssDNA binding domains: oligonucleotide / oligosaccharide- binding (OB) folds, K homology (KH) domains, RNA recognition motifs (RRMs), and whirly domains [2], The OB fold is a well-known ssDNA binding domain found in many SSBs, including A. coli SSB, bacteriophage T4 SSB, human RPA, human mtSSB, human SSB1 and SSB2, and the telomere-end protection family, such as TEBP from Oxytricha nova, Cdcl3 from Saccharomyces cerevisiae, and Poti from humans. SSB can interact with single-stranded DNA through hydrogen bonds, stacking, or electrostatic interactions. The interactions between SSB and ssDNA typically occur through the OB fold. The OB fold consists of a five-stranded antiparallel P-sheet that forms a characteristic P-barrel core with various lengths of loopsconnecting the strands. The number and organization of the OB domains that participate in ssDNA interaction can vary. For example, RPA is a heterotrimer with six OB folds, two for subunit interaction and four for ssDNA binding. E. coli SSB is a homotetramer with each unit having one OB domain. More detailed structural and functional information about SSB can be found, for example, in Guo et al. (Biomolecules. 2022 Sep; 12(9): 1187), the content of which is incorporated herein by reference. SSB proteins have been identified in many different organisms. The SSB used herein can be any suitable SSB known in the art. For example, the SSB used herein can be E. coli SSB, bacteriophage T4 SSB, Replication protein A (RPA), human RPA, hSSBl, hSSB2, and others identifiable to a person skilled in the art. In some embodiments, the SSB used herein is T4 gp32. Alternatively or additionally, the SSB used herein can be a DNA binding protein predicted from uncharacterized protein using machine learning-based methods based on the structural features and capable of maintain the transient unwinding of duplex DNA in a singlestranded state as will be understood to a person skilled in the art.
[0064] A splint oligonucleotide described herein can hybridize to a nucleic acid template via a portion that is complementary to the nucleic acid template. For example, a nucleic acid template can comprise a target nucleic acid and a primer binding site that is complementary to a portion of the splint oligonucleotide. In some embodiments, a splint oligonucleotide can hybridize to portions of a nucleic acid template that are present at opposite ends of a target nucleic acid. For example, the nucleic acid template can comprise a 5’ adapter and a 3’ adapter which can hybridize to a first portion and a second portion of the splint oligonucleotide, respectively. The splint can bring together the two ends of the target nucleic acid. The two ends can be ligated while hybridized to a splint nucleic acid to form a circular version of the target nucleic acid.
[0065] In some embodiments, the method can further comprise preparing a singlestranded nucleic acid template from a double-stranded nucleic acid template (e.g., dsDNA template). The method can comprise providing a double-stranded nucleic acid template, attaching the 5’ adapter and the 3’ adapter to the ends of the double-stranded nucleic acid template, and dehybridizing or denaturing the double-stranded nucleic acid template to form a single-stranded nucleic acid template. Denaturing the double-stranded nucleic acid template can be carried out using heat or chemical treatment (e.g., a denaturing buffer) as will be understood by a person skilled in the art. Suitable denaturing buffers are well known in the art. For example, alterations in pH and low ionic strength solutions can denature nucleic acids. Formamide and urea can be used for denaturation. Double-stranded nucleic acid molecules can be dehybridized by treatment with a solution of very low salt (e.g., less than 0.1 mM cationic conditions) and high pH (> 12) or by using chaotropic salt (e.g., guanidinium hydrochloride). In some embodiments, a base may be used, such as a basic chemical compound hat is able to deprotonate weak acids in an acid basereaction. Alternatively or additionally, dehybridization of the double-stranded nucleic acid template can occur in the presence of a helicase alone or in combination with a single-stranded DNA binding protein. Accordingly, dehybridization of the double-stranded DNA template can be performed prior to being in contact with a helicase or in the presence of the helicase (e.g., during the linearization step).
[0066] The method can take place in solution or on a solid surface. For example, linearizing the single-stranded nucleic acid template, maintaining linearity of the single-stranded nucleic acid template, and circularizing the linear single-stranded nucleic acid template can be carried out in a solution. The splint oligonucleotide and the nucleic acid template can be provided in a solution and then delivered to a vessel by any suitable means. The nucleic acid template can hybridize to the splint oligonucleotide in the solution. Following the hybridization, linearization and circularization can take place in the solution in the presence of a helicase and ligase. SSB can be added during linearization or circularization.
[0067] Alternatively, linearization of the single-stranded nucleic acid template and circularization of the linear single-stranded nucleic acid template can be carried out on a solid surface. For example, the splint oligonucleotide can be attached to a solid support, for example, using covalent or non-covalent attachment chemistries known in the art, prior to being in contact with the template nucleic acid and / or the circularization mixture (see, for example, FIG. 1). As used herein, the term “solid support” refers to a rigid substrate that is substantially insoluble in liquids that it contacts. The substrate can be non-porous or porous. The substrate can optionally be capable of taking up a liquid (e.g., due to porosity) but will typically be sufficiently rigid that the substrate does not swell substantially when taking up the liquid and does not contract substantially when the liquid is removed by drying. A nonporous solid support is generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, cyclic olefins, polyimides etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, optical fiber bundles, and polymers. Any of a variety of liquids, including but not limited to those set forth herein, can be contacted with a solid support.
[0068] The temperature and the length of incubation time for the linearizing and / or circularizing steps can vary in different embodiments, for example, depending on the reaction conditions, concentration of the reagents and the like. In some embodiments, the time duration of linearizing and / or circularizing the single-stranded DNA template can be about 5 minutes to 2 hours. For example, the time duration can be, be about, be at most, be at most about, be at least,be at least about, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120 minutes, or a number or a range between any two of these values. In some embodiments, the time duration can be about 20 minutes, 40 minutes, or 60 minutes. The ligase, the helicase, and optionally the single-stranded DNA binding protein can be incubated with the nucleic acid template and the splint oligonucleotide at any temperature conducive to the enzyme activity. Linearization of a singlestranded nucleic acid template and circularization of a linear single-stranded nucleic acid template can be performed at a same temperature or different temperatures.
[0069] In some embodiments, maintaining linearity of a single-stranded DNA template, linearizing a single-stranded DNA template, hybridizing a single-stranded DNA template to a splint oligonucleotide or a capture primer (which can be part of or separate from circularizing a linear single-stranded DNA template), and / or circularizing a linear single-stranded DNA template are performed between 20 °C and 70 °C, for example 20 °C, 25 °C, 30 °C, 35 °C, 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, 65 °C, 70 °C, or a number or a range between any two of these values. In some embodiments, maintaining linearity of the single-stranded DNA template and / or linearizing the single-stranded DNA template is performed at a temperature from about 30 °C to about 50 °C, for example 30 °C, 31 °C, 32 °C, 33 °C, 34 °C, 35 °C, 36 °C, 37 °C, 38 °C, 39 °C, 40 °C, 41 °C, 42 °C, 43 °C, 44 °C, 45 °C, 46 °C, 47 °C, 48 °C, 49 °C, 50 °C, or a number or a range between any two of these values. In some embodiments, maintaining linearity of the singlestranded DNA template and / or linearizing the single-stranded DNA template is performed at a temperature 40 °C. Circularization of a linear single-stranded DNA template can be performed at a temperature between 30°C and 60°C, such as 30 °C, 31 °C, 32 °C, 33 °C, 34 °C, 35 °C, 36 °C, 37 °C, 38 °C, 39 °C, 40 °C, 41 °C, 42 °C, 43 °C, 44 °C, 45 °C, 46 °C, 47 °C, 48 °C, 49 °C, 50 °C, 51 °C, 52 °C, 53 °C, 54 °C, 55 °C, 56 °C, 57 °C, 58 °C, 59 °C, 60 °C, or a number or a range between any two of these values. For example, the circularization temperature can be, or be about, 37°C. In some embodiments, the circularization temperature can be, or be about, 55°C. The linear single-stranded DNA template can hybridize to a splint oligonucleotide or a capture primer (e.g., prior to ligation) at a temperature between 30°C and 60°C, such as 30 °C, 31 °C, 32 °C, 33 °C, 34 °C, 35 °C, 36 °C, 37 °C, 38 °C, 39 °C, 40 °C, 41 °C, 42 °C, 43 °C, 44 °C, 45 °C, 46 °C, 47 °C, 48 °C, 49 °C, 50 °C, 51 °C, 52 °C, 53 °C, 54 °C, 55 °C, 56 °C, 57 °C, 58 °C, 59 °C, 60 °C, or a number or a range between any two of these values. For example, the linear single-stranded DNA template can hybridize to a splint oligonucleotide or a capture primer at, or at about, 37°C. In some embodiments, the linear single-stranded DNA template can hybridize to a splint oligonucleotide or a capture primer at, or at about, 55°C. In some embodiments, maintaininglinearity of a single-stranded DNA template, linearizing a single-stranded DNA template, hybridizing a single-stranded DNA template to a splint oligonucleotide or a capture primer (which can be part of or separate from circularizing a linear single-stranded DNA template), and / or circularizing a linear single-stranded DNA template are performed at a substantially isothermal temperature, for example, a temperature that does not vary more than by about 2-3 °C above or below a given temperature.
[0070] The helicase can be present at a concentration to facilitate the linearization of the single-stranded nucleic acid template as will be apparent to a skilled person. The concentration can vary in different embodiments, depending for example on the helicase activity. In some embodiments, the concentration of the helicase can be in a range of about 0.005 nM to about 5 nM. For example, the concentration of the helicase can be, be about, be at least, be at least about, be at most, or be at most about 0.005 nM, 0.006 nM, 0.007 nM, 0.008 nM, 0.009 nM, 0.01 nM, 0.02 nM, 0.03 nM, 0.04 nM, 0.05 nM, 0.06 nM, 0.07 nM, 0.08 nM, 0.09 nM, 0.1 nM, 0.5 nM, 1.0 nM, 1.5 nM, 2.0 nM, 2.5 nM, 3.0 nM, 3.5 nM, 4.0 nM, 4.5 nM, 5.0 nM, or a number or a range between any two of these values. In some embodiments, the helicase is in a range of about 0.005 nM to about 0.01 nM, optionally in a range of about 1 nM to about 2 nM. In some embodiments, the concentration of the helicase can be in a range of about 0.005 nM to about 5 pM. For example, the concentration of the helicase can be, be about, be at least, be at least about, be at most, or be at most about 0.005 pM, 0.006 pM, 0.007 pM, 0.008 pM, 0.009 pM, 0.01 pM, 0.02 pM, 0.03 pM, 0.04 pM, 0.05 pM, 0.06 pM, 0.07 pM, 0.08 pM, 0.09 pM, 0.1 pM, 0.5 pM, 1.0 pM, 1.5 pM, 2.0 pM, 2.5 pM, 3.0 pM, 3.5 pM, 4.0 pM, 4.5 pM, 5.0 pM, or a number or a range between any two of these values. In some embodiments, the concentration of the helicase is in a range of about 0.005 pM to about 0.01 pM, optionally in a range of about 1 pM to about 2 pM. In some embodiments, the helicase can be in a range of about 0.001 pg / mL to about 8 pg / mL. For example, the concentration of the helicase can be, be about, be at least, be at least about, be at most, or be at most about, 0.001 pg / mL, 0.002 pg / mL, 0.003 pg / mL, 0.004 pg / mL, 0.005 pg / mL, 0.006 pg / mL, 0.007 pg / mL, 0.008 pg / mL, 0.009 pg / mL, 0.010 pg / mL, 0.015 pg / mL, 0.020 pg / mL, 0.025 pg / mL, 0.030 pg / mL, 0.035 pg / mL, 0.040 pg / mL, 0.045 pg / mL, 0.050 pg / mL, 0.055 pg / mL, 0.060 pg / mL, 0.065 pg / mL, 0.070 pg / mL, 0.075 pg / mL, 0.080 pg / mL, 0.085 pg / mL, 0.090 pg / mL, 0.095 pg / mL, 0.10 pg / mL, 0.11 pg / mL, 0.12 pg / mL, 0.13 pg / mL, 0.14 pg / mL, 0.15 pg / mL, 0.16 pg / mL, 0.17 pg / mL, 0.18 pg / mL, 0.19 pg / mL, 0.2 pg / mL, 0.3 pg / mL, 0.4 pg / mL, 0.5 pg / mL, 0.6 pg / mL, 0.7 pg / mL, 0.8 pg / mL, 0.9 pg / mL, 1 pg / mL, 1.1 pg / mL, 1.2 pg / mL, 1.3 pg / mL, 1.4 pg / mL, 1.5 pg / mL, 1.6 pg / mL, 1.7 pg / mL, 1.8 pg / mL, 1.9 pg / mL, 2 pg / mL, 3 pg / mL, 4 pg / mL, 5 pg / mL, 6 pg / mL, 7 pg / mL, 8 pg / mL, or a number or a range between any two of these values. For example, for highly active helicase such as T4 DNA helicase (gp41), the concentration can be higher. For less active helicase such as Tte UvrD, theconcentration can be higher. Helicase activity can be evaluated using a helicase assay by, for example, incubating a helicase with ATP and a radio-labeled DNA duplex, terminating the reaction, and analyzing the products using non-denaturing polyacrylamide gel electrophoresis as will be understood by a person skilled in the art.
[0071] The single-stranded DNA binding protein can be present at a concentration to facilitate the linearization and / or circularization of the single-stranded nucleic acid template as will be apparent to a skilled person. The single-stranded DNA binding protein can be at a concentration of less than 1 pM. In some embodiments, the concentration of the single-stranded DNA binding protein is in a range from about 0.01 pM to about 1 pM. For example, the concentration of the single-stranded DNA binding protein can be about, at least, at least about, at most, or at most about 0.01 pM, 0.1 pM, 0.2 pM, 0.3 pM, 0.4 pM, 0.5 pM, 0.6 pM, 0.7 pM, 0.8 pM, 0.9 pM, 1.0 pM, or a number or a range between any two of these values. In some embodiments, the concentration of the single-stranded DNA binding protein is in a range from about 50 nM to about 200 nM, optionally about 50 nM, 55 nM, 60 nM, 65 nM, 70 nM, 75 nM, 80 nM, 85 nM, 90 nM, 95 nM, 100 nM, 105 nM, 110 nM, 115 nM, 120 nM, 125 nM, 130 nM, 135 nM, 140 nM, 145 nM, 150 nM, 155 nM, 160 nM, 165 nM, 170 nM, 175 nM, 180 nM, 185 nM, 190 nM, 195 nM, 200 nM, or a number or a range between any two of these values. In some embodiments, the concentration of the single-stranded DNA binding protein can be in a range of about 0.1 pg / mL to about 40 pg / mL. For example, the concentration of the single-stranded DNA binding protein can be, be about, be at least, be at least about, be at most, or be at most about, 0.1 pg / mL, 0.2 pg / mL, 0.3 pg / mL, 0.4 pg / mL, 0.5 pg / mL, 0.6 pg / mL, 0.7 pg / mL, 0.8 pg / mL, 0.9 pg / mL, 1 pg / mL, 1.1 pg / mL, 1.2 pg / mL, 1.3 pg / mL, 1.4 pg / mL, 1.5 pg / mL, 1.6 pg / mL, 1.7 pg / mL, 1.8 pg / mL, 1.9 pg / mL, 2.0 pg / mL, 2.1 pg / mL, 2.2 pg / mL, 2.3 pg / mL, 2.4 pg / mL, 2.5 pg / mL, 2.6 pg / mL, 2.7 pg / mL, 2.8 pg / mL, 2.9 pg / mL, 3.0 pg / mL, 3.1 pg / mL, 3.2 pg / mL, 3.3 pg / mL, 3.4 pg / mL, 3.5 pg / mL, 3.6 pg / mL, 3.7 pg / mL, 3.8 pg / mL, 3.9 pg / mL, 4.0 pg / mL, 4.1 pg / mL, 4.2 pg / mL, 4.3 pg / mL, 4.4 pg / mL, 4.5 pg / mL, 4.6 pg / mL, 4.7 pg / mL, 4.8 pg / mL, 4.9 pg / mL, 5.0 pg / mL, 5.1 pg / mL, 5.2 pg / mL, 5.3 pg / mL, 5.4 pg / mL, 5.5 pg / mL, 5.6 pg / mL, 5.7 pg / mL, 5.8 pg / mL, 5.9 pg / mL, 6.0 pg / mL, 6.5 pg / mL, 7.0 pg / mL, 7.5 pg / mL, 8.0 pg / mL, 8.5 pg / mL, 9.0 pg / mL, 9.5 pg / mL, 10 pg / mL, 11 pg / mL, 12 pg / mL, 13 pg / mL, 14 pg / mL, 15 pg / mL, 16 pg / mL, 17 pg / mL, 18 pg / mL, 19 pg / mL, 20 pg / mL, 25 pg / mL, 30 pg / mL, 35 pg / mL, 40 pg / mL, or a number or a range between any two of these values.
[0072] In an exemplary embodiment, a method of circularizing a single-stranded DNA template comprises maintaining linearity of a single-stranded DNA template in a circularization mixture through the activity of a helicase in the presence of a single-stranded DNA binding protein, wherein the single-stranded DNA binding protein binds to the single-stranded DNA template and facilitates association of the helicase with the single-stranded DNA template, and circularizing the linear single-stranded DNA template by ligation on a splint oligonucleotidein a ligase of the circularization mixture for a time duration to form a circularized DNA template, wherein the single-stranded DNA template comprises a 5’ adapter complimentary to a first portion of the splint oligonucleotide and a 3’ adapter complimentary to a second portion of the splint oligonucleotide, wherein the 5’ adapter and the 3’ adapter hybridize to the splint oligonucleotide and be ligated by the ligase to form the circularized DNA template. Maintaining linearity of the single-stranded DNA template and circularizing the linear single-stranded DNA template are performed in a single step.
[0073] The ligase can be present at a concentration to facilitate the circularization of the single-stranded nucleic acid template as will be apparent to a skilled person. The ligase can be at a concentration of less than 1 pM. In some embodiments, the concentration of the ligase is in a range from about 0.01 pM to about 1 pM. For example, the concentration of the ligase can be about, at least, at least about, at most, or at most about 0.01 pM, 0.1 pM, 0.2 pM, 0.3 pM, 0.4 pM, 0.5 pM, 0.6 pM, 0.7 pM, 0.8 pM, 0.9 pM, 1.0 pM, or a number or a range between any two of these values. In some embodiments, the concentration of the ligase is in a range from about 50 nM to about 200 nM, optionally about 50 nM, 55 nM, 60 nM, 65 nM, 70 nM, 75 nM, 80 nM, 85 nM, 90 nM, 95 nM, 100 nM, 105 nM, 110 nM, 115 nM, 120 nM, 125 nM, 130 nM, 135 nM, 140 nM, 145 nM, 150 nM, 155 nM, 160 nM, 165 nM, 170 nM, 175 nM, 180 nM, 185 nM, 190 nM, 195 nM, 200 nM, or a number or a range between any two of these values. In some embodiments, the concentration of the ligase can be in a range of about 0.1 U / mL to about 40 U / mL. For example, the concentration of the ligase can be, be about, be at least, be at least about, be at most, or be at most about, 0.1 U / mL, 0.2 U / mL, 0.3 U / mL, 0.4 U / mL, 0.5 U / mL, 0.6 U / mL, 0.7 U / mL, 0.8 U / mL, 0.9 U / mL, 1 U / mL, 1.1 U / mL, 1.2 U / mL, 1.3 U / mL, 1.4 U / mL, 1.5 U / mL, 1.6 U / mL, 1.7 U / mL, 1.8 U / mL, 1.9 U / mL, 2.0 U / mL, 2.1 U / mL, 2.2 U / mL, 2.3 U / mL, 2.4 U / mL, 2.5 U / mL, 2.6 U / mL, 2.7 U / mL, 2.8 U / mL, 2.9 U / mL, 3.0 U / mL, 3.1 U / mL, 3.2 U / mL, 3.3 U / mL, 3.4 U / mL, 3.5 U / mL, 3.6 U / mL, 3.7 U / mL, 3.8 U / mL, 3.9 U / mL, 4.0 U / mL, 4.1 U / mL, 4.2 U / mL, 4.3 U / mL, 4.4 U / mL, 4.5 U / mL, 4.6 U / mL, 4.7 U / mL, 4.8 U / mL, 4.9 U / mL, 5.0 U / mL, 5.1 U / mL, 5.2 U / mL, 5.3 U / mL, 5.4 U / mL, 5.5 U / mL, 5.6 U / mL, 5.7 U / mL, 5.8 U / mL, 5.9 U / mL, 6.0 U / mL, 6.5 U / mL, 7.0 U / mL, 7.5 U / mL, 8.0 U / mL, 8.5 U / mL, 9.0 U / mL, 9.5 U / mL, 10 U / mL, 11 U / mL, 12 U / mL, 13 U / mL, 14 U / mL, 15 U / mL, 16 U / mL, 17 U / mL, 18 U / mL, 19 U / mL, 20 U / mL, 25 U / mL, 30 U / mL, 35 U / mL, 40 U / mL, 45 U / mL, 50 U / mL, 55 U / mL, 60 U / mL, 65 U / mL, 70 U / mL, or a number or a range between any two of these values.
[0074] In some embodiments, the ssDNA may be provided with a 5’ phosphate for ligation (e.g., through ligation to adaptors with a 5’ phosphate and / or amplification with primers having a 5’ phosphate). Alternatively, a 5’ phosphate may be added to the ssDNA by a kinase, such as a T4 Polynucleotide Kinase (T4 PNK) in solution. The 5’ phosphate can be added by the kinase (for example, in the presence of the ligase in some embodiments). The kinase can be present at a concentration to facilitate the circularization of the single-stranded nucleic acid template aswill be apparent to a skilled person. The kinase can be at a concentration of less than 0.1 pM. In some embodiments, the concentration of the kinase is in a range from about 0.001 pM to about 0.1 pM. For example, the concentration of the kinase can be about, at least, at least about, at most, or at most about 0.001 pM, 0.01 pM, 0.02 pM, 0.03 pM, 0.04 pM, 0.05 pM, 0.06 pM, 0.07 pM, 0.08 pM, 0.09 pM, 0.1 pM, or a number or a range between any two of these values. In some embodiments, the concentration of the kinase is in a range from about 5 nM to about 20 nM, optionally about 5 nM, 5.5 nM, 6 nM, 6.5 nM, 7 nM, 7.5 nM, 8 nM, 8.5 nM, 9 nM, 9.5 nM, 10 nM, 10.5 nM, 11 nM, 11.5 nM, 12 nM, 12.5 nM, 13 nM, 13.5 nM, 14 nM, 14.5 nM, 15 nM, 15.5 nM, 16 nM, 16.5 nM, 17 nM, 17.5 nM, 18 nM, 18.5 nM, 19 nM, 19.5 nM, 20 nM, or a number or a range between any two of these values. In some embodiments, the concentration of the kinase can be in a range of about 0.01 U / mL to about 4 U / mL. For example, the concentration of the kinase can be, be about, be at least, be at least about, be at most, or be at most about, 0.01 U / mL, 0.02 U / mL, 0.03 U / mL, 0.04 U / mL, 0.05 U / mL, 0.06 U / mL, 0.07 U / mL, 0.08 U / mL, 0.09 U / mL, 0.1 U / mL, 0.15 U / mL, 0.2 U / mL, 0.25 U / mL, 0.3 U / mL, 0.31 U / mL, 0.32 U / mL, 0.33 U / mL, 0.34 U / mL, 0.35 U / mL, 0.36 U / mL, 0.37 U / mL, 0.38 U / mL, 0.39 U / mL, 0.4 U / mL, 0.41 U / mL, 0.42 U / mL, 0.43 U / mL, 0.44 U / mL, 0.45 U / mL, 0.46 U / mL, 0.47 U / mL, 0.48 U / mL, 0.49 U / mL, 0.5 U / mL, 0.51 U / mL, 0.52 U / mL, 0.53 U / mL, 0.54 U / mL, 0.55 U / mL, 0.56 U / mL, 0.57 U / mL, 0.58 U / mL, 0.59 U / mL, 0.6 U / mL, 0.61 U / mL, 0.62 U / mL, 0.63 U / mL, 0.64 U / mL, 0.65 U / mL, 0.66 U / mL, 0.67 U / mL, 0.68 U / mL, 0.69 U / mL, 0.7 U / mL, 0.75 U / mL, 0.80 U / mL, 0.85 U / mL, 0.90 U / mL, 0.95 U / mL, 10 U / mL, 11 U / mL, 12 U / mL, 13 U / mL, 14 U / mL, 15 U / mL, 16 U / mL, 17 U / mL, 18 U / mL, 19 U / mL, 20 U / mL, 25 U / mL, 30 U / mL, 35 U / mL, 40 U / mL, 45 U / mL, 50 U / mL, or a number or a range between any two of these values.
[0075] Circularization can occur in the presence of other components in addition to helicase, single-stranded DNA binding protein, and / or ligase (and / or kinase). For example, circularization can occur in the presence of betaine, such as 2 M of betaine. The concentration of betaine can be different in different implementations. In some embodiments, the concentration of betaine can be, be about, be at least, be at least about, be at most, or be at most about, 0.01 M, 0.02 M, 0.03 M, 0.04 M, 0.05 M, 0.06 M, 0.07 M, 0.08 M, 0.09 M, 0.1 M, 0.2 M, 0.3 M, 0.4 M, 0.5 M, 0.6 M, 0.7 M, 0.8 M, 0.9 M, 1 M, 1.1 M, 1.2 M, 1.3 M, 1.4 M, 1.5 M, 1.6 M, 1.7 M, 1.8 M, 1.9 M, 2 M, 2.1 M, 2.2 M, 2.3 M, 2.4 M, 2.5 M, 2.6 M, 2.7 M, 2.8 M, 2.9 M, 3 M, 4 M, 5 M, 6 M, 7 M, 8 M, 9 M, 10 M, 11 M, 12 M, 13 M, 14 M, 15 M, 20 M, 25 M, 30 M, 35 M, 40 M, 45 M, 50 M, 60 M, 70 M, 80 M, 90 M, 100 M, or a number or a range between any two of these values. For example, circularization can occur in the presence of MgCh, such as 0.01 M of MgCh. The concentration of MgCh can be different in different implementations. In some embodiments, the concentration of MgCh can be, be about, be at least, be at least about, be at most, or be at most about, 0.001 M, 0.002 M, 0.003 M, 0.004 M, 0.005 M, 0.006 M, 0.007 M, 0.008 M, 0.009 M, 0.01 M, 0.011 M, 0.012 M, 0.013 M, 0.014 M, 0.015 M, 0.016 M, 0.017 M, 0.018 M, 0.019 M, 0.02 M, 0.03 M, 0.04M, 0.05 M, 0.06 M, 0.07 M, 0.08 M, 0.09 M, 0.1 M, 0.11 M, 0.12 M, 0.13 M, 0.14 M, 0.15 M, 0.16 M, 0.17 M, 0.18 M, 0.19 M, 0.2 M, 0.25 M, 0.3 M, 0.35 M, 0.4 M, 0.45 M, 0.5 M, 0.6 M, 0.7 M, 0.8 M, 0.9 M, 1 M, or a number or a range between any two of these values. For example, circularization can occur in the presence of ATP, such as 5 pM of ATP. The concentration of ATP can be different in different implementations. In some embodiments, the concentration of ATP can be, be about, be at least, be at least about, be at most, or be at most about, 0.1 pM, 0.2 pM, 0.3 pM, 0.4 pM, 0.5 pM, 0.6 pM, 0.7 pM, 0.8 pM, 0.9 pM, 1 pM, 2 pM, 3 pM, 4 pM, 5 pM, 6 pM, 7 pM, 8 pM,9 pM, 10 pM, 11 pM, 12 pM, 13 pM, 14 pM, 15 pM, 16 pM, 17 pM, 18 pM, 19 pM, 20 pM, 25 pM, 30 pM, 35 pM, 40 pM, 45 pM, 50 pM, or a number or a range between any two of these values.
[0076] The nucleic acid template used herein can be DNA, including but not limited to genomic DNA, synthetic DNA, amplified DNA, complementary DNA (cDNA) or the like. The nucleic acid template used herein can also be RNA, including but not limited to, mRNA, ribosomal RNA, tRNA or the like. The nucleic acid template can be labeled or non-labeled (e.g., lack exogenous labels).
[0077] The nucleic acid template used herein can comprise nucleic acid analogs comprising modifications to the phosphate moiety, the sugar moiety and / or the nitrogenous base of a nucleotide analog. In some embodiments, the nucleic acid analog can include terminators that reversibly prevent subsequent nucleotide incorporation at the 3 ’-end of the primer. In some embodiments, such as in sequencing-by-binding or sequencing-by-synthesis, a reversible terminator moiety can be modified or removed from a primer, in a process known as “deblocking,” allowing for subsequent nucleotide incorporation.
[0078] In some embodiments, the nucleic acid template can be a library constituent having a distinct 5’ and 3’ adapter region flanking a target region. The adapter regions can have any of a variety functions including, for example, providing a binding site that complements portions of the splint oligonucleotide, providing a primer binding site for replicating the circularized nucleic acid template, providing a primer binding site for replicating a complement of the circularized template, providing a tag that is associated with the target region (e.g. a tag indicating the source of the target region such as a barcode or a tag used for identifying errors introduced during amplification of the target region etc.). The adaptor can have a length of about10 to about 250 nucleotides, depending on the number and size of the features included in the adaptors. In some embodiments, adaptors used herein can have a length pf about 10, 20, 30, 40, 50, or 60 nucleotides. The adapter regions or portions thereof can be common to a population of nucleic acid templates. The circularization method described herein can be performed for a plurality of nucleic acid templates, each comprising a 5’ adapter and / or a 3 adapter. The 5’ adapterand 3’ adapter are capable of hybridizing to a splint oligonucleotide or a primer. Whether or not the adapter regions have common sequences, the target regions in the nucleotide acid templates can have different sequences.
[0079] In some embodiments, the nucleic acid template has a high GC content or a high AT content. For example, the GC content in the nucleic acid template can be about, at least, or at least about 20%, 25%, 30%, 35%, 40%, 45%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or a number or a range between any two of these values. In some embodiments, the nucleic acid template has a GC content from about 50% to about 80%, optionally from about 70% to about 80%. In some embodiments, the nucleic acid template can have a high AT content. For example, the AT content in the nucleic acid template can be about, at least, or at least about 20%, 25%, 30%, 35%, 40%, 45%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or a number or a range between any two of these values. In some embodiments, the nucleic acid template has a AT content from about 50% to about 80%, optionally from about 70% to about 80%. In some embodiments, the single-stranded DNA can form a secondary structure, optionally the secondary structure comprises a stem, a hairpin structure, a pseudoknot, a bulge, an internal loop, a multiple loop, or a combination thereof. In some embodiments, the single-stranded DNA template is a library constituent.
[0080] The length of a nucleic acid template can be selected to suit a particular application of the methods set forth herein. For example, the length can be about, at least, at least about, at most, at most about 50, 100, 250, 500, 750, 1000, 1 x 104, 1 x 105or more nucleotides or base pairs. In some embodiments, the single-stranded DNA template is about 100 bps to 500 bps in length. For example, the length can be about, at least, at least about, at most, at most about 100, 150, 200, 250, 300, 350, 400, 450, 500 bps or a number or a range between any two of these values.
[0081] The target nucleic acids used herein can be derived from a biological source, a synthetic source or an amplification product. Exemplary organisms from which template nucleic acids can be derived include, for example, a mammal such as a rodent, mouse, rat, rabbit, guinea pig, ungulate, horse, sheep, pig, goat, cow, cat, dog, primate, human or non-human primate; a plant such as Arabidopsis thaliana, corn, sorghum, oat, wheat, rice, canola, or soybean; an algae such as Chlamydomonas reinhardlii: a nematode such as Caenorhabditis elegans an insect such as Drosophila melanogaster, mosquito, fruit fly, honey bee or spider; a fish such as zebrafish; a reptile; an amphibian such as a frog or Xenopus laevis: a dictyostelium discoideum: a fungi such as pneumocystis carinii. Takifugu rubripes. yeast, Saccharamoyces cerevisiae or Schizosaccharomyces pombe: or a plasmodium falciparum. Nucleic acids can be derived from a prokaryote such as a bacterium, Escherichia coli. staphylococci or mycoplasma pneumoniae,' anarchaea; a virus such as Hepatitis C virus, influenza virus, coronavirus or human immunodeficiency virus; or a viroid. Nucleic acids can be derived from a homogeneous culture or population of the above organisms or alternatively from a collection of several different organisms, for example, in a community or ecosystem. Nucleic acids can be isolated using methods known in the art including, for example, those described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rdedition, Cold Spring Harbor Laboratory, New York (2001) or in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1998), each of which is incorporated herein by reference.
[0082] In some embodiments, the circularized DNA template can be sequenced by, for example, single molecule real-time sequencing. In some embodiments, the amplified concatemers can be sequenced. The amplified concatemers can be sequenced by, for example, paired end sequencing. The amplified concatemers can be sequenced by, for example, SBB. The amplified concatemers can be sequenced by, for example, SBS. The amplified concatemers can be sequenced by, for example, single molecule real-time sequencing. The amplified concatemers can be sequenced by, for example, nanopore sequencing. In some embodiments, the amplified concatemers can be sequenced by SBB or SBS on (or in or using) a flow cell.
[0083] In some embodiments, GC bias of reads obtained by the sequencing is less than or less than about, 0.001, 0.002, 0.003, 0.004, 0.005, 0.006, 0.007, 0.008, 0.009, 0.01, 0.011, 0.012, 0.013, 0.014, 0.015, 0.016, 0.017, 0.018, 0.019, 0.02, 0.021, 0.022, 0.023, 0.024, 0.025, 0.026, 0.027, 0.028, 0.029, 0.03, 0.031, 0.032, 0.033, 0.034, 0.035, 0.036, 0.037, 0.038, 0.039, 0.04, 0.041, 0.042, 0.043, 0.044, 0.045, 0.046, 0.047, 0.048, 0.049, 0.05, 0.051, 0.052, 0.053, 0.054, 0.055, 0.056, 0.057, 0.058, 0.059, 0.061, 0.062, 0.063, 0.064, 0.065, 0.066, 0.067, 0.068, 0.069, 0.07, 0.071, 0.072, 0.073, 0.074, 0.075, 0.076, 0.077, 0.078, 0.079, 0.08, 0.081, 0.082, 0.083, 0.084, 0.085, 0.086, 0.087, 0.088, 0.089, 0.09, 0.091, 0.092, 0.093, 0.094, 0.095, 0.096, 0.097, 0.098, 0.099, 0.1, or a number or a range between any two of these values. For example, GC bias of reads obtained by the sequencing is less than 0.05. As another example, GC bias of reads obtained by the sequencing can be less than 0.03.
[0084] In some embodiments, the action of helicase and / or SSB reduces the GC bias to less than a percentage of GC bias in the absence of helicase and SSB, such as 85%, 84%, 83%,82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%,65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, 50%, 49%,48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%,31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, or a number or a range between any two of these values. For example, the action of helicase and / or SSB reduces the GC bias to less than 75% of GC bias in the absence of helicase and SSB. As another example, theaction of helicase and / or SSB reduces the GC bias to less than 50% of GC bias in the absence of helicase and SSB.
[0085] In some embodiments, the Fl score of the sequencing is greater than 0.9, 0.901, 0.902, 0.903, 0.904, 0.905, 0.906, 0.907, 0.908, 0.909, 0.91, 0.911, 0.912, 0.913, 0.914, 0.915, 0.916, 0.917, 0.918, 0.919, 0.92, 0.921, 0.922, 0.923, 0.924, 0.925, 0.926, 0.927, 0.928, 0.929, 0.93, 0.931, 0.932, 0.933, 0.934, 0.935, 0.936, 0.937, 0.938, 0.939, 0.94, 0.941, 0.942, 0.943, 0.944, 0.945, 0.946, 0.947, 0.948, 0.949, 0.95, 0.951, 0.952, 0.953, 0.954, 0.955, 0.956, 0.957, 0.958, 0.959, 0.96, 0.961, 0.962, 0.963, 0.964, 0.965, 0.966, 0.967, 0.968, 0.969, 0.97, 0.971, 0.972, 0.973, 0.974, 0.975, 0.976, 0.977, 0.978, 0.979, 0.98, 0.981, 0.982, 0.983, 0.984, 0.985, 0.986, 0.987, 0.988, 0.989, 0.99, 0.991, 0.992, 0.993, 0.994, 0.995, 0.996, 0.997, 0.998, 0.999, or a number or a range between any two of these values, for SNP. Alternatively or additionally, the Fl score of the sequencing is greater than 0.9, 0.901, 0.902, 0.903, 0.904, 0.905, 0.906, 0.907, 0.908, 0.909, 0.91, 0.911, 0.912, 0.913, 0.914, 0.915, 0.916, 0.917, 0.918, 0.919, 0.92, 0.921, 0.922, 0.923, 0.924, 0.925, 0.926, 0.927, 0.928, 0.929, 0.93, 0.931, 0.932, 0.933, 0.934, 0.935, 0.936, 0.937, 0.938, 0.939, 0.94, 0.941, 0.942, 0.943, 0.944, 0.945, 0.946, 0.947, 0.948, 0.949, 0.95, 0.951, 0.952, 0.953, 0.954, 0.955, 0.956, 0.957, 0.958, 0.959, 0.96, 0.961, 0.962, 0.963, 0.964, 0.965, 0.966, 0.967, 0.968, 0.969, 0.97, 0.971, 0.972, 0.973, 0.974, 0.975, 0.976, 0.977, 0.978, 0.979, 0.98, 0.981, 0.982, 0.983, 0.984, 0.985, 0.986, 0.987, 0.988, 0.989, 0.99, 0.991, 0.992, 0.993, 0.994, 0.995, 0.996, 0.997, 0.998, 0.999, or a number or a range between any two of these values, for indels. For example, the Fl score of the sequencing is greater than 0.99 for SNP. The Fl score of the sequencing is greater than 0.98 for indels. For example, the Fl score of the sequencing is greater than 0.993 for SNP. The Fl score of the sequencing is greater than 0.987 for indels. In some embodiments, the action of helicase and / or SSB maintains or increases at least one of average spot signal, uniformity of signal, and throughput for, for example, the sequencing.Rolling Circle Amplification (RCA)
[0086] Provided herein also include methods, kits, systems and compositions for rolling circle amplifications of nucleic acids. Generally, a rolling circle amplification or RCA method involves a polymerase extending a primer that is annealed to a circular template (or a linear template circularized to form a circular template) such that multiple laps of the polymerase around the circular template produces a concatemeric single stranded nucleic acid that contains multiple tandem repeats, each of the repeats being complementary to the circular template.
[0087] The nucleic acids can be circular nucleic acids or linear nucleic acids circularized to form circular nucleic acids. In some embodiments, the nucleic acid template is a linear nucleic acid that can be circularized to form a circular nucleic acid template, for example,using the circularization methods disclosed herein. Accordingly, in some embodiments, the method can comprise (a) circularizing a single-stranded DNA template on a splint oligonucleotide in a circularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration to form a circularized DNA template and (b) performing RCA by contacting the circularized DNA template with a RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template. The single-stranded DNA template can comprise a 5’ adapter and a 3’ adapter and the splint oligonucleotide can comprise a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter, thereby allowing the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template.
[0088] A variety of other methods can also be used to prepare a circular nucleic acid template from a linear nucleic acid template for an RCA. For example, circular nucleic acid can also be generated by chemical synthesis of suitable linear oligonucleotides following by circularization of the synthesized oligonucleotides. In some embodiments, a nucleic acid template is circularized prior to being hybridized to a primer for RCA. For example, in a configuration shown in FIG. 2A, a target nucleic acid is circularized prior to being hybridized to the immobilized primer on the solid support. Accordingly, in some embodiments, the method can comprise performing RCA by contacting the circularized DNA template with an RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template (see, for example, FIG. 2B).
[0089] Similar to the circularization process described herein, the RCA can be performed in solution or on a solid support. In some embodiments, circularizing a single-stranded DNA template can be performed in a solution. The splint oligonucleotide and the nucleic acid template can be provided in a solution and then delivered to a vessel by any suitable means. Following the circularization of the nucleic acid template in solution, the RCA of the circularized nucleic acid template can then be performed in the solution (e.g., in the same vessel where the circularization takes place) or on a solid support. In the embodiments wherein the RCA is performed in the solution, the splint oligonucleotide used for the circularization process can also be used as a primer for the RCA. For example, following the circularization of the single-stranded nucleic acid template in the solution, the splint oligonucleotide can function as a primer extending along the circularized nucleic acid template to produce a single-stranded concatemer via rolling circle amplification of the template that Is hybridized to the primer / splint oligonucleotide.
[0090] In some other embodiments, following the circularization of the nucleic acid template in solution, the RCA can then be performed on a solid support. A solid support having a capture primer can be provided. The capture primer can be attached to the solid support eitherdirectly or via a linker, using covalent or non-covalent attachment chemistries known in the art, prior to being contacted with the nucleic acid template and the RCA mixture. The capture primer can be used to capture the circularized nucleic acid template via at least a portion comprising a sequence complementary to that of the nucleic acid template. The method can further comprise contacting the circularized nucleic acid template with the RCA mixture in the presence of the capture primer immobilized on the solid support. The circularized nucleic acid template can be deposited on the surface of the solid support. The immobilized primer can hybridize to a portion or portions of a primer binding site present in the nucleic acid template and be extended along the circularized nucleic acid template to form a single-stranded concatemer.
[0091] In some embodiments, circularizing a single-stranded nucleic acid template and the RCA of the circularized nucleic acid template can be both performed on a solid support. FIGS. 2A-B provide a non-limiting schematic illustration of performing a rolling circle amplification for nucleic acids on a solid support. As shown in FIG. 2A, a nucleic acid primer (indicated by the open and lined rectangles) is attached to a solid support (indicated by the dotted rectangle) via a linker (indicated by the grey line). The primer can serve as a splint oligonucleotide for the circularization as well as a primer for the RCA. The primer can be used to capture a target nucleic acid via a primer binding site in the nucleic acid template (e.g., a linear nucleic acid template) that is complementary to the primer. In one configuration shown in FIG. 2A, the immobilized primer can hybridize to portions of the primer binding site that are present at opposite ends of a target sequence (the target sequence being indicated by a dotted line and the flanking primer binding site regions being indicated by open and lined rectangles, respectively). The immobilized primer thus functions as a splint that brings together the two ends of the target nucleic acid. The two ends can be ligated while hybridized to a splint nucleic acid to form a circular version of the target nucleic acid. The circularization process can occur in the presence of a helicase and optionally a single-stranded DNA binding protein as described herein.
[0092] FIG. 2B provides a non-limiting schematic illustration of a single-stranded concatemer being produced via rolling circle amplification of a primed circular template that is hybridized to an immobilized primer. The primer is immobilized in a way that the 3’ end is available for polymerase extension (e.g., the primer can be attached at or near its 5’ end). The product of the first sub-step is shown as having progressed to a point that two copies of the circular template (two sequence units) have already been produced and the circular template is hybridized to a portion of a third copy (third sequence unit) that is being replicated. Each of the sequence units includes a region that is complementary to the target sequence (indicated by the solid black line) and a region that is complementary to the primer (indicated by the open and lined rectangles). The product of the second sub-step has progressed to the point of having produced nearly sixcopies of the circular template. FIG. 2B shows the final product of the RCA reaction after the circular template is absent (e.g., has been removed) in the third sub-step. Two regions of the final product are shown for illustrative purposes: a region where the sequence units are delineated (indicative of the concatemeric primary structure of the amplified strand) and a region where the number and conformation of the sequence units is not specified (indicative of the dynamic and variable secondary structure for the cluster as a whole).
[0093] In the embodiments herein described, performing RCA comprises contacting the circularized nucleic acid template with an RCA mixture in a vessel for a time duration to form amplified concatemers of the nucleic acid template. In some embodiments, a vessel wherein the circularized nucleic acid template is being contacted with an RCA mixture is a flow cell. Accordingly, in some embodiments, the contacting step can be facilitated by the use of a flow cell. A typical flow cell includes microfluidic valving that permits delivery of liquid reagents (e.g., components of the RCA mixture) through an inlet and removal of liquid reagents from by exiting from an outlet. Flowing liquid reagents through a flow cell can permit reagent mixing and exchange. For example, contacting a nucleic acid template and a primer with an RCA mixture can comprise flowing the RCA mixture through a flow cell.
[0094] The nucleic acid template, the RCA mixture, and optionally the primer can be contacted simultaneously. Alternatively, the nucleic acid template, the RCA mixture, and optionally the primer can be contacted sequentially. For example, the nucleic acid template can be contacted with the primer to form a primed-template nucleic acid (e.g., in FIG. 2A), which is then contacted with the RCA mixture. In some embodiments, the nucleic acid template can be contacted with the RCA mixture and then with the capture primer.
[0095] The temperature for the RCA reaction can vary in different embodiments, for example, depending on the polymerase used, the reaction conditions, and the like. The RCA mixture can be incubated with the template nucleic acid in the presence of a primer at any temperature conducive to the polymerase activity. In some embodiments, contacting a template nucleic acid with a RCA mixture is performed at a substantially isothermal reaction temperature, for example, a temperature that does not vary more than by about 2-3 °C above or below a given temperature. The reaction temperature can be between 20°C and 70°C, for example 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, or a number or a range between any two of these values. In some embodiments, the reaction temperature is between 20°C and 60°C, for example between 20°C and 50°C. For example, the reaction temperature can be, or be about, 30°C. In some embodiments, the reaction temperature is or is about 37°C.
[0096] An RCA reaction can be terminated by denaturing the polymerase, for example, by heating the sample at 60°C, 65°C, 70°C, 75°C, 80°C, or higher. An RCA reactioncan also be terminated by removing one or more components of RCA, such as the polymerase, the dNTPs, or any combination thereof. Components of RCA can be removed by, for example, washing with a washing reagent.
[0097] In some embodiments, such seeding and amplification can result in a plurality of binding areas each containing an ensemble of essentially the same amplified molecules, and the ensemble of amplified molecules on the binding areas are different across the binding areas (that is, the ensemble of amplified molecules on a given binding area is different from the ensemble of amplified molecules on any other pad of the plurality of binding areas). The surface can be a flow cell surface.
[0098] In some embodiments, a surface (e.g., a structured surface) having a plurality of binding areas (e.g., pads) can be provided, wherein oligonucleotides A and B are immobilized on each of the plurality of binding areas. Oligonucleotides A comprises binding regions that are complementary and capable of capturing library elements, and oligonucleotide B comprises a binding region that is identical (or substantially identical) to a sequence of the library elements” A member nucleic acid of a linear library can comprise adapter A and adapter B at each end of the nucleic acid and can hybridize to a splint oligonucleotide and form ssDNA ring by ligation, for example, using the circularization method described herein. Upon ligation of adapter A to adapter B, a linear library constituent is converted to a circular molecule having an adapter A / adapter B ligated region that is reverse complementary to one set of oligonucleotides on the binding site, and locally identical in sequence to the second of the primers on the binding site. Consequently, the capturing oligonucleotide can serve to prime rolling circle amplification of the circularized library constituent, which in turn generates a linear concatemer reverse complementary to the original library constituent and having a segment that is reverse complementary to the second oligonucleotide on the surface. A flow cell surface contains a plurality of binding site or areas (e.g., pads) and each of the binding sites or areas has primers A and B attached (e.g., immobilized) thereon. Primers A and B each comprise sequence that spans the ligation event between adapters A and B of the original library constituent. Primers A and B are capable of functioning as primers in the subsequent reactions. Primer A comprises a first portion that is complementary (or substantially complementary) to a portion of adapter B, and a second portion that is complementary (or substantially complementary) to a portion of adapter A. Primer B comprises a sequence portion the same (or substantially the same) as a sequence portion in adapter A. Because of the sequence complementarity, primer A can bind to the ssDNA ring template to generate DNA concatemers in the presence of DNA polymerase. Primer B can then bind to the concatemers generated by extension of the primer A bound to the circular template,generating reverse complement concatemers. However, primer A and primer B are not reverse complementary to one another, such that they do not form dimers on the surface.
[0099] Because of the sequence complementarity that oligonucleotides A and B have with adapter A and adapter B on the member nucleic acid of the library, respectively, oligonucleotides A and B can be used as capture oligonucleotides and amplification oligonucleotides. For example, using oligonucleotide A as a primer, RCA can be carried out and generate concatemer copies of the template ssDNA. Oligonucleotide B can then bind to the concatemer DNA and functions as a primer to allow RCA to generate complementary concatemers. The seeding and amplification can result in a binding area having tethered thereto a plurality of concatemeric nucleic acid molecules, where a concatemeric strand can be partially hybridized to multiple complementary concatemeric strands. The concatemeric strands and complementary concatemeric strands can be subjected to next generation sequencing. The concatemeric strands and complementary concatemeric strands can be sequenced when they are partially hybridized to one another. Alternatively or additionally, the concatemeric strands and the complementary concatemeric strands can be separated by, for example, heat denaturation such that the strands are not partially hybridized. Detailed description of nucleic acid amplification and single-molecule seeding and amplification of nucleic acids on a surface can be found in U.S. Publication No. 2022 / 0356519 Al, the content of which is incorporated herein by reference in its entirety.Binding Sites
[0100] The methods, compositions and systems described herein can be used to isothermally seed and amplify nucleic acids on a surface (e.g., a structured surface) comprising a plurality of binding sites (or areas, spots or pads). Structured surfaces and methods of generating structured surfaces have been described in U.S. Application Publication No. 2022 / 02098568, the content of which is incorporated therein in its entirety. A structured surface can comprise a plurality of binding sites, such as at least 5,000, 7,500, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000 binding sites. A flow cell (or a flow cell surface) can comprise the structured surface. Each of the plurality of binding sites can be circular. Each of the plurality of binding sites can have a center point and a diameter. In some instances, the structured surface includes a plurality of binding sites of at least 10,000 ordered binding sites separated by disjunctions that are not predetermined and / or are randomly (or irregularly) distributed (or arranged). Alternatively or additionally, the structured surface contains a plurality of binding sites of at least 10,000 ordered binding sites separated by disjunctions. The configurations of the ordered binding sites and the disjunctions are not predetermined and / or are randomly distributed (or arranged). In someembodiments, the structured surface comprises a plurality of binding sites of at least 5,000 binding sites. The plurality of binding sites can include a first plurality of ordered binding sites and a second plurality of ordered binding sites separated by a disjunction that is not predetermined and / or is randomly distributed (or arranged). A first configuration of the first plurality of ordered binding sties and a second configuration of the second plurality of ordered binding sites can be different or the same.
[0101] The binding sites and / or the disjunctions can be at positions that are not predetermined. The binding sites and / or the disjunctions can be ordered, can be irregularly distributed, and / or can be randomly distributed. In some instances, configurations of the binding sites can be not predetermined. The disjunctions can be at positions that are not predetermined. The disjunctions (including binding sites therein if any) can be irregularly distributed and / or can be randomly distributed. Binding sites of the plurality of binding sites can be at positions that are not predetermined. The binding sites can be ordered, irregularly distributed (or arranged), and / or randomly distributed (or arranged). In some instances, the plurality of binding site can comprise a first plurality of binding sites and a second plurality of binding sites separated by a disjunction. The position, size, and / or shape of the disjunction can be not predetermined. The disjunction can be randomly distributed. Binding sites of the first plurality of binding sites and / or the second plurality of binding sites can be at positions that are not predetermined, can be ordered, can be irregularly distributed, and / or can be randomly distributed. A subset of binding sites can have a configuration that is not predetermined. A configuration can comprise the number binding sites in the subset of binding sites and / or positions of the binding sites in the subset of binding sites. The first subset of binding sties and the second subset of binding sites can have configurations that are different (or identical).
[0102] There can be different separations or pitches between any two neighbor binding sites. Separation between any binding site and any nearest neighbor binding site, measured from the center of the first binding site to the center of the nearest neighbor binding site, can be at least twice as large as the diameter of the first binding site. Alternatively or additionally, separation between any binding site and any nearest neighbor binding site, measured from an edge of the first binding site to a center of the nearest neighbor binding site is at least twice as large as the diameter of the first binding site. The two edges are closer than (or at least as close as) separation between any other edge of the first binding site and any edge of the second binding site. The separation can be at least two times, three times, four times, or more as large as the diameter of the first binding site diameter. The separation can be about 10'9m to about 10'4m, or shorter or longer.
[0103] Some, substantially all, or all binding sites can be randomly arrayed on thestructured surface. For example, at least a portion of the plurality of binding sites is randomly arrayed on the structured surface. As another example, the plurality of binding sites is not arrayed on the structured surface at a predetermined set of locations. For example, at least a portion of the plurality of binding sites does not share a common pattern. As another example, the plurality of binding sites comprises unpatterned binding sites.
[0104] The plurality of binding sites can include a first subset (or plurality) of binding sites and a second subset (or plurality) of binding sites. The first subset of binding sites and the second subset of binding sites can be separated by a disjunction that is not predetermined. A first configuration comprising the first plurality of binding sties (e.g., the relative positions of the binding sites) and a second configuration comprising the second plurality of ordered binding sites can be different. The first subset of binding sites can be in a first crystal lattice or a first irregular array. The second subset of binding sites can be in a second crystal lattice or a second irregular array. The first irregular array and / or the second irregular array can comprise no particles in a hexagonal configuration, an equilateral triangle configuration, a straight line configuration, a linear configuration, or a combination thereof. The first irregular array and / or the second irregular array can comprise no particles within at least a threshold distance of each other that are in a hexagonal configuration, an equilateral triangle configuration, a straight line configuration, a linear configuration, or a combination thereof. The threshold distance can be, for example, 5 pm.
[0105] There can be different local densities of the plurality of binding sites (or a subset of the binding sites). The plurality of binding sites (or a subset of binding sites) can be present at a local density of (about) 200k / mm2to (about) 8,000k / mm2. Non-limiting exemplary local densities of binding sites include at least 600k / mm2, at least 800k / mm2, at least l,000k / mm2, at least 2,000k / mm2, or more. A structured surface can comprise (or comprise at least, or comprise at most), 2,500, 5,000, 10,000, 25,000, 50,000, 75,000, 100,000 500,000, 1,000,000, 5,000,000, 10,000,000, or a number or a range between any two of these values, binding sites.
[0106] The material of the binding sites can be selected based on the desired properties or applications of the structured surface (or flow cell surface). The binding sites can be hydrophilic, hydrophobic, positively charged, negatively charged, uncharged, or a combination thereof. The material of the structured surface (or flow cell surface) can be selected based on the desired properties or applications of the structured surface (or flow cell surface). A material of the structured surface (or flow cell surface) comprises, for example, silicon, silicon nitride glass, borosilicate glass, quartz, fused quartz, silica, fused silica, a metal, a ceramic, plastic, or a combination thereof. The plurality of binding sites can comprise a plurality of nucleic acids. The plurality of nucleic acids can comprise at least one concatemeric nucleic acid (e.g., generated by rolling circle amplification following circularization). One, at least one, or each, of the pluralityof binding sites can comprise one, or at most one, of the plurality of nucleic acids. At least 30%, 35%, 40%, 45%, or 50% of the plurality of binding sites can comprise at least one nucleic acid of the plurality of nucleic acids and / or one bead of the plurality of beads.
[0107] In some embodiments, the binding sites (or reaction sites) comprise an acrylate functional silane, an aldehyde functional silane, an amino functional silane, an anhydride functional silane, an azide functional silane, a carboxylate functional silane, a phosphonate functional silane, a sulfonate functional silane, an epoxy functional silane, a thiol functional silane, an ester functional silane, a vinyl functional silane, an olefin functional silane, a halogen functional silane, a dipodal silane, or a combination thereof. For example, the reaction sites comprise an aminosilane, a glycidoxysilane, a mercaptosilanes, or a combination thereof. The binding sites can comprise nucleic acid concatemers.
[0108] Various binding site configurations that are random and / or not predetermined are consistent with the disclosure. For example, two structured surfaces (or flow cell surfaces) share a congruent binding site configuration comprising fewer than all of the binding sites on each of the two flow cell surfaces. Two flow cell surfaces of said plurality of flow cell surfaces can share a congruent binding site configuration comprising at most 1%, 2%, 3%, 4%, or 5% of the binding sites on each of the two flow cell surfaces. As an example, a structured surface, or every structured surface, comprises at least one plurality of at least three binding sites that is not congruent with any plurality of at least three binding sites on any other structured surface. As another example, a structured surface, or every structured surface, comprises at least one plurality of at least ten binding sites that is not congruent with any plurality of at least ten binding sites on any other structured surface. For example, a structured surface, or every structured surface, comprises at least 1%, 2%, 3%, 4%, or 5% of the plurality of binding sites on the structured surface that is not congruent with any 1%, 2%, 3%, 4%, 5% of the plurality of binding sites on any other structured surface. As an example, a structured surface, or every structured surface, comprises at least 5%, 10%, 15%, 20%, or 25% of the plurality of binding sites on the structured cell that is not congruent with any 5%, 10%, 15%, 20%, or 25% of the plurality of binding sites on any other structured surface.
[0109] Additional binding site configurations that are random and / or not predetermined are possible. As an example, a first structured surface comprises a first binding site array having two first regions each with a first regular, irregular, or random binding site array configuration and separated by a first region of irregular or random binding site array configuration. A second structured surface can comprise a second binding site array having two second regions each with a second regular, irregular, or random binding site array configuration and separated by a second region of irregular or random binding site array configuration. The firstbinding site array and the second binding site array can be distinct. The first binding site array can comprise all of the binding sites on the first structured surface. The second binding site array can comprise all of the binding sites on the second structured surface.
[0110] For example, the first region with the first irregular or random binding site array configuration and the second region with the second irregular or random binding site array configuration are not congruent. The first region with the first irregular or random binding site array configuration can comprise at least one plurality of at least three binding sites that is not congruent with any plurality of at least three binding sites of said second region with the second irregular or random binding site array configuration. The two first regions each with a first regular, irregular, or random binding site array configuration can comprise binding sites that are congruent. The wo first regions each with a first regular, irregular, or random binding site array configuration can comprise binding sites that are not congruent. One of the two first regions with a first regular, irregular, or random binding site array configuration. One of the second regions each with a second regular, irregular, or random binding site array configuration can comprise binding sites that are congruent (e.g., having binding sites at the same or substantially the same relative positions or locations). One of the two first regions each with a first regular, irregular, or random binding site array configuration and one of the second regions each with a second regular, irregular, or random binding site array configuration comprise binding sites that are not congruent.[OHl] The number of binding sites in a region of binding site array can vary. In some embodiments, each first region with a first regular, irregular, or random binding site array configuration can comprise at least 250, 500, 750, 1000, or 1,500 binding sites. Each second region with a second regular, irregular, or random binding site array configuration can comprise at least 250, 500, 750, 1000, or 1,500 binding sites. The first region with a first irregular or random binding site array configuration can comprise at least 500 binding sites. The second region with a second irregular or random binding site array configuration can comprise at least 250, 500, 750, 1000, or 1,500 binding sites. In some instances, each first region with a first regular, irregular, or random binding site array configuration comprises at least 1%, 2%, 3%, 4%, or 5% of the binding sites on the first structured surface and / or of the first binding array. Each second region with a second regular, irregular, or random binding site array configuration can comprise at least 1%, 2%, 3%, 4%, or 5% of the binding sites on the first structured surface and / or of the first binding array. The first region with an irregular or random binding site array configuration can comprise at least 1%, 2%, 3%, 4%, or 5% of the binding sites on the first structured surface and / or of the first binding array. The second region with an irregular or random binding site array configuration can comprise at least 1%, 2%, 3%, 4%, or 5% of the binding sites on the second structured surface and / or of the second binding array.
[0112] In some embodiments, the plurality of binding sites on one, one or more, or each, of the plurality of flow cell surfaces comprises a plurality of nucleic acids. Binding sites can be hydrophilic, hydrophobic, positively charged, negatively charged, uncharged, or a combination thereof. A nucleic acid can be a concatemeric nucleic acid. One, at least one, or each, of the plurality of binding sites on each of the plurality of flow cell surfaces can comprise one, or at most one, of the plurality of nucleic acids on the plurality of binding sites on the flow cell surface. The plurality of nucleic acids on the plurality of binding sites on each of the plurality of flow cell surfaces can be attached to a plurality of beads. One, at least one, or each, of the plurality of binding sites can comprise one, or at most one, of the plurality of beads. For a flow cell surface, one, one or more, or each of the plurality of nucleic acids on the plurality of binding sites on the flow cell surface comprises an amplification primer (or a splint primer) an amplification product from rolling circle amplification (RCA), a DNA tile, or a combination thereof. One, one or more, or each of the plurality of nucleic acids on the plurality of binding sites on the flow cell surface can comprise a first functional moiety capable of participating in a click chemistry reaction, such as methyltetrazine (MTz). One, one or more, or each of the plurality of binding sites on the flow cell surface comprises a second functional moiety capable of participating in the click chemistry reaction, such as trans-cyclooctene (TCO). The DNA tile can comprise an amplification primer or a splint primer, optionally the amplification primer or the splint primer is attached to the binding site, and optionally the amplification primer or the splint primer is attached to the binding site via the click chemistry reaction involving the first functional moiety and the second functional moiety.
[0113] At least 50% of the plurality of binding sites on one, one or more, or each, of the plurality of structured surfaces can comprise at least one nucleic acid of the pluralities of nucleic acids and / or one bead of the pluralities of beads.RCA Mixture
[0114] The RCA described herein is carried out by contacting a nucleic acid template with an RCA mixture. The RCA mixture can comprise a polymerase (e.g., a DNA polymerase), a dNTP mix (including dGTP, dATP, dTTP, dCTP, or any combination thereof), and a divalent metal cation.
[0115] The term “divalent metal cation” refers to a catalytic metal cation having a valence of two. The divalent metal cation is required for phosphodiester bond formation between the 3 ’-OH of a nucleic acid (e.g., a capture primer) and the phosphate of an incoming nucleotide. The divalent metal cation used herein can be, for example, a magnesium cation (Mg2+), a calcium cation (Ca2+), a strontium cation (Sr2+), manganese cation (Mn2+), copper cation (Cu2+), cadmium cation (Cd2+), Zinc cation (Zn2+), or a combination thereof.
[0116] The divalent metal cation can be present at different concentrations at different stages of a rolling circle amplification reaction, including but not limited to seeding / loading stage and amplification stage. In some embodiments, different types of divalent metal cations can be used in different stages of a rolling circle amplification. For example, one or more divalent metal cation can be used in a seeding / loading stage to facilitate efficient loading of a polymerase (e.g., a DNA polymerase) onto a template nucleic acid to form a polymerase-template nucleic acid complex, thereby initiating the amplification. Another divalent metal cation can be used in the amplification stage to facilitate the primer extension and nucleic acid amplification. In some embodiments, the divalent metal cation in seeding stage can be present at a low concentration necessary to facilitate efficient polymerase loading. In some embodiments, the divalent metal cation in an amplification buffer can be present at a higher concentration conducive to the nucleic acid amplification. The concentration of the divalent metal cation can also vary depending on the choice of the divalent metal cation, polymerase, and / or the template nucleic acid. The selection of the divalent metal cation may be based on the polymerase and / or the nucleotides in a reaction.
[0117] In some embodiments, the divalent metal cation used in an RCA mixture can be a calcium cation (Ca2+), a strontium cation (Sr2+) or a magnesium cation (Mg2+). In some embodiments, the divalent metal cation used in the seeding stage can be a calcium cation, a strontium cation and the divalent metal cation used in the amplification stage can be a magnesium cation. In some embodiments, the divalent metal cation used in the amplification stage is in a higher concentration than the divalent metal cation in the seeding stage. In some embodiments, the RCA mixture in the seeding stage does not comprise a magnesium cation and the magnesium cation is only introduced in the amplification stage. In some embodiments, the RCA mixture in the seeding stage and the amplification stage both comprises a magnesium cation. The Mg2+can be at different concentrations.
[0118] The divalent metal cation in the RCA mixture can be present at a concentration necessary to facilitate efficient loading of a polymerase onto a template nucleic acid to form a polymerase-nucleic acid complex, thereby initiating the amplification. The concentration of a divalent metal cation is from about 0.001 mM to about 1 M, about 0.5 mM to about 200 mM, or about 0.1 to 30 mM, or about 1 to 10 mM. For example, the concentration of the divalent metal cation can be about, at most, or at most about 0.001 mM, 0.05 mM, 0.1 mM, 0.5 mM, 1.0 mM, 1.5 mM, 2 mM, 2.5 mM, 3.0 mM, 3.5 mM, 4.0 mM, 4.5 mM, 5.0 mM, 5.5 mM, 6.0 mM, 6.5 mM, 7.0 mM, 7.5 mM, 8.0 mM, 8.5 mM, 9.0 mM, 9.5 mM, 10 mM, 20 mM, 40 mM, 60 mM, 80 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1 M, or a number or a range between any two of these values.
[0119] In some embodiments, the divalent metal cation is a calcium cation or a strontium cation. The concentration of the divalent metal cation is from about 0.001 mM to about 30 mM, about 0.1 to 20 mM, or about 1 to 10 mM. In some embodiments, the divalent metal cation is a magnesium cation. The concentration of the divalent metal cation is from about 0.001 mM to about 30 mM, about 0.1 to 20 mM, or about 1 to 10 mM. In some embodiments, the RCA mixture does not comprise a magnesium cation at the seeding stage.
[0120] The polymerase can be present at a concentration necessary to facilitate the nucleic acid amplification as will be apparent to a skilled person. In some embodiments, the mole ratio of primers to template nucleic acids can be about 1015: 1 or less. For example, the mole ratio of primers to template nucleic acids can be about 1015: 1, 5xl014: l, 1014: 1, 1013: 1, 1012: 1, 1011: 1, 1010: 1, 109: 1, 108: 1, 107: 1, or 106:1. In some embodiments, the molar ratio of polymerase to template nucleic acids can be about 1 ,5xl010: 1 or less. For example, the molar ratio of polymerase to template nucleic acids can be about 3xl09: l, 109: 1, 108: 1, 107: 1, 107: 1, 106: 1, 106: 1, 105: 1, 104:I,103:I, 102: l, or 50: 1.
[0121] The dNTPs in the RCA mixture can be in a range of about 10 pM to about lOmM as will be apparent to a skilled person. In some embodiments, the dNTP concentration is less than 10 mM to avoid hydrogel formation from the amplified concatemers and to remain at a concentration lower than or equal to the amount of divalent metal cation present in the RCA mixture.
[0122] In some embodiments, the RCA mixture further comprises a polymer. The term “polymer,” as used herein, can refer to a molecule comprising a plurality of repeating structural units (monomers), typically at least 3, linked together via covalent bonds. The polymer can be linear or branched. In some embodiments, the polymer used herein is hydrophilic. Suitable polymers do not interfere with the primer extension and nucleic acid amplification and can reduce the size of a cluster formed by a population of nucleic acids (e.g., a cluster formed by strands of nucleic acid concatemers).
[0123] A polymer used herein can comprise a natural polymer, including for example alginate, agarose, carrageenan, chitosan, dextran, carboxymethylcellulose, heparin, hyaluronic acid, polyamino acid, collagen, gelatin, fibrin, a fibrous protein-based biopolymer, or any combination thereof.
[0124] A polymer can comprise a synthetic polymer, such as, for example, alginic acid-polyethylene glycol copolymer, polyethylene glycol), poly(2-methyl-2-oxazoline), poly(ethylene oxide), poly(vinyl alcohol), poly(acrylamide), poly(n-butyl acrylate), poly-festers), poly(glycolic acid), poly(lactic-co-glycolic acid), poly(L-lactic acid), poly(N- isopropylacrylamide), butyryl-trihexyl-citrate, di(2-ethylhexyl)phthalate, di -iso-nonyl- 1,2-cyclohexanedicarboxylate, expanded polytetrafluoroethylene, ethylene vinyl alcohol copolymer, poly(hexamethylene diisocyanate), highly crosslinked poly(ethylene), poly(isophorone diisocyanate), poly(amide), poly(acrylonitrile), poly(carbonate), poly(caprolactone diol), poly(D- lactic acid), poly(dimethylsiloxane), poly(dioxanone), poly(ethylene), polyether ether ketone, polyester polymer alloy, polyether sulfone, poly(ethylene terephthalate), poly(hydroxyethyl methacrylate), poly(methyl methacrylate), poly(methylpentene), poly(propylene), polysulfone, poly(vinyl chloride), poly(vinylidene fluoride), poly(vinylpyrrolidone), poly(styrene-b- isobutylene-b-styrene), or any combination thereof.
[0125] A polymer can comprise a polymer selected from the group comprising collagen, gelatin, chitosan, alginate, hyaluronic acid, dextran, polylactic acid, polyglycolic acid, poly(lactic acid-co-glycolic acid), polycaprolactone, polyanhydride, polyorthoester, polyvinyl alcohol, polyethylene glycol, polyurethane, polyacrylic acid, poly-N-isopropylacrylamide, poly(ethylene oxide)-poly(propylene oxide)-poly(ethylene oxide) copolymer, copolymers thereof, or any combination thereof.
[0126] A polymer can comprise polystyrene, neoprene, poly etherether 10 ketone (PEEK), carbon reinforced PEEK, polyphenylene, PEKK, PAEK, polyphenyl sulphone, polysulphone, PET, polyurethane, polyethylene, low-density polyethylene (LDPE), linear low- density polyethylene (LLDPE), high-density polyethylene (HDPE), polypropylene, polyetherketoneetherketoneketone (PEKEKK), nylon, TEFLON® TFE, polyethylene terephthalate (PETE), TEFLON® FEP, TEFLON® PF A, and / or polymethylpentene (PMP) styrene maleic anhydride, styrene maleic acid, polyurethane, silicone, polymethyl methacrylate, polyacrylonitrile, poly (carbonate-urethane), poly (vinylacetate), nitrocellulose, cellulose acetate, urethane, urethane / carbonate, polylactic acid, polyacrylamide (PAAM), poly (N- isopropylacrylamine) (PNIPAM), poly (vinylmethylether), poly (ethylene oxide), poly (ethyl (hydroxyethyl) cellulose), poly(2-ethyl oxazoline), polylactide (PLA), polyglycolide (PGA), poly(lactide-co-glycolide) PLGA, poly(e-caprolactone), polydiaoxanone, polyanhydride, trimethylene carbonate, poly(P-ydroxybutyrate), poly(g-ethyl glutamate), poly(DTH- iminocarbonate), poly(bisphenol A iminocarbonate), poly(orthoester) (POE), polycyanoacrylate (PCA), polyphosphazene, polyethyleneoxide (PEO), polyethylene glycol (PEG), polyacrylacid (PAA), polyacrylonitrile (PAN), polyvinylacrylate (PVA), polyvinylpyrrolidone (PVP), polyglycolic lactic acid (PGLA), poly(2-hydroxypropyl methacrylamide) (pHPMAm), poly(vinyl alcohol) (PVOH), PEG diacrylate (PEGDA), poly(hydroxyethyl methacrylate) (pHEMA), N-I sopropylacrylamide (NIP A), poly(vinyl alcohol) poly(acrylic acid) (PVOH-PAA), collagen, silk, fibrin, gelatin, hyaluron, cellulose, chitin, dextran, casein, albumin, ovalbumin, heparin sulfate, starch, agar, heparin, alginate, fibronectin, fibrin, keratin, pectin, elastin, ethylene vinyl acetate,polyethylene oxide, PEG or any of its derivatives, PLLA, PDMS, PIPA, PEVA, PILA, PEG styrene, Teflon RFE, FLPE, Teflon FEP, methyl palmitate, NIPA, polycarbonate, polyethersulfone, polycaprolactone, polymethyl methacrylate, polyisobutylene, nitrocellulose, medical grade silicone, cellulose acetate, cellulose acetate butyrate, polyacrylonitrile, PLCL, and / or chitosan.
[0127] In some embodiments, the polymer in the RCA mixture can be a zwitterionic polymer. The term “zwitterionic polymer” refers to synthetic or natural polymers comprising one or more structural molecular unit comprising at least a pair of cationic and anionic groups. The cationic groups can include, for example, protonated amino, quaternary ammonium, and pyridine units, while anionic groups can include, for example, carboxylate, sulfonate, and phosphate. In some embodiments, cations in a zwitterionic polymer are quaternized ammonium, and the zwitterionic groups can be classified into sulfobetaine (SB), carboxybetaine (CB), and phosphorylcholine (PC) according to anion. For example, zwitterionic group is SB when anions are sulfonates, CB when anions are carboxylates, and PC when anions are phosphonates. According to the distribution of charge, zwitterionic polymers can be divided into two different structures, one has a positive and negative charge on the same side chain, while in the other, the charges are localized on different side chains. Zwitterionic polymers are generally pH-responsive, highly conductive and responsive to salt.
[0128] In some embodiments, the polymer in the RCA mixture can be a polyelectrolyte species. As used herein, the terms “polyelectrolyte species” or “polyelectrolytes” refer to polymers that, when dissolved in a polar solvent such as water, have a number of charged groups covalently linked to them. Polyelectrolytes can be polyanions, polycations, and polysalts. Branched polyelectrolytes refer to polyelectrolytes having secondary polymer chains linked to a primary backbone, resulting in a variety of polymer architectures such as spherical shaped, IT- shaped, pom-pom and comb-shaped polymers.
[0129] Branched polyelectrolytes include what are generally referred to as “dendrimers” that are repeatedly branched, roughly spherical three-dimensional molecules with nanometer-scale dimensions. Accordingly, the term “dendrimer” used herein refers to repetitively branched molecules having three basic architectural components namely a dendrimer core, repetitive branch cell units and terminal functional groups. A dendrimer core can be a chemical moiety presenting a backbone and at least two anchor atoms, each anchor atom defining a bonding position to a head attachment atom of a branch cell unit. In a dendrimer core, the backbone of the dendrimer core can be any stable chemical moiety having the capability to present anchoring positions for the attachment of branch cell units. In some embodiments, the core backbone structure can be one of aromatic, heteroaromatic rings, aliphatic, or heteroaliphatic rings or chains.In some embodiments, the backbone of the dendrimer core can be one single atom, including but not limited to, C, N, O, S, Si, or P. A “branch cell unit” is a chemical structure presenting one head attachment atom and at least two tail attachment atoms. The head attachment atom defines a bonding position to an anchor atom of a dendrimer core or a tail attachment atom of another branch cell unit. The tail attachment atom defines a bonding position to a head attachment atom of another branch cell unit or to a terminal functional group with the attachment possibly performed directly or indirectly. A generation of branch cell units within a dendrimer defines a shell of the dendrimer as will be understood by a skilled person (see “Dendrimers and other Dendritic polymers” by Jean M. J. Frechet and Donald A. Tomalia 2001). The branch cell units of a generation typically define an interior space inside the dendrimer herein also indicated as interior of shell as will be understood by a skilled person. A “terminal functional group” of a dendrimer is a functional group presented on the outermost part of the dendrimer attached to an end of a branch cell unit. The branch cell units attaching the terminal functional groups typically provide the outer shell or periphery of the dendrimer. Dendrimers can include globular dendrimers, dendrons, hyperbranched polymers, dendrigraft polymers, tecto-dendrimers, core-shell dendrimers, and other types of dendrimers identifiable to a person skilled in the art.
[0130] Dendrimers can be classified by a generation number. The common notation for this classification is GX, where X is a number referring to the generation number. For example, a zero-generation dendrimer is annotated as GO, a first-generation dendrimer is annotated as Gl, and so on. The total number of branch cell units (or number of branches) increases exponentially as a function of generation number. In some embodiments, species of dendrimer are available as Generation 0 (GO) up to Generation 10 (G10) with each generation having double the number of branches from the previous generation. For example, GO PAMAM has 4 branches, Gl PAMAM has 8 branches, and so on. In some embodiments, the dendrimers used herein are GO, Gl, G2, G3, G4, or G5 dendrimers, such as GO, Gl, G2, G3, G4, or G5 PAMAM.
[0131] Dendrimer species can comprise controlled terminal surface chemistry with one or more functional groups that include, but are not limited to, amines, carboxyl, and hydroxyl groups. With different terminal surface groups, dendrimers can be positively charged, negatively charged or neutral.
[0132] The dendrimer herein described can be modified by chemical reactions which modify their functional groups so that they have particular binding properties. For example, the high density of nitrogen or oxygen ligands in these dendrimers, along with the possibility of attaching various functional groups to them, make these dendrimers (e.g., PAMAM, PPI, and PEI) attractive as high capacity chelating agents for metal ions such as the metal ions used In RCA amplification reactions.
[0133] In some embodiments, the dendrimers used herein comprise positively charged terminal surface groups. For example, a branched polyamine that comprises a protonated structure can interact and form complexes with the negatively charged backbone of DNA. It will be appreciated that adaptor elements may be employed with branched polyamines in the embodiments described herein for alternative purposes or to provide improved binding characteristics for the dendrimer species to the nucleic acid.
[0134] Species of dendrimer that can be used in the methods, compositions and systems disclosed herein include, but are not limited to, poly(amidomine) (PAMAM), poly(propylenimine) (PPI), polyethyleneimine (PEI). In some embodiments, the branched poly electrolyte is PAMAM, for example, a G2 PAMAM dendrimer molecule with 16 branches having the amine (NH2) terminal surface chemistry or G3 PAMAM dendrimer molecule with 32 branches having the amine (NH2) terminal surface chemistry. Non-limiting examples of branched polyelectrolyte also include G4 (64 branches with the amine terminal group) and G5 (128 branches with the amine terminal group) PAMAM dendrimer species.
[0135] In some embodiments, a polymer (e.g., polyelectrolyte) in the RCA mixture can be present at a concentration from about 0.001 pM to about 1 M. For example, the concentration of the branched poly electrolyte in a loading buffer can be about, at most, or at most about 0.001 pM, 0.01 pM, 0.02 pM, 0.04 pM, 0.06 pM, 0.08 pM, 0.1 pM, 0.2 pM, 0.3 pM, 04. pM, 0.5 pM, 0.6 pM, 0.7 pM, 0.8 pM, 0.9 pM, 1.0 pM, 10 pM, 20 pM, 30 pM, 40 pM, 50 pM, 60 pM, 70 pM, 80 pM, 90 pM, 100 pM, 200 pM, 300 pM, 400 pM, 500 pM, 600 pM, 700 pM, 800 pM, 900 pM, 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 7 mM, 8 mM, 9 mM, 10 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or a number or a range between any two of these values. In some embodiments, the polymer (e.g., poly electrolyte) in the RCA mixture can be present at a concentration from about 1 mM to about 500 mM.
[0136] Similar to the divalent metal cation, the polymer (e.g., branched polyelectrolyte) can be present at different concentrations at different stages of a rolling circle amplification reaction. In some embodiments, the polymer (e.g., branched polyelectrolyte) in the loading stage can be present at a low concentration necessary to facilitate efficient loading of a polymerase onto a template nucleic acid to form the initial enzyme-template nucleic acid complexes. The polymer (e.g., branched polyelectrolyte) in the amplification stage can be present at a higher concentration that promotes the formation and stabilization of the nucleic acid clusters formed from primer extension reaction as described in great details below.
[0137] The polymer (e.g., poly electrolyte) in the RCA mixture can be present at a concentration necessary to initiate the RCA amplification. In some embodiments, the polymerused in the initial RCA mixture (e.g., the loading buffer) comprises branched polyelectrolyte such as PAMAM (e.g., G3 PAMAM). The concentration of the polymer can vary depending on the type of divalent metal cation in the RCA mixture. For example, when calcium or strontium is present in the RCA mixture, the polymer can be at a concentration from about 1 mM to about 1 M, optionally from about 5 mM to about 500 mM, optionally from about 10 mM to about 100 mM. When magnesium is present in the RCA mixture, the polymer can be at a concentration from about 0.001 pM to about 1 pM, from about 0.01 pM to about 1.0 pM, from about 0.01 pM to about 0.5 pM, or about 0.1 pM.
[0138] The RCA mixture can also include other auxiliary reagents necessary for carrying out a RCA reaction such as salts, buffers, small molecules, co-factors, metals and ion as will be apparent to a skilled person. For example, the RCA mixture can include Tris, Tricine, HEPES, MOPS, ACES, MES, phosphate-based buffers, and acetate-based buffers. In some embodiments, the RCA mixture can include one or more surfactants (e.g., Tween20, NP-40), one or more reducing agents (e.g., dithiothreitol), glycerol, or a combination thereof. The RCA mixture can include salts such as NaCl, KC1, potassium acetate, ammonium acetate, potassium glutamate, NFUCl, or NH4HSO4, which ionize in aqueous solution to yield monovalent cations. The RCA mixture can include chelating agents such as EDTA, EGTA, and the like.
[0139] In some embodiments, the RCA mixture can be replenished with an amplification buffer. The amplification buffer can be introduced under conditions that favor the nucleic acid extension reaction following polymerase loading. In some embodiments, introducing the amplification buffer can occur after a time duration of incubating the RCA mixture with the circularized nucleic acid template. For example, the amplification buffer can be introduced into the vessel 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 minutes after the RCA mixture is introduced into the vessel.
[0140] The amplification buffer can be introduced into the vessel by pumping, drawing, or otherwise flowing different liquid reagents in the amplification buffer sequentially or simultaneously in a combined or separated solution(s) through the vessel (e.g., flow cell). The amplification buffer can be replenished to the vessel as many times as desired to amplify the template nucleic acids to a sufficient copy number. In some embodiments, one or more additional amplification buffers can be sequentially introduced to the vessel. For example, one or more additional amplification buffers can be introduced to the vessel once, twice, or more times following the introduction of the first amplification buffer. Each introduction of an amplification buffer can be separated from the introduction of an additional amplification buffer by a time period, such as, by about 10 minutes to about 60 minutes, and optionally by about 30 minutes. Each of the one or more amplification buffers used herein can have a same or differentcomposition with respect to one another. For example, each of the one or more amplification buffers can have a same or different concentration of the divalent metal cation and / or a same or different concentration of the polymer (e.g., polyelectrolytes).
[0141] Each amplification buffer can be incubated in the vessel with the polymerasenucleic acid complexes for a desired amount of time (for example, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 minutes, or a number or a range between any two of these values) and at any temperature conducive to enzyme activity. The incubation temperature for the replenishment step can be the same as or different from the incubation temperature used for the loading step. The amplification buffer can be incubated in the vessel at a temperature between about 20°C and about 60°C or between about 20°C and about 50°C. Optionally, the reaction temperature is between about 30°C and 40°C (e.g., about 37°C). The reaction temperature can be, for example, 20 °C, 21°C, 22 °C, 23 °C, 24 °C, 25 °C, 26 °C, 27 °C, 28 °C, 29 °C, 30 °C, 31 °C, 32 °C, 33 °C, 34 °C, 35 °C, 36 °C, 37 °C, 38 °C, 39 °C, 40 °C, 41 °C, 42 °C, 43 °C, 44 °C, 45 °C, 46 °C, 47 °C, 48 °C, 49 °C, 50 °C, 51 °C, 52 °C, 53 °C, 54 °C, 55 °C, 56 °C, 57 °C, 58 °C, 59 °C, 60 °C, or a number or a range between any two of these values.
[0142] A reagent removal or wash procedure can be performed between any of a variety of steps set forth herein. For example, a washing step can be included following the loading step and prior to a replenishment step. A washing step can also be performed between any two replenishment steps. A washing step can be used to remove one or more of the reagents that are present in a reaction vessel. The one or more reagents to be removed from the vessel in the washing step can include reagents antagonistic to the polymerase activity, including, for example, enzyme seeding inhibitors and PCR inhibitors within the enzyme storage buffer, such as pyrophosphates and excess DTT, glycerol and surfactants. In some embodiments, the reagent to be removed from the vessel in a washing step can be the excess polymerase in solution. The polymerase can be removed from a vessel under conditions that will wash away the excess polymerase in solution without causing removal of the polymerase bound to the template nucleic acid.
[0143] In some embodiments, the reagent to be removed from the vessel in a washing step is the helicase, the single-stranded DNA binding protein, and / or the ligase. In some embodiments, the method can further comprise removing the helicase, the single-stranded DNA binding protein, and / or the ligase prior to performing RCA. The helicase, the single-stranded DNA binding protein and / or the ligase can be removed prior to introducing an amplification buffer, for example, following a loading step and prior to a replenishment step. In some embodiments, the helicase, the single-stranded DNA binding protein and / or the ligase can be removed prior to the loading step. The helicase, the single-stranded DNA binding protein, and / or the ligase can be removed by washing excess helicase, single-stranded DNA binding protein, and / or ligase. deAccordingly, in some embodiments, the RCA is performed in the absence of a helicase, a singlestranded DNA binding protein, and / or a ligase. Alternatively, the method does not comprise removing the excess helicase, the single-stranded DNA binding protein, and / or the ligase. Therefore, in some embodiments, the RCA is performed in the presence of the helicase, the singlestranded DNA binding protein, and / or the ligase. In some embodiments, helicase, single-stranded DNA binding protein can be added during the RCA process (e.g., during the loading step and / or the replenishment step).
[0144] Delivery of additional polymerase is not necessary in the replenishment step, which provides an advantage in reducing cost and time required to prepare additional polymerase. Therefore, the one or more amplification buffers including the one or more additional amplification buffers do not comprise a polymerase. For example, one or more of the amplification buffer and the additional amplification buffers can have the same composition as the loading buffer except that the amplification buffer and the additional amplification buffers do not comprise the polymerase (e.g. DNA polymerase), thus reducing the total amount of polymerase required for carrying out a RCA reaction. Replenishing an RCA reaction with an amplification buffer can also increase the processivity of a polymerase. Additional advantages of using a replenishment method are described in U.S. Publication No. 2022 / 0356515 Al, the content of which is incorporated herein by reference in its entirety.
[0145] In some embodiments, one or more of the amplification buffers and the additional amplification buffers introduced during a replenishment step have a different composition from the loading buffer as described above.
[0146] The divalent metal cation in the amplification buffer and / or the one or more additional amplification buffers can be present at a concentration in favor of nucleic acid amplification. In some embodiments, the concentration of a divalent metal cation (e.g., a Mg2+) in an amplification buffer is from about 1 mM to about 10 M. For example, the concentration of the divalent metal cation in an amplification buffer can be about, at least, or at least about 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 7 mM, 8 mM, 9 mM, 10 mM, 11 mM, 12 mM, 13 mM, 14 mM, 15 mM, 16 mM, 17 mM, 18 mM, 19 mM, 20 mM, 21 mM, 22 mM, 23 mM, 24 mM, 25 mM, 26 mM, 27 mM, 28 mM, 29 mM, 30 mM, 31 mM, 32 mM, 33 mM, 34 mM, 35 mM, 36 mM, 37 mM, 38 mM, 39 mM, 40 mM, 41 mM, 42 mM, 43 mM, 44 mM, 45 mM, 46 mM, 47 mM, 48 mM,49 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1 M, 2 M, 3 M, 4 M, 5 M, 6 M, 7 M, 8 M, 9 M, 10 M or a number or a range between any two of these values. In some embodiments, the concentration of the divalent metal cation in an amplification buffer is from about 10 mM to about50 mM. In some embodiments, the concentration of the divalent metal cation in an amplificationbuffer is from 10 mM to about 5 M, optionally from about 30 mM to about 1 M. The divalent metal cation can be provided at a series of concentrations in a descending order from, for example, about 1 M to about 30 mM.
[0147] In some embodiments, the concentration of the divalent metal cation in an amplification buffer is about, at least, or at least about 50, 75, 100, 150, 200, 250, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000%, or a number or a range between any two of these values, higher than the concentration of the divalent metal cation in a loading buffer. Optionally, the concentration of the divalent metal cation in an amplification buffer is about, at least, or at least about 100, 200, 300, 400 or 500% higher than the concentration of the divalent metal cation in a loading buffer. In some embodiments, the concentration of the divalent metal cation in an amplification buffer is about, at least, or at least about 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, or a number or a range between any two of these values, higher than the concentration of the divalent metal cation in a loading buffer.
[0148] In some embodiments, different types of divalent metal cations are utilized in the initial RCA mixture (e.g., the loading / seeding buffer) and the amplification buffer. For example, calcium or strontium is used in the initial RCA mixture while magnesium is used in the amplification buffer. In some embodiments, the loading buffer does not comprise magnesium. In some embodiments, calcium or strontium in the initial RCA mixture can be at a concentration from 0.001 mM to about 1 M, optionally from about 0.5 mM to about 200 mM, optionally from about 0.5 mM to about 10 mM. The magnesium in the amplification buffer can be at a concentration from about from about 1 mM to about 10 M, optionally from about 10 mM to about 5 M, optionally from about 30 mM to about 1 M.
[0149] The polymer (e.g., polyelectrolyte) in the amplification buffer and / or the one or more additional amplification buffers can be present at a concentration that favors the nucleic acid amplification and further stabilizes the formed nucleic acid clusters. In some embodiments, the concentration of a polymer (e.g., a branched poly electrolyte such as PAMAM) in an amplification buffer is from about 0.5 pM to about 1 M. In some embodiments, the concentration of a polymer in an amplification buffer is from about 0.5 pM to 20 pM, about 1 pM to 50 pM, or about 2 pM to 100 pM. For example, the concentration of the branched poly electrolyte in an amplification buffer can be about, at least, or at least about 0.5 pM, 1 pM, 2 pM, 3 pM, 4 pM, 5 pM, 6 pM, 7 pM, 8 pM, 9 pM, 10 pM, 11 pM, 12 pM, 13 pM, 14 pM, 15 pM, 16 pM, 17 pM, 18 pM, 19 pM, 20 pM, or a number or a range between any two of these values. In some embodiments, the concentration of a polymer in an amplification buffer is from 10 mM to about 1 M, optionally about 10 mM to about 500 mM. For example, the concentration of the branched poly electrolyte in an amplification buffer can be about, at least, or at least about 10 mM, 20 mM,40 mM, 60 mM, 80 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1 M, or a number or a range between any two of these values. In some embodiments, the concentration of a polymer in an amplification buffer is from about 20 mM to about 50 mM.
[0150] In some embodiments, the concentration of the polymer (e.g., branched polyelectrolyte) in an amplification buffer is about, at least, or at least about 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, or a number or a range between any of these values, higher than the concentration of the polymer (e.g., branched polyelectrolyte) in a loading buffer.
[0151] The amplification buffer used herein can further include reagents that may be exhausted through activity of an extending polymerase so as to facilitate ongoing rolling circle extension, including, for example, dNTPs.
[0152] More than one amplification buffer can be used in the methods, compositions and systems described herein. In some embodiments, the type of the branched polyelectrolyte and / or the concentration of the branched polyelectrolyte can be different in some (e.g., two, three, four, five, or six) or all of the amplification buffers.
[0153] Contacting a template nucleic acid with an RCA mixture in the presence of a primer for a time duration under a condition can form amplified concatemers of the template nucleic acid, such as the concatemers shown in FIG. 2B. In some embodiments, contacting a template nucleic acid with an RCA mixture in the presence of a primer can comprise hybridizing the template nucleic acid and the primer (e.g., through the corresponding primer binding region in the template nucleic acid) to form a primed-template nucleic acid and extending the capture primer along the template nucleic acid.
[0154] The amplified concatemers of the nucleic acid template can comprise two or more copies of the nucleic acid template. For example, the amplified concatemers can comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100 or more copies of the template nucleic acid. The amplified concatemers of the template nucleic acid can be generated as products of primer extension reactions using the nucleic acid template as a template, such as the RCA reaction illustrated in FIG. 2B.Making Concatemers
[0155] Briefly, the method of making (or generating) concatemers can include without limitation the following steps: 1) isolation and processing of nucleic acids, 2) attaching adapter sequence(s) to the nucleic acids and circularizing the template, 3) performing rolling circleamplification (RCA) to create a single stranded concatemer (a first strand or “sense strand”), and optionally 4) performing multiple displacement amplification (MDA) to create a plurality of seconds strands (“antisense strands” or “scales”) of the first (“sense”) strand of the concatemer for signal amplification and / or paired end sequencing. For paired end sequencing, the sense strand may be sequenced before MDA. Alternatively, antisense strands may sequenced and then digested to allow sequencing of the sense strand.1. Nucleic Acid Isolation, Fragmentation, and Size Capture:
[0156] Nucleic acids that are used in a method or composition herein can be deoxyribonucleic acids (DNA), such as, for example, genomic DNA, synthetic DNA, amplified DNA, complementary DNA (cDNA), or the like. Additionally, nucleic acids used in a method or composition herein can also be ribonucleic acids (RNA), such as, for example, mRNA, ribosomal RNA, tRNA, or the like. Further, nucleic acids used in a method or composition herein can also be nucleic acid analogs. For example, a nucleic acid analog can be used as a template for an amplification or sequencing process set forth herein. Nucleic acids used herein, for example, as a template to produce a concatemer or as a target for sequencing, can be derived from a biological source, synthetic source, or amplification product. Primers used herein can include, or can be, DNA, RNA, or analogs thereof.
[0157] A nucleic acid can be obtained from a preparative method such as genome, transcriptome or other nucleic acid isolation, genome fragmentation, gene cloning and / or amplification. One or more nucleic acids can be obtained from an amplification technique such as polymerase chain reaction (PCR), emulsion PCR, random prime amplification, RCA, MDA, or the like. RCA and MDA can be particularly useful for producing concatemeric products. Exemplary methods for isolating, amplifying and fragmenting nucleic acids to produce templates for analysis on an array are set forth in US Pat. Nos. 6,355,431 or 9,045,796, each of which is incorporated herein by reference. Amplification can also be carried out using a method set forth in Sambrook et al, Molecular Cloning: A Laboratory Manual, 3rdedition, Cold Spring Harbor Laboratory, New York (2001) or in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1998), each of which is incorporated herein by reference.
[0158] A nucleic acid template containing a target sequence subject to the methods described herein can be derived or generated from a sample. The sample can include one or more organisms. The nucleic acid template may be obtained or derived from the sample without performing polymerase chain reaction. The nucleic acid template may be obtained or derived from the sample by performing a few cycles of polymerase chain reaction, such as at most one cycle, two cycles, three cycles, four cycles, five cycles, six cycles, seven cycles, eight cycles, nine cycles, or ten cycles or more than ten cycles. Different lengths of nucleic acid templates (or targetsequences) are contemplated herein. Specific nucleic acid templates may be enriched, such as through hybridization capture. A nucleic acid can be at least, be at least about, be at most, or be at most about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410,420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600,610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790,800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980,990, 1000, or a number or a range between any two of these values, nucleotides in length.
[0159] Exemplary organisms from which nucleic acids can be derived include, for example, a mammal, such as, for example, a rodent, mouse, rat, rabbit, guinea pig, ungulate, horse, sheep, pig, goat, cow, cat, dog, primate, human or non-human primate; a plant, such as, for example, Arabidopsis thaliana, corn, sorghum, oat, wheat, rice, canola, or soybean; an algae, such as, for example, Chlamydomonas reinhardtii; a nematode, such as, for example, Caenorhabditis elegans; an insect, such as, for example, Drosophila melanogaster, mosquito, fruit fly, or honey bee; an arachnid, such as, for example, a spider; a fish, such as, for example, a zebrafish; a reptile; an amphibian, such as, for example, a frog or Xenopus laevis; a slime mold, such as, for example, a dictyostelium discoideum; a fungi, such as, for example, pneumocystis carinii, Takifugu rubripes, yeast, Saccharomyces cerevisiae or Schizosaccharomyces pombe; or a plasmodium falciparum. Nucleic acids can also be derived from a prokaryote such as a bacterium, such as, for example, Escherichia coli, Staphylococci or Mycoplasma pneumoniae; an archaea; a virus, such as, for example, Hepatitis C virus or human immunodeficiency virus; or a viroid. Nucleic acids can be derived from a homogeneous culture or population of the above organisms or alternatively from a collection of several different organisms, for example, in a community or ecosystem. Nucleic acid templates may be obtained from any suitable tissue, including blood and solid tissues.
[0160] Nucleic acids can be isolated using methods known in the art including, for example, those described in Sambrook et al, Molecular Cloning: A Laboratory Manual, 3rdedition, Cold Spring Harbor Laboratory, New York (2001) or in Ausubel et al, Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1998), each of which is incorporated herein by reference. Nucleic acids can be fragmented using methods known in the art including, for example, 1) physical methods, such as, for example, acoustic shearing, sonication, or hydrodynamic shearing, or 2) enzymatic methods, such as, for example, DNase I digestion, restriction endonucleases, or transposases, or 3) chemical fragmentation methods, such as, for example, chemical shearing using heat and divalent metal cation(s). Selection of nucleic acids based on ideal fragment length can be performed using methods known in the art including, for example, gel electrophoresis or bead-based size selection. The aforementioned steps lead to theformation of a linear nucleic acid template comprising a target sequence. RNA sample may be converted to cDNA by reverse transcription. Cell free DNA may be obtained from blood or other fluid biological samples. Whole genome samples may be obtained from cells or bulk tissue. In certain aspects, DNA template is a human genome sample.
[0161] DNA templates for the subject circularization methods may be produced as described above, and may be amplified, enriched, and / or ligated to adaptors in an order.2. Attaching Adapter Sequences:
[0162] One or more adapter sequences can be linked to fragmented nucleic acids, such as, for example, a linear nucleic acid template, using methods known in the art, such as, for example, enzymatic methods using a ligase enzyme, and / or commercially available library preparation kits, such as, for example, any of the Nextera DNA Flex Library Prep / Illumina DNA Prep kits (Cat. Nos. 20025519, 20025520, 20018704, and 20018705) or Qiagen QIAseq 1-Step Amplicon Library Kit (Cat. No. 180412). DNA fragments may be end repaired and / or A-tailed before ligation.
[0163] After ligation of at least one adapter sequences to the linear nucleic acid template, a splint oligonucleotide hybridizes to the 5’ and 3’ ends of the linear template and the linear template is ligated to form a circular nucleic acid template comprising a target sequence and one or more adapter sequences as described further herein. Alternatively, a splint oligonucleotide need not be used to ligate the ends, for example, when using CircLigase ™ (Epicenter, Madison WI) or other enzyme capable of splint-free ligation of nucleic acid ends. An exonuclease can be used to remove remaining linear fragments. Single strand DNA templates formed from ligation of adaptors to sample DNA (e.g., DNA fragments) may have 3’ and 5’ ends that hybridize to a splint oligonucleotide for circularization. In addition to regions that hybridized to the splint oligonucleotide, adaptors may provide one or more sequencing primer binding sites, MDA primer binding sites, sample barcodes, and / or molecular barcodes. During sequencing, sequencing primers are hybridized to sequence one or more regions of the sample DNA as well as barcodes of the adaptors.3. Rolling Circle Amplification:
[0164] RCA can be performed using methods known in the art including, for example, those described in Lizardi et al., Nat. Genet. 19:225-232 (1998). Generally, the method involves a polymerase, such as, for example, 029 (phi29) DNA polymerase, extending a primer that is annealed to a circular nucleic acid template such that multiple laps of the polymerase around the circular template produces a concatemeric single stranded nucleic acid (a “sense strand”) that contains multiple tandem repeats, each of the repeats being complementary to the circular nucleic acid template.
[0165] In one configuration, RCA can be performed initially in the presence of a low concentration of a polymer, such as dendrimers (e.g., polyamidoamine (PAMAM)), and subsequently in the presence of a polymer. In one configuration, an RCA reaction is stopped by denaturing the polymerase, for example, by heating the sample at 60°C, 65°C, 70°C, 75°C, 80°C, or more. In one configuration, an RCA reaction is stopped by removing one or more components of RCA, such as the polymerase, and dNTPs. Components of RCA can be removed by, for example, washing. Optionally, one or more antisense strands can be made by replicating the concatemeric sense strand, for example, using multiple displacement amplification (MDA).4. Multiple Displacement Amplification:
[0166] MDA can be performed using methods known in the art including, for example, those described in Lizardi et al., Nat. Genet. 19:225-232 (1998). Generally, primers are hybridized to one or more regions of the concatemeric single stranded nucleic acid, and a polymerase, such as, for example, 029 DNA polymerase, will extend the primers annealed to the concatemeric single stranded nucleic acid to produce a plurality of single stranded nucleic acids (a plurality of “antisense strands”).
[0167] RCA and MDA methods can be carried out isothermally. Generally, the polymerase used for RCA or MDA is a strand displacing polymerase. Methods and reagents that can be used for RCA, MDA, or some combination thereof are set forth, for example, in Lizardi et al., Nat. Genet. 19:225-232 (1998), US Pat. Nos. 6,830,884; 6,797,474; 6,670,126; 6,576,448; 6,323,009; 6,280,949 or US 2007 / 0099208 Al, each of which is incorporated herein by reference.
[0168] The RCA and / or MDA reaction can be carried out in the presence of deoxyribose adenosine triphosphate (dATP), deoxyribose thymidine triphosphate (dTTP), deoxyribose guanosine triphosphate (dGTP), and deoxyribose cytidine triphosphate (dCTP) (or analogues thereof). The antisense strands generated during an MDA reaction can include adenine, guanine, cytosine, and thymine bases. The MDA reaction can be carried out in the presence of deoxyribose uridine triphosphate (dUTP) (or analogues thereof). The antisense strands generated can include uracil bases (indicated by “stars” in the antisense strand) in addition to adenine, guanine, cytosine, and thymine bases when the MDA reaction is carried out in the presence of dUTP in addition to dATP, dTTP, dGTP, and dCTP. Antisense strands with uridine bases can be digested after the antisense strands are sequenced.
[0169] Alternatively, or additionally, the MDA reaction can be carried out in the presence of deoxyribonucleotide triphosphate with a modified base or a non-canonical base such that the antisense strands generated include one or more bases that are modified or non-canonical. Such modified bases or non-canonical bases can target the antisense strands for degradation, such as enzymatic digestion. For example, the MDA reaction can be carried out in the presence ofmodified or non-canonical deoxyribonucleotide triphosphate (e.g., deoxy pseudouridine triphosphate) such that the antisense strands generated include one or more nucleotides that are modified or non-canonical (e.g., deoxyribose pseudouridine monophosphate).
[0170] Whether a base in an antisense strand is a thymine or an uracil (when the corresponding base in the sense strand is adenine) depends on the relative concentration of dTTP and dUTP in the MDA reaction. The concentration of dUTP (or deoxyribonucleotide trisphosphate with a modified base or a non-canonical base, or deoxyribonucleotide trisphosphate that is modified or non-canonical) in an MDA reaction can be lower than the concentration of another deoxyribose nucleotide triphosphate in the MDA reaction such that the percentage of uracil bases (or modified bases or non-canonical bases or nucleotides that are modified or non- canonical) present in the antisense strand is low. The uracil bases (or modified bases or non- canonical bases or nucleotides that are modified or non-canonical) can be randomly distributed and present at a low percentage such that two antisense strands (any two antisense strands) include uracil bases (or modified bases or non-canonical bases or nucleotides that are modified or non- canonical) at different positions.
[0171] The concentration of a deoxyribonucleotide triphosphate (e.g., dATP, dTTP, dGTP, or dCTP), or the concentration of all deoxyribonucleotide triphosphates, in a RCA or MDA reaction can be about, be at least, be at least about, be at most, or be at most about, 0.1 mM, 0.2 mM, 0.3 mM, 0.4 mM, 0.5 mM, 0.6 mM, 0.7 mM, 0.8 mM, 0.9 mM, 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 7 mM, 8 mM, 9 mM, 10 mM, 11 mM, 12 mM, 13 mM, 14 mM, 15 mM, 16 mM, 17 mM, 18 mM, 19 mM, 20 mM, 25 mM, 30 mM, 35 mM, 40 mM, 45 mM, 50 mM, 55 mM, 60 mM, 65 mM, 70 mM, 75 mM, 80 mM, 85 mM, 90 mM, 95 mM, 100 mM, or a number or a range between any two of these values. The concentration of dUTP in an MDA reaction can be, be about, be at least, be at least about, be at most, or be at most about, 0.001 mM, 0.002 mM, 0.003 mM, 0.004 mM, 0.005 mM, 0.006 mM, 0.007 mM, 0.008 mM, 0.009 mM, 0.01 mM, 0.02 mM, 0.03 mM, 0.04 mM, 0.05 mM, 0.06 mM, 0.07 mM, 0.08 mM, 0.09 mM, 0.1 mM, 0.2 mM, 0.3 mM, 0.4 mM, 0.5 mM, 0.6 mM, 0.7 mM, 0.8 mM, 0.9 mM, 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 7 mM, 8 mM, 9 mM, 10 mM, 11 mM, 12 mM, 13 mM, 14 mM, 15 mM, 16 mM, 17 mM, 18 mM, 19 mM, 20 mM, or a number or a range between any two of these values.
[0172] The ratio of the concentration of dUTP (or deoxyribonucleotide trisphosphate with a modified base or a non- canonical base, or deoxyribonucleotide trisphosphate that is modified or non-canonical) relative to the concentration of dTTP (or the concentration of another deoxyribonucleotide triphosphate or the total concentration of deoxyribonucleotide triphosphate other than dUTP or deoxyribonucleotide trisphosphate with a modified base or a non-canonical base or deoxyribonucleotide trisphosphate that is modified or non-canonical) can be, be about, beat least, be at least about, be at most, or be at most about, 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77,1:76, 1:75, 1:74, 1:73, 1:72, 1:71, 1:70, 1:69, 1:68, 1:67, 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60,1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43,1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26,1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1 :6, 1:5, 1:4, 1 :3, 1 :2, or a number or a range between any two of these values.
[0173] The percentage of deoxyribonucleotide triphosphates in the MDA reaction that are dUTP (or deoxyribonucleotide trisphosphate with a modified base or a non-canonical base, or deoxyribonucleotide trisphosphate that is modified or non-canonical) can be, be about, be at least about, be at most, or be at most about, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or a number or a range between any two of these values. Other nucleotides that are cleavable are known to one of skill in the art and may be used in addition, or in place of, dUTP and my have similar ranges of concentration as those described for dUTP. dUTP and / or other cleavable nucleotides can be digested to remove antisense strands and allow for subsequent sequencing of sense strands.Solid-phase Synthesis versus Solution-phase Synthesis of Concatemers
[0174] Synthesizing concatemers can be performed in solution (“solution-phase synthesis”) or attached to a solid surface (“solid-phase synthesis”) as described further herein. Generally, any of the aspects and embodiments of the methods of use described herein can utilize concatemers provided from solution-phase or solid-phase synthesis.1. Solid-phase Synthesis:
[0175] In solid-phase synthesis, one starts with a linear nucleic acid and distinct adapter sequences are ligated to the 5’ and 3’ ends of the linear nucleic acid. Alternatively, one can start with a linear nucleic acid having distinct 5’ and 3’ adapter sequences. The linear nucleic acid having distinct 5’ and 3’ adapter sequences is annealed to a surface-bound oligo (such as a splint oligonucleotide) or primer having regions complementary to the distinct 5’ and 3’ adapter sequences of the linear nucleic acid, and oriented so as to position the linear nucleic 5’ and 3’ ends in close proximity. The 5’ and 3’ ends of the linear nucleic acid are ligated so as to circularize the linear nucleic acid. The surface-bound oligo is then subjected to an extension reaction, so as to have added at its 3 ’ end multiple monomer units of the originally linear nucleic acid via RCA. The result is a concatemer of multimers of the original linear nucleic acid being tethered to the surface.A plurality of second strands, referred to herein as scales, can subsequently be synthesized using multiple strand displacement (MDA).2. Solution-phase Synthesis:
[0176] In solution-phase synthesis, a concatemer can be completely synthesized in solution and then deposited onto a surface for sequencing, or a concatemer can be partially synthesized in solution and then deposited onto a surface to finalize synthesis of the concatemer for sequencing. Alternatively, a concatemer can be completely synthesized in solution and then deposited onto a surface for sequencing.
[0177] One can start with a linear nucleic acid and ligate distinct adapter sequences to the 5’ and 3’ ends of the linear nucleic acid. Subsequently, the linear nucleic acid sequence is circularized by binding to a splint oligo in solution, resulting in a circular template. Alternatively, a splint oligonucleotide need not be used to ligate the ends, for example, when using CircLigase ™ (Epicenter, Madison WI) or other enzyme capable of splint-free ligation of nucleic acid ends. In other configurations, one can start with a linear nucleic acid having distinct 5’ and 3’ adapter sequences and circularize the linear nucleic acid by binding to a splint oligo in solution, resulting in a circular template. In yet another alternate embodiment, one can start with a circular template. The circular template can be annealed to a surface-bound oligo, and RCA can be performed as described above, optionally followed by MDA (in the same solution or after depositing the RCA product on a surface). Alternatively, the circular template can be contacted to a primer and RCA can be performed in solution. Optionally, MDA can be performed to produce a plurality of second strands. Subsequently, the concatemer synthesized in solution can be deposited onto a surface for sequencing.
[0178] Combinations of in solution and solid phase synthesis, such as where the DNA template is circularized in solution and hybridized onto surface primers for surface RCA, are also within the scope of the subject methods, kits and systems. When circularization is performed in solution, remaining single strand DNA may be removed such as by an exonuclease cleanup step.Applications in Sequencing
[0179] The methods disclosed herein can be used in various sequencing platforms, including but not limited to, sequencing-by-synthesis or sequencing-by-binding (sometimes collectively referred to as sequencing-by-incorporation chemistries), pH-based sequencing, sequencing by polymerase monitoring, sequencing by hybridization, and other methods of massively parallel sequencing or next-generation sequencing. In some embodiments, the sequencing is carried out as described in US Pat. No. 10,077,470, which is incorporated byreference herein in its entirety. Suitable surfaces for carrying out sequencing include, but are not limited to, a planar substrate, a hydrogel, a nanohole array, a microparticle, or nanoparticle.
[0180] In some embodiments, RCA herein described can produce a linear concatemeric nucleic acid molecule, which takes the form of a random coil, commonly referred to as a “picosphere.” A picosphere can be immobilized to a surface suitable for sequencing (e.g., via hybridizing to a universal capture oligonucleotide on the surface of a sequencing substrate). The universal capture oligonucleotide has a sequence that is unrelated to any specific target sequence of interest and thus can be used to capture any target sequences. In some embodiments, the universal capture oligonucleotide can hybridize to the universal priming sequence in the picospheres. In some embodiments, the universal capture oligonucleotide is a barcode sequencing primer. In some embodiments, the picospheres is attached to the surface through ionic interactions, via covalent linkages, or mediated through binding of attached ligands (e.g., biotin and streptavidin). In some embodiments, one or several sequencing primers are hybridized to the picosphere before or after attachment to the surface for sequencing.
[0181] Therefore, the methods, compositions and systems disclosed herein for performing a rolling circle amplification can be used in nucleic acid sequencing, for example, in sequencing-by-binding (SBB) or in sequencing-by-synthesis (SBS) methods, compositions and systems.
[0182] “Sequencing-by-binding” refers to a sequencing technique wherein specific binding of a polymerase and cognate nucleotide to a primed template nucleic acid is used for identifying the next correct nucleotide to be incorporated into the primer strand of the primed template nucleic acid. The specific binding interaction need not result in chemical incorporation of the nucleotide into the primer. The specific binding interaction can precede chemical incorporation of the nucleotide into the primer strand or precedes chemical incorporation of an analogous, next correct nucleotide into the primer. Thus, identification of the next correct nucleotide can take place without incorporation of the next correct nucleotide.
[0183] Sequencing by binding has been described, for example, in US Pat. Nos. 10,443,098 and 10,246,744, and US Pat. App. Pub. No. 2018 / 0044727 published on February 15, 2018; the content of each is incorporated herein by reference in its entirety. Briefly, in SBB, the polymerase undergoes conformational transitions between open and closed conformations during discrete steps of a reaction. In one step, the polymerase binds to a primed template nucleic acid to form a binary complex, also referred to herein as the pre-insertion conformation. In a subsequent step, an incoming nucleotide is bound and the polymerase fingers close, forming a pre-chemistry conformation comprising the polymerase, a primed template nucleic acid and a nucleotide, wherein the bound nucleotide has not been incorporated. This step, also referred to herein as anexamination step, is followed by removal of the nucleotide (can be labeled or unlabeled nucleotide) without incorporation, and is then followed by de-blocking of the extended strand 3’ end so as to render it suitable for extension. Unlabeled, 3’ blocked nucleotides are then added, followed by a chemical incorporation step wherein a phosphodiester bond is formed with concomitant pyrophosphate cleavage from the nucleotide (nucleotide incorporation), to form an extension strand that has been extended by one base and that is not competent for further extension without modification. Unincorporated blocked extension bases are removed and labeled bases added, so that they can from ternary complexes at positions where they base pair with the template. These ternary complexes are assayed for fluorescence or other output to determine the identity of the paired base, and then the process is repeated through removal of the labeled base, chemical modification of the extending strand to reveal a 3’ OH, and contacting with a population of 3 ’blocked, unlabeled nucleotides for another single base extension.
[0184] The examination step can, for example, involve providing a primed template nucleic acid and contacting the primed template nucleic acid with a polymerase (e.g., a DNA polymerase) and one or more test nucleotides being investigated as the possible next correct nucleotide. The polymerase configuration and / or interaction with the primed template nucleic acid and further with a nucleotide can be monitored during an examination step to identify the next correct base in the template nucleic acid. In some embodiments, the SBB procedure includes a monitoring step that monitors or measures the interaction between the polymerase and the primed template nucleic acid in the presence of the test nucleotides. In some embodiments, the examination step determines the identity of the next correct nucleotide without requiring incorporation of that nucleotide (e.g., either without, or before chemical linkage of that nucleotide to the 3 ’-end of the primer through a covalent bond). For example, the primer of the primed template nucleic acid molecule can include a blocking group that precludes enzymatic incorporation of an incoming nucleotide into the primer. In some embodiments, the reaction mixture used in the examination step comprises catalytic metal ions at a low or deficient level to prevent the chemical incorporation of the nucleotide into the primer of the primed template nucleic acid. In some embodiments, the reaction mixture used in the examination step comprises a stabilizer that stabilize ternary complexes while precluding incorporation of any nucleotide into the primer, such as a non-catalytic metal ion that inhibits polymerization.
[0185] Generally, an examination step involves binding a polymerase to the polymerization initiation site of a primed template nucleic acid in a reaction mixture comprising one or more nucleotides, and monitoring the interaction. An examination step can, for example, include one or more of the following substeps: (1) providing a primed template nucleic acid (i.e., a template nucleic acid molecule hybridized with a primer that optionally may be blocked fromextension at its 3 ’-end); (2) contacting the primed template nucleic acid with a reaction mixture that includes a polymerase and at least one nucleotide; (3) monitoring the interaction of the polymerase with the primed template nucleic acid molecule in the presence of the nucleotide(s) and without chemical incorporation of any nucleotide into the primed template nucleic acid; and (4) determining from the monitored interaction the identity of the next base in the template nucleic acid (i.e., the next correct nucleotide). Examination typically involves detecting polymerase interaction with a template nucleic acid. Detection may include optical, electrical, thermal, acoustic, chemical and mechanical means. The examination step of the sequencing reaction can repeat 1, 2, 3, 4 or more times prior to the optional incorporation step.
[0186] In SBB, a reaction mixture used in the examination step can include 1, 2, 3, or 4 types of nucleotide molecules. The nucleotides can be selected from dATP, dTTP (or dUTP), dCTP, and dGTP. The examination reaction mixture can comprise one or more triphosphate nucleotides and one or more diphosphate nucleotides. A ternary complex can form between the primed template nucleic acid, the polymerase, and any one of the four nucleotide molecules so that four types of ternary complexes may be formed.
[0187] An incorporation step can be concurrent with or separate from the examination step. In some embodiments of an SBB procedure, the examination step is followed by an incorporation step that adds one or more complementary nucleotides to the 3’ end of the primer component of the primed template nucleic acid. The polymerase, primed template nucleic acid and newly incorporated nucleotide produce a post-chemistry conformation. Both pre-chemistry conformation and the post-chemistry conformation can be referred to as a ternary complex, each comprising a polymerase, a primed template nucleic acid and a nucleotide, wherein the polymerase is in a closed state and facilitates interaction between a next correct nucleotide and the primed template nucleic acid. During the incorporation step, divalent catalytic metal ions, such as Ca2+, Sr2, or Mg2+, mediate a chemical step involving nucleophilic displacement of a pyrophosphate (PPi) by the 3 ’ -hydroxyl of the primer terminus. The polymerase returns to an open state upon the release of PPi.
[0188] The incorporation step may be facilitated by an incorporation reaction mixture. The incorporation reaction mixture can have a different composition of nucleotides than the examination reaction. For example, the examination reaction can include one type of nucleotide and the incorporation reaction can include another type of nucleotide. By way of another example, the examination reaction comprises one type of nucleotide and the incorporation reaction comprises four types of nucleotides, or vice versa. The examination reaction mixture can be altered or replaced by the incorporation reaction mixture.
[0189] In some embodiments, an examination step is followed by removal of the labeled nucleotide without being incorporated, and is then followed by de-blocking of the 3’ end of the primer (or extended primer) of the primed template nucleic acid so as to render it suitable for extension. Unlabeled, 3’ blocked nucleotides are then added, followed by a chemical incorporation step wherein a phosphodiester bond is formed with concomitant pyrophosphate cleavage from the nucleotide (nucleotide incorporation), to form an extension strand that has been extended by one base and that is not competent for further extension without modification. Unincorporated blocked extension bases are removed and labeled bases added, so that they can from ternary complexes at positions where they base pair with the template. These ternary complexes are assayed for fluorescence or other output to determine the identity of the paired base, and then the process is repeated through removal of the labeled base, chemical modification of the extending strand to reveal a 3’ OH, and contacting with a population of 3’ blocked, unlabeled nucleotides for another single base extension.
[0190] In some embodiments, the methods, compositions and systems disclosed herein can be used in one or more steps of a SBB procedure that involves nucleotide incorporation. For example, the methods, systems, and compositions herein disclosed can be used in the incorporation step, either following or concurrent with the examination step, of a SBB procedure to allow nucleotide incorporation and primer extension.
[0191] In some embodiments, a SBB procedure uses two different reaction mixtures: an examination reaction mixture in the examination step and an incorporation reaction mixture in the incorporation step. The reaction mixtures typically include reagents that are commonly present in polymerase based nucleic acid synthesis reactions. Reaction mixture reagents can include, but are not limited to, enzymes (e.g., polymerase), dNTPs, template nucleic acids, primers, salts, buffers, small molecules, co-factors, metals, and ions.
[0192] The incorporation reaction mixture can, for example, comprise one or more nucleotides (e.g., same or different types) and polymerase extension reagents including a divalent metal cation and a polymer (e.g., a branched polyelectrolyte), in which one or both of the divalent metal cation and the branched polyelectrolyte is in a higher concentration as compared to the examination reaction mixture. For example, the concentration of the divalent metal cation in the incorporation reaction mixture is about, at least, or at least about 1-fold, 2-fold, 3-fold, 4-fold, 5- fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, or a number or a range between any two of these values, higher than the concentration of the divalent metal cation in the examination reaction mixture. In some embodiments, the concentration of the branched polyelectrolyte in the incorporation reaction mixture is about, at least, or at least about 2-fold, 3-fold, 4-fold, 5-fold, 6- fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold,90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, or a number or a range between any two of these values, higher than the concentration of the branched poly electrolyte in the examination reaction mixture. In some embodiments, the concentration of a divalent metal cation in the incorporation reaction mixture is from 0.001 mM to about 1 M, optionally from about 0.5 mM to about 200 mM. In some embodiments, the concentration of a branched poly electrolyte (e.g., a PAMAM) in the incorporation reaction mixture is from about 0.001 pM to about 1 M. In some embodiments, different divalent metal cation is used in the incorporation reaction mixture and in the examination reaction mixture. For example, the examination reaction mixture can comprise Ca2+or Sr2+, while the incorporation reaction mixture can comprise Mg2+. In some embodiments, the incorporation reaction mixture does not comprise a polymerase.
[0193] Accordingly, in some embodiments, the method of sequencing-by-binding can comprise replacing the examination reaction mixture used in the examination step with the incorporation reaction mixture herein described in the incorporation step. In some embodiments, the method of sequencing-by-binding can comprise washing the immobilized primed template nucleic acid molecule to remove one or more components of the examination reaction mixture (e.g. excess polymerase and reagents antagonistic to the polymerase activity) before introducing the incorporation reaction mixture.
[0194] The examination and incorporation reaction mixtures used in the methods, systems, and compositions of a SBB procedure can include other molecules or reagents generally present in a nucleic acid polymerization reaction. Description of SBB reaction mixtures and related methods and uses in a SBB procedure can be found, for example, in US Pat. Nos. 10,443,098 and 10,246,744, each of which is incorporated herein by reference.
[0195] In some embodiments, the methods, systems, and compositions herein disclosed can be used to produce (or synthesize) one or more strands of a nucleic acid concatemer for sequencing-by-binding. The one or more strands of a nucleic acid concatemer can be produced using a rolling circle amplification method herein described with a loading buffer and an amplification buffer having different compositions. For example, a rolling circle amplification reaction can be initiated with a loading buffer having a divalent metal cation and a branched polyelectrolyte in a low concentration and then replenished with an amplification buffer having the divalent metal cation and the branched polyelectrolyte in a high concentration. In some other embodiments, the loading buffer and the amplification buffer can use different divalent metal cation, such as Ca2+or Sr2+for the loading buffer and Mg2+for the amplification buffer. The branched polyelectrolyte can be provided at an increased concentration as described above in the rolling circle amplification section. Delivery of additional polymerase is not necessary in any of the replenishment steps, which provides an advantage in reducing cost and time required toprepare additional polymerase. The produced concatemers can then be sequenced by hybridizing a sequencing primer to a primer binding site in a sequence unit of a concatemer and extending the primer along the concatemer to determine the sequence. In some embodiments, one or more strands of a nucleic acid concatemer can form a nucleic acid cluster, which can be stabilized through the interaction between the positive charges carried by the branched polyelectrolyte and the negative charges carried by the nucleic acid cluster. The nucleic acid cluster can comprise a plurality of first strand concatemer strand along with a plurality of second strand concatemers that are complementary to the first strand concatemers. The second strand concatemers can be produced, for example, by rolling circle amplification or by multiple displacement amplification performed on the first strand concatemers. Any one of the strands produced can be sequenced. For example, the first strand of concatemer and the one or more second strands of concatemers can be sequenced sequentially or concurrently with different sequencing primers.
[0196] The methods, systems, and compositions herein disclosed can also be used in sequencing-by-synthesis (SBS). SBS generally involves the enzymatic extension of a nascent primer through the iterative addition of nucleotides against a template strand to which the primer is hybridized. SBS differs from SBB, above, in that labeled nucleotides are incorporated into the extending strand, assayed and then the label is removed or deactivated, and the 3’ block removed, to iteratively sequence a template. In SBB, a labeled base does not need to be incorporated into an extending strand. Rather, ternary complex formation is assayed, usually for the presence of a labeled base but sometimes for the presence of a labeled polymerase or other feature, after which point the complex is disassembled and a 3’ blocked, unlabeled base is used to extend the primer strand. Briefly, SBS can be initiated by contacting target nucleic acids, attached to sites in a flow cell, with one or more labeled nucleotides, DNA polymerase, etc. Those sites where a primer is extended using the target nucleic acid as template will incorporate a labeled nucleotide that can be detected. Detection can include scanning using an apparatus or method set forth herein. Optionally, the labeled nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analog having a reversible terminator moiety can be added to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the vessel (before or after detection occurs). Washes can be carried out between the various delivery steps. The cycle can be performed n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, reagents and detection components that can be readily adapted for use with a method, system or apparatus of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; WO91 / 06678; WO 07 / 123744; U.S. Pat. Nos. 7,057,026; 7,329,492; 7,211,414; 7,315,019 or 7,405,281, and US Pat. App. Pub. No. 2008 / 0108082 Al, each of which is incorporated herein by reference.
[0197] Accordingly, a method of sequencing-by-synthesis can comprise introducing into a vessel (e.g., a flow cell) a replenishment mixture. The replenishment mixture can comprise one or both of the divalent metal cation and the branched polyelectrolyte at a high concentration herein described (e.g., the concentrations used in an amplification buffer). Alternatively, the replenishment mixture can comprise a divalent metal cation different from the divalent metal cation used in the loading buffer. For example, a method of sequencing-by-synthesis can comprise hybridizing a sequencing primer to a primer binding site in a sequence unit of a concatemer and extending the sequencing primer along the concatemer to determine the sequence of a template nucleic acid of the concatemer. The replenishment mixture can be introduced after the sequencing primer is hybridized to the primer binding site in the sequence unit of the concatemer and before the primer extension takes place. In some embodiments, the replenishment mixture is introduced after the primer extension takes place. In some embodiments, the primer extension can include repeated cycles of adding a reversibly terminated nucleotide to the sequencing primer and deblocking the reversibly terminated nucleotide on the sequencing primer. The replenishment mixture can be delivered to the vessel as many times as desired to accommodate the number of cycles required to sequence the concatemer. Delivery of additional polymerase is not necessary in any of the replenishment steps.
[0198] In some embodiments, a SBS procedure is initiated with an initiation mixture comprising one or both of a divalent metal cation and a branched polyelectrolyte at a low concentration herein described (e.g., the concentrations used in a loading buffer), and then replenished with the replenishment mixture comprising one or both of the divalent metal cation and the branched poly electrolyte at a high concentration herein described (e.g., the concentrations used in an amplification buffer). For example, in some embodiments, the concentration of the divalent metal cation (e.g., Mg2+) of the replenishment mixture used in one or more steps of a SBS procedure can be from about 1 mM to about 10 M, optionally from about 10 mM to about 5 M. In some embodiments, the concentration of the branched polyelectrolyte (e.g., a G3 PAMAM) of the replenishment mixture used in one or more steps of a SBS procedure can be from about 1 pM to about 1 M, optionally from about 10 mM to about 1 M, more optionally about 10 mM to about 500 mM. In some embodiments, the replenishment mixture used in the SBB procedure does not comprise a polymerase. In some embodiments, the method of sequencing-by-synthesis can comprise, after the primer hybridization, replacing or altering the initiation mixture with the replenishment mixture or one or more components of the replenishment mixture to achieve adesired concentration of the one or more components (e.g., the divalent metal cation and / or the branched polyelectrolyte) as described herein.
[0199] Similar to SBB described above, the methods, systems, and compositions herein disclosed can also be used to produce (or synthesize) one or more strands of a nucleic acid concatemer for sequencing-by-synthesis.
[0200] Sequencing methods other than SBB and SBS are also in scope of the subject methods, kits and systems. For example, some alternative sequencing methods can be used for long read or for short read sequencing. Non-limiting examples of third generation sequencing include: single molecule real-time sequencing (e.g., Pacific Biosciences’ SMRT sequencing) and nanopore sequencing (e.g., Oxford Nanopore Technologies’ nanopore sequencing). In general, single molecule real-time sequencing can use a circular nucleic acid or a concatemer as a template while nanopore sequencing would generally use a concatemer but would not be compatible with direct sequencing of a circular template. Exemplary nanopore systems are described in Bayley, Hagan. “Nanopore sequencing: from imagination to reality.” Clinical chemistry 61.1 (2015): 25- 31. Exemplary single molecule real-time sequencing of a circular DNA template is described in US Publication No. 20090029385, the content of which is incorporated herein by reference.Systems and Kits
[0201] Provided herein also include systems and kits comprising reagents and components for performing one or more steps of the methods described herein. In some embodiments, provided herein include systems and kits for performing linearization and / or circularization of a single-stranded nucleic acid template and for performing rolling circle amplifications of nucleic acids. Systems disclosed herein can include a vessel, solid support or other apparatus for carrying out the methods. For example, the system can include an array, flow cell, multi-well plate, test tube, channel in a substrate, collection of droplets or vesicles, tray, centrifuge tube, tubing or other convenient apparatus. The apparatus can be removable, thereby allowing it to be placed into or removed from the system. As such, a system can be configured to process a plurality of apparatus (e.g., vessels or solid supports) sequentially or in parallel. The system can include a fluidic component configured to deliver one or more reagents set forth herein (e.g., polymerase, primer, template nucleic acid, nucleotides, loading buffer, amplification buffer, or mixtures of such components). The fluidic system can be configured to deliver reagents to a vessel or solid support, for example, via channels or droplet transfer apparatus (e.g., electrowetting apparatus). Any of a variety of detection apparatus can be configured to detect the vessel or solid support where reagents interact. Exemplary systems having fluidic and detection components those set forth in US Pat. App. Pub. No. 2018 / 0280975A1 published on October 4, 2018; U.S.Pat. Nos. 8,241,573; 7,329,860 or 8,039,817; or US Pat. App. Pub. Nos. 2009 / 0272914 Al published on November 5, 2009 or 2012 / 0270305 Al published on October 25, 2012, each of which is incorporated herein by reference.
[0202] The compositions described herein can be packaged together as a kit for performing any of the methods disclosed herein. In some embodiments herein disclosed, the kits can contain one or more components of the linearization reagents, the circularization mixture, the rolling circle amplification mixture, the loading buffer, and the amplification buffer as disclosed above. In some embodiments herein disclosed, the kits can contain one or more components for carrying out the linearization and / or circularization of a single-stranded nucleic acid template. For example, the kits can a helicase, a splint oligonucleotide, a ligase and optionally a single-stranded DNA binding protein. Kits may further include a kinase, such as a T4 Polynucleotide Kinase, for phosphorylating the 5 ’ end of the DNA template to allow for ligation. In some embodiments herein disclosed, the kits can contain one or more components for carrying out the rolling circle amplification. For example, the kits can contain a first buffer (e.g., a loading buffer) comprising a first divalent metal cation (e.g., Ca2+or Sr2+) and a polymer (e.g., branched polyelectrolyte) described herein. The loading buffer can further comprise polymerases, buffers, reagents and substrate solutions for carrying out a rolling circle amplification reaction. The kits can also contain a second buffer (e.g., an amplification buffer) comprising a second divalent metal cation (e.g., Mg2+) and a polymer (e.g., branched poly electrolyte) described herein. The kits may contain additional reagents suitable for the detection, purification, and further processing of the amplified nucleic acids (e.g., concatemers) in downstream applications (e.g., sequencing). The systems and kits can contain additional systems and reagents suitable for sequencing. The kits can contain the compositions in separate containers. The kits can include one or more of appropriate packaging materials, containers for holding the components of the kit, and instructional materials for practicing the methods herein disclosed instructions.EXAMPLES
[0203] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Sequencing was performed on the Pacific Biosciences Onso Sequencing System using the commercially available Onso Sequencing Kit, which includes a clustering reagent pack and a sequencing reagent pack. The clustering reagents of the Onso Sequencing Kit were modified with helicase and SSB as described and indicated below.Example 1Helicase and DNA binding protein in circularization improves GC bias
[0204] Previous bioinformatics analysis suggested a GC bias likely would be introduced during the RCA reactions. Without being bound by any particular theory, it is hypothesized that the ring making (i.e., circularization / CIR) step is likely a major contributing step to introduce Type I GC bias. Library molecules with high GC% is less likely to become rings due to secondary structures, hence high GC% spots are underrepresented in the cluster population. In this example, experiments were carried out to investigate if adding additives, such as helicase and DNA binding protein, to the circularization step can reduce Type I GC bias.
[0205] The Onso clustering reagent pack of an Onso Sequencing Kit was modified by spiking 2 nM helicase (Tie UvrD) and 1 pM SSB (T4 gp32) into a circularization reagent of the Onso clustering reagent pack. PA (Pseudomonas aeruginosa) library was sequenced on the Onso Sequencing System with modified and unmodified Onso clustering reagent packs for comparison. FIG. 3 illustrates IGV windows showing two regions of the Pseudomonas aeruginosa genome with especially high %GC (A: 75% and B: 72%). Gray, red and blue bars indicate sequenced reads that are mapped to these genome regions. The upper panels reflect sequencing data collected from clustering conditions that include a ligation mixture containing Helicase and single-stranded DNA binding protein (SSB), while the lower panels reflect sequenced reads collected under the same conditions, but without the Helicase and SSB. The presence of Helicase and SSB during ssDNA library ligation significantly improved sequencing coverage for both high %GC loci.
[0206] FIG. 4 provides plots demonstrating that helicase and single-stranded DNA binding protein can reduce background signal Upper panels - L01 / HS: circularization mixture with helicase (UvrD) and SSB (gp32) present and L02 / SOP: circularization mixture without helicase and SSB. Cycle 1 shows the raw sequencing images (Exam G) at the beginning of Read 1. Cycle 153 shows the raw sequencing images (Exam G) at the beginning of Read 2. The space between clusters (background) in L01 / HS panels on the left is darker than the L02 / SOP panels on the right for both Reads (Cycle 1 and 153). Lower panels - A histogram representation of the intensity along the yellow lines from the images above, confirming the background intensity is lower in the left L01 / HS panels than the L02 / SOP panels on the right. The data suggests that helicase and SSB can maintain the linearity of a single-stranded DNA template and reduce the likelihood for a single-strand DNA template to hybridize to a splint molecule with only one end (either 5’ or 3’) of the adapter, resulting in a reduction of the background signal.
[0207] An in-depth GC bias analysis was also performed for twO 30X mean coverage HG002 WGC data sets: without helicase and SSB (-Helicase and SSB) and with helicase and SSB (+ Helicase and SSB) conditions. The “+Helicase and SSB” condition comprises RCA conditionssuch as library capture / hybridization at 55 °C, library circulation for 40 min in the presence of helicase and SSB, and 30 mM MgCl2+during scaling (2ndstrand primer extension). An exemplary clustering condition is provided below in Table 1.Table 1. Exemplary Clustering conditions.
[0208] Conditions can be optimized with helicase and SSB to improve coverage and / or other sequencing metrics. In general, such conditions are base on the specific sequencing system. However, the below conditions have been found to be useful for achieving optimal sequencing metrics and could be further optimized for any sequencing system described herein. For example, optimization (such as increasing) temperature of DNA template capture on a surface splint oligonucleotide can increase coverage, reduce GC bias and / or reduce locus bias. Surface circularization may be at a temperature of more than 25 degrees, more than 37 degrees, or more than 45 degrees Celsius, such as at or between 37 degree and 65 degrees Celsius, at or between 45 degrees and 65 degrees, at or between 45 degrees and 60 degrees, at or between 50 degrees and 60 degrees, or at or between 50 and 60 degrees Celsius. In certain aspects, use of helicase and / or SSB and other optimizations may trade off sequencing metrics required input library with improved coverage. Optimizing the duration of incubation for circularization, for example to more than 20 minutes, more than 30 minutes, such as at or between 30 minutes and 60 minutes, or at or between 30 minutes and 50 minutes, may reduce average full width half maximum (FMHM), for example without reducing signal, may reduce GC bias, increase clonality, and / or reduce OFF intensity (background) increase over cycles. These optimizations may combine with the addition of helicase and SSB to provide an improvement in at least one of coverage, GC bias and locus bias, increased spot signal intensity, reduced CV of ON intensity (signal), and / or higher throughput (higher density of resolvable clusters and / or reads above a certain quality score, suchas reads with at least 90% Q30+ or Q40+). As such, the combination of addition of helicase and / or SSB along with additional condition optimizations may maintain or improve throughput, may reduce GC bias, increase clonality, and / or reduce OFF intensity (background) increase over cycles. These optimizations may combine can increase coverage, reduce GC bias and / or reduce locus bias and may further maintain or increase average spot signal intensity or uniformity and / or throughput.
[0209] In addition, conditions (e.g., reagent concentrations) of steps downstream of circularization such as rolling circle amplification and / or MDA can be optimized in a similar manner. For example, increasing the concentration of magnesium (MgCh) can improve spot intensity, reduce GC bias, reduce FWHM, and / or lower CV of signal across clusters (ON CV). Exemplary MgCh concentrations for amplification and / or scaling steps may be greater than 3 mM, such as 5-100 mM or 10-50 mM or 10-100 mM. Additional non-catalytic cations may also be screened for achieving preferred scale intensity. For example, inhibiting metal ions can reduce signal (e.g., reduce size of clusters or scales) and include, but are not limited to, calcium, strontium, scandium, titanium, vanadium, chromium, iron, cobalt, nickel, copper, zinc, gallium, germanium, arsenic, selenium, rhodium, europium and terbium ions. Non-inhibitory non-catalytic ions, such as lithium, my increase signal (e.g., increase size of clusters or scales). Various reagents for improving stability of ternary complex such as lithium, betaine and / or sucrose may be used in any of the clustering steps (e.g., circularization, rolling circle amplification, and / or MDA) and generally described in US patent publication no 20190345544, which is incorporated herein by reference. Exemplary RCA and replenishment for RCA are described in US patent publication no 20220356515, which is incorporated herein by reference. In an initial “seeding” step, low magnesium, or a non-catalytic cation such as calcium or strontium, may be used to reduce or prevent initial RCA and provide uniformity of cluster size (e.g., as measured by FWHM, ON signal). Subsequent replenishment with an RCA solution having increase magnesium can be performed after the seeding step. Exemplary MDA for sequencing applications are described in US patent publication no 20220349002, which is incorporated herein by reference. MDA formation of scales may be performed after sequencing the “sense” strand, or scales be sequenced first and then digested before sequencing of the “sense” strand. Other nucleotides that are cleavable are known to one of skill in the art and may be used in addition, or in place of, dUTP.
[0210] FIGS. 5A-B provide Picard GC bias plots under these two conditions. Picard GC bias plots show normalized sequencing coverage per %GC. The black distribution reflects the expected coverage distribution, while the red reflects the observed distribution. Blue circles show the normalized frequency of coverage, where a value =1 indicates a perfect correlation between observed and expected. Green line plots the observed base quality score for each GC bias bin.Figures were generated from HG002 PEI 50 sequencing data that used clustering conditions without (A) and with (B) helicase and SSB. FIG. 5C summarizes the sum difference between observed and expected base quality scores as an absolute value for both conditions. A significantly improved GC bias is observed in the clustering condition with helicase and SSB (0.028741401 vs. 0.066476851). GC bias is calculated as the combined absolute values of the difference of the reference normalized read frequency distribution of the sequenced compared to a reference such as a human genome reference (e.g., HG38, a published human reference genome). This is observed as the areas of non-overlap between the two bottom solid lines shown in FIGS. 5A-B. The solid black line is the read frequency of HG38, and red line is read frequency of obtained from human genome sample sequenced with the subject methods. Addition of helicase and SSB (FIG. 5B) provided a GC bias of less than 0.03, a more than 2-fold reduction in GC bias compared to the control (FIG. 5A).
[0211] FIGS. 6A-B are non-limiting Picard GC bias plots showing normalized sequencing coverage per %GC. The black distribution reflects the expected coverage distribution, while the red reflects the observed distribution. Blue circles show the normalized frequency of coverage, where a value =1 indicates a perfect correlation between observed and expected. Green line plots the observed base quality score for each GC bias bin. Picard figures (FIGS. 6A and 6B) demonstrate (FIG. 6B) (multiple reagent improvements in addition to Helicase / SSB) provides a significant improvement GC bias compared to SOP (FIG. 6A). FIG. 6C provides Fl scores derived from Hap.py analysis of sequencing coverage SNP and Indel sites from the human genome. Fl scores are reported from M10 (FIG. 6A), Mi l (FIG. 6B) and publicly available data from third party sequencing platforms. The F-score is a measure of the accuracy of a classifier. It is calculated from the precision and recall of the test, where the precision (or positive predictive value) is the number of true positive results divided by the number of all positive results, including those not identified correctly, and the recall (or sensitivity) is the number of true positive results divided by the number of all samples that should have been identified as positive. As applied to calling variants in the DNA sequence of a particular sample, a false positive corresponds to the case in which a sequence variant was incorrectly called relative to a genome reference. Likewise, a false negative call corresponds to the case where a sequence variant was incorrectly not called. Finally, a true positive call corresponds to the case where a sequence variant was correctly called. Fl score is commonly used to determine sequence quality. A Fl score of over 0.993 was observed for SNPs, and a score of over 0.987 was observed for indels using the subject methods.
[0212] Table 2 below provides the overall performance of variant calling under exemplary experimental conditions with (+helicase) or without helicase (-helicase). The resultsinclude performances of SNP and Indel calling and Fl score was computed from Precision and Recall (2 * Precision * Recall / (Precision + Recall)).Table 2. Precision, Recall, and Fl score metrics for variant calling performance.
[0213] The results demonstrated improved accuracy measured by higher overall Fl scores in the experimental conditions with helicase compared to the corresponding conditions without the helicase present. For example, the SNP Fl score of 20230528_Ml l_+Helicase_HG002_BB7_PE150_30X (with helicase; Fl score: 0.993021) outperformed 20230622 M1 l_-Helicase_HG002_BB7_PE150_30X (without helicase; Fl score: 0.991908). Similarly, the INDEL Fl score of 20230528_Ml l_+Helicase_HG002_BB7_PE150_30X (with helicase; Fl score: 0.938344) also outperformed 20230622 M1 l_-Helicase_HG002_BB7_PE150_30X (without helicase; Fl score: 0.935936).
[0214] Exemplary reagents and buffers used for the methods described herein are provided in Tables 3-6 below.Table 3. Exemplary circularization buffer.
[0215] In at least some of the previously described embodiments, one or more elements used in an embodiment can interchangeably be used in another embodiment unless such a replacement is not technically feasible. It will be appreciated by those skilled in the art that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter, as defined by the appended claims.
[0216] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0217] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.). It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, andC” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms.
[0218] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0219] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.
[0220] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
WHAT IS CLAIMED IS:
1. A method of circularizing a single-stranded DNA template, comprising: maintaining linearity of a single-stranded DNA template through the activity of a helicase; and circularizing the single-stranded DNA template by ligation on a splint oligonucleotide in a circularization mixture comprising a ligase for a time duration, wherein the single-stranded DNA template comprises a 5’ adapter and a 3’ adapter and the splint oligonucleotide comprises a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter, thereby allowing the 5 ’ adapter and the 3 ’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template.
2. The method of claim 1, wherein maintaining linearity of the single-stranded DNA template and circularizing the single-stranded DNA template are performed in a single step.
3. The method of claim 1 or 2, wherein circularizing the single-stranded DNA template occurs in the presence of a single-stranded DNA binding protein bound to the singlestranded DNA template.
4. The method of claim 3, wherein maintaining linearity of the single-stranded DNA template is performed in the presence of the single-stranded DNA binding protein, and wherein the single-stranded DNA binding protein binds to the single-stranded DNA template and facilitates association of the helicase with the single-stranded DNA template.
5. The method of claim 3 or 4, wherein circularizing the single-stranded DNA template is performed in the presence of the single-stranded DNA binding protein, and wherein the single-stranded DNA binding protein binds to the single-stranded DNA template.
6. The method of any one of claims 1-5, wherein the circularization mixture comprises the helicase.
7. The method of any one of claims 1-6, wherein maintaining linearity of the singlestranded DNA template further comprises preventing and / or disrupting at least one of secondary structure, coiling, reannealing, and hybridization of adaptors across DNA templates.
8. The method of any one of claims 1-7, further comprising, prior to maintaining linearity of the single-stranded DNA template, denaturing a DNA template to form the singlestranded DNA template, wherein the DNA template is at least partially double-stranded.
9. The method of any one of claims 1-8, wherein the method is performed for a plurality of DNA templates, each comprising a 5’ adapter and a 3’ adapter, at least a portion of which is complementary to the splint oligonucleotide.
10. A method of circularizing a single-stranded DNA template, comprising:circularizing a single-stranded DNA template on a splint oligonucleotide in a circularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration, wherein the single-stranded DNA template comprises a 5’ adapter and a 3’ adapter and the splint oligonucleotide comprises a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter, thereby allowing the 5 ’ adapter and the 3 ’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template.
11. The method of any one of claims 1-10, further comprising: performing rolling circle amplification (RCA) by contacting the circularized DNA template with a RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template.
12. A method of rolling circle amplification (RCA) for nucleic acids, comprising:(a) circularizing a single-stranded DNA template on a splint oligonucleotide in a circularization mixture comprising a helicase, a single-stranded DNA binding protein and a ligase for a time duration, wherein the single-stranded DNA template comprises a 5’ adapter and a 3’ adapter and the splint oligonucleotide comprises a first portion complementary to at least a portion of the 5’ adapter and a second portion complementary to at least a portion of the 3’ adapter, thereby allowing the 5’ adapter and the 3’ adapter to hybridize to the splint oligonucleotide and be ligated by the ligase to form a circularized DNA template; and(b) performing RCA by contacting the circularized DNA template with an RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template.
13. The method of any one of claims 1-12, wherein the single-stranded DNA template has a high GC content or a high AT content.
14. The method of any one of claims 1-13, wherein the single-stranded DNA forms a secondary structure prior to being in contact with the helicase and / or the single-stranded DNA binding protein, optionally the secondary structure comprises a stem, a hairpin structure, a pseudoknot, a bulge, an internal loop, a multiple loop, or a combination thereof.
15. The method of any one of claims 1-14, wherein the single-stranded DNA template is a library constituent.
16. The method of any one of claims 1-15, wherein the single-stranded DNA template comprises a target nucleic acid positioned between the 5’ adapter and the 3’ adapter.
17. The method of any one of claims 1-16, further comprising ligating the 5’ adapter and the 3’ adapter to a double-stranded DNA template and dehybridizing the double-stranded DNA template to form the single-stranded DNA template.
18. The method of claim 17, wherein dehybridizing the double-stranded DNA template is performed in the presence of the helicase or prior to being in contact with the helicase.
19. The method of claim 17 or 18, wherein dehybridizing the double-stranded DNA template comprises heat treatment, chemical treatment, enzymatic treatment by the helicase and the single-stranded DNA binding protein, or a combination thereof.
20. The method of any one of claims 1-19, wherein the ligation is performed in solution or on a solid support.
21. The method of any one of claims 1 -20, wherein the method is performed in solution or on a solid support.
22. The method of any one of claims 1-21, wherein circularizing the single-stranded DNA template is performed in solution or on a solid support.
23. The method of any one of claims 11-22, wherein the RCA is performed in solution or on a solid support.
24. The method of any one of claims 1-23, wherein the splint oligonucleotide is provided in solution, and maintaining linearity of the single-stranded DNA template and / or circularizing the single-stranded DNA template is performed in solution.
25. The method of any one of claims 11-24, wherein the circularization and RCA are performed in solution, and the splint oligonucleotide functions as a primer for the RCA.
26. The method of claim 25, comprising depositing the circularized DNA template on a surface.
27. The method of claim 25, comprising depositing the circularized DNA template on binding sites of a structured surface.
28. The method of any one of claims 11-24, wherein the circularization is performed in solution and the RCA is performed on a solid support, and contacting the circularized DNA template with the RCA mixture is in the presence of a capture primer immobilized on the solid support.
29. The method of claim 28, comprising providing the solid support having the capture primer.
30. The method of any one of claims 1-29, comprising hybridizing the circularized DNA template to the solid support or the capture primer.
31. The method of any one of claims 1-30, wherein the splint oligonucleotide is a capture primer immobilized on a solid support.
32. The method of claim 31, wherein linearizing the DNA template, circularizing the linear single-stranded DNA template, and / or performing RCA is performed on the solid support.ll. The method of any one of claims 11-32, wherein performing RCA comprises extending the primer or the capture primer along the circularized DNA template.
34. The method of any one of claims 1-33, wherein the time duration in the circularization is about 5 minutes to 2 hours, optionally, 20 minutes, 40 minutes, or 60 minutes.
35. The method of any one of claims 1-34, wherein the helicase is at a concentration of about 0.005 nM to about 5 nM, optionally about 0.005 nM to about 2 nM, optionally about 1 nM to about 2 nM, optionally about 0.005 nM to about 0.01 nM, and / or wherein the helicase is at a concentration of about 0.005 pM to about 5 pM, optionally about 0.005 pM to about 2 pM, optionally about 1 pM to about 2 pM, optionally about 0.005 pM to about 0.01 pM.
36. The method of any one of claims 1-35, wherein the helicase is Tte UvrD or T4 helicase (gp41).
37. The method of any one of claims 1-36, wherein the single-stranded DNA binding protein is at a concentration of less than 1 pM, optionally about 0.01 pM to about 1 pM, optionally 100 nM, 125 nM, 150 nM, or 200 nM.
38. The method of any one of claims 1-37, wherein the single-stranded DNA binding protein is gp32.
39. The method of any one of claims 1-38, wherein maintaining linearity of the singlestranded DNA template is performed at a temperature from about 30 °C to about 50 °C, optionally 40 °C.
40. The method of any one of claims 1-39, wherein the circularizing is performed at a temperature from about 37 °C to about 60 °C.
41. The method of any one of claims 1-40, wherein the ligase is at a concentration of less than 1 pM, optionally about 0.01 pM to about 1 pM, optionally 100 nM, 125 nM, 150 nM, or 200 nM.
42. The method of any one of claims 1-41, wherein the ligase is T4 DNA Ligase.
43. The method of any one of claims 1-42, wherein the single-stranded DNA template is about 100 bps to 500 bps in length.
44. The method of any one of claims 11-43, wherein the RCA is performed at about 37 °C.
45. The method of any one of claims 11-44, wherein the RCA mixture comprises a divalent metal cation, a DNA polymerase, and a dNTP mix.
46. The method of claim 45, wherein the divalent metal cation is Ca2+, Mg2+, or Sr2+.
47. The method of claim 45 or 46, wherein the divalent metal cation is Ca2+or Sr2+.
48. The method of any one of claims 45-47, wherein the RCA mixture does not comprise Mg2+.-Ti49. The method of any one of claims 45-48, wherein the divalent metal cation is at a concentration from about 0.001 mM to about 1 M, optionally from about 0.5 mM to about 200 mM.
50. The method of any one of claims 45-49, wherein the RCA mixture further comprises a polymer.
51. The method of claim 50, wherein the polymer is linear or branched.
52. The method of claim 50 or 51, wherein the polymer is a polyelectrolyte species, optionally the polyelectrolyte species is branched.
53. The method of claim 52, wherein the branched poly electrolyte species is a dendrimer species.
54. The method of claim 53, wherein the dendrimer species is poly(amidoamine) (PAMAM) dendrimer.
55. The method of claim 54, wherein the PAMAM dendrimer is a G1 PAMAM, G2 PAMAM, G3 PAMAM, G4 PAMAM, G5 PAMAM, or a combination thereof.
56. The method of any one of claims 50-55, wherein the polymer is at a concentration from about 0.001 pM to about 1 M.
57. The method of any one of claims 45-56, wherein the divalent metal cation is Ca2+or Sr2+and the polymer is at a concentration from about 1 mM to about 500 mM.
58. The method of any one of claims 11-57, wherein performing RCA further comprises introducing an amplification buffer into the vessel, wherein the amplification buffer does not comprise any DNA polymerase and comprises a divalent metal cation and a polymer, following contacting the circularized DNA template with the RCA mixture for about 10 minutes to about 60 minutes, optionally the polymer is branched polyelectrolyte species.
59. The method of claim 58, wherein the divalent metal cation in the amplification buffer is Mg2+.
60. The method of claim 58 or 59, wherein the divalent metal cation in the RCA mixture is Ca2+or Sr2+and the divalent metal cation in the amplification buffer is Mg2+, optionally the RCA is performed on the solid support.
61. The method of any one of claims 58-60, wherein the divalent metal cation in the amplification buffer is in a concentration of at least 10 mM, optionally about 10 mM to about 10 M.
62. The method of any one of claims 58-61, wherein the branched polyelectrolyte in the amplification buffer is in a concentration of at least 5 pM.
63. The method of any one of claims 58-62, comprising removing the helicase, the single-stranded DNA binding protein, and / or the ligase from the vessel prior to introducing the amplification buffer.
64. The method of any one of claims 11-63, comprising: removing the helicase, the single-stranded DNA binding protein, and / or the ligase prior to performing RCA, optionally wherein the helicase, the single-stranded DNA binding protein, and / or the ligase is removed by washing excess helicase, single-stranded DNA binding protein, and / or ligase.
65. The method of any one of claims 11-64, wherein the RCA is performed in the absence of the helicase, the single-stranded DNA binding protein, and / or the ligase,66. The method of any one of claims 11-62, wherein the RCA is performed in the presence of the helicase, the single-stranded DNA binding protein, and / or the ligase.
67. The method of claim 66, comprising adding helicase and / or the single-stranded DNA binding protein during the RCA.
68. The method of any one of claims 1-67, comprising preparing a DNA library comprising a plurality of library constituents in the presence of the helicase and / or the singlestranded DNA binding protein, prior to maintaining linearity of the single-stranded DNA template and circularizing the single-stranded DNA template.
69. A method of rolling circle amplification (RCA) for nucleic acids, comprising: contacting a circular DNA template and a capture primer with a RCA mixture in a vessel for a time duration to form amplified concatemers of the DNA template, wherein the RCA mixture comprises a DNA polymerase, a dNTP mix, a divalent metal cation, and a branched polyelectrolyte species, wherein the branded polyelectrolyte species is at a concentration from about 1 pM to about 1 M and the divalent metal cation is Ca2+or Sr2+.
70. The method of claim 69, wherein the RCA mixture does not comprise Mg2+.
71. The method of claim 69 or 70, wherein the divalent metal cation is at a concentration from about 0.001 mM to about 1 M, optionally from about 0.5 mM to about 200 mM.
72. The method of any one of claims 69-71, wherein the branded poly electrolyte species is at a concentration from about 1 mM to about 500 mM.
73. The method of any one of claims 69-72, wherein the branched poly electrolyte species is a dendrimer species, optionally the dendrimer species is poly(amidoamine) (PAMAM) dendrimer, optionally the PAMAM dendrimer is a G1 PAMAM, G2 PAMAM, G3 PAMAM, G4 PAMAM, G5 PAMAM, or a combination thereof.
74. The method of any one of claims 69-73, comprising: introducing an amplification buffer into the vessel, wherein the amplification buffer does not comprise any DNA polymeraseand comprises a divalent metal cation and the branched polyelectrolyte species, following contacting the circular DNA template with the RCA mixture for about 10 minutes to about 60 minutes.
75. The method of claim 74, wherein the divalent metal cation in the amplification buffer is Mg2+.
76. The method of claim 74 or 75, wherein the divalent metal cation in the amplification buffer is in a concentration of from about 1 mM to about 10 M, optionally from about 10 mM to about 5 M.
77. The method of any one of claims 74-76, wherein the branched polyelectrolyte in the amplification buffer is in a concentration from about 10 mM to about 1 M, optionally about 10 mM to about 500 mM.
78. The method of any one of claims 69-77, wherein the capture primer is immobilized on a solid surface.
79. The method of any one of claims 11-78, wherein the vessel is a flow cell.
80. The method of claim 1-79, further comprising sequencing the circularized DNA template by single molecule real-time sequencing.
81. The method of any one of claims 11-80, further comprising sequencing the amplified concatemers.
82. The method of claim 81, wherein the sequencing comprises paired end sequencing.
83. The method of claim 81, wherein the sequencing comprises SBB, SBS, single molecule real-time sequencing, or nanopore sequencing.
84. The method of claim 81, wherein the sequencing comprises SBB or SBS and is performed on a flow cell.
85. The method of any one of claims 81-84, wherein the GC bias of reads obtained by the sequencing is less than 0.05.
86. The method of claim 85, wherein the GC bias of reads obtained by the sequencing is less than 0.03.
87. The method of claim 85, wherein the action of helicase and / or SSB reduces the GC bias to less than 75% of GC bias in the absence of helicase and SSB.
88. The method of claim 85, wherein the action of helicase and / or SSB reduces the GC bias to less than 50% of GC bias in the absence of helicase and SSB.
89. The method of any one of claims 85-88, wherein the Fl score of the sequencing is greater than 0.99 for SNP or greater than for 0.98 indels.
90. The method of any one of claims 85-88, wherein the Fl score of the sequencing is greater than 0.993 for SNP or greater than for 0.987 indels.
91. The method of any one of claims 80-90, wherein the action of helicase and / or SSB maintains or increases at least one of average spot signal, uniformity of signal, and throughput.
92. A kit for rolling circle amplification, comprising: a first buffer comprising a divalent metal cation of Ca2+or Sr2+, a branched polyelectrolyte species, and a DNA polymerase, and a second buffer that does not comprise any DNA polymerase and comprises Mg2+and the branched polyelectrolyte species.
93. The kit of claim 80, wherein the branched poly electrolyte species in the first buffer is at a concentration from about 1 pM to about 1 M, optionally at a concentration from about 1 mM to about 500 mM.
94. The kit of claim 80 or 81, wherein the branched poly electrolyte species is a dendrimer species.
95. A kit for circularizing a single-stranded DNA template, comprising: a helicase; a splint oligonucleotide; a ligase; and optionally a single-stranded DNA binding protein.
96. A method of circularizing a single-stranded DNA template, comprising: maintaining linearity of a single-stranded DNA template in a circularization mixture through the activity of a helicase in the presence of a single-stranded DNA binding protein, wherein the single-stranded DNA binding protein binds to the single-stranded DNA template and facilitates association of the helicase with the single-stranded DNA template, and circularizing the linear single-stranded DNA template by ligation on a splint oligonucleotide in a ligase of the circularization mixture for a time duration to form a circularized DNA template, wherein the single-stranded DNA template comprises a 5’ adapter complimentary to a first portion of the splint oligonucleotide and a 3’ adapter complimentary to a second portion of the splint oligonucleotide, wherein the 5’ adapter and the 3’ adapter hybridize to the splint oligonucleotide to enable ligation by the ligase to form the circularized DNA template, wherein maintaining linearity of the single-stranded DNA template and circularizing the linear single-stranded DNA template are performed in a single step.
Citation Information
Patent Citations
Nucleic acid sequencing methods and systems
US10077470B2
Method and system for sequencing nucleic acids
US10246744B2
Method and system for sequencing nucleic acids
US10443098B2
Single molecule arrays for genetic and chemical analysis
US20070099208A1
Polymerase enzymes and reagents for enhanced nucleic acid sequencing
US20080108082A1