Methods and compositions for stabilizing concatemers
By utilizing the hybridization of staple molecules and adapter sequences to form a stable tandem structure during nucleic acid amplification and sequencing, and combining this with rolling circle amplification technology, the problem of poor tandem stability was solved, thus improving the accuracy and efficiency of sequencing.
Patent Information
- Application Number
- CN202480056185.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-06
- Filing Date
- 2024-06-25
- Publication Date
- 2026-06-05
AI Technical Summary
In existing technologies, the stability of tandem sequences is poor during nucleic acid amplification and sequencing, which affects the accuracy and efficiency of sequencing.
By providing multiple tandem strands to contact the staple molecule, a stable tandem strand structure is formed. The first and second regions of the staple molecule hybridize with the adaptor sequence. Combined with the rolling circle amplification (RCA) technique of strand displacement polymerase, a circular nucleic acid template is formed and the tandem strand is deposited on the structured surface. The sense strand is amplified and the antisense strand is generated. The stability of the tandem strand is enhanced by the specific hybridization of the staple molecule with the adaptor sequence.
It improves the stability of the tandem and reduces the increase in full width at half maximum (FWHM) during sequencing, ensuring the accuracy and efficiency of sequencing, especially after 100 sequencing cycles.
Smart Images

Figure CN122161943A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 525,275, filed July 6, 2023, the disclosure of which is incorporated herein by reference in its entirety. Background Technology
[0003] This application generally relates to molecular biology, and more specifically to compositions and methods for enhancing the stability of nucleic acid amplification and sequencing, as well as methods of using them. Summary of the Invention
[0004] In general aspect A1, a method of forming a stabilized tandem assembly includes: (a) providing a plurality of tandem assemblies, wherein each tandem assembly includes a plurality of instances of a target sequence and a plurality of instances of at least one adaptor sequence; (b) contacting the plurality of tandem assemblies with staple molecules to form a plurality of stabilized tandem assemblies; wherein each staple molecule includes at least a first region and a second region; and wherein the first region and the second region of the staple molecule each hybridize with different instances of the at least one adaptor sequence.
[0005] The method of this application may further include any of the following aspects:
[0006] A2. The method according to aspect A1, wherein the provision of the plurality of tandem sequences in step (a) comprises rolling circle amplification (RCA) with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises the target sequence and the at least one adaptor sequence.
[0007] A3. The method according to aspect A2, further comprising circularizing a linear nucleic acid template comprising the target sequence and the at least one adaptor sequence to generate the circular nucleic acid template.
[0008] A4. The method according to aspect A3, wherein the linear nucleic acid template comprises a first adaptor sequence of the target sequence 3' and a second adaptor sequence of the target sequence 5'.
[0009] A5. The method according to aspect A4, wherein the first adaptor sequence and the second adaptor sequence of the linear nucleic acid template are ligated after hybridization with the splint oligonucleotide.
[0010] A6. The method according to aspect A5, wherein the splint oligonucleotide is the primer that hybridizes with the circulated nucleic acid.
[0011] A7. The method according to aspect A5, wherein the splint oligonucleotide is removed prior to the RCA.
[0012] A8. The method according to aspect A2, wherein the primer is fixed to the surface during RCA.
[0013] A9. The method according to aspect A2, wherein the primer is in solution during RCA.
[0014] A10. The method according to aspect A9, further comprising depositing the plurality of stabilized tandem bodies on the surface after contact step (b).
[0015] A11. The method according to aspect A10, wherein the surface is a structured surface.
[0016] A12. The method according to aspect A9, further comprising depositing the plurality of tandem bodies on the surface prior to contact step (b).
[0017] A13. The method according to aspect A12, wherein the surface is a structured surface.
[0018] A14. The method according to aspect A2, wherein the RCA generates a sense strand, and the method further includes amplifying the sense strand to generate multiple antisense strands.
[0019] A15. The method according to aspect A14, wherein at least some of the staple molecules hybridize with the connector sequences of the plurality of antisense strands.
[0020] A16. The method according to aspect A14, wherein step (b) further comprises hybridizing the staple molecule with the sense strand.
[0021] A17. The method according to aspect A16, wherein at least some of the staple molecules are staple primers for amplifying the sense strand to generate the plurality of antisense strands.
[0022] A18. The method according to aspect A2, wherein the contact step (b) occurs during the RCA step (a).
[0023] A19. The method according to aspect A1, wherein each instance of the at least one adaptor sequence includes a primer binding site.
[0024] A20. The method according to aspect A19, wherein each instance of the at least one connector sequence further includes a label region, optionally wherein the label region is a sample index.
[0025] A21. The method according to aspect A19, wherein each instance of the at least one adaptor sequence further includes a splint binding site.
[0026] A22. The method according to aspect A19, wherein each instance of the at least one adaptor sequence further includes a variable region, optionally wherein the variable region is a unique molecular identifier (UMI).
[0027] A23. The method according to aspect A1, wherein each tandem strand comprises a sense strand that hybridizes with a plurality of antisense strands.
[0028] A24. The method according to aspect A23, wherein the first region and the second region of the staple molecule each hybridize with different instances of the at least one adaptor sequence on different antisense strands.
[0029] A25. The method according to aspect A23, wherein the sense chain does not include uracil and the plurality of antisense chains include uracil.
[0030] A26. The method according to aspect A1, wherein all instances of the target sequence in each of the tandem bodies are identical.
[0031] A27. The method according to aspect A1, wherein instances of the target sequence include sense sequences or antisense sequences.
[0032] A28. The method according to aspect A1, wherein the plurality of series bodies are provided fixed to a surface in a flow cell.
[0033] A29. The method according to aspect A28, wherein the surface is a continuous surface.
[0034] A30. The method according to aspect A28, wherein the plurality of tandem bodies are fixed to the binding sites on the structured surface of the flow cell.
[0035] A31. The method according to aspect A1, wherein the plurality of series bodies are provided in solution in step (a) and deposited on the surface of the flow cell in step (b).
[0036] A32. The method according to aspect A1, wherein the sequences of the first region and the second region are identical.
[0037] A33. The method according to aspect A1, wherein at least one of the first and second regions of the staple molecule includes a 3' end that allows extension.
[0038] A34. The method according to aspect A33, wherein the 3' end of the staple molecule hybridizes within 40 nucleotides from the 3' end of the target sequence.
[0039] A35. The method according to aspect A1, wherein at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes with the primer binding site of the at least one adaptor sequence.
[0040] A36. The method according to aspect A1, wherein at least one of the first and second regions of the staple molecule includes a 3' end that allows for the formation of a ternary complex.
[0041] A37. The method according to aspect A36, wherein the 3' end that allows the formation of the ternary complex is reversibly terminated.
[0042] A38. The method according to aspect A1, wherein neither the first region nor the second region of the staple molecule includes a 3' end that allows the formation or extension of the ternary complex.
[0043] A39. The method according to aspect A1, wherein the plurality of tandem units includes a repeating sequence unit, the repeating sequence unit including the target sequence and the at least one linker sequence.
[0044] A40. The method according to aspect A39, wherein the at least one connector sequence within the sequence unit includes a first connector sequence of the target sequence 3' and a second connector sequence of the target sequence 5'.
[0045] A41. The method according to aspect A40, wherein the first region of the staple molecule hybridizes with the first adaptor sequence and the second region of the staple molecule hybridizes with the second adaptor sequence.
[0046] A42. The method according to aspect A40, wherein at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes with a 3' adaptor.
[0047] A43. The method according to aspect A1, wherein the length of each of the first and second regions of the staple molecule is at least 10 nucleotides.
[0048] A44. The method according to aspect A1, wherein the first and second regions of the staple molecule each hybridize with different instances of the at least one adaptor sequence spaced at least 100 nucleotides apart.
[0049] A45. The method according to aspect A1, wherein the 3' end of the staple molecule includes at least one mismatch, a nucleotide incorporation blocker, or a ternary complex formation cap.
[0050] A46. The method according to aspect A45, wherein the 3' end of the staple molecule includes a blocking agent to prevent nucleotide incorporation.
[0051] A47. The method according to aspect A45, wherein the 3' end of the staple molecule prevents the formation of a ternary complex.
[0052] A48. The method according to aspect A47, wherein the 3' end of the staple molecule is capped with a portion preventing the formation of a ternary complex.
[0053] A49. The method according to aspect A47, wherein the 3' end of the staple molecule comprises a mismatch and a blocking agent, optionally wherein the blocking agent is a 3' phosphate blocking agent.
[0054] A50. The method according to aspect A1, wherein at least some of the first and second regions of the staple molecules are separated by spacers.
[0055] A51. The method according to aspect A50, wherein at least some of the staple molecules do not include spacers.
[0056] A52. The method according to aspect A50, wherein the spacer comprises a polynucleotide sequence.
[0057] A53. The method according to aspect A52, wherein the length of the polynucleotide sequence is variable.
[0058] A54. The method according to aspect A52, wherein the polynucleotide sequence comprises a double-stranded DNA sequence.
[0059] A55. The method according to aspect A54, wherein the first region and the second region each include a 3' end that hybridizes with the tandem.
[0060] A56. The method according to aspect A50, wherein the spacer has a variable length in different staple molecules.
[0061] A57. The method according to aspect A50, wherein the spacer comprises a nonnucleotide polymer linker.
[0062] A58. The method according to aspect A57, wherein the nonnucleotide polymeric linker comprises polyethylene glycol (PEG).
[0063] A59. The method according to aspect A57, wherein the nonnucleotide polymer linker comprises a dendritic macromolecule.
[0064] A60. The method according to aspect A59, wherein the dendritic macromolecule comprises polyamide amine (PAMAM).
[0065] A61. The method according to aspect A59, wherein the staple molecule hybridizes with three or more instances of the at least one adaptor sequence.
[0066] A62. The method according to aspect A61, wherein the staple molecule hybridizes with 10 or more instances of the at least one adaptor sequence.
[0067] A63. The method according to aspect A61, wherein a majority of the staple molecule hybridizes with an adaptor sequence side-attached to at least 10 different target sequences.
[0068] A64. The method according to aspect A59, wherein neither the first region nor the second region of the staple molecule includes a 3' end that allows the ternary complex to form or extend.
[0069] A65. The method according to aspect A50, wherein the spacer is coupled to the 3' end of the first region.
[0070] A66. The method according to aspect A65, wherein the spacer is coupled to the 5' end of the second region.
[0071] A67. The method according to aspect A65, wherein the spacer is coupled to the 3' end of the second region and wherein the staple molecule cannot act as a primer.
[0072] A68. The method according to aspect A1, wherein multiple staple molecules hybridize with the same tandem.
[0073] A69. The method according to aspect A1, wherein providing the plurality of tandem bodies comprises forming the tandem bodies in the presence of a compaction additive.
[0074] A70. The method according to aspect A69, wherein the compaction additive is a polymer.
[0075] A71. The method according to aspect A70, wherein the polymer is a PAMAM dendritic macromolecule.
[0076] A72. The method according to any one of aspects A1 or A69, wherein the staple primer does not compact the tandem.
[0077] A73. The method according to aspect A1 or A72, wherein the first region of the staple primer is the 3' of the second region of the staple primer.
[0078] A74. The method according to aspect 73, further comprising sequencing a portion of the tandem from the 3' end of the first region of the staple primer.
[0079] A75. The method according to aspect A74, wherein the portion of the tandem body is at least a first portion of the target sequence.
[0080] A76. The method according to aspect A75, wherein the portion of the tandem is a tag region of the at least one connector sequence, optionally wherein the tag region is a sample index.
[0081] A77. The method according to any one of aspects A74 to A76, further comprising blocking the staple primer to prevent further sequencing from the staple primer.
[0082] A78. The method described in aspect A77, wherein the blockade is performed using a ternary complex inhibitor.
[0083] A79. The method according to aspect A77 or A78, further comprising hybridizing another primer with a region of the adapter sequence that has not hybridized with the staple primer, wherein the other primer is not a staple primer.
[0084] A80. The method according to aspect A79, further comprising sequencing a second portion of the tandem from the other primer while the staple primer remains hybridized with the tandem.
[0085] A81. The method according to any one of aspects A74 to A80, comprising at least 100 sequencing cycles.
[0086] A82. The method according to aspect 81, wherein after the at least 100 sequencing cycles, the full width at half maximum (FWHM) of the tandem does not increase by more than 5%.
[0087] A83. The method according to aspect 82, wherein if the second region of the staple primer is not present, the full width at half maximum (FWHM) of the tandem does not increase by more than 5% after the at least 100 sequencing cycles.
[0088] A84. The method according to aspect 83, wherein if the second region of the staple primer is not present, the full width at half maximum (FWHM) of the tandem does not increase by more than 10% after the at least 100 sequencing cycles.
[0089] A85. The method according to any one of aspects A73 to A84, wherein the tandem strand is a sense strand, and wherein providing a plurality of tandem strands includes generating a plurality of antisense strands from the tandem strands by multi-strand substitution (MDA), sequencing portions of the antisense strands, and digesting the antisense strands.
[0090] Composition
[0091] The composition may typically include an aspect B1 of a stabilized tandem, comprising:
[0092] (a) a tandem comprising multiple instances of a target sequence and multiple instances of at least one adaptor sequence; and (b) one or more staple molecules; wherein each staple molecule comprises at least a first region and a second region; and wherein the first region and the second region of the staple molecule each hybridize with different instances of the at least one adaptor sequence.
[0093] The composition may further include any of the following:
[0094] B2. The composition according to aspect B1, wherein the tandem polymerase is the product of RCA performed by a primer hybridized to a circular nucleic acid template using a chain displacement polymerase, wherein the circular nucleic acid template comprises the target sequence and the at least one adaptor sequence.
[0095] B3. The composition according to aspect B2, wherein the circular nucleic acid template is a product of circularizing a linear nucleic acid template comprising the target sequence and at least one adaptor sequence.
[0096] B4. The composition according to aspect B3, wherein the linear nucleic acid template comprises a first adaptor sequence of the target sequence 3' and a second adaptor sequence of the target sequence 5'.
[0097] B5. The composition according to aspect B4, wherein the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with the splint oligonucleotide.
[0098] B6. The composition according to aspect B5, wherein the splice oligonucleotide is the primer that hybridizes with the circulated nucleic acid.
[0099] B7. The composition according to aspect B5, wherein the splint oligonucleotide is removed prior to the RCA.
[0100] B8. The composition according to aspect B2, wherein the primers are immobilized on the surface during RCA.
[0101] B9. The composition according to aspect B2, wherein the primer is in solution during RCA.
[0102] B10. The composition according to aspect B9, wherein the tandem body is attached to the surface after hybridization with one or more staple molecules.
[0103] B11. The composition according to aspect B10, wherein the surface is a structured surface.
[0104] B12. The composition according to aspect B9, wherein the tandem body is attached to the surface before hybridizing with one or more staple molecules.
[0105] B13. The composition according to aspect B12, wherein the surface is a structured surface.
[0106] B14. The composition according to aspect B2, wherein the tandem strand is a sense strand, and the composition further comprises a plurality of antisense strands generated by the amplification of the sense strand.
[0107] B15. The composition according to aspect B14, wherein at least some of the one or more staple molecules hybridize with the connective sequences of the plurality of antisense strands.
[0108] B16. The composition according to aspect B14, wherein one or more staple molecules hybridize with a sense chain.
[0109] B17. The composition according to aspect B16, wherein at least some of the one or more staple molecules are one or more staple primers for amplifying the sense strand to produce the plurality of antisense strands.
[0110] B18. The method according to aspect B2, wherein the one or more staple molecules hybridize with the tandem strand during RCA.
[0111] B19. The composition according to aspect B1, wherein each instance of the at least one adaptor sequence includes a primer binding site.
[0112] B20. The composition according to aspect B19, wherein each instance of the at least one linker sequence further includes a tag region, optionally wherein the tag region is a sample index.
[0113] B21. The composition according to aspect B19, wherein each instance of the at least one adaptor sequence further includes a splint binding site.
[0114] B22. The composition according to aspect B19, wherein each instance of the at least one adaptor sequence further includes a variable region, optionally wherein the variable region is a unique molecular identifier (UMI).
[0115] B23. The composition according to aspect B1, wherein the tandem strand comprises a sense strand that hybridizes with a plurality of antisense strands.
[0116] B24. The composition according to aspect B23, wherein the first and second regions of the staple molecule each hybridize with different instances of the at least one adaptor sequence on different antisense strands.
[0117] B25. The composition according to aspect B23, wherein the sense chain does not include uracil and the plurality of antisense chains include uracil.
[0118] B26. The composition according to aspect B1, wherein all instances of the target sequence within the tandem body are identical.
[0119] B27. The composition according to aspect B1, wherein examples of the target sequence include sense sequences or antisense sequences.
[0120] B28. The composition according to aspect B1, wherein the series of bodies are provided to be fixed to a surface in a flow cell.
[0121] B29. The composition according to aspect B28, wherein the surface is a continuous surface.
[0122] B30. The composition according to aspect B28, wherein the tandem body is fixed to the binding site on the structured surface of the flow cell.
[0123] B31. The composition according to aspect B1, wherein the series body is generated in solution and subsequently attached to the surface of the flow cell.
[0124] B32. The composition according to aspect B1, wherein the sequences of the first region and the second region are identical.
[0125] B33. The composition according to aspect B1, wherein at least one of the first and second regions of the staple molecule includes a 3' end that allows extension.
[0126] B34. The composition according to aspect B33, wherein the 3' end of the staple molecule hybridizes within 40 nucleotides from the 3' end of the target sequence.
[0127] B35. The composition according to aspect B1, wherein at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes with the primer binding site of the at least one adaptor sequence.
[0128] B36. The composition according to aspect B1, wherein at least one of the first and second regions of the staple molecule includes a 3' end that allows for the formation of a ternary complex.
[0129] B37. The composition according to aspect B36, wherein the 3' end that allows the formation of the ternary complex is reversibly terminated.
[0130] B38. The composition according to aspect B1, wherein neither the first nor the second region of the staple molecule includes a 3' end that allows the formation or extension of the ternary complex.
[0131] B39. The composition according to aspect B1, wherein the tandem comprises a repeating sequence unit, the repeating sequence unit comprising the target sequence and the at least one linker sequence.
[0132] B40. The composition according to aspect B39, wherein the at least one linker sequence within the sequence unit comprises a first linker sequence of target sequence 3' and a second linker sequence of target sequence 5'.
[0133] B41. The composition according to aspect B40, wherein the first region of the staple molecule hybridizes with the first adaptor sequence and the second region of the staple molecule hybridizes with the second adaptor sequence.
[0134] B42. The composition according to aspect B40, wherein at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes with a 3' adaptor.
[0135] B43. The composition according to aspect B1, wherein the length of each of the first and second regions of the staple molecule is at least 10 nucleotides.
[0136] B44. The composition according to aspect B1, wherein the first and second regions of the staple molecule each hybridize with different instances of the at least one adaptor sequence spaced at least 100 nucleotides apart.
[0137] B45. The composition according to aspect B1, wherein the 3' end of the staple molecule includes at least one mismatch, a nucleotide incorporation blocker, or a ternary complex formation preventer cap.
[0138] B46. The composition according to aspect B45, wherein the 3' end of the staple molecule includes a blocker to prevent nucleotide incorporation.
[0139] B47. The composition according to aspect B45, wherein the 3' end of the staple molecule prevents the formation of a ternary complex.
[0140] B48. The composition according to aspect B47, wherein the 3' end of the staple molecule is capped with a portion preventing the formation of a ternary complex.
[0141] B49. The composition according to aspect B47, wherein the 3' end of the staple molecule comprises a mismatch and a blocking agent, optionally wherein the blocking agent is a 3' phosphate blocking agent.
[0142] B50. The composition according to aspect B1, wherein at least some of the first and second regions of the staple molecules are separated by spacers.
[0143] B51. The composition according to aspect B50, wherein at least some of the one or more staple molecules do not include spacers.
[0144] B52. The composition according to aspect B50, wherein the spacer comprises a polynucleotide sequence.
[0145] B53. The composition according to aspect B52, wherein the length of the polynucleotide sequence is variable.
[0146] B54. The composition according to aspect B52, wherein the polynucleotide sequence comprises a double-stranded DNA sequence.
[0147] B55. The composition according to aspect B54, wherein the first region and the second region each include a 3' end that hybridizes with the tandem.
[0148] B56. The composition according to aspect B50, wherein the spacers have variable lengths in different staple molecules.
[0149] B57. The composition according to aspect B50, wherein the spacer comprises a nonnucleotide polymer linker.
[0150] B58. The composition according to aspect B57, wherein the nonnucleotide polymer linker comprises polyethylene glycol (PEG).
[0151] B59. The composition according to aspect B57, wherein the nonnucleotide polymer linker comprises a dendritic macromolecule.
[0152] B60. The composition according to aspect B59, wherein the dendritic macromolecule comprises polyamide amine (PAMAM).
[0153] B61. The composition according to aspect B59, wherein the one or more staple molecules hybridize with three or more instances of the at least one adaptor sequence.
[0154] B62. The composition according to aspect B61, wherein the one or more staple molecules hybridize with 10 or more instances of the at least one adaptor sequence.
[0155] B63. The composition according to aspect A61, wherein a majority of the one or more staple molecules hybridize with an adaptor sequence side-attached to at least 10 different target sequences.
[0156] B64. The composition according to aspect B59, wherein neither the first nor the second region of the staple molecule includes a 3' end that allows the formation or extension of the ternary complex.
[0157] B65. The composition according to aspect B50, wherein the spacer is coupled to the 3' end of the first region.
[0158] B66. The composition according to aspect B65, wherein the spacer is coupled to the 5' end of the second region.
[0159] B67. The composition according to aspect B65, wherein the spacer is coupled to the 3' end of the second region and wherein the staple molecule cannot act as a primer.
[0160] B68. The composition according to aspect B1, wherein a plurality of staple molecules hybridize with the same tandem strand.
[0161] sequencing methods
[0162] The sequencing method may generally include aspect C1 of a method for sequencing a target sequence, the method comprising: (i) providing a plurality of stable tandems according to any one of aspects B1 to B68 or by any one of the methods A1 to A68; and (ii) sequencing at least a first portion of the target sequence.
[0163] Sequencing methods may further include any of the following:
[0164] C2. The method according to aspect C1, wherein sequencing step (ii) uses a reversibly terminated nucleotide.
[0165] C3. The method according to aspect C1, wherein the staple molecule includes the staple primer in sequencing step (ii).
[0166] C4. The method according to aspect C3, wherein after sequencing from the staple primers, the staple primers are blocked or capped, and additional primers are hybridized to the tandem, followed by sequencing from the additional primers.
[0167] C31. The method according to aspect C4, wherein the additional primer is a non-staple primer.
[0168] C32. The method according to aspect C4, wherein the additional primer is a staple primer.
[0169] C5. The method according to aspect C1, wherein the sequencing step (ii) includes sequencing-by-synthesis (SBS).
[0170] C6. The method according to aspect C1, wherein the sequencing step (ii) includes sequencing-by-binding (SBB).
[0171] C7. The method according to aspect C1, wherein the sequencing step (ii) includes sequencing by binding (SBB) using the staple molecule as a staple primer.
[0172] C8. The method according to aspect C7, wherein the SBB comprises the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer, (B) check: forming and detecting a stable ternary complex comprising the staple primer, and (C) activation: cleaving the reversible terminator from the staple primer.
[0173] C9. The method according to aspect C8, wherein the stabilized ternary complex comprises at least one of a labeled nucleotide and a labeled polymerase.
[0174] C10. The method according to aspect C8 further includes capping the staple primer.
[0175] C11. The method according to aspect C10, further comprising hybridizing a non-staple primer with the stabilized tandem and sequencing a second portion of the target sequence during the SBB process.
[0176] C12. The method according to aspect C11, wherein the second portion of the target sequence overlaps with at least a portion of the first portion.
[0177] C13. The method according to aspect C11, wherein the second portion of the target sequence is upstream of the first portion.
[0178] C14. The method according to aspect C11, wherein the second portion of the target sequence is downstream of the first portion.
[0179] C15. The method according to aspect C11, wherein the sequencing of the second portion of the target sequence is repeated multiple times.
[0180] C16. The method according to aspect C8, wherein the cyclic step of SBB is repeated 50 times or more.
[0181] C17. The method according to aspect C1, wherein each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands.
[0182] C18. The method according to aspect C17, wherein the first and second regions of the staple molecule each hybridize with different instances of at least one connective sequence on different antisense strands of the plurality of antisense strands.
[0183] C19. The method according to aspect C18, further comprising sequencing at least some of the plurality of antisense strands of the plurality of stabilized tandem strands.
[0184] C20. The method according to aspect C19, further comprising degrading the plurality of antisense strands of the plurality of stabilized tandem strands, thereby releasing the staple molecules, optionally wherein the degradation comprises digesting the uracil nucleotides of the antisense strands.
[0185] C21. The method according to aspect C20, wherein the degradation comprises digesting uracil with uracil-DNA glycosidase.
[0186] C22. The method according to aspect C20, further comprising contacting the plurality of tandem strands with staple molecules, each staple molecule hybridizing with a different instance of the at least one connective sequence on the sense strand.
[0187] C23. The method according to aspect C22 further includes sequencing the sense strand of each tandem.
[0188] C24. The method according to aspect C23, wherein the sequencing of the sense strand begins with a staple molecule that hybridizes with the sense strand.
[0189] C25. The method according to aspect C17, wherein the staple molecule hybridizes with the sense strand.
[0190] C26. The method according to aspect C25, wherein the stabilized tandem comprises a plurality of staple molecules hybridizing with the sense strand, and wherein the plurality of staple molecules extend to form the plurality of antisense strands.
[0191] C27. The method according to aspect C1, wherein after at least 100 sequencing cycles, the average full width at half maximum (FWHM) of the stabilized tandem, measured in at least one dimension, does not increase by more than 5%.
[0192] C28. The method according to aspect C1, wherein after at least 100 sequencing cycles, the average increase in FWHM of the stabilized tandem, measured in at least one dimension, is reduced by at least 5-fold by the staple molecules.
[0193] C29. The method according to aspect C1, wherein the stabilized tandem is sequenced with a read length of more than 150 cycles.
[0194] C30. The method according to aspect C29, wherein the average signal obtained from the tandem stabilized by the staple molecules in the final sequencing cycle is 50% higher than the average signal obtained from the tandem stabilized by the staple molecules.
[0195] Reagent test kit
[0196] The kit may typically include aspect D1 of the kit, which includes staple molecules according to any one of aspects 1 to 166.
[0197] The kit may further include any of the following:
[0198] D2. The kit according to aspect D1 further comprises one or more of the following: (i) one or more stapled primers, (ii) one or more non-stapled primers, (iii) one or more adaptors, (iv) one or more polymerases, (v) one or more ligases, (vi) one or more splice oligonucleotides, (vii) a flow cell, (viii) multiple labeled nucleotides, (ix) multiple reversibly terminated nucleotides, (x) a capped portion or any combination thereof.
[0199] E1. A kit comprising: (i) a first adapter, (ii) a second adapter, and (iii) a staple molecule, wherein a first region and a second region of the staple molecule hybridize with different instances of the first adapter and / or the second adapter.
[0200] E2. The kit according to aspect E1, further comprising: (iv) a reagent sufficient to form a stable tandem from the target nucleic acid, wherein the stable tandem comprises the first adaptor, the second adaptor, and the staple molecule. Attached Figure Description
[0201] The novel features of the invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of illustrative embodiments and the accompanying drawings (also referred to herein as figures (“Fig.”, “FIG.”, “Figure”, “Figures”, “Figs.”, and “FIGs.”)) in which the principles of the invention are utilized, and in the drawings:
[0202] Figure 1 A schematic diagram of nucleic acid clusters on the surface of a flow cell and a schematic diagram of multiple staple molecules bound to two adapter sequences of the nucleic acid clusters are shown.
[0203] Figure 2 A schematic diagram is shown of staple molecules (with and without polynucleotide spacers) bound to two separate instances of the “A” adaptor sequence of a nucleic acid cluster.
[0204] Figure 3 A schematic diagram of staple molecules (with and without polynucleotide spacers) is shown, which are mismatched at the 3' end and blocked from binding to two separate instances of the “P” adaptor sequence of the nucleic acid cluster.
[0205] Figure 4 A schematic diagram of a set of two staple molecules is shown, wherein the first region of the first staple molecule and the second staple molecule is designed to bind to each other, thereby forming a double-stranded region of DNA, and the second region of the first staple molecule and the second staple molecule is designed to bind to a single instance of the “A” adaptor sequence of the nucleic acid cluster.
[0206] Figure 5 A schematic diagram of a set of two staple molecules is shown, wherein the first region of the first staple molecule and the second staple molecule is designed to bind to each other, thereby forming a double-stranded region of DNA, wherein the double-stranded region is extended by including a polynucleotide spacer, and the second region of the first staple molecule and the second staple molecule is designed to bind to a separate instance of the “A” adaptor sequence of the nucleic acid cluster.
[0207] Figure 6 A schematic diagram showing multiple staple molecules binding to non-DNA adapters and their interactions with multiple “P” adaptor sequences of nucleic acid clusters is shown, along with schematic diagrams of two exemplary non-DNA adapters.
[0208] Figure 7 The diagram shows a nucleic acid cluster and a mixture of staple molecules (with and without polynucleotide spacers) that bind to the “A” and “P” adaptor sequences of the nucleic acid cluster.
[0209] Figure 8The images shown are inspection images from the tandem sequencing experiment. The left column shows the inspection images after the first sequencing cycle, and the right column shows the inspection images after the last (177th) sequencing cycle. The top row shows the inspection images of the tandem stabilized with staple molecules ("+staple"), and the bottom row shows the inspection images of the tandem stabilized without staple molecules ("no staple").
[0210] Figure 9 The results of the sequencing experiments are shown, with one tandem array including staple molecules (left column; "+staple") and another array omitting staple molecules (right column; "SOP"). The top row shows the 50th percentile of the ratio of the raw intensity value to the observed background (or "OFF intensity") collected over 177 sequencing cycles (ATGC) across the entire lane of the flow cell ("ON / OFF P50"). The middle row shows a plot of the average fluorescence intensity of pixel clusters in the column direction for each examination ("Column FWHM"). The bottom row shows a plot of the average fluorescence intensity of pixel clusters in the row direction for each examination ("Row FWHM").
[0211] Figure 10 This displays four plots of bands collected during sequencing of stapled tandems (left column; "+staple") and unstapled tandems (right column; "no staple"). The top row represents the total number of individual tandems sequenced in each plot. The middle row shows the 50th percentile of the mean read length (over 177 cycles) for tandems within each plot. The bottom row shows the 25th percentile of the mean read length for tandems within each plot.
[0212] Figure 11 Results of single-end RCA clustering / sequencing experiments using stapled molecularly stable tandems are shown, with the stable tandems conjugated to different fluorophores. The 25th percentile of read lengths and 90% accurate O scores are displayed for various fluorophore configurations.
[0213] Figure 12 The results of paired-end RCA clustering / sequencing experiments demonstrate the feasibility of dual indexing of reads on paired-end tandem structures with stapled molecular stability.
[0214] Figure 13 A schematic diagram of buffer exchange during RCA is shown.
[0215] Figure 14 An example of a paired-end sequencing workflow that begins at the end of a paired-end RCA clustering workflow is shown.
[0216] Figure 15 The clonalness of stable tandem clones under different inoculation densities was demonstrated.
[0217] Figure 16 The comparison of the FWHM of the stabilized tandem strand with that of the unstapled tandem strand is shown, thus allowing for a comparison of the size of the tandem strand. Detailed Implementation
[0218] In the following detailed description, reference is made to the accompanying drawings, which form a part of this document. In the drawings, similar symbols generally identify similar components unless the context otherwise requires. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that, as generally described herein and illustrated in the accompanying drawings, aspects of this disclosure can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein and are part of the disclosure herein.
[0219] With respect to the relevant technology, all patents, published patent applications, other publications, and sequences from GenBank and other databases mentioned herein are incorporated herein by reference in their entirety.
[0220] This disclosure provides methods, compositions, and kits for forming stable tandem molecules using staple molecules. Furthermore, this disclosure provides methods of using the methods, compositions, and kits for forming stable tandem molecules using staple molecules. The methods herein involve one or more of the steps described below, such that practice of the disclosed methods may include some or all of the steps disclosed herein. Upon review of this disclosure as a whole, it will be contemplated that those skilled in the art can practice the disclosed invention at the “start” or midway through a given method without departing from the scope of the invention. Thus, for example, those skilled in the art will understand that they may begin with a circularized library component having adjacent first and second adaptor sequences, rather than with a linear nucleic acid having 5' and 3' adaptors.
[0221] I. Definition
[0222] Unless otherwise specified, the technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this disclosure pertains. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology, 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory, New York (2001). For the purposes of this disclosure, the following terms are defined below.
[0223] Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein include plural indicators. Thus, for example, reference to “antigen” includes a mixture of antigens; reference to “pharmaceutically acceptable carrier” includes a mixture of two or more such carriers, etc. Thus, the terms “a”, “one or more,” and “at least one” are used interchangeably herein.
[0224] Furthermore, the term “and / or” as used herein should be considered to specifically disclose that each of the two specified features or components exists with or without the other. Therefore, the term “and / or” as used herein in phrases such as “A and / or B” is intended to include “A and B”, “A or B”, “A (alone)”, and “B (alone)”.
[0225] As used herein, the term "approximately" refers to ±10% of a specified value. When referring to a range of values (or parameters), the term "approximately" means +10% of the upper limit and -10% of the lower limit of the specified range. When providing ranges of values, it should be understood that every intermediate value between the upper and lower limits of the range, as well as any other stated value or intermediate value within the stated range, is included within the scope of this disclosure. Where a stated range includes the upper and / or lower limits, the range excluding any of these included limitations is also included in this disclosure.
[0226] As described herein, the term "nucleotide" may be used to refer to natural nucleotides or their analogues. Examples include, but are not limited to, nucleotide triphosphates (NTPs), such as ribonucleotide triphosphates (rNTPs), deoxyribonucleotide triphosphates (dNTPs), or their non-natural analogues, such as dioxynucleotide methylphosphates (ddNTPs) or reversibly terminated nucleotide triphosphates (rtNTPs).
[0227] As used herein, the term "template" refers to a nucleic acid or a portion thereof having a nucleotide base sequence that serves as a guide for producing complementary copies of that sequence. The template can be replicated via extension of a primer that hybridizes at or near the template. Extension can be mediated by a polymerase or ligase. The template may include or may be DNA, RNA, or analogues thereof. The template may include a linear nucleic acid template, a circular nucleic acid template, or a circularized linear nucleic acid template.
[0228] As used herein, the term "tandem" when referring to a nucleic acid molecule means a continuous nucleic acid molecule containing multiple copies of a tandemly linked common sequence, such as, for example, multiple copies of a tandemly linked target sequence and one or more adaptor sequences. Similarly, the term "tandem" when referring to a nucleotide sequence means a continuous nucleotide sequence containing multiple copies of a tandemly linked common sequence. Each copy of the sequence may be referred to as a "sequence unit" of the tandem. Sequence units may have a length of at least 10 bases, 50 bases, 100 bases, 250 bases, 500 bases, or more. A tandem may include at least 2, 5, 10, 50, 100, or more sequence units. Sequence units may include subregions having any of a variety of functions, such as adaptor sequence regions, primer-binding regions, target sequence regions, tag regions, unique molecular identifiers (UMIs), etc.
[0229] As used herein, the term "common sequence" refers to a nucleotide sequence that is identical for two or more nucleic acid molecules. The common sequence of two or more nucleic acids may include all or part of the nucleic acids being compared. The common sequence may have a length of at least 5, 10, 25, 50, 100, 250, 500, 1000, or more nucleotides. Alternatively or additionally, the length may be up to 1000, 500, 250, 100, 50, 25, 10, or 5 nucleotides. A population of nucleic acid molecules may include individual molecules having inter-individual common sequence regions (e.g., "universal primers" or "universal primer binding sites") and inter-individual variable sequence regions (e.g., "target regions") that differ from each other.
[0230] The terms “sense” and “antense” are used in this paper to distinguish members of a pair of complementary nucleic acid molecules or sequences. These terms are intended as context-specific identifiers. These terms are interchangeable depending on their use in the field of molecular biology. Thus, a strand identified as a “sense strand” in one context may be called an “antense strand” in another context. This is unrelated to how similar or different the first context is compared to the second.
[0231] As used herein, the term "circular" when referring to a nucleic acid strand means that the strand lacks ends (i.e., it lacks both 3' and 5' ends). Therefore, the 3' oxygen and 5' phosphate portions of each nucleotide monomer in a circular strand are covalently attached to adjacent nucleotide monomers in the strand. Circular DNA strands can be used as templates for generating tandem amplicones via rolling circle amplification (RCA), where each sequence unit of the tandem amplicon is the reverse complement of the circular nucleic acid strand. Circular nucleic acids can be double-stranded. One or both strands of a double-stranded nucleic acid may lack both 3' and 5' ends. One strand of a double-stranded nucleic acid may have a gap (the absence of at least one nucleotide monomer relative to the other strand) or a nick (the absence of a phosphodiester bond between two nucleotide monomers), provided that the other strand is circular.
[0232] As used herein, the term "polymerase" can refer to nucleic acid synthases, including but not limited to DNA polymerases, RNA polymerases, reverse transcriptases, primases, and transferases. Typically, a polymerase has one or more active sites at which nucleotide binding and / or nucleotide polymerization catalysis can occur. A polymerase can catalyze the polymerization of a nucleotide to the 3' end of the first strand of a double-stranded nucleic acid molecule. For example, a polymerase catalyzes the addition of the next correct nucleotide to the 3' oxygen atom of the first strand of a double-stranded nucleic acid molecule via a phosphodiester bond, thereby covalently incorporating a nucleotide into the first strand of the double-stranded nucleic acid molecule. Optionally, the polymerase does not need to be able to incorporate nucleotides under one or more conditions used in the methods described herein. For example, a mutant polymerase may be able to form a ternary complex but cannot catalyze nucleotide incorporation. A polymerase may have strand substitution activity, such as Phi29. A polymerase may lack strand substitution activity. A polymerase may have 5' 3' exonuclease activity. A polymerase may lack 5' 3' exonuclease activity.
[0233] As used herein, the term "primer" refers to a nucleic acid having a sequence that binds to a nucleic acid at or near a template sequence. Typically, primers bind in a conformation that allows template replication, for example, via primer polymerase extension. A primer can be a first part of a nucleic acid molecule that binds to a second part of the nucleic acid molecule, the first part being the primer sequence and the second part being the primer-binding sequence (e.g., a hairpin primer). Alternatively, a primer can be a first nucleic acid molecule that binds to a second nucleic acid molecule having a template sequence. Primers can consist of DNA, RNA, or analogues thereof. Primers can have an extendable 3' end or a 3' end that prevents primer extension.
[0234] As used herein, the term "adapter sequence" refers to a known synthetic nucleic acid sequence. An adaptor sequence can serve as a starting point for reading bases at multiple locations beyond the adaptor sequence-target sequence linking point, and optionally, bases can be read in both directions starting from the adaptor sequence. Adaptor sequences can be engineered to include one or more of the following: 1) a length of about 10 to about 100 nucleotides, 2) features for linking to the 5' and / or 3' ends of the target sequence, 3) distinct and unique anchoring binding sites at the 5' and / or 3' ends of the adaptor sequence for sequencing adjacent target sequences, and 4) optionally one or more restriction sites. Adaptor sequences can be inserted or added to predetermined positions relative to the target sequence, including but not limited to using molecular engineering via primer extension (e.g., where the adaptor is the 5' tail of the primer), by linking, or by insertion into the target sequence. In the case of tandem sequences, one or more adaptor sequences can be regularly spaced apart, such as by rolling circle amplification of circularized oligonucleotides containing the target sequence and one or more adaptor sequences.
[0235] As used herein, the term "staple molecule" refers to a composition containing a polynucleotide sequence that enables structural stability of a tandem molecule. The staple molecule of the present invention has a nucleic acid sequence comprising at least a first region and a second region, the first and second regions being designed to have sequences that allow hybridization with different sequences within the tandem molecule, particularly with different instances of at least one adaptor sequence of the tandem molecule. In other words, the staple molecule of the present invention comprises a first region and a second region, each having a sequence complementary to an adaptor sequence within the tandem molecule. When such a staple molecule is applied to a tandem molecule, the first and second regions are able to hybridize with different instances of an adaptor sequence within the tandem molecule by virtue of their sequence design, thereby compacting or otherwise stabilizing the tandem molecule by connecting different portions of the tandem sequence across the three-dimensional structure of the tandem molecule. Compared to other mechanisms that rely on non-specific interactions with the tandem, the staple molecule provides a greater level of tandem stabilization and compaction. The staple molecule specifically interacts with and hybridizes with the tandem molecule, and this structural interaction prevents the staple molecule from being washed away during sequencing. As discussed herein, typically the sequences of the first and second regions of a staple molecule are configured to hybridize with an adaptor sequence in a tandem. Furthermore, the adaptor sequences with which the first and second regions of the staple molecule hybridize can be identical (e.g., the first and second regions of the staple molecule hybridize with different instances of the same adaptor sequence in the tandem), different adaptor sequences, or adaptor sequences that are at least partially conserved in sequence. In some aspects, the staple molecule contains a primer sequence at its 3' end that hybridizes with the primer binding site of the tandem and a tail sequence at its 5' end that hybridizes with the adaptor of the tandem. Such staple molecules may be referred to herein as "staple primers". Primer sequences can be used to sequence a portion of the tandem, such as the target sequence of the tandem or at least a portion of a sample index. Staple molecules can be specifically engineered to compact and / or stabilize the tandem in different ways depending on the nature and sequence of the adaptor and the first and second regions of the staple molecule. Staple molecules can stabilize the tandem (e.g., such as the full width at half maximum of the tandem) without compacting the tandem, such as when compaction is provided by an additive. The additive can be an alcohol such as ethanol, or a polymer such as PEG or PAMAM dendrimer. Furthermore, the sequences of the first and second regions can have the same length, or the first region can be greater than or equal to the length of the second region, or the second region can be greater than or equal to the length of the first region.As described further in detail herein, the staple molecule may further comprise, but is not limited to, one or more of the following components: a reversible terminator portion, a blocking agent or blocking portion, a capping or capping portion, a double-stranded nucleic acid region, at least one nucleotide mismatch in the optional double-stranded portion, a 3' phosphate blocker, a spacer between the first and second regions, or any combination thereof. The spacer may be a polynucleotide sequence or a non-nucleotide polymer linker. The non-nucleotide polymer linker may be a branched polyelectrolyte species, such as polyethylene glycol, or a dendritic macromolecule, such as poly(amidoamine). In some configurations, at least one of the first and second regions of the staple molecule may further comprise one or more of the following components: a 3' end that allows extension, a 3' end that allows the formation of a ternary complex, a reversibly terminated 3' end, or any combination thereof.
[0236] As used herein, the term "primer-template nucleic acid hybrid" or "primer-template hybrid" refers to a nucleic acid having double-stranded regions such that one strand is the primer and the other is the template. The two strands can be multiple parts of a continuous nucleic acid molecule (e.g., a hairpin structure), or the two strands can be separable molecules that are not covalently attached to each other.
[0237] As used herein, the term "cluster," when referring to nucleic acids, means a group of nucleic acids attached to a solid support, such as attachment at sites in a site array on a solid support. The term "clonal population" refers to a homogeneous group of nucleic acids with respect to a particular nucleic acid sequence. Homogeneous sequences are typically at least 10 nucleotides long, but can be even longer, including, for example, at least 50, 100, 500, 1000, or 2500 nucleotides long. Clonal populations can be derived from a single template nucleic acid. Clonal populations can include at least 2, 10, 100, 1000, or more copies of a particular nucleic acid sequence. These copies can be present in a single nucleic acid molecule, for example, as tandem, or these copies can be present on separate nucleic acid molecules. Typically, all nucleic acids in a cluster will have the same nucleotide sequence. It should be understood that a negligible number of contaminating nucleic acids or mutations (e.g., due to amplification artifacts) may be present in a cluster without deviating from epigenetic clonalness. Clusters can be at least 80%, 90%, 95%, or 99% cloned. Optionally, a cluster can be 100% cloned.
[0238] As used herein, the term "branched polyelectrolyte" includes substances commonly referred to as "dendritic macromolecules," which are generally spherical three-dimensional molecules with repeating branching at nanoscale dimensions. Dendritic macromolecule species may contain controlled terminal surface chemistry having one or more functional groups, including but not limited to amine, carboxyl, and hydroxyl groups. Dendritic macromolecule species can be obtained as generation 0 (G0) through generation 10 (G10), with each generation having twice the number of branches as the previous generation. Thus, G0 = 4 branches, G1 = 8 branches, and so on. Dendritic macromolecule species that can be used in the methods, compositions, and systems disclosed herein include, but are not limited to, branched polyamines comprising a protonated structure that interacts with the negatively charged DNA backbone to form a complex. It should be understood that, in the embodiments described herein, adaptor elements may still be used in conjunction with branched polyamines for alternative purposes or to provide improved binding properties of dendritic macromolecule species to nucleic acids. In some embodiments, the branched polyelectrolyte is a poly(amidoamine) dendritic macromolecule species (also known as PAMAM), such as the G2 PAMAM dendritic macromolecule having 16 branches and an amine (NH2) terminal surface chemistry. Non-limiting examples of branched polyelectrolytes also include G4 (64 branches with amine terminal groups) and G5 (128 branches with amine terminal groups) PAMAM dendritic macromolecule species.
[0239] As used herein, the term "blocking moiety," when used to refer to a nucleotide, means a portion of a nucleotide that inhibits or prevents the 3' oxygen of the nucleotide from forming a covalent bond with the next correct nucleotide during nucleic acid polymerization. The blocking moiety of a "reversible termination" nucleotide may be removed from a nucleotide analogue or otherwise modified to allow the 3' oxygen of the nucleotide to covalently attach to the next correct nucleotide. Such a blocking moiety is referred to herein as a "reversible termination moiety." Exemplary reversible termination moieties are set forth in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference. A nucleotide having a blocking moiety or a reversible termination moieties may be located at the 3' end of a nucleic acid, such as a primer, or the nucleotide may be a monomer that is not covalently attached to the nucleic acid. The blocking portion does not need to prevent or inhibit the formation of a ternary complex at the 3' end of the nucleic acid to which the blocking portion is attached. A particularly useful blocking portion will be located at the 3' end of the nucleic acid involved in the formation of the ternary complex.
[0240] As used herein, the term "capped portion," when referring to nucleic acids, means a portion present in nucleic acids that prevents or inhibits the binding of the 3' end of the nucleic acid to polymerase and the next correct nucleotide to form a ternary complex. Portions that generate spatial blocks to form ternary complexes are particularly useful and include, for example, polymeric or ligation products that extend a primer to the end of a template that hybridizes with the primer. Another example of a spatial block is a mismatched nucleotide. Capped portions can have a positive or negative charge that prevents or inhibits ternary complex formation. Capped portions can include ligands that bind to receptors to prevent or inhibit ternary complex formation, such as biotin (or its analogues) that binds to streptavidin (or its analogues), epitopes that bind to antibodies (or their functional fragments), carbohydrates that bind to lectins, etc. Therefore, a ternary complex inhibitor can be a ligand-receptor complex that inhibits ternary complex formation. Further examples of portions that can be used as inhibitors of ternary complexes include base modifications and nucleotide analogs described in U.S. Patent Application Publication No. 2020 / 0032322 A1 or Turcatti et al. Nucl. Acids. Res. 36(4)e25 (2008), each of which is incorporated herein by reference.
[0241] As used herein, the term "array" refers to a group of molecules attached to one or more solid supports in a manner that allows molecules to be distinguished from one another. An array may include distinct molecules, each located at a different addressable site on the solid support. An array may include individual solid supports, each acting as a site with a different molecule, which can be identified based on the position of the solid support on the surface to which it is attached, or based on the position of the solid support in a liquid (such as a fluid flow). The molecules in an array may be, for example, nucleotides, nucleic acid primers, nucleic acid templates, or nucleases such as polymerases, ligases, exonucleases, or combinations thereof.
[0242] As used herein, the term "site" when referring to an array means the location within the array where a specific molecule is present. A site may contain only a single molecule, or it may contain a group (collection of molecules) of several molecules from the same species. Alternatively, a site may include a group of molecules from different species (e.g., a group of ternary complexes with different template sequences). Sites in an array are typically discrete. Discrete sites may be continuous or may have gaps between them. Arrays used herein may have sites spaced, for example, less than 100 micrometers, 50 micrometers, 10 micrometers, 5 micrometers, 1 micrometer, or 0.5 micrometers apart. Alternatively or additionally, an array may have sites spaced greater than 0.5 micrometers, 1 micrometer, 5 micrometers, 10 micrometers, 50 micrometers, or 100 micrometers apart. Sites may each have an area of less than 1 square millimeter, 500 square micrometers, 100 square micrometers, 25 square micrometers, 1 square micrometer, or smaller. Sites may also be referred to as "features" of the array.
[0243] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquids. The substrate may be non-porous or porous. The substrate may optionally be able to absorb liquids (e.g., due to its porosity), but generally possesses sufficient rigidity such that the substrate does not substantially expand when absorbing liquids and does not substantially shrink when the liquid is removed by drying. Non-porous solid supports are generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylic, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutene, polyurethane, Teflon™, cycloolefins, polyimide, etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials (including silicon and modified silicon), carbon, metals, inorganic glasses, fiber bundles, and polymers.
[0244] As used herein, the term "attachment" refers to a state in which two objects are joined, fastened, adhered, connected, or bound together. For example, a reactive component, such as an initiating template nucleic acid or polymerase, can be attached to a solid component via covalent or non-covalent bonds. Covalent bonds are characterized by the sharing of electron pairs between atoms. Non-covalent bonds are chemical bonds that do not involve the sharing of electron pairs and can include, for example, hydrogen bonds, ionic bonds, van der Waals forces, hydrophilic interactions, and hydrophobic interactions.
[0245] As used herein, a "container" is a vessel used to isolate one chemical process (e.g., a binding event; an incorporation reaction; etc.) from another chemical process, or to provide space in which a chemical process may occur. Examples of containers used in conjunction with the disclosed techniques include, but are not limited to, flow cells, pores in multi-well plates; microscope slides; tubes (e.g., capillaries); droplets, vesicles, test tubes, trays, centrifuge tubes, features in arrays, pipes, channels in substrates, etc.
[0246] As used herein, the term "cycle," when used to refer to a sequencing procedure, refers to a portion of a sequencing run that is repeated to indicate the presence of nucleotides. Typically, a cycle includes several steps, such as a reagent delivery step, a washing away of unreacted reagents step, and a detection step that indicates a change in response to the added reagents.
[0247] As used herein, the term "deblocking" means the removal or modification of the reversible terminator portion of a nucleotide to make the nucleotide extendable. For example, a nucleotide may be present at the 3' end of a primer, such that deblocking makes the primer extendable. Exemplary deblocking reagents and methods are set forth in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference.
[0248] As used herein, the term "exogenous" when referring to a part of a molecule means a chemical part that is not present in the molecule's natural analogues. For example, an exogenous label on a nucleotide is a label that is not present on naturally occurring nucleotides. Similarly, an exogenous label present on a polymerase is not found on polymerases in their natural environment.
[0249] As used herein, the term "extension," when referring to nucleic acids, means the process of adding at least one nucleotide to the 3' end of a nucleic acid. The term "polymerase extension," when referring to nucleic acids, refers to the polymerase-catalyzed process of adding at least one nucleotide to the 3' end of a nucleic acid. The nucleotide or oligonucleotide added to a nucleic acid through extension is called incorporation into the nucleic acid. Therefore, the term "incorporation" can be used to refer to the process of linking a nucleotide or oligonucleotide to the 3' end of a nucleic acid by forming a phosphodiester bond.
[0250] As used herein, the term "extendable" when referring to a nucleotide means that the nucleotide has an oxygen or hydroxyl moiety at the 3' position and is capable of forming a covalent bond with the next correct nucleotide. An extendable nucleotide can be at the 3' position of a polymeric nucleic acid, or it can be a monomeric nucleotide. An extendable nucleotide will lack a blocking moiety, such as a reversible terminator moiety.
[0251] As used herein, the term “immobilization” when referring to molecules means the direct or indirect, covalent or non-covalent attachment of molecules to a surface, such as the surface of a solid support. In some configurations, covalent attachment may be preferred, but generally all that is needed is for the molecules (e.g., nucleic acids) to remain immobilized or attached to the surface under conditions intended for surface retention.
[0252] As used herein, the term "label" refers to a molecule or portion thereof that provides a detectable characteristic. Detectable characteristics can be, for example, optical signals such as radiative absorbance, fluorescence emission, cold light emission, fluorescence lifetime, fluorescence polarization, etc.; Rayleigh and / or Mie scattering; binding affinity to a ligand or acceptor; magnetic properties; electrical properties; charge; mass; radioactivity, etc. Exemplary labels include, but are not limited to, fluorophores, chromophores, nanoparticles (e.g., gold, silver, carbon nanotubes), heavy atoms, radioactive isotopes, mass labels, charge labels, spin labels, acceptors, ligands, etc.
[0253] As used herein, the term "next correct nucleotide" refers to a nucleotide or type of nucleotide that will bind to and / or be incorporated at the 3' end of a primer to be complementary to a base in the template strand that the primer hybridizes with. The base in the template strand is called the "next base" and is immediately adjacent to the 5' end of the base in the template that hybridizes with the 3' end of the primer. The next correct nucleotide can be called a "homolog" of the next base, and vice versa. Homologous nucleotides that interact with each other in a ternary complex or double-stranded nucleic acid are said to be "paired." Nucleotides with bases that are not complementary to the next template base are called "incorrect," "mismatched," or "non-homologous" nucleotides.
[0254] As used herein, the term "non-catalytic metal ion" refers to a metal ion that, in the presence of polymerase, does not promote the formation of the phosphodiester bonds required for the chemical incorporation of nucleotides into primers. Non-catalytic metal ions can interact with polymerases, for example, through competitive binding compared to catalytic metal ions. Therefore, non-catalytic metal ions can act as inhibitory metal ions. A "divalent non-catalytic metal ion" is a non-catalytic metal ion having a divalent oxidation state. Examples of divalent non-catalytic metal ions include, but are not limited to, Ca²⁺, Zn²⁺, Co²⁺, Ni²⁺, and Sr²⁺. Trivalent Eu³⁺ and Tb³⁺ ions are non-catalytic metal ions having a trivalent oxidation state.
[0255] As used herein, the term "ternary complex" refers to the intermolecular association between a polymerase, a double-stranded nucleic acid, and a nucleotide. Typically, the polymerase promotes the interaction between the next correct nucleotide and the template strand of the initiating nucleic acid. The next correct nucleotide can interact with the template strand via Watson-Crick hydrogen bonds. The term "stabilized ternary complex" refers to a ternary complex with a promoted or elongated presence, or a ternary complex whose destruction has been inhibited. Generally, stabilization of the ternary complex prevents the covalent incorporation of the nucleotide component of the ternary complex into the initiating nucleic acid component.
[0256] II. Overview
[0257] This article provides methods and compositions for forming stabilized tandem molecules using staple molecules. The staple molecules specifically interact with and hybridize to the tandem, and this structural interaction serves to stabilize the tandem, which further simplifies downstream processes, including but not limited to loading the tandem onto arrays and utilizing the stabilized tandem in applications such as sequencing reactions.
[0258] As discussed further in detail herein, a tandem mass is a continuous nucleic acid molecule containing multiple copies of a common sequence. In some embodiments, the tandem masses used in the aspects discussed herein and in the embodiments contain repeating copies of a target sequence and one or more adaptor sequences. As the volume of tandem masses increases, they may become increasingly difficult to load onto surfaces such as arrays, and such tandem masses may often exhibit further signal strength reduction in applications such as sequencing or other reactions utilizing signal transduction molecules. These problems that can occur with tandem masses are generally referred to as tandem masses becoming unstable or becoming unstable. Additives used to compact or stabilize tandem masses can alleviate these problems. However, many additives become gradually washed away from tandem masses in downstream processes such as sequencing reactions because such additives often rely on nonspecific interactions with the tandem masses to exert a stabilizing effect.
[0259] Conversely, the staple molecule described herein is configured such that the sequences of the first and second regions of the staple molecule hybridize with the connective sequences within the tandem. As a result, the regions of the tandem pull closer to each other in their respective directions, thereby compacting the structure of the entire molecule and stabilizing it in other ways.
[0260] In some embodiments, the methods and compositions described herein increase the compaction of the tandem sequence by using staple molecules to connect different portions of the tandem sequence across the three-dimensional structure of the tandem molecule. In other embodiments, the methods and compositions of the present invention suppress molecular dispersion of the tandem during sequencing.
[0261] Generally, staple oligonucleotides (i.e., staple molecules containing oligonucleotide sequences as described herein) bind directly or indirectly to two or more regions of a tandem. In various designs, staple oligonucleotides can hybridize directly to two or more connective regions, hybridize with each other, and / or bind to intermediates. For example, a portion of a staple oligonucleotide (e.g., one end of a staple oligonucleotide) can hybridize with a single-stranded portion of a first connective region, while another portion (e.g., its other end) can hybridize with a single-stranded portion of a second connective region or with a single-stranded portion of another instance of the first connective region. In some aspects, portions of staple oligonucleotides can hybridize with each other or with intermediates. For example, portions of two or more staple oligonucleotides can each hybridize with the same single intermediate oligonucleotide or different intermediate oligonucleotides (presented by particles). Such hybridization can bridge different fragments, allowing fragments generated by a given tandem to aggregate tightly. Staple oligonucleotides can be functionalized (e.g., at their 3' and / or 5' ends) to crosslink with each other or via intermediates. For example, biotinylated staple oligonucleotides can be cross-linked with streptavidin (e.g., by adding streptavidin after hybridizing the staple oligonucleotide with a first and / or second adaptor region). In another example, staple oligonucleotides can be functionalized with click chemistry groups such as strain-promoted click chemistry groups, and can be cross-linked by presenting multivalent intermediates of complementary click chemistry partners in various cases. One exemplary strain-promoted click chemistry partner includes dibenzocyclooctylene (DBCO) and azides, and another is transcyclooctene (TCO) and tetrazides and their derivatives. For example, DBCO-functionalized staple oligonucleotides can be cross-linked by presenting dendritic macromolecules of multiple azides.
[0262] A single oligonucleotide can optionally be used as both a primer and a staple oligonucleotide. For example, the 5' end of a first sequencing primer can hybridize with the second adaptor region, while the 3' end hybridizes with the first adaptor region (making the first sequencing primer also a staple oligonucleotide), and / or the 5' end of a sequencing primer can hybridize with the first adaptor region, while the 3' end hybridizes with the second adaptor region (making the sequencing primer also a staple oligonucleotide). Clearly, the portions of the adaptor regions complementary to the sequencing and staple portions of such primers typically do not overlap, preventing the primers from competing for their binding sites.
[0263] In further embodiments, the methods and compositions described herein stabilize tandem conjugates, resulting in an observed increase in cycle number during sequencing. In yet another embodiment, stabilization of tandem conjugates according to the methods and compositions described herein results in an increase in read length during the sequencing reaction. Further examples of downstream effects of tandem conjugate stabilization include, but are not limited to, enhanced signal detection and increased ability of sequencing instruments and / or the determination of tandem conjugate clonality, even at increased dot density.
[0264] III. Methods for forming stable series circuits
[0265] In one aspect, this disclosure provides a method for forming stabilized tandems. During sequencing, tandems typically spread out, become less compacted, or “stripe”, resulting in reduced quality, intensity, and read length during sequencing. As tandem volume increases and they exhibit reduced signal intensity, they are termed unstable or unstable. Additives used to compact or “stabilize” tandems can alleviate these problems. However, these additives are gradually washed away from the tandems during sequencing due to their nonspecific interactions with the tandems. As discussed further in detail herein, stabilized tandems are resistant to molecular diffusion and spectral bleeding due to the hybridization of staple molecules with the tandems. Unlike additives, staple molecules interact specifically with the tandems and are therefore resistant to being washed away during sequencing. Furthermore, staple molecules promote tandem stabilization by bridging segments of the tandems, resulting in reduced tandem volume and increased signal intensity. It should be understood that any aspect and embodiment of the method for forming a stable tandem body described herein may utilize any aspect and / or embodiment of the composition described below, and may be further used in any of the methods of use also described herein.
[0266] A. Methods for preparing tandem conductors
[0267] In short, the method for preparing tandem strands may include, but is not limited to, the following steps: 1) isolating and processing nucleic acids, 2) attaching an adapter sequence to the nucleic acid and circularizing the template, 3) performing rolling circle amplification (RCA) to generate a single-stranded tandem strand (first strand or “sense strand”), and optionally 4) performing multiple substitution amplification (MDA) to generate multiple second strands (“antisense strands”) of the first strand (“sense strand”) of the tandem strand.
[0268] 1. Nucleic acid isolation, fragmentation, and size capture :
[0269] Nucleic acids can be isolated using methods known in the art, including, for example, those described in Sambrook et al., *Molecular Cloning: A Laboratory Manual*, 3rd edition, Cold Spring Harbor Laboratory, New York (2001) or Ausubel et al., *Current Protocols in Molecular Biology*, John Wiley and Sons, Baltimore, Md. (1998), each of which is incorporated herein by reference. Nucleic acids can be fragmented using methods known in the art, including, for example, 1) physical methods such as acoustic shearing, sonication, or hydrodynamic shearing; or 2) enzymatic methods such as DNase I digestion, restriction endonucleases, or transposases; or 3) chemical fragmentation methods such as, for example, chemical shearing using heat and divalent metal cations. Nucleic acid selection based on an ideal fragment length can be performed using methods known in the art, including, for example, gel electrophoresis or bead-based size selection. The above steps result in the formation of a linear nucleic acid template containing the target sequence.
[0270] 2. Attachment Connector Sequence :
[0271] One or more adaptor sequences can be ligated to fragmented nucleic acids, such as linear nucleic acid templates, using methods known in the art, such as enzymatic methods using ligases, and / or commercially available library preparation kits, such as Nextera DNA Flex Library Preparation / Illumina DNA Preparation Kits (catalog numbers 20025519, 20025520, 20018704 and 20018705) or Qiagen QIAseq 1-Step Amplicon Library Kit (catalog number 180412).
[0272] After at least one adaptor sequence is ligated to a linear nucleic acid template, the splice oligonucleotide hybridizes to the 5' and 3' ends of the linear template to form a circular nucleic acid template containing the target sequence and one or more adaptor sequences. Alternatively, splice oligonucleotides are not required for end ligation, for example, when using CircLigase™ (Epicenter, Madison WI) or other enzymes capable of splice-free end ligation. Exonucleases can be used to remove residual linear fragments.
[0273] 3. Rolling circle amplification :
[0274] RCA can be performed using methods known in the art, including, for example, those described in Lizardi et al., Nat. Genet. 19:225-232 (1998). Generally, the method involves a polymerase, such as, for example, Φ29 (phi29) DNA polymerase, whose extension is annealed to a circular nucleic acid template, such that the polymerase generates a tandem single-stranded nucleic acid (“sense strand”) containing multiple tandem repeat sequences, each of which is complementary to the circular nucleic acid template.
[0275] In one configuration, RCA can be carried out initially in the presence of a low concentration of a polymer, such as a dendritic macromolecule (e.g., polyamidoamine (PAMAM)), and subsequently in the presence of the polymer. In one configuration, the RCA reaction is stopped by denaturing the polymerase, for example by heating the sample at 60°C, 65°C, 70°C, 75°C, 80°C, or higher. In one configuration, the RCA reaction is stopped by removing one or more components of the RCA, such as the polymerase and dNTPs. Components of the RCA can be removed, for example, by washing. Optionally, one or more antisense strands can be prepared by, for example, replicating the sense strand of the tandem strand using multiple substitution amplification (MDA).
[0276] 4. Multiple substitution amplification :
[0277] MDA can be performed using methods known in the art, including, for example, those described in Lizardi et al., Nat. Genet. 19:225-232 (1998). Generally, primers hybridize to one or more regions of a tandem single-stranded nucleic acid, and a polymerase (such as, for example, Φ29 DNA polymerase) extends the primers annealed to the tandem single-stranded nucleic acid to produce multiple single-stranded nucleic acids (multiple “antisense strands”).
[0278] RCA and MDA methods can be performed isothermally. Generally, the polymerase used for RCA or MDA is a chain displacement polymerase. Methods and reagents that can be used for RCA, MDA, or some combination thereof are described, for example, in Lizardi et al., Nat. Genet. 19:225-232 (1998); U.S. Patent Nos. 6,830,884; 6,797,474; 6,670,126; 6,576,448; 6,323,009; 6,280,949 or US 2007 / 0099208 A1, each of which is incorporated herein by reference.
[0279] RCA and / or MDA reactions can be performed in the presence of deoxyribonucleoside triphosphate (dATP), deoxyribonucleoside triphosphate (dTTP), deoxyribonucleoside triphosphate (dGTP), and deoxyribonucleoside triphosphate (dCTP) (or their analogues). The antisense strand generated during the MDA reaction can include adenine, guanine, cytosine, and thymine bases. MDA reactions can also be performed in the presence of deoxyribonucleoside triphosphate (dUTP) (or its analogues). When the MDA reaction is performed in the presence of dUTP as well as dATP, dTTP, dGTP, and dCTP, the generated antisense strand may also include uracil bases (indicated by an asterisk in the antisense strand) in addition to adenine, guanine, cytosine, and thymine bases. Antisense strands containing uracil bases can be digested after antisense strand sequencing.
[0280] Alternatively or additionally, the MDA reaction can be carried out in the presence of deoxyribonucleotide triphosphates having modified or uncanonical bases, such that the resulting antisense strand comprises one or more modified or uncanonical bases. Such modified or uncanonical bases can target the antisense strand for degradation, such as enzymatic digestion. For example, the MDA reaction can be carried out in the presence of modified or uncanonical deoxyribonucleotide triphosphates (e.g., deoxypseudouridine triphosphates), such that the resulting antisense strand comprises one or more modified or uncanonical nucleotides (e.g., deoxypseudouridine monophosphates).
[0281] Whether the base in the antisense strand is thymine or uracil (when the corresponding base in the sense strand is adenine) depends on the relative concentrations of dTTP and dUTP in the MDA reaction. The concentration of dUTP (or a deoxyribonucleotide triphosphate with modified or uncanonical bases, or a modified or uncanonical deoxyribonucleotide triphosphate) in the MDA reaction can be lower than the concentration of the other deoxyribonucleotide triphosphate in the MDA reaction, resulting in a low percentage of uracil bases (or modified or uncanonical bases, or modified or uncanonical nucleotides) present in the antisense strand. Uracil bases (or modified or uncanonical bases, or modified or uncanonical nucleotides) can be randomly distributed and present in a low percentage, such that the two antisense strands (any two antisense strands) include uracil bases (or modified or uncanonical bases, or modified or uncanonical nucleotides) at different positions.
[0282] The concentration of deoxyribonucleic acid triphosphates (e.g., dATP, dTTP, dGTP, or dCTP) or all deoxyribonucleic acid triphosphates in the RCA or MDA reaction may be about, at least, at least about, at most, or at most about 0.1 mM, 0.2 mM, 0.3 mM, 0.4 mM, 0.5 mM, 0.6 mM, 0.7 mM, 0.8 mM, 0.9 mM, 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 7 mM, 8 mM, 9 mM, 10 mM, 11 mM, 12 mM, 13 mM, 14 mM, 15 mM, 16 mM, 17 mM, 18 mM, 19 mM, 20 mM, 25 mM, 30 mM, 35 mM, 40 mM, 45 mM, 50 mM, 55 mM, 60 mM, 65 mM, 70 mM, 75 mM, 80 mM, 85 mM, 90 mM, 95 mM, 100 mM, or any value or range between any two of these values. The concentration of dUTP in the MDA reaction can be approximately, at least, at least about, at most or at most about 0.001 mM, 0.002 mM, 0.003 mM, 0.004 mM, 0.005 mM, 0.006 mM, 0.007 mM, 0.008 mM, 0.009 mM, 0.01 mM, 0.02 mM, 0.03 mM, 0.04 mM, 0.05 mM, 0.06 mM, 0.07 mM, 0.08 mM, 0.09 mM, 0.1 mM, 0.2 mM, 0.3 mM, 0.4 mM, 0.5 mM, 0.6 mM, 0.7 mM, 0.8 mM, 0.9 mM, 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 7 mM, 8 mM, ... mM, 10mM, 11 mM, 12 mM, 13 mM, 14 mM, 15 mM, 16 mM, 17 mM, 18 mM, 19 mM, 20 mM, or any value or range between any two of these values.
[0283] The ratio of the concentration of dUTP (or a deoxyribonucleotide triphosphate with modified or non-canonical bases, or a modified or non-canonical deoxyribonucleotide triphosphate) to the concentration of dTTP (or the concentration of another deoxyribonucleotide triphosphate, or the total concentration of deoxyribonucleotide triphosphates other than dUTP or a deoxyribonucleotide triphosphate with modified or non-canonical bases, or a modified or non-canonical deoxyribonucleotide triphosphate) can be approximately, at least At least about, at most or at most about 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1:73, 1:72, 1:71, 1:70, 1:69 1:68, 1:67, 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35, 1: 34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, or a value or range between any two of these values.
[0284] The percentage of deoxyribonucleic acid triphosphates (dUTP, or deoxyribonucleic acid triphosphates with modified or uncanonical bases, or modified or uncanonical deoxyribonucleic acid triphosphates) in the MDA reaction may be, about, at least about, at most or at most about 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any value or range between these values.
[0285] In the MDA reaction, the concentration of dUTP (or a deoxyribonucleotide triphosphate with modified or uncanonical bases, or a modified or uncanonical deoxyribonucleotide triphosphate) relative to another deoxyribonucleotide triphosphate such as dTTP can be low, resulting in a low percentage of uracil bases (or modified or uncanonical bases, or modified or uncanonical deoxyribonucleotides) present in the antisense strand. The ratio of a nucleotide having a uracil base (or a modified or non-canonical base, or a modified or non-canonical nucleotide) to a thymine base (or another base, or all non-uracil bases) can be, about, at least, at least about, at most or at most about 1:10000, 1:9000, 1:8000, 1:7000, 1:6000, 1:5000, 1:6000, 1:5000, 1:4000, 1:3000, 1:2000, 1:1000, 1:9 00, 1:800, 1:700, 1:600, 1:500, 1:400, 1:300, 1:200, 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1: 73, 1:72, 1:71, 1:70, 1:69, 1:68, 1:67, 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, or a value or range between any two of these values.The percentage of uracil bases in the antisense strand can be, about, at least about, at most or at most about 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any value or range between these values.
[0286] B. Solid-phase synthesis and solution-phase synthesis of tandems
[0287] Synthesizing tandem bodies can be carried out in solution (“solution-phase synthesis”) or attached to a solid surface (“solid-phase synthesis”). Generally, the stabilization of tandem bodies is not significantly affected by the method used to synthesize the tandem body requiring stabilization. Therefore, specifically, aspects of the method for forming a stabilized tandem body and any of the examples, aspects of the composition of a stabilized tandem body and any of the examples, or aspects of the method of use described herein and any of the examples, can utilize tandem bodies provided by solution-phase or solid-phase synthesis.
[0288] 1. solid phase synthesis :
[0289] In solid-phase synthesis, a linear nucleic acid is started, and different adaptor sequences are attached to its 5' and 3' ends. Alternatively, a linear nucleic acid with different 5' and 3' adaptor sequences can be started. The linear nucleic acid with different 5' and 3' adaptor sequences is annealed with a surface-bound oligonucleotide (such as a splint oligonucleotide) or a primer having regions complementary to the different 5' and 3' adaptor sequences of the linear nucleic acid, and oriented to position the 5' and 3' ends of the linear nucleic acid in close proximity. The 5' and 3' ends of the linear nucleic acid are then joined to circularize the linear nucleic acid. The surface-bound oligonucleotide is then extended to add multiple monomeric units of the original linear nucleic acid at its 3' end via RCA. The result is a tandem of polymers of the original linear nucleic acid tethered to the surface. Multiple second strands can then be synthesized using MDA.
[0290] 2. solution phase synthesis :
[0291] In solution-phase synthesis, the tandem polymer can be fully synthesized in solution and then deposited onto a surface for sequencing, or the tandem polymer can be partially synthesized in solution and then deposited onto a surface to complete the tandem polymer synthesis for sequencing. Alternatively, the tandem polymer can be fully synthesized in solution, stabilized with staple molecules in solution, and then deposited onto a surface for sequencing. Alternatively, the tandem polymer can be fully synthesized in solution, deposited onto a surface, and then stabilized with staple molecules prior to sequencing.
[0292] A linear nucleic acid can be started with, and different adaptor sequences can be ligated to the 5' and 3' ends of the linear nucleic acid. The linear nucleic acid sequence is then circularized by binding to a splice oligonucleotide in solution, thereby generating a circular template. Alternatively, splice oligonucleotides are not required for end-ligation, for example, when using CircLigase™ (Epicenter, Madison WI) or other enzymes capable of splice-free end-ligation of nucleic acids. In other configurations, a linear nucleic acid with different 5' and 3' adaptor sequences can be started with, and the linear nucleic acid can be circularized by binding to a splice oligonucleotide in solution, thereby generating a circular template. In yet another alternative embodiment, a circular template can be started. The circular template can be annealed with surface-bound oligonucleotides and can be RCA performed as described above, optionally followed by MDA. Alternatively, the circular template can be contacted with primers and RCA can be performed in solution. Optionally, MDA can be performed to generate multiple second strands. Subsequently, the tandem synthesized in solution can be deposited onto a surface for sequencing. Alternatively, staple molecules can be used to stabilize the tandem and then deposited onto a surface for sequencing.
[0293] C. Stabilizing the tandem structure using staple molecules.
[0294] In a first aspect of the invention, a method for forming a stabilized tandem strand is provided, comprising: (a) providing a plurality of tandem strands, wherein each tandem strand includes a plurality of instances of a target sequence and a plurality of instances of at least one adaptor sequence; and (b) contacting the plurality of tandem strands with staple molecules to form a plurality of stabilized tandem strands; wherein each staple molecule includes at least a first region and a second region; and wherein the first region and the second region of the staple molecule each hybridize with different instances of the at least one adaptor sequence. In any aspect and embodiment of this disclosure, it is particularly contemplated that references to the “target sequence” of a tandem strand may further encompass sequences that are reversed and complementary to the target sequence, and references to “at least one adaptor sequence” may further encompass sequences that are reversed and complementary to the at least one adaptor sequence.
[0295] As described in further detail herein, such as in the method and composition sections for forming a stable tandem, the staple molecule comprises at least a first region and a second region, wherein each of the first and second regions of the staple molecule hybridizes with different instances of at least one adaptor sequence of the tandem. The sequences of the first and second regions may be identical, partially conserved, or completely different. Furthermore, the sequences of the first and second regions may have the same length, or the first region may be greater than or equal to the length of the second region, or the second region may be greater than or equal to the length of the first region. The staple molecule may further comprise one or more of the following components: a reversible terminator portion, a blocking agent or blocking portion, a capping or capping portion, a double-stranded nucleic acid region, at least one nucleotide mismatch in an optional double-stranded portion, a 3' phosphate blocking agent, a spacer between the first and second regions, or any combination thereof. The spacer may be a polynucleotide sequence or a non-nucleotide polymer linker. The non-nucleotide polymer linker may be a branched polyelectrolyte species, such as polyethylene glycol, or a dendritic macromolecule, such as poly(amide amine). In some configurations, at least one of the first and second regions of the staple molecule may further include one or more of the following components: a 3' end that allows extension, a 3' end that allows the formation of a ternary complex, a 3' end that allows reversible termination, or any combination thereof. Other aspects, embodiments, and considerations are discussed below.
[0296] In some embodiments, providing multiple tandem polymerases in step (a) includes performing RCA with a polymerase such as a chain replacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template contains a target sequence and at least one adaptor sequence. The circular nucleic acid template (such as, for example, a circular nucleic acid template containing a target sequence and at least one adaptor sequence) can be single-stranded or double-stranded. One or both strands of a double-stranded nucleic acid may lack a 3' and a 5' end. One strand of a double-stranded nucleic acid may have a gap (the absence of at least one nucleotide monomer relative to the other strand) or a nick (the absence of a phosphodiester bond between two nucleotide monomers), provided that the other strand is circular. Any of a variety of polymerases may be used in the methods or compositions described herein. Non-limiting examples of polymerases that may be used include naturally occurring polymerases and modified versions thereof, including but not limited to mutants, recombinants, fusions, genetically modified, chemically modified, synthetics, analogs, etc.
[0297] In some embodiments, providing multiple tandem sequences in step (a) includes RCA with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the method further includes circularizing a linear nucleic acid template comprising the target sequence and at least one adaptor sequence. Circularization of the linear nucleic acid template can be performed using any method known in the art. In some embodiments, the circularization of the linear nucleic acid template is performed using a ligase. Circularization of the linear nucleic acid template can be performed using any suitable ligase, such as, for example, T4 DNA ligase.
[0298] In some embodiments, providing multiple tandem sequences in step (a) includes performing RCA with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the method further comprises circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template comprises a first adaptor sequence of target sequence 3' and a second adaptor sequence of target sequence 5'.
[0299] In some embodiments, providing multiple tandem sequences in step (a) includes RCA with a chain displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template contains a target sequence and at least one adaptor sequence. In a further embodiment, the method further includes circularizing a linear nucleic acid template containing a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template contains a first adaptor sequence at 3' of the target sequence and a second adaptor sequence at 5' of the target sequence. In even further embodiments, the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with a splint oligonucleotide. A “splint oligonucleotide” is an oligonucleotide that, when hybridized with other polynucleotides such as, for example, a first adaptor sequence, or a second adaptor sequence, or a (nucleic acid) template, or a linear (nucleic acid) template, acts as a “splice” to position the polynucleotides adjacent to each other so that they can be ligated together. In some embodiments, the splint oligonucleotide is DNA or RNA. The splint oligonucleotide may include a nucleotide sequence partially complementary to the nucleotide sequence of two or more different oligonucleotides. In some embodiments, the splint oligonucleotide facilitates the ligation of a “donor” oligonucleotide and a “recipient” oligonucleotide. Generally, RNA ligase, DNA ligase, or another type of ligase is used to join two nucleotide sequences together. In some embodiments, the length of the splice oligonucleotide is between 10 and 50 nucleotides, for example, between 10 and 45 nucleotides, 10 and 40 nucleotides, 10 and 35 nucleotides, 10 and 30 nucleotides, 10 and 25 nucleotides, or 10 and 20 nucleotides. In some embodiments, the length of the splice oligonucleotide is between 15 and 50 nucleotides, 15 and 45 nucleotides, 15 and 40 nucleotides, 15 and 35 nucleotides, 15 and 30 nucleotides, or 15 and 25 nucleotides. Alternatively, splice oligonucleotides are not required for end-joining, for example, when using CircLigase™ (Epicenter, Madison WI) or other enzymes capable of splice-free end-joining of nucleic acids.
[0300] In some embodiments, providing multiple tandem sequences in step (a) includes RCA with a chain displacement polymerase using primers hybridizing to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the method further comprises circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template comprises a first adaptor sequence at 3' of the target sequence and a second adaptor sequence at 5' of the target sequence. In even further embodiments, the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with a splice oligonucleotide. In a particular embodiment of even further embodiments, the splice oligonucleotide is a primer hybridizing to the circularized nucleic acid.
[0301] In some embodiments, providing multiple tandem sequences in step (a) includes performing RCA with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the method further includes circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template comprises a first adaptor sequence of target sequence 3' and a second adaptor sequence of target sequence 5'. In even further embodiments, the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with a splint oligonucleotide. In a particular embodiment of even further embodiments, the splint oligonucleotide is removed prior to RCA. Alternatively, the splint oligonucleotide may be removed during or after RCA.
[0302] In some embodiments, providing multiple tandems in step (a) includes performing RCA with a chain-displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the primers are immobilized on a surface during RCA. Suitable surfaces include, but are not limited to, structured surfaces, planar substrates, hydrogels, nanopore arrays, microparticles, nanoparticles, flow cell surfaces, surfaces of solid supports, or surfaces of solid supports within a flow cell. The surface may be planar or curved. As discussed in further detail below, the solid support may be made of any of a variety of materials used for analytical biochemistry. Suitable materials may include, for example, glass, polymeric materials, silicon, quartz (fused silica), borofloat glass, silica, silica-based materials, carbon, metals, optical fibers or fiber bundles, sapphire, or plastic materials. Materials may be selected based on properties desired for a particular application. For example, materials transparent to a desired wavelength of radiation may be used in analytical techniques that utilize radiation at that wavelength. Conversely, it may be desirable to select materials that do not transmit radiation at a particular wavelength (e.g., opaque, absorptive, or reflective materials). Wavelength regions that may or may not pass through a particular material include, for example, UV, VIS (e.g., red, yellow, green, or blue), or IR. Other properties of materials that can be utilized include inertness or reactivity to certain reagents used in downstream processes (such as those described herein), ease of handling, or low manufacturing cost.
[0303] In some embodiments, providing multiple tandem sequences in step (a) includes performing RCA with a chain displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template contains a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA.
[0304] In some embodiments, providing multiple tandem polymerases in step (a) includes performing RCA with a chain displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template contains a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the method further includes depositing multiple stabilized tandem polymerases onto a surface after contact step (b).
[0305] In some embodiments, providing multiple tandem polymerases in step (a) includes performing RCA with a chain displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template contains a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the method further includes depositing multiple stabilized tandem polymerases onto a surface after contact step (b). In even further embodiments, the surface is a structured surface.
[0306] In some embodiments, providing multiple tandem polymerases in step (a) includes performing RCA with a chain displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template contains a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the method further includes depositing multiple tandem polymerases on a surface prior to contact step (b).
[0307] In some embodiments, providing multiple tandem polymerases in step (a) includes performing RCA with a chain displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template contains a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the method further includes depositing multiple tandem polymerases on a surface prior to contact step (b). In even further embodiments, the surface is a structured surface.
[0308] In some embodiments, step (a) of providing multiple tandem strands includes performing RCA with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the RCA generates a sense strand, and the method further includes amplifying the sense strand to generate multiple antisense strands.
[0309] In some embodiments, step (a) of providing multiple tandem strands includes RCA with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, RCA generates a sense strand, and the method further includes amplifying the sense strand to generate multiple antisense strands. In yet another further embodiment, at least some of the staple molecules hybridize with adaptor sequences of multiple antisense strands.
[0310] In some embodiments, step (a) of providing multiple tandem strands includes RCA with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, RCA generates a sense strand, and the method further includes amplifying the sense strand to generate multiple antisense strands. In yet another further embodiment, step (b) further includes hybridizing staple molecules with the sense strands.
[0311] In some embodiments, step (a) of providing multiple tandem strands includes RCA with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, RCA generates a sense strand, and the method further includes amplifying the sense strand to generate multiple antisense strands. In yet another embodiment, step (b) further includes hybridizing staple molecules with the sense strand. In even further embodiments, at least some of the staple molecules are staple primers for amplifying the sense strand to generate multiple antisense strands.
[0312] In some embodiments, providing multiple tandem sequences in step (a) includes RCA with a chain displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, contact step (b) occurs during RCA step (a).
[0313] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site.
[0314] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site. In a further embodiment, each instance of at least one adaptor sequence further includes a tag region, optionally wherein the tag region is a sample index.
[0315] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site. In a further embodiment, each instance of at least one adaptor sequence further includes a splint binding site.
[0316] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site. In further embodiments, each instance of at least one adaptor sequence further includes a variable region, optionally wherein the variable region is a unique molecular identifier (UMI). Generally, a UMI is a nucleotide sequence applied to or identified in a polynucleotide that can be used to distinguish individual nucleic acid molecules present in an initial reaction from one another. In some cases, a UMI may contain about 5 to about 20 nucleotides. Alternatively, a UMI may contain fewer than about 5 or more than 20 nucleotides. A UMI can be a unique sequence that varies among individual nucleic acid molecules. In some cases, a UMI can be a random sequence. In some cases, a UMI can be a predetermined sequence. In a sequencing reaction, a UMI can be sequenced along with the nucleic acid molecules associated with it to determine whether the read sequence is a sequence of one source nucleic acid molecule or a sequence of another source nucleic acid molecule. The term “UMI” is used herein to refer to both the sequence information of the polynucleotide and the physical polynucleotide containing that sequence information. Other examples of UMIs and their uses are provided, for example, in US 2016 / 0319345 A1, which is incorporated herein by reference.
[0317] In some embodiments, each tandem strand includes a sense strand that hybridizes with multiple antisense strands.
[0318] In some embodiments, each tandem strand includes a sense strand that hybridizes with multiple antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands.
[0319] In some embodiments, each tandem strand includes a sense strand that hybridizes with multiple antisense strands. In a further embodiment, the sense strand does not contain uracil and the multiple antisense strands do contain uracil.
[0320] In some embodiments, all instances of the target sequence within each tandem body are identical.
[0321] In some embodiments, instances of the target sequence may include sense sequences or antisense sequences.
[0322] In some embodiments, a plurality of series bodies are provided fixed to a surface in a flow cell.
[0323] In some embodiments, a plurality of series of bodies are provided fixed to a surface in a flow cell. In a further embodiment, the surface is a continuous surface.
[0324] In some embodiments, a plurality of series members are provided and fixed to a surface in the flow cell. In a further embodiment, the plurality of series members are fixed to bonding sites on a structured surface of the flow cell.
[0325] In some embodiments, a plurality of tandem bodies are provided in solution in step (a) and deposited on the surface of a flow cell in step (b).
[0326] In some embodiments, the sequences of the first region and the second region are the same.
[0327] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows for extension. The extension can be mediated by a polymerase or a ligase.
[0328] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows for extension. In further embodiments, the 3' end of the staple molecule hybridizes within 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 10, 5, or fewer nucleotides from the 3' end of the target sequence. In an exemplary embodiment, the 3' end of the staple molecule hybridizes within 40 nucleotides from the 3' end of the target sequence.
[0329] In some embodiments, at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes to a primer binding site of at least one adaptor sequence.
[0330] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows the formation of a ternary complex.
[0331] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows the formation of the ternary complex. In some embodiments, the 3' end that allows the formation of the ternary complex is reversibly terminated. Reversible termination can be performed using any reversible terminator. Exemplary reversible terminators, such as those in which the 3'-OH group is partially replaced by a 3'-ONH2 portion, are set forth in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference.
[0332] In some embodiments, neither the first nor the second region of the staple molecule contains a 3' end that allows the ternary complex to form or extend.
[0333] In some embodiments, the plurality of tandem units include repeating sequence units, each of which includes a target sequence and at least one adapter sequence.
[0334] In some embodiments, the plurality of tandem units include repeating sequence units, each repeating sequence unit comprising a target sequence and at least one interleave sequence. In a further embodiment, the at least one interleave sequence within the sequence unit comprises a first interleave sequence of target sequence 3' and a second interleave sequence of target sequence 5'.
[0335] In some embodiments, the plurality of tandem units include repeating sequence units, each repeating sequence unit comprising a target sequence and at least one adaptor sequence. In a further embodiment, the at least one adaptor sequence within the sequence unit comprises a first adaptor sequence of target sequence 3' and a second adaptor sequence of target sequence 5'. In yet another further embodiment, a first region of the staple molecule hybridizes with the first adaptor sequence and a second region of the staple molecule hybridizes with the second adaptor sequence.
[0336] In some embodiments, the plurality of tandem units comprise repeating sequence units, each repeating sequence unit comprising a target sequence and at least one adaptor sequence. In a further embodiment, at least one adaptor sequence within the sequence unit comprises a first adaptor sequence of the target sequence 3' and a second adaptor sequence of the target sequence 5'. In yet another further embodiment, at least one of the first and second regions of the staple molecule comprises the 3' end of the staple molecule and hybridizes with a 3' adaptor.
[0337] In some embodiments, the first and second regions of the staple molecule are of equal length. Alternatively, the first region of the staple molecule may be longer than the second region. In other embodiments, the second region of the staple molecule may be longer than the first region. In some embodiments, the length of the first region of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. In some embodiments, the length of the second region of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.
[0338] In some embodiments, the length of each of the first and second regions of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. In an exemplary embodiment, the length of each of the first and second regions of the staple molecule is at least 10 nucleotides.
[0339] In some embodiments, the first and second regions of the staple molecule each hybridize with different instances of at least one adaptor sequence spaced at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, or 200 or more nucleotides apart. In an exemplary embodiment, the first and second regions of the staple molecule each hybridize with different instances of at least one adaptor sequence spaced at least 100 nucleotides apart.
[0340] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, nucleotide incorporation-preventing, or ternary complex formation-preventing cap. The blocking component can be any blocking part. A blocking part is a portion of a nucleotide that inhibits or prevents the 3' oxygen of a nucleotide from forming a covalent bond with the next correct nucleotide during nucleic acid polymerization. The blocking part of a “reversible termination” nucleotide may be removed from a nucleotide analog or otherwise modified to allow the 3' oxygen of the nucleotide to covalently attach to the next correct nucleotide. Such a blocking part is referred to herein as a “reversible termination part.” The blocking part does not need to prevent or inhibit the formation of a ternary complex at the 3' end of the nucleic acid to which the blocking part is attached. The cap can be any capping part. The capping part can have a positive or negative charge that inhibits or prevents ternary complex formation. The capping part may include a ligand that binds to a receptor to inhibit or prevent ternary complex formation, such as biotin (or its analogues) that binds to streptavidin (or its analogues), an epitope that binds to an antibody (or its functional fragment), a carbohydrate that binds to a lectin, etc. Further examples of the capped portion are described in U.S. Patent Application Publication No. 2020 / 0032322A1 or Turcatti et al. Nucl. Acids. Res. 36(4) e25 (2008), each of which is incorporated herein by reference.
[0341] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a nucleotide incorporation blocker, or a ternary complex formation cap. In a further embodiment, the 3' end of the staple molecule includes a nucleotide incorporation blocker.
[0342] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a blocking agent to prevent nucleotide incorporation, or a cap to prevent ternary complex formation. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation.
[0343] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, nucleotide incorporation blocking agent, or ternary complex formation prevention cap. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation. In yet another embodiment, the 3' end of the staple molecule is capped by a portion preventing ternary complex formation.
[0344] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a blocking agent to prevent nucleotide incorporation, or a cap to prevent ternary complex formation. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation. In yet another embodiment, the 3' end of the staple molecule includes a mismatch and a blocking agent, optionally wherein the blocking agent is a 3' phosphate blocking agent.
[0345] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. Any suitable spacer known in the art can be used. In some instances, the spacer may contain a polynucleotide sequence. Alternatively, the spacer may contain a non-nucleotide polymer linker, such as, for example, a branched polyelectrolyte species. Non-limiting examples of branched polyelectrolyte species include polyethylene glycol (PEG) and dendritic macromolecules.
[0346] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, at least some of the staple molecules do not contain spacers.
[0347] In some embodiments, at least some of the first and second regions in the staple molecule are separated by spacers. In a further embodiment, the spacers comprise a polynucleotide sequence. The polynucleotide sequence has a length of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.
[0348] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise polynucleotide sequences. In yet another embodiment, the length of the polynucleotide sequence is variable.
[0349] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise a polynucleotide sequence. In yet another embodiment, the polynucleotide sequence comprises a double-stranded DNA sequence.
[0350] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacer comprises a polynucleotide sequence. In yet another embodiment, the polynucleotide sequence comprises a double-stranded DNA sequence. In still a further embodiment, the first and second regions each include a 3' end that hybridizes to the tandem strand.
[0351] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, the spacers have variable lengths on different staple molecules.
[0352] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, the spacers comprise non-nucleotide polymer connectors.
[0353] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise polyethylene glycol (PEG). PEG may comprise PEG with an average molecular weight of about 200 Daltons (e.g., PEG-200) to about 8000 Daltons (e.g., PEG-8000).
[0354] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. Species of dendritic macromolecules that can be used in the methods, compositions, and systems disclosed herein include, but are not limited to, branched polyamines comprising protonated structures that interact with the negatively charged DNA backbone to form a complex. It should be understood that, in the embodiments described herein, the adaptor element can still be used in conjunction with the branched polyamine for alternative purposes or to provide improved binding properties of the dendritic macromolecule species to nucleic acids. Dendritic macromolecule species may comprise controlled terminal surface chemistry having one or more functional groups, including but not limited to amine, carboxyl, and hydroxyl groups. Dendritic macromolecule species can be obtained as generation 0 (G0) through generation 10 (G10), with the number of branches in each generation being twice that of the previous generation. Thus, G0 = 4 branches, G1 = 8 branches, and so on. In some embodiments, the branched polyelectrolyte is a poly(amidoamine) dendritic macromolecule species (also known as PAMAM), such as the G2 PAMAM dendritic macromolecule having 16 branches and an amine (NH2) terminal surface chemistry. Non-limiting examples of branched polyelectrolytes also include G4 (64 branches with amine terminal groups) and G5 (128 branches with amine terminal groups) PAMAM dendritic macromolecule species.
[0355] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer connectors. In yet another embodiment, the nonnucleotide polymer connectors comprise dendritic macromolecules. In even further embodiments, the dendritic macromolecules comprise polyamidoamine (PAMAM). Any species of PAMAM described above, both those specifically described and those implied by the entirety of this disclosure, may be used.
[0356] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. In even further embodiments, the staple molecule hybridizes with 3, 4, 5, 6, 7, 8, 9, or 10 or more instances of at least one adaptor. In an exemplary embodiment, the staple molecule hybridizes with 3 or more instances of at least one adaptor. In another exemplary embodiment, the staple molecule hybridizes with 10 or more instances of at least one adaptor sequence.
[0357] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. In even further embodiments, the staple molecule hybridizes with three or more instances of at least one adaptor. In a particular embodiment of even further embodiments, a majority of the staple molecule hybridizes with adaptor sequences flanking at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, or 200 or more different target sequences. In an exemplary embodiment, a majority of the staple molecule hybridizes with adaptor sequences flanking at least 100 different target sequences.
[0358] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. In even further embodiments, neither the first nor second region of the staple molecule contains a 3' end that allows the formation or extension of the ternary complex.
[0359] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, the spacers are coupled to the 3' end of the first region.
[0360] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, the spacers are coupled to the 3' end of the first region. In yet another further embodiment, the spacers are coupled to the 5' end of the second region.
[0361] In some embodiments, at least some of the first and second regions of the staple molecules are separated by spacers. In a further embodiment, the spacers are coupled to the 3' end of the first region. In yet another further embodiment, the spacers are coupled to the 3' end of the second region, and the staple molecules do not act as primers.
[0362] In some embodiments, multiple staple molecules hybridize with the same tandem strand.
[0363] D. Stabilization indicators
[0364] As mentioned above, tandems typically exhibit a structure that tends to diffuse or become blurred during sequencing, and thus show a decrease in quality, intensity, and read length during tandem sequencing. Therefore, tandem stabilization can be measured using various metrics related to 1) changes in tandem volume or size, 2) signal intensity and resolution from the tandem during sequencing, and 3) changes in read length during tandem sequencing. During sequencing, individual fields of view are collected and called patches. Single-row patches collected across flow cell lanes are called samples. Metrics related to stabilization are collected during the sequencing reaction.
[0365] 1. Changes in volume or size :
[0366] Although tandems and nucleic acid clusters are inherently three-dimensional structures and therefore can exhibit volume changes, the images acquired during sequencing runs are typically two-dimensional. The different dimensions of a sequencing image are described as either the "column direction" (i.e., the vertical plane) or the "row direction" (i.e., the horizontal plane). Full width at half maximum (FWHM) refers to the normal distribution of the fluorescence signal of each tandem or cluster across a certain number of pixels. FWHM can be quantized in the column and / or row directions. Unstabilized tandems will exhibit changes in shape / size, resulting in an increase in FWHM in one or both directions during sequencing runs. Tandems stabilized with staple molecules are resistant to changes in FWHM.
[0367] 2. Signal strength and signal resolution :
[0368] During sequencing, raw intensity values for each nucleotide check (ATGC) are collected for all cycles of the sequencing run and are referred to as “ON intensity”. Additionally, raw intensity values for the observed background, or “OFF intensity”, are collected. The 50th percentile of the ratio of the raw intensity value of each check to the raw intensity value of the observed background is calculated as the average over the entire lane of the flow cell (“ON / OFF intensity”). When tandem particles become dispersed, molecular dispersion causes spectral bleed-out, interfering with the distinction between tandem particle intensity values and background intensity.
[0369] The density of tandems on the surface significantly affects signal resolution during sequencing. Lower dot density makes the signal more likely to originate from a single tandem (also known as a "monoclonal" signal). As dot density increases, spectral bleeding between adjacent tandems can cause multiple dots to be "treated" as monoclonal signals by the sequencer, even though these dots originate from separate tandems. Further complicating matters, the probability of multiple dots being treated as monoclonal signals increases as tandems become more dispersed. The inability to separate one dot from another inhibits longer sequencing reactions, resulting in shorter read lengths.
[0370] 3. Segment length :
[0371] Sequencing read lengths can vary considerably based on the sequencing method employed and various other considerations, such as those described above and others known in the art. When tandem sequences become unstable, the length of the sequencing run should be shortened to maintain the integrity of the sequencing reads. However, this results in shorter read lengths, which may prevent the capture of at least some of the variant sequences present. In terms of read length, stabilization can therefore be measured by the change in sequencing run length (i.e., the number of cycles in the sequencing reaction) or by the change in the length of the resulting reads.
[0372] E. Conditions for tandem formation and sequencing
[0373] Tandem formation and sequencing of tandems are generally described in U.S. Patent Publication No. 2022 / 0349002, which is incorporated herein by reference. In some aspects, tandems can be formed by rolling circle amplification (RCA), such as in solution prior to deposition on a solid surface, or by primers bound to a solid surface. In some aspects, the primers are clip primers, on which a template containing the target sequence and the adaptor sequence is circularized prior to RCA. During tandem formation, compaction additives such as PAMAM can be added or their concentration increased.
[0374] IV. Composition
[0375] Any aspect of the compositions described herein and any of the examples can be used in the above-described method of forming a stabilized tandem or in any aspect of the following methods of use and / or examples.
[0376] A. Nucleic acid
[0377] The nucleic acids used in the methods or compositions described herein may be deoxyribonucleic acid (DNA), such as genomic DNA, synthetic DNA, amplified DNA, complementary DNA (cDNA), etc. Alternatively, the nucleic acids used in the methods or compositions described herein may also be ribonucleic acid (RNA), such as mRNA, ribosomal RNA, tRNA, etc. Furthermore, the nucleic acids used in the methods or compositions described herein may also be nucleic acid analogs. For example, nucleic acid analogs can be used as templates for the amplification or sequencing processes described herein. The nucleic acids used herein, for example, as templates for generating tandem complexes or as targets for sequencing, may be derived from biological sources, synthetic sources, or amplification products. The primers used herein may include or may be DNA, RNA, or analogs thereof.
[0378] Nucleic acids can be obtained from preparative methods such as genomic, transcriptomic, or other nucleic acid isolation; genomic fragmentation; gene cloning and / or amplification. One or more nucleic acids can be obtained from amplification techniques such as polymerase chain reaction (PCR), emulsion PCR, random primer amplification, RCA, MDA, etc. RCA and MDA may be particularly useful for producing tandem products. Exemplary methods for isolating, amplifying, and fragmenting nucleic acids to produce templates for analysis on arrays are described in U.S. Patent Nos. 6,355,431 or 9,045,796, each of which is incorporated herein by reference. Amplification can also be performed using methods described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition, Cold SpringHarbor Laboratory, New York (2001) or Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1998), each of which is incorporated herein by reference.
[0379] Nucleic acid templates containing the target sequence subjected to the methods described herein can be derived from or generated from a sample. The sample may include one or more organisms. Nucleic acid templates can be obtained or derived from the sample without performing a polymerase chain reaction (PCR). Nucleic acid templates can be obtained or derived from the sample by performing a polymerase chain reaction for several cycles, such as up to one cycle, two cycles, three cycles, four cycles, five cycles, six cycles, seven cycles, eight cycles, nine cycles, or ten cycles or more. This document considers nucleic acid templates (or target sequences) of different lengths. The length of the nucleic acid can be at least, at least about, at most, or at most about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, or 270. 1, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 53 0, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 7 Nucleotides of 90, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000 or any of these values.
[0380] Exemplary organisms from which nucleic acids can be derived include, for example, mammals such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cattle, cats, dogs, primates, humans, or non-human primates; plants such as Arabidopsis thaliana, maize, sorghum, oats, wheat, rice, rapeseed, or soybeans; algae such as Chlamydomonas reinhardtii; nematodes such as Caenorhabditis elegans; insects such as Drosophila melanogaster, mosquitoes, fruit flies, or bees; arachnids such as spiders; fish such as zebrafish; reptiles; and amphibians such as frogs or Xenopus clawed frogs. Nucleic acids can also be derived from prokaryotes, such as bacteria, such as *Dictyophora laevis*; slime molds, such as *dictyostelium discoideum*; fungi, such as *Pneumocystis carinii*, *Takifugu rubripes*, yeasts, *Saccharomyces cerevisiae*, or *Schizosaccharomyces pombe*; or *Plasmodium falciparum*. Nucleic acids can also be derived from prokaryotes, such as bacteria, such as *Escherichia coli*, *Staphylococci*, or *Mycoplasma pneumoniae*; archaea; viruses, such as hepatitis C virus or human immunodeficiency virus; or viroids. Nucleic acids can be derived from a homogeneous culture or a group of organisms or alternatively from a collection of several different organisms (e.g., a community or ecosystem).
[0381] B. Primers
[0382] As described above, a primer is a nucleic acid having a sequence that binds to a nucleic acid at or near a template sequence. Typically, primers bind in a conformation that allows template replication, for example, by polymerase extension via the primer. In some embodiments of this disclosure, the staple molecule is configured such that a first and / or second region of the staple molecule is annealed with one or more adaptor sequences and then extended to the target sequence by polymerase. In such embodiments, the staple molecule is referred to as a "staple primer" or acts as a "staple primer." In some embodiments, the staple molecule and "additional primers" are used in one or more steps of a given method or composition. "Additional primers" may refer to primers as previously defined above, or a single staple molecule configured as a staple primer. To depict embodiments where the additional primer is not a staple primer, the additional primer may be referred to as a "non-staple primer."
[0383] In some embodiments, primers may have the same length. In other embodiments, primers may have different lengths. The length of primers (or two or more primers of one type or group, or each primer of one type or group) can be: about, at least, at least about, at most, at most about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92 The lengths of 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, and 1000 nucleotides. The length of the primers (or two or more primers of a type or population, or each primer of a type or population) can be, about, at least, at least about, at most, at most about 100 A, 110 A, 120 A, 130 A, 140 A, 150 A, 160 A, 170 A, 180 A, 190 A, 200 A, 300 A, 400 A, 500 A, 600 A, 700 A, 800 A, 900 A, 1000 A, 0.2 pm, 0.3 pm, 0.4 pm, 0.5 pm, 0.6 pm, 0.7 pm, 0.8 pm, 0.9 pm, 1 pm, 2 pm, 3 pm, 4 pm, 5 pm, 6 pm, 7 pm, 8 pm, 9 pm, 10 pm, or a value or range between any two of these values.
[0384] The ratio of the length of a primer of a certain type or population (e.g., capture primer) to the length of primers of the same type or population (e.g., capture primer), or the ratio of the length of a primer of a certain type or population (e.g., capture primer) to the length of primers of another type or population (e.g., amplification primer), can vary. The ratio of the lengths of two primers of one type or population, or the ratio of the lengths of two primers of different types or populations, can be approximately, at least, at least about, at most, or at most about 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1:73, 1:72, 1:71, 1:70, 1:69, 1:6 8, 1:67, 1:66, 1:65, 1:64, 1:63, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1 :23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 2 7:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:172:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, or a value or range between any two of these values.
[0385] C. Polymerase
[0386] Any of a variety of polymerases can be used in the methods described herein, for example, to replicate nucleic acid templates, form tandem complexes, form nucleic acid clusters, form stabilized tandem complexes, etc. Polymerases that can be used include naturally occurring polymerases and their modified variants, including but not limited to mutants, recombinants, fusions, genetically modified, chemically modified, synthetics, and analogs. Naturally occurring polymerases and their modified variants are not limited to polymerases capable of catalyzing polymerization reactions. Optionally, their naturally occurring and / or modified variants have the ability to catalyze polymerization reactions under at least one condition not used during the formation or inspection of the stabilized ternary complex. Optionally, the naturally occurring and / or modified variants participating in the stabilized ternary complex have modified properties, such as enhanced binding affinity for nucleic acids, reduced binding affinity for nucleic acids, enhanced binding affinity for nucleotides, reduced binding affinity for nucleotides, enhanced specificity for the next correct nucleotide, reduced specificity for the next correct nucleotide, reduced catalytic rate, catalytic inactivation, etc. Mutant polymerases include, for example, polymerases in which one or more amino acids are replaced by other amino acids or inserted or deleted one or more amino acids. Exemplary polymerase mutants that can be used to form stable ternary complexes include, for example, those set forth in U.S. Patent Application Serial No. 15 / 866,353, disclosed in U.S. Patent Application Publication No. 2018 / 0155698 A1; U.S. Patent Application Publication No. 2017 / 0314072 A1 or 2020 / 0087637 A1, each of which is incorporated herein by reference.
[0387] Modified polymerases include polymerases containing an exogenous labeled moiety (e.g., an exogenous fluorophore) that can be used to detect polymerases. Optionally, the labeled moiety can be attached after the polymerase has been at least partially purified using protein separation techniques. For example, the exogenous labeled moiety can be covalently linked to the polymerase using a free thiol or free amine moiety of the polymerase. This may involve covalently linking the polymerase via a side chain of a cysteine residue or via a free amino group at the N-terminus. The exogenous labeled moiety can also be attached to the polymerase via protein fusion. Exemplary labeled moiety that can be attached via protein fusion include, for example, green fluorescent protein (GFP), phycobiliproteins (e.g., phycocyanin and phycoerythrin), or wavelength-shifted variants of GFP or phycobiliproteins. In some embodiments, the exogenous label on the polymerase can act as a member of a FRET pair. The other member of the FRET pair can be an exogenous label that is attached to a nucleotide that binds to the polymerase in a stable ternary complex. Thus, the stable ternary complex can be detected or identified via FRET.
[0388] Alternatively, polymerases involved in stable ternary complexes or used for extending or modifying primers do not require attachment to a foreign label. For example, polymerases do not require covalent attachment to a foreign label. Instead, polymerases may lack any label until they associate with labeled nucleotides and / or labeled nucleic acids (e.g., labeled primers and / or labeled templates).
[0389] Different activities of polymerases can be utilized in the methods described herein. Polymerases can be used, for example, in template amplification processes, primer modification processes such as primer extension or primer capping steps, inspection steps, or combinations thereof. Different activities may arise from structural differences (e.g., via native activity, mutation, or chemical modification). However, polymerases can be obtained from a variety of known sources and applied in accordance with the teachings set forth herein and the accepted activities of polymerases. Available DNA polymerases include, but are not limited to, bacterial DNA polymerases, eukaryotic DNA polymerases, archaea DNA polymerases, viral DNA polymerases, and bacteriophage DNA polymerases. Bacterial DNA polymerases include *Escherichia coli* DNA polymerases I, II, and III, IV, and V; the Klenow fragment of *Escherichia coli* DNA polymerase; *Clostridium stercorarium* (Cst) DNA polymerase; *Clostridium thermocellum* (Cth) DNA polymerase; and *Sulfolobus sofataricus* (Sso) DNA polymerase. Eukaryotic DNA polymerases include DNA polymerases a, b, g, d, ε, h, z, l, s, m, and k, as well as Rev1 polymerase (terminal deoxycytidine transferase) and terminal deoxynucleotide transferase (TdT). Viral DNA polymerases include T4 DNA polymerase, phi-29 DNA polymerase, GA-1, phi-29-like DNA polymerase, PZA DNA polymerase, phi-15 DNA polymerase, Cp1 DNA polymerase, Cp7 DNA polymerase, T7 DNA polymerase, and T4 polymerase. Other available DNA polymerases include thermostable and / or thermophilic DNA polymerases, such as *Thermus aquaticus* (Taq) DNA polymerase, *Thermus filiformis* (Tfi) DNA polymerase, *Thermococcus zilligi* (Tzi) DNA polymerase, *Thermus thermophilus* (Tth) DNA polymerase, *Thermus flavus* (Tfl) DNA polymerase, *Pyrococcus woesei* (Pwo) DNA polymerase, *Pyrococcus furiosus* (Pfu) DNA polymerase and Turbo Pfu DNA polymerase, *Thermococcus litoralis* (Tli) DNA polymerase, and species of the genus *Pyrococcus* (sp.).GB-D polymerase, *Thermotoga maritima* (Tma) DNA polymerase, *Bacillus stearothermophilus* (Bst) DNA polymerase, *Pyrococcus Kodakaraensis* (KOD) DNA polymerase, Pfx DNA polymerase, *Thermococcus sp.* JDF-3 (JDF-3) DNA polymerase, *Thermococcus gorgonarius* (Tgo) DNA polymerase, *Thermococcus acidophilium* DNA polymerase, *Sulfolobus acidocaldarius* DNA polymerase, *Thermococcus* sp. *go* N-7 DNA polymerase, *Pyrodictium ocultum* DNA polymerase, *Methanococcus voltae* DNA polymerase, *Methanococcus thermoautotrophicum* DNA polymerase, *Methanococcus johnsonii* DNA polymerase. DNA polymerases including *Jannaschii*, *Desulfurococcus* strain TOK DNA polymerase (D. Tok Pol), *Pyrococcus abyssi* DNA polymerase, *Pyrococcus horikoshii* DNA polymerase, *Pyrococcus islandicum* DNA polymerase, *Thermococcus fumicolans* DNA polymerase, *Aeropyrum pernix* DNA polymerase, and heterodimeric DNA polymerases DP1 / DP2. Engineered and modified polymerases may also be used in conjunction with the disclosed techniques. For example, a modified version of the extreme thermophilic marine archaea *Thermococcus* species 9°N (e.g., Therminator DNA polymerase from New England BioLabs Inc.; Ipswich, MA) may be used. Other available DNA polymerases, including 3PDX polymerase, are disclosed in US 8,703,461, the disclosure of which is incorporated herein by reference.
[0390] Available RNA polymerases include, but are not limited to, viral RNA polymerases such as T7 RNA polymerase, T3 polymerase, SP6 polymerase and Kll polymerase; eukaryotic RNA polymerases such as RNA polymerase I, RNA polymerase II, RNA polymerase III, RNA polymerase IV and RNA polymerase V; and archaea RNA polymerases.
[0391] Another type of polymerase available is reverse transcriptase. Exemplary reverse transcriptases include, but are not limited to, HIV-1 reverse transcriptase from human immunodeficiency virus type 1 (PDB 1HMV), HIV-2 reverse transcriptase from human immunodeficiency virus type 2, M-MLV reverse transcriptase from Moloney murine leukemia virus, AMV reverse transcriptase from avian myeloblastosis virus, and telomerase reverse transcriptase for maintaining eukaryotic chromosome telomeres.
[0392] Polymerases with inherent 3'-5' proofreading exonuclease activity are available in some embodiments. Polymerases that substantially lack 3'-5' proofreading exonuclease activity are also available in some embodiments, for example, in some sequencing embodiments. The absence of exonuclease activity can be a wild-type characteristic or a characteristic conferred by a variant or engineered polymerase structure. For example, the exo-Klenow fragment is a mutant version of the Klenow fragment that lacks 3'-5' proofreading exonuclease activity. The Klenow fragment and its exo-variants can be used in the methods or compositions described herein.
[0393] D. Series
[0394] A tandem strand (or “tandem strand of a nucleic acid molecule” or “tandem strand” or other derived linguistic phrases) comprises multiple instances of a target sequence and multiple instances of at least one adaptor sequence. In any aspect and embodiment of this disclosure, it is specifically contemplated that references to “target sequence” of a tandem strand may further encompass sequences that are reversed and complementary to the target sequence, and references to “at least one adaptor sequence” may further encompass sequences that are reversed and complementary to the at least one adaptor sequence. The tandem strand may be prepared prior to or as part of the methods described herein.
[0395] A tandem strand can include multiple copies of tandemly linked sequence units (such as, for example, sequence units containing a target sequence and at least one adaptor sequence). For example, a tandem strand of a nucleic acid molecule can include at least 2, 10, 25, 100, or more sequence units. The number of sequence units in the tandem strand can be, for example, at most 100, 25, 10, or 2 sequence units. The number of sequence units in the tandem strand produced by RCA will be a function of the number of times the polymerase completes one revolution around the circular template during replication. The contents of each sequence unit produced by RCA will be the reverse complement of the contents of the replicated circular template. For example, the circular template can contain two adaptor sequences and a target sequence, wherein the first adaptor sequence is at the 3' of the target sequence and the second adaptor sequence is at the 5' of the target sequence.
[0396] As described above, an adaptor sequence is a known synthetic sequence of nucleic acids arranged at regular intervals. The adaptor sequence can act as a starting point for reading bases at multiple locations beyond the adaptor sequence-target sequence linking point, and optionally, bases can be read in both directions starting from the adaptor sequence. The adaptor sequence can have any of a variety of functions, including, but not limited to, providing a binding site complementary to a capture probe (e.g., a capture probe attached to a solid support), providing a primer binding site for replicating a circular template, providing a primer binding site for replicating a complement to a circular template, providing a tag associated with the target region (e.g., a tag indicating the origin of the target region or a tag for identifying errors introduced during target region amplification, etc.). The adaptor sequence can be engineered to include one or more of the following: 1) a length of about 10 to about 100 nucleotides, 2) features for linking to the 5' and / or 3' ends of the target sequence, 3) distinct and unique anchoring binding sites at the 5' and / or 3' ends of the adaptor sequence for sequencing adjacent target sequences, and 4) optional one or more restriction sites. The linker sequence or a portion thereof may be common to either a circular template group or a tandem group. Regardless of whether the linker sequences have a common sequence, the target region within a circular template group or a tandem group may have different sequences.
[0397] Therefore, when comparing sequence units between two or more tandems or between two or more circular templates, the sequence units may have common sequence regions (e.g., universal primer binding sites or universal capture probe binding sites) and / or the sequence units may have regions with different sequences (e.g., different target sequences). The length of the sequence units in the tandem or the length of the circular template can be selected to suit the specific application of the method described herein. For example, the length can be at least about 50, 100, 250, 500, 1000, or 1 x 10^6. 4 1 x 10 5One or more nucleotides. Alternatively or additionally, the length may not exceed 1 x 10^6 nucleotides. 5 1 x 10 4 1, 1000, 500, 250, 100, or 50 nucleotides. It should be understood that the above length ranges can be applied to target regions of sequence units or circular templates, but do not include adaptor sequences (e.g., common or universal adaptor sequences) that may also be present in sequence units. For example, length ranges can describe the size of genomic fragments or other nucleic acid fragments used to generate clusters or otherwise exist within clusters.
[0398] E. Nucleic acid clusters
[0399] Nucleic acid clusters may contain one or more strands of a nucleic acid tandem. For example, a cluster may contain only a single strand of a nucleic acid tandem. A single tandem strand can be generated by an RCA reaction. Alternatively, a nucleic acid cluster may contain a first strand (e.g., a sense strand) as a tandem and one or more second strands (e.g., antisense strands) complementary to the first strand. For example, one or more second strands can be generated by multiple substitution amplification (MDA) on a tandem template. Nucleic acid clusters may be prepared prior to or as part of the methods described herein.
[0400] A cluster may include one or more chains of tandem bodies. In some configurations, a cluster contains no more than one chain of tandem bodies. Alternatively, a cluster may include multiple chains of tandem bodies, for example, at least 2, 4, 10, 50, 100 or more meaningful chains including tandem bodies. Alternatively or additionally, the number of chains of tandem bodies in a cluster may be, for example, at most 100, 50, 10, 4, 2 or 1 meaningful chains including tandem bodies. In some embodiments, the number of tandem chains can be approximately, at least, at least about, at most, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8 The values can be 000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or any value or range between these values. Chains of tandem sequences within a particular cluster can have the same target sequence; for example, they can be meaningful chains of the same tandem sequence. Alternatively, a cluster can have multiple distinct chains of tandem sequences, and such clusters are non-clonal.
[0401] A cluster containing at least one sense chain of a tandem may further contain at least one antisense chain of the tandem. For example, a cluster containing one or more sense chains of a tandem may further contain multiple antisense chains. Multiple antisense chains may include at least 2, 4, 10, 50, 100 or more antisense chains of a particular tandem. Alternatively or additionally, the number of antisense chains in a cluster may be, for example, at most 100, 50, 10, 4, 2 or 1 antisense chains of a particular tandem. In some embodiments, the number of antisense chains can be approximately, at least, at least about, at most, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8 The values can be 000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or any value or range between these values. In a particular configuration, a cluster may contain a single meaningful concatenation chain (no more than one meaningful chain) and a single antisense concatenation chain (no more than one antisense chain). Alternatively, a cluster may contain at least one meaningful chain of a concatenation and multiple antisense chains of a concatenation. The number of antisense chains in a cluster may exceed the number of meaningful chains in the cluster. Alternatively, the number of meaningful chains in a cluster may exceed the number of antisense chains in the cluster. Note that the length of the antisense chain of a tandem complex need not be the same as the length of the sense chain. For example, the antisense chain can have more sequence units than the sense chain, or it can have fewer sequence units than the sense chain. The number of sequence units in the antisense chain can fall within the range described in this paper for the sense chain of a tandem complex. The antisense chain of a tandem complex need not have more than one sequence unit. In fact, the antisense chain does not need to have a complete sequence unit.
[0402] A cluster containing at least one sense strand of a tandem hybrid may further contain at least one antisense strand of the tandem hybrid that hybridizes with the sense strand via Watson-Crick base pairing. For example, the sense strand of a particular tandem hybridizes with at least 2, 4, 10, 50, 100 or more antisense strands of the particular tandem hybridizes. Alternatively or additionally, the sense strand of a particular tandem hybridizes with, for example, up to 100, 50, 10, 4, 2 or 1 antisense strand of the particular tandem hybridizes.
[0403] F. Staple molecule
[0404] According to the first aspect and embodiments described above and / or the second aspect and embodiments described below, the staple molecule comprises at least a first region and a second region, wherein the first region and the second region of the staple molecule each hybridize with different instances of at least one adaptor sequence of a tandem strand. In any of the aspects and embodiments of this disclosure, it is specifically contemplated that references to a “target sequence” of a tandem strand may further encompass sequences that are reversed and complementary to the target sequence, and references to “at least one adaptor sequence” may further encompass sequences that are reversed and complementary to the at least one adaptor sequence. Therefore, in some embodiments, the first region and / or the second region of the staple molecule will hybridize with at least one adaptor sequence and / or sequences that are reversed and complementary to the at least one adaptor sequence.
[0405] In some embodiments, the sequences of the first and second regions are identical. In other embodiments, the sequences of the first and second regions are partially conservative, while in still other embodiments, the sequences are discrete.
[0406] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows for extension. The extension can be mediated by a polymerase or a ligase.
[0407] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows for extension. In further embodiments, the 3' end of the staple molecule hybridizes within 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 10, 5, or fewer nucleotides from the 3' end of the target sequence. In an exemplary embodiment, the 3' end of the staple molecule hybridizes within 40 nucleotides from the 3' end of the target sequence.
[0408] In some embodiments, at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes to a primer binding site of at least one adaptor sequence.
[0409] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows the formation of a ternary complex.
[0410] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows the formation of the ternary complex. In a further embodiment, the 3' end that allows the formation of the ternary complex is reversibly terminated. Reversible termination can be performed using any reversible terminator. Exemplary reversible terminators, such as those in which the 3'-OH group is partially replaced by a 3'-ONH2 portion, are set forth in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference.
[0411] In some embodiments, neither the first nor the second region of the staple molecule contains a 3' end that allows the ternary complex to form or extend.
[0412] In some embodiments, the tandem mass includes a repeating sequence unit that includes a target sequence and at least one adaptor sequence. In a further embodiment, the at least one adaptor sequence within the sequence unit includes a first adaptor sequence of target sequence 3' and a second adaptor sequence of target sequence 5'. In yet another further embodiment, a first region of the staple molecule hybridizes with the first adaptor sequence and a second region of the staple molecule hybridizes with the second adaptor sequence.
[0413] In some embodiments, the tandem mass includes a repeating sequence unit that includes a target sequence and at least one adaptor sequence. In a further embodiment, the at least one adaptor sequence within the sequence unit includes a first adaptor sequence of the target sequence 3' and a second adaptor sequence of the target sequence 5'. In yet another further embodiment, at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes with a 3' adaptor.
[0414] In some embodiments, the first and second regions of the staple molecule are of equal length. Alternatively, the first region of the staple molecule may be longer than the second region. In other embodiments, the second region of the staple molecule may be longer than the first region. In some embodiments, the length of the first region of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, the length of the second region of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides.
[0415] In some embodiments, the length of each of the first and second regions of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In an exemplary embodiment, the length of each of the first and second regions of the staple molecule is at least 10 nucleotides.
[0416] In some embodiments, the first and second regions of the staple molecule each hybridize with different instances of at least one adaptor sequence spaced at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or more nucleotides apart. In an exemplary embodiment, the first and second regions of the staple molecule each hybridize with different instances of at least one adaptor sequence spaced at least 100 nucleotides apart.
[0417] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, nucleotide incorporation-preventing, or ternary complex formation-preventing cap. The blocking component can be any blocking part. A blocking part is a portion of a nucleotide that inhibits or prevents the 3' oxygen of a nucleotide from forming a covalent bond with the next correct nucleotide during nucleic acid polymerization. The blocking part of a “reversible termination” nucleotide may be removed from a nucleotide analog or otherwise modified to allow the 3' oxygen of the nucleotide to covalently attach to the next correct nucleotide. Such a blocking part is referred to herein as a “reversible termination part.” The blocking part does not need to prevent or inhibit the formation of a ternary complex at the 3' end of the nucleic acid to which the blocking part is attached. The cap can be any capping part. The capping part can have a positive or negative charge that inhibits or prevents ternary complex formation. The capping part may include a ligand that binds to a receptor to inhibit or prevent ternary complex formation, such as biotin (or its analogues) that binds to streptavidin (or its analogues), an epitope that binds to an antibody (or its functional fragment), a carbohydrate that binds to a lectin, etc. Further examples of the capped portion are described in U.S. Patent Application Publication No. 2020 / 0032322A1 or Turcatti et al. Nucl. Acids. Res. 36(4) e25 (2008), each of which is incorporated herein by reference.
[0418] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a nucleotide incorporation blocker, or a ternary complex formation cap. In a further embodiment, the 3' end of the staple molecule includes a nucleotide incorporation blocker.
[0419] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a blocking agent to prevent nucleotide incorporation, or a cap to prevent ternary complex formation. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation.
[0420] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, nucleotide incorporation blocking agent, or ternary complex formation prevention cap. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation. In yet another embodiment, the 3' end of the staple molecule is capped by a portion preventing ternary complex formation.
[0421] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a blocking agent to prevent nucleotide incorporation, or a cap to prevent ternary complex formation. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation. In yet another embodiment, the 3' end of the staple molecule includes a mismatch and a blocking agent, optionally wherein the blocking agent is a 3' phosphate blocking agent.
[0422] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. Any suitable spacer known in the art can be used. In some instances, the spacer may contain a polynucleotide sequence. Alternatively, the spacer may contain a non-nucleotide polymer linker, such as, for example, a branched polyelectrolyte species. Non-limiting examples of branched polyelectrolyte species include polyethylene glycol (PEG) and dendritic macromolecules.
[0423] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, at least some of the one or more staple molecules do not contain spacers.
[0424] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise polynucleotide sequences. The length of the polynucleotide sequence is at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides.
[0425] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise polynucleotide sequences. In yet another embodiment, the length of the polynucleotide sequence is variable.
[0426] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacer comprises a polynucleotide sequence. In yet another embodiment, the polynucleotide sequence comprises a double-stranded DNA sequence.
[0427] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by a spacer. In a further embodiment, the spacer comprises a polynucleotide sequence. In yet another embodiment, the polynucleotide sequence comprises a double-stranded DNA sequence. In still a further embodiment, the first and second regions each include a 3' end that hybridizes to the tandem strand.
[0428] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers have variable lengths on different staple molecules.
[0429] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise non-nucleotide polymer connectors.
[0430] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise polyethylene glycol (PEG). PEG may comprise PEG with an average molecular weight of about 200 Daltons (e.g., PEG-200) to about 8000 Daltons (e.g., PEG-8000).
[0431] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. Species of dendritic macromolecules that can be used in the methods, compositions, and systems disclosed herein include, but are not limited to, branched polyamines comprising protonated structures that interact with the negatively charged DNA backbone to form complexes. It should be understood that, in the embodiments described herein, the adaptor element can still be used in conjunction with the branched polyamine for alternative purposes or to provide improved binding properties of the dendritic macromolecule species to nucleic acids. Dendritic macromolecule species may comprise controlled terminal surface chemistry having one or more functional groups, including but not limited to amine, carboxyl, and hydroxyl groups. Dendritic macromolecule species can be obtained as generation 0 (G0) through generation 10 (G10), with the number of branches in each generation being twice that of the previous generation. Thus, G0 = 4 branches, G1 = 8 branches, and so on. In some embodiments, the branched polyelectrolyte is a poly(amidoamine) dendritic macromolecule species (also known as PAMAM), such as the G2 PAMAM dendritic macromolecule having 16 branches and an amine (NH2) terminal surface chemistry. Non-limiting examples of branched polyelectrolytes also include G4 (64 branches with amine terminal groups) and G5 (128 branches with amine terminal groups) PAMAM dendritic macromolecule species.
[0432] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer connectors. In yet another embodiment, the nonnucleotide polymer connectors comprise dendritic macromolecules. In even further embodiments, the dendritic macromolecules comprise polyamidoamine (PAMAM). Any species of PAMAM described above, both those specifically described and those implied by the entirety of this disclosure, may be used.
[0433] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. In even further embodiments, the staple molecule hybridizes with 3, 4, 5, 6, 7, 8, 9, 10, or more instances of at least one adaptor. In an exemplary embodiment, the staple molecule hybridizes with 3 or more instances of at least one adaptor. In another exemplary embodiment, the staple molecule hybridizes with 10 or more instances of at least one adaptor sequence.
[0434] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. In even further embodiments, the staple molecule hybridizes with three or more instances of at least one adaptor. In specific embodiments of even further embodiments, a majority of the staple molecule hybridizes with adaptor sequences flanking at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200 or more different target sequences. In an exemplary embodiment, a majority of the staple molecule hybridizes with adaptor sequences flanking at least 100 different target sequences.
[0435] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer connectors. In yet another embodiment, the nonnucleotide polymer connectors comprise dendritic macromolecules. In even further embodiments, neither the first nor second region of the staple molecule contains a 3' end that allows the formation or extension of the ternary complex.
[0436] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers are coupled to the 3' end of the first region.
[0437] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers are coupled to the 3' end of the first region. In yet another further embodiment, the spacers are coupled to the 5' end of the second region.
[0438] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by a spacer. In a further embodiment, the spacer is coupled to the 3' end of the first region. In yet another further embodiment, the spacer is coupled to the 3' end of the second region and wherein the staple molecule cannot act as a primer.
[0439] In some embodiments, multiple staple molecules hybridize with the same tandem strand.
[0440] G. Solid supports and arrays
[0441] Tandems, nucleic acid clusters, and / or nucleic acids can be attached to the surface of a solid support. The solid support can be made from any of a variety of materials used in analytical biochemistry. Suitable materials may include, for example, glass, polymer materials, silicon, quartz (fused silica), Borofloat glass, silica, silica-based materials, carbon, metals, optical fibers or fiber bundles, sapphire, or plastic materials. Materials can be selected based on the properties desired for a particular application. For example, materials that are transparent to a desired radiation wavelength may be used in analytical techniques that utilize radiation at that wavelength. Conversely, it may be desirable to select materials that do not transmit radiation at a particular wavelength (e.g., opaque, absorptive, or reflective materials). Wavelength regions that may or may not transmit through a particular material include, for example, UV, VIS (e.g., red, yellow, green, or blue), or IR. Other properties of materials that can be utilized include inertness or reactivity to certain reagents used in downstream processes (such as those described herein), ease of handling, or low manufacturing cost.
[0442] A particularly useful solid support is particles, such as beads or microspheres. Populations of beads can be used to attach nucleic acid populations. In some embodiments, it may be useful to use a configuration where each bead has a single target sequence. Individual beads may have a single nucleic acid molecule with the target sequence, or alternatively, individual beads may have multiple nucleic acid molecules, each of which has the target sequence. In some configurations, beads may be attached to a nucleic acid tandem, where multiple copies of the target sequence are present in a single nucleic acid molecule. Beads in a population may have different target sequences from one another. When compared to each other, beads in a population may have a common nucleic acid sequence. For example, a population of beads may be attached to universal primers such that the same primer sequence is present on multiple beads in the population.
[0443] The composition of the beads can vary depending on, for example, the format, chemistry, and / or method of attachment to be used. Exemplary bead compositions include a solid support, optionally comprising chemical functional groups for protein and nucleic acid capture methods. Such compositions include, for example, plastics, ceramics, glass, polystyrene, melamine, methylstyrene, acrylic polymers, paramagnetic materials, thorium oxide sol, carbon graphite, titanium dioxide, latex, or cross-linked dextran such as Sepharose™, cellulose, nylon, cross-linked micelles, and Teflon™, as well as other materials described in the “Microsphere Detection Guide” from Bangs Laboratories, Fishers Ind. (incorporated herein by reference).
[0444] The geometry of the particles (such as beads or microspheres) can also correspond to a variety of different forms and shapes. For example, particles can be symmetrical (e.g., spherical or cylindrical) or irregular (e.g., controlled-pore glass). Additionally, particles can be porous, thereby increasing the surface area available for trapping ternary composites or their components. Exemplary sizes of beads used herein can range from nanometers to millimeters or from about 10 nm to about 1 mm.
[0445] Exemplary bead-based arrays that may be used include, but are not limited to, BeadChip™ arrays or arrays such as those described in U.S. Patent Nos. 6,266,459; 6,355,431; 6,770,441; 6,859,570; or 7,622,294; or PCT Publication No. WO 00 / 63437, each of which is incorporated herein by reference. Beads may be located at discrete locations on a solid support, such as holes, where each location accommodates a single bead. Alternatively, the discrete locations where beads reside may each comprise multiple beads, as described, for example, in U.S. Patent Application Publication Nos. 2004 / 0263923 A1, 2004 / 0233485 A1, 2004 / 0132205 A1, or 2004 / 0125424 A1, each of which is incorporated herein by reference.
[0446] In some embodiments of the sequencing methods using stapled primers described below, the beads may be arranged or otherwise spatially differentiated. Exemplary bead-based arrays that may be used include, but are not limited to, the BeadChip™ arrays available from Illumina, Inc. (San Diego, CA) or arrays such as those described in U.S. Patent Nos. 6,266,459; 6,355,431; 6,770,441; 6,859,570; or 7,622,294; or PCT Publication No. WO 00 / 63437, each of which is incorporated herein by reference. The beads may be located at discrete locations on a solid support, such as holes, with each location accommodating a single bead. Alternatively, the discrete locations where the beads reside may each comprise multiple beads, as described, for example, in U.S. Patent Application Publication Nos. 2004 / 0263923 A1, 2004 / 0233485 A1, 2004 / 0132205 A1, or 2004 / 0125424A1, each of which is incorporated herein by reference.
[0447] In some embodiments of the sequencing method using stapled primers described below, the method can be performed in multiplex format, thereby allowing parallel detection of multiple different types of nucleic acids. Other types of arrays can be used instead of bead arrays, including, for example, those described in further detail below. Although different types of nucleic acids can also be processed sequentially using one or more steps of the sequencing method using stapled primers, parallel processing can provide cost savings, time savings, and consistency of conditions. The arrays or methods of this disclosure can be configured to include at least 2, 10, 100, or 1 x 10^6 primers. 3 Seed, 1 x 10 4 Seed, 1 x 10 5 Seed, 1 x 10 6 Seed, 1 x 10 9 One or more different nucleic acids. Alternatively or additionally, the arrays or methods of this disclosure can be configured to include up to 1 x 103 9 Seed, 1 x 10 6 Seed, 1 x 10 5 Seed, 1 x 10 4 Seed, 1 x 10 3 There can be one, 100, 10, 2, or fewer different nucleic acids. Nucleic acids can attach to different sites in the array. Therefore, the number of sites in the array can be within the range exemplified here for different nucleic acids. Furthermore, the various reagents or products described herein (e.g., primer-template nucleic acid hybrids or stabilized ternary complexes) can be multiplexed to have different types or species within these ranges.
[0448] Further examples of commercially available arrays that can be used with the compositional methods described herein include arrays prepared by photolithographic synthesis of nucleic acids, such as the Affymetrix GeneChip™ array. According to some embodiments, dot arrays can also be used to attach pre-synthesized nucleic acids to array sites. An exemplary dot array is the CodeLink™ array, available from Amersham Biosciences. Another available array is one fabricated using inkjet printing methods, such as SurePrint™ Technology, available from Agilent Technologies. The methods used to attach nucleic acid probes to these arrays can be modified to attach nucleic acid primers for amplification (e.g., via RCA) and / or sequencing of target nucleic acids hybridized to the primers.
[0449] Other available arrays include those for nucleic acid sequencing applications. For example, methods and compositions for attaching amplicon to genomic fragments (often referred to as clusters) to form arrays can be particularly useful. Examples are described in Bentley et al., Nature 456:53-59 (2008); PCT Publication Nos. WO 91 / 06678; WO 04 / 018497 or WO 07 / 123744; U.S. Patent Nos. 7,057,026; 7,211,414; 7,315,019; 7,329,492 or 7,405,281; or U.S. Patent Application Publication No. 2008 / 0108082 A1, each of which is incorporated herein by reference.
[0450] Nucleic acids, tandem nucleic acids, and / or clusters of nucleic acids can be attached to a solid support (e.g., sites of an array) via covalent or non-covalent bonds. For example, the solid support can be covalently or non-covalently attached to the 5' end of a tandem nucleic acid. This configuration may occur, for example, when a tandem nucleic acid has been generated by an RCA (Reactive Carbon Acetate) performed by a primer that extends to or near the 5' end of the solid support. The attachment of nucleic acids to a solid support can be mediated by any of a variety of surface chemistry, such as the reaction of a carboxylic acid ester or succinimidyl ester moiety on the solid support with an amine-modified nucleic acid, the reaction of an alkylating agent (e.g., iodoacetamide or maleimide) on the solid support with a thiol-modified nucleic acid, the reaction of an epoxysilane or isothiocyanate-modified solid support with an amine-modified nucleic acid, the reaction of an aminophenyl or aminopropyl-modified solid support with a succinylated nucleic acid, the reaction of an aldehyde or epoxide-modified solid support with an acylhydrazine-modified nucleic acid, or the reaction of a thiol-modified solid support with a thiol-modified nucleic acid. The members of the aforementioned reaction pairs can be switched depending on whether they are present on a solid support or on nucleic acids. Click chemistry can be used to attach nucleic acids to a solid support. Exemplary reagents and methods for use in click chemistry are set forth in U.S. Patent Nos. 6,737,236; 7,375,234; 7,427,678 and 7,763,736, each of which is incorporated herein by reference.
[0451] Solid supports may include two (or more, such as three, four, five, six, seven, eight, nine, ten or more) types or groups of primers. Two or more types of primers may, for example, serve as multiple capture primers and multiple amplification primers. Alternatively or additionally, the primer types may include multiple capture primers and multiple sequencing primers for sequencing, for example, a first strand generated by extending capture primers. The density of one type or group of primers (e.g., capture primers) on the solid support may be higher than the density of another type or group of primers (e.g., amplification primers) on the solid support. The density of one type or group of primers on the solid support may be the same as the density of another type or group of primers on the solid support. The density of a certain type or group of primers (or all primers) may vary. The density of a certain type or group of primers (or all primers) on the solid support may be, about, at least, at least about, at most, or at most about 1 x 10⁻⁶. 10 1, 2 x 10 10 1, 3 x 10 10 1, 4 x 10 10 5 x 10 10 6 x 10 10 7 x 10 10 8 x 10 10 9 x 10 10 1 x 10 11 1, 2 x 10 11 1, 3 x 10 11 1, 4 x 10 11 5 x 10 11 6 x 10 11 7 x 10 11 8 x 10 11 9 x 10 11 1 x 10 12 1, 2 x 10 12 1, 3 x 10 12 1 piece, 4 x 10 12 5 x 10 12 6 x 10 12 7 x 10 12 8 x 10 12 9 x 10 12 1 x 10 13 1, 2 x 10 13 1, 3 x 10 13 1, 4 x 10 13 5 x 10 13 6 x 1013 7 x 10 13 8 x 10 13 9 x 10 13 1 x 10 14 1, 2 x 10 14 1, 3 x 10 14 1, 4 x 10 14 5 x 10 14 6 x 10 14 7 x 10 14 8 x 10 14 9 x 10 14 1 x 10 15 1, 2 x 10 15 1, 3 x 10 15 1, 4 x 10 15 5 x 10 15 6 x 10 15 7 x 10 15 8 x 10 15 9 x 10 15 1 x 10 16 1, 2 x 10 16 1, 3 x 10 16 1, 4 x 10 16 5 x 10 16 6 x 10 16 7 x 10 16 8 x 10 16 9 x 10 16 primers / m 2 , or a value or range between any two of these values.
[0452] This paper considers various separation distances or average separation distances between two adjacent primers of the same type or population (or two different types or populations). The spacing or average spacing between two adjacent primers of the same type or population (or two different types or populations) can be, about, at least, at least about, at most, or at most about 10 nm, 11 nm, 12 nm, 13 nm, 14 nm, 15 nm, 16 nm, 17 nm, 18 nm, 19 nm, 20 nm, 21 nm, 22 nm, 23 nm, 24 nm, 25 nm, 26 nm, 27 nm, 28 nm, 29 nm, 30 nm, 31 nm, 32 nm, 33 nm, 34 nm, 35 nm, 36 nm, 37 nm, 38 nm, 39 nm, 40 nm, 41 nm, 42 nm, 43 nm, 44 nm, 45 nm, 46 nm, 47 nm, 48 nm, 49 nm, 50 nm, 51 nm, 52 nm, 53 nm, 54 nm, 55 nm, 56 nm, 57 nm, 58 nm, 59 nm, 60 nm, 61 nm, or 10 nm. nm, 62 nm, 63 nm, 64nm, 65 nm, 66 nm, 67 nm, 68 nm, 69 nm, 70 nm, 71 nm, 72 nm, 73 nm, 74 nm, 75 nm, 76 nm, 77nm, 78 nm, 79 nm, 80 nm, 81 nm, 82 nm, 83 nm, 84 nm, 85 nm, 86 nm, 87 nm, 88 nm, 89 nm, 90nm, 91 nm, 92 nm, 93 nm, 94 nm, 95 nm, 96 nm, 97 nm, 98 nm, 99 nm, 100 nm, 110 nm, 120nm, 130 nm, 140 nm, 150 nm, 160 nm, 170 nm, 180 nm, 190 nm, 200 nm, 210 nm, 220 nm, 230nm, 240 nm, 250 nm, 260 nm, 270 nm, 280 nm, 290 nm, 300 nm, 310 nm, 320 nm, 330 nm, 340nm, 350 nm, 360 nm, 370 nm, 380 nm, 390 nm, 400 nm, 410 nm, 420 nm, 430 nm, 440 nm, 450nm, 460 nm, 470 nm, 480 nm, 490 nm, 500 nm, 510 nm, 520 nm, 530 nm, 540 nm, 550 nm, 560nm, 570 nm, 580 nm, 590 nm, 600 nm, 610nm, 620 nm, 630 nm, 640 nm, 650 nm, 660 nm, 670 nm, 680 nm, 690 nm, 700 nm, 710 nm, 720 nm, 730 nm, 740 nm, 750 nm, 760 nm, 770 nm, 780 nm, 790 nm, 800 nm, 810 nm, 820 nm, 830 nm, 840 nm, 850 nm, 860 nm, 870 nm, 880 nm, 890 nm, 900 nm, 910 nm, 920 nm, 930 nm, 940 nm, 950 nm, 960 nm, 970 nm, 980 nm, 990 nm, 1000 nm, or a value or range between any two of these values.
[0453] This disclosure considers various ratios of the number of primers of one type or population to the number of primers of another type or population. The ratio of the number of primers of one type or population to the number of primers of another type or population can be approximately, at least, at least about, at most, or at most about 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1:73, 1:72, 1:71, 1:70, 1:69, 1:68, 1:67, 1:66, 1:65, 1:64, 1:6 3, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:1 5, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:187:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, or a value or range between any two of these values.
[0454] The average distance between the positions on the solid support to which two adjacent or closest capture primers are attached can be greater than (or less than, or equal to) the length of one of the two capture primers, the length of the two capture primers, the average length of the two capture primers, or the total length of the two capture primers (or 0.1 x, 0.2 x, 0.3 x, 0.4 x, 0.5 x, 0.6 x, 0.7 x, 0.8 x, 0.9 x of the length). The average distance between the positions on the solid support to which two adjacent or closest amplification primers of a plurality of amplification primers are attached may be greater than (or less than, or equal to) the length of one of the two amplification primers, the length of the two amplification primers, the average length of the two amplification primers, or the total length of the two capture primers (or 0.1 x, 0.2 x, 0.3 x, 0.4 x, 0.5 x, 0.6 x, 0.7 x, 0.8 x, 0.9 x, 1.0 x, 1.1 x, 1.2 x, 1.3 x, 1.4 x, 1.5 x, 1.6 x, 1.7 x, 1.8 x, 1.9 x, 2 x, 3 x, 4 x, 5 x, 6 x, 7 x, 8 x, 9 x, 10 x).
[0455] H. Flow cell
[0456] Flow cells facilitate fluid manipulation by introducing and withdrawing a solution into a fluid chamber that contacts an analyte bound to a support. Flow cells also provide the ability to detect fluidly manipulated components. Detectors can be positioned to detect signals from the solid support, such as signals from tags recruited to the solid support during the sequencing process. Exemplary flow cells that can be used are described, for example, in U.S. Patent Application Publication No. 2010 / 0111768 A1, WO 05 / 065814, or U.S. Patent Application Publication No. 2012 / 0270305 A1, each of which is incorporated herein by reference.
[0457] In certain configurations, nucleic acids, tandem molecules, and / or nucleic acid clusters may be attached to the surface of the flow cell and / or to a solid support within the flow cell. Tandem molecules and / or nucleic acid clusters may comprise a first strand (sense strand) and / or multiple second strands (antisense strands). In some configurations, nucleic acids, tandem molecules, and / or nucleic acid clusters are attached to the surface of the flow cell and / or to the solid support within the flow cell via covalent attachment. In some configurations, the first strand is covalently attached to the surface of the flow cell and / or to the solid support within the flow cell, and multiple second strands are not covalently attached. For example, multiple second strands may remain on the surface of the flow cell or to the solid support within the flow cell due to Watson-Crick base pairing with the first strand.
[0458] I. via a stable series
[0459] In a second aspect of the invention, a stabilized tandem composition is provided, comprising: (a) a tandem comprising a plurality of instances of a target sequence and a plurality of instances of at least one adaptor sequence; and (b) one or more staple molecules; wherein each staple molecule comprises at least a first region and a second region; and wherein the first region and the second region of the staple molecule each hybridize with different instances of at least one adaptor sequence. In any aspect and embodiment of this disclosure, it is specifically contemplated that references to “target sequence” of a tandem may further cover sequences that are reversed and complementary to the target sequence, and references to “at least one adaptor sequence” may further cover sequences that are reversed and complementary to the at least one adaptor sequence. Therefore, in some embodiments, the first region and / or the second region of the staple molecule will hybridize with at least one adaptor sequence and / or sequences that are reversed and complementary to the at least one adaptor sequence.
[0460] In some embodiments, the tandem polymerase is the product of RCA performed by a primer hybridized to a circular nucleic acid template using a strand substitution polymerase, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. The circular nucleic acid template (such as, for example, a circular nucleic acid template comprising a target sequence and at least one adaptor sequence) can be single-stranded or double-stranded. One or both strands of the double-stranded nucleic acid may lack a 3' and a 5' end. One strand of the double-stranded nucleic acid may have a gap (the absence of at least one nucleotide monomer relative to the other strand) or a cleavage (the absence of a phosphodiester bond between two nucleotide monomers), provided that the other strand is circular. Any of a variety of polymerases may be used in the methods or compositions described herein. Non-limiting examples of polymerases that may be used include naturally occurring polymerases and modified versions thereof, including but not limited to mutants, recombinants, fusions, genetically modified, chemically modified, synthetics, analogs, etc.
[0461] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Genetic Alternating Current Acute Restriction) of primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the circular nucleic acid template is the product of circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. Circularization of the linear nucleic acid template can be performed using any method known in the art. In some embodiments, a ligase is used to circularize the linear nucleic acid template. Circularization of the linear nucleic acid template can be performed using any suitable ligase, such as, for example, T4 DNA ligase.
[0462] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Genetic Alternative) with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the circular nucleic acid template is the product of circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template comprises a first adaptor sequence at 3' of the target sequence and a second adaptor sequence at 5' of the target sequence.
[0463] In some embodiments, the tandem is the product of RCA by a primer hybridized to a circular nucleic acid template using a strand substitution polymerase, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the circular nucleic acid template is the product of circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template comprises a first adaptor sequence at 3' of the target sequence and a second adaptor sequence at 5' of the target sequence. In even further embodiments, the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with a splint oligonucleotide. As described above, a "splint oligonucleotide" is an oligonucleotide that, when hybridized with other polynucleotides (such as, for example, a first adaptor sequence, or a second adaptor sequence, or a (nucleic acid) template, acts as a "splice" to position the polynucleotides adjacent to each other so that they can be ligated together. In some embodiments, the splint oligonucleotide is DNA or RNA. The splint oligonucleotide may include a nucleotide sequence partially complementary to the nucleotide sequence of two or more different oligonucleotides. In some embodiments, the splint oligonucleotide facilitates the ligation of a "donor" oligonucleotide and a "recipient" oligonucleotide. Generally, RNA ligase, DNA ligase, or another type of ligase is used to join two nucleotide sequences together. In some embodiments, the length of the splice oligonucleotide is between 10 and 50 nucleotides, for example, between 10 and 45 nucleotides, 10 and 40 nucleotides, 10 and 35 nucleotides, 10 and 30 nucleotides, 10 and 25 nucleotides, or 10 and 20 nucleotides. In some embodiments, the length of the splice oligonucleotide is between 15 and 50 nucleotides, 15 and 45 nucleotides, 15 and 40 nucleotides, 15 and 35 nucleotides, 15 and 30 nucleotides, or 15 and 25 nucleotides. Alternatively, splice oligonucleotides are not required for end-joining, for example, when using CircLigase™ (Epicenter, Madison WI) or other enzymes capable of splice-free end-joining of nucleic acids.
[0464] In some embodiments, the tandem polymerase is the product of RCA with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the circular nucleic acid template is the product of circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template comprises a first adaptor sequence at 3' of the target sequence and a second adaptor sequence at 5' of the target sequence. In even further embodiments, the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with a splice oligonucleotide. In a particular embodiment of even further embodiments, the splice oligonucleotide is a primer hybridized to a circularized nucleic acid.
[0465] In some embodiments, the tandem polymerase is the product of RCA with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the circular nucleic acid template is the product of circularizing a linear nucleic acid template comprising a target sequence and at least one adaptor sequence. In yet another embodiment, the linear nucleic acid template comprises a first adaptor sequence at 3' of the target sequence and a second adaptor sequence at 5' of the target sequence. In even further embodiments, the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with a splint oligonucleotide. In a particular embodiment of even further embodiments, the splint oligonucleotide is removed prior to RCA. Alternatively, the splint oligonucleotide may be removed during or after RCA.
[0466] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Genetic Alternating Current) of primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In further embodiments, the primers are immobilized on a surface during RCA. Suitable surfaces include, but are not limited to, structured surfaces, planar substrates, hydrogels, nanopore arrays, microparticles, nanoparticles, flow cell surfaces, surfaces of solid supports, or surfaces of solid supports within a flow cell. The surface may be planar or curved. As discussed in further detail above, the solid support may be made of any of a variety of materials used for analytical biochemistry. Suitable materials may include, for example, glass, polymeric materials, silicon, quartz (fused silica), borofloat glass, silica, silica-based materials, carbon, metals, optical fibers or fiber bundles, sapphire, or plastic materials. Materials may be selected based on properties desired for a particular application. For example, materials transparent to a desired wavelength of radiation may be used in analytical techniques that utilize radiation at that wavelength. Conversely, it may be desirable to select materials that do not transmit radiation at a particular wavelength (e.g., opaque, absorptive, or reflective materials). Wavelength regions that may or may not pass through a particular material include, for example, UV, VIS (e.g., red, yellow, green, or blue), or IR. Other properties of materials that can be utilized include inertness or reactivity to certain reagents used in downstream processes (such as those described herein), ease of handling, or low manufacturing cost.
[0467] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Chain Alignment) of primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA.
[0468] In some embodiments, the tandem polymerase is the product of RCA (Reactive Carbon Alginate) of primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the tandem polymerase is attached to a surface after hybridization with one or more staple molecules.
[0469] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Genetic Alternating Current Acetate) with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the tandem polymerase is attached to a surface after hybridization with one or more staple molecules. In even further embodiments, the surface is a structured surface.
[0470] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Chain Alignment) of primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the tandem polymerase is attached to a surface before hybridizing with one or more staple molecules.
[0471] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Genetic Alternating Current Acetate) with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the primers are in solution during RCA. In yet another embodiment, the tandem polymerase is attached to a surface before hybridizing with one or more staple molecules. In even further embodiments, the surface is a structured surface.
[0472] In some embodiments, the tandem strand is the product of RCA (Recombinant Genetic Alternative) using a strand substitution polymerase with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the tandem strand is a sense strand, and the composition further comprises multiple antisense strands generated by amplification of the sense strand.
[0473] In some embodiments, the tandem strand is the product of RCA (Recombinant Genetic Alternative) using a strand substitution polymerase with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the tandem strand is a sense strand, and the composition further comprises multiple antisense strands generated by the amplification of the sense strand. In yet another further embodiment, at least some of one or more staple molecules hybridize with adaptor sequences of multiple antisense strands.
[0474] In some embodiments, the tandem strand is the product of RCA (Recombinant Genetic Alternative) using a strand substitution polymerase with primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the tandem strand is a sense strand, and the composition further comprises multiple antisense strands generated by amplification of the sense strand. In yet another further embodiment, one or more staple molecules hybridize with the sense strand.
[0475] In some embodiments, the tandem strand is the product of RCA by a strand substitution polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, the tandem strand is a sense strand, and the composition further comprises multiple antisense strands generated by the amplification of the sense strand. In yet another embodiment, one or more staple molecules hybridize to the sense strand. In even further embodiments, at least some of the one or more staple molecules are one or more staple primers for amplifying the sense strand to generate multiple antisense strands.
[0476] In some embodiments, the tandem polymerase is the product of RCA (Recombinant Chain Alignment) with primers hybridizing to a circular nucleic acid template, wherein the circular nucleic acid template comprises a target sequence and at least one adaptor sequence. In a further embodiment, one or more staple molecules hybridize with the tandem polymerase during RCA.
[0477] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site.
[0478] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site. In a further embodiment, each instance of at least one adaptor sequence further includes a tag region, optionally wherein the tag region is a sample index.
[0479] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site. In a further embodiment, each instance of at least one adaptor sequence further includes a splint binding site.
[0480] In some embodiments, each instance of at least one adaptor sequence includes a primer binding site. In further embodiments, each instance of at least one adaptor sequence further includes a variable region, optionally wherein the variable region is a unique molecular identifier (UMI). Generally, a UMI is a nucleotide sequence applied to or identified in a polynucleotide that can be used to distinguish individual nucleic acid molecules present in an initial reaction from one another. In some cases, a UMI may contain about 5 to about 20 nucleotides. Alternatively, a UMI may contain fewer than about 5 or more than 20 nucleotides. A UMI can be a unique sequence that varies among individual nucleic acid molecules. In some cases, a UMI can be a random sequence. In some cases, a UMI can be a predetermined sequence. In a sequencing reaction, a UMI can be sequenced along with the nucleic acid molecules associated with it to determine whether the read sequence is a sequence of one source nucleic acid molecule or a sequence of another source nucleic acid molecule. The term “UMI” is used herein to refer to both the sequence information of the polynucleotide and the physical polynucleotide containing that sequence information. Other examples of UMIs and their uses are provided, for example, in US 2016 / 0319345 A1, which is incorporated herein by reference.
[0481] In some embodiments, the tandem strand includes a sense strand that hybridizes with multiple antisense strands.
[0482] In some embodiments, the tandem strand comprises a sense strand that hybridizes with multiple antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands.
[0483] In some embodiments, the tandem strand comprises a sense strand that hybridizes with multiple antisense strands. In a further embodiment, the sense strand does not contain uracil and the multiple antisense strands do contain uracil.
[0484] In some embodiments, all instances of the tandem in vivo target sequence are identical.
[0485] In some embodiments, instances of the target sequence may include sense sequences or antisense sequences.
[0486] In some embodiments, a series of elements are provided that are fixed to a surface in a flow cell.
[0487] In some embodiments, a series of surfaces fixed to a flow cell are provided. In a further embodiment, the surfaces are continuous surfaces.
[0488] In some embodiments, a series of structures are provided that are fixed to a surface in the flow cell. In a further embodiment, the series of structures are fixed to bonding sites on a structured surface of the flow cell.
[0489] In some embodiments, the series bodies are generated in the solution and subsequently attached to the surface of the flow cell.
[0490] According to the first and / or second aspects and the above embodiments, the staple molecule includes at least a first region and a second region, wherein the first region and the second region of the staple molecule each hybridize with different instances of at least one linker sequence of the tandem.
[0491] In some embodiments, the sequences of the first region and the second region are the same.
[0492] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows for extension. The extension can be mediated by a polymerase or a ligase.
[0493] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows for extension. In further embodiments, the 3' end of the staple molecule hybridizes within 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 10, 5, or fewer nucleotides from the 3' end of the target sequence. In an exemplary embodiment, the 3' end of the staple molecule hybridizes within 40 nucleotides from the 3' end of the target sequence.
[0494] In some embodiments, at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes to a primer binding site of at least one adaptor sequence.
[0495] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows the formation of a ternary complex.
[0496] In some embodiments, at least one of the first and second regions of the staple molecule includes a 3' end that allows the formation of the ternary complex. In a further embodiment, the 3' end that allows the formation of the ternary complex is reversibly terminated. Reversible termination can be performed using any reversible terminator. Exemplary reversible terminators, such as those in which the 3'-OH group is partially replaced by a 3'-ONH2 portion, are set forth in U.S. Patent Nos. 7,427,673; 7,414,116; 7,057,026; 7,544,794 or 8,034,923; or PCT Publications WO 91 / 06678 or WO 07 / 123744, each of which is incorporated herein by reference.
[0497] In some embodiments, neither the first nor the second region of the staple molecule contains a 3' end that allows the ternary complex to form or extend.
[0498] In some embodiments, the tandem mass includes a repeating sequence unit that includes a target sequence and at least one adapter sequence.
[0499] In some embodiments, the tandem unit includes a repeating sequence unit that includes a target sequence and at least one interleave sequence. In a further embodiment, the at least one interleave sequence within the sequence unit includes a first interleave sequence of target sequence 3' and a second interleave sequence of target sequence 5'.
[0500] In some embodiments, the tandem mass includes a repeating sequence unit that includes a target sequence and at least one adaptor sequence. In a further embodiment, the at least one adaptor sequence within the sequence unit includes a first adaptor sequence of target sequence 3' and a second adaptor sequence of target sequence 5'. In yet another further embodiment, a first region of the staple molecule hybridizes with the first adaptor sequence and a second region of the staple molecule hybridizes with the second adaptor sequence.
[0501] In some embodiments, the tandem mass includes a repeating sequence unit that includes a target sequence and at least one adaptor sequence. In a further embodiment, the at least one adaptor sequence within the sequence unit includes a first adaptor sequence of the target sequence 3' and a second adaptor sequence of the target sequence 5'. In yet another further embodiment, at least one of the first and second regions of the staple molecule includes the 3' end of the staple molecule and hybridizes with a 3' adaptor.
[0502] In some embodiments, the first and second regions of the staple molecule are of equal length. Alternatively, the first region of the staple molecule may be longer than the second region. In other embodiments, the second region of the staple molecule may be longer than the first region. In some embodiments, the length of the first region of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, the length of the second region of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides.
[0503] In some embodiments, the length of each of the first and second regions of the staple molecule is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In an exemplary embodiment, the length of each of the first and second regions of the staple molecule is at least 10 nucleotides.
[0504] In some embodiments, the first and second regions of the staple molecule each hybridize with different instances of at least one adaptor sequence spaced at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or more nucleotides apart. In an exemplary embodiment, the first and second regions of the staple molecule each hybridize with different instances of at least one adaptor sequence spaced at least 100 nucleotides apart.
[0505] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, nucleotide incorporation-preventing, or ternary complex formation-preventing cap. The blocking component can be any blocking part. A blocking part is a portion of a nucleotide that inhibits or prevents the 3' oxygen of a nucleotide from forming a covalent bond with the next correct nucleotide during nucleic acid polymerization. The blocking part of a “reversible termination” nucleotide may be removed from a nucleotide analog or otherwise modified to allow the 3' oxygen of the nucleotide to covalently attach to the next correct nucleotide. Such a blocking part is referred to herein as a “reversible termination part.” The blocking part does not need to prevent or inhibit the formation of a ternary complex at the 3' end of the nucleic acid to which the blocking part is attached. The cap can be any capping part. The capping part can have a positive or negative charge that inhibits or prevents ternary complex formation. The capping part may include a ligand that binds to a receptor to inhibit or prevent ternary complex formation, such as biotin (or its analogues) that binds to streptavidin (or its analogues), an epitope that binds to an antibody (or its functional fragment), a carbohydrate that binds to a lectin, etc. Further examples of the capped portion are described in U.S. Patent Application Publication No. 2020 / 0032322A1 or Turcatti et al. Nucl. Acids. Res. 36(4) e25 (2008), each of which is incorporated herein by reference.
[0506] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a nucleotide incorporation blocker, or a ternary complex formation cap. In a further embodiment, the 3' end of the staple molecule includes a nucleotide incorporation blocker.
[0507] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a blocking agent to prevent nucleotide incorporation, or a cap to prevent ternary complex formation. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation.
[0508] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, nucleotide incorporation blocking agent, or ternary complex formation prevention cap. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation. In yet another embodiment, the 3' end of the staple molecule is capped by a portion preventing ternary complex formation.
[0509] In some embodiments, the 3' end of the staple molecule includes at least one mismatch, a blocking agent to prevent nucleotide incorporation, or a cap to prevent ternary complex formation. In a further embodiment, the 3' end of the staple molecule prevents ternary complex formation. In yet another embodiment, the 3' end of the staple molecule includes a mismatch and a blocking agent, optionally wherein the blocking agent is a 3' phosphate blocking agent.
[0510] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. Any suitable spacer known in the art can be used. In some instances, the spacer may contain a polynucleotide sequence. Alternatively, the spacer may contain a non-nucleotide polymer linker, such as, for example, a branched polyelectrolyte species. Non-limiting examples of branched polyelectrolyte species include polyethylene glycol (PEG) and dendritic macromolecules.
[0511] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, at least some of the one or more staple molecules do not contain spacers.
[0512] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise polynucleotide sequences. The length of the polynucleotide sequence is at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides.
[0513] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise polynucleotide sequences. In yet another embodiment, the length of the polynucleotide sequence is variable.
[0514] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacer comprises a polynucleotide sequence. In yet another embodiment, the polynucleotide sequence comprises a double-stranded DNA sequence.
[0515] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by a spacer. In a further embodiment, the spacer comprises a polynucleotide sequence. In yet another embodiment, the polynucleotide sequence comprises a double-stranded DNA sequence. In still a further embodiment, the first and second regions each include a 3' end that hybridizes to the tandem strand.
[0516] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers have variable lengths on different staple molecules.
[0517] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise non-nucleotide polymer connectors.
[0518] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise polyethylene glycol (PEG). PEG may comprise PEG with an average molecular weight of about 200 Daltons (e.g., PEG-200) to about 8000 Daltons (e.g., PEG-8000).
[0519] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. Species of dendritic macromolecules that can be used in the methods, compositions, and systems disclosed herein include, but are not limited to, branched polyamines comprising protonated structures that interact with the negatively charged DNA backbone to form complexes. It should be understood that, in the embodiments described herein, the adaptor element can still be used in conjunction with the branched polyamine for alternative purposes or to provide improved binding properties of the dendritic macromolecule species to nucleic acids. Dendritic macromolecule species may comprise controlled terminal surface chemistry having one or more functional groups, including but not limited to amine, carboxyl, and hydroxyl groups. Dendritic macromolecule species can be obtained as generation 0 (G0) through generation 10 (G10), with the number of branches in each generation being twice that of the previous generation. Thus, G0 = 4 branches, G1 = 8 branches, and so on. In some embodiments, the branched polyelectrolyte is a poly(amidoamine) dendritic macromolecule species (also known as PAMAM), such as the G2 PAMAM dendritic macromolecule having 16 branches and an amine (NH2) terminal surface chemistry. Non-limiting examples of branched polyelectrolytes also include G4 (64 branches with amine terminal groups) and G5 (128 branches with amine terminal groups) PAMAM dendritic macromolecule species.
[0520] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer connectors. In yet another embodiment, the nonnucleotide polymer connectors comprise dendritic macromolecules. In even further embodiments, the dendritic macromolecules comprise polyamidoamine (PAMAM). Any species of PAMAM described above, both those specifically described and those implied by the entirety of this disclosure, may be used.
[0521] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. In even further embodiments, the staple molecule hybridizes with 3, 4, 5, 6, 7, 8, 9, 10, or more instances of at least one adaptor. In an exemplary embodiment, the staple molecule hybridizes with 3 or more instances of at least one adaptor. In another exemplary embodiment, the staple molecule hybridizes with 10 or more instances of at least one adaptor sequence.
[0522] In some embodiments, at least some of the first and second regions of the staple molecule are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer linkers. In yet another embodiment, the nonnucleotide polymer linkers comprise dendritic macromolecules. In even further embodiments, the staple molecule hybridizes with three or more instances of at least one adaptor. In specific embodiments of even further embodiments, a majority of the staple molecule hybridizes with adaptor sequences flanking at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200 or more different target sequences. In an exemplary embodiment, a majority of the staple molecule hybridizes with adaptor sequences flanking at least 100 different target sequences.
[0523] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers comprise nonnucleotide polymer connectors. In yet another embodiment, the nonnucleotide polymer connectors comprise dendritic macromolecules. In even further embodiments, neither the first nor second region of the staple molecule contains a 3' end that allows the formation or extension of the ternary complex.
[0524] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers are coupled to the 3' end of the first region.
[0525] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by spacers. In a further embodiment, the spacers are coupled to the 3' end of the first region. In yet another further embodiment, the spacers are coupled to the 5' end of the second region.
[0526] In some embodiments, at least some of the first and second regions of one or more staple molecules are separated by a spacer. In a further embodiment, the spacer is coupled to the 3' end of the first region. In yet another further embodiment, the spacer is coupled to the 3' end of the second region and wherein the staple molecule cannot act as a primer.
[0527] In some embodiments, multiple staple molecules hybridize with the same tandem strand.
[0528] J. Reagent kit
[0529] In another aspect disclosed herein, a kit is provided comprising staple molecules of any one of the first aspect and embodiments, the second aspect and embodiments, or the fifth aspect and embodiments.
[0530] In some embodiments, the kit further comprises one or more of the following: (i) one or more stapled primers, (ii) one or more non-stapled primers, (iii) one or more adaptors, (iv) one or more polymerases, (v) one or more ligases, (vi) one or more splice oligonucleotides, (vii) a flow cell, (viii) multiple labeled nucleotides, (ix) multiple reversibly terminated nucleotides, (x) a capped portion or any combination thereof.
[0531] On the other hand, a kit is provided comprising: (i) a first adaptor sequence, (ii) a second adaptor sequence, and (iii) a staple molecule, wherein a first region and a second region of the staple molecule are each hybridized with different instances of the first adaptor sequence and / or the second adaptor sequence.
[0532] In some embodiments, the kit further includes (iv) a reagent sufficient to form a stable tandem from the target nucleic acid, wherein the stable tandem comprises a first adaptor, a second adaptor, and a staple molecule.
[0533] V. Usage Instructions
[0534] A. Sequencing methods
[0535] Various methods have been developed to determine the partial sequence of nucleic acid templates, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) molecules. The result of any sequencing method is the production of a set of sequence reads (or other derived language phrases). Therefore, any sequencing method can also be referred to as a method for producing a set of sequence reads. It is anticipated that all current sequencing methods, as well as future sequencing methods, can be improved using the aspects and examples of the compositions and methods for preparing stable tandems described above. Sequencing methods can be stratified at a high level based on factors such as: 1) the type of nucleic acid molecule used as a template (DNA sequencing or RNA sequencing), 2) the number of template ends being sequenced (single-end sequencing or paired-end sequencing), and 3) the read length produced by the sequencing method (short-read sequencing or long-read sequencing). Alternatively or additionally, sequencing methods can be stratified at a high level based on sequencing "generations".
[0536] 1. DNA sequencing and RNA sequencing :
[0537] The difference between DNA sequencing and RNA sequencing methods is primarily based on the type of nucleic acid molecule used as a template or starting material in a given workflow. DNA sequencing is the process of determining the nucleotide sequence of a given DNA fragment. RNA sequencing is the process of determining the nucleotide sequence of a given RNA fragment. RNA is less stable than DNA and is more susceptible to attack by nucleases in experiments. Therefore, RNA sequencing methods typically involve reverse transcription of RNA from a sample to generate complementary DNA (cDNA) fragments, and then using a given sequencing method to determine their sequence.
[0538] 2. Single-end sequencing and paired-end sequencing :
[0539] The main difference between single-end sequencing and paired-end sequencing methods is that single-end sequencing is configured to determine the sequence at one end of the nucleic acid template, while paired-end sequencing is configured to determine the sequence at the opposite ends of the nucleic acid template. Typically, in both single-end and paired-end sequencing, the sequence at the first end is determined by extending a first primer along the first strand of the nucleic acid template. Then, in paired-end sequencing, the sequence at the second end is determined by extending a second primer along the second strand of the nucleic acid template. Because the nucleotide orientation in one strand is opposite to that in the other strand (these strands are called "antiparallel"), the first and second primers extend in opposite directions and towards each other. Due to this orientation, some paired-end sequencing methods, in particular, require sequencing each strand in the absence of the other strand. Therefore, some paired-end methods typically require a strand removal or strand synthesis step between the two sequencing reads. Other paired-end sequencing methods allow sequencing of the first strand (sense strand) of the target nucleic acid in the presence of the second strand (antisense strand).
[0540] 3. Short read sequencing and long read sequencing :
[0541] High-throughput sequencing refers to "short read" and "long read" sequencing methods (as well as exome sequencing, genome sequencing, genome resequencing, transcriptome analysis (RNA-Seq), DNA-protein interaction sequencing (ChIP-sequencing), and epigenome characterization). The key difference between short read and long read sequencing lies primarily in the "read length" of the nucleic acid template. Short read sequencing methods result in a large number of copies (ranging from one million to several billion) of reads with an average length of 50 to 400 base pairs. Conversely, long read sequencing methods have the ability to sequence an average of 5,000 to 30,000 base pairs within a single read.
[0542] 4. First-generation sequencing :
[0543] Early iterations of sequencing methods were based on the chain termination method developed by Sanger and Coulson in 1975 (“Sanger sequencing”) or the chemical method (chain degradation) developed by Maxam and Gilbert between 1976 and 1977 (“Maxam-Gilbert sequencing”). Later, a modified version of Sanger sequencing was developed, known as “shotgun sequencing.” Newer sequencing methods that allow for large-scale parallel determination of nucleic acid molecular sequences have been developed and have significantly replaced first-generation sequencing methods.
[0544] 5. Second-generation sequencing :
[0545] Next-generation sequencing, also known as second-generation sequencing, massively parallel sequencing, or massively parallel sequencing, is any of several high-throughput methods that use the concept of massively parallel processing to sequence nucleic acid molecules. Next-generation sequencing methods are well-known in the field and primarily include short-read sequencing methods. Non-limiting examples of second-generation sequencing include: 454 pyrosequencing, sequencing-by-synthesis (SBS), sequencing-by-ligation, sequencing-by-hybridization, sequencing-by-binding (SBB), Illumina dye sequencing, Solexa sequencing, ion semiconductor sequencing, ABI SOLiD sequencing, sequencing using combined probes anchored ligation (cPAL sequencing), RNA-Seq, etc.
[0546] 6. Third-generation sequencing :
[0547] Long-read sequencing, also known as third-generation sequencing, is a class of sequencing methods currently under active development. Although the field is still relatively new, some third-generation sequencing methods are well-known in the field. Non-limiting examples of third-generation sequencing include: single-molecule real-time (SMRT) sequencing by Pacific Biosciences, nanopore sequencing by Oxford Nanopore Technologies, and single-molecule fluorescence sequencing by Helicos.
[0548] 7. sequencing by binding :
[0549] Sequencing by binding (SBB) is a sequencing technique in which the specific binding of polymerase and homologous nucleotides to the priming template nucleic acid is used to identify the next correct nucleotide in the primer strand to be incorporated into the priming template nucleic acid. This specific binding interaction does not require the incorporation of a nucleotide chemical into the primer. The specific binding interaction can occur before or before a similar next correct nucleotide chemical is incorporated into the primer strand. Therefore, the identification of the next correct nucleotide can be performed without incorporating it.
[0550] SBB has been described in U.S. Patent Nos. 10,443,098, 10,246,744, and U.S. Patent Application Publication No. 2018 / 0044727 A1; the contents of each of these references are incorporated herein by reference in their entirety. In short, in SBB, the polymerase undergoes a conformational transition between an open and closed conformation during discrete steps of the reaction. In one step, the polymerase binds to the initiating template nucleic acid to form a binary complex, also referred herein to as the pre-insertion conformation. In a subsequent step, the incoming nucleotide binds and the polymerase finger closes, resulting in a pre-chemical conformation comprising the polymerase, the initiating template nucleic acid, and the nucleotide; in which the bound nucleotide has not yet been incorporated. This step, also referred herein to as the check step, may be followed by a chemical incorporation step, in which a phosphodiester bond is formed, accompanied by the cleavage of pyrophosphate from the nucleotide (nucleotide incorporation). The polymerase, the initiating template nucleic acid, and the newly incorporated nucleotide produce a chemically post-translational pre-conformation. Since both the pre-chemical conformation and the pre-translocation conformation contain a polymerase, an initiating template nucleic acid, and a nucleotide, with the polymerase in a closed state, either conformation can be referred to herein as a ternary complex. The polymerase conformation and / or its interaction with the nucleic acid can be monitored during the inspection step to identify the next correct base in the nucleic acid sequence. Before or after incorporation, reaction conditions can be altered to dissociate the polymerase from the initiating template nucleic acid, and again to remove any reagents that inhibit polymerase binding from the local environment.
[0551] Generally, the SBB procedure includes a “checking” step to identify the next template base, and an optional “incorporation” step to add one or more complementary nucleotides to the 3' end of the primer component of the priming template nucleic acid. The identity of the next correct nucleotide to be added is determined either before or without chemically linking the nucleotide to the 3' end of the primer via covalent bonding. The checking step may involve providing the priming template nucleic acid to be used in the procedure, and contacting the priming template nucleic acid with a polymerase (e.g., DNA polymerase) and one or more test nucleotides investigated as possible next correct nucleotides. Furthermore, there are steps involving monitoring or measuring the interaction between the polymerase and the priming template nucleic acid in the presence of the test nucleotides.
[0552] The inspection process typically includes the following sub-steps: (1) providing a template nucleic acid for initiation (i.e., a template nucleic acid molecule that hybridizes with a primer, which may optionally be blocked from extending at its 3' end); (2) contacting the template nucleic acid for initiation with a reaction mixture comprising a polymerase and at least one nucleotide; (3) monitoring the interaction between the polymerase and the template nucleic acid molecule in the presence of a nucleotide and without any nucleotide chemical incorporation into the template nucleic acid for initiation; and (4) determining the identity of the next base in the template nucleic acid (i.e., the next correct nucleotide) based on the monitored interaction.
[0553] The check step can be controlled to reduce or complete nucleotide incorporation. If nucleotide incorporation is reduced, a separate incorporation step can be performed. A separate incorporation step can be performed without monitoring because the bases have already been identified during the check step. If nucleotide incorporation is performed during the check step, subsequent nucleotide incorporation may be weakened by the stabilizers that trap the polymerase on the nucleic acid after incorporation. A reversibly terminated nucleotide can be used in the incorporation step to prevent the addition of more than one nucleotide during a single cycle.
[0554] The SBB method allows for controlled determination of template nucleic acid bases without the need for labeled nucleotides, as the interaction between the polymerase and the template nucleic acid can be monitored without labeling the nucleotides. Controlled nucleotide incorporation can also provide accurate sequence information for repetitive and homopolymer regions without the use of labeled nucleotides. Furthermore, template nucleic acid molecules can be sequenced under inspection conditions that do not require attachment of the template nucleic acid or polymerase to a solid support. However, in some preferred embodiments, the priming template nucleic acid to be sequenced is attached to a solid support, such as the inner surface of a flow cell.
[0555] The check procedure can be partially controlled by providing reaction conditions to prevent the chemical incorporation of nucleotides, while allowing the identification of the next correct base on the initiated template nucleic acid molecule. These reaction conditions can be called check reaction conditions.
[0556] The inspection typically involves detecting the interaction between the polymerase and the template nucleic acid. Detection can include optical, electrical, thermal, acoustic, chemical, and mechanical methods. Generally, the inspection step involves binding the polymerase to the polymerization initiation site of the initiated template nucleic acid in a reaction mixture containing one or more nucleotides and monitoring the interaction. The inspection step of the sequencing reaction can be repeated one, two, three, four, or more times before the incorporation step. The inspection and incorporation steps can be repeated until the desired template nucleic acid sequence is obtained.
[0557] SBB involves the contact of a template nucleic acid molecule with a reaction mixture containing a polymerase and one or more nucleotide molecules, preferably under conditions that stabilize the formation of a ternary complex without stabilizing the formation of a binary complex. The formation of a ternary complex, or a stabilized ternary complex, can be used to ensure that only one nucleotide is added to the template nucleic acid primer per sequencing cycle, wherein the added nucleotide is isolated within the ternary complex. Controlled incorporation of a single nucleotide per sequencing cycle improves sequencing accuracy for nucleic acid regions containing homopolymer repeats.
[0558] In SBB, the reaction mixture used in the examination step can include one, two, three, or four types of nucleotide molecules. The nucleotides can be selected from dATP, dTTP (or dUTP), dCTP, and dGTP. The reaction mixture can contain one or more triphosphate nucleotides and one or more diphosphate nucleotides. A ternary complex can be formed between the initiating template nucleic acid, the polymerase, and any of the four types of nucleotide molecules, thus allowing for the formation of four types of ternary complexes.
[0559] Monitoring or measuring the interaction between a polymerase and the initiated template nucleic acid molecule in the presence of nucleotide molecules can be accomplished in many different ways. For example, monitoring may include measuring the association kinetics of the interaction between the initiated template nucleic acid, the polymerase, and any of the four nucleotide molecules. Monitoring the interaction between the polymerase and the initiated template nucleic acid molecule in the presence of nucleotide molecules may include measuring the equilibrium binding constant between the polymerase and the initiated template nucleic acid molecule (i.e., the equilibrium binding constant between the polymerase and the template nucleic acid in the presence of any one or all four nucleotides). Thus, for example, monitoring includes measuring the equilibrium binding constant between the polymerase and the initiated template nucleic acid in the presence of any of the four nucleotides. Monitoring the interaction between the polymerase and the initiated template nucleic acid molecule in the presence of nucleotide molecules includes measuring the kinetics of the dissociation of the polymerase from the initiated template nucleic acid in the presence of any of the four nucleotides.
[0560] Monitoring steps may include monitoring the steady-state interaction between the polymerase and the initiated template nucleic acid molecule in the presence of a first nucleotide molecule, without chemically incorporating the first nucleotide molecule into the primer of the initiated template nucleic acid molecule. Monitoring may include monitoring the dissociation of the polymerase from the initiated template nucleic acid molecule in the presence of a first nucleotide molecule, without chemically incorporating the first nucleotide molecule into the primer of the initiated template nucleic acid molecule. Monitoring may include monitoring the association between the polymerase and the initiated template nucleic acid molecule in the presence of a first nucleotide molecule, without chemically incorporating the first nucleotide molecule into the primer of the initiated template nucleic acid molecule. Similarly, the test nucleotide in these procedures may be a natural nucleotide (i.e., unlabeled), a labeled nucleotide (e.g., a fluorescently labeled nucleotide), or a nucleotide analog (e.g., a nucleotide modified to include a reversible terminator motif).
[0561] In SBB, chemical blocking on the 3' nucleotide of the primer of the initiating template nucleic acid molecule (e.g., a reversible terminator moiety on the base or sugar of the nucleotide), or the absence of catalytic metal ions in the reaction mixture, or the absence of catalytic metal ions in the active site of the polymerase, prevents the chemical incorporation of nucleotides into the primer of the initiating template nucleic acid.
[0562] The identity of the next correct base or nucleotide can be determined by monitoring the presence, formation, and / or dissociation of the ternary complex. The identity of the next base can be determined without chemically incorporating the next correct nucleotide into the 3' end of the primer. The identity of the next base can also be determined by monitoring the affinity of the polymerase for the initiated nucleic acid template in the presence of an added nucleotide.
[0563] SBB may include an incorporation step. The incorporation step involves chemically incorporating one or more nucleotides at the 3' end of a primer that binds to the template nucleic acid. The incorporation reaction can be promoted by incorporating a reaction mixture. The incorporation reaction mixture may have a different nucleotide composition than the test reaction. For example, the test reaction may include one type of nucleotide, and the incorporation reaction may include another type of nucleotide. As another example, the test reaction may contain one type of nucleotide, and the incorporation reaction may contain four types of nucleotides, or vice versa. The test reaction mixture may be altered or replaced by the incorporation reaction mixture.
[0564] Nucleotides present in the reaction mixture but not isolated within the ternary complex may cause multiple nucleotide insertions. A washing step can be performed prior to the chemical incorporation step to ensure that only nucleotides isolated within the captured ternary complex are available for incorporation during the incorporation step. The captured polymerase complex can be a ternary complex, a stabilized ternary complex, or a ternary complex involving the polymerase, the initiated template nucleic acid, and the next correct nucleotide.
[0565] 8. Synthesis-sequencing :
[0566] "Sequencing-by-synthesis" (SBS) typically involves the enzymatic extension of nascent primers by iteratively adding nucleotides to the template strand that hybridizes with the primers. SBS has been described, for example, in U.S. Patent Nos. 5,302,509 and 6,828,100 and U.S. Patent Application Publication No. 2009 / 0247414 A1; the contents of each of these references are incorporated herein by reference in their entirety. SBS differs from the aforementioned SBB in that a labeled nucleotide is incorporated into the extended strand, measured, and then the tag is removed or inactivated, and the 3' block is removed, to iteratively sequence the template. In SBB, the labeled base is not incorporated into the extended strand. Instead, the formation of the ternary complex is typically measured in the presence of the labeled base, but sometimes in the presence of a labeled polymerase or other signature, after which the complex is broken down, and the primer strand is extended using unlabeled bases with a 3' block.
[0567] In short, SBS can be initiated by contacting a target nucleic acid attached to a site in a flow cell with one or more labeled nucleotides, DNA polymerase, etc. Those sites extending the primer using the target nucleic acid as a template will incorporate detectable labeled nucleotides. Detection can include scanning using the devices or methods described herein. Optionally, the labeled nucleotides may further include a reversible termination property that terminates further primer extension once the nucleotide has been added to the primer. For example, a nucleotide analog with a reversible termination moiety can be added to the primer, preventing subsequent extension until a deblocking agent is delivered to remove the moiety. Thus, in embodiments using reversible termination, the deblocking agent can be delivered to the container (before or after detection). Washing can be performed between the various delivery steps. The cycle can be performed n times to extend the primer by n nucleotides, thereby detecting a sequence of length n.
[0568] Exemplary SBS procedures, reagents, and detection components readily applicable to use with the methods or compositions of this disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008); WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Patent Nos. 7,057,026; 7,329,492; 7,211,414; 7,315,019 or 7,405,281; and U.S. Patent Application Publication No. 2008 / 0108082 A1, each of which is incorporated herein by reference. Also available are SBS methods commercially available from Illumina, Inc. (San Diego, Calif.). One or more reagents used in the SBS process may optionally be delivered via a mixed-phase fluid (e.g., a fluid foam, fluid slurry, or fluid emulsion), contacted with a mixed-phase fluid, and / or removed by a mixed-phase fluid. During the SBS process, the mixed-phase fluid can be taken out of the flow cell for testing.
[0569] B. Advantages of using stable tandem sequencing
[0570] As mentioned above, tandems are typically unstable during sequencing reactions, negatively impacting the integrity of the sequencing results. Sequencing methods using stabilized tandems demonstrate several unexpected benefits and improvements compared to methods known in the art. Non-limiting examples of these unexpected benefits and improvements compared to methods known in the art are described below.
[0571] 1. Resisting dimensional changes :
[0572] Variations in FWHM along the column and / or row directions during sequencing runs indicate instability of the tandem. Tandems stabilized with staple molecules are resistant to FWHM variations. Small changes in FWHM (i.e., reductions in FWHM increases) allow for additional sequencing cycles on the stabilized tandems. This contributes to increased read integrity and length.
[0573] In some embodiments, after at least 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 150, 175, or 200 sequencing cycles, the average full width at half maximum (FWHM) of the stabilized tandem does not increase by more than 5% in at least one dimension. In an exemplary embodiment, after at least 100 sequencing cycles, the average FWHM of the stabilized tandem does not increase by more than 5% in at least one dimension. In some embodiments, after at least 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 150, 175, or 200 sequencing cycles, the average increase in FWHM of the stabilized tandem, measured in at least one dimension, is reduced by at least 2, 3, 4, or 5 times by the staple molecules. In an exemplary embodiment, after at least 100 sequencing cycles, the average increase in FWHM of the stabilized tandem molecules, measured in at least one dimension, is reduced by at least 5-fold by the staple molecules.
[0574] 2. Increase sequencing cycle number and read length :
[0575] When tandem stabilizes, the length of sequencing runs should be shortened to maintain the integrity of the sequencing reads. However, this results in shorter read lengths, which may prevent the capture of at least some of the present variant sequences. Staple molecular stabilization of tandems allows for increased sequencing cycle numbers and longer read lengths compared to sequencing methods that do not stabilize the tandem.
[0576] In some embodiments, sequencing is performed on a stable tandem with read lengths greater than 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150 cycles. In an exemplary embodiment, sequencing is performed on a stable tandem with read lengths greater than 150 cycles.
[0577] 3. Improve signal strength and resolution :
[0578] When tandems become dispersed, molecular dispersion causes spectral bleeding, interfering with the distinction between tandem intensity values and background intensity. Furthermore, as tandems become dispersed, the probability of multiple points being perceived as monoclonal signals increases. Staple-based molecular stabilization of tandems suppresses molecular dispersion, allowing for better distinction between ON and OFF intensity values. Moreover, staple-based molecular stabilization of tandems produces a larger ratio of monoclonal to multiple clonal points at higher point densities, overcoming a significant obstacle in short-read sequencing methods.
[0579] In some embodiments, the stabilized tandems are sequenced with read lengths greater than 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150 cycles. In further embodiments, the average signal obtained from the stapled tandems in the final sequencing cycle is 10%, 20%, 30%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50% higher than the average signal obtained from tandems not stabilized by stapled molecules. In one exemplary embodiment, the stabilized tandem sequence is sequenced with a read length greater than 150 cycles, and the average signal obtained from the tandem sequence stabilized by the stapled molecule is 50% higher in the final sequencing cycle compared to the average signal obtained from the tandem sequence stabilized by the stapled molecule. The read length can be defined by a signal-to-noise ratio below which the bases it no longer calls. For example, the read length can be a cycle at which the signal-to-noise ratio drops below 2, 3, or 4. Alternatively, the read length can be defined by a precision below which the bases it no longer calls. For example, the read length can be a cycle at which the base-calling precision drops below 99%, 99.5%, 99.9%, 99.95%, or 99.99%.
[0580] C. Single-ended clustering workflow
[0581] Various methods have been developed to determine the sequences of different portions of a nucleic acid template. In some cases, these methods are configured to determine the sequence at one end of the nucleic acid template, and are therefore referred to as "single-end" sequencing. While any single-end sequencing method can be utilized in the aspects and embodiments described above and below, exemplary workflows for synthesizing tandem copies requiring single-end sequencing are provided, such as... Figure 13 The examples shown are for further understanding of the contents of this document.
[0582] The single-end clustering workflow can begin with the synthesis of tandems from a linear nucleic acid template containing the target sequence. As described above, tandems can be synthesized in solution in a container or attached to a surface. The linear nucleic acid template is prepared by extracting nucleic acids from a sample, followed by fragmentation and size selection. Methods for nucleic acid isolation, fragmentation, and size capture are described above. As described above, one or more adaptor sequences are ligated to the linear nucleic acid template. The linear nucleic acid template now containing the target sequence and one or more adaptor sequences can be annealed with a guide oligonucleotide having regions complementary to one or more adaptors of the linear nucleic acid. In some configurations, the guide oligonucleotide binds to the surface. The guide oligonucleotide acts as a clamp (or "clamp oligonucleotide"), causing a conformational change in the linear nucleic acid such that the 5' and 3' ends are brought close together. The 5' and 3' ends of the linear nucleic acid template are joined to circularize the template. The ends can be joined by hybridization with clamp oligonucleotides in solution or on a surface, such as, for example, the surface of a solid support, or the surface of a flow cell, or the surface of a solid support in a flow cell. Alternatively, CircLigase™ (Epicenter, Madison WI) or other enzymes capable of splinting the ends of nucleic acids can be used to circularize linear nucleic acid templates without splints.
[0583] The guide oligonucleotide is then extended to add multiple monomeric units of the original linear nucleic acid at its 3' end via RCA. During the extension / RCA, an initial inoculation / extension phase is initiated, allowing the formation of nascent tandems and the commencement of amplification. This phase can be carried out in the presence of a first reaction mixture containing: non-catalytic metal ions, dendritic macromolecules such as, for example, PAMAM, polymerase, and various dNTPs. Following initial inoculation and extension, buffer exchange is performed, such as... Figure 15 As shown. Subsequently, the synthesis of tandem polymers continues during a second extension phase. The second extension phase is carried out in the presence of a second reaction mixture containing a higher concentration of dendritic macromolecules and non-catalytic metal ions, but without the addition of additional polymerase to the mixture. Tandem polymers of the previously linear nucleic acids are produced. The tandem polymers can be attached to the surface via guiding oligonucleotides, or the tandem polymers can be synthesized in solution and then attached to the surface.
[0584] RCA reactions can be terminated via thermal inactivation of the polymerase or by washing (alone or in combination with chemical inactivation). In one configuration, the RCA reaction is stopped by denaturing the polymerase, for example by heating the sample at 60°C, 65°C, 70°C, 75°C, 80°C, or higher. In another configuration, the RCA reaction is stopped by removing one or more components of the RCA, such as the polymerase and dNTPs. Components of the RCA can be removed, for example, by washing. The tandem can be stabilized by providing staple molecules before, during, or after polymerase inactivation, as further described below. After the RCA reaction is terminated, the tandem can be subjected to any number of single-end sequencing methods, or the tandem can be provided for a paired-end clustering workflow, as described below.
[0585] D. Two-terminal clustering and sequencing workflow
[0586] Various methods have been developed to determine the sequences of different portions of a nucleic acid template. In some cases, these methods are configured to determine the sequences at opposite ends of the nucleic acid template and are therefore referred to as "paired-end" sequencing. Some methods for paired-end sequencing of nucleic acid templates require the degradation of at least one strand of the tandem. In other configurations, paired-end sequencing does not require the degradation of at least one strand of the tandem. Although any paired-end sequencing method can be utilized in the aspects and embodiments described above and below, exemplary workflows for paired-end clustering and sequencing of tandems are provided, such as... Figure 14 The examples shown are intended to facilitate a further understanding of the contents disclosed herein.
[0587] An exemplary paired-end clustering / sequencing workflow can begin from any of the steps in the single-end clustering workflow described above. Therefore, the description of the paired-end clustering / sequencing workflow begins with the first strand of the synthesized / provided tandem strand and the termination of the RCA reaction. Subsequently, the second reaction mixture is removed by washing and replaced with an appropriate buffer / reaction mixture. Multiple primers hybridize with the first strand of the tandem strand, generating multiple second strands of the tandem strand via MDA. The 3' end of the second strand of the tandem strand is blocked or capped to prevent further amplification during sequencing. Additional primers are then hybridized with multiple second strands, and the multiple second strands are sequenced. These additional primers may contain non-staple primers and / or staple molecules configured to act as staple primers. Figure 14 As shown, after sequencing multiple second strands, multiple second strands are removed, and the first strand of the tandem primer is hybridized with additional primers to allow the sequencing reaction to occur. The additional primers may contain non-staple primers and / or staple molecules configured to act as staple primers.
[0588] E. Sequencing using staple molecules
[0589] In a third aspect of the invention, a sequencing method is provided, comprising: (i) providing a plurality of stable tandem sequences of the second aspect and / or any of the embodiments of the second aspect, or providing a plurality of stable tandem sequences by any of the methods of the first aspect and / or the embodiments of the first aspect; and (ii) sequencing at least a first portion of the target sequence.
[0590] In some embodiments, sequencing step (ii) uses reversibly terminated nucleotides.
[0591] In some embodiments, the staple molecule comprises the staple primer from the sequencing step (ii).
[0592] In some embodiments, the staple molecule comprises the staple primer from sequencing step (ii). In a further embodiment, after sequencing from the staple primer, the staple primer is blocked or capped, and additional primers are hybridized to the tandem, followed by sequencing from the additional primers. As described above, in some instances, the additional primers may be non-staple primers. Alternatively or additionally, the additional primers may be staple molecules serving as staple primers.
[0593] In some embodiments, sequencing step (ii) includes sequencing-by-synthesis (SBS).
[0594] In some embodiments, sequencing step (ii) includes sequencing-by-binding (SBB).
[0595] In some embodiments, sequencing step (ii) includes using a staple molecule as a staple primer (SBB).
[0596] In some embodiments, sequencing step (ii) includes an SBB using a staple molecule as a staple primer. In a further embodiment, the SBB includes the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer, (B) check: forming and detecting a stable ternary complex containing the staple primer, and (C) activation: cleaving the reversible terminator from the staple primer.
[0597] As described above, the examination steps may involve providing a template nucleic acid for initiation to be used in the procedure, and contacting the template nucleic acid with a polymerase (e.g., DNA polymerase) and one or more test nucleotides to be studied as possible next correct nucleotides. Furthermore, there are steps involving monitoring or measuring the interaction between the polymerase and the template nucleic acid in the presence of the test nucleotides.
[0598] The check step typically includes the following sub-steps: (1) providing a template nucleic acid for initiation (i.e., a template nucleic acid molecule hybridized to a primer, which may optionally be blocked at its 3' end); (2) contacting the template nucleic acid for initiation with a reaction mixture comprising a polymerase and at least one nucleotide; (3) monitoring the interaction between the polymerase and the template nucleic acid molecule in the presence of a nucleotide and without any nucleotide chemical incorporation into the template nucleic acid for initiation; and (4) determining the identity of the next base in the template nucleic acid (i.e., the next correct nucleotide) based on the monitored interaction. The check step can be controlled to reduce or complete nucleotide incorporation. If nucleotide incorporation is reduced, a separate incorporation step can be performed. A separate incorporation step can be performed without monitoring because the base has already been identified during the check step. If nucleotide incorporation is performed during the check step, subsequent nucleotide incorporation may be reduced by the stabilizers that trap the polymerase on the nucleic acid after incorporation. A reversibly terminated nucleotide can be used in the incorporation step to prevent the addition of more than one nucleotide during a single cycle.
[0599] In some embodiments, sequencing step (ii) includes an SBB using a staple molecule as a staple primer. In a further embodiment, the SBB includes the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer; (B) check: forming and detecting a stable ternary complex containing the staple primer; and (C) activation: cleaving the reversibly terminated nucleotide from the staple primer. In yet another embodiment, the stable ternary complex comprises at least one of a labeled nucleotide and a labeled polymerase.
[0600] Any suitable marker can be used. Non-limiting examples of markers that can be used in some embodiments include: acridine orange (+DNA), acridine orange (+RNA), Alexa Fluor® 350, Alexa Fluor® 430, Alexa Fluor® 488, Alexa Fluor® 532, Alexa Fluor® 546, Alexa Fluor® 555, Alexa Fluor® 568, Alexa Fluor® 594, Alexa Fluor® 633, Alexa Fluor® 647, Alexa Fluor® 660, Alexa Fluor® 680, Alexa Fluor® 700, Alexa Fluor® 750, allophycocyanin (APC), AMCA / AMCA-X, 7-aminoactinomycin D (7-AAD), 7-amino-4-methylcoumarin, 6-aminoquinoline, aniline blue, ANS, APC-Cy7, ATTO-TAG™ CBQCA, ATTO-TAG™ FQ, Auramine O-Folgen, BCECF (High pH), BFP (Blue Fluorescent Protein), BFP / GFP FRET, BOBO™-1 / BO-PRO™-1, BOBO™-3 / BO-PRO™-3, BODIPY® FL, BODIPY® TMR, BODIPY® TR-X, BODIPY® 530 / 550, BODIPY® 558 / 568, BODIPY® 564 / 570, BODIPY® 581 / 591, BODIPY® 630 / 650-X, BODIPY® 650-665-X, BTC, Calcein, Calcein Blue, CalciumCrimson™, Calcium Green-1™, Calcium Orange™, Calcofluor® White, 5-Carboxyfluorescein (5-FAM), 5-Carboxynaphthylfluorescein, 6-Carboxyrhodamine 6G, 5-Carboxytetramethylrhodamine (5-TAMRA), Carboxy-X-Rhodamine (5-ROX), Cascade Blue®, Cascade Yellow™, CCF2 (GeneBLAzer™), CFP (Cyan Fluorescent Protein), CFP / YFP FRET, Chromomycin A3, Cl-NERF (Low pH), CPM, 6-CR 6G, CTC Formazan, Cy2®, Cy3®, Cy3.5®, Cy5®, Cy5.5®, Cy7®, Phycoerythrin-Cy5 conjugate (PE-Cy5), Dansylamide, Dansyl cadaverine, Dansyl chloride, DAPI, Dapoxyl, DCFH, DHR, DiA (4-Di-16-ASP), DiD (DilC18(5)), DIDS, Dil (DilC18(3)), DiO (DiOC18(3)), DiR (DilC18(7)), Di-4 ANEPPS, Di-8 ANEPPS, DM-NERF (4.5-6.5pH), DsRed (red fluorescent protein), EBFP, ECFP, EGFP, ELF®-97 alcohol, Eosin, Erythrosine, Ethidium bromide, Ethidium homodimer-1 (EthD-1), Europium(III) chloride, 5-FAM (5-carboxyfluorescein), Glyceryl Blue, Fluorescein-dT phosphorimide, FITC, Fluo-3, Fluo-4, FluorX®, Fluoro-Gold™ (high pH), Fluoro-Gold™ (low pH), Fluoro-Jade, FM® 1-43, Fura-2 (high calcium), Fura-2 / BCECF, Fura Red™ (high calcium), Fura Red™ / Fluo-3, GeneBLAzer™ (CCF2), Redshifted GFP (rsGFP), Wild-type GFP, GFP / BFP FRET, GFP / DsRed FRET, Hoechst 33342 & 33258, 7-hydroxy-4-methylcoumarin (pH 9), 1,5 IAEDANS, Indo-1 (High Calcium), Indo-1 (Low Calcium), Indo-dicarbonylcyanine, Indo-tricarbonylcyanine, JC-1, 6-JOE, JOJO™-1 / JO-PRO™-1, LDS 751 (+DNA), LDS 751 (+RNA), LOLO™-1 / LO-PRO™-1, Fluorescent Yellow, LysoSensor™ Blue (pH 5), LysoSensor™ Green (pH 5), LysoSensor™ Yellow / Blue pH 4.2) LysoTracker® Green, LysoTracker® Red, LysoTracker® Yellow, Mag-Fura-2, Mag-Indo-1, Magnesium Green™ Marina Blue®, 4-Methylumbelliferone, scintillans, MitoTracker® Green, MitoTracker® Orange, MitoTracker® Red, NBD (amine), Nile Red, Oregon Green® 488, Oregon Green® 500, Oregon Green® 514, Pacific Blue, PBF1, PE (R-phycoerythrin), PE-Cy5, PE-Cy7, PE-Texas Red, PerCP (polydiophycoxanthin-chlorophyll-protein), PerCP-Cy5.5 (TruRed), PharRed (APC-Cy7), C-phycocyanin, R-phycocyanin, R-phycoerythrin (PE), PI (Propidium) Iodide), PKH26, PKH67, POPO™-1 / PO-PRO™-1, POPO™-3 / PO-PRO™-3, propidium iodide (PI), PyMPO, Pyrene, Pyronine Y, Quantum Red (PE-Cy5), Nitrogen mustard quinacrine, R670 (PE-Cy5), Red 613 (PE-Texas Red), Red Fluorescent Protein (DsRed), Halogen, RH 414, Rhod-2, Rhodamine B, Rhodamine Green™, Rhodamine Red™, Rhodamine-labeled phalloidin, Rhodamine 110, Rhodamine 123, 5-ROX (carboxy-X-rhodamine), S65A, S65C, S65L, S65T, SBFI, SITS, SNAFL®-1 (high pH), SNAFL®-2, SNARF®-1 (high pH), SNARF®-1 (low pH), Sodium Green™, Spectrum Aqua®, Spectrum Green® #1, Spectrum Green® #2, Spectrum Orange®, Spectrum Red®, SYTO® 11, SYTO® 13, SYTO® 17, SYTO® 45. SYTOX® Blue, SYTOX® Green, SYTOX® Orange, 5-TAMRA (5-carboxytetramethylrhodamine), Tetramethylrhodamine (TRITC), Texas Red® / Texas Red®-X, Texas Red®-X (NHS ester), Thiadicyanine, Thiazole Orange, TOTO®-1 / TO-PRO®-1, TOTO®-3 / TO-PRO®-3, TO-PRO®-5, Tri-color (PE-Cy5), TRITC (Tetramethylrhodamine), TruRed (PerCP-Cy5).5) WW 781, X-Rhodamine (XRITC), Y66F, Y66H, Y66W, YFP (yellow fluorescent protein), YOYO®-1 / YO-PRO®-1, YOYO®-3 / YO-PRO®-3, 6-FAM (fluorescein), 6-FAM (NHS ester), 6-FAM (azide), HEX, TAMRA (NHS ester), Yakima Yellow, MAX, TET, TEX615, ATTO 488, ATTO 532, ATTO 550, ATTO 565, ATTO Rho101, ATTO 590, ATTO 633, ATTO 647N, TYE 563, TYE 665, TYE 705, 5' IRDye® 700, 5' IRDye® 800, 5' IRDye® 800CW (NHS ester), WellRED D4 dye, WellRED D3 dye, WellRED D2 dye, Lightcycler® 640 (NHS ester), Dy 750 (NHS ester), horseradish peroxidase (HRP), soybean peroxidase (SP), alkaline phosphatase, luciferase, 2,4,5-triphenylimidazolium, p-dimethylamino and p-methoxy substituents, oxalates such as oxaloyl active ester, p-nitrophenyl, N-alkyl acridine onion ester, luciferin, gloss enhancer, and acridine onion ester.
[0601] In some embodiments, sequencing step (ii) includes an SBB using a staple molecule as a staple primer. In a further embodiment, the SBB includes the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer; (B) check: forming and detecting a stable ternary complex containing the staple primer; and (C) activation: cleaving the reversible terminator from the staple primer. In yet another embodiment, the SBB cyclic steps are repeated at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 times or more. In an exemplary embodiment, the SBB cyclic steps are repeated 50 times or more.
[0602] In some embodiments, sequencing step (ii) includes an SBB using a staple molecule as a staple primer. In a further embodiment, the SBB includes the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer; (B) check: forming and detecting a stable ternary complex containing the staple primer; and (C) activation: cleaving the reversible terminator from the staple primer. In yet another embodiment, the method further includes capping the staple primer.
[0603] In some embodiments, sequencing step (ii) includes SBB using staple molecules as staple primers. In a further embodiment, SBB includes the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer, (B) check: forming and detecting a stable ternary complex containing the staple primer, and (C) activation: cleaving the reversible terminator from the staple primer. In yet another embodiment, the method further includes capping the staple primer. In even further embodiments, the method further includes hybridizing a non-staple primer with a stable tandem and sequencing a second portion of the target sequence during the SBB process.
[0604] In some embodiments, sequencing step (ii) includes SBB using a staple molecule as a staple primer. In a further embodiment, SBB includes the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer, (B) check: forming and detecting a stable ternary complex containing the staple primer, and (C) activation: cleaving the reversible terminator from the staple primer. In yet another embodiment, the method further includes capping the staple primer. In even further embodiments, the method further includes hybridizing a non-staple primer with a stable tandem and sequencing a second portion of the target sequence during the SBB process. In a particular embodiment of even further embodiments, the second portion of the target sequence (a) overlaps with at least a portion of the first portion, (b) is upstream of the first portion, or (c) is downstream of the first portion.
[0605] In some embodiments, sequencing step (ii) includes SBB using staple molecules as staple primers. In a further embodiment, SBB includes the following cyclic steps: (A) extension: adding a reversibly terminated nucleotide to the staple primer, (B) check: forming and detecting a stable ternary complex containing the staple primer, and (C) activation: cleaving the reversible terminator from the staple primer. In yet another embodiment, the method further includes capping the staple primer. In even further embodiments, the method further includes hybridizing a non-staple primer with a stable tandem and sequencing a second portion of the target sequence during the SBB process. In a particular embodiment of even further embodiments, sequencing of the second portion of the target sequence is repeated multiple times.
[0606] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands.
[0607] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands of the plurality of antisense strands.
[0608] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands of the plurality of antisense strands. In yet another embodiment, the method further comprises sequencing at least some of the plurality of antisense strands of the plurality of stabilized tandem strands.
[0609] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands of the plurality of antisense strands. In yet another embodiment, the method further comprises sequencing at least some of the plurality of antisense strands of the plurality of stabilized tandem strands. In even further embodiments, the method further comprises degrading the plurality of antisense strands of the plurality of stabilized tandem strands, thereby releasing the staple molecule, optionally wherein the degradation comprises uracil nucleotides that digest the antisense strands.
[0610] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands of the plurality of antisense strands. In yet another embodiment, the method further comprises sequencing at least some of the plurality of antisense strands of the plurality of stabilized tandem strands. In even further embodiments, the method further comprises degrading the plurality of antisense strands of the plurality of stabilized tandem strands, thereby releasing the staple molecule, optionally wherein the degradation comprises digesting uracil nucleotides of the antisense strands. In a particular embodiment of even further embodiments, the degradation comprises digesting uracil with a uracil-DNA glycosylase.
[0611] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands of the plurality of antisense strands. In yet another embodiment, the method further comprises sequencing at least some of the plurality of antisense strands of the plurality of stabilized tandem strands. In even further embodiments, the method further comprises degrading the plurality of antisense strands of the plurality of stabilized tandem strands, thereby releasing the staple molecule, optionally wherein the degradation comprises uracil nucleotides digesting the antisense strands. In a particular embodiment of even further embodiments, the method further comprises contacting the plurality of tandem strands with the staple molecule, each of the staple molecule hybridizing with different instances of at least one adaptor sequence on the sense strand.
[0612] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands of the plurality of antisense strands. In yet another embodiment, the method further comprises sequencing at least some of the plurality of antisense strands of the plurality of stabilized tandem strands. In even further embodiments, the method further comprises degrading the plurality of antisense strands of the plurality of stabilized tandem strands, thereby releasing the staple molecule, optionally wherein the degradation comprises uracil nucleotides digesting the antisense strands. In a particular embodiment of even further embodiments, the method further comprises contacting the plurality of tandem strands with the staple molecule, each of the staple molecule hybridizing with different instances of at least one adaptor sequence on the sense strand. In even more specific embodiments, the method further comprises sequencing the sense strand of each tandem strand.
[0613] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, a first region and a second region of the staple molecule each hybridize with different instances of at least one adaptor sequence on different antisense strands of the plurality of antisense strands. In yet another embodiment, the method further comprises sequencing at least some of the plurality of antisense strands of the plurality of stabilized tandem strands. In even further embodiments, the method further comprises degrading the plurality of antisense strands of the plurality of stabilized tandem strands, thereby releasing the staple molecule, optionally wherein the degradation comprises uracil nucleotides digesting the antisense strands. In a particular embodiment of even further embodiments, the method further comprises contacting the plurality of tandem strands with the staple molecule, each of the staple molecule hybridizing with different instances of at least one adaptor sequence on the sense strand. In even more specific embodiments, the method further comprises sequencing the sense strand of each tandem strand. In even more specific embodiments, sequencing of the sense strand begins with the staple molecule hybridizing with the sense strand.
[0614] In some embodiments, each of the plurality of stabilized tandem molecules comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, the staple molecule hybridizes with the sense strand.
[0615] In some embodiments, each of the plurality of stabilized tandem strands comprises a sense strand that hybridizes with a plurality of antisense strands. In a further embodiment, staple molecules hybridize with sense strands. In yet another embodiment, the stabilized tandem strand comprises a plurality of staple molecules hybridizing with sense strands, and wherein the plurality of staple molecules extend to form a plurality of antisense strands.
[0616] In some embodiments, after at least 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 150, 175, or 200 sequencing cycles, the average full width at half maximum (FWHM) of the stabilized tandem does not increase by more than 5% in at least one dimension. In an exemplary embodiment, after at least 100 sequencing cycles, the average FWHM of the stabilized tandem does not increase by more than 5% in at least one dimension.
[0617] In some embodiments, after at least 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 150, 175, or 200 sequencing cycles, the average increase in FWHM of the stabilized tandem, measured in at least one dimension, is reduced by at least 2, 3, 4, or 5-fold by staple molecules. In an exemplary embodiment, after at least 100 sequencing cycles, the average increase in FWHM of the stabilized tandem, measured in at least one dimension, is reduced by at least 5-fold by staple molecules.
[0618] In some embodiments, sequencing is performed on a stable tandem with read lengths greater than 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150 cycles. In an exemplary embodiment, sequencing is performed on a stable tandem with read lengths greater than 150 cycles.
[0619] In some embodiments, the stabilized tandems are sequenced with read lengths greater than 20, 30, 40, 50, 60, 70, 80, 80, 100, 125, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150 cycles. In a further embodiment, the average signal per pixel obtained from the tandems stabilized by the staple molecules in the final sequencing cycle is 10%, 20%, 30%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50% higher than the average signal obtained from tandems never stabilized by the staple molecules. In one exemplary embodiment, the stabilized tandem sequence is sequenced with a read length greater than 150 cycles, and the average signal per pixel obtained from the tandem sequence stabilized by the stapled molecule is 50% higher in the final sequencing cycle compared to the average signal obtained from the tandem sequence stabilized by the stapled molecule. The read length can be defined by a signal-to-noise ratio below which the bases it no longer calls. For example, the read length can be a cycle at which the signal-to-noise ratio drops below 2, 3, or 4. Alternatively, the read length can be defined by a precision below which the bases it no longer calls. For example, the read length can be a cycle at which the base-calling precision drops below 99%, 99.5%, 99.9%, 99.95%, or 99.99%.
[0620] F. Load onto the array
[0621] One of the difficulties that plagues various sequencing methods involves efficiently and densely loading target analyte molecules (such as, for example, tandem or stabilized tandems) onto an array. Methods for loading target analyte molecules onto arrays are described, for example, in U.S. Patent Nos. 8,906,831 and 10,300,452, and all teachings in these references, for all purposes and particularly in relation to methods for loading analytes onto arrays, are incorporated herein by reference in their entirety.
[0622] As described above, tandem structures can be attached to the surface of a solid support. The solid support can be made of any of a variety of materials used for biochemical analysis. Particularly usable solid supports are particles, such as beads or microspheres. Populations of beads can be used to attach nucleic acid populations. The composition of the beads can vary depending on, for example, the format, chemistry, and / or method of attachment to be used. The geometry of the particles (such as beads or microspheres) can also correspond to a variety of different forms and shapes. In some embodiments of the sequencing method using stapled primers as described above, the beads can be arranged or otherwise spatially differentiated. Exemplary bead-based arrays that can be used include, but are not limited to, the BeadChip™ arrays available from Illumina, Inc. (San Diego, CA) or arrays such as those described in U.S. Patent Nos. 6,266,459; 6,355,431; 6,770,441; 6,859,570; or 7,622,294; or PCT Publication No. WO 00 / 63437, each of which is incorporated herein by reference. The beads can be located at discrete positions on a solid support, such as holes, with each position accommodating a single bead. Alternatively, the discrete positions where the beads reside can each comprise multiple beads, as described, for example, in U.S. Patent Application Publication Nos. 2004 / 0263923 A1, 2004 / 0233485 A1, 2004 / 0132205 A1, or 2004 / 0125424 A1, each of which is incorporated herein by reference.
[0623] In some embodiments of the sequencing method using stapled primers described above, the method can be performed in multiplex format, thereby allowing parallel detection of multiple different types of nucleic acids. Other types of arrays can be used instead of bead arrays, including, for example, those further detailed above. Although different types of nucleic acids can also be processed sequentially using one or more steps of the sequencing method using stapled primers, parallel processing can provide cost savings, time savings, and consistency of conditions. The arrays or methods of this disclosure can be configured to include at least 2, 10, 100, or 1 x 10^6 primers. 3 Seed, 1 x 10 4 Seed, 1 x 10 5 Seed, 1 x 106 Seed, 1 x 10 9 One or more different nucleic acids. Alternatively or additionally, the arrays or methods of this disclosure can be configured to include up to 1 x 103 9 Seed, 1 x 10 6 Seed, 1 x 10 5 Seed, 1 x 10 4 Seed, 1 x 10 3 There can be one, 100, 10, 2, or fewer different nucleic acids. Nucleic acids can attach to different sites in the array. Therefore, the number of sites in the array can be within the range exemplified here for different nucleic acids. Furthermore, the various reagents or products described herein (e.g., primer-template nucleic acid hybrids or stabilized ternary complexes) can be multiplexed to have different types or species within these ranges.
[0624] The density of primers of a certain type or population (or all primers) on the array can be, about, at least, at least about, at most, or at most about 1 x 10⁻⁶. 10 1, 2 x 10 10 1, 3 x 10 10 1, 4 x 10 10 5 x 10 10 6 x 10 10 7 x 10 10 8 x 10 10 9 x 10 10 1 x 10 11 1, 2 x 10 11 1, 3 x 10 11 1, 4 x 10 11 5 x 10 11 6 x 10 11 7 x 10 11 8 x 10 11 9 x 10 11 1 x 10 12 1, 2 x 10 12 1, 3 x 10 12 1, 4 x 10 12 5 x 10 12 6 x 10 12 7 x 10 12 8 x 10 12 9 x 10 12 1 x 10 13 1 piece, 2 x 10 13 1, 3 x 10 131, 4 x 10 13 5 x 10 13 6 x 10 13 7 x 10 13 8 x 10 13 9 x 10 13 1 x 10 14 1, 2 x 10 14 1, 3 x 10 14 1, 4 x 10 14 5 x 10 14 6 x 10 14 7 x 10 14 8 x 10 14 9 x 10 14 1 x 10 15 1, 2 x 10 15 1, 3 x 10 15 1, 4 x 10 15 5 x 10 15 6 x 10 15 7 x 10 15 8 x 10 15 9 x 10 15 1 x 10 16 1, 2 x 10 16 1, 3 x 10 16 1, 4 x 10 16 5 x 10 16 6 x 10 16 7 x 10 16 8 x 10 16 9 x 10 16 primers / m 2 , or a value or range between any two of these values.
[0625] This paper considers various separation distances or average separation distances between two adjacent primers of the same type or population (or two different types or populations). The spacing or average spacing between two adjacent primers of the same type or population (or two different types or populations) can be, about, at least, at least about, at most, or at most about 10 nm, 11 nm, 12 nm, 13 nm, 14 nm, 15 nm, 16 nm, 17 nm, 18 nm, 19 nm, 20 nm, 21 nm, 22 nm, 23 nm, 24 nm, 25 nm, 26 nm, 27 nm, 28 nm, 29 nm, 30 nm, 31 nm, 32 nm, 33 nm, 34 nm, 35 nm, 36 nm, 37 nm, 38 nm, 39 nm, 40 nm, 41 nm, 42 nm, 43 nm, 44 nm, 45 nm, 46 nm, 47 nm, 48 nm, 49 nm, 50 nm, 51 nm, 52 nm, 53 nm, 54 nm, 55 nm, 56 nm, 57 nm, 58 nm, 59 nm, 60 nm, 61 nm, or 10 nm. nm, 62 nm, 63 nm, 64nm, 65 nm, 66 nm, 67 nm, 68 nm, 69 nm, 70 nm, 71 nm, 72 nm, 73 nm, 74 nm, 75 nm, 76 nm, 77nm, 78 nm, 79 nm, 80 nm, 81 nm, 82 nm, 83 nm, 84 nm, 85 nm, 86 nm, 87 nm, 88 nm, 89 nm, 90nm, 91 nm, 92 nm, 93 nm, 94 nm, 95 nm, 96 nm, 97 nm, 98 nm, 99 nm, 100 nm, 110 nm, 120nm, 130 nm, 140 nm, 150 nm, 160 nm, 170 nm, 180 nm, 190 nm, 200 nm, 210 nm, 220 nm, 230nm, 240 nm, 250 nm, 260 nm, 270 nm, 280 nm, 290 nm, 300 nm, 310 nm, 320 nm, 330 nm, 340nm, 350 nm, 360 nm, 370 nm, 380 nm, 390 nm, 400 nm, 410 nm, 420 nm, 430 nm, 440 nm, 450nm, 460 nm, 470 nm, 480 nm, 490 nm, 500 nm, 510 nm, 520 nm, 530 nm, 540 nm, 550 nm, 560nm, 570 nm, 580 nm, 590 nm, 600 nm, 610nm, 620 nm, 630 nm, 640 nm, 650 nm, 660 nm, 670 nm, 680 nm, 690 nm, 700 nm, 710 nm, 720 nm, 730 nm, 740 nm, 750 nm, 760 nm, 770 nm, 780 nm, 790 nm, 800 nm, 810 nm, 820 nm, 830 nm, 840 nm, 850 nm, 860 nm, 870 nm, 880 nm, 890 nm, 900 nm, 910 nm, 920 nm, 930 nm, 940 nm, 950 nm, 960 nm, 970 nm, 980 nm, 990 nm, 1000 nm, or a value or range between any two of these values.
[0626] This disclosure considers various ratios of the number of primers of one type or population to the number of primers of another type or population. The ratio of the number of primers of one type or population to the number of primers of another type or population can be approximately, at least, at least about, at most, or at most about 1:100, 1:99, 1:98, 1:97, 1:96, 1:95, 1:94, 1:93, 1:92, 1:91, 1:90, 1:89, 1:88, 1:87, 1:86, 1:85, 1:84, 1:83, 1:82, 1:81, 1:80, 1:79, 1:78, 1:77, 1:76, 1:75, 1:74, 1:73, 1:72, 1:71, 1:70, 1:69, 1:68, 1:67, 1:66, 1:65, 1:64, 1:6 3, 1:62, 1:61, 1:60, 1:59, 1:58, 1:57, 1:56, 1:55, 1:54, 1:53, 1:52, 1:51, 1:50, 1:49, 1:48, 1:47, 1:46, 1:45, 1:44, 1:43, 1:42, 1:41, 1:40, 1:39, 1:38, 1:37, 1:36, 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:1 5, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:187:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, or a value or range between any two of these values.
[0627] The average distance between the positions on the solid support to which two adjacent or closest capture primers are attached can be greater than (or less than, or equal to) the length of one of the two capture primers, the length of the two capture primers, the average length of the two capture primers, or the total length of the two capture primers (or 0.1 x, 0.2 x, 0.3 x, 0.4 x, 0.5 x, 0.6 x, 0.7 x, 0.8 x, 0.9 x of the length). The average distance between the positions on the solid support to which two adjacent or closest amplification primers of a plurality of amplification primers are attached may be greater than (or less than, or equal to) the length of one of the two amplification primers, the length of the two amplification primers, the average length of the two amplification primers, or the total length of the two capture primers (or 0.1 x, 0.2 x, 0.3 x, 0.4 x, 0.5 x, 0.6 x, 0.7 x, 0.8 x, 0.9 x, 1.0 x, 1.1 x, 1.2 x, 1.3 x, 1.4 x, 1.5 x, 1.6 x, 1.7 x, 1.8 x, 1.9 x, 2 x, 3 x, 4 x, 5 x, 6 x, 7 x, 8 x, 9 x, 10 x).
[0628] In some embodiments, the tandem mass is synthesized via solid-state synthesis. As described above, solid-state synthesis of the tandem mass results in the formation of a tandem mass tethered to a surface (such as, for example, the surface of an array). The solid-state synthesis of the tandem mass can be performed according to any of the above configurations.
[0629] In some embodiments, the tandem bodies can be synthesized in solution, stabilized in solution using staple molecules, and then deposited onto a surface. Embodiments characterized by stabilizing the tandem bodies prior to deposition onto a surface are prone to tandem body aggregation. To overcome this potential limitation, these embodiments may further include providing a congesting agent, such as, for example, PEG, to inhibit the formation of aggregates of the (stabilized) tandem bodies. Any suitable congesting agent known in the art can be used. According to any of the above configurations, the stabilized tandem bodies synthesized via solution synthesis can be deposited onto the array.
[0630] In other embodiments, the tandem bodies are synthesized in solution, deposited onto a surface, and then stabilized using staple molecules. A key feature of embodiments where the tandem bodies are stabilized after deposition onto the surface is that aggregation of the tandem bodies is less likely to occur. However, these embodiments may further include providing a congesting agent, such as, for example, PEG, to inhibit the formation of aggregates of the (stabilized) tandem bodies. Any suitable congesting agent known in the art can be used. According to any of the above configurations, the tandem bodies synthesized via solution synthesis can be deposited onto an array.
[0631] Example
[0632] The following examples are provided to illustrate some embodiments of the invention more fully. However, they should not in any way be construed as limiting the broad scope of the invention. Those skilled in the art can readily devise many variations and modifications of the principles disclosed herein without departing from the scope of the invention.
[0633] A. Example 1 – Molecular Configuration of Staples
[0634] Various working examples of staple molecules have been designed to facilitate tandem stabilization and / or sequencing. Due to the microscopic nature of this invention, the working examples described in this section are shown in illustrative form, beginning with a general description of the stabilized tandem. Figure 1 A schematic diagram of a tandem mass on the surface of a flow cell and an enlarged cross-section of the tandem mass are shown, illustrating the stabilization of the tandem mass using various staple molecules combined with different instances of linker sequences of the tandem mass.
[0635] 1. Sequencing was performed using staple molecules as staple primers. :
[0636] like Figure 2 As shown, staple molecules can be designed such that the first and second regions of the staple molecule hybridize with different instances of the adaptor sequence of the tandem. Optionally, the staple molecule may further include a variable-length multinucleotide spacer between the first and second regions of the staple molecule. In this working example, the staple molecule is configured to hybridize with a single instance of the “A” adaptor sequence of the tandem. Neither the first nor the second region of the staple molecule contains a mismatched and blocking 3' end, thereby allowing the staple molecule to act as a staple primer to facilitate sequencing of the first portion of the stable tandem. It is also considered that in some cases, it may be advantageous to f...
Claims
1. A method for forming a stable series of interconnects, the method comprising: (a) Provide a plurality of tandems, wherein each tandem contains a plurality of instances of a target sequence and a plurality of instances of at least one interleaver sequence; (b) Contact the plurality of serial links with staple molecules to form a plurality of stable serial links; Each staple element contains at least a first region and a second region; and The first and second regions of the staple molecule each hybridize with different instances of the at least one adaptor sequence.
2. The method of claim 1, wherein the provision of the plurality of tandem sequences in step (a) comprises rolling circle amplification (RCA) with a strand displacement polymerase using primers hybridized to a circular nucleic acid template, wherein the circular nucleic acid template comprises the target sequence and the at least one adaptor sequence.
3. The method of claim 2, further comprising circularizing a linear nucleic acid template comprising the target sequence and the at least one adaptor sequence to generate the circular nucleic acid template.
4. The method according to claim 3, wherein the linear nucleic acid template comprises a first adaptor sequence of the target sequence 3' and a second adaptor sequence of the target sequence 5'.
5. The method of claim 4, wherein the first and second adaptor sequences of the linear nucleic acid template are ligated after hybridization with the splint oligonucleotide.
6. The method of claim 5, wherein the splint oligonucleotide is the primer that hybridizes with the circulated nucleic acid.
7. The method of claim 5, wherein the splint oligonucleotides are removed prior to the RCA.
8. The method of claim 2, wherein the primers are fixed to the surface during RCA.
9. The method of claim 2, wherein the primer is in solution during RCA.
10. The method of claim 9, further comprising depositing the plurality of stabilized tandem bodies on the surface after contact step (b).
11. The method of claim 10, wherein the surface is a structured surface.
12. The method of claim 9, further comprising depositing the plurality of tandem bodies on the surface prior to contact step (b).
13. The method of claim 12, wherein the surface is a structured surface.
14. The method of claim 2, wherein the RCA generates a sense strand, and the method further comprises amplifying the sense strand to generate a plurality of antisense strands.
15. The method of claim 14, wherein at least some of the staple molecules hybridize with the connector sequences of the plurality of antisense strands.
16. The method of claim 14, wherein step (b) further comprises hybridizing the staple molecule with the sense strand.
17. The method of claim 16, wherein at least some of the staple molecules are staple primers for amplifying the sense strand to generate the plurality of antisense strands.
18. The method of claim 2, wherein the contact step (b) occurs during the RCA step (a).
19. The method of claim 1, wherein each instance of the at least one adaptor sequence comprises a primer binding site.
20. The method of claim 19, wherein each instance of the at least one connector sequence further comprises a label region, optionally wherein the label region is a sample index.
21. The method of claim 19, wherein each instance of the at least one adaptor sequence further comprises a splint binding site.
22. The method of claim 19, wherein each instance of the at least one adaptor sequence further comprises a variable region, optionally wherein the variable region is a unique molecular identifier (UMI).
23. The method of claim 1, wherein each tandem strand comprises a sense strand that hybridizes with a plurality of antisense strands.
24. The method of claim 23, wherein the first region and the second region of the staple molecule each hybridize with different instances of the at least one adaptor sequence on different antisense strands.
25. The method of claim 23, wherein the sense strand does not contain uracil and the plurality of antisense strands contain uracil.
26. The method of claim 1, wherein all instances of the target sequence within each tandem body are identical.
27. The method of claim 1, wherein an instance of the target sequence comprises a sense sequence or an antisense sequence.
28. The method of claim 1, wherein the plurality of series bodies are provided fixed to the surface in the flow cell.
29. The method of claim 28, wherein the surface is a continuous surface.
30. The method of claim 28, wherein the plurality of tandem bodies are fixed to the binding sites on the structured surface of the flow cell.
31. The method of claim 1, wherein the plurality of series bodies are provided in solution in step (a) and deposited on the surface of the flow cell in step (b).
32. The method of claim 1, wherein the sequences of the first region and the second region are identical.
33. The method of claim 1, wherein at least one of the first and second regions of the staple molecule includes a 3' end that allows extension.
34. The method of claim 33, wherein the 3' end of the staple molecule hybridizes within 40 nucleotides from the 3' end of the target sequence.
35. The method of claim 1, wherein at least one of the first and second regions of the staple molecule comprises the 3' end of the staple molecule and hybridizes with the primer binding site of the at least one adaptor sequence.
36. The method of claim 1, wherein at least one of the first and second regions of the staple molecule comprises a 3' end that allows for the formation of a ternary complex.
37. The method of claim 36, wherein the 3' end that allows the formation of the ternary complex is reversibly terminated.
38. The method of claim 1, wherein neither the first region nor the second region of the staple molecule contains a 3' end that allows the ternary complex to form or extend.
39. The method of claim 1, wherein the plurality of tandem units comprises repeating sequence units, the repeating sequence units comprising the target sequence and the at least one linker sequence.
40. The method of claim 39, wherein the at least one connector sequence within the sequence unit comprises a first connector sequence of the target sequence 3' and a second connector sequence of the target sequence 5'.
41. The method of claim 40, wherein the first region of the staple molecule hybridizes with the first adaptor sequence and the second region of the staple molecule hybridizes with the second adaptor sequence.
42. The method of claim 40, wherein at least one of the first and second regions of the staple molecule comprises the 3' end of the staple molecule and hybridizes with a 3' adaptor.
43. The method of claim 1, wherein the length of each of the first and second regions of the staple molecule is at least 10 nucleotides.
44. The method of claim 1, wherein the first and second regions of the staple molecule each hybridize with different instances of the at least one adaptor sequence spaced at least 100 nucleotides apart.
45. The method of claim 1, wherein the 3' end of the staple molecule comprises at least one mismatch, a nucleotide incorporation blocker, or a ternary complex formation preventer cap.
46. The method of claim 45, wherein the 3' end of the staple molecule comprises a blocker that prevents nucleotide incorporation.
47. The method of claim 45, wherein the 3' end of the staple molecule prevents the formation of a ternary complex.
48. The method of claim 47, wherein the 3' end of the staple molecule is capped with a portion to prevent the formation of a ternary complex.
49. The method of claim 47, wherein the 3' end of the staple molecule comprises a mismatch and a blocking agent, optionally wherein the blocking agent is a 3' phosphate blocking agent.
50. The method of claim 1, wherein at least some of the first and second regions of the staple molecules are separated by spacers.
51. The method of claim 50, wherein at least some of the staple molecules do not contain spacers.
52. The method of claim 50, wherein the spacer comprises a polynucleotide sequence.
53. The method of claim 52, wherein the length of the polynucleotide sequence is variable.
54. The method of claim 52, wherein the polynucleotide sequence comprises a double-stranded DNA sequence.
55. The method of claim 54, wherein the first region and the second region each comprise a 3' end that hybridizes with the tandem.
56. The method of claim 50, wherein the spacers have variable lengths in different staple molecules.
57. The method of claim 50, wherein the spacer comprises a nonnucleotide polymer linker.
58. The method of claim 57, wherein the nonnucleotide polymer linker comprises polyethylene glycol (PEG).
59. The method of claim 57, wherein the nonnucleotide polymer linker comprises a dendritic macromolecule.
60. The method of claim 59, wherein the dendritic macromolecule comprises polyamidoamine (PAMAM).
61. The method of claim 59, wherein the staple molecule hybridizes with three or more instances of the at least one adaptor sequence.
62. The method of claim 61, wherein the staple molecule hybridizes with 10 or more instances of the at least one adaptor sequence.
63. The method of claim 61, wherein a majority of the staple molecule hybridizes with an adaptor sequence side-attached to at least 10 different target sequences.
64. The method of claim 59, wherein neither the first region nor the second region of the staple molecule contains a 3' end that allows the ternary complex to form or extend.
65. The method of claim 50, wherein the spacer is coupled to the 3' end of the first region.
66. The method of claim 65, wherein the spacer is coupled to the 5' end of the second region.
67. The method of claim 65, wherein the spacer is coupled to the 3' end of the second region and wherein the staple molecule cannot act as a primer.
68. The method of claim 1, wherein a plurality of staple molecules hybridize with the same tandem strand.
69. The method of claim 1, wherein providing the plurality of tandem bodies comprises forming the tandem bodies in the presence of a compaction additive.
70. The method of claim 69, wherein the compaction additive is a polymer.
71. The method of claim 70, wherein the polymer is a PAMAM dendrimer.
72. The method according to any one of claims 1 or 69, wherein the staple primer does not compact the tandem.
73. The method according to claim 1 or 72, wherein the first region of the staple primer is the 3' of the second region of the staple primer.
74. The method of claim 73, further comprising sequencing a portion of the tandem from the 3' end of the first region of the staple primer.
75. The method of claim 74, wherein the portion of the tandem is at least a first portion of the target sequence.
76. The method of claim 75, wherein the portion of the tandem is a tag region of the at least one connector sequence, optionally wherein the tag region is a sample index.
77. The method according to any one of claims 74 to 76, further comprising blocking the staple primer to prevent further sequencing from the staple primer.
78. The method of claim 77, wherein the blockade is performed using a ternary complex inhibitor.
79. The method of claim 77 or 78, further comprising hybridizing another primer with a region of the adapter sequence that has not hybridized with the staple primer, wherein the other primer is not a staple primer.
80. The method of claim 79, further comprising sequencing a second portion of the tandem from the other primer while the staple primer remains hybridized with the tandem.
81. The method according to any one of claims 74 to 80, comprising at least 100 sequencing cycles.
82. The method of claim 81, wherein after the at least 100 sequencing cycles, the full width at half maximum (FWHM) of the tandem does not increase by more than 5%.
83. The method of claim 82, wherein if the second region of the staple primer is absent, the full width at half maximum (FWHM) of the tandem does not increase by more than 5% after the at least 100 sequencing cycles.
84. The method of claim 83, wherein if the second region of the staple primer is absent, the full width at half maximum (FWHM) of the tandem does not increase by more than 10% after the at least 100 sequencing cycles.
85. The method according to any one of claims 73 to 84, wherein the tandem strand is a sense strand, and wherein providing a plurality of tandem strands includes generating a plurality of antisense strands from the tandem strands by multi-strand substitution (MDA), sequencing portions of the antisense strands, and digesting the antisense strands.
Citation Information
Patent Citations
Method and system for sequencing nucleic acids
US10246744B2
Methods and compositions for single molecule composition loading
US10300452B2
Method and system for sequencing nucleic acids
US10443098B2
Diffraction grating-based encoded micro-particles for multiplexed experiments
US20040125424A1
Method and apparatus for aligning microbeads in order to interrogate the same
US20040132205A1