Spatially resolved surface capture of nucleic acids
The method uses a support with a low non-specific binding coating and immobilized primers to capture and sequence nucleic acids, addressing the challenge of spatially resolved analysis by reducing assembly and computational requirements.
Patent Information
- Application Number
- JP2025508428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-15
- Filing Date
- 2023-08-15
- Publication Date
- 2025-09-02
AI Technical Summary
Existing methods for spatially resolved analysis of nucleic acids in biological samples lack efficient techniques for capturing and sequencing nucleic acids while preserving their spatial location information, leading to increased assembly and computational requirements.
A method involving a support with a low non-specific binding coating and immobilized surface capture primers, followed by reverse transcription, template switching, and rolling circle amplification to generate spatially resolved nucleic acid concatemers, allowing for sequencing and spatial positioning of nucleic acids on the support.
Enables efficient spatially resolved capture and sequencing of nucleic acids, reducing assembly and computational requirements while maintaining spatial location information.
Smart Images

Figure 2025528821000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 398,183, filed August 15, 2022, the contents of which are incorporated herein by reference in their entirety.
[0002] Electronic Sequence Listing Reference The contents of the Electronic Sequence Listing (ELEM_014_001WO_SeqList_ST26.xml, size: 9,927 bytes, and creation date: August 9, 2023) are incorporated herein by reference in their entirety.
[0003] The present disclosure provides compositions and methods for capturing nucleic acids from a cellular sample on a support, preparing library molecules on a support, amplifying the library molecules on a support to generate nucleic acid template molecules, and analyzing the immobilized nucleic acid template molecules (including detecting and / or sequencing the immobilized nucleic acid template molecules). The immobilized nucleic acid template molecules correspond to nucleic acids from the cellular sample. The immobilized nucleic acid template molecules are spatially located on the support in an arrangement similar to their spatial locations in the cellular sample. [Background technology]
[0004] Cells within a tissue of interest have differences in cell morphology and / or function due to various analyte levels (e.g., gene and / or protein expression) within different cells. The specific location of a cell within a tissue (e.g., the location of a cell relative to neighboring cells or the location of a cell relative to the tissue microenvironment) can affect, for example, cell morphology, differentiation, fate, viability, proliferation, behavior, and signaling and crosstalk with other cells within the tissue.
[0005] Spatial heterogeneity has previously been studied using techniques that only provide data for a few analytes in the context of intact tissue or tissue portions, or that provide many analyte data for single cells but do not provide information about the location of single cells in the parent biological sample (e.g., tissue sample).
[0006] Spatial analysis of analytes within a biological sample may require determining the sequence of the analyte sequence or its complement and the spatial barcode or its complement to identify the location of the analyte. When a biological sample is analyzed to identify or characterize an analyte, such as DNA, RNA, or other genetic material within the sample, it may be placed on a solid support to improve specificity and efficiency. There is a need for improved spatially resolved surface capture and downstream assembly and sequencing techniques for nucleic acids on a support that reduce the assembly, time, and / or computational requirements required to obtain spatial sequencing information. Compositions, devices, and methods that address this need are provided herein. Summary of the Invention
[0007] In one aspect, the disclosure provides a method for preparing spatially resolved nucleic acids, comprising: a) providing a support comprising: (i) a low non-specific binding coating comprising at least one hydrophilic polymer, wherein the low non-specific binding coating has a water contact angle of 45 degrees or less; and (ii) a plurality of immobilized surface capture primers covalently tethered to the low non-specific binding coating, wherein each individual surface capture primer comprises a universal surface capture primer sequence (110), a universal binding site for a reverse sequencing primer (120), and an RNA capture sequence; and b) positioning a cellular sample on the coated support under conditions suitable for the cellular sample to remain in a fixed position on the coated support, disrupting the cellular sample, and releasing a plurality of RNA molecules of the cellular sample from the cellular sample under conditions suitable for preserving spatial location information of the RNA molecules, and hybridizing the individual RNA molecules to individual immobilized surface capture primers to form a plurality of capture primer-RNA hybridization sequences. c) collecting a plurality of RNA molecules onto a plurality of immobilized surface capture primers on a support under conditions suitable for generating duplexes, each primer-RNA duplex comprising an immobilized surface capture primer hybridized to an RNA molecule; and c) performing a reverse transcription reaction on the coated support under conditions suitable for extending the 3' ends of the immobilized surface capture primers, using the hybridized RNA as a template strand, under conditions suitable for generating a non-template poly-C tail at the 3' end, thereby generating a plurality of first-strand cDNA molecules having a non-template poly-C tail at the 3' end, the reverse transcription reaction comprising a plurality of template switching oligonucleotides, each of the template switching oligonucleotides comprising a poly-G region capable of hybridizing to the non-template poly-C tail of a first-strand cDNA molecule, a universal binding site for a forward sequencing primer (130), and a universal binding site for a surface pinning primer (140).d) performing a primer extension reaction from the 3' ends of the non-template poly-C tails of the plurality of first-strand cDNA molecules using the template switching oligonucleotide as a template strand, thereby generating a plurality of immobilized full-length first-strand cDNAs, each full-length first-strand cDNA comprising a universal surface capture primer sequence (110), a universal binding site for a reverse sequencing primer (120), an RNA capture sequence, a cDNA insert region, a poly-C tail region, a universal binding site for a forward sequencing primer (130), and a universal binding site for a surface pinning primer (140); and e) removing the RNA molecules while retaining the plurality of immobilized full-length first-strand cDNAs. f) contacting the retained plurality of immobilized full-length first-strand cDNAs with a plurality of single-stranded circularization oligonucleotides under conditions suitable for hybridizing each circularization oligonucleotide to the immobilized full-length first-strand cDNAs, thereby forming gapped single-stranded circular molecules, wherein each single-stranded circularization oligonucleotide comprises: (i) a sequence complementary to the universal binding site (120) for the reverse sequencing primer; (ii) a sequence complementary to the universal surface capture primer sequence (110); (iii) a linker region; (iv) a sequence complementary to the universal binding site (140) for the surface pinning primer; and (v) a sequence complementary to the universal binding site (130) for the forward sequencing primer; and g) performing a polymerase-catalyzed extension reaction to fill the gaps using the poly-C tail region, the cDNA insert region, and the RNA capture sequence of the immobilized full-length first-strand cDNAs as template strands, wherein the polymerase-catalyzed extension reaction is performed to fill the gaps using the poly-C tail region, the cDNA insert region, and the RNA capture sequence of the immobilized full-length first-strand cDNAs as template strands. h) performing a rolling circle amplification reaction using the 3' end of the immobilized full-length first-strand cDNA as an initiation site and the covalently closed circular molecule as a template strand, thereby producing a plurality of immobilized nucleic acid concatemers that are spatially resolved on the support; i) sequencing the plurality of individual immobilized nucleic acid concatemers, wherein the sequencing comprises at least a portion of the cDNA insert region of each individual nucleic acid concatemer that corresponds to an individual RNA molecule eluted from the cell sample; and j) determining a position of each individual nucleic acid concatemer on the coated support that corresponds to the spatial position of each RNA molecule eluted from the cell sample.
[0008] In some embodiments, the cell sample of step b) comprises a single cell, multiple cells, a tissue, an organ, an organism, or a sectioned cell sample, hi some embodiments, the cell sample of step b) comprises a fresh sample, a frozen sample, a fresh frozen sample, or a formalin-fixed, paraffin-embedded sample.
[0009] In some embodiments, the plurality of template switching oligonucleotides in step c) comprises chimeric DNA and / or RNA oligonucleotides.
[0010] In some embodiments, the linker region of the single-stranded circularization oligonucleotide of step f) comprises at least one sample index sequence for multiplexing, at least one unique molecular index (UMI) sequence for molecular tagging, and / or at least one universal binding site for a compaction oligonucleotide.
[0011] In some embodiments, the rolling circle amplification reaction of step h) is carried out in the presence of a plurality of compaction oligonucleotides, each comprising a single-stranded oligonucleotide having a first region at a first distal end that hybridizes to one portion of a concatemeric molecule and a second region at a second distal end that hybridizes to a second portion of the same concatemeric molecule, bringing the first portion of the concatemeric molecule and the second portion of the concatemeric molecule into close proximity, whereby compaction of the concatemeric molecules forms DNA nanoballs.
[0012] In some embodiments, the sequencing in step i) comprises: a. contacting a plurality of concatemer molecules with a plurality of sequencing polymerases and a plurality of nucleic acid sequencing primers, wherein the contacting is performed under conditions suitable to form a plurality of hybrid sequencing polymerases, each hybrid sequencing polymerase comprising a sequencing polymerase bound to a nucleic acid duplex, wherein the nucleic acid duplex comprises a portion of the concatemer molecules hybridized to the nucleic acid sequencing primer; and b. contacting the plurality of hybrid sequencing polymerases with a plurality of detectably labeled nucleotides comprising blocking moieties at the 2' or 3' sugar position. wherein the contacting is performed under conditions suitable for binding of the at least one nucleotide to at least one of the multiplexed sequencing polymerases, the conditions being suitable for promoting polymerase-catalyzed nucleotide incorporation; c. incorporating the nucleotide into the 3' end of the sequencing primer of the at least one multiplexed sequencing polymerase; d. detecting the incorporated nucleotide and identifying the nucleobase of the incorporated nucleotide; e. removing the blocking moiety from the incorporated nucleotide; and f. repeating steps b. to e. at least once.
[0013] In some embodiments, the sequencing in step i) comprises: a. contacting a plurality of concatemer molecules with a plurality of sequencing polymerases and a plurality of nucleic acid sequencing primers, wherein the contacting is performed under conditions suitable to form a plurality of hybrid sequencing polymerases, each hybrid sequencing polymerase comprising a sequencing polymerase bound to a nucleic acid duplex, the nucleic acid duplex comprising a portion of the concatemer molecules hybridized to the nucleic acid sequencing primer; and b. contacting the plurality of hybrid sequencing polymerases with a detectable label attached to a phosphate moiety of the phosphate strand, each hybrid sequencing polymerase comprising a detectable label attached to a phosphate moiety of the phosphate strand. a. contacting a sequence primer with a plurality of nucleotides comprising a nucleotide sequence comprising at least one nucleotide selected from the group consisting of nucleotides 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 110, 111, 112, 113, 114, 115, 116, 117, 118, 120, 121,
[0014] In some embodiments, the sequencing in step i) comprises: a. contacting the plurality of concatemer molecules with a plurality of first sequencing polymerases and a plurality of nucleic acid sequencing primers, wherein the contacting is performed under conditions suitable for binding the plurality of first polymerases to the plurality of nucleic acid template molecules and the plurality of nucleic acid primers, thereby forming a plurality of first multiplexed polymerases, each first multiplexed polymerase comprising the first polymerase bound to a nucleic acid duplex, and the nucleic acid duplex comprising the nucleic acid template molecule hybridized to the nucleic acid primer; and b. contacting the plurality of first multiplexed polymerases with a plurality of detectably labeled multivalent molecules to form a plurality of multivalent binding complexes. forming a plurality of detectably labeled multivalent molecules, each detectably labeled multivalent molecule comprising a core attached to a plurality of nucleotide arms, each nucleotide arm being attached to a nucleotide unit, and contacting the plurality of detectably labeled multivalent molecules under conditions suitable for binding complementary nucleotide units of the multivalent molecule to at least two of the plurality of first multiplexed polymerases, thereby forming a plurality of multivalent binding complexes, wherein the conditions are suitable for inhibiting incorporation of the complementary nucleotide units into primers of the plurality of multivalent binding complexes; c. detecting the plurality of multivalent binding complexes; and d. identifying the nucleobases of the complementary nucleotide units in the plurality of multivalent binding complexes, thereby determining the sequence of the nucleic acid template molecule.
[0015] In some embodiments, the sequencing in step i) comprises: a. dissociating the plurality of multivalent binding complexes by removing the plurality of first sequencing polymerases and their bound multivalent molecules, retaining a plurality of nucleic acid duplexes; b. contacting the retained plurality of nucleic acid duplexes of step e) with a plurality of second sequencing polymerases under conditions suitable for binding the plurality of second polymerases to the retained plurality of nucleic acid duplexes, thereby forming a plurality of second multiplexed polymerases, each comprising a second polymerase bound to a nucleic acid duplex; and c. contacting the plurality of second multiplexed polymerases with a plurality of nucleotides, wherein the contacting is performed under conditions suitable for binding complementary nucleotides from the plurality of nucleotides to at least two of the second multiplexed polymerases, thereby forming a plurality of nucleotide binding complexes, wherein the conditions are suitable for promoting nucleotide incorporation of the bound complementary nucleotides into primers of the nucleotide binding complexes.
[0016] In some embodiments, the method further comprises detecting a complementary nucleotide incorporated into the primer for the nucleotide-complexed polymerase. In some embodiments, the method further comprises: d. detecting a complementary nucleotide incorporated into the primer for the nucleotide-complexed polymerase; and e. identifying the nucleobase of the complementary nucleotide incorporated into the primer for the nucleotide-complexed polymerase.
[0017] In some embodiments, contacting the plurality of first complexed polymerases with the plurality of multivalent molecules in step b. is performed in the presence of a non-catalytic divalent cation that inhibits polymerase-catalyzed nucleotide incorporation, wherein the non-catalytic divalent cation comprises strontium, barium, or calcium.
[0018] In some embodiments, each multivalent molecule in the plurality of multivalent molecules comprises (a) a core and (b) a plurality of nucleotide arms, wherein the plurality of nucleotide arms comprises (i) a core attachment moiety, (ii) a spacer, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms via their core attachment moieties, the spacer is attached to the linker, and the linker is attached to the nucleotide unit.
[0019] In some embodiments, the linker comprises an aliphatic chain having 2 to 6 subunits or an oligoethylene glycol chain having 2 to 6 subunits.
[0020] In some embodiments, multiple nucleotide arms attached to a given core have the same type of nucleotide unit, where the type of nucleotide unit comprises dATP, dGTP, dCTP, dTTP, or dUTP.
[0021] In some embodiments, the plurality of multivalent molecules comprises one type of multivalent molecule, where each multivalent molecule in the plurality of multivalent molecules has the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. In some embodiments, the plurality of multivalent molecules comprises a mixture of any combination of two or more types of multivalent molecules, where each type has a nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
[0022] In some embodiments, the method comprises forming a plurality of binding complexes: a) binding a first nucleic acid sequencing primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemeric molecule, thereby forming a first binding complex, wherein a first nucleotide unit of the first multivalent molecule binds to the first polymerase; and b) binding a second nucleic acid sequencing primer, a second sequencing polymerase, and the first multivalent molecule to a second portion of the same concatemeric molecule, thereby forming a second binding complex, wherein a second nucleotide unit of the first multivalent molecule binds to the second polymerase; The method further includes forming a plurality of binding complexes, wherein the first binding complex and the second binding complex comprise the same multivalent molecule, forming an avidity complex.
[0023] In some embodiments, the method comprises forming an avidity complex comprising: a) contacting a plurality of first sequencing polymerases and a plurality of nucleic acid sequencing primers with different portions of a concatemeric nucleic acid template molecule to form at least a first multiplexed polymerase and a second multiplexed polymerase on the same concatemeric molecule; and b) contacting a plurality of detectably labeled multivalent molecules with at least the first multiplexed polymerase and the second multiplexed polymerase on the same concatemeric molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first multiplexed polymerase and the second multiplexed polymerase, wherein at least a first nucleotide unit of the single multivalent molecule binds to the first multiplexed polymerase comprising a first primer hybridized to a first portion of the concatemeric molecule, thereby forming a first binding complex and at least a second nucleotide unit of the single multivalent molecule. a) contacting the first and second nucleotide units in the first and second binding complexes under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second binding complexes, whereby the first and second binding complexes bound to the same multivalent molecule form an avidity complex; b) detecting the first and second binding complexes on the same concatemeric molecule; and c) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the concatemeric molecule, and identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the concatemeric molecule.
[0024] In some embodiments, contacting the plurality of second hybrid polymerases with the plurality of nucleotides in step c. is performed in the presence of a catalytic divalent cation that promotes polymerase-catalyzed nucleotide incorporation, wherein the catalytic divalent cation comprises magnesium or manganese.
[0025] In some embodiments, each nucleotide in the plurality of nucleotides of step c. comprises an aromatic base, a 5-carbon sugar, and 1 to 10 phosphate groups. In some embodiments, the plurality of nucleotides of step c. comprises one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP, or a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
[0026] In some embodiments, at least one of the nucleotides in the plurality of nucleotides in step c. is labeled with a fluorophore. In some embodiments, the plurality of nucleotides in step c. lacks a fluorophore label.
[0027] In some embodiments, at least one of the nucleotides in the plurality of nucleotides of step c. comprises a removable chain-terminating moiety attached to the 3' carbon position of the sugar group, wherein the removable chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an azido group, an O-azidomethyl group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group, wherein the removable chain-terminating moiety is cleavable with a chemical compound to generate an extendable 3' OH moiety on the sugar group.
[0028] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief explanation of the drawings]
[0029] [Figure 1]1 is a schematic diagram of an exemplary low-binding support comprising alternating layers of a glass substrate and a hydrophilic coating covalently or non-covalently adhered to the glass, further comprising chemically reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides). In alternative embodiments, the support can be made from any material, such as glass, plastic, or a polymeric material. [Figure 2] Schematic diagrams of various exemplary configurations of multivalent molecules. Left (Class I): Schematic diagram of a multivalent molecule with a "starburst" or "helter-skelter" configuration. Center (Class II): Schematic diagram of a multivalent molecule with a dendrimer configuration. Right (Class III): Schematic diagram of multiple multivalent molecules formed by reacting streptavidin with a 4-arm or 8-arm PEG-NHS bearing biotin and dNTPs. Nucleotide units are represented as "N," biotin is represented as "B," and streptavidin is represented as "SA." [Figure 3] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a generic core attached to multiple nucleotide arms. [Figure 4] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms. [Figure 5] 1 shows a schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, the nucleotide arms comprising biotin, spacers, linkers, and nucleotide units. [Figure 6] FIG. 1 is a schematic diagram of an exemplary nucleotide arm comprising a core attachment moiety, a spacer, a linker, and a nucleotide unit. [Figure 7] The chemical structures of an exemplary spacer (top) and various exemplary linkers (bottom) are shown, including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker. [Figure 8] 1 shows the chemical structures of various exemplary linkers, including linkers 1-9. [Figure 9A]1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 9B] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 9C] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 9D] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 10] 1 shows the chemical structure of an exemplary biotinylated nucleotide arm. In this example, the nucleotide unit is connected to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base. [Figure 11] FIG. 1 is a schematic diagram of a support containing immobilized surface capture primers. [Figure 12] FIG. 1 is a schematic diagram showing a cell sample positioned on a support and RNA molecules eluted from the cell sample on a surface capture primer immobilized on the support. [Figure 13] FIG. 1 is a schematic diagram showing the synthesis of first strand cDNA (dashed line) and a non-templated poly-C tail at the 3′ end of the first strand cDNA. [Figure 14] 1 is a schematic diagram showing a template switching oligonucleotide hybridized to a non-template poly-C tail at the 3' end of a first-strand cDNA. rGrGrG represents an exemplary poly-G region. The poly-G region of the template switching oligonucleotide may comprise ribonucleotide bases. [Figure 15] FIG. 1 is a schematic diagram showing the extension of first-strand cDNA (upper dashed arrow) using a template switching oligonucleotide as the template strand to generate a full-length first-strand cDNA. [Figure 16] FIG. 1 is a schematic showing that the RNA strand is removed / degraded while retaining the full-length first strand cDNA immobilized on the support. [Figure 17]FIG. 1 is a schematic diagram showing hybridization of a circularization oligonucleotide to immobilized full-length first strand cDNA to form a gapped single-stranded circular molecule. [Figure 18] 1 is a schematic diagram showing a polymerase-catalyzed extension reaction to fill the gap resulting in a nick. A ligation reaction can be performed to ligate the nick. [Figure 19] FIG. 1 is a schematic diagram showing the initial steps of a rolling circle amplification reaction on a support. [Figure 20] 1 is a schematic diagram of a G-quadruplex (e.g., a G-quadruplex). [Figure 21] FIG. 1 is a schematic diagram of an exemplary intramolecular G-quadruplex structure. DETAILED DESCRIPTION OF THE INVENTION
[0030] Definition: The headings provided herein are not limitations of various aspects of the disclosure, which aspects can be understood by reference to the specification as a whole.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art. Generally, terms relating to molecular biology, nucleic acid chemistry, protein chemistry, genetics, microbiology, transgenic cell production, and hybridization techniques described herein are well known and commonly used in the art. The techniques and procedures described herein are generally performed according to conventional methods well known in the art and as described in various general and more specific references cited and discussed throughout the specification. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY 2000). See also Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992). The nomenclature used in connection with, and the experimental procedures and techniques described herein are well known and commonly used in the art.
[0032] Unless otherwise required by context herein, singular terms include pluralities and plural terms include the singular. The singular forms "a," "an," and "the," as well as use of the singular form of any word, include plural references unless expressly and unambiguously limited to a reference to one.
[0033] The use of alternative terms (eg, "or") is understood to mean either one or both of the alternatives, or any combination thereof.
[0034] As used herein, the term "and / or" should be understood to mean a specific disclosure that each of the specified features or components may or may not have the other. For example, when used herein in phrases such as "A and / or B," the term "and / or" is intended to include "A and B," "A or B," "A" (A alone), and "B" (B alone). In a similar manner, when used in phrases such as "A, B, and / or C," the term "and / or" is intended to encompass each of the following embodiments: "A, B, and C," "A, B, or C," "A or C," "A or B," "B or C," "A and B," "B and C," "A and C," "A" (alone), "B" (alone), and "C" (alone).
[0035] As used in this specification and the appended claims, the terms "comprising," "including," "having," and "containing," and grammatical variations thereof, as used herein, are intended to be open-ended so that one or more items in a list do not exclude other items that may be substituted for or added to the listed items. Wherever an embodiment is described herein with the term "comprising," it is understood that alternative, similar embodiments described with the terms "consisting of" and / or "consisting essentially of" are also provided.
[0036] As used herein, the term "about" or "approximately" refers to a value or composition that is within an acceptable error range for a particular value or composition, as determined by one of ordinary skill in the art. The acceptable error range depends in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" or "approximately" can mean within one or more standard deviations per practice in the art. Alternatively, "about" or "approximately" can mean a range of up to 10% (i.e., ±10%) or more, depending on the limitations of the measurement system. For example, about 5 mg can include any number between 4.5 mg and 5.5 mg. Furthermore, particularly with respect to biological systems or processes, the term can mean up to one order of magnitude or up to five times the value. When a particular value or composition is provided in this disclosure, unless otherwise specified, the meaning of "about" or "approximately" should be considered to be within an acceptable error range for that particular value or composition. Additionally, when ranges and / or subranges of values are provided, the ranges and / or subranges can include the endpoints of the ranges and / or subranges.
[0037] As used herein, the term "polymerase" and variations thereof include enzymes that contain a nucleotide (or nucleoside)-binding domain, and the polymerase can form a complex with a template nucleic acid and a complementary nucleotide. A polymerase can have one or more activities, including, but not limited to, base analog detection activity, DNA polymerization activity, reverse transcriptase activity, DNA binding, strand displacement activity, and nucleotide binding and recognition. A polymerase can be any enzyme that can catalyze the polymerization of nucleotides (including their analogs) into a nucleic acid strand. Typically, but not necessarily, such nucleotide polymerization can occur in a template-dependent manner. Typically, a polymerase contains one or more active sites, and nucleotide binding and / or catalysis of nucleotide polymerization can occur in one or more active sites. In some embodiments, a polymerase includes other enzymatic activities, such as 3' to 5' exonuclease activity or 5' to 3' exonuclease activity. In some embodiments, a polymerase has strand displacement activity. Polymerases can include, but are not limited to, naturally occurring polymerases and any subunits and truncations thereof, mutant polymerases, variant polymerases, recombinant, fused, or otherwise engineered polymerases, chemically modified polymerases, synthetic molecules or assemblies, and any analogs, derivatives, or fragments thereof (e.g., catalytically active fragments) that retain the ability to catalyze nucleotide polymerization. Polymerases include catalytically inactive polymerases, catalytically active polymerases, reverse transcriptases, and other enzymes that contain a nucleotide-binding domain. In some embodiments, polymerases may be isolated from cells or produced using recombinant DNA technology or chemical synthesis methods. In some embodiments, polymerases may be expressed in prokaryotic, eukaryotic, viral, or phage organisms. In some embodiments, polymerases may be post-translationally modified proteins or fragments thereof. Polymerases may be derived from prokaryotic, eukaryotic, viral, or phage organisms.Polymerases include DNA-directed DNA polymerases and RNA-directed DNA polymerases.
[0038] As used herein, the term "strand displacement" refers to the ability of a polymerase to locally separate strands of double-stranded nucleic acid and synthesize a new strand in a template-based manner. Strand-displacing polymerases displace a complementary strand from the template strand and catalyze new strand synthesis. Strand-displacing polymerases include mesophilic and thermophilic polymerases. Strand-displacing polymerases include wild-type enzymes and variants, including exonuclease-minus mutants, mutated versions, chimeric enzymes, and truncated enzymes. Examples of strand-displacing polymerases include, but are not limited to, phi29 DNA polymerase, large fragment of Bst DNA polymerase, large fragment of Bsu DNA polymerase (exo-), Bca DNA polymerase (exo-), Klenow fragment of E. coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV viral reverse transcriptase, Deep Vent DNA polymerase, and KOD DNA polymerase. The phi29 DNA polymerase can be a wild-type phi29 DNA polymerase (e.g., MagniPhi from Expedeon), or a variant EquiPhi29 DNA polymerase (e.g., from Thermo Fisher Scientific), or a chimeric QualiPhi DNA polymerase (e.g., from 4basebio).
[0039] As used herein, the terms "nucleic acid," "polynucleotide," and "oligonucleotide," as well as other related terms, are used interchangeably and refer to a polymer of nucleotides and are not limited to any particular length. Nucleic acids include recombinant or chemically synthesized forms. Nucleic acids can be isolated. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of DNA or RNA produced using nucleotide analogs (e.g., peptide nucleic acid (PNA) and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids comprise polymers of nucleotides, and the nucleotides can contain natural or non-natural bases and / or sugars. Nucleic acids contain naturally occurring internucleoside linkages, such as phosphodiester linkages. Nucleic acids can lack phosphate groups. Nucleic acids may contain non-natural internucleoside linkages, including phosphorothioate, phosphorothiolate, and / or peptide nucleic acid (PNA) linkages. In some embodiments, the nucleic acid comprises one type of polynucleotide or a mixture of two or more different types of polynucleotides.
[0040] As used herein, the terms "operably linked" and "operably linked," or related terms, refer to the juxtaposition of components. Juxtaposed components can be covalently linked together. For example, two nucleic acid components can be enzymatically ligated together, where the bond linking the two components together comprises a phosphodiester bond. A first and second nucleic acid component can be linked together, where the first nucleic acid component can confer a function to the second nucleic acid component. For example, the bond between a primer binding sequence and a sequence of interest forms a nucleic acid library molecule having a portion capable of binding to the primer. In another example, a transgene (e.g., a nucleic acid encoding a polypeptide or nucleic acid sequence of interest) can be ligated into a vector, where the bond allows for expression or function of the transgene sequence contained within the vector. In some embodiments, the transgene is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects expression of the transgene. In some embodiments, the vector comprises at least one host cell regulatory sequence, including a promoter sequence, an enhancer, a transcriptional and / or translational initiation sequence, a transcriptional and / or translational termination sequence, a polypeptide secretion signal sequence, etc. In some embodiments, the host cell regulatory sequence controls the level, timing, and / or location of expression of the transgene.
[0041] The terms "linked," "coupled," "attached," and "appended," and variations thereof, include any type of fusion, bond, adhesion, or association between any combination of compounds or molecules that is stable enough to withstand use in a particular procedure. Procedures can include, but are not limited to, nucleotide binding, nucleotide incorporation, deblocking (e.g., removal of chain-terminating moieties), washing, removal, flow, detection, imaging, and / or identification. Such binding can include, for example, covalent, ionic, hydrogen, dipole-dipole, hydrophilic, hydrophobic, or affinity binding, bonds or associations involving van der Waals forces, and mechanical binding. In some embodiments, such binding occurs intramolecularly, for example, by joining the ends of a single- or double-stranded linear nucleic acid molecule together to form a circular molecule. In some embodiments, such binding can occur between different molecular combinations or between molecules and non-molecules, including, but not limited to, binding of a nucleic acid molecule to a solid surface, binding of a protein to a detectable reporter moiety, binding of a nucleotide to a detectable reporter moiety, and the like. Some examples of conjugation can be found, for example, in Hermanson, G., "Bioconjugate Techniques", Second Edition (2008), Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998), Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998).
[0042] As used herein, the term "primer" and related terms refer to an oligonucleotide capable of hybridizing to a DNA and / or RNA polynucleotide template to form a duplex molecule. Primers contain natural nucleotides and / or nucleotide analogs. Primers can be recombinant nucleic acid molecules. Primers can be of any length, but typically range from 4 to 50 nucleotides. A typical primer contains a 5' end and a 3' end. The 3' end of a primer can contain a 3' OH moiety that functions as a nucleotide polymerization initiation site in a polymerase-catalyzed primer extension reaction. Alternatively, the 3' end of a primer can lack a 3' OH moiety or can contain a terminal 3' blocking group that inhibits nucleotide polymerization in a polymerase-catalyzed reaction. Any one or more nucleotides along the length of a primer can be labeled with a detectable reporter moiety. Primers can be in solution (e.g., soluble primers) or immobilized on a support (e.g., capture primers).
[0043] The terms "template nucleic acid," "template polynucleotide," "target nucleic acid," "target polynucleotide," "template strand," and other variations thereof, refer to a nucleic acid strand that serves as the base nucleic acid molecule for any of the iterative sequencing methods described herein. A template nucleic acid may be single-stranded or double-stranded, or it may have single-stranded or double-stranded portions. A template nucleic acid may be obtained from a naturally occurring source or recombinantly, or may be chemically synthesized to contain any type of nucleic acid analog. A template nucleic acid may be linear, circular, or in other forms. A template nucleic acid may include an insert having an insert sequence. A template nucleic acid may also include at least one adapter sequence. An insert may be isolated in any form, including a chromosome, a genome, an organelle (e.g., a mitochondrion, a chloroplast, or a ribosome), a recombinant molecule, cloned, amplified, cDNA, RNA such as precursor mRNA or mRNA, an oligonucleotide, total genomic DNA obtained from fresh-frozen paraffin-embedded tissue, a needle biopsy, circulating tumor cells, cell-free circulating DNA, or any type of nucleic acid library. The insert may be isolated from any source, including organisms such as prokaryotes, eukaryotes (e.g., human, plant, and animal), fungi, viral cells, tissues, normal or diseased cells or tissues, bodily fluids including blood, urine, serum, lymph, tumors, saliva, anal and vaginal secretions, amniotic fluid samples, sweat, and semen, environmental samples, culture samples, or synthetic nucleic acid molecules prepared using recombinant molecular biology or chemical synthesis methods. The insert may be isolated from any organ, including the head, neck, brain, breast, ovaries, cervix, colon, rectum, endometrium, gallbladder, intestine, bladder, prostate, testes, liver, lung, kidney, esophagus, pancreas, thyroid, pituitary, thymus, skin, heart, larynx, or other organs. The template nucleic acid may be subjected to nucleic acid analysis, including sequencing and compositional analysis.
[0044] The term "adapter" and related terms refer to an oligonucleotide capable of operably binding to a target polynucleotide, where the adapter confers a function to the co-ligated adapter-target molecule. Adapters include DNA, RNA, chimeric DNA / RNA, or analogs thereof. Adapters can contain at least one ribonucleoside residue. Adapters can be single-stranded or double-stranded, or can have single-stranded and / or double-stranded portions. Adapters can be configured to be linear, stem-loop, hairpin, or Y-shaped. Adapters can be any length, from 4 to 100 or more nucleotides. Adapters can have blunt ends, overhanging ends, or a combination of both. Overhanging ends include 5' overhangs and 3' overhanging ends. The 5' end of a single-stranded adapter, or one strand of a double-stranded adapter, can have a 5' phosphate group or lack a 5' phosphate group. The adapter may include a 5' tail that does not hybridize to the target polynucleotide (e.g., a tailed adapter), or the adapter may be tailless. The adapter may include a sequence complementary to at least a portion of a primer, such as an amplification primer, a sequencing primer, or a capture primer (e.g., a soluble or immobilized capture primer). The adapter may include a random or degenerate sequence. The adapter may include at least one inosine residue. The adapter may include at least one phosphorothioate, phosphorothiolate, and / or phosphoramidate linkage. The adapter may include a barcode sequence, which can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in multiplex assays. The adapter may include a unique identification sequence (e.g., a unique molecular index, UMI, or unique molecular tag), which can be used to uniquely identify the nucleic acid molecule to which the adapter is attached.In some embodiments, the unique identification sequence can be used to increase error correction and accuracy, reduce the rate of false-positive variant calls, and / or increase the sensitivity of variant detection. The adapter can comprise at least one restriction enzyme recognition sequence, wherein the at least one restriction enzyme recognition sequence comprises any one or any combination of two or more selected from the group consisting of Type I, Type II, Type III, Type IV, Hs, or Type IIB.
[0045] In some embodiments, any of the amplification primer sequence, sequencing primer sequence, capture primer sequence, target capture sequence, circularization anchor sequence, sample barcode sequence, spatial barcode sequence, or anchor region sequence can be about 3 to 50 nucleotides in length, or about 5 to 40 nucleotides in length, or about 5 to 25 nucleotides in length.
[0046] The term "universal sequence" and related terms refer to a sequence in a nucleic acid molecule that is common between two or more polynucleotide molecules. For example, an adapter having a universal sequence can be operably linked to multiple polynucleotides, such that a population of co-linked molecules possesses the same universal adapter sequence. Examples of universal adapter sequences include amplification primer sequences, sequencing primer sequences, or capture primer sequences (e.g., soluble or immobilized capture primers).
[0047] When used in reference to nucleic acid molecules, the terms "hybridize" or "hybridizing" or "hybridization," or other related terms, refer to hydrogen bonding between two different nucleic acids to form a double-stranded nucleic acid. Hybridization also includes hydrogen bonding between two different regions of a single nucleic acid molecule to form a self-hybridizing molecule having a double-stranded region. Hybridization can involve Watson-Crick or Hoogsteen binding to form a double-stranded double-stranded nucleic acid or a double-stranded region within a nucleic acid molecule. The double-stranded nucleic acid, or two different regions of a single nucleic acid, can be fully complementary or partially complementary. Complementary nucleic acid strands need not hybridize to each other throughout their entire length. Complementary base pairing can be standard AT or CG base pairing, or other forms of base pairing interactions. Double-stranded nucleic acids can contain mismatched base-pairing nucleotides.
[0048] When used with respect to nucleic acids, the terms "extend," "extending," "extension," and other variations refer to the incorporation of one or more nucleotides into a nucleic acid molecule. Nucleotide incorporation involves the polymerization of one or more nucleotides onto the terminal 3'OH terminus of a nucleic acid chain, resulting in the elongation of the nucleic acid chain. Nucleotide incorporation can be performed with natural nucleotides and / or nucleotide analogs. Typically, although not necessarily, nucleotide incorporation occurs in a template-dependent manner. Any suitable method for extending a nucleic acid molecule may be used, including primer extension catalyzed by DNA polymerase or RNA polymerase.
[0049] The term "nucleotide" and related terms refer to a molecule comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and at least one phosphate group. Standard or non-standard nucleotides are consistent with the use of this term. In some embodiments, the phosphate comprises a monophosphate, diphosphate, or triphosphate, or a corresponding phosphate analog. The term "nucleoside" refers to a molecule comprising an aromatic base and a sugar. Nucleotides and nucleosides can be unlabeled or labeled with a detectable reporter moiety.
[0050] Nucleotides (and nucleosides) typically contain heterocyclic bases containing a substituted or unsubstituted nitrogen-containing parent heteroaromatic ring, which are commonly found in nucleic acids, including naturally occurring, substituted, modified, or engineered variants, or analogs thereof. The base of a nucleotide (or nucleoside) is capable of forming Watson-Crick and / or Hoogsteen hydrogen bonds with an appropriate complementary base. Exemplary bases are purines and pyrimidines, such as 2-aminopurine, 2,6-diaminopurine, adenine (A), ethenoadenine, N, N-acetylglucosamine ... 6 -Δ 2 -Isopentenyladenine (6iA), N 6 -Δ 2 -Isopentenyl-2-methylthioadenine (2ms6iA), N 6 -Methyladenine, guanine (G), isoguanine, N 2 -dimethylguanine (dmG), 7-methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine, and O 6 -methylguanine; 7-deaza-purines, such as 7-deazaadenine (7-deaza-A) and 7-deazaguanine (7-deaza-G); pyrimidines, such as cytosine (C), 5-propynylcytosine, isocytosine, thymine (T), 4-thiothymine (4sT), 5,6-dihydrothymine, O 4Examples of bases include, but are not limited to, methylthymine, uracil (U), 4-thiouracil (4sU), and 5,6-dihydrouracil (dihydrouracil; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; nebularine; inosine; hydroxymethylcytosine; 5-methycytosine; base (Y); and methylated, glycosylated, and acylated base moieties. Additional exemplary bases can be found in Fasman, 1989, "Practical Handbook of Biochemistry and Molecular Biology," pp. 385-394, CRC Press, Boca Raton, Fla.
[0051] Nucleotides (and nucleosides) typically include a sugar moiety, such as a carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100:4319-48), an acyclic moiety (Martinez, et al., 1999 Nucleic Acids Research 27:1271-1274; Martinez, et al., 1997 Bioorganic & Medicinal Chemistry Letters vol. 7:3013-3016), and another sugar moiety (Joeng, et al., 1993 J. Med. Chem. 36:2627-2638; Kim, et al., 1993 J. Med. Chem. 36:30-7; Eschenmosser 1999 Science 284:2118-2124; and U.S. Pat. No. 5,558,991). Sugar moieties include ribosyl; 2'-deoxyribosyl; 3'-deoxyribosyl; 2',3'-dideoxyribosyl; 2',3'-didehydrodideoxyribosyl; 2'-alkoxyribosyl; 2'-azidoribosyl; 2'-aminoribosyl; 2'-fluororibosyl; 2'-mercaptoriboxyl; 2'-alkylthioribosyl; 3'-alkoxyribosyl; 3'-azidoribosyl; 3'-aminoribosyl; 3'-fluororibosyl; 3'-mercaptoriboxyl; 3'-alkylthioribosyl carbocyclic; acyclic, or other modified sugars.
[0052] In some embodiments, the nucleotide comprises a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramido linkage. In some embodiments, the nucleotide is an analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise substituted side chain groups including O, S, or BH3. In some embodiments, the chain comprises phosphate groups substituted with analogs including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoramidite groups.
[0053] As used herein, "nucleotide unit" or "nucleotide moiety" refers to a nucleotide (e.g., dATP, dTTP, dGTP, dCTP, or dUTP), or an analog thereof, that comprises a base, a sugar, and at least one phosphate group. The nucleotide unit can be attached to a multivalent molecule used in the sequencing reactions described herein. Generally, all nucleotide units attached to the same multivalent molecule will have the same identity (e.g., all A's, all T's, all C's, or all G's), although one of skill in the art will understand that there may be situations in which multivalent molecules comprising nucleotide units of different identities are advantageous.
[0054] The terms "reporter moiety," "reporter moieties," or related terms refer to a compound that produces or can be caused to produce a detectable signal. Reporter moieties are often referred to as "labels." Any suitable reporter moiety can be used, and suitable reporter moieties include luminescence, photoluminescence, electroluminescence, bioluminescence, chemiluminescence, fluorescence, phosphorescence, chromophores, radioisotopes, electrochemistry, mass spectrometry, Raman, haptens, affinity tags, atoms, or enzymes. A reporter moiety produces a detectable signal that results from a chemical or physical change (e.g., heat, light, electricity, pH, salt concentration, enzymatic activity, or a proximity event). A proximity event involves two reporter moieties coming into close proximity with, associating with, or binding to each other. It is well known to those skilled in the art to select reporter moieties so that each absorbs excitation radiation and / or emits fluorescence at a wavelength distinguishable from other reporter moieties, allowing for the monitoring of the presence of different reporter moieties in the same or different reactions. Two or more different reporter moieties may be selected that have spectrally distinct emission profiles or that have minimal overlapping spectral emission profiles. The reporter moiety may be bound (e.g., operably bound) to a nucleotide, a nucleoside, a nucleic acid, an enzyme (e.g., a polymerase or reverse transcriptase), or a support (e.g., a surface).
[0055] The reporter moiety (or label) comprises a fluorescent label or fluorophore. Exemplary fluorescent moieties that can function as fluorescent labels or fluorophores include fluorescein and fluorescein derivatives, such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynapthofluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamidofluorescein, fluorescein maleimide, SAMSA-fluorescein, fluorescein thiosemicarbazide, carbohydrazinomethylthioacetyl-aminofluorescein, rhodamine and rhodamine derivatives, such as TRITC, TMR, lissamine rhodamine, Texas Red, rhodamine B, rhodamine 6G, rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, lissamine rhodamine B sulfonyl chloride, lissamine rhodamine B sulfonylhydrazine, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550 C3, BODIPY 530 / 550 C3-SE, BODIPY 530 / 550 C3 hydrazide, BODIPY 493 / 503 C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FL IA, BODIPY 530 / 551 IA, Br-BODIPY 493 / 503, Cascade Blue and derivatives such as Cascade Blue acetyl azide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Yellow and derivatives, such as Lucifer Yellow iodoacetamide, Lucifer Yellow CH, cyanines and derivatives, such as indolium-based cyanine dyes, benzo-indolium-based cyanine dyes, pyridium-based cyanine dyes, thiozolium-based cyanine dyes, quinolinium-based cyanine dyes, imidazolium-based cyanine dyes, Cy3, Cy5,Lanthanide chelates and derivatives, such as BCPDA, TBP, TMT, BHHCT, BCOT, europium chelates, terbium chelates, Alexa Fluor dyes, DyLight dyes, Atto dyes, LightCycler Red dyes, CAL Flour dyes, JOE and its derivatives, Oregon Green dyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, malachite green, stilbenes, DEG dyes, NR dyes, near-infrared dyes, and others known in the art, such as those described in Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th Edition, Lakowicz, Principles of Fluorescence Spectroscopy, 2nd Ed., Plenum Press New York (1999), or Hermanson, Bioconjugate Techniques, 2nd Edition, or derivatives thereof, or any combination thereof. Cyanine dyes can exist in either sulfonated or non-sulfonated form and consist of two indolenine, benzoindolium, pyridium, thiozolium, and / or quinolinium groups separated by a polymethine bridge between the two nitrogen atoms. Commercially available cyanine fluorophores include, for example, Cy3 (which is 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium, may include 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium-5-sulfonate), Cy5 (which may include1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-indolin-2-ylidene)penta-1,3-dien-1-yl)-3,3-dimethyl-3H-yne dol-1-ium, or 1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-sulfoindolin-2-ylidene)penta-1,3-dien-1-yl) Cy7 (which may include 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), and Cy8 (which may include 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), where "Cy" stands for "cyanine" and the first number identifies the number of carbon atoms between the two indolenine groups. Cy2, which is an oxazole derivative rather than an indolenine, and benzo-derivatized Cy3.5, Cy5.5, and Cy7.5 are exceptions to this rule.
[0056] In some embodiments, the reporter moieties may be FRET pairs, allowing multiple classifications to be performed under a single excitation and imaging step. As used herein, FRET may include excitation exchange (Förster) transfer or electron exchange (Dexter) transfer.
[0057] The terms "amplify," "amplifying," "amplification," and other related terms, when used with respect to nucleic acids, include producing multiple copies of an original polynucleotide template molecule, where the copies contain a sequence that is complementary to the template sequence or where the copies contain a sequence that is identical to the template sequence. In some embodiments, the copies contain a sequence that is substantially identical to the template sequence or a sequence that is substantially identical to the sequence that is complementary to the template sequence.
[0058] As used herein, the term "support" refers to a substrate designed for the deposition of biomolecules or biological samples for assay and / or analysis. Examples of biomolecules deposited on a support include nucleic acids (e.g., DNA, RNA), polypeptides, sugars, lipids, single cells, or multiple cells. Examples of biological samples include, but are not limited to, saliva, sputum, mucus, blood, plasma, serum, urine, feces, sweat, tears, and fluids from tissues or organs.
[0059] In some embodiments, the support is solid, semi-solid, or a combination of both. In some embodiments, the support is porous, semi-porous, non-porous, or any combination of porous. In some embodiments, the support can be substantially planar, concave, convex, or any combination thereof. In some embodiments, the support can be cylindrical, for example, comprising a capillary or the inner surface of a capillary.
[0060] In some embodiments, the surface of the support can be substantially smooth, hi some embodiments, the support can be regularly or irregularly textured, including ridges, etchings, pores, three-dimensional scaffolds, or any combination thereof.
[0061] In some embodiments, the support comprises beads having any shape, including spherical, hemispherical, cylindrical, barrel-shaped, toroidal, disk-shaped, rod-shaped, conical, triangular, cubic, polygonal, tubular, or wire-shaped.
[0062] The support can be made of any material, including, but not limited to, glass, fused silica, silicon, polymer (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic substrates are contemplated.
[0063] In some aspects, a support can have a plurality (e.g., two or more) of nucleic acid templates immobilized thereon. In some embodiments, the plurality of immobilized nucleic acid templates have the same sequence. In some embodiments, the plurality of immobilized nucleic acid templates have different sequences. In some embodiments, individual nucleic acid template molecules in the plurality of nucleic acid templates are immobilized at different sites on the support. In some embodiments, two or more individual nucleic acid template molecules in the plurality of nucleic acid templates are immobilized at sites on the support.
[0064] The term "array" refers to a support comprising a plurality of sites located at predetermined locations on the support, forming an array of sites. The sites may be dispersed and separated by interstitial regions. In some embodiments, the predetermined sites on the support may be arranged in rows or columns in one dimension, or in rows or columns in two dimensions. In some embodiments, the plurality of predetermined sites are arranged in an organized manner on the support. In some embodiments, the plurality of predetermined sites are arranged in any organized pattern, including linear, hexagonal, lattice, patterns with reflection symmetry, patterns with rotational symmetry, etc. The pitch between different pairs of sites may be the same or may vary. In some embodiments, the support has at least 10 2 at least 10 sites 3 at least 10 sites 4 at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites 9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13 at least 10 sites 14 sites, or at least 10 15 In some embodiments, the substrate comprises a plurality of predetermined sites (e.g., 10 or more sites), the sites being located at predetermined locations on the substrate. 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 Nucleic acid templates are immobilized at a plurality of predetermined sites (e.g., 10 or more sites) to form a nucleic acid template array. In some embodiments, the nucleic acid templates are immobilized at a plurality of predetermined sites by hybridization to the immobilized surface capture primers, or the nucleic acid templates are covalently attached to the surface capture primers. In some embodiments, the nucleic acid templates immobilized at a plurality of predetermined sites (e.g., 10 or more sites) form a nucleic acid template array. 2 ~10 15 More than one site (e.g., 10 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 In some embodiments, the immobilized nucleic acid template is clonally amplified to generate immobilized nucleic acid clusters at multiple predetermined sites. In some embodiments, the individual immobilized nucleic acid clusters comprise linear clusters or single- or double-stranded concatemers.
[0065] In some embodiments, a support comprising a plurality of sites located at random positions on the support is referred to herein as a support having randomly located sites thereon. In such embodiments, the locations of the randomly located sites on the support are not predetermined locations. As a result, the plurality of randomly located sites are arranged in an irregular and / or unpredictable manner on the support. In some embodiments, the support comprises at least 10 2 at least 10 sites 3 at least 10 sites 4 at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites 9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13 at least 10 sites 14 sites, or at least 10 15 In some embodiments, the substrate comprises a plurality of randomly positioned sites (e.g., 10 or more sites), where the sites are randomly located on the substrate. 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15In some embodiments, nucleic acid template molecules are immobilized at a plurality of randomly positioned sites (e.g., 10 or more sites) to form an immobilized nucleic acid template array. In some embodiments, the nucleic acid templates are immobilized at a plurality of randomly positioned sites by hybridization to the immobilized surface capture primers, or the nucleic acid templates are covalently attached to the surface capture primers. In some embodiments, the nucleic acid templates are immobilized at a plurality of randomly positioned sites (e.g., 10 or more sites), forming an immobilized nucleic acid template array. 2 ~10 15 More than one site (e.g., 10 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 In some embodiments, the immobilized nucleic acid template is clonally amplified to generate clusters of nucleic acids immobilized at multiple randomly located sites. In some embodiments, the individual immobilized nucleic acid clusters comprise linear clusters or single- or double-stranded concatemers.
[0066] In some embodiments, multiple immobilized surface capture primers on a support are in fluid communication with each other to allow solutions of reagents (e.g., nucleic acid template molecules, soluble primers, enzymes, nucleotides, divalent cations, buffers, etc.) to flow over the support, thereby allowing multiple immobilized surface capture primers on a support to react essentially simultaneously with reagents in a massively parallel manner. In some embodiments, the fluid communication of multiple immobilized surface capture primers can be used to perform nucleic acid amplification reactions (e.g., RCA, MDA, PCR, and bridge amplification) essentially simultaneously on multiple immobilized surface capture primers.
[0067] In some embodiments, multiple immobilized nucleic acid clusters on a support are in fluid communication with each other, allowing solutions of reagents (e.g., enzymes, nucleotides, divalent cations, etc.) to flow over the support, thereby allowing multiple immobilized nucleic acid clusters on a support to react with reagents essentially simultaneously in a massively parallel manner. In some embodiments, the fluid communication of multiple immobilized nucleic acid clusters can be used to perform nucleotide binding assays and / or nucleotide polymerization reactions (e.g., primer extension or sequencing) substantially simultaneously on multiple immobilized nucleic acid clusters, and optionally, to perform detection and imaging for massively parallel sequencing.
[0068] When used in reference to an immobilized enzyme, the term "immobilized" and related terms refer to an enzyme (e.g., a polymerase) that is attached to a support via covalent or non-covalent interactions, or that is attached to a coating on a support, or that is embedded within a matrix formed by a coating on a support.
[0069] The term "immobilized" and related terms, when used with respect to immobilized nucleic acids, refer to nucleic acid molecules that are attached to a support via covalent or non-covalent interactions, or that are attached to a coating on a support, or that are embedded within a matrix formed by a coating on a support, where the nucleic acid molecules include a surface capture primer, a nucleic acid template molecule, and an extension product of the capture primer. The extension product of the capture primer can include a nucleic acid concatemer (e.g., a nucleic acid cluster).
[0070] In some embodiments, one or more nucleic acid templates are immobilized on a support, e.g., immobilized at sites on a support. In some embodiments, one or more nucleic acid templates are clonally amplified. In some embodiments, one or more nucleic acid templates are clonally amplified off the support (e.g., in solution), then deposited on a support and immobilized on the support. In some embodiments, a clonal amplification reaction of one or more nucleic acid templates is performed on the support, resulting in immobilization on the support. In some embodiments, one or more nucleic acid templates are clonally amplified using a nucleic acid amplification reaction (e.g., in solution or on the support), where the nucleic acid amplification reaction includes any one of polymerase chain reaction (PCR), multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, bridge amplification, isothermal bridge amplification, rolling circle amplification (RCA), circle-circle amplification, helicase-dependent amplification, recombinase-dependent amplification, and / or single-strand binding (SSB) protein-dependent amplification, or any combination thereof.
[0071] The term "duration" and related terms refer to the length of time that a binding complex formed between a target nucleic acid, a polymerase, and a conjugated or unconjugated nucleotide remains stable without any binding components dissociating from the binding complex. Duration indicates the stability of the binding complex and the strength of the binding interaction. Duration can be measured by observing the onset and / or duration of the binding complex, for example, by observing a signal from a labeled component of the binding complex. For example, a labeled nucleotide or a labeling reagent comprising one or more nucleotides can be present in the binding complex, thus allowing a signal from the label to be detected during the duration of the binding complex. One exemplary label is a fluorescent label.
[0072] Introduction The present disclosure provides compositions, devices, and methods for capturing nucleic acids from a cellular sample on a support, preparing library molecules on the support, amplifying the library molecules on the support to generate nucleic acid template molecules, and analyzing the immobilized nucleic acid template molecules (including detecting and / or sequencing the immobilized nucleic acid template molecules). The immobilized nucleic acid template molecules correspond to nucleic acids from the cellular sample. The immobilized nucleic acid template molecules are spatially located on the support in an arrangement similar to their spatial locations in the cellular sample.
[0073] Methods for preparing spatially resolved nucleic acids The present disclosure provides a method for preparing spatially resolved nucleic acids, comprising step (a) providing a support passivated with at least one layer of a low nonspecific binding coating. In some embodiments, the low nonspecific binding coating comprises a high contrast-to-noise (CNR) coating that provides a low-binding / low-scattering base layer. In some embodiments, the low nonspecific binding coating comprises at least one hydrophilic polymer layer. In some embodiments, the low nonspecific binding coating has a water contact angle of 45 degrees or less. In some embodiments, the low nonspecific binding coating can form a continuous layer on the support, or the low nonspecific binding coating can be disposed on the support in an organized pattern across the support, for example, as spots, grids, and / or lines. In some embodiments, the low nonspecific binding coating can provide a surface with a high contrast-to-noise (CNR) ratio. Additional description of low nonspecific binding coatings is provided below. In some embodiments, the support is passivated with a cell-binding coating comprising polylysine, polyarginine, other polycations, tissue-specific antibodies, detergents, amphiphilic peptides, tissue-specific or general receptor ligands, particularly cell surface receptor ligands, and / or lectins. In some embodiments, the cell-binding coating may be disposed on the support in an organized pattern across the support, for example, as spots, grids, and / or lines. In some embodiments, the cell-binding coating may be present at a low density optimized to provide cell or tissue binding without interfering with the low binding / high CNR characteristics of the low nonspecific binding coating(s).
[0074] In some embodiments, in step (a), the support is solid, semi-solid, or a combination of both. In some embodiments, the support is porous, semi-porous, non-porous, or any combination of porous. In some embodiments, the support can be substantially planar, concave, convex, or any combination thereof.
[0075] In some embodiments, the low non-specific binding coating in step (a) comprises a plurality of surface capture primers that selectively hybridize (capture) nucleic acids, such as RNA and / or DNA, from a cell sample. The surface capture primers can hybridize to any portion of the RNA or DNA from a cell sample. In some embodiments, the surface capture primers comprise an RNA capture sequence comprising a poly-T sequence, a target-specific sequence, and / or a random sequence. In some embodiments, the surface capture primers comprise a universal adapter sequence, such as a universal surface capture primer sequence, a universal surface pinning primer sequence, a universal binding site for a reverse sequencing primer, a universal binding site for a forward sequencing primer, and / or a universal binding site for a compaction oligonucleotide. In some embodiments, the poly-T sequence comprises 3 to 30 nucleotides having thymine bases. In some embodiments, each surface capture primer comprises a universal surface capture primer sequence (110) (e.g., arranged in a 5' to 3' direction), a universal binding site for a reverse sequencing primer (120), and an RNA capture sequence (see, e.g., Figure 11). In some embodiments, the RNA capture sequence comprises a poly-T sequence and / or an RNA target-specific sequence. In some embodiments, the 5' end of the surface capture primer is tethered to a low non-specific binding coating. In some embodiments, the support further comprises a plurality of surface pinning primers tethered to the low non-specific binding coating. In some embodiments, the immobilized surface pinning primer functions to pin at least one portion of the concatemer molecule to the support (e.g., concatemer molecules are described below). In some embodiments, the immobilized surface pinning primer has a non-extendible 3' end and cannot be used for amplification. In some embodiments, the surface capture primer and the surface pinning primer can be covalently tethered to the low non-specific binding coating.
[0076] In some embodiments, the universal surface capture primer sequence (110) comprises the following sequence: 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 1).
[0077] In some embodiments, the universal surface capture primer sequence (110) comprises the following sequence: 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 2).
[0078] In some embodiments, the universal binding site for the reverse sequencing primer (120) comprises the following sequence: 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO: 3).
[0079] In some embodiments, the universal binding site for the reverse sequencing primer (120) comprises the following sequence: 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO: 4).
[0080] In some embodiments, the universal binding site for the reverse sequencing primer (120) comprises the following sequence: 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO: 5).
[0081] In some embodiments, the method for preparing spatially resolved nucleic acids further includes step (b): positioning the cell sample on the support under conditions suitable for the cell sample to remain in a fixed position on the coated support. In some embodiments, the cell sample can be manually placed on the support using a tool such as forceps. In some embodiments, the cell sample can be a tissue slice that flows or is placed on the support. In some embodiments, the support includes a flow cell, which can be a closed flow cell or a resealable flow cell. In some embodiments, the cell sample can be permeabilized to allow nucleic acids to elute from inside the cells to outside the cells and contact the support. In some embodiments, RNA can be eluted from inside the cells onto a surface capture primer immobilized on the coated support. In some embodiments, the eluted RNA can hybridize to the surface capture primer to generate multiple capture primer-RNA duplexes, each comprising a surface capture primer hybridized to an RNA molecule (see, e.g., Figure 12). In some embodiments, RNA can be eluted from inside the cells onto the support by diffusion, heat-assisted diffusion, radiation-assisted diffusion, electrophoresis, centrifugation, or other methods. In some embodiments, RNA can be eluted from inside the cells onto the surface capture primer in a manner that preserves the spatial location information of the RNA molecules in the cell sample. In some embodiments, the positions of the RNA molecules hybridized to the surface capture primer on the coated support correspond to the spatial locations of the RNA molecules when they were previously located inside the cell sample. In some embodiments, the cell sample can be permeabilized before, after, or simultaneously with positioning the cell sample on the support. In some embodiments, the cell sample can be treated with a chemical fixation reagent, or the cell sample can be unfixed chemically. In some embodiments, the cell sample can be stained, destained, or unstained. In some embodiments, the cell sample can be imaged in step (b), or stained and imaged.Cell samples were collected from “Histological and Histochemical Methods: Theory and Practice”, 4. th Edition, Kiernan, JA, Ed. (2008). In some embodiments, imaging can be performed using negative staining, or by fluorescence or interferometry. In some embodiments, in step (b), the cell sample (e.g., a tissue slice) can be positioned on a support for automated or semi-automated scanning. In some embodiments, the cell sample can be removed after eluting nucleic acids from the cell sample onto the immobilized surface capture primers on the support in step (b). In some embodiments, the cell sample can be removed from the support after any of steps (c)-(i) described below. In some embodiments, the cell sample can be imaged, or stained and imaged, during any of steps (c)-(i).
[0082] In some embodiments, the method for preparing spatially resolved nucleic acids further includes step (c): performing a reverse transcription reaction on the support by contacting the capture primer-RNA duplex with Moloney murine leukemia virus (MMLV) reverse transcriptase and a plurality of nucleotides (e.g., dATP, dGTP, dCTP, dTTP, and / or dUTP). In some embodiments, the reverse transcription reaction generates a plurality of first-strand cDNA molecules by extending the 3' end of the immobilized surface capture primer and using the hybridized RNA as a template strand. In some embodiments, the reverse transcription reaction can be performed under conditions suitable for MMLV reverse transcriptase to generate a non-template poly-C tail at the 3' end of the first-strand cDNA (see, e.g., Figure 13). Alternatively, a non-template poly-C tail can be added to the 3' end of the first-strand cDNA using terminal deoxynucleotidyl transferase and dCTP.
[0083] In some embodiments, the reverse transcription reaction of step (c) can be performed in the presence of multiple template switching oligonucleotides. Each template switching oligonucleotide can hybridize to a non-template poly-C tail at the 3' end of the first strand of the cDNA (see, e.g., Figure 14). In some embodiments, the template switching oligonucleotide comprises a DNA / RNA chimeric oligonucleotide. In some embodiments, the template switching oligonucleotide comprises a poly-G region that can hybridize to a poly-C tail at the 3' end of the first strand of the cDNA. In some embodiments, the poly-G region of the template switching oligonucleotide comprises ribonucleotide bases (see, e.g., Figure 14). In some embodiments, the template switching oligonucleotide further comprises at least one universal adapter sequence, e.g., a universal binding site (140) for a surface pinning primer and a universal binding site (130) for a forward sequencing primer (see, e.g., Figure 14).
[0084] In some embodiments, the universal binding site (130) of the forward sequencing primer comprises the following sequence: 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO: 6).
[0085] In some embodiments, the universal binding site (130) of the forward sequencing primer comprises the following sequence: 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 7).
[0086] In some embodiments, the universal binding site (130) of the forward sequencing primer comprises the following sequence: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 8).
[0087] In some embodiments, the universal binding site for the surface pinning primer (140) comprises the following sequence: 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 9).
[0088] In some embodiments, the universal binding site for the surface pinning primer (140) comprises the following sequence: 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 10).
[0089] In some embodiments, the method for preparing spatially resolved nucleic acids further comprises step (d): extending the first strand of the cDNA using MMLV reverse transcriptase, a plurality of nucleotides (e.g., dATP, dGTP, dCTP, dTTP, and / or dUTP), and at least a portion of a template switching oligonucleotide (TSO) as the template strand (see the upper dashed line with an arrow in Figure 15 labeled "First Strand cDNA (TSO)"). In some embodiments, the extension reaction of step (d) extends from the 3' end of the poly-C tail of the first strand cDNA. The extension reaction of step (d) may produce a full-length first strand of cDNA that includes (in the 5' to 3' direction) the universal surface capture primer sequence (110), a universal binding site for the reverse sequencing primer (120), the RNA capture sequence, the cDNA insert corresponding to the RNA, a poly-C tail, a universal binding site for the forward sequencing primer (130), and a universal binding site for the surface pinning primer (140).
[0090] In some embodiments, the method for preparing spatially resolved nucleic acids further comprises step (e): removing RNA while retaining the immobilized full-length first-strand cDNA. In some embodiments, the RNA can be removed using RNase. In some embodiments, the retained immobilized full-length first-strand cDNA comprises (in the 5' to 3' direction) a universal surface capture primer sequence (110), a universal binding site for a reverse sequencing primer (120), an RNA capture sequence, a cDNA insert corresponding to the RNA, a poly-C tail, a universal binding site for a forward sequencing primer (130), and a universal binding site for a surface pinning primer (140) (see, e.g., Figure 16).
[0091] In some embodiments, the method for preparing spatially resolved nucleic acids further comprises step (f): contacting the retained, immobilized, full-length, first-strand cDNA with a plurality of single-stranded circularization oligonucleotides. In some embodiments, each circularization oligonucleotide comprises an adapter sequence complementary to the universal binding site (120) for the reverse sequencing primer, an adapter sequence complementary to the universal surface capture primer sequence (110), a linker region, an adapter sequence complementary to the universal binding site (140) for the surface pinning primer, and an adapter sequence complementary to the universal binding site (130) for the forward sequencing primer (see, e.g., Figure 17). In some embodiments, the circularization oligonucleotide lacks the adapter sequence complementary to the universal surface capture primer sequence (110). The retained, immobilized, full-length first-strand cDNA is contacted with multiple single-stranded circularization oligonucleotides under conditions suitable for the circularization oligonucleotides to hybridize to the universal adapter sequence of the immobilized first-strand cDNA but not to the RNA capture sequence, the cDNA insert region, or the poly-C tail. Hybridization of the circularization oligonucleotides to the immobilized full-length first-strand cDNA forms gapped, single-stranded circular molecules. In some embodiments, the linker region may include any one or any combination of two or more of at least one sample index sequence for multiplexing, at least one unique molecular index (UMI) sequence for molecular tagging, and / or a universal binding site for a compaction oligonucleotide. In some embodiments, the sample index comprises 5 to 20 bases and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, a unique molecular index (UMI) comprises 3-20 bases that can be used to uniquely identify an individual nucleic acid molecule to which the UMI sequence is attached (e.g., molecularly tagged).In some embodiments, the unique molecular index (UMI) comprises a random sequence.
[0092] In some embodiments, the method for preparing spatially resolved nucleic acids further comprises step (g): performing a polymerase-catalyzed extension reaction to fill in the gaps using the poly-C tail region, the first-strand cDNA insert region, and the RNA capture sequence region as template strands. The polymerase-catalyzed extension reaction forms a nicked, single-stranded, circularized molecule. In some embodiments, step (g) further comprises performing an enzymatic ligation reaction to generate a single-stranded, covalently closed circular molecule hybridized to the immobilized full-length first-strand cDNA (see, e.g., Figure 18).
[0093] In some embodiments, the method for preparing spatially resolved nucleic acids further comprises step (h): performing a rolling circle amplification reaction using the 3' end of the immobilized full-length first strand cDNA as an initiation site to generate nucleic acid concatemers immobilized on a support. In some embodiments, the rolling circle amplification reaction uses a DNA polymerase with strand displacement activity and a plurality of nucleotides (e.g., dATP, dGTP, dCTP, dTTP, and / or dUTP). A rolling circle amplification reaction can use a covalently closed circular molecule as a template strand to generate concatemers with multiple tandem repeat units, each unit comprising a linker region (or its complementary sequence), a universal surface capture primer sequence (110) (or its complementary sequence), a universal binding site for a reverse sequencing primer (120) (or its complementary sequence), an RNA capture sequence (or its complementary sequence), a first strand cDNA insert region (or its complementary sequence), a polyC region (or its complementary sequence), a universal binding site for a forward sequencing primer (130) (or its complementary sequence), and a universal binding site for a surface pinning primer (140) (or its complementary sequence) (see, e.g., Figure 19).
[0094] In some embodiments, the rolling circle amplification reaction of step (h) can be performed in the presence or absence of multiple compaction oligonucleotides. In some embodiments, the compaction oligonucleotides comprise single-stranded oligonucleotides having a first region at one end that hybridizes to a portion of a concatemer molecule and a second region at the other end that hybridizes to another portion of the same concatemer molecule, such that hybridization of the compaction oligonucleotide to a given concatemer compacts the size and / or shape of the concatemer. In some embodiments, the 5' and 3' regions of the compaction oligonucleotide can hybridize to different portions of the same concatemer, bringing distal portions of the concatemer together and causing compaction of the concatemer to form DNA nanoballs. Additional description of compaction oligonucleotides can be found below.
[0095] In some embodiments, the DNA polymerase with strand displacement activity in step (h) can be selected from the group consisting of phi29 DNA polymerase, large fragment of Bst DNA polymerase, large fragment of Bsu DNA polymerase, and Bca(exo)DNA polymerase, Klenow fragment of E. coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV viral reverse transcriptase, or Deep Vent DNA polymerase. In some embodiments, the phi29 DNA polymerase can be wild-type phi29 DNA polymerase (e.g., MagniPhi from Expedeon), or variant EquiPhi29 DNA polymerase (e.g., from Thermo Fisher Scientific), and chimeric QualiPhi DNA polymerase (e.g., from 4basebio).
[0096] In some embodiments, the method for preparing spatially resolved nucleic acids further includes step (i): sequencing a plurality of immobilized concatemers. In some embodiments, the sequencing in step (i) includes sequencing at least a portion of each immobilized concatemer to identify each captured RNA. In some embodiments, the location of each sequenced concatemer on the coated support corresponds to the spatial location of each RNA from the cell sample. In some embodiments, the sequencing in step (i) includes sequencing at least a portion of each immobilized concatemer, including sequencing at least a portion of the cDNA insert of the concatemer corresponding to the captured RNA. In some embodiments, sequencing the concatemers immobilized on the low nonspecific binding coating in step (a) can provide a surface with a high contrast-to-noise (CNR) ratio during sequencing. A high contrast-to-noise (CNR) ratio can provide improved sensitivity and specificity for the precise spatial location of the concatemers on the support, which corresponds to the precise location of the RNA transcript of interest from the cell sample. In some embodiments, the sequencing in step (i) comprises the use of fluorescently labeled nucleotide reagents and a sequencing polymerase, hi some embodiments, the sequencing in step (i) comprises imaging to detect fluorescent signals emitted from immobilized concatemers during a sequencing reaction using fluorescently labeled nucleotide reagents and a sequencing polymerase.
[0097] In some embodiments, multiple immobilized concatemers can be sequenced using any nucleic acid sequencing method using labeled or unlabeled chain-terminating nucleotides, where the chain-terminating nucleotides contain a 3'-O-azido group (or a 3'-O-methyl azido group) or any other type of bulky blocking group at the 3' position of the sugar. In some embodiments, concatemer template molecules can be sequenced using a two-step sequencing method using labeled multivalent molecules and unlabeled chain-terminating nucleotides. In some embodiments, concatemer template molecules can be sequenced using a sequencing-by-synthesis (SBS) method using labeled chain-terminating nucleotides. In some embodiments, concatemer template molecules can be sequenced using a sequencing-by-binding (SBB) method using unlabeled chain-terminating nucleotides. In some embodiments, concatemer template molecules can be sequenced using phosphate-chain-labeled nucleotides. These various sequencing methods are described below.
[0098] In some embodiments, the immobilized concatemers can be detected by contacting the immobilized concatemers with a plurality of oligonucleotide probes labeled with detectable reporter moieties under conditions suitable for selective hybridization of the probes to the immobilized concatemers. In some embodiments, the contacting step can further include an imaging step. In some embodiments, the oligonucleotide probes comprise target-specific sequences, such as sequences complementary to any portion of the cDNA insert region. In some embodiments, the oligonucleotide probes comprise sequences that can hybridize to a universal adapter sequence comprising any of a universal binding site for a surface capture primer, a universal binding site for a surface pinning primer, a universal binding sequence for a forward sequencing primer, a universal binding sequence for a reverse sequencing primer, a sample barcode sequence, and / or a unique molecular index sequence.
[0099] In some embodiments, the method for preparing spatially resolved nucleic acids further comprises step (j): determining the positions of individual concatemers on the coated support that correspond to the spatial positions of individual RNAs eluted from the cell sample. In some embodiments, the sequencing and imaging of step (i) can be used to determine the positions of individual concatemers immobilized on the coated support that correspond to the spatial positions of individual RNAs eluted from the cell sample.
[0100] Cell samples In any of the methods described herein, a cell sample refers to a single cell, a whole cell, multiple cells, a tissue, an organ, an organism, or a section of any of these cell samples. Cell samples can be extracted from an organism (e.g., a biopsy) or obtained from cell cultures grown in liquid or on a culture dish. Cell samples include fresh samples, frozen samples, fresh frozen samples, or archived (e.g., formalin-fixed paraffin-embedded, FFPE) samples. Cell samples can be embedded in wax, paraffin, resin, epoxy, or agar. Cell samples can be fixed, for example, in any combination of one or more of acetone, ethanol, methanol, formaldehyde, paraformaldehyde-Triton, or glutaraldehyde. Cell samples may or may not be sectioned. Cell samples may be stained, destained, or unstained.
[0101] In some embodiments, the cell sample may be obtained from a virus, a fungus, a prokaryote, or a eukaryote. In some embodiments, the cell sample may be obtained from an animal, an insect, or a plant. In some embodiments, the cell sample comprises one or more virus-infected cells.
[0102] In some embodiments, the cell sample may be obtained from any organism, including a human, monkey, ape, dog, cat, cow, horse, mouse, pig, goat, wolf, frog, fish, plant, insect, or bacterium.
[0103] In some embodiments, cell samples may be obtained from any organ, including the head, neck, brain, breast, ovaries, cervix, colon, rectum, endometrium, gallbladder, intestines, bladder, prostate, testes, liver, lung, kidney, esophagus, pancreas, thyroid, pituitary, thymus, skin, heart, larynx, or other organs.
[0104] In any of the methods described herein, a cell sample carries a plurality of RNAs that can be eluted onto a plurality of surface capture primers immobilized on a coated support. In some embodiments, the eluted RNA comprises target RNA and / or non-target RNA. In some embodiments, the eluted RNA comprises wild-type RNA, mutant RNA, and / or splice variant RNA. In some embodiments, the eluted RNA comprises pre-spliced RNA, partially spliced RNA, and / or fully spliced RNA. In some embodiments, the eluted RNA comprises coding RNA, non-coding RNA, mRNA, tRNA, alternative rRNA, rRNA, microRNA (miRNA), mature microRNA, immature microRNA, lncRNA, ncRNA, and / or intronic RNA. In some embodiments, the eluted RNA comprises housekeeping RNA, cell-specific RNA, tissue-specific RNA, or disease-specific RNA. In some embodiments, the eluted RNA comprises RNA expressed by one or more cells in response to a stimulus, such as heat, light, a chemical, or a drug. In some embodiments, the eluted RNA comprises RNA found in healthy or diseased cells. In some embodiments, the eluted RNA comprises RNA transcribed from a transgenic DNA sequence introduced into the cell sample using recombinant DNA procedures. For example, the RNA can be transcribed from a transgenic DNA sequence controlled by an inducible or constitutive promoter sequence. In some embodiments, the eluted RNA comprises RNA transcribed from a non-transgenic DNA sequence.
[0105] Transparency In any of the methods described herein, a cell sample can be permeabilized to allow nucleic acids within the sample, including target nucleic acid molecules (e.g., RNA and / or DNA), to migrate from inside the cell(s) to multiple capture oligonucleotides immobilized on a support. The cell sample can be contacted with one or more permeabilizing agents, including organic solvents, detergents, compounds, crosslinking agents, and / or enzymes. In some embodiments, the organic solvent includes acetone, ethanol, and methanol. In some embodiments, the detergent includes saponin, Triton X-100, Tween-20, sodium dodecyl sulfate (SDS), or N-lauroyl sarcosine sodium salt solution. In some embodiments, the crosslinking agent includes paraformaldehyde. In some embodiments, the enzyme includes trypsin, pepsin, or a protease (e.g., proteinase K). In some embodiments, the cell sample can be permeabilized using alkaline conditions or acidic conditions with a protease enzyme. In some embodiments, target nucleic acid molecules from a cell sample are hybridized (captured) onto capture oligonucleotides immobilized on a support (or support coating) in a manner that preserves the spatial location of the target nucleic acid molecules in the cell sample.
[0106] In any of the methods described herein, the cell sample comprises a permeabilized cell sample. In some embodiments, the method comprises treating the cell sample with a permeabilization reagent that alters the cell membrane to allow experimental reagents to penetrate the cells. For example, the permeabilization reagent removes membrane lipids from the cell membrane. In some embodiments, the cell sample can be treated with a permeabilization reagent comprising any combination of an organic solvent, a detergent, a compound, a crosslinking agent, and / or an enzyme. In some embodiments, the organic solvent comprises acetone, ethanol, and methanol. In some embodiments, the detergent comprises saponin, Triton X-100, Tween-20, sodium dodecyl sulfate (SDS), N-lauroyl sarcosine sodium salt solution, or a non-ionic polyoxyethylene surfactant (e.g., NP40). In some embodiments, the crosslinking agent comprises paraformaldehyde. In some embodiments, the enzyme comprises trypsin, pepsin, or a protease (e.g., proteinase K). In some embodiments, the cells can be permeabilized using alkaline conditions or acidic conditions with a protease enzyme. In some embodiments, the permeabilization reagent comprises water and / or PBS.
[0107] For example, fixed cells can be permeabilized with 70% ethanol for about 30-60 minutes, and the permeabilization reagent can be exchanged for PBS-T (e.g., PBS with 0.05% Tween-20). In some embodiments, cells can be post-fixed with 3% paraformaldehyde and 0.1% glutaraldehyde for about 30-60 minutes and washed multiple times with PBS-T.
[0108] cell fixation In any of the methods described herein, the cell sample comprises a fixed cell sample. In some embodiments, the cell sample may be treated with a fixation reagent (e.g., a fixation reagent) that can preserve cells and their contents, inhibiting degradation and inhibiting cell lysis. For example, the fixation reagent can preserve RNA carried by the cell sample. In some embodiments, the fixation reagent inhibits loss of nucleic acids from the cell sample.
[0109] In some embodiments, the fixation reagent can cross-link RNA to prevent it from escaping from the cell sample. In some embodiments, the cross-linking fixation reagent comprises any combination of aldehydes, formaldehyde, paraformaldehyde, formalin, glutaraldehyde, imidoesters, N-hydroxysuccinimide esters (NHS), and / or glyoxal (a bifunctional aldehyde).
[0110] In some embodiments, the fixation reagent comprises at least one alcohol, including methanol or ethanol. In some embodiments, the fixation reagent comprises at least one ketone, including acetone. In some embodiments, the fixation reagent comprises acetic acid, glacial acetic acid, and / or picric acid. In some embodiments, the fixation reagent comprises mercuric chloride. In some embodiments, the fixation reagent comprises a zinc salt, including zinc sulfate or zinc chloride. In some embodiments, the fixation reagent is capable of denaturing the polypeptide.
[0111] In some embodiments, the fixation reagent comprises 4% w / v paraformaldehyde in water / PBS. In some embodiments, the fixation reagent comprises 10% 35% formaldehyde at neutral pH. In some embodiments, the fixation reagent comprises 2% v / v glutaraldehyde in water / PBS. In some embodiments, the fixation reagent comprises 25% 37% formaldehyde solution, 70% picric acid, and 5% acetic acid.
[0112] In some embodiments, cell samples may be fixed on a support with 4% paraformaldehyde for approximately 30-60 minutes and washed with PBS.
[0113] Compaction Oligonucleotides In some embodiments, the rolling circle amplification reaction can be performed in the presence of a plurality of compaction oligonucleotides that, when hybridized to concatemer molecules, compact the size and / or shape of the concatemers to form compact nanoballs. In some embodiments, the compaction oligonucleotides comprise single-stranded oligonucleotides having a first region at one end that hybridizes to a portion of the concatemer molecule and a second region at the other end that hybridizes to another portion of the same concatemer molecule, such that hybridization of the compaction oligonucleotide to a given concatemer compacts the size and / or shape of the concatemer.
[0114] In some embodiments, the compaction oligonucleotide comprises a 5' region, an optional internal region (intervening region), and a 3' region. The 5' and 3' regions of the compaction oligonucleotide can hybridize to different portions of the same concatemer, bringing the distal portions of the concatemer together and causing compaction of the concatemer to form DNA nanoballs. For example, without limitation, the 5' region of the compaction oligonucleotide is designed to hybridize to a first portion of the concatemer molecule (e.g., a universal compaction oligonucleotide binding site), and the 3' region of the compaction oligonucleotide is designed to hybridize to a second portion of the concatemer molecule (e.g., a universal compaction oligonucleotide binding site). The inclusion of the compaction oligonucleotide during RCA can promote the formation of DNA nanoballs with a more compact size and shape compared to concatemers generated in the absence of the compaction oligonucleotide. The compact and stable characteristics of DNA nanoballs improve sequencing accuracy, for example, by increasing signal intensity, and the nanoballs retain their shape and size during multiple sequencing cycles.
[0115] In some embodiments, the compaction oligonucleotide comprises a single-stranded oligonucleotide comprising DNA, RNA, or a combination of DNA and RNA. The compaction oligonucleotide can be any length, including 20 to 150 nucleotides, e.g., 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 nucleotides, or any range therebetween. In some embodiments, the compaction oligonucleotide can be 30 to 100 nucleotides in length. In some embodiments, the compaction oligonucleotide can be 40 to 80 nucleotides in length. In some embodiments, the compaction oligonucleotide can be 80 to 1200 nucleotides in length.
[0116] In some embodiments, the compaction oligonucleotide comprises a 5' region and a 3' region, and optionally an intermediate region between the 5' and 3' regions. The intervening region can be any length, e.g., about 2 to 20 nucleotides in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, or any range therebetween. The intervening region comprises a homopolymer of consecutive identical bases (e.g., AAA, GGG, CCC, TTT, or UUU). The intervening region comprises a non-homopolymer sequence.
[0117] In some embodiments, the 5' region of the compaction oligonucleotide may be fully or partially complementary along its length to a first portion of the concatemer molecule. In some embodiments, the 3' region of the compaction oligonucleotide may be fully or partially complementary along its length to a second portion of the concatemer molecule. In some embodiments, the 5' region of the compaction oligonucleotide can hybridize to a first universal sequence portion of the concatemer molecule. In some embodiments, the 3' region of the compaction oligonucleotide can hybridize to a second universal sequence portion of the concatemer molecule.
[0118] In some embodiments, the 5' region of the compaction oligonucleotide may have the same sequence as the 3' region. The 5' region of the compaction oligonucleotide may have a different sequence than the 3' region. In some embodiments, the 3' region of the compaction oligonucleotide may have a sequence that is the reverse of the 5' region. In some embodiments, the 5' region of the compaction oligonucleotide may have a sequence that is the reverse of the 3' region.
[0119] In some embodiments, the 3' region of either of the compaction oligonucleotides can include an additional three bases at the 3' end of the terminal that includes a 2'-O-methyl RNA base (e.g., designated mUmUmU), or the 3' end of the terminal lacks the additional 2'-O-methyl RNA base.
[0120] In some embodiments, the compaction oligonucleotide contains one or more modified bases or linkages at its 5' or 3' end to confer specific functionality. In some embodiments, the compaction oligonucleotide contains at least one phosphorothioate linkage at its 5' and / or 3' end to confer exonuclease resistance. In some embodiments, at least one nucleotide at or near the 3' end contains a 2' fluoro base, which confers exonuclease resistance. In some embodiments, the 3' end of the compaction oligonucleotide is non-extendable. In some embodiments, the 3' end of the compaction oligonucleotide contains at least one 2'-O-methyl RNA base that blocks polymerase-catalyzed extension. For example, the 3' end of the compaction oligonucleotide contains three bases, including a 2'-O-methyl RNA base (e.g., designated mUmUmU). In some embodiments, the compaction oligonucleotide contains a 3' inverted dT at its 3' end to block polymerase-catalyzed extension. In some embodiments, the compaction oligonucleotide comprises a 3' phosphorylation that blocks polymerase-catalyzed elongation. In some embodiments, the internal region of the compaction oligonucleotide comprises at least one locked nucleic acid (LNA) that increases the thermal stability of the duplex formed by hybridizing the compaction oligonucleotide to the concatemer molecule. In some embodiments, the compaction oligonucleotide comprises a 5' end that is phosphorylated (e.g., using polynucleotide kinase).
[0121] In some embodiments, the compaction oligonucleotide comprises an additional three bases at the 3' end of the end that includes a 2'-O-methyl RNA base (e.g., designated mUmUmU), or the terminal 3' end lacks the additional 2'-O-methyl RNA base.
[0122] In some embodiments, the compaction oligonucleotide may comprise at least one region having consecutive guanines. For example, the compaction oligonucleotide may comprise at least one region having 2, 3, 4, 5, or more consecutive guanines. In some embodiments, the compaction oligonucleotide comprises four consecutive guanines that can form a G-quadruplex structure (see Figure 20). The G-quadruplex structure may be stabilized through Hoogsteen hydrogen bonding. The G-quadruplex structure may be stabilized by a central cation including potassium, sodium, lithium, rubidium, or cesium.
[0123] At least one compaction oligonucleotide can form a G-quadruplex (Figure 20) and hybridize to the universal binding sequence in the concatemer, allowing the concatemer to fold and form an intramolecular G-quadruplex structure (Figure 21). The concatemer can self-collapse to form a compact nanoball. The formation of the G-quadruplex and G-quadruplex in the nanoball can increase the stability of the nanoball, allowing it to retain a compact size and shape that can withstand changes in pH, temperature, and / or repeated flow of reagents during sequencing in cell samples.
[0124] In some embodiments, the rolling circle amplification reaction includes multiple compaction oligonucleotides having the same sequence. Alternatively, the rolling circle amplification reaction includes multiple compaction oligonucleotides having a mixture of two or more different sequences.
[0125] In some embodiments, the immobilized concatemeric template molecules can self-collapse into compact nucleic acid nanoballs, which can be imaged and FWHM measurements obtained to obtain the shape / size of the nanoballs.
[0126] In some embodiments, the inclusion of compaction oligonucleotides in rolling circle amplification reactions can facilitate the collapse of concatemers into DNA nanoballs. Performing RCA using compaction oligonucleotides helps preserve the compact size and shape of DNA nanoballs during multiple sequencing cycles, which can improve the FWHM (full width at half maximum) of spot images of DNA nanoballs within a cell sample. In some embodiments, DNA nanoballs do not resolve during multiple sequencing cycles. In some embodiments, DNA nanoball spot images do not expand during multiple sequencing cycles. In some embodiments, DNA nanoball spot images remain discrete spots during multiple sequencing cycles. Spot images can be represented as Gaussian spots, and size can be measured as FWHM. Smaller spot sizes, indicated by smaller FWHMs, typically correlate with improved spot images. In some embodiments, the FWHM of nanoball spots can be approximately 10 μm or less.
[0127] Methods for sequencing The present disclosure provides methods for sequencing a plurality of concatemeric template molecules immobilized on a support, such as any of the immobilized concatemeric molecules described herein. In some embodiments, the sequencing reaction uses a plurality of sequencing primers, a plurality of sequencing polymerases, and a nucleotide reagent comprising any one or any combination of nucleotides and / or multivalent molecules. In some embodiments, the nucleotide reagent comprises a standard nucleotide. In some embodiments, the nucleotide reagent comprises a nucleotide analog comprising a detectably labeled nucleotide. In some embodiments, the nucleotide reagent comprises a nucleotide bearing a removable or non-removable chain-terminating moiety. In some embodiments, the nucleotide reagent comprises a multivalent molecule comprising a central core attached to a plurality of polymer arms, each having a nucleotide unit at the end of the arm. In some embodiments, the sequencing reaction uses binding of an unlabeled nucleotide without incorporation. In some embodiments, the sequencing reaction uses incorporation of an unlabeled nucleotide. In some embodiments, the sequencing reaction uses incorporation of a detectably labeled nucleotide with a removable chain-terminating moiety. In some embodiments, the sequencing reaction uses a two-step sequencing reaction that includes binding to a detectably labeled multivalent molecule without incorporation and incorporating a nucleotide analog, hi some embodiments, the sequencing reaction uses a phosphate-chain-labeled nucleotide.
[0128] Methods for sequencing using nucleotide analogs - Patent Application 20070122997 In some aspects, the disclosure provides a method for sequencing, comprising step (a): contacting a sequencing polymerase with (i) nucleic acid concatemer molecules and (ii) a nucleic acid primer (e.g., a forward or reverse sequencing primer), wherein the contacting is performed under conditions suitable for binding of the sequencing polymerase to the nucleic acid concatemer molecules hybridized to the nucleic acid primer, such that the nucleic acid concatemer molecules hybridized to the nucleic acid primer form a nucleic acid duplex. In some embodiments, the sequencing polymerase comprises a recombinant mutant sequencing polymerase. In some embodiments, the primer comprises a 3' extendable end.
[0129] In some embodiments, the method for sequencing further comprises step (b): contacting a sequencing polymerase with a plurality of nucleotides under conditions suitable for binding at least one nucleotide to the sequencing polymerase bound to the nucleic acid duplex and suitable for polymerase-catalyzed nucleotide incorporation. In some embodiments, the sequencing polymerase is contacted with the plurality of nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises at least one nucleotide analog having a chain-terminating moiety at the sugar 2' or 3' position. In some embodiments, the plurality of nucleotides comprises at least one nucleotide lacking a chain-terminating moiety.
[0130] In some embodiments, the method for sequencing further comprises step (c): incorporating at least one nucleotide into the 3' end of the extendable primer under conditions suitable for incorporation of the at least one nucleotide. In some embodiments, the conditions suitable for nucleotide binding to the polymerase and nucleotide incorporation may be the same or different. In some embodiments, the conditions suitable for nucleotide incorporation comprise including at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, at least one nucleotide binds to the sequencing polymerase and is incorporated into the 3' end of the extendable primer. In some embodiments, incorporating a nucleotide into the 3' end of the primer in step (c) comprises a primer extension reaction.
[0131] In some embodiments, the method for sequencing further comprises step (d): repeating the incorporation of at least one nucleotide into the 3' end of the extendable primer of step (c) at least once. In some embodiments, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base with a linker, the linker being cleavable / removable from the base. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to enable detection and identification of the nucleotide base. In some embodiments, the method further comprises detecting the at least one incorporated nucleotide in steps (c) and / or (d). In some embodiments, the method further comprises identifying at least one incorporated nucleotide in steps (c) and / or (d). In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined by detecting and identifying the nucleotide bound to the sequencing polymerase, thereby determining the sequence of the concatemer molecule. In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined by detecting and identifying the nucleotide incorporated within the 3' end of the primer, thereby determining the sequence of the concatemer molecule.
[0132] In some embodiments, in the method for sequencing, the plurality of sequencing polymerases bound to the nucleic acid duplex comprises a plurality of multiplexed polymerases having at least first and second multiplexed polymerases, wherein (a) the first multiplexed polymerase comprises a first sequencing polymerase bound to a first nucleic acid duplex comprising a first nucleic acid template sequence hybridized to a first nucleic acid primer, (b) the second multiplexed polymerase comprises a second sequencing polymerase bound to a second nucleic acid duplex comprising a second nucleic acid template sequence hybridized to a second nucleic acid primer, (c) the first and second nucleic acid template sequences comprise the same or different sequences, (d) the first and second nucleic acid concatemers are clonally amplified, (e) the first and second primers comprise an extendable 3' end or a non-extendable 3' end, and (f) the plurality of multiplexed polymerases are immobilized on a support. In some embodiments, the density of the plurality of multiplexed polymerases is greater than 1 mm 2 immobilized on the support. 2 Approximately 10 per 2 ~10 15 (e.g., 10 2 ~10 15 or more, e.g., 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 ) complex polymerases.
[0133] Two-step sequencing method for nucleic acids In some aspects, the present disclosure provides a two-step method for sequencing a nucleic acid molecule. In some embodiments, the first step generally comprises binding a multivalent molecule to a composite polymerase to form a multivalent composite polymerase and detecting the multivalent composite polymerase.
[0134] In some embodiments, the first stage comprises step (a): contacting a plurality of first sequencing polymerases with (i) a plurality of nucleic acid concatemer molecules and (ii) a plurality of nucleic acid sequencing primers (e.g., forward or reverse sequencing primers), wherein the contacting is performed under conditions suitable for binding the plurality of first sequencing polymerases to the plurality of nucleic acid concatemer molecules and the plurality of nucleic acid primers, thereby forming a plurality of first hybrid polymerases each comprising the first sequencing polymerase bound to a nucleic acid duplex, wherein the nucleic acid duplex comprises the nucleic acid concatemer molecules hybridized to the nucleic acid primers. In some embodiments, the first polymerase comprises a recombinant mutant sequencing polymerase.
[0135] In some embodiments, in the method for sequencing concatemer molecules, the primer comprises a 3' extendable end or a 3' non-extendable end. In some embodiments, the plurality of nucleic acid concatemer molecules comprises amplified template molecules (e.g., clonally amplified template molecules). In some embodiments, the plurality of nucleic acid concatemer molecules comprises one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules comprises (two or more tandem copies, e.g., concatemers, of) a target sequence of interest. In some embodiments, the nucleic acid concatemer molecules in the plurality of nucleic acid concatemer molecules comprise the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid concatemer molecules and / or the plurality of nucleic acid primers are in solution or immobilized on a support. In some embodiments, when the plurality of nucleic acid concatemer molecules and / or the plurality of nucleic acid primers are immobilized on a support, binding with a first sequencing polymerase generates a plurality of immobilized first multiplexed polymerases. In some embodiments, the plurality of nucleic acid concatemer molecules and nucleic acid primers are 10 on support 2 ~10 15 different sites (e.g., 10 2 ~1015 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 In some embodiments, the binding of the plurality of concatemer molecules and nucleic acid primers to the plurality of first sequencing polymerases is performed at 10 sites on the support. 2 ~10 15 different sites (e.g., 10 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 In some embodiments, the plurality of immobilized first hybrid polymerases on the support are immobilized at predetermined or random sites on the support. In some embodiments, the plurality of immobilized first hybrid polymerases are in fluid communication with each other, allowing a solution of reagents (e.g., enzymes, including sequencing polymerases, multivalent molecules, nucleotides, and / or divalent cations) to flow over the support, whereby the plurality of immobilized hybrid polymerases on the support react with the solution of reagents in a massively parallel manner.
[0136] In some embodiments, the method for sequencing further comprises step (b): contacting a plurality of first multiplexed polymerases with a plurality of multivalent molecules to form a plurality of multivalent multiplexed polymerases (e.g., bound complexes). In some embodiments, each multivalent molecule in the plurality of multivalent molecules comprises a core attached to a plurality of nucleotide arms, each nucleotide arm being attached to a nucleotide (e.g., a nucleotide unit) (e.g., Figures 2-6). In some embodiments, the contacting in step (b) is performed under conditions suitable for binding complementary nucleotide units of the multivalent molecule to at least two of the plurality of first multiplexed polymerases, thereby forming a plurality of multivalent multiplexed polymerases. In some embodiments, the conditions are suitable for inhibiting polymerase-catalyzed incorporation of complementary nucleotide units into primers of the plurality of multivalent multiplexed polymerases. In some embodiments, the plurality of multivalent molecules comprises at least one multivalent molecule having multiple nucleotide arms (e.g., Figures 2-6), each of the multiple nucleotide arms being attached to a nucleotide analog (e.g., a nucleotide analog unit) that comprises a chain-terminating moiety at the sugar 2' and / or 3' position. In some embodiments, the plurality of multivalent molecules comprises at least one multivalent molecule comprising multiple nucleotide arms, each of the multiple nucleotide arms being attached to a nucleotide unit lacking a chain-terminating moiety. In some embodiments, at least one of the multivalent molecules in the plurality of multivalent molecules is labeled with a detectable reporter moiety. Any portion of the multivalent molecule can be labeled, including the core, the nucleotide arms, or the nucleobase. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the contacting in step (b) is performed in the presence of at least one non-catalytic cation comprising strontium, barium, and / or calcium.
[0137] In some embodiments, the method for sequencing further comprises step (c): detecting the plurality of multivalent composite polymerases. In some embodiments, detecting comprises detecting multivalent molecules bound to the composite polymerases, wherein complementary nucleotide units of the multivalent molecules are bound to the primers but incorporation of the complementary nucleotide units is inhibited. In some embodiments, the multivalent molecules are labeled with a detectable reporter moiety to enable detection. In some embodiments, the labeled multivalent molecule comprises a fluorophore attached to the core, linker, and / or nucleotide units of the multivalent molecule.
[0138] In some embodiments, the method for sequencing further comprises step (d): identifying bases of complementary nucleotide units bound to the plurality of first multiplexed polymerases, thereby determining the sequence of the concatemeric molecule. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety corresponding to a particular nucleotide unit attached to the nucleotide arm, so as to allow identification of the complementary nucleotide units (e.g., the nucleotide bases adenine, guanine, cytosine, thymine, or uracil) bound to the plurality of first multiplexed polymerases.
[0139] In some embodiments, the second step of the two-step sequencing method generally comprises nucleotide incorporation. In some embodiments, the method for sequencing further comprises step (e): dissociating the plurality of multivalent composite polymerases, removing the plurality of first sequencing polymerases and their bound multivalent molecules, and retaining the plurality of nucleic acid duplexes.
[0140] In some embodiments, the method for sequencing further comprises the step (f): contacting the retained plurality of nucleic acid duplexes of step (e) with a plurality of second sequencing polymerases, wherein the contacting is performed under conditions suitable for binding the plurality of second sequencing polymerases to the retained plurality of nucleic acid duplexes, thereby forming a plurality of second multiplexed polymerases, each comprising a second sequencing polymerase bound to a nucleic acid duplex. In some embodiments, the second sequencing polymerase comprises a recombinant mutant sequencing polymerase.
[0141] In some embodiments, the plurality of first sequencing polymerases in step (a) have amino acid sequences that are 100% identical to the amino acid sequence of the plurality of second sequencing polymerases in step (f). In some embodiments, the plurality of first sequencing polymerases in step (a) have amino acid sequences that are different from the amino acid sequences of the plurality of second sequencing polymerases in step (f).
[0142] In some embodiments, the method for sequencing further comprises step (g): contacting a plurality of second hybrid polymerases with a plurality of nucleotides, wherein the contacting is performed under conditions suitable for binding complementary nucleotides from the plurality of nucleotides to at least two of the second hybrid polymerases, thereby forming a plurality of nucleotide-hybridized polymerases. In some embodiments, the contacting in step (g) is performed under conditions suitable for promoting polymerase-catalyzed incorporation of the bound complementary nucleotides into the primer by the nucleotide-hybridized polymerase, thereby forming a plurality of nucleotide-hybridized polymerases. In some embodiments, incorporating the nucleotide into the 3' end of the primer in step (g) comprises a primer extension reaction. In some embodiments, the contacting in step (g) is performed in the presence of at least one catalytic cation, including magnesium and / or manganese. In some embodiments, the contacting in step (g) is performed in the presence of magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises natural nucleotides (e.g., non-analog nucleotides) or nucleotide analogs. In some embodiments, the plurality of nucleotides comprises removable or non-removable 2' and / or 3' chain terminating moieties. In some embodiments, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety may comprise a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base with a linker, which may be cleavable / removable from the base or may not be removable from the base. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and identification of the nucleotide base.In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the plurality of nucleotides is labeled with a detectable reporter moiety.
[0143] In some embodiments, the method for sequencing further comprises step (h): if the nucleotide is labeled with a detectable reporter moiety, step (h) comprises detecting a complementary nucleotide incorporated into the primer of the nucleotide-conjugated polymerase. In some embodiments, the plurality of nucleotides are labeled with a detectable reporter moiety to enable detection. In some embodiments, in the method for sequencing concatemeric molecules, if the nucleotide is not labeled, the detection step is omitted.
[0144] In some embodiments, the method for sequencing further comprises step (i): if the nucleotide is labeled with a detectable reporter moiety, step (i) comprises identifying the base of the complementary nucleotide incorporated into the primer of the nucleotide-complexed polymerase. In some embodiments, the identification of the incorporated complementary nucleotide in step (i) can be used to confirm the identity of the complementary nucleotide of the multivalent molecule bound to the plurality of first composite polymerases in step (d). In some embodiments, the identifying in step (i) can be used to determine the sequence of the nucleic acid concatemer molecule. In some embodiments, in the method for sequencing concatemer molecules, if the nucleotide is not labeled, the identifying step is omitted.
[0145] In some embodiments, the method for sequencing further comprises step (j): if step (g) is carried out by contacting a plurality of second hybrid polymerases with a plurality of nucleotides comprising at least one nucleotide having a 2' and / or 3' chain terminating moiety, removing the chain terminating moiety from the incorporated nucleotides.
[0146] In some embodiments, the method for sequencing further comprises step (k): repeating steps (a) through (j) at least once. In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined in steps (c) and (d) by detecting and identifying multivalent molecules that bind to the sequencing polymerase but are not incorporated into the 3' end of the primer. In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined (or confirmed) in steps (h) and (i) by detecting and identifying nucleotides that are incorporated into the 3' end of the primer.
[0147] In some embodiments, in any of the methods for sequencing nucleic acid molecules, binding of a plurality of first multiplexed polymerases to a plurality of multivalent molecules forms at least one avidity complex, the method comprising the steps of: (a) binding a first nucleic acid primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemeric template molecule, thereby forming a first binding complex, wherein a first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (b) binding a second nucleic acid primer, a second sequencing polymerase, and a first multivalent molecule to a second portion of the same concatemeric template molecule, thereby forming a second binding complex, wherein a second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase; and the first and second binding complexes comprising the same multivalent molecule form an avidity complex. In some embodiments, the first sequencing polymerase comprises any wild-type or mutant polymerase described herein. In some embodiments, the second sequencing polymerase comprises any wild-type or mutant polymerase described herein. The concatemeric template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. First and second nucleic acid primers can bind to the sequencing primer binding sites along the concatemeric template molecule. Exemplary multivalent molecules are shown in Figures 2-6.
[0148] In some embodiments, in any of the methods for sequencing nucleic acid molecules, the method comprises combining a plurality of first multiplexed polymerases with a plurality of multivalent molecules to form at least one avidity complex, the method comprising the steps of: (a) contacting a plurality of sequencing polymerases and a plurality of nucleic acid primers with different portions of concatemeric nucleic acid concatemeric molecules to form at least first and second multiplexed polymerases on the same concatemeric molecule; and (b) contacting the plurality of multivalent molecules with at least first and second multiplexed polymerases on the same concatemeric template molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first and second multiplexed polymerases, wherein at least a first nucleotide unit of the single multivalent molecule comprises a first primer that hybridizes to a first portion of the concatemeric template molecule, thereby forming a first binding complex (e.g., a first ternary complex). (c) contacting a first multiplexed polymerase comprising a second primer that binds to the first and second multiplexed polymerase and that comprises a second primer, wherein at least a second nucleotide unit of the single multivalent molecule hybridizes to a second portion of the concatemeric template molecule, thereby forming a second binding complex (e.g., a second ternary complex), under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second binding complexes, such that the first and second binding complexes that bind to the same multivalent molecule form an avidity complex; (d) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the concatemeric template molecule. In some embodiments, the plurality of sequencing polymerases comprises any wild-type or mutant sequencing polymerase described herein.The concatemer template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. Multiple nucleic acid primers can bind to the sequencing primer binding sites along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 2-6.
[0149] In some embodiments, two-stage sequencing can use multivalent molecules labeled with fluorophores, and the detecting and / or identifying step involves the use of fluorescent imaging. In some embodiments, the fluorescent imaging involves dual-wavelength excitation / four-wavelength emission fluorescent imaging. In some embodiments, four different types of multivalent molecules are used, each containing a different nucleotide unit (or nucleotide unit analog). For example, a first type of multivalent molecule contains dATP nucleotide units, a second type of multivalent molecule contains dGTP nucleotide units, a third type of multivalent molecule contains dCTP nucleotide units, and a fourth type of multivalent molecule contains dTTP nucleotide units. In some embodiments, the four different types of multivalent molecules are labeled with different fluorophores corresponding to the nucleotide units attached to a given multivalent molecule, allowing for identification of the nucleotide units. In some embodiments, the detecting step involves simultaneous or single excitation at wavelengths sufficient to excite all four fluorophores and imaging of the fluorescent emissions at wavelengths sufficient to detect each respective fluorophore. In some embodiments, four labeled multivalent molecules are used to determine the identity of a terminal nucleotide in a nucleic acid template molecule. In some embodiments, the four types of multivalent molecules are labeled with different fluorophores, including fluorophores that emit different visible colors, such as blue, green, yellow, orange, or red. In some embodiments, the four types of multivalent molecules are labeled with different fluorophores, including, for example, Cy2 or a dye or fluorophore with similar excitation or emission properties, Cy3 or a dye or fluorophore with similar excitation or emission properties, Cy3.5 or a dye or fluorophore with similar excitation or emission properties, Cy5 or a dye or fluorophore with similar excitation or emission properties, Cy5.5 or a dye or fluorophore with similar excitation or emission properties, and Cy7 or a dye or fluorophore with similar excitation or emission properties.In some embodiments, the detecting step comprises simultaneous excitation at any two of 532 nm, 568 nm, and 633 nm, respectively, and imaging of fluorescence emission at about 570 nm, 592 nm, 670 nm, and 702 nm. In some embodiments, the fluorescence imaging comprises dual-wavelength excitation / dual-wavelength emission fluorescence imaging. In some embodiments, four different types of multivalent molecules are labeled with distinguishable fluorophores (or sets of fluorophores), and the detecting step comprises simultaneous or single excitation at wavelengths sufficient to excite one, two, three, or four fluorophores or sets of fluorophores, and imaging of fluorescence emission at wavelengths sufficient to detect each respective fluorophore.
[0150] In some embodiments, a two-step sequencing method can be performed with three different types of labeled multivalent molecules and one type of unlabeled multivalent molecule (e.g., a "dark" multivalent molecule), where the labeled multivalent molecules are labeled with different fluorophores corresponding to the nucleotide units attached to a given multivalent molecule, allowing for identification of the nucleotide units. In some embodiments, the detection step involves simultaneous excitation at wavelengths sufficient to excite the three types of fluorophores, imaging of the fluorescence emissions at wavelengths sufficient to detect each respective fluorophore, and detection of the fourth type of multivalent molecule is determined or determinable with reference to the location of the "dark" or unlabeled spot.
[0151] In some embodiments, the fluorophores comprise FRET donor and acceptor pairs, allowing multiple detections and identifications to be performed under a single excitation and imaging step. In some embodiments, a sequencing cycle comprises forming multiple hybridized polymerases, contacting the hybridized polymerases with multiple different types of fluorescently labeled multivalent molecules, and detecting the fluorescently labeled multivalent molecules bound to the hybridized polymerases. In some embodiments, a sequencing cycle can be performed in less than 30 minutes, less than 20 minutes, or less than 10 minutes. In some embodiments, performing a sequencing reaction using labeled multivalent molecules results in an average Q-score of base calling accuracy across the sequencing run of 30 or greater and / or 40 or greater. In some embodiments, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the base calls have a Q-score greater than 30 and / or 40 or greater. In some embodiments, the present disclosure provides methods herein, wherein at least 95% of the base calls have a Q-score greater than 30.
[0152] Sequencing by Binding Methods In some aspects, the present disclosure provides methods for sequencing any of the immobilized concatemeric molecules described herein, wherein the sequencing method comprises a sequencing by binding (SBB) procedure using unlabeled chain-terminating nucleotides. In some embodiments, the sequencing by binding (SBB) method includes the steps of: (a) sequentially contacting a primed template nucleic acid molecule with at least two separate mixtures under ternary complex-stabilizing conditions, each of the at least two separate mixtures comprising a polymerase and a nucleotide, whereby the sequential contacting results in a primed template nucleic acid that has been contacted with nucleotide analogs for first, second, and third base types in the template under ternary complex-stabilizing conditions; (b) examining the at least two separate mixtures to determine whether a ternary complex has formed; and (c) examining the primed template nucleic acid molecule with at least two separate mixtures to determine whether a ternary complex has formed. (d) adding the next correct nucleotide to the primer of the primed template nucleic acid molecule after step (b), thereby generating an extended primer; and (e) repeating steps (a)-(d) at least once on the primed template nucleic acid molecule containing the extended primer. Exemplary sequencing-by-binding methods are described in U.S. Patent Nos. 10,246,744 and 10,731,141, the contents of both patents being incorporated herein by reference in their entireties.
[0153] Methods for sequencing using phosphate chain-labeled nucleotides The present disclosure provides a method for sequencing any of the immobilized concatemeric molecules described herein, wherein the sequencing method comprises step (a): contacting (i) a plurality of sequencing polymerases, (ii) a plurality of concatemeric template molecules immobilized on a support, and (iii) a plurality of nucleic acid sequencing primers (e.g., forward or reverse sequencing primers) with phosphate-chain-labeled nucleotides, wherein the contacting is performed under conditions suitable to form a plurality of multiplexed sequencing polymerase complexes, each complex comprising a sequencing polymerase bound to a nucleic acid duplex, wherein the nucleic acid duplex comprises a portion of the concatemeric template molecules hybridized to the nucleic acid sequencing primer. In some embodiments, the sequencing polymerase comprises a recombinant mutant sequencing polymerase capable of binding to and incorporating nucleotide analogs. In some embodiments, the sequencing polymerase is immobilized on the same support as the concatemeric template molecules. In some embodiments, the sequencing polymerase is not immobilized on a support. In some embodiments, the sequencing primer comprises a 3' extendable end or a 3' blocked end that can be converted to a 3' extendable end.
[0154] In some embodiments, the method for sequencing concatemeric template molecules further includes step (b): contacting a plurality of hybrid sequencing polymerases with a plurality of phosphate-chain-labeled nucleotides under conditions suitable for binding at least one phosphate-chain-labeled nucleotide to one of the hybrid sequencing polymerases, the conditions being suitable for promoting polymerase-catalyzed nucleotide incorporation. In some embodiments, the hybrid sequencing polymerase is contacted with the plurality of nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, each phosphate-chain-labeled nucleotide in the plurality of phosphate-chain-labeled nucleotides comprises a phosphate chain comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and 3 to 20 (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) phosphate groups, wherein the terminal phosphate group is attached to a detectable reporter moiety (e.g., a fluorophore). The first, second, and third phosphate groups may be referred to as alpha, beta, and gamma phosphate groups. In some embodiments, a specific detectable reporter moiety attached to the terminal phosphate group corresponds to a nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to enable detection and identification of the nucleic acid base. In some embodiments, a sequencing polymerase is capable of binding a complementary phosphate-strand-labeled nucleotide and incorporating the complementary nucleotide opposite the nucleotide in the template molecule. In some embodiments, a polymerase-catalyzed nucleotide incorporation reaction cleaves between the alpha and beta phosphate groups, thereby releasing a polyphosphate strand attached to the detectable reporter moiety. In some embodiments, the multiple phosphate-strand-labeled nucleotides comprise one type or a mixture of any two or more types of nucleotides, including dATP, dGTP, dCTP, dTTP, and / or dUTP.
[0155] In some embodiments, the sequencing method further comprises step (c): detecting a fluorescent signal emitted by a phosphate-chain-labeled nucleotide bound by the sequencing polymerase and incorporated onto the end of the sequencing primer. In some embodiments, step (c) further comprises identifying the phosphate-chain-labeled nucleotide bound by the sequencing polymerase and incorporated onto the end of the sequencing primer.
[0156] In some embodiments, the sequencing method further comprises step (d): repeating steps (b)-(c) at least once. In some embodiments, the sequencing method using phosphate-chain-labeled nucleotides can be performed according to the methods described in U.S. Patent Nos. 7,170,050, 7,302,146, and / or 7,405,281.
[0157] In some embodiments, in step (a), the plurality of concatemeric template molecules are immobilized on a support comprising a plurality of separate compartments. In some embodiments, the plurality of sequencing polymerases are in solution in the compartments. In some embodiments, at least one sequencing polymerase is immobilized on the bottom of an individual compartment. In some embodiments, the separate compartments comprise a silica bottom through which light can pass. In some embodiments, the separate compartments comprise a silica bottom comprised of a nanophotonic confinement structure comprising a hole in a metal cladding film (e.g., an aluminum cladding film). In some embodiments, the hole in the metal cladding has a small opening, for example, approximately 70 nm. In some embodiments, the height of the nanophotonic confinement structure is approximately 100 nm. In some embodiments, the nanophotonic confinement structure comprises a zero-mode waveguide (ZMW). In some embodiments, the nanophotonic confinement structure contains a liquid.
[0158] In some embodiments, the covalently closed circular library molecules (600) can function as non-immobilized template molecules. In some embodiments, the sequencing method includes step (a): providing a support having a plurality of sequencing polymerases immobilized thereon. In some embodiments, the sequencing polymerase comprises a processive DNA polymerase. In some embodiments, the sequencing polymerase comprises a wild-type or mutant DNA polymerase, including, for example, Phi29 DNA polymerase. In some embodiments, the support comprises a plurality of separate compartments, and the sequencing polymerase is immobilized at the bottom of the compartment. In some embodiments, the separate compartments comprise a silica bottom through which light can pass. In some embodiments, the separate compartments comprise a silica bottom comprised of a nanophotonic confinement structure comprising a hole in a metal cladding film (e.g., an aluminum cladding film). In some embodiments, the hole in the metal cladding has a small opening, for example, approximately 70 nm. In some embodiments, the height of the nanophotonic confinement structure is approximately 100 nm. In some embodiments, the nanophotonic confinement structure comprises a zero-mode waveguide (ZMW). In some embodiments, the nanophotonic confinement structure contains a liquid.
[0159] In some embodiments, the sequencing method further comprises step (b): contacting a plurality of immobilized sequencing polymerases with a plurality of single-stranded circular nucleic acid template molecules (e.g., covalently closed circular library molecules (600)) and a plurality of oligonucleotide sequencing primers under conditions suitable for each immobilized sequencing polymerase to bind to the single-stranded circular template molecule and for each sequencing primer to hybridize to each single-stranded circular template molecule, thereby generating a plurality of polymerase / template / primer complexes. In some embodiments, each sequencing primer hybridizes to a universal sequencing primer binding site on the single-stranded circular template molecule.
[0160] In some embodiments, the sequencing method further includes step (c): contacting the plurality of polymerase / template / primer complexes with a plurality of phosphate-chain-labeled nucleotides, each nucleotide comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and a phosphate chain comprising 3 to 20 phosphate groups, wherein the terminal phosphate group is attached to a detectable reporter moiety (e.g., a fluorophore). The first, second, and third phosphate groups may be referred to as alpha, beta, and gamma phosphate groups. In some embodiments, the specific detectable reporter moieties attached to the terminal phosphate groups correspond to the nucleotide bases (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to enable detection and identification of the nucleobase. In some embodiments, the plurality of polymerase / template / primer complexes are contacted with the plurality of phosphate-chain-labeled nucleotides under conditions suitable for polymerase-catalyzed nucleotide incorporation. In some embodiments, a sequencing polymerase can bind to a complementary phosphate-chain-labeled nucleotide and incorporate a complementary nucleotide opposite the nucleotide in the template molecule. In some embodiments, a polymerase-catalyzed nucleotide incorporation reaction cleaves between the alpha and beta phosphate groups, thereby releasing multiple phosphate chains attached to fluorophores.
[0161] In some embodiments, the sequencing method further comprises step (d): detecting a fluorescent signal emitted by a phosphate-chain-labeled nucleotide bound by the sequencing polymerase and incorporated onto the end of the sequencing primer. In some embodiments, step (d) further comprises identifying the phosphate-chain-labeled nucleotide bound by the sequencing polymerase and incorporated onto the end of the sequencing primer.
[0162] In some embodiments, the sequencing method further comprises step (d): repeating steps (c) through (d) at least once. In some embodiments, the sequencing method using phosphate-chain-labeled nucleotides can be performed according to the methods described in U.S. Patent Nos. 7,170,050, 7,302,146, and / or 7,405,281.
[0163] Sequencing polymerase In any of the methods described herein, a sequencing polymerase can be used to perform the sequencing reaction. In some embodiments, the sequencing polymerase(s) are capable of binding to and incorporating complementary nucleotides opposite the nucleotides in the concatemeric template molecule. In some embodiments, the sequencing polymerase(s) are capable of binding to complementary nucleotide units of the multivalent molecule opposite the nucleotides in the concatemeric template molecule. In some embodiments, the plurality of sequencing polymerases comprises recombinant mutant polymerases.
[0164] Examples of polymerases suitable for use in sequencing nucleotides and / or polyvalent molecules include Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales archaea; Candidatus Hadarchaeum Yellowstonense; Hadesarchaea archaea; Euryarchaeota archaea; Thermoplasmata archaea; Thermococcus polymerases, e.g., Thermococcus litoralis, bacteriophage T7 DNA polymerase; human alpha, delta, and epsilon DNA polymerases; bacteriophage polymerases, e.g., T4, RB69, and phi29 bacteriophage DNA polymerases; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; E. coli DNA polymerase III alpha and epsilon; 9 degree Examples of DNA polymerases include, but are not limited to, N polymerase; reverse transcriptases such as HIV-type M or O reverse transcriptase; avian myeloblastosis virus reverse transcriptase; Moloney murine leukemia virus (MMLV) reverse transcriptase; or telomerase. Further non-limiting examples of DNA polymerases include those from various archaeal genera, such as Aeropyrum, Archaeglobus, Desulfurococcus, Pyrobaculum, Pyrococcus, Pyrolobus, Pyrodictium, Staphylothermus, Stetteria, Sulfolobus, Thermococcus, and Vulcanisaeta, or variants thereof, including such polymerases known in the art, such as 9 degrees N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo, and RB69 polymerases.
[0165] Nucleotides and Chain-Terminating Nucleotides In any of the methods described herein, any of the sequencing methods described herein can use at least one nucleotide. A nucleotide comprises a base, a sugar, and at least one phosphate group. In some embodiments, at least one nucleotide in the plurality of nucleotides comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 to 10 phosphate groups). The plurality of nucleotides can comprise at least one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The plurality of nucleotides can comprise a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, at least one nucleotide in the plurality of nucleotides is not a nucleotide analog. In some embodiments, at least one nucleotide in the plurality of nucleotides comprises a nucleotide analog.
[0166] In some embodiments, in any of the sequencing methods described herein, at least one nucleotide of the plurality of nucleotides comprises a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramide linkage. In some embodiments, at least one nucleotide in the plurality of nucleotides is an analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise a substituted side chain group comprising O, S, or BH3. In some embodiments, the chain comprises a phosphate group substituted with an analog comprising a phosphoramidate, phosphorothioate, phosphorodithioate, and O-methylphosphoramidite group.
[0167] In some embodiments, in any of the methods for sequencing described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog, wherein the terminator nucleotide analog has a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain-terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide in the nascent strand during a primer extension reaction. In some embodiments, the chain-terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain-terminating moiety is removable / cleavable from the 3' sugar hydroxyl position to generate a nucleotide having a 3' OH sugar group that is extendable with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, an acetal group, or a silyl group. In some embodiments, the chain-terminating moiety is cleavable / removable from the nucleotide, for example, by reacting the chain-terminating moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the chain-terminating moieties alkyl, alkenyl, alkynyl, and aryl can be cleaved using tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or 2,3-dichloro-5,6-dicyano-1,4-benzo-quinone (DDQ). In some embodiments, the chain-terminating moieties aryl and benzyl can be cleaved with HPd / C. In some embodiments, the chain terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphines or thiol groups, including beta-mercaptoethanol or dithiothritol (DTT).In some embodiments, the chain-terminating carbonate moiety can be cleaved using potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn in acetic acid (AcOH). In some embodiments, the chain-terminating urea and silyl moieties can be cleaved with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0168] In some embodiments, in any of the methods for sequencing described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog, wherein the terminator nucleotide analog has a chain-terminating moiety (e.g., a blocking moiety) at the 2' sugar position, the 3' sugar position, or the 2' and 3' sugar positions. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or 3'-O-azidomethyl group. In some embodiments, the chain-terminating azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0169] In some embodiments, in any of the methods for sequencing described herein, the nucleotide comprises a chain-terminating moiety selected from the group consisting of 3'-deoxynucleotides, 2',3'-dideoxynucleotides, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino groups, 3'-phosphorothioates, and 3-O-benzyl, or derivatives thereof.
[0170] In some embodiments, in any of the methods for sequencing described herein, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base with a linker, the linker being cleavable / removable from the base. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to enable detection and identification of the nucleotide base.
[0171] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, the cleavable linker on the nucleotide base comprises a cleavable moiety comprising an alkyl group, an alkenyl group, an alkynyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, an acetal group, or a silyl group. In some embodiments, the cleavable linker on the base is cleavable / removable from the base by reacting the cleavable moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the cleavable moieties alkyl, alkenyl, alkynyl, and aryl are cleavable using tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or 2,3-dichloro-5,6-dicyano-1,4-benzo-quinone (DDQ). In some embodiments, aryl and benzyl cleavable moieties are cleavable with HPd / C. In some embodiments, amine, amide, keto, isocyanate, phosphate, thio, and disulfide cleavable moieties are cleavable with phosphines or thiol groups, including beta-mercaptoethanol or dithiothritol (DTT). In some embodiments, carbonate cleavable moieties are cleavable with potassium carbonate (KCO) in MeOH, triethylamine in pyridine, or Zn in acetic acid (AcOH). In some embodiments, urea and silyl cleavable moieties are cleavable with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0172] In some embodiments, in any of the methods for sequencing described herein, the cleavable linker on the nucleotide base comprises a cleavable moiety including an azide, azido, or azidomethyl group. In some embodiments, the cleavable moieties azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0173] In some embodiments, in any of the methods for sequencing described herein, the chain-terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the cleavable linker on the nucleotide base have the same or different cleavable moieties. In some embodiments, the chain-terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the detectable reporter moiety attached to the base are chemically cleavable / removable with the same chemical agent. In some embodiments, the chain-terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the detectable reporter moiety attached to the base are chemically cleavable / removable with different chemical agents.
[0174] Multivalent molecules In any of the methods described herein, sequencing uses at least one multivalent molecule comprising a plurality of nucleotide arms attached to a core and having any configuration, including starburst, helter-skelter, or bottlebrush configurations (e.g., Figure 2). In some embodiments, the multivalent molecule comprises (1) a core and (2) a plurality of nucleotide arms, the plurality of nucleotide arms comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit. In some embodiments, the nucleotide unit comprises a base, a sugar, and at least one phosphate group, and the linker is attached to the nucleotide unit via the base. In some embodiments, the linker comprises an aliphatic chain or an oligoethylene glycol chain, and both linker chains have 2 to 6 subunits. In some embodiments, the linker also comprises an aromatic moiety. Exemplary nucleotide arms are shown in Figure 6. Exemplary multivalent molecules are shown in Figures 2-5. An exemplary spacer is shown in Figure 7 (top), and an exemplary linker is shown in Figure 7 (bottom) and Figure 8. Exemplary nucleotides attached to linkers are shown in Figures 9A-9D. An exemplary biotinylated nucleotide arm is shown in Figure 10.
[0175] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, the plurality of nucleotide arms having the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
[0176] In some embodiments, a multivalent molecule comprises a core attached to multiple nucleotide arms, each arm comprising a nucleotide unit. The nucleotide unit comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., about 1-10 phosphate groups). The multiple multivalent molecules can comprise one type of multivalent molecule having one type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The multiple multivalent molecules can be comprised in any combination of mixtures of two or more types of multivalent molecules, each individual multivalent molecule in the mixture comprising a nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP.
[0177] In some embodiments, the nucleotide unit comprises a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramidate bond. In some embodiments, at least one nucleotide unit is a nucleotide analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise substituted side chain groups including O, S, or BH. In some embodiments, the chain comprises phosphate groups substituted with analogs including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoramidite groups.
[0178] In some embodiments, the multivalent molecule comprises a core attached to multiple nucleotide arms, each nucleotide arm comprising a nucleotide unit, the nucleotide unit being a nucleotide analog having a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the nucleotide unit comprises a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain-terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide into a nascent chain during a primer extension reaction. In some embodiments, the chain-terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain-terminating moiety is removable / cleavable from the 3' sugar hydroxyl position to generate a nucleotide having a 3'OH sugar group that is extendable with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, an acetal group, or a silyl group. In some embodiments, the chain-terminating moiety is cleavable / removable from the nucleotide unit, for example, by reacting the chain-terminating moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the chain-terminating moieties alkyl, alkenyl, alkynyl, and aryl can be cleaved using tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or 2,3-dichloro-5,6-dicyano-1,4-benzo-quinone (DDQ). In some embodiments, the chain-terminating moieties aryl and benzyl can be cleaved with HPd / C. In some embodiments, the chain terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphines or thiol groups, including beta-mercaptoethanol or dithiothritol (DTT).In some embodiments, the chain-terminating carbonate moiety can be cleaved using potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn in acetic acid (AcOH). In some embodiments, the chain-terminating urea and silyl moieties can be cleaved with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0179] In some embodiments, the nucleotide unit comprises a chain-terminating moiety (e.g., a blocking moiety) at the 2' sugar position, the 3' sugar position, or the 2' and 3' sugar positions. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or 3'-O-azidomethyl group. In some embodiments, the chain-terminating azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP).
[0180] In some embodiments, a nucleotide unit comprising a chain-terminating moiety selected from the group consisting of 3'-deoxynucleotide, 2',3'-dideoxynucleotide, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino group, 3'-phosphorothioate, and 3-O-benzyl, or a derivative thereof.
[0181] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, the nucleotide arms comprising spacers, linkers, and nucleotide units, and the core, linkers, and / or nucleotide units are labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) attached to the multivalent molecule can correspond to a base of a nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and identification of the nucleotide base.
[0182] In some embodiments, at least one nucleotide arm of the multivalent molecule has a nucleotide unit attached to a detectable reporter moiety. In some embodiments, the detectable reporter moiety is attached to a nucleotide base. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the multivalent molecule can correspond to the base of the nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and identification of the nucleotide base.
[0183] In some embodiments, the core of the multivalent molecule comprises an avidin-like or streptavidin-like moiety, and the core-attached moiety comprises biotin. In some embodiments, the core comprises a streptavidin-type or avidin-type moiety, which includes avidin protein, as well as any derivatives, analogs, and other non-natural forms of avidin, capable of binding to at least one biotin moiety. Other forms of avidin moieties include native and recombinant avidin and streptavidin, as well as derivatized molecules, such as non-glycosylated avidin and truncated streptavidin. For example, the avidin moiety can include deglycosylated forms of avidin, bacterial streptavidin produced by Streptomyces (e.g., Streptomyces avidinii), as well as derivatized forms such as N-acylavidins, e.g., N-acetyl, N-phthalyl, and N-succinyl avidin, and the commercially available products EXTRAVIDIN™, CAPTAVIDIN™, NEUTRAVIDIN™, and NEUTRALITE AVIDIN™.
[0184] In some embodiments, any of the methods for sequencing a nucleic acid molecule described herein can include forming a binding complex, wherein the binding complex comprises (i) a polymerase, a nucleic acid concatemer molecule duplexed with a primer, and a nucleotide, or the binding complex comprises (ii) a polymerase, a nucleic acid concatemer molecule duplexed with a primer, and a nucleotide unit of a multivalent molecule. In some embodiments, the binding complex has a duration of greater than about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. The binding complex has a duration of greater than about 0.1-0.25 seconds, or about 0.25-0.5 seconds, or about 0.5-0.75 seconds, or about 0.75-1 second, or about 1-2 seconds, or about 2-3 seconds, or about 3-4 seconds, or about 4-5 seconds, and / or the method is performed or can be performed at a temperature of 15°C or higher, 20°C or higher, 25°C or higher, 35°C or higher, 37°C or higher, 42°C or higher, 55°C or higher, 60°C or higher, 72°C or higher, or 80°C or higher, or within a range defined by any of the foregoing. The binding complex (e.g., ternary complex) remains stable until subjected to conditions that cause dissociation of interactions between the polymerase, template molecule, primer, and / or any of the nucleotide units or nucleotides. For example, dissociation conditions include contacting the binding complex with any one of detergent, EDTA, and / or water, or any combination thereof. In some embodiments, the disclosure provides the method, wherein the binding complex is deposited on, attached to, or hybridized to a surface that, in the detecting step, exhibits a contrast-to-noise ratio of greater than 20. In some embodiments, the disclosure provides the method, wherein the contacting is performed under conditions that stabilize the binding complex when the nucleotide or nucleotide unit is complementary to the next base of the template nucleic acid and destabilize the binding complex when the nucleotide or nucleotide unit is not complementary to the next base of the template nucleic acid.
[0185] Automatic Mode In any of the methods described herein, any combination of steps for preparing spatially resolved nucleic acids can be performed in an automated mode using a fluid delivery system, including cell seeding, cell fixation, cell permeabilization, reverse transcription reaction, circularized oligonucleotide hybridization, circularized oligonucleotide ligation reaction, rolling circle amplification, and sequencing.
[0186] In some embodiments, the cell sample is deposited on a flow cell (e.g., a support). The flow cell may be coated with a reagent that promotes cell adhesion or fixation to the flow cell. The flow cell with the cell sample attached thereto may be placed on a sequencing instrument having a flow cell holder / cradle fluidically connected to an automated fluid distribution system and configured on a fluorescence microscope. In some embodiments, the sequencing instrument may be configured with at least one fluid delivery device, at least one fluidic device (e.g., a microfluidic device), at least one imaging device, and / or at least one sensor to detect signals from the sequencing reaction.
[0187] In some embodiments, an automated fluid distribution system can be used to deliver fixation reagents to the cell sample on the flow cell, and the cell sample can be incubated under conditions suitable for cell fixation.
[0188] In some embodiments, an automated fluid distribution system can be used to deliver permeabilization reagents to a fixed cell sample on a flow cell, and the cell sample can be incubated under conditions suitable for cell permeabilization.
[0189] In some embodiments, an automated fluid distribution system can be used to deliver reagents for reverse transcription of RNA onto a flow cell under conditions suitable for generating multiple cDNAs immobilized on the flow cell.
[0190] In some embodiments, an automated fluid distribution system can be used to deliver reagents for circularizing oligonucleotide hybridization and ligation under conditions suitable for generating multiple covalently closed circular molecules on a flow cell.
[0191] In some embodiments, an automated fluid distribution system can be used to deliver reagents for performing rolling circle amplification under conditions suitable for generating a plurality of concatemeric molecules immobilized on a flow cell.
[0192] In some embodiments, an automated fluid distribution system can be used to deliver reagents for sequencing cycles of concatemeric molecules under conditions suitable for generating multiple sequencing lead products on a flow cell. In some embodiments, individual cycle times can be achieved in less than 30 minutes. In some embodiments, the field of view (FOV) can be 1 mm or less. 2 It can exceed the limit for large areas (>10mm 2 The cycle time for scanning the ) can be less than 5 minutes.
[0193] In some embodiments, an automated fluid distribution system can be used to deliver reagents to remove multiple sequencing lead products from concatemer molecules and retain the concatemer molecules on the flow cell.
[0194] Support and low non-specific coating In some aspects, the present disclosure provides pairwise sequencing compositions and methods, which use a support comprising a plurality of oligonucleotide surface primers immobilized thereon. In some embodiments, the support is passivated with a low non-specific binding coating. In some embodiments, the surface coatings described herein exhibit very low non-specific binding to reagents typically used for nucleic acid capture, amplification, and sequencing workflows, such as dyes, nucleotides, enzymes, and nucleic acid primers. The surface coatings may exhibit a low background fluorescence signal or a high contrast-to-noise (CNR) ratio compared to conventional surface coatings.
[0195] In some embodiments, the low non-specific binding coating comprises one or more layers (FIG. 1). In some embodiments, multiple surface primers are immobilized on the low non-specific binding coating. In some embodiments, at least one surface primer is embedded within the low non-specific binding coating. In some embodiments, the low non-specific binding coating enables improved nucleic acid hybridization and amplification performance. In some embodiments, the support comprises a substrate (or support structure), one or more covalently or non-covalently attached low-binding chemically modified layers, e.g., a silane layer, a polymer film, and one or more covalently or non-covalently attached surface primers that can be used to designate single-stranded nucleic acid library molecules to the support. In some embodiments, the coating formulation, e.g., the chemical composition of one or more layers, the coupling chemistry used to crosslink one or more layers to the support and / or to each other, and the total number of layers, can be varied to minimize or reduce non-specific binding of proteins, nucleic acid molecules, and other hybridization and amplification reaction components to the coating relative to an equivalent monolayer. The coating formulations described herein can be varied to minimize or reduce nonspecific hybridization on the coating relative to a comparable monolayer. The coating formulations can be varied to minimize or reduce nonspecific amplification on the coating relative to a comparable monolayer. The coating formulations can be varied to maximize specific amplification rate and / or yield on the coating. Amplification levels suitable for detection are achieved in 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30 or fewer amplification cycles, or more than 30, in some cases disclosed herein.
[0196] The support structure comprising one or more chemically modified layers, e.g., layers of low nonspecific binding polymers, may be freestanding or integrated within another structure or assembly. For example, in some embodiments, the support structure may comprise one or more surfaces within an integrated or assembled microfluidic flow cell. The support structure may comprise one or more surfaces within a microplate format, e.g., the bottom surface of a well in a microplate. In some embodiments, the support structure comprises the inner surface (e.g., lumen surface) of a capillary. In some embodiments, the support structure comprises the inner surface (e.g., lumen surface) of a capillary etched into a planar chip.
[0197] The bonding chemistry used to graft the first chemically modified layer onto the surface of the support generally depends on both the material from which the surface is fabricated and the chemical nature of the layer. In some embodiments, the first layer may be covalently attached to the surface. In some embodiments, the first layer may be noncovalently attached, e.g., adsorbed, to the support via noncovalent interactions between the support and the molecular components of the first layer, e.g., electrostatic interactions, hydrogen bonding, or van der Waals interactions. In either case, the support may be treated prior to the attachment or deposition of the first layer. Any of a variety of surface preparation techniques known to those skilled in the art may be used to clean or treat the surface. For example, glass or silicon surfaces may be acid-cleaned using piranha solution (a mixture of sulfuric acid (H2SO4) and hydrogen peroxide (H2O2)), base treatment in KOH and NaOH, and / or cleaned using an oxygen plasma treatment method.
[0198] Silane chemistry provides a non-limiting approach for covalently modifying silanol groups on glass or silicon surfaces to attach more reactive functional groups (e.g., amine or carboxyl groups), which can then be used in coupling linker molecules (e.g., linear hydrocarbon molecules of various lengths, e.g., C6, C12, C18 hydrocarbons, or linear polyethylene glycol (PEG) molecules) or layer molecules (e.g., branched PEG molecules or other polymers) to the surface. Examples of suitable silanes that can be used in making any of the disclosed low-binding coatings include, but are not limited to, (3-aminopropyl)trimethoxysilane (APTMS), (3-aminopropyl)triethoxysilane (APTES), any of the various PEG silanes (e.g., those with molecular weights of 1K, 2K, 5K, 10K, 20K, etc.), amino-PEG silane (i.e., those containing a free amino functional group), maleimide-PEG silane, and biotin-PEG silane.
[0199] Any of a variety of molecules known to those of skill in the art can be used in creating one or more chemically modified layers on a support, including, but not limited to, amino acids, peptides, nucleotides, oligonucleotides, other monomers, or polymers, or combinations thereof, and the selection of components used can be varied to modify one or more properties of the layer, such as the surface density of functional groups and / or tethered oligonucleotide primers, the hydrophilicity / hydrophobicity of the layer, or the three-dimensional nature (i.e., "thickness") of the layer. Examples of polymers that can be used to create one or more layers of low nonspecific binding material in any of the disclosed coatings include, but are not limited to, polyethylene glycol (PEG) of various molecular weights and branched structures, streptavidin, polyacrylamide, polyester, dextran, polylysine, and polylysine copolymers, or any combination thereof. Examples of conjugation chemistries that can be used to graft one or more layers of material (e.g., polymer layers) onto a surface and / or crosslink layers to one another include, but are not limited to, biotin-streptavidin interactions (or variations thereof), his-tag-Ni / NTA conjugation chemistry, methoxy ether conjugation chemistry, carboxylate conjugation chemistry, amine conjugation chemistry, NHS esters, maleimides, thiols, epoxies, azides, hydrazides, alkynes, isocyanates, and silanes.
[0200] The low non-specific binding surface coating can be applied uniformly across the entire support. Alternatively, the surface coating can be patterned so that the chemically modified layer is limited to one or more individual regions of the support. For example, the coating can be patterned using photolithography techniques to create an ordered array or random pattern of chemically modified regions on the support. Alternatively or in combination, the coating can be patterned using, for example, contact printing and / or inkjet printing techniques. In some embodiments, the ordered array or random pattern of chemically modified regions can include at least 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 or more individual regions.
[0201] In some embodiments, the low nonspecific binding coating comprises a hydrophilic polymer that is nonspecifically adsorbed or covalently grafted to the support. Typically, passivation is performed using poly(ethylene glycol) (PEG, also known as polyethylene oxide (PEO) or polyoxyethylene) or other hydrophilic polymers with different molecular weights and end groups attached to the support, for example, using silane chemistry. End groups distal to the surface can include, but are not limited to, biotin, methoxy ether, carboxylate, amine, NHS ester, maleimide, and bis-silane. In some embodiments, two or more layers of hydrophilic polymers, such as linear, branched, or hyperbranched polymers, can be deposited on the surface. In some embodiments, the two or more layers can be covalently coupled to each other or internally crosslinked to improve the stability of the resulting coating. In some embodiments, surface primers with different nucleotide sequences and / or base modifications (or other biomolecules, such as enzymes or antibodies) can be tethered to the resulting layer at various surface densities. In some embodiments, for example, both the surface functional group density and the surface primer concentration can be varied to achieve a desired surface primer density range. In addition, the surface primer density can be controlled by diluting the surface primer with other molecules that carry the same functional group. For example, but not limited to, amine-labeled surface primers can be diluted with amine-labeled polyethylene glycol in reaction with an NHS-ester-coated surface to reduce the final primer density. Also, surface primers with different linker lengths between the hybridization region and the surface-attached functional group can be applied to control the surface density.Examples of suitable linkers include poly-T and poly-A tethers (e.g., 0-20 bases, e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) at the 5' end of the primer, PEG linkers (e.g., 3-20 monomer units, e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 monomer units), and carbon chains (e.g., C6, C12, C18, etc.) To measure primer density, fluorescently labeled primers can be tethered to a surface and the fluorescent readings can then be compared to those for dye solutions of known concentration.
[0202] In some embodiments, the low nonspecific binding coating comprises a functionalized polymer coating layer covalently bonded to at least a portion of the substrate via chemical groups on the substrate, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some embodiments, the functionalized polymer coating comprises poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide (PAZAM).
[0203] To tailor primer surface density and add additional dimensionality to hydrophilic or amphoteric coatings, supports containing multilayer coatings of PEG and other hydrophilic polymers have been developed. By using hydrophilic and amphoteric surface layering techniques, including but not limited to the polymer / copolymer materials described below, it is possible to significantly increase primer loading density on supports. Conventional PEG coating techniques use monolayer primer deposition, which has generally been reported for single molecule applications but does not result in high copy numbers for nucleic acid amplification applications. As described herein, "layering" can be achieved with any compatible polymer or monomer subunit, such that surfaces containing two or more highly crosslinked layers can be sequentially constructed using conventional crosslinking techniques. Examples of suitable polymers include, but are not limited to, streptavidin, polyacrylamide, polyester, dextran, polylysine, and copolymers of polylysine and PEG. In some embodiments, the different layers may be attached to one another via any of a variety of conjugation reactions, including, but not limited to, biotin-streptavidin binding, azide-alkyne click reactions, amine-NHS ester reactions, thiol-maleimide reactions, and ionic interactions between positively and negatively charged polymers. In some embodiments, the high primer density material may be constructed in solution and then layered onto a surface in multiple steps.
[0204] Examples of materials from which the support structure may be fabricated include, but are not limited to, glass, fused silica, silicon, polymers (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic support structures are contemplated.
[0205] The support structure can be any of a variety of geometries and dimensions known to those of skill in the art and can comprise any of a variety of materials known to those of skill in the art. For example, the support structure can be locally planar (e.g., including a microscope slide or the surface of a microscope slide). Generally, the support structure can be cylindrical (e.g., including a capillary or the interior surface of a capillary), spherical (e.g., including the exterior surface of a non-porous bead), or irregular (e.g., including the exterior surface of an irregularly shaped non-porous bead or particle). In some embodiments, the surface of the support structure used for nucleic acid hybridization and amplification can be a solid, non-porous surface. In some embodiments, the surface of the support structure used for nucleic acid hybridization and amplification can be porous, such that the coatings described herein can permeate the porous surface and nucleic acid hybridization and amplification reactions performed thereon can occur within the pores.
[0206] The support structure comprising one or more chemically modified layers, e.g., layers of low nonspecific binding polymers, can be freestanding or integrated within another structure or assembly. For example, the support structure can comprise one or more surfaces within an integrated or assembled microfluidic flow cell. The support structure can comprise one or more surfaces within a microplate format, e.g., the bottom surface of a well in a microplate. In some embodiments, the support structure comprises the inner surface (e.g., lumen surface) of a capillary. In some embodiments, the support structure comprises the inner surface (e.g., lumen surface) of a capillary etched into a planar chip.
[0207] As described herein, the low non-specific binding supports of the present disclosure exhibit reduced non-specific binding of proteins, nucleic acids, and other components of hybridization and / or amplification formulations used for solid-phase nucleic acid amplification. The degree of non-specific binding exhibited by a given support surface can be assessed either qualitatively or quantitatively. Exposing a surface to, for example, a fluorescent dye (e.g., a cyanine, such as Cy3 or Cy5, fluorescein, coumarin, rhodamine, or other dyes disclosed herein), fluorescently labeled nucleotides, fluorescently labeled oligonucleotides, and / or fluorescently labeled proteins (e.g., polymerase) under a standardized set of conditions, followed by a specific rinsing protocol and fluorescent imaging, can be used as a qualitative tool for comparing non-specific binding on supports containing different surface formulations. In some embodiments, exposing a surface under a standardized set of conditions, e.g., to a fluorescent dye, fluorescently labeled nucleotide, fluorescently labeled oligonucleotide, and / or fluorescently labeled protein (e.g., polymerase), followed by a specific rinsing protocol and fluorescent imaging, can be used as a quantitative tool for comparing nonspecific binding on supports comprising different surface formulations, provided that care is taken to ensure that the fluorescent imaging is performed under conditions where the fluorescent signal is linearly related (or related in a predictable manner) to the number of fluorophores on the support (e.g., under conditions where signal saturation and / or self-quenching of the fluorophores is not an issue) and that suitable calibration standards are used. In some embodiments, other techniques known to those skilled in the art, e.g., radioisotope labeling and counting methods, can be used for quantitative assessment of the degree to which nonspecific binding is exhibited by different support surface formulations of the present disclosure.
[0208] Some surfaces disclosed herein exhibit a ratio of specific to nonspecific binding of a fluorophore, such as Cy3, of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value encompassed by the ranges described herein. Some surfaces disclosed herein exhibit a ratio of specific to nonspecific fluorescence of a fluorophore, such as Cy3, of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value encompassed by the ranges described herein.
[0209] In some embodiments, the degree of nonspecific binding exhibited by the disclosed low-binding supports can be assessed using a standardized protocol for contacting the surface with a labeled protein (e.g., bovine serum albumin (BSA), streptavidin, DNA polymerase, reverse transcriptase, helicase, single-stranded binding protein (SSB), etc., or any combination thereof), labeled nucleotide, labeled oligonucleotide, etc., under a standardized set of incubation and rinsing conditions, followed by detection of the amount of label remaining on the surface and comparison of the resulting signal to an appropriate calibration standard. In some embodiments, the label can comprise a fluorescent label. In some embodiments, the label can comprise a radioisotope. In some embodiments, the label can comprise any other detectable label known to those of skill in the art. Thus, in some embodiments, the degree of nonspecific binding exhibited by a given support surface formulation can be assessed in terms of the number of nonspecifically bound protein molecules (or nucleic acid molecules, or other molecules) per unit area. In some embodiments, the low-binding supports of the present disclosure can be assessed in terms of the number of nonspecifically bound protein molecules (or nucleic acid molecules, or other molecules) per unit area. 2 Less than 0.001 molecules per 1 μm 2 Less than 0.01 molecules per 1 μm 2 Less than 0.1 molecules per 1 μm 2Less than 0.25 molecules per 1 μm 2 Less than 0.5 molecules per 1 μm 2 Less than one molecule per 1 μm 2 Less than 10 molecules per 1 μm 2 Less than 100 molecules per 1 μm 2 A given support surface of the present disclosure may exhibit nonspecific protein binding (or nonspecific binding of other specific molecules (e.g., cyanines, such as Cy3 or Cy5, fluoresceins, coumarins, rhodamines, or other dyes disclosed herein) of less than 1,000 molecules per μm. 2 Those skilled in the art will recognize that the nonspecific binding of some modified surfaces may be less than 86 molecules / μm. For example, some modified surfaces disclosed herein exhibit nonspecific binding of 0.5 molecules / μm after contact with a 1 μM solution of Cy3-labeled streptavidin (GE Amersham) in phosphate-buffered saline (PBS) buffer for 15 minutes, followed by rinsing three times with deionized water. 2 Some modified surfaces disclosed herein exhibit non-specific protein binding of less than 1 μm 2The nonspecific binding of less than 0.25 molecules of Cy3 dye per sample was demonstrated in an independent nonspecific binding assay using 1 μM labeled Cy3 SA (ThermoFisher), 1 μM Cy5 SA dye (ThermoFisher), 10 μM aminoallyl-dUTP-ATTO-647N (Jena Biosciences), 10 μM aminoallyl-dUTP-ATTO-Rhol 1 (Jena Biosciences), 10 μM aminoallyl-dUTP-ATTO-Rhol 1 (Jena Biosciences), 10 μM 7-propargylamino-7-deaza-dGTP-Cy5 (Jena Biosciences), and 10 μM 7-propargylamino-7-deaza-dGTP-Cy3 (Jena Biosciences). Biosciences) was incubated in a 384-well plate format on a low-binding coated support for 15 minutes at 37°C. Each well was rinsed two to three times with 50 μL of deionized RNase / DNase-free water and two to three times with 25 mM ACES buffer, pH 7.4. The 384-well plates were imaged in a GE Typhoon instrument at a PMT gain setting of 800 and a resolution of 50–100 μm using a Cy3, AF555, or Cy5 filter set as specified by the manufacturer (according to the dye test performed). For higher-resolution imaging, images were captured using a total internal reflection fluorescence (TIRF) objective (100×, 1.5 NA, Olympus), a CCD camera (e.g., Olympus EM-CCD monochrome camera, Olympus XM-10 monochrome camera, or Olympus DP80 color and monochrome camera), and an illumination source (e.g., Olympus 100W Images were collected on an Olympus IX83 microscope (e.g., an inverted fluorescence microscope) (Olympus Corp., Center Valley, Pa.) with a 532 nm or 635 nm excitation wavelength and a 500 W Xe lamp, an Olympus 75 W Xe lamp, or an Olympus U-HGLGPS fluorescent light source.Dichroic mirrors were purchased from Semrock (IDEX Health & Science, LLC, Rochester, NY) and were dichroic reflectors / beam splitters of, for example, 405, 488, 532, or 633 nm, and bandpass filters were selected as 532LP or 645LP to match the appropriate excitation wavelength. Some modified surfaces disclosed herein have a 1 μm bandpass filter. 2 This shows non-specific binding of less than 0.25 dye molecules per 1000. In some embodiments, the coated substrate was immersed in a buffer solution (e.g., 25 mM ACES, pH 7.4) while the image was being acquired.
[0210] In some embodiments, surfaces disclosed herein exhibit a ratio of specific to nonspecific binding of a fluorophore, such as Cy3, of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value encompassed by the ranges described herein. In some embodiments, surfaces disclosed herein exhibit a ratio of specific to nonspecific fluorescent signal for a fluorophore, such as Cy3, of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 75, 100, or greater than 100, or any intermediate value encompassed by the ranges described herein.
[0211] Low background surfaces consistent with the disclosure herein may exhibit a ratio of specific dye attachment (e.g., Cy3 attachment) to nonspecific dye adsorption (e.g., Cy3 dye adsorption) of at least 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 30:1, 40:1, 50:1, or greater than 50 attached specific dye molecules per nonspecifically adsorbed molecule. Similarly, when subjected to excitation energy, low background surfaces consistent with the disclosure herein having attached fluorophores, e.g., Cy3, may exhibit a ratio of specific fluorescent signal (e.g., resulting from Cy3-labeled oligonucleotides attached to the surface) to nonspecifically adsorbed dye fluorescent signal of at least 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 30:1, 40:1, 50:1, or greater than 50:1.
[0212] In some embodiments, the degree of hydrophilicity (or "wettability" with aqueous solutions) of the disclosed support surfaces can be assessed, for example, via water contact angle measurements, in which a small drop of water is placed on the surface and its contact angle with the surface is measured, for example, using an optical tensiometer. In some embodiments, a static contact angle can be determined. In some embodiments, an advancing or receding contact angle can be determined. In some embodiments, the water contact angle for a surface-treated hydrophilic low-binding support disclosed herein can range from about 0 degrees to about 30 degrees. In some embodiments, the water contact angle for a surface-treated hydrophilic low-binding support disclosed herein can be 50 degrees, 40 degrees, 30 degrees, 25 degrees, 20 degrees, 18 degrees, 16 degrees, 14 degrees, 12 degrees, 10 degrees, 8 degrees, 6 degrees, 4 degrees, 2 degrees, or 1 degree or less. In many cases, the contact angle will be 40 degrees or less. One of skill in the art will recognize that a given hydrophilic low-binding support surface of the present disclosure can exhibit a water contact angle having a value anywhere within this range.
[0213] In some embodiments, the hydrophilic surfaces disclosed herein often facilitate reduced wash times for bioassays due to reduced nonspecific binding of biomolecules to the low-binding surface. In some embodiments, a suitable wash step can be performed in less than 60, 50, 40, 30, 20, 15, 10 seconds, or less than 10 seconds. For example, a suitable wash step can be performed in less than 30 seconds.
[0214] The low-binding surfaces of the present disclosure may exhibit significant improvements in stability or durability to long-term exposure to solvents and elevated temperatures, or to repeated cycles of solvent exposure or temperature changes. For example, the stability of the disclosed surfaces can be tested by fluorescently labeling functional groups on the surface or tethered biomolecules (e.g., oligonucleotide primers) on the surface and monitoring the fluorescent signal before, during, and after long-term exposure to solvents and elevated temperatures, or to repeated cycles of solvent exposure or temperature changes. In some embodiments, the degree of change in fluorescence used to assess surface quality may be less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, or 25% (or any combination of these percentages when measured over these time periods) over a 1 minute, 2 minutes, 3 minutes, 4 minutes, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 40 minutes, 50 minutes, 60 minutes, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 15 hours, 20 hours, 25 hours, 30 hours, 35 hours, 40 hours, 45 hours, 50 hours, or 100 hours period of exposure to solvent and / or elevated temperature. In some embodiments, the degree of change in fluorescence used to assess surface quality may be less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, or 25% (or any combination of these percentages when measured over this range of cycles) over 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 cycles of repeated exposure to solvent changes and / or changes in temperature.
[0215] In some embodiments, the surfaces disclosed herein may exhibit a high ratio of specific signal to nonspecific signal or other background. For example, when used for nucleic acid amplification, some surfaces may exhibit an amplification signal that is at least 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 75, 100, or more than 100 times greater than the signal in adjacent non-populated areas of the surface. Similarly, some surfaces exhibit an amplification signal that is at least 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 75, 100, or more than 100 times greater than the signal in adjacent amplified nucleic acid-populated areas of the surface.
[0216] In some embodiments, fluorescent images of the disclosed low background surfaces when used to generate polonies of hybridized or clonally amplified nucleic acid molecules (e.g., labeled directly or indirectly with a fluorophore) in nucleic acid hybridization or amplification applications exhibit a contrast-to-noise ratio (CNR) of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 20, 210, 220, 230, 240, 250, or greater than 250.
[0217] One or more types of primers may be attached or tethered to the support surface. In some embodiments, the one or more types of adapters or primers may include a spacer sequence, an adapter sequence for hybridization to an adapter-ligated target library nucleic acid sequence, a forward amplification primer, a reverse amplification primer, a sequencing primer, and / or a molecular barcoding sequence, or any combination thereof. In some embodiments, one primer or adapter sequence may be tethered to at least one layer of the surface. In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 different primer or adapter sequences may be tethered to at least one layer of the surface.
[0218] In some embodiments, the tethered adapter and / or primer sequence can range from about 10 nucleotides to about 100 nucleotides in length. In some embodiments, the tethered adapter and / or primer sequence can be at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 nucleotides in length. In some embodiments, the tethered adapter and / or primer sequence can be up to 100, up to 90, up to 80, up to 70, up to 60, up to 50, up to 40, up to 30, up to 20, or up to 10 nucleotides in length. Any of the lower and upper limits described in this paragraph may be combined to form a range encompassed within the present disclosure; for example, in some embodiments, the length of the tethered adapter and / or primer sequence can range from about 20 nucleotides to about 80 nucleotides. One of skill in the art will recognize that the length of the tethered adapter and / or primer sequence can have any value within this range, e.g., about 24 nucleotides.
[0219] In some embodiments, the resulting surface density of primers (e.g., capture primers) on the low binding support surface of the present disclosure is less than 1 μm 2 Approximately 100 primer molecules per ~1 μm 2 In some embodiments, the resulting surface density of primers on the low binding support surface of the present disclosure can range from about 100,000 primer molecules per μm 2 Approximately 1,000 primer molecules per ~1 μm 2 In some embodiments, the surface density of primers can range from about 1,000,000 primer molecules per 1 μm 2 In some embodiments, the surface density of the primers can be at least 1,000, at least 10,000, at least 100,000, or at least 1,000,000 molecules per μm. 2Any of the lower and upper limits in this paragraph may be combined to form ranges within the disclosure; for example, in some embodiments, the surface density of the primers is greater than or equal to 1,000,000, up to 100,000, up to 10,000, or up to 1,000 molecules per μm. 2 Approximately 10,000 molecules per ~1 μm 2 The surface density of primer molecules can range from about 100,000 molecules per μm. 2 Those skilled in the art will recognize that the surface density of target library nucleic acid sequences initially hybridized to adapter or primer sequences on the support surface may be equal to or less than that indicated for the surface density of tethered primers. In some embodiments, the surface density of clonally amplified target library nucleic acid sequences hybridized to adapter or primer sequences on the support surface may encompass the same range as that indicated for the surface density of tethered primers.
[0220] The local densities listed above are for surfaces with a density of, for example, 500,000 / μm 2 It is not excluded that the density may vary across the surface, so that one region may include a region with a local oligo density of 1000, while also including at least a second region with a substantially different local density.
[0221] In some embodiments, the performance of nucleic acid hybridization and / or amplification reactions using the disclosed reaction formulations and low-binding supports can be evaluated using fluorescence imaging techniques, with the image contrast-to-noise ratio (CNR) providing an important metric for assessing amplification specificity and nonspecific binding on the support. CNR is generally defined as CNR = (signal - background) / noise. The background term is generally considered to be the signal measured over an interstitial region surrounding a specific feature (diffraction-limited spot, DLS) in a specific region of interest (ROI). While signal-to-noise ratio (SNR) is often considered a benchmark for overall signal quality, as shown in the following examples, improved CNR can be shown to offer significant advantages over SNR as a benchmark for signal quality in applications requiring rapid image capture (e.g., sequencing applications where cycle time must be minimized). With a high CNR, the imaging time required to reach accurate discrimination (and therefore accurate base calling in the case of sequencing applications) can be significantly reduced with even a moderate improvement in CNR. Improved CNR in imaging data over imaging integration time provides a method for more accurately detecting features such as clonally amplified nucleic acid colonies on a support surface.
[0222] In most ensemble-based sequencing approaches, the background term is typically measured as the signal associated with the "interstitial" regions. "Interstitial" background (B inter ), plus the "intrastitial" background (B intra ) resides within the region occupied by the amplified DNA colonies. The combination of these two background signals determines the achievable CNR, which then directly impacts optical equipment requirements, architecture costs, reagent costs, run times, cost per genome, and ultimately, accuracy and data quality for circular array-based sequencing applications. Binter Background signals arise from a variety of sources, some examples of which include autofluorescence from the exhaust flow cell, nonspecific adsorption of detection molecules resulting in spurious fluorescent signals that can obscure signals from the ROI, and the presence of nonspecific DNA amplification products (e.g., those resulting from primer dimers). In a typical next-generation sequencing (NGS) application, this background signal in the current field of view (FOV) is averaged and subtracted over time. Signals arising from individual DNA colonies (i.e., (signal)-B(interstitial) in the FOV) provide distinguishable features that can be classified. In some embodiments, intrastitial background (B(intrastitial)) may contribute to confounding fluorescent signals that are not specific to the target of interest but are present in the same ROI, thus making averaging and subtraction much more difficult.
[0223] Nucleic acid amplification on the low-binding coated supports described herein can reduce B (interstitial) background signal by reducing nonspecific binding, resulting in an improvement in specific nucleic acid amplification, which can affect background signal arising from both interstitial and intrastitial regions. In some embodiments, the disclosed low-binding coated supports, optionally used in combination with the disclosed hybridization and / or amplification reaction formulations, can provide an improvement in CNR of 2, 5, 10, 100, 250, 500, or 1000 times greater than that achieved using conventional supports and hybridization, amplification, and / or sequencing protocols. While described herein in the context of using fluorescent imaging as a readout or detection mode, the same principles apply to the use of the disclosed low-binding coated supports and nucleic acid hybridization and amplification formulations for other detection modes, including both optical and non-optical detection modes.
[0224] Incorporation by Reference Throughout this application, various publications, patents, and / or patent applications are referenced. The disclosures of these publications, patents, and / or patent applications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this disclosure pertains.
[0225] equivalent The details of one or more embodiments of the present disclosure are set forth in the accompanying description above. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, the preferred methods and materials are described herein. Other features, objects, and advantages of the present disclosure will become apparent from the description and claims. As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All patents and publications cited herein are incorporated by reference.
[0226] The foregoing description has been presented for illustrative purposes only and is intended to limit the disclosure not to the precise form disclosed, but rather by the scope of the claims appended hereto.
Claims
1. 1. A method for preparing spatially resolved nucleic acids, comprising: a) providing a support comprising: (i) a low non-specific binding coating comprising at least one hydrophilic polymer, said low non-specific binding coating having a water contact angle of 45 degrees or less; and (ii) a plurality of immobilized surface capture primers covalently tethered to said low non-specific binding coating, each individual surface capture primer comprising a universal surface capture primer sequence (110), a universal binding site for a reverse sequencing primer (120), and an RNA capture sequence; b) positioning a cell sample on the coated support under conditions suitable for the cell sample to remain in a fixed position on the coated support, disrupting the cell sample and releasing a plurality of RNA molecules of the cell sample from the cell sample under conditions suitable for preserving spatial location information of the RNA molecules, and collecting the plurality of RNA molecules onto a plurality of immobilized surface capture primers on the support under conditions suitable for hybridizing individual RNA molecules to individual immobilized surface capture primers to generate a plurality of capture primer-RNA duplexes, each primer-RNA duplex comprising an immobilized surface capture primer hybridized to an RNA molecule; c) performing a reverse transcription reaction on the coated support under conditions suitable for extending the 3' end of the immobilized surface capture primer using the hybridized RNA as a template strand, and under conditions suitable for generating a non-template poly-C tail at the 3' end, thereby generating a plurality of first-strand cDNA molecules having a non-template poly-C tail at their 3' end, wherein the reverse transcription reaction comprises a plurality of template switching oligonucleotides, each of the template switching oligonucleotides comprising a poly-G region capable of hybridizing to the non-template poly-C tail of the first-strand cDNA molecule, a universal binding site for a forward sequencing primer (130), and a universal binding site for a surface pinning primer (140); d) performing a primer extension reaction from the 3'-end of the non-template poly-C tail of the plurality of first-strand cDNA molecules using the template switching oligonucleotide as a template strand, thereby generating a plurality of immobilized full-length first-strand cDNAs, each full-length first-strand cDNA comprising a universal surface capture primer sequence (110), a universal binding site for a reverse sequencing primer (120), an RNA capture sequence, a cDNA insert region, a poly-C tail region, a universal binding site for a forward sequencing primer (130), and a universal binding site for a surface pinning primer (140); e) removing the RNA molecules while retaining the plurality of immobilized full-length first strand cDNAs; f) contacting the retained plurality of immobilized full-length first-strand cDNAs with a plurality of single-stranded circularization oligonucleotides under conditions suitable for hybridizing each circularization oligonucleotide to the immobilized full-length first-strand cDNAs, thereby forming gapped single-stranded circular molecules, wherein each single-stranded circularization oligonucleotide comprises: (i) a sequence complementary to the universal binding site (120) for the reverse sequencing primer; (ii) a sequence complementary to the universal surface capture primer sequence (110); (iii) a linker region; (iv) a sequence complementary to the universal binding site (140) for the surface pinning primer; and (v) a sequence complementary to the universal binding site (130) for the forward sequencing primer; g) performing a polymerase-catalyzed extension reaction to fill the gap using the poly-C tail region, the cDNA insert region, and the RNA capture sequence of the immobilized full-length first-strand cDNA as a template strand, wherein the polymerase-catalyzed extension reaction forms a single-stranded circularized molecule with a nick, which is closed by performing an enzymatic ligation reaction to generate a single-stranded covalently closed circular molecule that is hybridized to the immobilized full-length first-strand cDNA; h) performing a rolling circle amplification reaction using the 3' end of the immobilized full-length first strand cDNA as an initiation site and the covalently closed circular molecule as a template strand, thereby generating a plurality of immobilized nucleic acid concatemers that are spatially resolved on the support; i) sequencing the plurality of individual immobilized nucleic acid concatemers, wherein said sequencing comprises at least a portion of the cDNA insert region of each individual nucleic acid concatemer corresponding to the individual RNA molecules eluted from the cell sample; j) determining the location of individual nucleic acid concatemers on said coated support which correspond to the spatial location of said individual RNA molecules eluted from said cell sample.
2. 10. The method of claim 1, wherein the cell sample of step b) comprises a single cell, multiple cells, a tissue, an organ, an organism, or a sectioned cell sample.
3. 10. The method of claim 1, wherein the cell sample of step b) comprises a fresh sample, a frozen sample, a fresh frozen sample, or a formalin-fixed paraffin-embedded sample.
4. 2. The method of claim 1, wherein the plurality of template switching oligonucleotides in step c) comprises chimeric DNA and / or RNA oligonucleotides.
5. 2. The method of claim 1, wherein the linker region of the single-stranded circularization oligonucleotide in step f) comprises at least one sample index sequence for multiplexing, at least one unique molecular index (UMI) sequence for molecular tagging, and / or at least one universal binding site for a compaction oligonucleotide.
6. 2. The method of claim 1, wherein the rolling circle amplification reaction of step h) is carried out in the presence of a plurality of compaction oligonucleotides, each of the compaction oligonucleotides comprising a single-stranded oligonucleotide having a first region at a first distal end that hybridizes to one portion of a concatemeric molecule and a second region at a second distal end that hybridizes to a second portion of the same concatemeric molecule, and wherein the first portion of the concatemeric molecule and the second portion of the concatemeric molecule are brought into close proximity, whereby compaction of the concatemeric molecules forms DNA nanoballs.
7. said sequencing of step i) a. contacting a plurality of concatemer molecules with a plurality of sequencing polymerases and a plurality of nucleic acid sequencing primers, said contacting being carried out under conditions suitable to form a plurality of hybrid sequencing polymerases, each hybrid sequencing polymerase comprising a sequencing polymerase bound to a nucleic acid duplex, said nucleic acid duplex comprising a portion of the concatemer molecules hybridized to said nucleic acid sequencing primer; b. contacting the multiple hybrid sequencing polymerases with multiple detectably labeled nucleotides comprising a blocking moiety at the 2' or 3' sugar position, wherein the contacting is performed under conditions suitable for binding at least one nucleotide to at least one of the hybrid sequencing polymerases, the conditions being suitable for promoting polymerase-catalyzed nucleotide incorporation; c. incorporating nucleotides into the 3' end of the sequencing primer of at least one multiplexed sequencing polymerase; d. detecting the incorporated nucleotide and identifying the nucleobase of the incorporated nucleotide; e. removing the blocking moiety from the incorporated nucleotide; and f. repeating steps b. through e. at least one time.
8. said sequencing of step i) a. contacting a plurality of concatemer molecules with a plurality of sequencing polymerases and a plurality of nucleic acid sequencing primers, said contacting being carried out under conditions suitable to form a plurality of hybrid sequencing polymerases, each hybrid sequencing polymerase comprising a sequencing polymerase bound to a nucleic acid duplex, said nucleic acid duplex comprising a portion of the concatemer molecules hybridized to said nucleic acid sequencing primer; b. contacting the multiple hybrid sequencing polymerases with a multiple nucleotides each comprising a detectable label attached to a phosphate moiety of a phosphate strand, said contacting being carried out under conditions suitable for binding at least one nucleotide to at least one of the hybrid sequencing polymerases, said conditions being suitable for promoting polymerase-catalyzed nucleotide incorporation; c. incorporating nucleotides into the 3' end of the sequencing primer of at least one multiplexed sequencing polymerase; d. detecting the incorporated nucleotide and identifying the nucleobase of the incorporated nucleotide; and e. repeating steps b. through d. at least one time.
9. said sequencing of step i) contacting a plurality of concatemer molecules with a plurality of first sequencing polymerases and a plurality of nucleic acid sequencing primers, said contacting being carried out under conditions suitable for binding the plurality of first polymerases to the plurality of nucleic acid template molecules and the plurality of nucleic acid primers, thereby forming a plurality of first multiplexed polymerases, each first multiplexed polymerase comprising a first polymerase bound to a nucleic acid duplex, said nucleic acid duplex comprising a nucleic acid template molecule hybridized to a nucleic acid primer; b) contacting the plurality of first multiplexed polymerases with a plurality of detectably labeled multivalent molecules to form a plurality of multivalent binding complexes, wherein each detectably labeled multivalent molecule in the plurality of detectably labeled multivalent molecules comprises a core attached to a plurality of nucleotide arms, each nucleotide arm being attached to a nucleotide unit, said contacting being carried out under conditions suitable for binding complementary nucleotide units of the multivalent molecules to at least two of the plurality of first multiplexed polymerases, thereby forming a plurality of multivalent binding complexes, said conditions being suitable for inhibiting incorporation of the complementary nucleotide units into the primers of the plurality of multivalent binding complexes; c. detecting said plurality of multivalent binding complexes; and (d) identifying the nucleobases of the complementary nucleotide units in the plurality of multivalent binding complexes, thereby determining the sequence of the nucleic acid template molecule.
10. said sequencing of step i) a. dissociating the plurality of multivalent binding complexes by removing the plurality of first sequencing polymerases and their bound multivalent molecules, retaining a plurality of nucleic acid duplexes; b. contacting the retained plurality of nucleic acid duplexes of step e) with a plurality of second sequencing polymerases under conditions suitable for binding a plurality of second polymerases to the retained plurality of nucleic acid duplexes, thereby forming a plurality of second hybrid polymerases each comprising a second polymerase bound to a nucleic acid duplex; c) contacting the plurality of second multiplexed polymerases with a plurality of nucleotides, said contacting being carried out under conditions suitable for binding complementary nucleotides from the plurality of nucleotides to at least two of the second multiplexed polymerases, thereby forming a plurality of nucleotide-binding complexes, said conditions being suitable for promoting nucleotide incorporation of the bound complementary nucleotides into the primers of the nucleotide-binding complexes.
11. d. The method of claim 10, further comprising detecting the complementary nucleotide incorporated into the primer by a nucleotide-complexed polymerase.
12. d. detecting the complementary nucleotide incorporated into the primer by the nucleotide-complexed polymerase; e. identifying the nucleobase of the complementary nucleotide incorporated into the primer of the nucleotide-conjugated polymerase.
13. 10. The method of claim 9, wherein said contacting said plurality of first complexed polymerases with said plurality of multivalent molecules in step b. is performed in the presence of non-catalytic divalent cations that inhibit polymerase-catalyzed nucleotide incorporation, said non-catalytic divalent cations comprising strontium, barium, or calcium.
14. 10. The method of claim 9, wherein each multivalent molecule in the plurality of multivalent molecules comprises (a) a core and (b) a plurality of nucleotide arms, wherein the plurality of nucleotide arms comprises (i) a core attachment moiety, (ii) a spacer, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms via their core attachment moieties, the spacer is attached to the linker, and the linker is attached to the nucleotide unit.
15. 15. The method of claim 14, wherein the linker comprises an aliphatic chain having 2 to 6 subunits or an oligoethylene glycol chain having 2 to 6 subunits.
16. 15. The method of claim 14, wherein the multiple nucleotide arms attached to a given core have the same type of nucleotide unit, and the type of nucleotide unit comprises dATP, dGTP, dCTP, dTTP, or dUTP.
17. 10. The method of claim 9, wherein the plurality of multivalent molecules comprises one type of multivalent molecule, and each multivalent molecule in the plurality of multivalent molecules has the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
18. 10. The method of claim 9, wherein the plurality of multivalent molecules comprises a mixture of any combination of two or more types of multivalent molecules, each type having a nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
19. forming a plurality of binding complexes, a) binding a first nucleic acid sequencing primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemeric molecule, thereby forming a first binding complex, wherein a first nucleotide unit of the first multivalent molecule binds to the first polymerase; b) binding a second nucleic acid sequencing primer, a second sequencing polymerase, and the first multivalent molecule to a second portion of the same concatemeric molecule, thereby forming a second binding complex, wherein a second nucleotide unit of the first multivalent molecule binds to a second polymerase; 10. The method of claim 9, further comprising forming a plurality of binding complexes, wherein the first binding complex and the second binding complex comprise the same multivalent molecule, forming an avidity complex.
20. forming an avidity complex, a) contacting the plurality of first sequencing polymerases and the plurality of nucleic acid sequencing primers with different portions of a concatemeric nucleic acid template molecule to form at least a first and a second multiplexed polymerase on the same concatemeric molecule; b) contacting a plurality of detectably labeled multivalent molecules with said at least a first multiplexed polymerase and said second multiplexed polymerase on the same concatemeric molecule under conditions suitable for binding of a single multivalent molecule from said plurality of multivalent molecules to said first multiplexed polymerase and said second multiplexed polymerase, wherein at least a first nucleotide unit of said single multivalent molecule binds to said first multiplexed polymerase comprising a first primer hybridized to a first portion of said concatemeric molecule, thereby forming a first binding complex, and at least a second nucleotide unit of said single multivalent molecule binds to said second multiplexed polymerase comprising a second primer hybridized to a second portion of said concatemeric molecule, thereby forming a second binding complex; said contacting is performed under conditions suitable to inhibit polymerase-catalyzed incorporation of the linked first and second nucleotide units in said first binding complex and said second binding complex; contacting the first binding complex and the second binding complex bound to the same multivalent molecule to form an avidity complex; c) detecting the first binding complex and the second binding complex on the same concatemeric molecule; 10. The method of claim 9, further comprising forming an avidity complex comprising the steps of: (a) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the concatemeric molecule; and (b) identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the concatemeric molecule.
21. 11. The method of Claim 10, wherein said contacting of said plurality of second hybrid polymerases with said plurality of nucleotides in step c. is performed in the presence of a catalytic divalent cation that promotes polymerase-catalyzed nucleotide incorporation, said catalytic divalent cation comprising magnesium or manganese.
22. 11. The method of claim 10, wherein each nucleotide in the plurality of nucleotides of step c. comprises an aromatic base, a 5-carbon sugar, and 1 to 10 phosphate groups.
23. 11. The method of claim 10, wherein the plurality of nucleotides in step c. comprises one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP, or a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
24. 11. The method of claim 10, wherein at least one of the nucleotides in the plurality of nucleotides in step c. is labeled with a fluorophore.
25. 11. The method of claim 10, wherein the plurality of nucleotides in step c. lacks a fluorophore label.
26. 11. The method of claim 10, wherein at least one of the nucleotides in the plurality of nucleotides of step c. comprises a removable chain-terminating moiety attached to a 3' carbon position of a sugar group, wherein the removable chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an O-azidomethyl group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group, and wherein the removable chain-terminating moiety is cleavable with a chemical compound to generate an extendable 3' OH moiety on the sugar group.