PCR-free library preparation and use using double-stranded splint adapters
Double-stranded splint adapters facilitate efficient sequencing by forming covalently closed circular molecules, addressing the challenges of short read assembly in next-generation sequencing and improving haplotype recovery.
Patent Information
- Application Number
- JP2024576632
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-16
- Filing Date
- 2023-07-05
- Publication Date
- 2025-08-05
AI Technical Summary
Next-generation sequencing methods generate numerous short reads that require laborious and computationally demanding assembly, leading to challenges in recovering haplotype information and straining computer clusters, particularly for complex genomes.
A method using double-stranded splint adapters to hybridize with single-stranded nucleic acid library molecules, forming covalently closed circular molecules for downstream amplification and sequencing, reducing computational requirements and enabling efficient haplotype recovery.
The method simplifies sequencing workflows by circularizing library molecules, reducing computational demands and improving haplotype information recovery, thereby enhancing sequencing efficiency and accuracy.
Smart Images

Figure 2025525422000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 358,491, filed July 5, 2022, and U.S. Provisional Patent Application No. 63 / 508,833, filed June 16, 2023, the entire contents of each of which are incorporated herein by reference.
[0002] Electronic Sequence Listing Reference The contents of the Electronic Sequence Listing (ELEM-013_001WO_SeqListing_ST26.xml, size 50,242 bytes, and created on June 30, 2023) are incorporated herein by reference in their entirety.
[0003] The present disclosure provides compositions comprising nucleic acid double-stranded splint adaptors and methods for preparing nucleic acid libraries using the double-stranded splint adaptors, which can hybridize to a portion of a library molecule to form a library-splint complex having a nick, which can be ligated to form a covalently closed circular molecule that can be subjected to downstream amplification and sequencing workflows. [Background technology]
[0004] While the transition from traditional Sanger-style sequencing methods to next-generation sequencing methods has reduced the cost of sequencing, next-generation sequencing methods remain subject to significant limitations. In one respect, available sequencing platforms generate numerous but relatively short sequencing reads that may require computational reassembly into the complete sequence of interest. Available assembly methods may be slow, laborious, expensive, computationally demanding, and / or unsuitable for populations of similar individuals (e.g., viruses). This is particularly true for sequencing complex genomes. Assembly is challenging, in part, due to the ever-expanding sequencing datasets associated with assembling short reads. Such datasets can place a significant strain on computer clusters. For example, de novo assembly may require simultaneously storing sequencing reads (or k-mers derived from them) in random access memory (RAM). For large datasets, this requirement is non-trivial. Furthermore, even when assembly is possible, important haplotype information often cannot be recovered. Indeed, the inherent limitations of available technologies hinder improvements to overcome the shortcomings of current sequencing technologies. Thus, there is a need for improved sequencing methods and associated assembly techniques that reduce the time and / or computational requirements necessary to obtain accurate sequences. Summary of the Invention
[0005] The present disclosure provides a method for forming a plurality of library-sprint complexes (500), comprising providing a plurality of double-stranded splint adapters (200), each double-stranded splint adapter (200) in the plurality of double-stranded splint adapters (200) comprising a first splint strand (300) hybridized to a second splint strand (400), the double-stranded splint adapter comprising a double-stranded region and two adjacent single-stranded regions, the first splint strand comprising a first region (320), an internal region (310), and a second region (330), the internal region (310) of the first splint strand hybridizing to the second splint strand (400). and hybridizing the plurality of double-stranded splint adapters to a plurality of single-stranded nucleic acid library molecules (100), each of the library molecules comprising a sequence of interest (110) flanked on a first side by a universal adapter sequence (120) for a forward sequencing primer binding site and on a second side by a universal adapter sequence (130) for a reverse sequencing primer binding site, thereby circularizing the plurality of library molecules to form a plurality of library-splint complexes (500), each having two nicks.
[0006] In some embodiments, the method further comprises (c) contacting a plurality of library-splint complexes (500) with a ligase to generate a plurality of covalently closed circular library molecules (600).
[0007] In some embodiments, the hybridizing is performed under conditions suitable for hybridizing the first region (320) of the first splint strand to the universal adapter sequence (120) for the forward sequencing primer binding site of the library molecule. In some embodiments, the conditions are suitable for hybridizing the second region (330) of the first splint strand to the universal adapter sequence (130) for the reverse sequencing primer binding site of the library molecule.
[0008] In some embodiments, the internal region (310) of the first splint strand (300) comprises at least three subregions. In some embodiments, the at least three subregions comprise subregion (311), subregion (312), and subregion (313). In some embodiments, subregion (311) comprises a universal adapter sequence for the surface capture primer binding site, a universal adapter sequence for the surface pinning primer binding site, a sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). In some embodiments, subregion (312) comprises a universal adapter sequence for the surface capture primer binding site, a universal adapter sequence for the surface pinning primer binding site, a sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). In some embodiments, subregion (313) comprises a universal adapter sequence for the surface capture primer binding site, a universal adapter sequence for the surface pinning primer binding site, a sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI).
[0009] In some embodiments, subregion (311), (312), or (313) comprises a sample index sequence, the sample index sequence comprising: a sample index sequence lacking a short random sequence (NNN); a sample index sequence and a short random sequence (NNN); a sample index flanked on both sides by nucleotide bases that can be converted to an abasic base; at least one nucleotide base that can be converted to an abasic base; at least one deoxyinosine; an 18-carbon spacer; and / or an 18-carbon spacer and at least one deoxyinosine.
[0010] In some embodiments, the method further comprises distributing the plurality of covalently closed circular library molecules (600) onto a support having a plurality of surface capture primers immobilized thereon under conditions suitable for hybridizing each of the covalently closed circular library molecules (600) to each of the immobilized surface capture primers, thereby immobilizing the plurality of covalently closed circular library molecules (600) to the support. In some embodiments, the support further comprises a plurality of surface pinning primers immobilized thereon.
[0011] In some embodiments, the method further comprises contacting the plurality of immobilized covalently closed circular library molecules (600) with a plurality of strand-displacing polymerases and a plurality of nucleotides under conditions suitable for performing a rolling circle amplification reaction on a support using a plurality of surface capture primers as immobilized amplification primers and a plurality of covalently closed circular library molecules (600) as template molecules, thereby generating a plurality of nucleic acid concatemer molecules immobilized to the surface capture primers.
[0012] In some embodiments, the method further comprises iii) sequencing the plurality of nucleic acid concatemer molecules immobilized on the surface capture primers, wherein the sequencing comprises (i) sequencing the sample index and (ii) sequencing the sequence of interest (110).
[0013] In some embodiments, the method further comprises iv) sequencing the plurality of nucleic acid concatemer molecules immobilized on the surface capture primers, wherein the sequencing comprises (A) sequencing one or more short random sequences NNN; (B) sequencing one or more sample indices; and (C) sequencing the sequence of interest (110).
[0014] In some embodiments, the interior region (310) of the first splint strand (300) comprises one sample index.
[0015] In some embodiments, the interior region (310) of the first splint strand (300) comprises one sample index and a short random sequence (NNN).
[0016] In some embodiments, the library molecule (100) comprises one or more nucleotide sequences selected from Table 1.
[0017] In some embodiments, the first splint strand (300) comprises one or more nucleotide sequences selected from Table 2.
[0018] In some embodiments, the second splint strand (400) comprises one or more nucleotide sequences selected from Table 3.
[0019] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a schematic diagram illustrating an exemplary linear nucleic acid library molecule (100) comprising an insert region (110) (e.g., a sequence of interest) flanked on one side by a universal adapter sequence (120) for a forward sequencing primer binding site, and the insert region (110) is flanked on the other side by a universal adapter sequence (130) for a reverse sequencing primer binding site. [Figure 2] FIG. 1 is a schematic diagram of an exemplary double-stranded splint adapter (200) comprising a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). [Figure 3]FIG. 1 is a schematic diagram illustrating an exemplary library circularization workflow, which involves hybridizing a linear, single-stranded library molecule (100) with a double-stranded splint adaptor (200), thereby circularizing the library molecule and forming a library-splint complex (500) with two nicks. [Figure 4] Schematic diagram showing an exemplary ligation reaction involving performing an enzymatic ligation reaction on a nick in a library-sprint complex (500), thereby closing the nick and forming a covalently closed circular library molecule (600) that hybridizes to the first splint strand (300). [Figure 5] 6 is a schematic diagram showing an exemplary covalently closed circular library molecule (600) hybridized to an amplification primer, with the dotted line representing the nascent extension product. [Figure 6] 1 is a schematic diagram illustrating several embodiments of a nucleic acid sequence of a first splint strand (300, long splint strand), which includes an external first region (320), an external second region (330), and three internal subregions, including subregion (311), subregion (312), and subregion (313). The schematic diagram also illustrates an exemplary nucleic acid sequence of a second splint strand (400, short splint strand), which includes three subregions, including subregion (411), subregion (412), and subregion (413). [Figure 7] Schematic diagram showing an exemplary nucleic acid sequence of a first splint strand (300, long splint strand), which includes an external first region (320), an external second region (330), and two internal subregions, including subregion (311) and subregion (313). The schematic diagram also shows an exemplary nucleic acid sequence of a second splint strand (400, short splint strand), which includes three subregions, including subregion (411), subregion (412), and subregion (413). The loop in subregion (412) represents loop formation due to the absence of subregion (312) in the first splint strand (300). [Figure 8A]1A-1C are a series of schematic diagrams showing various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (120) for a forward sequencing primer binding site, and a double-stranded adapter bearing the full-length sequence of a universal adapter sequence (120) for a forward sequencing primer binding site. [Figure 8B] 1A-1C are a series of schematic diagrams illustrating various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (120) for a forward sequencing primer binding site, and a double-stranded adapter bearing a cleavage sequence of the universal adapter sequence (120) for a forward sequencing primer binding site. [Figure 8C] 1 is a series of schematic diagrams showing various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (120) for a forward sequencing primer binding site. A schematic diagram of a double-stranded adapter with a 5' overhanging end, where one adapter strand bears the full-length sequence of the universal adapter sequence (120) for a forward sequencing primer binding site, and the other adapter strand bears a truncated sequence of the universal adapter sequence (120) for a forward sequencing primer binding site. [Figure 8D] 1 is a series of schematic diagrams showing various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (120) for a forward sequencing primer binding site. A schematic diagram of a double-stranded adapter with a 3' overhanging end, where one adapter strand bears a cleavage sequence of the universal adapter sequence (120) for a forward sequencing primer binding site, and the other adapter strand bears the full-length sequence of the universal adapter sequence (120) for a forward sequencing primer binding site. [Figure 9A]1A-1C are a series of schematic diagrams showing various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (130) for a reverse sequencing primer binding site; and FIG. 1D is a schematic diagram of a double-stranded adapter bearing the full-length sequence of a universal adapter sequence (130) for a reverse sequencing primer binding site. [Figure 9B] 1A-1C are a series of schematic diagrams illustrating various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (130) for a reverse sequencing primer binding site, and a double-stranded adapter bearing a cleavage sequence of the universal adapter sequence (130) for a reverse sequencing primer binding site. [Figure 9C] 1 is a series of schematic diagrams showing various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (130) for a reverse sequencing primer binding site. A schematic diagram of a double-stranded adapter with a 5' overhanging end, where one adapter strand bears the full-length sequence of the universal adapter sequence (130) for a reverse sequencing primer binding site, and the other adapter strand bears a truncated sequence of the universal adapter sequence (130) for a reverse sequencing primer binding site. [Figure 9D] 1 is a series of schematic diagrams showing various embodiments of double-stranded nucleic acid adapters, each bearing a universal adapter sequence (130) for a reverse sequencing primer binding site. A schematic diagram of a double-stranded adapter with a 3' overhanging end, where one adapter strand bears a cleavage sequence of the universal adapter sequence (130) for a reverse sequencing primer binding site, and the other adapter strand bears the full-length sequence of the universal adapter sequence (130) for a reverse sequencing primer binding site. [Figure 10A]1 is a series of schematic diagrams illustrating various embodiments of Y-shaped adapters, each comprising two oligonucleotides hybridized together and having a double-stranded annealing region and a mismatch portion. A schematic diagram of a Y-shaped adapter comprising a first oligonucleotide carrying the full-length sequence of a universal adapter sequence (120) for a forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of a universal adapter sequence (130) for a reverse sequencing primer binding site. [Figure 10B] 1 is a series of schematic diagrams illustrating various embodiments of Y-shaped adapters, each comprising two oligonucleotides hybridized together and having a double-stranded annealing region and a mismatch portion. A schematic diagram of a Y-shaped adapter comprising a first oligonucleotide carrying the cleavage sequence of the universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the cleavage sequence of the universal adapter sequence (130) for the reverse sequencing primer binding site. [Figure 10C] 1 is a series of schematic diagrams illustrating various embodiments of Y-shaped adapters, each comprising two oligonucleotides hybridized together and having a double-stranded annealing region and a mismatch portion. A schematic diagram of a Y-shaped adapter comprising a first oligonucleotide carrying the full-length sequence of the universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the cleavage sequence of the universal adapter sequence (130) for the reverse sequencing primer binding site. [Figure 10D]1 is a series of schematic diagrams illustrating various embodiments of Y-shaped adapters, each comprising two oligonucleotides hybridized together and having a double-stranded annealing region and a mismatch portion. A schematic diagram of a Y-shaped adapter comprising a first oligonucleotide carrying the cleavage sequence of the universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of the universal adapter sequence (130) for the reverse sequencing primer binding site. [Figure 11A]
[0023] Figure 1 is a series of schematic diagrams illustrating embodiments of transpososomes. These diagrams show two transpososomes. The first transpososome (top) contains a transposase bound to a double-stranded polynucleotide containing a transposon end sequence and a universal adapter sequence (120) for a forward sequencing primer binding site. The transposon end sequence specifically binds to the transposase. The second transpososome (bottom) contains a transposase bound to a double-stranded polynucleotide containing a transposon end sequence and a universal adapter sequence (130) for a reverse sequencing primer binding site. The transposon end sequence specifically binds to the transposase. [Figure 11B]
[0023] Figure 1 is a series of schematic diagrams illustrating embodiments of a transpososome. A schematic diagram of an exemplary transpososome includes a transposase bound to a first double-stranded polynucleotide comprising a transposon end sequence and a universal adapter sequence for a forward sequencing primer binding site (120), and the transposase is bound to a second double-stranded polynucleotide comprising a transposon end sequence and a universal adapter sequence for a reverse sequencing primer binding site (130). The transposon end sequence specifically binds to the transposase. [Figure 12]Schematic diagram showing one embodiment of a transpososome comprising a transposase bound to Y-shaped adapters, each of which comprises two oligonucleotides hybridized together and has a double-stranded annealing region and a mismatch portion. [Figure 13] 1 is a schematic diagram showing an exemplary adapter ligation workflow for generating double-stranded linear nucleic acid library molecules. Double-stranded nucleic acid fragments are enzymatically ligated on one side to double-stranded adapters carrying a universal adapter sequence (120) for a forward sequencing primer binding site. Double-stranded nucleic acid fragments are enzymatically ligated on the other side to double-stranded adapters carrying a universal adapter sequence (130) for a reverse sequencing primer binding site. [Figure 14] 1 is a schematic diagram illustrating an exemplary adapter ligation workflow for generating double-stranded linear nucleic acid library molecules. Double-stranded nucleic acid fragments are enzymatically ligated to a first Y-shaped adapter comprising, on one side, a first oligonucleotide carrying the full-length sequence of a universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of a universal adapter sequence (130) for the reverse sequencing primer binding site. The double-stranded nucleic acid fragments are enzymatically ligated to a second Y-shaped adapter comprising, on the other side, a first oligonucleotide carrying the full-length sequence of a universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of a universal adapter sequence (130) for the reverse sequencing primer binding site. [Figure 15]1 is a schematic diagram illustrating an exemplary adapter ligation workflow for generating double-stranded linear nucleic acid library molecules. Double-stranded nucleic acid fragments are enzymatically ligated to a first Y-shaped adapter comprising, on one side, a first oligonucleotide carrying the full-length sequence of a universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of a universal adapter sequence (130) for the reverse sequencing primer binding site. The double-stranded nucleic acid fragments are enzymatically ligated to a second Y-shaped adapter comprising, on the other side, a first oligonucleotide carrying the full-length sequence of a universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of a universal adapter sequence (130) for the reverse sequencing primer binding site. The resulting double-stranded linear nucleic acid library molecules are subjected to primer extension using a primer having a region that hybridizes to one of the mismatched portions (e.g., (130)), which also carries a sample index sequence and a universal surface primer binding site (SurfPBS) at its other end. A primer extension reaction is performed using the hybridized primer as a template to generate extended library molecules containing the forward sequencing primer binding site (120), the insert sequence (110), the reverse sequencing primer binding site (130), the sample index sequence, and the universal surface primer binding site (SurfPBS). [Figure 16]Schematic diagram showing an exemplary adapter ligation workflow for generating double-stranded linear nucleic acid library molecules. Double-stranded nucleic acid fragments are enzymatically ligated to a first Y-shaped adapter, which includes, on one side, a first oligonucleotide carrying the full-length sequence of a universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of a universal adapter sequence (130) for the reverse sequencing primer binding site. The double-stranded nucleic acid fragments are enzymatically ligated to a second Y-shaped adapter, which includes, on the other side, a first oligonucleotide carrying the full-length sequence of a universal adapter sequence (120) for the forward sequencing primer binding site and a second oligonucleotide carrying the full-length sequence of a universal adapter sequence (130) for the reverse sequencing primer binding site. The resulting double-stranded linear nucleic acid library molecules are hybridized with linear double-stranded adapters having 3' overhanging ends and blunt ends. The 3' overhanging end contains a sequence that can hybridize to a mismatched portion with a reverse sequencing primer binding site (130), generating a partially double-stranded region with a nick (black triangle). The nick can be ligated, and the unligated strand can be removed. [Figure 17] Schematic of an exemplary tagging workflow using multiple transpososomes: Input double-stranded DNA is contacted with multiple transpososomes, each containing a transposase bound to a first double-stranded polynucleotide comprising a transposon end sequence and a universal adapter sequence for a forward sequencing primer binding site (120), and the transposase is bound to a second double-stranded polynucleotide comprising a transposon end sequence and a universal adapter sequence for a reverse sequencing primer binding site (130). [Figure 18]1 is a schematic diagram of an exemplary library circularization workflow, which includes hybridizing a linear, single-stranded library molecule (100) with a first splint strand (200), thereby circularizing the library molecule and forming a library single-splint complex (700) having a gap between the ends of the library molecule. A linear nucleic acid library molecule (100) includes an insert region (110) (e.g., a sequence of interest) flanked on one side by a universal adapter sequence (120) for a forward sequencing primer binding site, and the insert region (110) is flanked on the other side by a universal adapter sequence (130) for a reverse sequencing primer binding site. The first splint strand (200) contains a first region (320) that hybridizes with a universal adapter sequence (120) for a forward sequencing primer binding site on one end of the linear single-stranded library molecule (100), and the first splint strand contains a second region (330) that hybridizes with a universal adapter sequence (130) for a reverse sequencing primer binding site on the other end of the linear single-stranded library molecule. The gap is closed by a polymerase-catalyzed fill-in reaction and an enzymatic ligation reaction to generate a covalently closed circular molecule. [Figure 19] 1 is a graph showing the nucleotide base diversity of sample index sequences containing short 3-mer random sequences (NNN). The graph shows the nucleotide diversity of the 3-mer random sequences (NNN) of approximately 30% for A and T base calls and approximately 20% for C and G base calls. [Figure 20] 1 is a graph showing the nucleotide base diversity of sample index sequences lacking short 3-mer random sequences (NNN). The graph shows a nucleotide diversity of approximately 40% for A and T base calls, approximately 15% for C base calls, and approximately 5% for G base calls. [Figure 21]FIG. 1 is a schematic diagram of an exemplary low-binding support comprising a glass substrate and alternating layers of a hydrophilic coating covalently or non-covalently adhered to the glass, and further comprising chemically reactive functional groups that serve as binding sites for oligonucleotide primers (e.g., capture oligonucleotides and circularization oligonucleotides). [Figure 22] Schematic diagrams of various exemplary configurations of multivalent molecules. Left: Schematic diagram of a multivalent molecule with a starburst or helter-skelter configuration. Center: Schematic diagram of a multivalent molecule with a dendrimer configuration. Right: Schematic diagram of multiple multivalent molecules formed by reacting streptavidin with 4-arm or 8-arm PEG-NHS bearing biotin and dNTPs. Nucleotide units are represented as "N," biotin is represented as "B," and streptavidin is represented as "SA." [Figure 23] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a generic core attached to multiple nucleotide arms. [Figure 24] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms. [Figure 25] 1 shows a schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, the nucleotide arms comprising biotin, a spacer, a linker, and a nucleotide unit. [Figure 26] FIG. 1 is a schematic diagram of an exemplary nucleotide arm comprising a core-binding moiety, a spacer, a linker, and a nucleotide unit. [Figure 27] Shown are the chemical structures of an exemplary spacer (top) and various exemplary linkers (bottom), including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker. [Figure 28] 1 shows the chemical structures of various exemplary linkers, including linkers 1-9. [Figure 29] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 30] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 31] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 32] The chemical structure of an exemplary biotinylated nucleotide arm is shown, in which the nucleotide unit is connected to the linker via a propargylamine bond at the 5-position of the pyrimidine base or the 7-position of the purine base. DETAILED DESCRIPTION OF THE INVENTION
[0021] Definition: The headings provided herein are not limitations of various aspects of the disclosure, which aspects can be understood by reference to the specification as a whole.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art. Generally, terms relating to molecular biology, nucleic acid chemistry, protein chemistry, genetics, microbiology, transgenic cell production, and hybridization techniques described herein are well known and commonly used in the art. The techniques and procedures described herein are generally carried out according to conventional methods well known in the art and as described in various general and more specific references cited and discussed throughout the specification. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY 2000). See also Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992). The nomenclature used in connection with, and the experimental procedures and techniques described herein are well known and commonly used in the art.
[0023] Unless otherwise required by context herein, singular terms include plurals and plural terms include the singular. The singular forms "a," "an," and "the," and use of any word in the singular, include plural referents unless clearly and unambiguously limited to one referent.
[0024] The use of alternative terms (eg, "or") is understood to mean either one or both of the alternatives, or any combination thereof.
[0025] As used herein, the term "and / or" should be understood to mean a specific disclosure of each of the specified features or components with or without the other. For example, the term "and / or" used in phrases such as "A and / or B" is intended to include "A and B," "A or B," "A" (A alone), and "B" (B alone). In a similar manner, the term "and / or" used in phrases such as "A, B, and / or C" is intended to encompass each of the following aspects: "A, B, and C," "A, B, or C," "A or C," "A or B," "B or C," "A and B," "B and C," "A and C," "A" (alone), "B" (alone), and "C" (alone).
[0026] As used in this specification and the appended claims, the terms "comprising," "including," "having," and "containing," and grammatical variations thereof, as used herein, are intended to be open-ended so that one or more items in a list do not exclude other items that may be substituted for or added to the listed items. Wherever embodiments are described herein with the term "comprising," it is understood that alternatively similar embodiments described with the terms "consisting of" and / or "consisting essentially of" are also provided.
[0027] As used herein, the term "about" or "approximately" refers to a value or composition that is within an acceptable error range for a particular value or composition, as determined by one of ordinary skill in the art. The acceptable error range depends in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" or "approximately" can mean within one or more standard deviations per practice in the art. Alternatively, "about" or "approximately" can mean a range of up to 10% (i.e., ±10%) or more, depending on the limitations of the measurement system. For example, about 5 mg can include any number between 4.5 mg and 5.5 mg. Furthermore, particularly with respect to biological systems or processes, the term can mean up to one order of magnitude or up to five times the value. When a particular value or composition is provided in this disclosure, unless otherwise specified, the meaning of "about" or "approximately" should be considered to be within an acceptable error range for that particular value or composition. Additionally, when ranges and / or subranges of values are provided, the ranges and / or subranges can include the endpoints of the ranges and / or subranges.
[0028] The terms "peptide," "polypeptide," and "protein," as well as other related terms used herein, are used interchangeably and refer to a polymer of amino acids and are not limited to any particular length. Polypeptides can contain natural and unnatural amino acids. Polypeptides include recombinant or chemically synthesized forms. Polypeptides also include precursor molecules that have not yet undergone post-translational modifications, such as proteolytic cleavage, ribosomal skipping cleavage, hydroxylation, methylation, lipidation, acetylation, sumoylation, ubiquitination, glycosylation, phosphorylation, and / or disulfide bond formation. These terms encompass natural and artificial proteins, protein fragments, and polypeptide analogs of protein sequences (such as muteins, variants, chimeric proteins, and fusion proteins), as well as proteins that are post-translationally or otherwise covalently or non-covalently modified.
[0029] The term "cellular biological sample" refers to a single cell, multiple cells, tissue, organ, organism, or section of any of these cellular biological samples. A cellular biological sample may be extracted from an organism (e.g., biopsy) or obtained from a cell culture growing in liquid or on a culture dish. A cellular biological sample includes a fresh sample, a frozen sample, a fresh frozen sample, or an archived sample (e.g., formalin-fixed paraffin-embedded; FFPE) sample. A cellular biological sample may be embedded in wax, resin, epoxy, or agar. A cellular biological sample may be fixed, for example, in any one or any combination of two or more of acetone, ethanol, methanol, formaldehyde, paraformaldehyde-Triton, or glutaraldehyde. A cellular biological sample may or may not be sectioned. A cellular biological sample may be stained, destained, or unstained.
[0030] Nucleic acids of interest can be extracted from cells or cellular biological samples using any of several techniques known to those skilled in the art. For example, a typical DNA extraction procedure includes: (i) collecting a cell or tissue sample from which DNA is to be extracted; (ii) disrupting cell membranes (i.e., lysing cells) to release DNA and other cytoplasmic components; (iii) treating the lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA; and (iv) purifying DNA from the supernatant to remove detergents, proteins, salts, or other reagents used during cell membrane lysis. A variety of suitable commercially available nucleic acid extraction and purification kits are consistent with the disclosure herein. Examples include, but are not limited to, the QIAamp kit (for isolating genomic DNA from human samples) and the DNAeasy kit (for isolating genomic DNA from animal or plant samples) from Qiagen (Germantown, MD), or the Maxwell® and ReliaPrep™ series of kits from Promega (Madison, WI).
[0031] As used herein, the term "polymerase" and variations thereof include enzymes that contain a nucleotide (or nucleoside)-binding domain, and the polymerase can form a complex with a template nucleic acid and a complementary nucleotide. A polymerase can have one or more activities, including, but not limited to, base analog detection activity, DNA polymerization activity, reverse transcriptase activity, DNA binding, strand displacement activity, and nucleotide binding and recognition. A polymerase can be any enzyme that can catalyze the polymerization of nucleotides (including their analogs) into a nucleic acid strand. Typically, although not necessarily, such nucleotide polymerization can occur in a template-dependent manner. Typically, a polymerase contains one or more active sites at which nucleotide binding and / or catalysis of nucleotide polymerization can occur. In some embodiments, a polymerase includes other enzymatic activities, such as 3' to 5' exonuclease activity or 5' to 3' exonuclease activity. In some embodiments, a polymerase has strand displacement activity. Polymerases can include, but are not limited to, naturally occurring polymerases and any subunits and truncations thereof, mutant polymerases, variant polymerases, recombinant, fused, or otherwise engineered polymerases, chemically modified polymerases, synthetic molecules or assemblies, and any analogs, derivatives, or fragments thereof (e.g., catalytically active fragments) that retain the ability to catalyze nucleotide polymerization. Polymerases include catalytically inactive polymerases, catalytically active polymerases, reverse transcriptases, and other enzymes that contain a nucleotide-binding domain. In some embodiments, polymerases may be isolated from cells or produced using recombinant DNA technology or chemical synthesis methods. In some embodiments, polymerases may be expressed in prokaryotic, eukaryotic, viral, or phage organisms. In some embodiments, polymerases may be post-translationally modified proteins or fragments thereof. Polymerases may be derived from prokaryotic, eukaryotic, viral, or phage organisms. Polymerases include DNA-directed DNA polymerases and RNA-directed DNA polymerases.
[0032] The term "strand displacement" refers to the ability of a polymerase to locally separate strands of double-stranded nucleic acid and synthesize a new strand in a template-based manner. Strand-displacing polymerases displace a complementary strand from the template strand and catalyze new strand synthesis. Strand-displacing polymerases include mesophilic and thermophilic polymerases. Strand-displacing polymerases include wild-type enzymes and variants, including exonuclease-minus mutants, mutated versions, chimeric enzymes, and truncated enzymes. Examples of strand-displacing polymerases include, but are not limited to, phi29 DNA polymerase, large fragment of Bst DNA polymerase, large fragment of Bsu DNA polymerase (exo-), Bca DNA polymerase (exo-), Klenow fragment of E. coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV viral reverse transcriptase, Deep Vent DNA polymerase, and KOD DNA polymerase. The phi29 DNA polymerase can be a wild-type phi29 DNA polymerase (e.g., MagniPhi™ from Expedeon™), or a variant EquiPhi29™ DNA polymerase (e.g., from Thermo Fisher Scientific™), or a chimeric QualiPhi™ DNA polymerase (e.g., from 4basebio™).
[0033] As used herein, the terms "nucleic acid," "polynucleotide," and "oligonucleotide," as well as other related terms, are used interchangeably and refer to a polymer of nucleotides and are not limited to any particular length. Nucleic acids include recombinant or chemically synthesized forms. Nucleic acids can be isolated. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of DNA or RNA produced using nucleotide analogs (e.g., peptide nucleic acids and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids comprise polymers of nucleotides, and the nucleotides can contain natural or non-natural bases and / or sugars. Nucleic acids contain naturally occurring internucleoside linkages, such as, but not limited to, phosphodiester linkages. Nucleic acids may also contain non-natural internucleoside linkages, including phosphorothioate, phosphorothiolate, and / or peptide nucleic acid (PNA) linkages. In some embodiments, nucleic acids comprise one type of polynucleotide or a mixture of two or more different types of polynucleotides.
[0034] As used herein, the terms "operably linked" and "operably associated," or related terms, refer to the juxtaposition of components. The juxtaposed components can be covalently linked to one another. For example, two nucleic acid components can be enzymatically ligated, where the bond linking the two components comprises a phosphodiester bond. A first and second nucleic acid component can be linked to one another, where the first nucleic acid component can confer a function to the second nucleic acid component. For example, a bond between a primer binding sequence and a sequence of interest forms a nucleic acid library molecule having a portion capable of binding to a primer. In another example, a transgene (e.g., a nucleic acid encoding a polypeptide or nucleic acid sequence of interest) can be ligated to a vector, where the bond allows for expression or function of the transgene sequence contained within the vector. In some embodiments, the transgene is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects expression of the transgene. In some embodiments, the vector comprises at least one host cell regulatory sequence, including a promoter sequence, an enhancer, a transcription and / or translation initiation sequence, a transcription and / or translation termination sequence, and a polypeptide secretion signal sequence. In some embodiments, host cell regulatory sequences control the level, timing, and / or location of expression of the transgene.
[0035] The terms "linked," "joined," "attached," "appended," and variations thereof, include any type of fusion, bond, adhesion, or association between any combination of compounds or molecules that is stable enough to withstand use in a particular procedure. Procedures can include, but are not limited to, nucleotide binding, nucleotide incorporation, deblocking (e.g., removal of chain-terminating moieties), washing, removal, flow, detection, imaging, and / or identification. Such binding can include, for example, covalent, ionic, hydrogen, dipole-dipole, hydrophilic, hydrophobic, or affinity binding, van der Waals binding or association, and mechanical binding. In some embodiments, such binding occurs intramolecularly, for example, by joining the ends of single- or double-stranded linear nucleic acid molecules together to form a circular molecule. In some embodiments, such binding can occur between different molecular combinations or between molecules and non-molecules, including, but not limited to, binding between nucleic acid molecules and solid surfaces, binding between proteins and detectable reporter moieties, and binding between nucleotides and detectable reporter moieties. Some examples of binding can be found, for example, in Hermanson, G., "Bioconjugate Techniques", Second Edition (2008); Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998); Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998).
[0036] As used herein, the term "primer" and related terms refer to an oligonucleotide capable of hybridizing to a DNA and / or RNA polynucleotide template to form a duplex molecule. Primers may be single-stranded along their entire length or may have single-stranded and double-stranded portions. Primers contain natural nucleotides and / or nucleotide analogs. Primers can be recombinant nucleic acid molecules. Primers can be of any length but typically range from 4 to 50 nucleotides. Typical primers contain a 5' end and a 3' end. The 3' end of a primer can contain a 3' OH moiety that functions as a nucleotide polymerization initiation site in a polymerase-catalyzed primer extension reaction. Alternatively, the 3' end of a primer can lack a 3' OH moiety or contain a terminal 3' blocking group that inhibits nucleotide polymerization in a polymerase-catalyzed reaction. Any one or more nucleotides along the length of a primer can be labeled with a detectable reporter moiety. Primers can be in solution (e.g., soluble primers) or immobilized on a support (e.g., capture primers).
[0037] The terms "template nucleic acid," "template polynucleotide," "target nucleic acid," "target polynucleotide," "template strand," and other variations thereof, refer to a nucleic acid strand that serves as the base nucleic acid molecule for any of the amplification and / or sequencing methods described herein. A template nucleic acid may be single-stranded or double-stranded, or may have single-stranded or double-stranded portions. A template nucleic acid may be obtained from naturally occurring sources, recombinant forms, or chemically synthesized to include any type of nucleic acid analog. A template nucleic acid may be linear, concatemeric, circular, or in other forms.
[0038] When used in reference to nucleic acid molecules, the terms "hybridize" or "hybridizing" or "hybridization," or other related terms, refer to hydrogen bonding between two different nucleic acids to form a double-stranded nucleic acid. Hybridization also includes hydrogen bonding between two different regions of a single nucleic acid molecule to form a self-hybridizing molecule having a double-stranded region. Hybridization can involve Watson-Crick or Hoogsteen binding to form a double-stranded double-stranded nucleic acid or a double-stranded region within a nucleic acid molecule. A double-stranded nucleic acid, or two different regions of a single nucleic acid, can be fully complementary or partially complementary. Complementary nucleic acid strands need not hybridize to each other throughout their entire length. Complementary base pairing can be standard AT or CG base pairing, or other forms of base pairing interactions. A double-stranded nucleic acid can contain mismatched base-pairing nucleotides.
[0039] When used in reference to nucleic acids, the terms "extend," "extending," "extension," and other variations refer to the incorporation of one or more nucleotides into a nucleic acid molecule. Nucleotide incorporation involves the polymerization of one or more nucleotides to the terminal 3'OH terminus of a nucleic acid chain, resulting in the elongation of the nucleic acid chain. Nucleotide incorporation can be performed with natural nucleotides and / or nucleotide analogs. Typically, although not necessarily, nucleotide incorporation occurs in a template-dependent manner. Any suitable method for extending a nucleic acid molecule may be used, including primer extension catalyzed by DNA polymerase or RNA polymerase.
[0040] The term "nucleotide" and related terms refer to a molecule comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and at least one phosphate group. Standard or non-standard nucleotides are consistent with the use of this term. In some embodiments, a nucleotide comprises a monophosphate, diphosphate, or triphosphate, or the corresponding phosphate analog. The term "nucleoside" refers to a molecule comprising an aromatic base and a sugar. Nucleotides and nucleosides can be unlabeled or labeled with a detectable reporter moiety.
[0041] Nucleotides (and nucleosides) typically contain heterocyclic bases containing a substituted or unsubstituted nitrogen-containing parent heteroaromatic ring, which are commonly found in nucleic acids, including naturally occurring, substituted, modified, or engineered variants, or analogs thereof. The base of a nucleotide (or nucleoside) is capable of forming Watson-Crick and / or Hoogsteen hydrogen bonds with an appropriate complementary base. Exemplary bases are purines and pyrimidines, such as 2-aminopurine, 2,6-diaminopurine, adenine (A), ethenoadenine, N, N-acetylglucosamine ... 6 -Δ 2 -Isopentenyladenine (6iA), N 6 -Δ 2 -Isopentenyl-2-methylthioadenine (2ms6iA), N 6 -Methyladenine, guanine (G), isoguanine, N 2 -dimethylguanine (dmG), 7-methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine, and O 6 -methylguanine; 7-deaza-purines, such as 7-deazaadenine (7-deaza-A) and 7-deazaguanine (7-deaza-G); pyrimidines, such as cytosine (C), 5-propynylcytosine, isocytosine, thymine (T), 4-thiothymine (4sT), 5,6-dihydrothymine, O 4Examples of bases include, but are not limited to, methylthymine, uracil (U), 4-thiouracil (4sU), and 5,6-dihydrouracil (dihydrouracil; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; nebularine; inosine; hydroxymethylcytosine; 5-methycytosine; base (Y); and methylated, glycosylated, and acylated base moieties. Additional exemplary bases can be found in Fasman, 1989, "Practical Handbook of Biochemistry and Molecular Biology," pp. 385-394, CRC Press, Boca Raton, Fla.
[0042] Nucleotides (and nucleosides) typically include a sugar moiety, e.g., a carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100:4319-48), an acyclic moiety (Martinez, et al., 1999 Nucleic Acids Research 27:1271-1274; Martinez, et al., 1997 Bioorganic & Medicinal Chemistry Letters vol. 7:3013-3016), and another sugar moiety (Joeng, et al., 1993 J. Med. Chem. 36:2627-2638; Kim, et al., 1993 J. Med. Chem. 36:30-7; Eschenmosser 1999 Science 284:2118-2124; and U.S. Pat. No. 5,558,991). Sugar moieties include ribosyl; 2'-deoxyribosyl; 3'-deoxyribosyl; 2',3'-dideoxyribosyl; 2',3'-didehydrodideoxyribosyl; 2'-alkoxyribosyl; 2'-azidoribosyl; 2'-aminoribosyl; 2'-fluororibosyl; 2'-mercaptoriboxyl; 2'-alkylthioribosyl; 3'-alkoxyribosyl; 3'-azidoribosyl; 3'-aminoribosyl; 3'-fluororibosyl; 3'-mercaptoriboxyl; 3'-alkylthioribosyl carbocyclic; acyclic, or other modified sugars.
[0043] In some embodiments, the nucleotide comprises a chain of one, two, or three phosphorus atoms, typically linked to the 5' carbon of the sugar moiety via an ester or phosphoramido linkage. In some embodiments, the nucleotide is an analog having a phosphorus chain, in which the phosphorus atoms are linked to each other via O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise a substituted side chain group, including O, S, or BH3. In some embodiments, the chain comprises a phosphate group substituted with an analog, including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoramidite groups.
[0044] As used herein, "nucleotide unit" or "nucleotide moiety" refers to a nucleotide (e.g., dATP, dTTP, dGTP, dCTP, or dUTP), or an analog thereof, that comprises a base, a sugar, and at least one phosphate group. The nucleotide unit can be attached to a multivalent molecule used in the sequencing reactions described herein. Generally, all nucleotide units attached to the same multivalent molecule will have the same identity (e.g., all A, all T, all C, or all G), although one skilled in the art will understand that there may be situations in which multivalent molecules comprising nucleotide units of different identities are advantageous.
[0045] The term "rolling circle amplification" generally refers to an amplification method using a circularized nucleic acid template molecule containing a target sequence of interest, an amplification primer binding sequence, and, optionally, one or more adapter sequences, such as a sequencing primer binding sequence and / or a sample index sequence. A rolling circle amplification reaction may be performed under isothermal amplification conditions and includes a circularized nucleic acid template molecule, an amplification primer, a strand-displacing polymerase, and a plurality of nucleotides to generate concatemers containing tandem repeat sequences of any adapter sequences present in the circularized template molecule and the original circularized nucleic acid template molecule. The concatemers can self-collapse to form nucleic acid nanoballs. The shape and size of the nanoballs can be further compacted by including a pair of inverted repeat sequences within the circular template molecule or by performing the rolling circle amplification reaction with one or more compaction oligonucleotides. One advantage of using rolling circle amplification to generate clonal amplicons for sequencing workflows is that the repeat copies of the target sequence in the nanoballs can be simultaneously sequenced, increasing signal intensity. In some embodiments, a rolling circle amplification reaction can be performed in the presence of multiple compaction oligonucleotides having at least four consecutive guanines (e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, or more than 12 guanines). The rolling circle amplification reaction can generate concatemers containing repeated copies of the universal binding sequence for the compaction oligonucleotides. At least one compaction oligonucleotide can form a G-quadruplex and hybridize to the universal binding sequence for the compaction oligonucleotide, and the resulting concatemers can fold and form an intramolecular G-quadruplex structure. The concatemers can self-collapse to form compact nanoballs.Without wishing to be bound by theory, it is hypothesized that the formation of G-quadruplexes and G-quadruplexes in nanoballs can increase the stability of the nanoballs and allow them to retain their compact size and shape, which can withstand repeated flows of reagents to perform any of the sequencing workflows described herein.
[0046] The terms "amplify," "amplifying," "amplification," and other related terms, when used with respect to nucleic acids, include producing multiple copies of an original polynucleotide template molecule, where the copies contain sequences that are complementary to the template sequence and / or where the copies contain sequences that are identical to the template sequence. In some embodiments, the copies contain sequences that are substantially identical to the template sequence and / or sequences that are substantially identical to the sequence that is complementary to the template sequence.
[0047] The terms "reporter moiety," "reporter moieties," or related terms refer to a compound that produces or can be caused to produce a detectable signal. Reporter moieties are often referred to as "labels." Any suitable reporter moiety can be used, including luminescence, photoluminescence, electroluminescence, bioluminescence, chemiluminescence, fluorescence, phosphorescence, chromophores, radioisotopes, electrochemistry, mass spectrometry, Raman, haptens, affinity tags, atoms, or enzymes. A reporter moiety produces a detectable signal that results from a chemical or physical change (e.g., heat, light, electricity, pH, salt concentration, enzymatic activity, or a proximity event). A proximity event involves two reporter moieties coming into close proximity to, associating with, or binding to each other. It is well known to those skilled in the art to select reporter moieties so that each absorbs excitation radiation and / or emits fluorescence at a wavelength distinguishable from other reporter moieties, allowing for the monitoring of the presence of different reporter moieties in the same or different reactions. Two or more different reporter moieties may be selected that have spectrally distinct emission profiles or that have minimal overlapping spectral emission profiles. The reporter moiety may be bound (e.g., operably bound) to a nucleotide, a nucleoside, a nucleic acid, an enzyme (e.g., a polymerase or reverse transcriptase), or a support (e.g., a surface).
[0048] The reporter moiety (or label) comprises a fluorescent label or fluorophore. Exemplary fluorescent moieties that can function as fluorescent labels or fluorophores include fluorescein and fluorescein derivatives, such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynapthofluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamidofluorescein, fluorescein maleimide, SAMSA-fluorescein, fluorescein thiosemicarbazide, carbohydrazinomethylthioacetyl-aminofluorescein, rhodamine and rhodamine derivatives, such as TRITC, TMR, lissamine rhodamine, Texas Red, rhodamine B, rhodamine 6G, rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, lissamine rhodamine B sulfonyl chloride, lissamine rhodamine B sulfonylhydrazine, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550 C3, BODIPY 530 / 550 C3-SE, BODIPY 530 / 550 C3 hydrazide, BODIPY 493 / 503 C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FL IA, BODIPY 530 / 551 IA, Br-BODIPY 493 / 503, Cascade Blue and derivatives such as Cascade Blue acetyl azide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Yellow and derivatives, such as Lucifer Yellow iodoacetamide, Lucifer Yellow CH, cyanines and derivatives, such as indolium-based cyanine dyes, benzo-indolium-based cyanine dyes, pyridium-based cyanine dyes, thiozolium-based cyanine dyes, quinolinium-based cyanine dyes, imidazolium-based cyanine dyes, Cy3, Cy5,Lanthanide chelates and derivatives, such as BCPDA, TBP, TMT, BHHCT, BCOT, europium chelates, terbium chelates, Alexa Fluor dyes, DyLight dyes, Atto dyes, LightCycler Red dyes, CAL Flour dyes, JOE and its derivatives, Oregon Green dyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, malachite green, stilbenes, DEG dyes, NR dyes, near-infrared dyes, and others known in the art, such as those described in Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th Edition, Lakowicz, Principles of Fluorescence Spectroscopy, 2nd Ed., Plenum Press New York (1999), or Hermanson, Bioconjugate Techniques, 2nd Edition, or derivatives thereof, or any combination thereof. Cyanine dyes can exist in either sulfonated or non-sulfonated form and consist of two indolenine, benzoindolium, pyridium, thiozolium, and / or quinolinium groups separated by a polymethine bridge between the two nitrogen atoms. Commercially available cyanine fluorophores include, for example, Cy3 (which is 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium, may include 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium-5-sulfonate), Cy5 (which may include1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-indolin-2-ylidene)penta-1,3-dien-1-yl)-3,3-dimethyl-3H-yne dol-1-ium, or 1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-sulfoindolin-2-ylidene)penta-1,3-dien-1-yl) Cy7 (which may include 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), and Cy8 (which may include 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), where "Cy" stands for "cyanine" and the first number identifies the number of carbon atoms between the two indolenine groups. Cy2, which is an oxazole derivative rather than an indolenine, and benzo-derivatized Cy3.5, Cy5.5, and Cy7.5 are exceptions to this rule.
[0049] In some embodiments, the reporter moieties may be FRET pairs, allowing multiple classifications to be performed under a single excitation and imaging step. As used herein, FRET may include excitation exchange (Förster) transfer or electron exchange (Dexter) transfer.
[0050] As used herein, the term "support" refers to a substrate designed for the deposition of biomolecules or biological samples for assay and / or analysis. Examples of biomolecules deposited on a support include nucleic acids (e.g., DNA, RNA), polypeptides, sugars, lipids, single cells, or multiple cells. Examples of biological samples include, but are not limited to, saliva, sputum, mucus, blood, plasma, serum, urine, feces, sweat, tears, and fluids from tissues or organs.
[0051] In some embodiments, the support is solid, semi-solid, or a combination of both. In some embodiments, the support is porous, semi-porous, non-porous, or any combination of porous. In some embodiments, the support can be substantially planar, concave, convex, or any combination thereof. In some embodiments, the support can be cylindrical, for example, comprising a capillary or the interior surface of a capillary.
[0052] In some embodiments, the surface of the support can be substantially smooth, hi some embodiments, the support can have a regular or irregular texture, for example, including ridges, etchings, pores, a three-dimensional scaffold, or any combination thereof.
[0053] In some embodiments, the support comprises beads having any shape, including spherical, hemispherical, cylindrical, barrel-shaped, toroidal, disk-shaped, rod-shaped, conical, triangular, cubic, polygonal, tubular, or wire-shaped.
[0054] The support can be made of any material, including, but not limited to, glass, fused silica, silicon, polymers (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic substrates are contemplated.
[0055] In some embodiments, the present disclosure provides a plurality (e.g., two or more) of nucleic acid template molecules immobilized on a support. In some embodiments, the immobilized plurality of nucleic acid template molecules have the same sequence. In some embodiments, the immobilized plurality of nucleic acid template molecules have different sequences. In some embodiments, individual nucleic acid template molecules in the plurality of nucleic acid template molecules are immobilized at different sites on the support. In some embodiments, two or more individual nucleic acid template molecules in the plurality of nucleic acid templates are immobilized at sites on the support.
[0056] The term "array" refers to a support comprising a plurality of sites located at predetermined locations on the support, forming an array of sites. The sites may be dispersed and separated by interstitial regions. In some embodiments, the predetermined sites on the support may be arranged in rows or columns in one dimension, or in rows and columns in two dimensions. In some embodiments, the plurality of predetermined sites are arranged in an organized manner on the support. In some embodiments, the plurality of predetermined sites are arranged in any organized pattern, including linear, hexagonal, lattice, patterns with reflection symmetry, patterns with rotational symmetry, etc. The pitch between different pairs of sites may be the same or may vary. In some embodiments, the support has a pitch of at least 10 2 at least 10 sites3 at least 10 sites 4 at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites 9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13 at least 10 sites 14 at least 10 sites 15 In some embodiments, the substrate comprises a plurality of predetermined sites (e.g., 10 sites, 100 sites, or more), the sites being located at predetermined locations on the substrate. 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 At a plurality of predetermined sites (e.g., 10 or more sites), nucleic acid template molecules are immobilized to form a nucleic acid template array. In some embodiments, at a plurality of predetermined sites, nucleic acid template molecules are immobilized by hybridization to immobilized surface capture primers, or the nucleic acid template molecules are covalently attached to the surface capture primers. In some embodiments, at a plurality of predetermined sites, nucleic acid template molecules are immobilized by hybridization to immobilized surface capture primers, or the nucleic acid template molecules are covalently attached to the surface capture primers. 2 ~10 15 More than one site (e.g., 10 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 103 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 In some embodiments, the immobilized nucleic acid template molecules are clonally amplified to generate immobilized nucleic acid clusters at a plurality of predetermined sites. In some embodiments, the individual immobilized nucleic acid clusters comprise linear clusters or single- or double-stranded concatemers.
[0057] In some embodiments, a support comprising a plurality of sites located at random positions on the support is referred to herein as a support having randomly located sites thereon. In such embodiments, the locations of the randomly located sites on the support are not predetermined. As a result, the plurality of randomly located sites are arranged in an irregular and / or unpredictable manner on the support. In some embodiments, the support has at least 10 2 at least 10 sites 3 at least 10 sites 4 at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites 9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13 at least 10 sites 14 sites, or at least 10 15In some embodiments, the substrate comprises a plurality of randomly located sites (e.g., 10 sites, 100 sites, or more), where the sites are randomly located on the substrate. 2 ~10 15 More than 10 sites, e.g., 10 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 At a plurality of randomly located sites (e.g., 10 or more sites), a nucleic acid template molecule is immobilized. In some embodiments, the nucleic acid template molecule is immobilized at a plurality of randomly located sites by hybridization to an immobilized surface capture primer, or the nucleic acid template molecule is covalently attached to a surface capture primer. In some embodiments, the nucleic acid template molecule is immobilized at a plurality of randomly located sites, e.g., 10 or more sites. 2 ~10 15 More than 10 sites, e.g., 10 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15In some embodiments, the immobilized nucleic acid template molecule is clonally amplified to generate immobilized nucleic acid clusters at multiple randomly located sites. In some embodiments, the individual immobilized nucleic acid clusters comprise linear clusters or single- or double-stranded concatemers.
[0058] In some embodiments, multiple immobilized surface capture primers on a support (e.g., located at predetermined or random locations on the support) are in fluid communication with one another, allowing solutions of reagents (e.g., nucleic acid template molecules, soluble primers, enzymes, nucleotides, divalent cations, buffers, etc.) to flow over the support so that multiple immobilized surface capture primers on the support can react with the reagents essentially simultaneously in a massively parallel manner. In some embodiments, the fluid communication of multiple immobilized surface capture primers can be used to perform nucleic acid amplification reactions (e.g., RCA, MDA, PCR, and bridge amplification) essentially simultaneously on multiple immobilized surface capture primers.
[0059] In some embodiments, multiple immobilized nucleic acid clusters on a support are in fluid communication with each other, allowing solutions of reagents (e.g., enzymes, nucleotides, divalent cations, etc.) to flow over the support, thereby allowing multiple immobilized nucleic acid clusters on a support to react with reagents essentially simultaneously in a massively parallel manner. In some embodiments, the fluid communication of multiple immobilized nucleic acid clusters can be used to perform nucleotide binding assays and / or nucleotide polymerization reactions (e.g., primer extension or sequencing) substantially simultaneously on multiple immobilized nucleic acid clusters, and optionally, to perform detection and imaging for massively parallel sequencing.
[0060] In some embodiments, the term "immobilized" and related terms refer to nucleic acid molecules that are bound to a support via covalent or non-covalent interactions, or that are attached to a coating on a support, or that are embedded within a matrix formed by a coating on a support, where the nucleic acid molecules include a surface capture primer, a nucleic acid template molecule, and an extension product of the capture primer. The extension product of the capture primer may include a nucleic acid concatemer (e.g., a nucleic acid cluster). The nucleic acid molecules may be immobilized at predetermined or random positions on the support. The nucleic acid molecules may be immobilized at predetermined or random positions on or within a passivated coating on the support.
[0061] In some embodiments, the term "immobilized" and related terms refer to an enzyme (e.g., a polymerase) that is bound to a support through covalent or non-covalent interactions, or that is attached to a coating on a support, or that is embedded within a matrix formed by a coating on a support. The enzyme can be immobilized at predetermined or random locations on the support. The enzyme can be immobilized at predetermined or random locations on or within a passivated coating on the support.
[0062] In some embodiments, one or more nucleic acid template molecules are immobilized on a support, e.g., immobilized at a site on a support. In some embodiments, one or more nucleic acid template molecules are clonally amplified. In some embodiments, one or more nucleic acid template molecules are clonally amplified off-support (e.g., in solution) and then deposited on a support and immobilized thereon. In some embodiments, a clonal amplification reaction of one or more nucleic acid template molecules is performed on the support, resulting in immobilization on the support. In some embodiments, one or more nucleic acid template molecules are clonally amplified (e.g., in solution or on a support) using a nucleic acid amplification reaction including any one or any combination of polymerase chain reaction (PCR), multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, bridge amplification, isothermal bridge amplification, rolling circle amplification (RCA), circle-circle amplification, helicase-dependent amplification, recombinase-dependent amplification, and / or single-strand binding (SSB) protein-dependent amplification.
[0063] The term "surface primer" and related terms refer to a single-stranded oligonucleotide that is immobilized to a support and includes a sequence that can hybridize to at least a portion of a nucleic acid template molecule. Surface capture primers can be used to immobilize template molecules to a support via hybridization. Surface capture primers can be immobilized to a support in a manner that resists primer removal during flow, washing, suction, and changes in temperature, pH, salt, chemical, and / or enzyme conditions. Typically, although not necessarily, the 5' end of a surface capture primer can be immobilized to (or embedded within) a coating on) the support or a coating on the support. Alternatively, an internal portion or 3' end of a surface capture primer can be immobilized to the support.
[0064] The sequences of the surface capture primers can be wholly or partially complementary to at least a portion of the nucleic acid template molecule along their length. The support can contain multiple immobilized surface capture primers with the same sequence or with two or more different sequences. Surface capture primers can be any length, for example, 4-50 nucleotides, 50-100 nucleotides, 100-150 nucleotides, or longer.
[0065] A surface capture primer can have a terminal 3' nucleotide with a sugar 3' OH moiety that is extendible for nucleotide polymerization (e.g., polymerase-catalyzed polymerization). A surface capture primer can have a terminal 3' nucleotide with a 3' sugar position attached to a chain-terminating moiety that inhibits nucleotide polymerization. The 3' chain-terminating moiety can be removed (e.g., deblocked) using a deblocking agent to convert the 3' end to an extendible 3' OH end. Examples of chain-terminating moieties include, but are not limited to, alkyl, alkenyl, alkynyl, allyl, aryl, benzyl, azide, amine, amide, keto, isocyanate, phosphate, thio, disulfide, carbonate, urea, cetal, or silyl groups. Azido-type chain-terminating moieties include azide, azido, and azidomethyl groups. Examples of deblocking agents include phosphine compounds, such as tris(2-carboxyethyl)phosphine (TCEP) and bis-sulfotriphenylphosphine (BS-TPP), for chain-terminating azide, azido, and azidomethyl groups. Examples of deblocking agents include tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ) for chain-terminating alkyl, alkenyl, alkynyl, and aryl groups. Examples of deblocking agents include Pd / C for chain-terminating aryl and benzyl groups. Examples of deblocking agents include phosphines, beta-mercaptoethanol, or dithiothreitol (DTT) for chain-terminating amine, amide, keto, isocyanate, phosphate, thio, and disulfide groups. Examples of deblocking agents include potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, and Zn in acetic acid (AcOH) for carbonate chain terminating groups.Examples of deblocking agents include tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, and triethylamine trihydrofluoride for urea and silyl chain terminating groups.
[0066] The term "sequencing" and related terms refer to methods for obtaining nucleotide sequence information from a nucleic acid molecule, typically by determining the identities of at least some nucleotides (including their nucleobase components) within the nucleic acid molecule. In some embodiments, sequence information for a given region of a nucleic acid molecule includes identifying each and every nucleotide within the sequenced region. In some embodiments, the sequencing information determines only some of the region of nucleotides, with the identities of some nucleotides remaining undetermined or incorrectly determined. Any suitable sequencing method can be used. In exemplary embodiments, sequencing can include label-free or ion-based sequencing methods. In some embodiments, sequencing can include labeled, dye-containing, or fluorescent-based nucleotide sequencing methods. In some embodiments, sequencing can include polony-based sequencing or bridge sequencing methods. In some embodiments, sequencing uses a polymerase and a multivalent molecule to generate at least one avidity complex, with each multivalent molecule comprising multiple nucleotide units tethered to a core. In some embodiments, sequencing uses a polymerase and free nucleotides to perform sequencing by synthesis, hi some embodiments, sequencing uses a ligase enzyme and multiple sequence-specific oligonucleotides to perform sequencing by ligation.
[0067] In some aspects, the present disclosure provides various reagents for nucleic acid denaturation (dehybridization) and sequencing, and methods using the reagents. The various reagents may include at least one pH buffer. The full names of exemplary, non-limiting pH buffers are listed herein.
[0068] The term "Tris" refers to the pH buffer tris(hydroxymethyl)-aminomethane. The term "TrisHCl" refers to the pH buffer tris(hydroxymethyl)-aminomethane hydrochloride. The term "Trisacetate" refers to a pH buffer containing the acetate salt of tris(hydroxymethyl)-aminomethane.
[0069] The term "tricine" refers to the pH buffering agent N-[tris(hydroxymethyl)methyl]glycine.
[0070] The term "bicine" refers to the pH buffering agent N,N-bis(2-hydroxyethyl)glycine. The term "bis-trispropane" refers to the pH buffering agent 1,3 bis[tris(hydroxymethyl)methylamino]propane.
[0071] The term "HEPES" refers to the pH buffering agent 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid.
[0072] The term "MES" refers to 2-(N-morpholino)ethanesulfonic acid), a pH buffering agent.
[0073] The term "MOPS" refers to the pH buffering agent 3-(N-morpholino)propanesulfonic acid.
[0074] The term "MOPSO" refers to the pH buffering agent 3-(N-morpholino)-2-hydroxypropanesulfonic acid.
[0075] The term "BES" refers to the pH buffering agent N,N-bis(2-hydroxyethyl)-2-aminoethanesulfonic acid.
[0076] The term "TES" refers to the pH buffering agent 2-[(2-hydroxy-1,1 bis(hydroxymethyl)ethyl)amino]ethanesulfonic acid).
[0077] The term "CAPS" refers to the pH buffering agent 3-(cyclohexylamino)-1 propanesuhinic acid.
[0078] The term "TAPS" refers to the pH buffering agent N-[tris(hydroxymethyl)methyl]-3-aminopropanesulfonic acid.
[0079] The term "TAPSO" refers to the pH buffering agent N-[tris(hydroxymethyl)methyl]-3-amino-2-hydroxypropanesulfonic acid.
[0080] The term "ACES" refers to the pH buffering agent N-(2-acetamido)-2-aminoethanesulfonic acid.
[0081] The term "PIPES" refers to piperazine-1,4-bis(2-ethanesulfonic acid), a pH buffering agent.
[0082] The term "ethanolamine" refers to the pH buffering agent also known as 2-aminoethanol.
[0083] Introduction: Two-strand splint adapter In some aspects, the disclosure provides compositions, including kits, comprising nucleic acid double-stranded splint adaptors, and methods of using double-stranded splint adaptors.
[0084] The double-stranded splint adapter (200) can be used in a one-pot multi-enzyme reaction to introduce one or more new adapter sequences into a library molecule. In some embodiments, the double-stranded splint adapter (200) comprises a first splint strand (long splint strand (300)) and a second splint strand (short splint strand (400)). In certain embodiments, the first and second splint strands are hybridized together to form a double-stranded splint adapter (200) having a double-stranded region and two adjacent single-stranded regions (see, e.g., Figures 2 and 3). In some embodiments, the first splint strand (long splint strand (300)) and the second splint strand (short splint strand (400)) hybridize to each other along their entire lengths. In some embodiments, the first splint strand (long splint strand (300)) and the second splint strand (short splint strand (400)) comprise portions that are not hybridized together. The second splint strand (400) may carry the new adapter sequence(s) introduced, such as a new universal binding sequence and / or a new index sequence. The first splint strand may include a first region (320), an internal region (310), and a second region (330). The internal region (310) of the first splint strand may hybridize to the second splint strand (400). In some embodiments, the two adjacent single-stranded regions (e.g., (320) and (330)) of the double-stranded splint adapter are designed to hybridize to the universal adapter sequences at the ends of the single-stranded linear library molecule (100) having the sequence of interest (110). For example, the first region (320) of the first splint strand may hybridize to one end (120) of the library molecule, and the second region (330) of the first splint strand may hybridize to the other end (130) of the library molecule, thereby circularizing the library molecule and generating a library-sprint complex (500) containing two nicks (see, e.g., Figure 3).
[0085] In some embodiments, the linear nucleic acid library molecule (100) comprises an insert region (110) (e.g., a sequence of interest) flanked on one side by a universal adapter sequence (120) for a forward sequencing primer binding site, and the insert region (110) is flanked on the other side by a universal adapter sequence (130) for a reverse sequencing primer binding site. In some embodiments, the double-stranded splint adapter (200) comprises a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand may include a first region (320) that hybridizes with a universal adapter sequence (120) for a forward sequencing primer binding site on one end of the linear single-stranded library molecule (100), and the first splint strand may include a second region (330) that hybridizes with a universal adapter sequence (130) for a reverse sequencing primer binding site on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint strand may hybridize to the second splint strand (400). The nick may be enzymatically ligated to generate a covalently closed circular molecule (600) in which the second splint strand (400) is covalently attached to the library molecule at both ends, thereby introducing a new adapter sequence into the library molecule (see, e.g., Figure 4). A ligation reaction can link sequences from the second splint strand (400) to the end of the library molecule (100).
[0086] In some embodiments, any of the subregions of the second splint strand (400) can be designed to hybridize to an amplification primer. In some embodiments, the amplification primer has an extendable 3' end that can be used to initiate a primer extension reaction, for example, as shown in Figure 5. In some embodiments, the amplification primer can be a soluble primer or immobilized on a support (e.g., a surface capture primer). In some embodiments, the amplification primer can be used to perform a rolling circle amplification reaction to generate nucleic acid concatemer molecules that are complementary to the covalently closed circular library molecule (600).
[0087] Thus, any linear library molecule can be converted into a covalently closed circular molecule using the double-stranded splint adapters and methods described herein. The second splint strand (400) can include at least one new universal adapter sequence (e.g., a new surface primer sequence), thereby enabling the covalently closed circular molecule (600) to bind to a support having multiple surface capture primers immobilized thereon. The new universal adapter sequence(s) in the second splint strand (400) enables the use of the covalently closed circular molecule (600) in amplification and sequencing workflows.
[0088] The methods described herein also offer the advantage of introducing new adapter sequences using a ligation reaction rather than a gap-fill reaction, which results in highly efficient circularization with as few as 0.25 pmol of library molecules.
[0089] Because annealing and multi-enzyme reactions can be performed in a single reaction vessel (one-pot) by combining several enzymatic reactions (e.g., phosphorylation and ligation) and adding subsequent enzymes (e.g., exonucleases) without alcohol precipitation or organic extraction, the methods described herein can be performed manually or easily adapted for automation.
[0090] Two-strand splint adapter In some embodiments, the present disclosure provides a nucleic acid double-stranded splint adapter (200) comprising (i) a first splint strand (long splint strand (300)) that hybridizes to (ii) a second splint strand (short splint strand (400)) (see, e.g., Figures 2 and 3). The first splint strand can comprise a first region (320), an internal region (310), and a second region (330). The internal region (310) of the first splint strand can hybridize to the second splint strand (400) to form a double-stranded splint adapter (200) having a double-stranded region and two adjacent single-stranded regions. The two adjacent single-stranded regions of the double-stranded splint adapter (200) can be designed to hybridize to end sequences of a linear nucleic acid library molecule (100). The terminal sequences of the linear nucleic acid library molecule can comprise first (120) and second (130) universal adaptor sequences, respectively. In some embodiments, the first and second universal adaptor sequences of the linear nucleic acid library molecule comprise binding sequences for forward (120) and reverse (130) sequencing primer binding sites, respectively.
[0091] As illustrated in Figure 2, in some embodiments, the first splint strand comprises a first sequence (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second sequence (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint strand may hybridize to the second splint strand (400). The internal region (310) of the first splint strand (300) may comprise at least three subregions, including subregions (311), (312), and (313). The second splint strand (400) may comprise at least three subregions, including subregions (411), (412), and (413). For example, subregion (311) hybridizes to subregion (411). In another example, subregion 312 hybridizes to subregion 412. In another example, subregion 313 hybridizes to subregion 413. The subregions of second splint strand 400 may comprise, in any combination and in any order, any one or more of a universal primer binding sequence for a surface capture primer, a universal primer binding sequence for a surface pinning primer, a sample index sequence, a short random sequence, and / or a unique molecular index (UMI) sequence.
[0092] In some embodiments, the first region (320) of the first splint strand comprises a first universal adaptor sequence capable of hybridizing to a first universal binding sequence at one end of a linear nucleic acid library molecule (see, e.g., FIG. 2). The second region (330) of the first splint strand may comprise a second universal adaptor sequence capable of hybridizing to a second universal binding sequence at the other end of a linear nucleic acid library molecule (see, e.g., FIG. 2). In some embodiments, the first region (320) of the first splint strand comprises a first universal adaptor sequence comprising a universal binding sequence for a forward or reverse sequencing primer. In some embodiments, the second region (330) of the first splint strand comprises a second universal adaptor sequence comprising a universal binding sequence for a forward or reverse sequencing primer. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0093] In some embodiments, the second splint strand (400) comprises at least three subregions, including a first, second, and third subregion (see, e.g., Figures 2 and 3). The first subregion (411) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The second subregion (412) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The third subregion (413) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5-20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3-20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule to which the adapter is attached (e.g., molecular tagging). In some embodiments, the unique molecular index (UMI) comprises a random sequence. In some embodiments, the second splint strand (400) is designed to exhibit reduced or no hybridization to the insert sequence (110) of the library molecule (100).
[0094] An exemplary arrangement of sequences within the subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) includes a universal sequence for binding a surface pinning primer]-[(412) includes a sample index sequence and an optional short random sequence NNN]-[(413) includes a universal sequence for binding a surface capture primer].
[0095] An exemplary arrangement of sequences within the subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) includes a universal sequence for binding a surface pinning primer]-[(412) includes a universal sequence for binding a surface capture primer]-[(413) includes a sample index sequence and an optional short random sequence NNN].
[0096] An exemplary arrangement of sequences within the subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) includes a sample index sequence and an optional short random sequence NNN] - [(412) includes a universal sequence for binding a surface pinning primer] - [(413) includes a universal sequence for binding a surface capture primer].
[0097] In some embodiments, the second splint strand (400) comprises an additional subregion carrying a second sample index sequence and an optional short random sequence NNN. For example, the additional subregion can be located between subregions (411) and (412) or between subregions (412) and (413).
[0098] In some embodiments, the second splint strand (400) can be 20 to 100 (e.g., about 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100) nucleotides in length. In some embodiments, the second splint strand (400) can be 30 to 80 (e.g., about 30, 35, 40, 45, 50, 60, 70, or 80) nucleotides in length, or 40 to 60 (e.g., about 40, 45, 50, or 60) nucleotides in length. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated; alternatively, the 5' end of the second splint strand (400) is not phosphorylated. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group; alternatively, the 3' end of the second splint strand (400) comprises a terminal 3' blocking group. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position, for example, to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.
[0099] In some embodiments, the first splint strand (300) comprises an internal region (310) comprising at least three subregions, including a fourth subregion (311), a fifth subregion (312), and a sixth subregion (313). The fourth subregion (311) can hybridize to a first subregion (411) of the second splint strand (400). The fourth subregion (311) can be fully or partially complementary to the first subregion (411) of the second splint strand (400). The fifth subregion (312) can hybridize to a second subregion (412) of the second splint strand (400). The fifth subregion (312) can be fully or partially complementary to the second subregion (412) of the second splint strand (400). The sixth subregion (313) can hybridize to the third subregion (413) of the second splint strand (400). The sixth subregion (313) can be fully or partially complementary to the third subregion (413) of the second splint strand (400). The fourth, fifth, and sixth subregions do not hybridize (or at least show very little hybridization) to the sequence of interest, surface capture primer, or surface pinning primer.
[0100] In some embodiments, one of the subregions of the first splint strand (300) comprises an index or random sequence, such as a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). For example, subregions (311), (312), or (313) comprise a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5 to 20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3 to 20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule to which the adaptor is attached (e.g., molecular tagging). In some embodiments, the unique molecular index (UMI) comprises a random sequence.
[0101] As shown in Figure 6, in some embodiments, subregion (312) comprises any one or any combination of two or more of a sample index sequence (indicated by "S"), a random sequence (indicated by "N"), at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8-oxoG), or deoxyinosine), deoxyinosine (indicated by "I"), and / or a spacer (e.g., an 18-carbon spacer). Figure 6 shows an alternative exemplary nucleic acid sequence for the second splint strand (400, short splint strand), which comprises three subregions, including subregion (411), subregion (412), and subregion (413).
[0102] In some embodiments, the first splint strand subregion comprising the index or random sequence comprises any one or any combination of two or more of the following: a sample index sequence (denoted by "S"), a random sequence (denoted by "N"), at least one nucleotide (denoted by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8-oxoG), or deoxyinosine), deoxyinosine (denoted by "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer. See Figure 6 for non-limiting examples.
[0103] In some embodiments, the double-stranded splint adapter (200) comprises a first splint strand (300) partially hybridized to a second splint strand (400), where the first splint strand (300) comprises a universal length splint strand. In some embodiments, the universal length splint strand comprises at least one subregion only partially hybridized to the second splint strand (400). For example, but not limited to, the universal length splint strand (300) comprises a subregion bearing at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8oxoG), or deoxyinosine), deoxyinosine (indicated by an "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer. An exemplary universal long splint chain (300) is shown in Figure 6.
[0104] In some embodiments, the first splint strand lacks subregion (312), as shown in Figure 7. Figure 7 also shows an exemplary nucleic acid sequence of the second splint strand (400, short splint strand), which comprises three subregions, including subregion (411), subregion (412), and subregion (413). In certain embodiments, when the first splint strand (300) hybridizes to the second splint strand (400), subregion (412) of the second splint strand (400) loops out because the first splint strand (300) lacks subregion (312).
[0105] In some embodiments, the first splint strand (300) lacks subregions (311), (312), or (313), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has a portion of the second splint strand looped out. For example, without limitation, the first splint strand (300) lacks subregion (312), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has subregion (412) of the second splint strand looped out (see, e.g., FIG. 7). In some embodiments, a first splint strand (300) lacking a subregion is an example of a universal length splint strand (300).
[0106] In some embodiments, the first splint strand (300) can be 50-150 nucleotides in length, or 60-100 nucleotides in length, or 70-90 nucleotides in length. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at the 5' and / or 3' ends to confer exonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' ends. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at an internal position.
[0107] Tables 1-6 below list various embodiments of the universal adapter sequence in the library molecule (100), the sequence in the first splint strand (300), the sequence in the second splint strand (400), the sequence of the immobilized surface primer, and the sequence of the sequencing primer.
[0108] In some embodiments, any of the universal adaptor sequences in the library molecule (100) listed in Table 1 can be truncated at the 5' and / or 3' end, and the truncation can be 1 to 12 nucleotides. In some embodiments, the truncation can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated.
[0109] In some embodiments, the sequencing primer comprises a sequence that is complementary to any of the sequences listed in Table 1 below.
[0110] In some embodiments, the sequencing primer comprises a sequence complementary to any of the sequences listed in Table 1 below, and the sequencing primer is truncated at the 5' and / or 3' end, and the truncation can be 1 to 12 nucleotides. In some embodiments, the truncation can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated.
[0111] In some embodiments, any of the sequences in the first splint strand (300) and / or the second splint strand (400) comprises a sequence that is complementary to any of the sequences listed in Tables 1-6 below.
[0112] In some embodiments, any of the sequences in the first splint strand (300) and / or second splint strand (400) listed in Tables 2-3 can be truncated at the 5' and / or 3' end, and the truncation can be 1 to 12 nucleotides. In some embodiments, the truncation can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated. [Table 1] [Table 2] [Table 3-1] [Table 3-2] [Table 4] [Table 5] [Table 6]
[0113] Library-Sprint Complex In some aspects, the present disclosure provides a library-sprint complex (500) comprising: (i) a single-stranded nucleic acid library molecule (100) comprising an insert region (110) (e.g., a sequence of interest) flanked on one side by a universal adapter sequence (120) for a forward sequencing primer binding site, the insert region (110) being flanked on the other side by a universal adapter sequence (130) for a reverse sequencing primer binding site; and (ii) a double-stranded splint adapter (200) comprising a first splint strand (long splint strand (300)) hybridized to a second splint strand (short splint strand (400)), the first splint strand comprising a first region (320), an internal region (310), and a second region (330). In certain embodiments, the internal region (310) of the first splint strand hybridizes to the second splint strand (400) to form a double-stranded splint adapter (200) having a double-stranded region and two adjacent single-stranded regions.
[0114] In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a universal adapter sequence for a surface capture primer binding site. In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a universal adapter sequence for a surface pinning primer binding site. In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a universal adapter sequence for a sample index sequence. In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a unique molecular index (UMI) sequence.
[0115] In the library-sprint complex (500), the first region (320) of the first splint strand can hybridize to the universal adapter sequence (120) for the forward sequencing primer binding site of the library molecule, and the second region (330) of the first splint strand can hybridize to the universal adapter sequence (130) for the reverse sequencing primer binding site of the library molecule, thereby circularizing the library molecule and generating the library-sprint complex (500) (see, for example, Figures 3 and 4).
[0116] In the library-sprint complex (500), the second splint strand (400) can bring the library molecule (100) into proximity with at least a universal adapter sequence for a surface pinning primer binding site, a universal adapter sequence for a surface capture primer binding site, and a sample index sequence. In some embodiments, the second splint strand (400) brings the library molecule (100) into proximity with other sequences, including short random sequences (e.g., NNN) and / or unique molecular index (UMI) sequences.
[0117] The library-sprint complex (500) can include a first nick between the 5' end of the library molecule and the 3' end of the second splint strand. In certain embodiments, the library-sprint complex (500) also includes a second nick between the 5' end of the second splint strand and the 3' end of the library molecule (see, e.g., Figures 3 and 4). In some embodiments, the first and second nicks are enzymatically ligatable. The ligation reaction joins the sequence from the second splint strand (400) to the end of the library molecule (100).
[0118] In the library-sprint complex (500), the first region (320) of the first splint strand can hybridize to either the sense or antisense strand of the double-stranded nucleic acid library molecule. In the library-sprint complex (500), the second region (330) of the first splint strand can hybridize to either the sense or antisense strand of the double-stranded nucleic acid library molecule. The double-stranded nucleic acid library molecule can be denatured to generate single-stranded sense and antisense library strands.
[0119] In the library-sprint complex (500), the first region (320) of the first splint strand does not hybridize to the sequence of interest (110), and the second region (330) of the first splint strand does not hybridize to the sequence of interest (110).
[0120] In the library-sprint complex (500), the second splint strand (400) does not hybridize to the sequence of interest (110), and the internal region (310) of the first splint strand does not hybridize to the sequence of interest (110).
[0121] In some embodiments, in the library-sprint complex (500), the internal region (310) of the first splint strand (300) hybridizes to the second splint strand (400). The internal region (310) of the first splint strand (300) can include at least three subregions, including subregions (311), (312), and (313). The second splint strand (400) can include at least three subregions, including subregions (411), (412), and (413). For example, subregion (311) hybridizes to subregion (411). In another example, subregion (312) hybridizes to subregion (412). In another example, subregion (313) hybridizes to subregion (413).
[0122] In some embodiments, any of the library-sprint complexes (500) described herein includes a plurality of library-sprint complexes (500), and the target sequences (110) of the individual library-sprint complexes in the plurality of library-sprint complexes (500) include the same target sequence or different target sequences.
[0123] In some embodiments, in the library-sprint complex (500), the 5' ends of the single-stranded library molecules (100) are phosphorylated. In some embodiments, in the library-sprint complex (500), the 5' ends of the single-stranded library molecules (100) lack a phosphate group. In some embodiments, the 3' ends of the single-stranded library molecules comprise a terminal 3' OH group or a terminal 3' blocking group.
[0124] In some embodiments, in the library-sprint complex (500), the first splint strand (300) can be 50 to 150 (e.g., about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150) nucleotides in length. In some embodiments, in the library-sprint complex (500), the first splint strand (300) can be 60 to 100 (e.g., about 60, 70, 80, 90, or 100) nucleotides in length. In some embodiments, in the library-sprint complex (500), the first splint strand (300) can be 70 to 90 (e.g., about 70, 75, 80, 85, or 90) nucleotides in length. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at an internal position. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated. In some embodiments, the 5' end of the first splint strand (300) lacks a phosphate group. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0125] In some embodiments, in the library-sprint complex (500), the first splint strand (300) comprises an internal region (310) comprising at least three subregions, including a fourth subregion (311), a fifth subregion (312), and a sixth subregion (313). The fourth subregion (311) can hybridize to a first subregion (411) of the second splint strand (400). The fourth subregion (311) can be fully or partially complementary to the first subregion (411) of the second splint strand (400). The fifth subregion (312) can hybridize to a second subregion (412) of the second splint strand (400). The fifth subregion (312) can be fully or partially complementary to the second subregion (412) of the second splint strand (400). The sixth subregion (313) can hybridize to the third subregion (413) of the second splint strand (400). The sixth subregion (313) can be fully or partially complementary to the third subregion (413) of the second splint strand (400). The fourth, fifth, and sixth subregions do not hybridize (e.g., minimally exhibit very little hybridization) to the sequence of interest, surface capture primer, or surface pinning primer.
[0126] In some embodiments, one of the subregions of the first splint strand (300) comprises an index or random sequence, such as a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). For example, without limitation, subregions (311), (312), or (313) comprise a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5 to 20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3 to 20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule to which the adaptor is attached (e.g., molecular tagging). In some embodiments, the unique molecular index (UMI) comprises a random sequence.
[0127] In some embodiments, the first splint strand subregion comprising the index or random sequence comprises any one or any combination of two or more of the following: a sample index sequence (denoted by "S"), a random sequence (denoted by "N"), at least one nucleotide (denoted by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8-oxoG), or deoxyinosine), deoxyinosine (denoted by "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer.
[0128] In some embodiments, the double-stranded splint adapter (200) comprises a first splint strand (300) partially hybridized to a second splint strand (400), where the first splint strand (300) comprises a universal length splint strand. In some embodiments, the universal length splint strand comprises at least one subregion only partially hybridized to the second splint strand (400). For example, the universal length splint strand (300) can comprise a subregion bearing at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8oxoG), or deoxyinosine), deoxyinosine (indicated by an "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer. An exemplary universal long splint chain (300) is shown in Figure 6.
[0129] In some embodiments, in the library-sprint complex (500), the first splint strand (300) lacks subregions (311), (312), or (313), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has a portion of the second splint strand looped out. For example, the first splint strand (300) lacks subregion (312), and in the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400), subregion (412) of the second splint strand (400) loops out (see, e.g., Figure 7). In some embodiments, the first splint strand (300) lacking a subregion is an example of a universal length splint strand (300).
[0130] In some embodiments, in the library-sprint complex (500), the second splint strand (400) comprises at least three subregions, including a first, second, and third subregion (see, e.g., Figures 3 and 4). The first subregion (411) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The second subregion (412) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The third subregion (413) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5-20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3-20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule (e.g., molecular tagging) to which the adapter is attached. In some embodiments, the unique molecular index (UMI) comprises a random sequence. The second splint strand (400) is designed to exhibit reduced or no hybridization to the insert sequence (110) of the library molecule (100).
[0131] In some embodiments, in the library-sprint complex (500), an exemplary arrangement within a subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) comprising a universal sequence for binding a surface pinning primer]-[(412) comprising a sample index sequence and an optional short random sequence NNN]-[(413) comprising a universal sequence for binding a surface capture primer].
[0132] In some embodiments, in the library-sprint complex (500), an exemplary arrangement within a subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) includes a universal sequence for binding a surface pinning primer]-[(412) includes a universal sequence for binding a surface capture primer]-[(413) includes a sample index sequence and an optional short random sequence NNN].
[0133] In some embodiments, in the library-sprint complex (500), an exemplary arrangement within a subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) comprising a sample index sequence and an optional short random sequence NNN]-[(412) comprising a universal sequence for binding a surface pinning primer]-[(413) comprising a universal sequence for binding a surface capture primer].
[0134] In some embodiments, in library-sprint complex (500), second splint strand (400) includes an additional subregion carrying a second sample index sequence and an optional short random sequence NNN. For example, but not limited to, the additional subregion can be located between subregions (411) and (412) or between subregions (412) and (413).
[0135] In some embodiments, in the library-sprint complex (500), the second splint strand (400) can be 20 to 100 (e.g., about 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100) nucleotides in length. In some embodiments, the second splint strand (400) can be 30 to 80 nucleotides in length, 80 (e.g., about 30, 35, 40, 45, 50, 60, 70, or 80) nucleotides in length, or 40 to 60 nucleotides in length, 60 (e.g., about 40, 45, 50, or 60) nucleotides in length. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated; alternatively, the 5' end of the second splint strand (400) is not phosphorylated. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group; alternatively, the 3' end of the second splint strand (400) comprises a terminal 3' blocking group. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position, for example, to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at an internal position.
[0136] In the library-sprint complex (500), the second splint strand (400) brings the library molecule (100) into proximity with at least a universal adapter sequence for the surface pinning primer binding site, a universal adapter sequence for the surface capture primer binding site, and a sample index sequence. In some embodiments, the second splint strand (400) brings the library molecule (100) into proximity with other sequences, including short random sequences (e.g., NNN) and / or unique molecular index (UMI) sequences.
[0137] Tables 1-4 above list various embodiments of the sequences of the various molecules that form the library-sprint complex (500), including the library molecule (100), the first splint strand (300), and the second splint strand (400). In some embodiments, the sequences in the first splint strand (300) and / or the second splint strand (400) can include sequences that are complementary to sequences listed in Tables 2-3. In some embodiments, any of the sequences in the first splint strand (300) and / or the second splint strand (400) listed in Tables 2-3 can be truncated at the 5' and / or 3' ends, and the truncation can be 1 to 12 nucleotides. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated. In some embodiments, the library molecules have forward sequencing primer binding site (120) and / or reverse sequencing primer binding site (130) sequences listed in Table 1, which may be truncated at the 5' and / or 3' ends, and the truncation may be 1 to 12 nucleotides. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated.
[0138] In some aspects, the present disclosure provides a reaction mixture comprising any of the plurality of library-sprint complexes (500) described herein. In some embodiments, the reaction mixture comprises any of the plurality of library-sprint complexes (500) described herein and T4 polynucleotide kinase. In some embodiments, the reaction mixture comprises any of the plurality of library-sprint complexes (500) described herein and lacks T4 polynucleotide kinase. In some embodiments, the reaction mixture comprises any of the plurality of library-sprint complexes (500) described herein and a ligase enzyme. In some embodiments, the reaction mixture comprises any of the plurality of library-sprint complexes (500) described herein and T4 polynucleotide kinase and a ligase enzyme. In some embodiments, the reaction mixture comprises any of the plurality of library-sprint complexes (500) described herein and a ligase enzyme and lacks T4 polynucleotide kinase. In some embodiments, the ligase enzyme comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase.
[0139] Covalent closed ring molecule In some aspects, the present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), a forward sequencing primer binding site (120), a reverse sequencing primer binding site (130), and a sequence forming a second splint strand (400). In some embodiments, the covalently closed circular library molecule (600) comprises the sequence of interest (110) flanked on one side by the forward sequencing primer binding site (120) and on the other side by the reverse sequencing primer binding site (130), wherein the forward sequencing primer binding site (120) and the reverse sequencing primer binding site (130) are covalently linked to the second splint strand (400). Exemplary covalently closed circular library molecules (600) are shown in Figures 4 and 5.
[0140] In some embodiments, any of the covalently closed circular library molecules (600) described herein further comprises a plurality of covalently closed circular library molecules (600), wherein the sequences of interest (110) of each of the covalently closed circular library molecules (600) of the plurality of covalently closed circular library molecules comprise the same sequence of interest. Alternatively, in some embodiments, the sequences of interest (110) comprise different sequences of interest.
[0141] In some embodiments, in the covalently closed circular library molecule (600), the second splint strand (400) comprises at least three subregions, including a first, second, and third subregion (see, e.g., Figures 4 and 5). The first subregion (411) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The second subregion (412) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The third subregion (413) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5-20 bases and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3-20 bases and can be used to uniquely identify the nucleic acid molecule to which the adapter is attached (e.g., molecular tagging). In some embodiments, the unique molecular index (UMI) comprises a random sequence. The second splint strand (400) is designed to exhibit reduced or no hybridization to the insert sequence (110) of the library molecule (100).
[0142] In some embodiments, in the covalently closed circular library molecule (600), the arrangement of sequences within the subregion of the second splint strand (400) comprises, in a 3' to 5' orientation, [(411) comprising a universal sequence for binding a surface pinning primer]-[(412) comprising a sample index sequence and an optional short random sequence NNN]-[(413) comprising a universal sequence for binding a surface capture primer].
[0143] In some embodiments, in the covalently closed circular library molecule (600), the arrangement of sequences within the subregion of the second splint strand (400) comprises, in a 3' to 5' orientation, [(411) comprising a universal sequence for binding a surface pinning primer]-[(412) comprising a universal sequence for binding a surface capture primer]-[(413) comprising a sample index sequence and an optional short random sequence NNN].
[0144] In some embodiments, in the covalently closed circular library molecule (600), the arrangement of sequences within the subregion of the second splint strand (400) comprises, in a 3' to 5' direction, [(411) comprising a sample index sequence and an optional short random sequence NNN]-[(412) comprising a universal sequence for binding a surface pinning primer]-[(413) comprising a universal sequence for binding a surface capture primer].
[0145] In some embodiments, in the covalently closed circular library molecule (600), the second splint strand (400) comprises an additional subregion bearing a second sample index sequence and an optional short random sequence NNN. For example, the additional subregion can be located between subregions (411) and (412) or between subregions (412) and (413).
[0146] In some embodiments, in the covalently closed circular library molecule (600), the second splint strand (400) can be 20 to 100 nucleotides in length, or 30 to 80 nucleotides in length, or 40 to 60 nucleotides in length. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.
[0147] In some embodiments, the covalently closed circular library molecule (600) is hybridized to the first splint strand (300), or the first splint strand (300) is absent. In some embodiments, the first splint strand (300) can be 50-150 (e.g., about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150) nuclease length. In some embodiments, the first splint strand (300) can be 60-100 (e.g., about 60, 70, 80, 90, or 100) nucleotides long. In some embodiments, the first splint strand (300) can be 70-90 (e.g., about 70, 75, 80, 85, or 90) nucleotides long. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at the 5' and / or 3' end, e.g., to confer exonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at an internal position, e.g., to confer endonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at an internal position. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated; alternatively, the 5' end of the first splint strand (300) lacks a phosphate group. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group; alternatively, the 3' end of the first splint strand (300) comprises a terminal 3' blocking group.
[0148] In some embodiments, in the covalently closed circular library molecule (600), the first splint strand (300) comprises an internal region (310) comprising at least three subregions, including a fourth subregion (311), a fifth subregion (312), and a sixth subregion (313). The fourth subregion (311) can hybridize to a first subregion (411) of the second splint strand (400). The fourth subregion (311) can be fully or partially complementary to the first subregion (411) of the second splint strand (400). The fifth subregion (312) can hybridize to a second subregion (412) of the second splint strand (400). The fifth subregion (312) can be fully or partially complementary to the second subregion (412) of the second splint strand (400). The sixth subregion (313) can hybridize to the third subregion (413) of the second splint strand (400). The sixth subregion (313) can be fully or partially complementary to the third subregion (413) of the second splint strand (400). The fourth, fifth, and sixth subregions do not hybridize (e.g., show very little hybridization) to the sequence of interest, surface capture primer, or surface pinning primer.
[0149] In some embodiments, in the covalently closed circular library molecule (600), one of the subregions of the first splint strand (300) comprises an index or random sequence, such as a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). For example, subregions (311), (312), or (313) comprise a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5-20 bases and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3-20 bases and can be used to uniquely identify the nucleic acid molecule to which the adapter is attached (e.g., molecular tagging). In some embodiments, the unique molecular index (UMI) comprises a random sequence.
[0150] In some embodiments, in the covalently closed circular library molecule (600), the subregion (312) comprises any one or any combination of two or more of a sample index sequence (indicated by "S"), a random sequence (indicated by "N"), at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8-oxoG), or deoxyinosine), deoxyinosine (indicated by "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 including a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer.
[0151] In some embodiments, when a covalently closed circular library molecule (600) hybridizes to a first splint strand (300), the first splint strand partially hybridizes to a sequence from a second splint strand (400), and the first splint strand (300) comprises a universal length splint strand. In some embodiments, the universal length splint strand can comprise at least one subregion that is only partially hybridized to a sequence from the second splint strand (400). For example, the universal length splint strand (300) can comprise a subregion bearing at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8oxoG), or deoxyinosine), deoxyinosine (indicated by an "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer. An exemplary universal long splint chain (300) is shown in Figure 6.
[0152] In some embodiments, in a covalently closed circular library molecule (600), the first splint strand (300) lacks subregions (311), (312), or (313), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has a portion of the second splint strand looped out. For example, the first splint strand (300) may lack subregion (312), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has subregion (412) of the second splint strand looped out (see, e.g., Figure 7). In some embodiments, a first splint strand (300) lacking a subregion is an example of a universal length splint strand (300).
[0153] In some embodiments, the exemplary covalently closed circular molecule (600) comprises (i) a sequence of interest (110), (ii) a universal adapter sequence (140) having a binding sequence for a forward sequencing primer, (iii) a universal adapter sequence (150) having a binding sequence for a reverse sequencing primer, (iv) a universal adapter sequence (subregion 411) having a binding sequence for a surface pinning primer, (v) a universal adapter sequence (subregion 413) having a binding sequence for a surface capture primer, and (vi) a sample index, a short random sequence (NNN), and / or a unique molecular index (UMI) (subregion 412), wherein the covalently closed circular molecule (600) is optionally hybridized to the first splint strand (300).
[0154] In some embodiments, the exemplary covalently closed circular molecule (600) comprises (i) a sequence of interest (110), (ii) a universal adapter sequence (140) having a binding sequence for a forward sequencing primer, (iii) a universal adapter sequence (150) having a binding sequence for a reverse sequencing primer, (iv) a universal adapter sequence (subregion 411) having a binding sequence for a surface pinning primer, (v) a universal adapter sequence (subregion 412) having a binding sequence for a surface capture primer, and (vi) a sample index, a short random sequence (NNN), and / or a unique molecular index (UMI) (subregion 413), wherein the covalently closed circular molecule (600) is optionally hybridized to the first splint strand (300).
[0155] In some embodiments, the exemplary covalently closed circular molecule (600) comprises (i) a sequence of interest (110), (ii) a universal adapter sequence (140) having a binding sequence for a forward sequencing primer, (iii) a universal adapter sequence (150) having a binding sequence for a reverse sequencing primer, (iv) a universal adapter sequence (subregion 412) having a binding sequence for a surface pinning primer, (v) a universal adapter sequence (subregion 413) having a binding sequence for a surface capture primer, and (vi) a sample index, a short random sequence (NNN), and / or a unique molecular index (UMI) (subregion 411), wherein the covalently closed circular molecule (600) is optionally hybridized to the first splint strand (300).
[0156] In some embodiments, any of the covalently closed circular molecules (600) described herein comprises a plurality of covalently closed circular molecules (600), wherein the sequence of interest (110) of each of the covalently closed circular molecules (600) of the plurality of covalently closed circular molecules comprises the same sequence of interest. In some embodiments, the sequence of interest (110) of each of the covalently closed circular molecules (600) of the plurality of covalently closed circular molecules comprises a different sequence of interest.
[0157] Tables 1-4 above list various embodiments of sequences forming a covalently closed circular library molecule (600), including a library molecule (100), a first splint strand (300), and a second splint strand (400). In some embodiments, the sequences in the first splint strand (300) and / or the second splint strand (400) comprise sequences complementary to sequences listed in Tables 2-3. In some embodiments, any of the sequences in the first splint strand (300) and / or the second splint strand (400) listed in Tables 2-3 can be truncated at the 5' and / or 3' ends, and the truncation can be 1 to 12 nucleotides. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated. In some embodiments, the library molecules have forward sequencing primer binding site (120) and / or reverse sequencing primer binding site (130) sequences listed in Table 1, which may be truncated at the 5' and / or 3' ends, and the truncation may be 1 to 12 nucleotides. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated.
[0158] The present disclosure provides a reaction mixture comprising a plurality of any of the covalently closed circular molecules (600) described herein and at least one exonuclease enzyme, in some embodiments, the exonuclease enzyme comprises any one or any combination of two or more of Exonuclease I, Thermolabile Exonuclease I, and / or T7 Exonuclease.
[0159] A multiplex workflow is enabled by preparing sample-indexed covalently closed circular library molecules (600) using a double-stranded splint adapter carrying at least one sample index sequence. Using the sample index sequence, separate batches of sample-indexed covalently closed circular library molecules (600) can be prepared using input nucleic acids isolated from different sources. The sample-indexed covalently closed circular library molecules (600) can be pooled together to generate a multiplexed covalently closed circular library molecule (600) mixture, and the pooled covalently closed circular library molecules (600) can be amplified and / or sequenced. The sequence of the insert region, along with the sample index sequence, can be used to identify the source of the input nucleic acid. In some embodiments, any number of batches of sample-indexed covalently closed circular library molecules (600) can be pooled together, for example, 2 to 10, or 10 to 50, or 50 to 100, or 100 to 200, or more than 200 batches of sample-indexed covalently closed circular library molecules (600) can be pooled. Exemplary nucleic acid sources include naturally occurring sources, recombinant sources, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or whole organisms. Exemplary nucleic acid sources include fresh, frozen, fresh frozen, or archived (e.g., formalin-fixed, paraffin-embedded; FFPE) sources. Those skilled in the art will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.
[0160] Kit containing a double-stranded splint adapter The present disclosure provides kits for use in introducing one or more new adapter sequences into linear nucleic acid library molecules. In some embodiments, the kits can be used to circularize single-stranded nucleic acid library molecules having a sequence of interest (110) flanked on one side by a universal adapter sequence (120) for a forward sequencing primer binding site and on the other side by a universal adapter sequence (130) for a reverse sequencing primer binding site. In some embodiments, the circularized library molecules can be converted into covalently closed circular molecules, which can be subjected to a rolling circle amplification (RCA) reaction to generate nucleic acid concatemers. The concatemers can be immobilized on a support for massively parallel sequencing.
[0161] The present disclosure provides a kit including a nucleic acid double-stranded splint adaptor (200), wherein the nucleic acid double-stranded splint adaptor (200) includes (i) a first splint strand (long splint strand (300)) and (ii) a second splint strand (short splint strand (400)). The first splint strand includes a first region (320), an internal region (310), and a second region (330). The internal region (310) of the first splint strand can hybridize to the second splint strand (400) to form a double-stranded splint adaptor (200) having a double-stranded region and two adjacent single-stranded regions. The two adjacent single-stranded regions of the double-stranded splint adaptor (200) are designed to hybridize to end sequences of a linear nucleic acid library molecule (100). The terminal sequences of the linear nucleic acid library molecule can comprise first (120) and second (130) universal adapter sequences, respectively. In some embodiments, the first and second universal adapter sequences of the linear nucleic acid library molecule comprise binding sequences for forward (120) and reverse (130) sequencing primer binding sites, respectively.
[0162] At least a portion of the internal region (310) of the first splint strand can hybridize to the second splint strand (400) to form a double-stranded splint adapter (200) having a double-stranded region and two adjacent single-stranded regions. The second splint strand (400) includes at least one new adapter sequence. Formation of the library-sprint complex (500) brings the library molecule (100) into proximity with at least one new universal adapter sequence, including, for example, a universal sequence for a surface pinning primer binding site, a universal sequence for a surface capture primer binding site, a sample index sequence, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI) sequence. Exemplary double-stranded splint adapters are shown in Figures 2 and 3. The kit can include containers containing the first splint strand (300) and the second splint strand (400) in hybridized or unhybridized form. The kit may include a first container containing a first splint strand (300) and a second container containing a second splint strand (400).
[0163] In some embodiments, in the kit, the first region (320) of the first splint strand comprises a first universal adapter sequence capable of hybridizing to a first universal binding sequence at one end of a linear nucleic acid library molecule (see, e.g., Figures 2 and 3). The second region (330) of the first splint strand comprises a second universal adapter sequence capable of hybridizing to a second universal binding sequence at the other end of the linear nucleic acid library molecule (see, e.g., Figures 2 and 3). In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence comprising a universal binding sequence for a forward or reverse sequencing primer. In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a forward or reverse sequencing primer. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated, or alternatively, the 5' end of the first splint strand (300) is not phosphorylated. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group, or alternatively, the 5' end of the first splint strand (300) is a terminal 3' blocking group.
[0164] In some embodiments, in the kit, the second splint strand (400) comprises at least three subregions, including a first, second, and third subregion (see, e.g., Figures 2 and 3). The first subregion (411) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The second subregion (412) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI). The third subregion (413) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5-20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3-20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule (e.g., molecular tagging) to which the adapter is attached. In some embodiments, the unique molecular index (UMI) comprises a random sequence. The second splint strand (400) is designed to exhibit reduced or no hybridization to the insert sequence (110) of the library molecule (100).
[0165] An exemplary arrangement of sequences within the subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) includes a universal sequence for binding a surface pinning primer]-[(412) includes a sample index sequence and an optional short random sequence NNN]-[(413) includes a universal sequence for binding a surface capture primer].
[0166] An exemplary arrangement of sequences within the subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) includes a universal sequence for binding a surface pinning primer]-[(412) includes a universal sequence for binding a surface capture primer]-[(413) includes a sample index sequence and an optional short random sequence NNN].
[0167] An exemplary arrangement of sequences within the subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) includes a sample index sequence and an optional short random sequence NNN] - [(412) includes a universal sequence for binding a surface pinning primer] - [(413) includes a universal sequence for binding a surface capture primer].
[0168] In some embodiments, in the kit, the second splint strand (400) can be 20 to 100 (e.g., about 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100) nucleotides in length. In some embodiments, the second splint strand (400) can be 30 to 80 nucleotides in length, 80 (e.g., about 30, 35, 40, 45, 50, 60, 70, or 80) nucleotides in length, or 40 to 60 nucleotides in length, 60 (e.g., about 40, 45, 50, or 60) nucleotides in length. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated; alternatively, the 5' end of the second splint strand (400) is not phosphorylated. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group; alternatively, the 3' end of the second splint strand (400) comprises a terminal 3' blocking group. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position, for example, to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.
[0169] In some embodiments, in the kit, the first splint strand (300) comprises an internal region (310) comprising at least three subregions, including a fourth subregion (311), a fifth subregion (312), and a sixth subregion (313). The fourth subregion (311) can hybridize to a first subregion (411) of the second splint strand (400). The fourth subregion (311) can be fully or partially complementary to the first subregion (411) of the second splint strand (400). The fifth subregion (312) can hybridize to a second subregion (412) of the second splint strand (400). The fifth subregion (312) can be fully or partially complementary to the second subregion (412) of the second splint strand (400). The sixth subregion (313) can hybridize to the third subregion (413) of the second splint strand (400). The sixth subregion (313) can be fully or partially complementary to the third subregion (413) of the second splint strand (400). The fourth, fifth, and sixth subregions do not hybridize (or at least show very little hybridization) to the sequence of interest, surface capture primer, or surface pinning primer.
[0170] In some embodiments, in the kit, one of the subregions of the first splint strand (300) comprises an index or random sequence, such as a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). For example, subregions (311), (312), or (313) comprise a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5 to 20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3 to 20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule to which the adaptor is attached (e.g., molecular tagging). In some embodiments, the unique molecular index (UMI) comprises a random sequence.
[0171] In some embodiments, in the kit, the first splint strand subregion comprising the index or random sequence comprises any one or any combination of two or more of the following: a sample index sequence (indicated by "S"), a random sequence (indicated by "N"), at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8-oxoG), or deoxyinosine), deoxyinosine (indicated by "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer.
[0172] In some embodiments, the kit includes a double-stranded splint adapter (200) that includes a first splint strand (300) that is partially hybridized to a second splint strand (400), where the first splint strand (300) includes a universal length splint strand. In some embodiments, the universal length splint strand includes at least one subregion that is only partially hybridized to the second splint strand (400). For example, but not limited to, the universal length splint strand (300) includes a subregion that possesses at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8oxoG), or deoxyinosine), deoxyinosine (indicated by an "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), multiple C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer. An exemplary universal long splint chain (300) is shown in Figure 6.
[0173] In some embodiments, in the kit, the first splint strand (300) lacks subregion (311), (312), or (313), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has a portion of the second splint strand looped out. For example, but not limited to, the first splint strand (300) lacks subregion (312), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has subregion (412) of the second splint strand looped out (see, e.g., FIG. 7). In some embodiments, the first splint strand (300) lacking the subregion is an example of a universal length splint strand (300).
[0174] In some embodiments, in the kit, the first splint strand (300) can be 50-150 nucleotides in length, or 60-100 nucleotides in length, or 70-90 nucleotides in length. In some embodiments, the first splint strand (300) includes one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the first splint strand (300) includes one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the first splint strand (300) includes one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.
[0175] Tables 1-4 above list various embodiments of universal adaptor sequences, first splint strands (300), second splint strands (400), and immobilized surface primers within the library molecule (100) that may be included in the kits. In some embodiments, in any of the kits described herein, the sequences within the first splint strand (300) and / or second splint strand (400) comprise sequences complementary to sequences listed in Tables 2-3. In some embodiments, in any of the kits described herein, any of the sequences within the first splint strand (300) and / or second splint strand (400) listed in Tables 2-3 may be truncated at the 5' and / or 3' end, and the truncation may be 1 to 12 nucleotides. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated.
[0176] In some embodiments, the kit includes a nucleic acid double-stranded splint adaptor (200) and further includes T4 polynucleotide kinase. In some embodiments, the kit further includes a ligase enzyme, optionally the ligase enzyme includes T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase. In some embodiments, the kit further includes at least one endonuclease, including any one of Exonuclease I, Thermolabile Exonuclease I, and / or T7 Exonuclease, or any combination of two or more thereof.
[0177] In some embodiments, the kit includes at least one buffer for hybridizing a plurality of double-stranded splint adaptors (200) and a plurality of nucleic acid library molecules (100). In some embodiments, the kit includes a buffer for performing multiple enzymatic reactions in a single reaction vessel, the multiple enzymatic reactions including any combination of (i) phosphorylating the 5' ends of the first splint strand and / or the second splint strand (e.g., (300) and / or (400)), (ii) ligating nicks within the library-sprint complex (500), and / or (iii) exonuclease digestion of the first splint strand (300) from the covalently closed circular molecule (600). Alternatively, the kit includes two or more separate buffers, e.g., a first buffer can be used to perform a phosphorylation reaction, a second buffer can be used to perform a ligation reaction, and a third buffer can be used to perform an exonuclease digestion reaction.
[0178] In some embodiments, the kit includes one or more containers containing any of the double-stranded splint adapters (200) described herein, or any of the first and second splint strands (300) and (400) described herein. The kit may further include one or more containers containing T4 polynucleotide kinase, at least one ligase, and / or at least one exonuclease. The kit can include any of these components in any combination, and may be contained in a single container, or in separate containers, or any combination thereof.
[0179] The kit can include instructions for using the kit, for example, instructions for carrying out a reaction to introduce one or more new adapter sequences into a linear nucleic acid library molecule.
[0180] Methods for forming multiple library-splint complexes In some aspects, the present disclosure provides a method for forming a plurality of library-sprint complexes (500), comprising: (a) providing a plurality of double-stranded splint adapters (200), each double-stranded splint adapter (200) in the plurality of double-stranded splint adapters (200) comprising a first splint strand (300) hybridized to a second splint strand (400), the double-stranded splint adapter comprising a double-stranded region and two adjacent single-stranded regions, the first splint strand comprising a first region (320), an internal region (310), and a second region (330), the internal region (310) of the first splint strand hybridizing to the second splint strand (400). Exemplary double-stranded splint adapters (200) are shown in Figures 2 and 3. In some embodiments, a portion of the interior region (310) of the first splint strand (300) does not hybridize to a portion of the second splint strand (400).
[0181] In some embodiments, the method for forming a plurality of library-sprint complexes (500) further comprises step (b) hybridizing a plurality of double-stranded splint adapters to a plurality of single-stranded nucleic acid library molecules (100), each library molecule comprising a sequence of interest (110) flanked on one side by a universal adapter sequence (120) for a forward sequencing primer binding site and on the other side by a universal adapter sequence (130) for a reverse sequencing primer binding site (e.g., FIG. 3). The hybridizing can be performed under conditions suitable for hybridizing a first region (320) of the first splint strand to the portion (120) of the library molecule, and conditions suitable for hybridizing a second region (330) of the first splint strand to the portion (130) of the library molecule, thereby circularizing the plurality of library molecules to form a plurality of library-sprint complexes (500).
[0182] In some embodiments, in a method for forming a plurality of library-sprint complexes (500), the first region (320) of the first splint strand comprises a sequence that can hybridize to a universal adapter sequence (120) for a forward sequencing primer binding site at one end of a linear nucleic acid library molecule.
[0183] In some embodiments, in a method for forming a plurality of library-sprint complexes (500), the second region (330) of the first splint strand comprises a sequence capable of hybridizing to a universal adapter sequence (130) for a reverse sequencing primer binding site at the other end of the linear nucleic acid library molecule.
[0184] In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0185] In some embodiments, in a method for forming a plurality of library-sprint complexes (500), a first region (320) of a first splint strand hybridizes to a universal adapter sequence (120) for a forward sequencing primer binding site of a library molecule, and a second region (330) of the first splint strand hybridizes to a universal adapter sequence (130) for a reverse sequencing primer binding site of a library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500). The library-sprint complex (500) may include a first nick between the 5' end of the library molecule and the 3' end of the second splint strand (e.g., Figures 3 and 4). The library-sprint complex (500) may also include a second nick between the 5' end of the second splint strand and the 3' end of the library molecule (e.g., Figures 3 and 4). In some embodiments, the first and second nicks are enzymatically ligatable. A ligation reaction joins the sequence from the second splint strand (400) to the end of the library molecule (100).
[0186] In some embodiments, in the method for forming a plurality of library-sprint complexes (500), the first region (320) of the first splint strand can hybridize to the sense strand or the antisense strand of the double-stranded nucleic acid library molecule. In the library-sprint complexes (500), the second region (330) of the first splint strand can hybridize to the sense strand or the antisense strand of the double-stranded nucleic acid library molecule. The double-stranded nucleic acid library molecule can be denatured to generate single-stranded sense and antisense library strands.
[0187] In some embodiments, in a method for forming a plurality of library-sprint complexes (500), the second splint strand (400) does not hybridize to the sequence of interest (110) and the internal region (310) of the first splint strand does not hybridize to the sequence of interest (110).
[0188] In some embodiments, in a method for forming a plurality of library-sprint complexes (500), a first region (320) of a first splint strand does not hybridize to a sequence of interest (110), and a second region (330) of the first splint strand does not hybridize to a sequence of interest (110).
[0189] In some embodiments, in the method for forming a plurality of library-splint complexes (500), the 5' ends of the single-stranded library molecules (100) are phosphorylated, or alternatively, the 5' ends of the single-stranded library molecules (100) lack a phosphate group. In some embodiments, the 3' ends of the single-stranded library molecules comprise a terminal 3' OH group, or alternatively, the 3' ends of the single-stranded library molecules comprise a terminal 3' blocking group.
[0190] In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a universal adapter sequence for a surface capture primer binding site. In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a universal adapter sequence for a surface pinning primer binding site. In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a universal adapter sequence for a sample index sequence. In some embodiments, the single-stranded nucleic acid library molecule (100) lacks a unique molecular index (UMI) sequence.
[0191] In some embodiments, in a method for forming a plurality of library-splint complexes (500), the second splint strand (400) comprises at least three subregions, including a first, second, and third subregion (see, e.g., Figures 2 and 3). The first subregion (411) comprises a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, and / or a unique molecular index (UMI). The second subregion (412) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, and / or a unique molecular index (UMI). The third subregion (413) may comprise a universal binding sequence for the immobilized surface pinning primer, an immobilized surface capture primer, at least one sample index sequence, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index comprises 5-20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) comprises 3-20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule (e.g., molecular tagging) to which the adapter is attached. In some embodiments, the unique molecular index (UMI) comprises a random sequence. The second splint strand (400) is designed to exhibit reduced or no hybridization to the insert sequence (110) of the library molecule (100).
[0192] In some embodiments, in a method for forming a plurality of library-splint complexes (500), an exemplary arrangement within a subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) comprising a universal sequence for binding a surface pinning primer]-[(412) comprising a sample index sequence and an optional short random sequence NNN]-[(413) comprising a universal sequence for binding a surface capture primer].
[0193] In some embodiments, in a method for forming a plurality of library-splint complexes (500), an exemplary arrangement within a subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) comprising a universal sequence for binding a surface pinning primer]-[(412) comprising a universal sequence for binding a surface capture primer]-[(413) comprising a sample index sequence and an optional short random sequence NNN].
[0194] In some embodiments, in a method for forming a plurality of library-splint complexes (500), an exemplary arrangement within a subregion of the second splint strand (400) includes, in a 3' to 5' orientation, [(411) comprising a sample index sequence and an optional short random sequence NNN]-[(412) comprising a universal sequence for binding a surface pinning primer]-[(413) comprising a universal sequence for binding a surface capture primer].
[0195] In some embodiments, in the method for forming a plurality of library-splint complexes (500), the second splint strand (400) comprises an additional subregion carrying a second sample index sequence and an optional short random sequence NNN. For example, the additional subregion can be located between subregions (411) and (412) or between subregions (412) and (413).
[0196] In some embodiments, in the method for forming a plurality of library-sprint complexes (500), the second splint strand (400) can be 20 to 100 (e.g., about 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100) nucleotides in length. In some embodiments, the second splint strand (400) can be 30 to 80 nucleotides in length, 80 (e.g., about 30, 35, 40, 45, 50, 60, 70, or 80) nucleotides in length, or 40 to 60 nucleotides in length, 60 (e.g., about 40, 45, 50, or 60) nucleotides in length. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated; alternatively, the 5' end of the second splint strand (400) is not phosphorylated. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group; alternatively, the 3' end of the second splint strand (400) comprises a terminal 3' blocking group. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position, for example, to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.
[0197] In some embodiments, in the method for forming a plurality of library-splint complexes (500), the first splint strand (300) can be 50-150 nucleotides in length, or 60-100 nucleotides in length, or 70-90 nucleotides in length. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0198] In some embodiments, in a method for forming a plurality of library-sprint complexes (500), the first splint strand (300) comprises an internal region (310) comprising at least three subregions, including a fourth subregion (311), a fifth subregion (312), and a sixth subregion (313). The fourth subregion (311) can hybridize to a first subregion (411) of the second splint strand (400). The fourth subregion (311) can be fully or partially complementary to the first subregion (411) of the second splint strand (400). The fifth subregion (312) can hybridize to a second subregion (412) of the second splint strand (400). The fifth subregion (312) can be fully or partially complementary to the second subregion (412) of the second splint strand (400). The sixth subregion (313) can hybridize to the third subregion (413) of the second splint strand (400). The sixth subregion (313) can be fully or partially complementary to the third subregion (413) of the second splint strand (400). The fourth, fifth, and sixth subregions do not hybridize (or at least show very little hybridization) to the sequence of interest, surface capture primer, or surface pinning primer.
[0199] In some embodiments, in a method for forming a plurality of library-sprint complexes (500), the first splint strand (300) includes a subregion (312) that includes a sample index, a short random sequence (e.g., NNN), and / or a unique molecular index (UMI). In some embodiments, the sample index includes 5-20 bases (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. In some embodiments, the unique molecular index (UMI) includes 3-20 bases (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases) and can be used to uniquely identify the nucleic acid molecule to which the adaptor is attached (e.g., molecular tagging). In some embodiments, the unique molecular index (UMI) comprises a random sequence.
[0200] In some embodiments, in the method for forming a plurality of library-sprint complexes (500), the subregions (312) comprise any one or any combination of two or more of a sample index sequence (indicated by "S"), a random sequence (indicated by "N"), at least one nucleotide (indicated by a black rectangle) that can be converted to an abasic base (e.g., uridine, 8-oxo-7,8-dihydroguanine (e.g., 8-oxoG), or deoxyinosine), deoxyinosine (indicated by "I"), and / or a spacer. In some embodiments, the spacer comprises an 18-carbon spacer (e.g., including a hexa-ethylene glycol spacer), a plurality of C3 spacer phosphoramidites, or a spacer 9 comprising a trimethylene glycol spacer. In some embodiments, the spacer comprises a polyethylene glycol spacer, e.g., a PEG2, PEG3, or PEG4 spacer.
[0201] In some embodiments, in the method for forming a plurality of library-sprint complexes (500), the first splint strand (300) lacks subregions (311), (312), or (313), such that the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400) has a portion of the second splint strand looped out. For example, the first splint strand (300) lacks subregion (312), and in the duplex formed by hybridization between the first splint strand (300) and the second splint strand (400), subregion (412) of the second splint strand (400) loops out (see, e.g., Figure 7).
[0202] Tables 1-4 above list various embodiments of the sequences of the various molecules that form the library-sprint complex (500), including the library molecule (100), the first splint strand (300), and the second splint strand (400). In some embodiments, in any of the methods for forming a plurality of library-sprint complexes (500) described herein, the sequences in the first splint strand (300) and / or the second splint strand (400) comprise sequences that are complementary to sequences listed in Tables 2-3. In some embodiments, in any of the methods for forming a plurality of library-sprint complexes (500) described herein, the sequences in the first splint strand (300) and / or the second splint strand (400) listed in Tables 2-3 may be truncated at the 5' and / or 3' end, and the truncation may be from 1 to 12 nucleotides. In some embodiments, the truncation can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated. In some embodiments, in library molecules, the sequence of the forward sequencing primer binding site (120) and / or the sequence of the reverse sequencing primer binding site (130) are listed in Table 1 and may be truncated at the 5' end and / or the 3' end, and the truncation can be 1 to 12 nucleotides. In some embodiments, the truncation can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In certain embodiments, the 5' end is truncated. In certain embodiments, the 3' end is truncated.
[0203] In some embodiments, any of the methods for forming multiple library-splint complexes (500) described herein can further include at least one enzymatic reaction, including a phosphorylation reaction, a ligation reaction, and / or an exonuclease reaction. The enzymatic reactions can be performed sequentially or essentially simultaneously. The enzymatic reactions can be performed in a single reaction vessel. Alternatively, a first enzymatic reaction can be performed in a first reaction vessel, then transferred to a second reaction vessel, a second enzymatic reaction can be performed in the second reaction vessel, then transferred to a third reaction vessel, and a third enzymatic reaction can be performed in the third reaction vessel.
[0204] In some embodiments, any of the methods for forming a plurality of library-splint complexes (500) described herein further comprises performing separate and sequential phosphorylation and ligation reactions, wherein the separate and sequential phosphorylation and ligation reactions are performed in separate reaction vessels. In some embodiments, the method for forming a plurality of library-splint complexes (500) further comprises step (c1): contacting, in a first reaction vessel, a plurality of double-stranded splint adaptors (200) and a plurality of single-stranded nucleic acid library molecules (100) with T4 polynucleotide kinase enzyme under conditions suitable for phosphorylating the 5' ends of the plurality of double-stranded splint adaptors (200) and / or the plurality of single-stranded nucleic acid library molecules (100), and transferring the phosphorylation reaction to a second reaction vessel. In some embodiments, the method for forming a plurality of library-splint complexes (500) further comprises step (d1): contacting, in a second reaction vessel, a plurality of phosphorylated double-stranded splint adapters (200) and a plurality of phosphorylated single-stranded nucleic acid library molecules (100) with a ligase under conditions suitable for enzymatic ligation of the first and second nicks, thereby generating a plurality of covalently closed circular library molecules (600), each hybridized to the first splint strand (300). In some embodiments, the ligase enzyme comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase. An exemplary method for generating a covalently closed circular library molecule (600) is shown in FIG. 4. The enzymatic ligation reaction can link sequences from the second splint strand (400) to the end of the library molecule (100). The double-stranded adapters (200) allow new universal adapter sequences and sample index sequences to be ligated to both ends of the library molecules (100) without the need for primer extension or PCR reactions. Thus, a PCR-free workflow using double-stranded adapters (200) can be used to generate covalently circularized library molecules (600) with the adapter sequences required for downstream workflows such as amplification and sequencing.
[0205] As illustrated in Figure 4, in some embodiments, in the covalently closed circular library molecule (600), a subregion of the second splint strand (400) is covalently linked to a universal forward sequencing primer binding site (120) and a universal reverse sequencing primer binding site (130). In some embodiments, the first splint strand (300) has an extendable 3' end that can be used to initiate a primer extension reaction. In some embodiments, a rolling circle amplification reaction can be performed using the first splint strand (300) as an amplification primer to generate nucleic acid concatemer molecules that are complementary to the covalently closed circular library molecule (600).
[0206] In some embodiments, any of the methods for forming a plurality of library-sprint complexes (500) described herein further comprises performing sequential phosphorylation and ligation reactions, wherein the sequential phosphorylation and ligation reactions are performed sequentially in the same reaction vessel. In some embodiments, the method for forming a plurality of library-sprint complexes (500) further comprises step (c2): contacting, in a first reaction vessel, the plurality of double-stranded splint adaptors (200) and the plurality of single-stranded nucleic acid library molecules (100) with T4 polynucleotide kinase enzyme under conditions suitable for phosphorylating the 5' ends of the plurality of double-stranded splint adaptors (200) and the plurality of single-stranded nucleic acid library molecules (100). In some embodiments, the method for forming a plurality of library-splint complexes (500) further comprises step (d2): contacting, in the same first reaction vessel, the phosphorylated double-stranded splint adapters (200) and the phosphorylated single-stranded nucleic acid library molecules (100) with a ligase under conditions suitable for enzymatic ligation of the first and second nicks, thereby generating a plurality of covalently closed circular library molecules (600), each hybridized to the first splint strand (300). In some embodiments, the ligase enzyme comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase. An exemplary method for generating covalently closed circular library molecules (600) is shown in FIG. 4.
[0207] In some embodiments, any of the methods for forming a plurality of library-splint complexes (500) described herein further includes performing essentially simultaneous phosphorylation and ligation reactions, wherein the essentially simultaneous phosphorylation and ligation reactions are performed together in the same reaction vessel. In some embodiments, the method for forming a plurality of library-splint complexes (500) further includes step (c3): contacting, in a first reaction vessel, a plurality of double-stranded splint adapters (200) and a plurality of single-stranded nucleic acid library molecules (100) with (i) a T4 polynucleotide kinase enzyme and (ii) a ligase enzyme under conditions suitable for phosphorylating the 5' ends of the plurality of double-stranded splint adapters (200) and the plurality of single-stranded nucleic acid library molecules (100), wherein the conditions are suitable for enzymatically ligating the first and second nicks, thereby generating a plurality of covalently closed circular library molecules (600), each hybridized to a first splint strand (300). In some embodiments, the ligase enzyme comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase. An exemplary method for generating covalently closed circular library molecules (600) is shown in FIG.
[0208] In some embodiments, any of the methods for forming a plurality of library-splint complexes (500) described herein further includes an optional step of enzymatically removing a plurality of first splint strands (300) from a plurality of covalently closed circular library molecules (600), comprising contacting the plurality of covalently closed circular library molecules (600) with at least one exonuclease enzyme to remove the plurality of first splint strands (300) and retain the plurality of covalently closed circular library molecules (600). In some embodiments, the exonuclease reaction can be performed in the same reaction buffer used to perform the phosphorylation and / or ligation reactions, or in a different reaction buffer. In some embodiments, the phosphorylation reaction can be performed in a first reaction vessel (c1) and the ligation reaction can be performed in a second reaction vessel (d1), after which the exonuclease reaction can be performed in a third reaction vessel. In some embodiments, a phosphorylation reaction may be performed in a first reaction vessel (c2), a sequential ligation reaction may be performed in the first reaction vessel (d2), and then an exonuclease reaction may be performed in the first reaction vessel. In some embodiments, an exonuclease reaction may be performed in the first reaction vessel (c3), after essentially simultaneous phosphorylation and ligation reactions are performed in the first reaction vessel. In some embodiments, the at least one exonuclease enzyme comprises any combination of two or more of exonuclease I, thermolabile exonuclease I, and / or T7 exonuclease.
[0209] A multiplex workflow is enabled by preparing sample-indexed covalently closed circular library molecules (600) using a double-stranded splint adapter carrying at least one sample index sequence. Using the sample index sequence, separate batches of sample-indexed covalently closed circular library molecules (600) can be prepared using input nucleic acids isolated from different sources. The sample-indexed covalently closed circular library molecules (600) can be pooled together to generate a multiplexed covalently closed circular library molecule (600) mixture, and the pooled covalently closed circular library molecules (600) can be amplified and / or sequenced. The sequence of the insert region, along with the sample index sequence, can be used to identify the source of the input nucleic acid. In some embodiments, any number of batches of sample-indexed covalently closed circular library molecules (600) can be pooled together, for example, 2 to 10, or 10 to 50, or 50 to 100, or 100 to 200, or more than 200 batches of sample-indexed covalently closed circular library molecules (600) can be pooled. Exemplary nucleic acid sources include naturally occurring sources, recombinant sources, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or whole organisms. Exemplary nucleic acid sources include fresh, frozen, fresh frozen, or archived (e.g., formalin-fixed, paraffin-embedded; FFPE) sources. Those skilled in the art will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.
[0210] Methods for Rolling Circle Amplification The present disclosure provides methods for performing a rolling circle amplification reaction on a covalently closed circular library molecule (600). The rolling circle amplification reaction can be performed after the phosphorylation and ligation reaction or after the ligation reaction. In some embodiments, the rolling circle amplification reaction can be performed on a covalently closed circular library molecule (600) that is no longer hybridized to the first splint strand (300) after the exonuclease reaction. In some embodiments, the rolling circle amplification reaction can be performed on a covalently closed circular library molecule (600) that is hybridized to the first splint strand (300). In some embodiments, the covalently closed circular library molecule (600) can be dispensed onto a support, and then a rolling circle amplification reaction can be performed on it. In some embodiments, the covalently closed circular library molecule (600) can be subjected to a rolling circle amplification reaction in solution, and then dispensed onto a support. In some embodiments, the rolling circle amplification reaction may use the retained first splint strand (300) as an amplification primer (e.g., Figure 4), or the first splint strand (300) may be removed (e.g., via exonuclease digestion) and replaced with a soluble amplification primer (e.g., Figure 5).
[0211] On-support rolling circle amplification In some embodiments, a method for performing a rolling circle amplification reaction on a plurality of covalently closed circular library molecules (600) lacking a hybridized first splint strand (300), wherein each covalently closed circular library molecule (600) in the plurality of covalently closed circular library molecules comprises a universal binding sequence for a surface capture primer, the method comprising the step (a): distributing the plurality of covalently closed circular library molecules (600) onto a support having a plurality of surface capture primers immobilized thereon under conditions suitable for hybridizing each of the covalently closed circular library molecules (600) to each of the immobilized surface capture primers, thereby immobilizing the plurality of covalently closed circular library molecules (600) to the support.
[0212] In some embodiments, the plurality of surface capture primers comprises any one of the sequences listed in Table 4, or a complementary sequence thereof.
[0213] Each surface capture primer can hybridize to a covalently closed circular library molecule (600) that has a universal binding sequence for the surface capture primer.
[0214] In some embodiments, the method for performing a rolling circle amplification reaction further comprises step (b): contacting a plurality of immobilized covalently closed circular library molecules (600) with a plurality of strand-displacing polymerases and a plurality of nucleotides under conditions suitable for performing a rolling circle amplification reaction on a support using a plurality of surface capture primers as immobilized amplification primers and a plurality of covalently closed circular library molecules (600) as template molecules, thereby generating a plurality of nucleic acid concatemer molecules immobilized to the surface capture primers. In some embodiments, the plurality of nucleotides comprises any combination of two or more of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, each immobilized concatemer is covalently linked to an individual surface capture primer. In some embodiments, each covalently closed circular library molecule (600) in the plurality of covalently closed circular library molecules comprises a universal binding sequence for a surface capture primer and a surface pinning primer, such that a rolling circle amplification reaction generates concatemer molecules having multiple tandem copies of the sequence carried by the covalently closed circular library molecule (600) comprising the universal binding sequence for the surface capture primer and the surface pinning primer. In some embodiments, the support further comprises a plurality of surface pinning primers. In some embodiments, the immobilized surface pinning primer functions to pin at least one portion of the concatemer molecule to the support. In some embodiments, the immobilized surface pinning primer has a non-extendable 3' end and cannot be used for amplification. In some embodiments, the plurality of surface pinning primers comprises any one of the sequences listed in Table 4, or a complementary sequence thereof. In some embodiments, a sequencing reaction may be performed on the immobilized concatemers.
[0215] Each surface pinning primer can hybridize to a portion of the concatemer molecule that has a universal binding sequence for the surface pinning primer. In some embodiments, the immobilized surface pinning primer functions to pin at least one portion of the concatemer molecule to the support. In some embodiments, the immobilized surface pinning primer has a non-extendable 3' end and cannot be used for amplification. In some embodiments, a sequencing reaction can be performed on the immobilized concatemer.
[0216] In some embodiments, in a method for performing a rolling circle amplification reaction, a plurality of covalently closed circular library molecules (600) can be distributed on a support coated with one or more compounds that generate a passivated layer on the support (e.g., FIG. 21). In some embodiments, the passivated layer forms a porous or semi-porous layer. In some embodiments, one or more types of surface primers, concatemer template molecules, and / or polymerase can be attached to the passivated layer for immobilization to the support. In some embodiments, the support comprises a low non-specific binding surface, which allows for improved nucleic acid hybridization and amplification performance on the support. Generally, the support can comprise one or more layers of covalently or non-covalently bound low-binding chemically modified layers, e.g., silane layers, polymer films, and one or more covalently or non-covalently bound oligonucleotides that can be used to immobilize a plurality of nucleic acid concatemer molecules to the support. In some embodiments, the support may comprise a functionalized polymer coating layer covalently bonded to at least a portion of the support via chemical groups on the support, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some embodiments, the functionalized polymer coating comprises poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide (PAZAM). In some embodiments, the support comprises a surface coating having at least one hydrophilic polymer coating layer and at least one layer of a plurality of oligonucleotides. The hydrophilic polymer coating layer may comprise polyethylene glycol (PEG). The hydrophilic polymer coating layer may comprise a branched PEG having at least four branches. In some embodiments, the low nonspecific binding coating has a degree of hydrophilicity that can be measured as a water contact angle, where the water contact angle is 45 degrees or less. In some embodiments, the density of the covalently closed circular library molecules (600) immobilized to the support or to the coating on the support is less than 1 mm 2 Approximately 10 per 2~10 6 , (e.g., 10 2 , 10 3 , 10 4 , 10 5 , or 10 6 In some embodiments, the covalently closed circular library molecule (600) immobilized on a support or on a coating on a support is 1 mm 2 Approximately 10 per 6 ~10 9 (e.g., about 10 6 , 10 7 , 10 8 , or 10 9 In some embodiments, the density of the covalently closed circular library molecules (600) immobilized on the support or on the coating on the support is 1 mm 2 Approximately 10 per 9 ~10 12 (e.g., about 10 9 , 10 10 , 10 11 , or 10 12 ). In some embodiments, the plurality of covalently closed circular library molecules (600) are immobilized to a support or a coating on a support at predetermined sites on the support (or a coating on the support). In some embodiments, the plurality of covalently closed circular library molecules (600) are immobilized to a support or a coating on a support at random sites on the support (or a coating on the support).
[0217] In-solution rolling circle amplification using soluble amplification primers In some embodiments, a method for performing a rolling circle amplification reaction on a plurality of covalently closed circular library molecules (600) lacking a hybridized first splint strand (300), wherein each covalently closed circular library molecule (600) in the plurality of covalently closed circular library molecules comprises a universal binding sequence for a forward amplification primer and a universal binding sequence for a surface capture primer, the method comprising: (a) hybridizing the plurality of covalently closed circular library molecules and a plurality of soluble forward amplification primers in solution; and (b) performing a first rolling circle amplification reaction by contacting the plurality of covalently closed circular library molecules (600) with a plurality of strand displacement polymerases and a plurality of nucleotides using the plurality of forward amplification primers and the plurality of covalently closed circular library molecules (600) as template molecules under conditions suitable for performing a rolling circle amplification reaction in solution, thereby generating a plurality of nucleic acid concatemer molecules having portions still hybridized to the covalently closed circular library molecules (600). In some embodiments, the method for performing a rolling circle amplification reaction further comprises step (c): distributing the plurality of concatemer molecules onto a support having a plurality of surface capture primers immobilized thereon under conditions suitable for hybridizing at least a portion of the concatemers to the plurality of immobilized surface capture primers, thereby immobilizing the plurality of concatemer molecules. The plurality of immobilized concatemer molecules remain hybridized to their covalently closed circular library molecules (600). In some embodiments, the method for performing a rolling circle amplification reaction further comprises step (d): contacting the immobilized plurality of concatemer molecules with a plurality of strand-displacing polymerases and a plurality of nucleotides under conditions suitable for performing a second rolling circle amplification reaction on the support using the plurality of covalently closed circular library molecules (600) as template molecules, thereby extending the plurality of immobilized nucleic acid concatemer molecules.In some embodiments, the first and / or second rolling circle amplification reaction can be performed with a plurality of nucleotides comprising any combination of two or more of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, each immobilized concatemer is hybridized to an individual surface capture primer. In some embodiments, each covalently closed circular library molecule (600) in the plurality of covalently closed circular library molecules comprises universal binding sequences for the surface capture primer and the surface pinning primer, such that a rolling circle amplification reaction in solution generates concatemer molecules having multiple tandem copies of the sequence carried by the covalently closed circular library molecule (600) comprising the universal binding sequences for the surface capture primer and the surface pinning primer. In some embodiments, the support further comprises a plurality of surface pinning primers. In some embodiments, the immobilized surface pinning primer functions to pin at least one portion of the concatemer molecule to the support. In some embodiments, the immobilized surface pinning primer has a non-extendable 3' end and cannot be used for amplification. In some embodiments, a sequencing reaction may be performed on the immobilized concatemers. In some embodiments, the soluble forward amplification primer has the same sequence as the surface capture primer.
[0218] In some embodiments, the plurality of surface capture primers comprises any one of the sequences listed in Table 4, or complementary sequences thereof. In some embodiments, the plurality of surface pinning primers comprises any one of the sequences listed in Table 4, or complementary sequences thereof.
[0219] In some embodiments, in a method for performing a rolling circle amplification reaction, the plurality of surface capture primers immobilized on a support comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO: 34) (or a complementary sequence thereof).
[0220] In some embodiments, in a method for performing a rolling circle amplification reaction, the plurality of surface capture primers immobilized on a support comprises the sequence 5'-ATTACATGGATCAGGTGAGGCT-3' (SEQ ID NO: 35) (or a complementary sequence thereof).
[0221] In some embodiments, the immobilized surface pinning primer functions to pin at least one portion of the concatemer molecule to the support. In some embodiments, the immobilized surface pinning primer has a non-extendable 3' end and cannot be used for amplification. In some embodiments, a sequencing reaction can be performed on the immobilized concatemer.
[0222] In some embodiments, the plurality of concatemer molecules of step (c) can be distributed onto a support coated with one or more compounds that generate a passivated layer on the support, e.g., glass (see, e.g., Figure 21). In alternative embodiments, the support can be made of any material, such as glass, plastic, or polymeric material. In some embodiments, the passivated layer forms a porous or semi-porous layer. In some embodiments, one or more types of surface primers, concatemer template molecules, and / or polymerases can be bound to the passivated layer for immobilization to the support. In some embodiments, the support comprises a low non-specific binding surface, which allows for improved nucleic acid hybridization and amplification performance on the support. Generally, the support can comprise one or more layers of covalently or non-covalently bound low-binding chemically modified layers, e.g., silane layers, polymer films, and one or more covalently or non-covalently bound oligonucleotides that can be used to immobilize the plurality of nucleic acid concatemer molecules to the support. In some embodiments, the support may include a functionalized polymer coating layer covalently bonded to at least a portion of the support via chemical groups on the support, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some embodiments, the functionalized polymer coating includes poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide (PAZAM). In some embodiments, the support includes a surface coating having at least one hydrophilic polymer coating layer and at least one layer of a plurality of oligonucleotides. The hydrophilic polymer coating layer may include polyethylene glycol (PEG). The hydrophilic polymer coating layer may include a branched PEG having at least four branches (e.g., 4, 5, 6, 7, 8, 9, 10, or more branches).In some embodiments, the low nonspecific binding coating has a degree of hydrophilicity that can be measured as a water contact angle of 45 degrees or less (e.g., 5 degrees or less, 10 degrees or less, 15 degrees or less, 20 degrees or less, 25 degrees or less, 30 degrees or less, 35 degrees or less, 40 degrees or less, or 45 degrees or less). In some embodiments, the density of concatemeric molecules immobilized to the support or to the coating on the support is 1 mm. 2 Approximately 10 per 2 ~10 6 (e.g., about 10 2 , 10 3 , 10 4 , 10 5 , or 10 6 In some embodiments, the density of the concatemeric molecules immobilized on the support or on the coating on the support is 1 mm 2 Approximately 10 per 6 ~10 9 (e.g., about 10 6 , 10 7 , 10 8 , or 10 9 In some embodiments, the density of the concatemeric molecules immobilized on the support or on the coating on the support is 1 mm 2 Approximately 10 per 9 ~10 12 (e.g., about 10 9 , 10 10 , 10 11 , or 10 12 In some embodiments, the plurality of concatemeric molecules are immobilized to the support or to a coating on the support at predetermined sites on the support (or coating on the support). In some embodiments, the plurality of concatemeric molecules are immobilized to a coating on the support at random sites on the support (or coating on the support).
[0223] In-solution rolling circle amplification using the first splint strand In some embodiments, a method for performing a rolling circle amplification reaction on a plurality of covalently closed circular library molecules hybridized to a first splint strand (300), wherein each covalently closed circular library molecule (600) in the plurality of covalently closed circular library molecules comprises a universal binding sequence for a surface capture primer, the method comprising: (a) contacting the plurality of covalently closed circular library molecules (600) hybridized to the first splint strand (300) with a plurality of strand displacement polymerases and a plurality of nucleotides in solution under conditions suitable for performing a first rolling circle amplification reaction using the first splint strand (300) as an amplification primer, thereby generating a plurality of concatemeric molecules still hybridized to the covalently closed circular library molecules (600) (see, e.g., FIG. 4 ).
[0224] In some embodiments, the method for performing a rolling circle amplification reaction further comprises step (b): distributing the plurality of concatemer molecules hybridized to their covalently closed circular library molecules (600) onto a support having a plurality of surface capture primers immobilized thereon under conditions suitable for hybridizing at least a portion of the concatemers to the plurality of immobilized surface capture primers, thereby immobilizing the plurality of concatemer molecules, the plurality of immobilized concatemer molecules still hybridized to their covalently closed circular library molecules (600).
[0225] In some embodiments, the method for performing a rolling circle amplification reaction further comprises step (c): contacting the plurality of immobilized concatemer molecules with a plurality of strand-displacing polymerases and a plurality of nucleotides under conditions suitable for performing a second rolling circle amplification reaction on the support using the plurality of covalently closed circular library molecules (600) as template molecules, thereby extending the plurality of immobilized nucleic acid concatemer molecules.
[0226] In some embodiments, the first and / or second rolling circle amplification reaction can be performed with a plurality of nucleotides comprising any combination of two or more of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, each immobilized concatemer is hybridized to an individual surface capture primer. In some embodiments, each covalently closed circular library molecule (600) in the plurality of covalently closed circular library molecules comprises universal binding sequences for the surface capture primer and the surface pinning primer, such that a rolling circle amplification reaction in solution generates concatemer molecules having multiple tandem copies of the sequence carried by the covalently closed circular library molecule (600) comprising the universal binding sequences for the surface capture primer and the surface pinning primer. In some embodiments, the support further comprises a plurality of surface pinning primers. In some embodiments, the immobilized surface pinning primer functions to pin at least one portion of the concatemer molecule to the support. In some embodiments, the immobilized surface pinning primer has a non-extendable 3' end and cannot be used for amplification. In some embodiments, a sequencing reaction may be performed on the immobilized concatemers.
[0227] In some embodiments, the plurality of surface capture primers comprises any one of the sequences listed in Table 4, or complementary sequences thereof. In some embodiments, the plurality of surface pinning primers comprises any one of the sequences listed in Table 4, or complementary sequences thereof.
[0228] In some embodiments, the immobilized surface pinning primer functions to pin at least one portion of the concatemer molecule to the support. In some embodiments, the immobilized surface pinning primer has a non-extendable 3' end and cannot be used for amplification. In some embodiments, the immobilized concatemer may be subjected to a sequencing reaction.
[0229] In some embodiments, the plurality of concatemer molecules of step (b) can be distributed on a support coated with one or more compounds that generate a passivated layer on the support (e.g., Figure 21). In some embodiments, the passivated layer forms a porous or semi-porous layer. In some embodiments, surface primers, concatemer template molecules, and / or polymerase can be bound to the passivated layer for immobilization to the support. In some embodiments, the support comprises a low-nonspecific binding surface that allows for improved nucleic acid hybridization and amplification performance on the support. In some embodiments, the support can comprise one or more layers of covalently or non-covalently bound low-binding chemically modified layers, e.g., silane layers, polymer films, and one or more covalently or non-covalently bound oligonucleotides that can be used to immobilize the plurality of nucleic acid concatemer molecules to the support. In some embodiments, the support can comprise, at least in part, a functionalized polymer coating layer covalently bound via chemical groups on the support, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and functionalized polymer coating. In some embodiments, the functionalized polymer coating comprises poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide (PAZAM). In some embodiments, the support comprises a surface coating having at least one hydrophilic polymer coating layer and at least one layer of a plurality of oligonucleotides. The hydrophilic polymer coating layer can comprise polyethylene glycol (PEG). The hydrophilic polymer coating layer can comprise a branched PEG having at least four branches (e.g., 4, 5, 6, 7, 8, 9, 10 or more branches). In some embodiments, the low nonspecific binding coating has a degree of hydrophilicity that can be measured as a water contact angle of 45 degrees or less (e.g., 5 degrees or less, 10 degrees or less, 15 degrees or less, 20 degrees or less, 25 degrees or less, 30 degrees or less, 35 degrees or less, 40 degrees or less, or 45 degrees or less).In some embodiments, the density of the concatemeric molecules immobilized to the support or to the coating on the support is 1 mm. 2 Approximately 10 per 2 ~10 6 (e.g., about 10 2 , 10 3 , 10 4 , 10 5 , or 10 6 In some embodiments, the density of the concatemeric molecules immobilized on the support or on the coating on the support is 1 mm 2 Approximately 10 per 6 ~10 9 (e.g., about 10 6 , 10 7 , 10 8 , or 10 9 In some embodiments, the density of the concatemeric molecules immobilized on the support or on the coating on the support is 1 mm 2 Approximately 10 per 9 ~10 12 (e.g., about 10 9 , 10 10 , 10 11 , or 10 12 In some embodiments, the plurality of concatemeric molecules are immobilized to the support or to a coating on the support at predetermined sites on the support (or coating on the support). In some embodiments, the plurality of concatemeric molecules are immobilized to a coating on the support at random sites on the support (or coating on the support).
[0230] Compaction Oligonucleotides In some aspects, the present disclosure provides compositions and methods for performing rolling circle amplification with or without multiple compaction oligonucleotides. The compaction oligonucleotides are single-stranded and can include a 5' region, an optional internal region, and a 3' region. The 5' and 3' regions of the compaction oligonucleotides can hybridize to concatemers, pulling the distal portions of the concatemers together and causing compaction of the concatemers to form nanostructures. For example, but not limited to, the 5' region of the compaction oligonucleotide can be designed to hybridize to a first portion of the concatemer molecule (e.g., a first universal adapter sequence), and the 3' region of the compaction oligonucleotide can be designed to hybridize to a second portion of the concatemer molecule (e.g., a second universal adapter sequence). Including compaction oligonucleotides during rolling circle amplification can promote the formation of nanostructures with a more compact size and shape compared to concatemers generated in the absence of the compaction oligonucleotides. The compact and stable characteristics of nucleic acid nanostructures improve sequencing accuracy, for example, by increasing signal intensity, and retain their shape and size through multiple sequencing cycles. In some embodiments, the 3' end of the compaction oligonucleotide is a designed blocked primer extension.
[0231] In some embodiments, the compaction oligonucleotide is a single-stranded oligonucleotide comprising DNA, RNA, or a combination of DNA and RNA. The compaction oligonucleotide can be any length, including 20 to 150 nucleotides, e.g., 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 nucleotides, or any range therebetween. In some embodiments, the compaction oligonucleotide can be 30 to 100 nucleotides in length. In some embodiments, the compaction oligonucleotide can be 40 to 80 nucleotides in length.
[0232] In some embodiments, the compaction oligonucleotide comprises a 5' region and a 3' region, and optionally an intervening region between the 5' and 3' regions. The intervening region can be any length, e.g., about 2-20 nucleotides in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, or any range therebetween. In some embodiments, the intervening region comprises a homopolymer having consecutive identical bases (e.g., AAA, GGG, CCC, TTT, or UUU). In some embodiments, the intervening region comprises a non-homopolymer sequence.
[0233] In some embodiments, the 5' region of the compaction oligonucleotide may be fully or partially complementary along its length to the first portion of the concatemer molecule. For example, but not limited to, the 5' region of the compaction oligonucleotide may comprise a sequence capable of hybridizing to a universal adapter sequence listed in Tables 1-3, including a surface capture primer binding site, a surface pinning primer binding site, a forward sequencing primer binding site, or a reverse sequencing primer binding site.
[0234] In some embodiments, the 3' region of the compaction oligonucleotide may be fully or partially complementary along its length to the second portion of the concatemer molecule. For example, but not limited to, the 3' region of the compaction oligonucleotide may comprise a sequence capable of hybridizing to a universal adapter sequence listed in Tables 1-3, including a surface capture primer binding site, a surface pinning primer binding site, a forward sequencing primer binding site, or a reverse sequencing primer binding site.
[0235] The 5' region of the compaction oligonucleotide can hybridize to a first universal sequence portion of the concatemer molecule. The 3' region of the compaction oligonucleotide can hybridize to a second universal sequence portion of the concatemer molecule. The 5' and 3' regions of the compaction oligonucleotide can hybridize to the concatemer and pull the distal portions of the concatemer together, causing the concatemer to compact and form a nanostructure.
[0236] In some embodiments, the 5' region of the compaction oligonucleotide may have the same sequence as the 3' region. In some embodiments, the 5' region of the compaction oligonucleotide may have a sequence that is different from the 3' region. In some embodiments, the 3' region of the compaction oligonucleotide may have a sequence that is the reverse of the 5' region.
[0237] In some embodiments, the compaction oligonucleotide contains one or more modified bases or linkages at its 5' or 3' end to confer specific functionality. In some embodiments, the compaction oligonucleotide contains at least one phosphorothioate linkage at its 5' and / or 3' end, for example, to confer exonuclease resistance. In some embodiments, at least one nucleotide at or near the 3' end contains a 2' fluoro base, which confers exonuclease resistance. In some embodiments, the 3' end of the compaction oligonucleotide contains at least one 2'-O-methyl RNA base, which blocks polymerase-catalyzed elongation. In some embodiments, the compaction oligonucleotide contains a 3' inverted dT at its 3' end, which blocks polymerase-catalyzed elongation. In some embodiments, the compaction oligonucleotide contains a 3' phosphorylation, which blocks polymerase-catalyzed elongation. In some embodiments, the internal region of the compaction oligonucleotide comprises at least one locked nucleic acid (LNA), which increases the thermal stability of the duplex formed by hybridizing the compaction oligonucleotide to the concatemer molecule.
[0238] The compaction oligonucleotide may contain at least one region having consecutive guanines. For example, the compaction oligonucleotide may contain at least one region having 2, 3, 4, 5, 6, or more consecutive guanines. In some embodiments, the compaction oligonucleotide contains four consecutive guanines that can form a G-quadruplex structure. The G-quadruplex structure may be stabilized, for example, through Hoogsteen hydrogen bonding. The G-quadruplex structure may be stabilized by a central cation, for example, potassium, sodium, lithium, rubidium, or cesium.
[0239] The rolling circle amplification reaction can be performed in the presence of multiple compaction oligonucleotides having at least four consecutive guanines. The resulting concatemers contain repeated copies of the universal binding sequence for the compaction oligonucleotides. At least one compaction oligonucleotide can form a G-quadruplex and hybridize to the universal binding sequence for the compaction oligonucleotide, and the resulting concatemers can fold and form an intramolecular G-quadruplex structure. The concatemers can self-collapse to form compact nanostructures. The formation of G-quadruplexes and G-quadruplexes in the nanostructures increases the stability of the nanostructures, allowing them to maintain their compact size and shape, which can withstand repeated changes in pH, temperature, and / or reagent flow.
[0240] Methods for sequencing The present disclosure provides a method for sequencing any of the immobilized concatemer molecules described herein. Any of the methods for performing a rolling circle amplification reaction described herein can be used to generate a plurality of concatemer molecules immobilized on a support, and a sequencing reaction can be performed on the immobilized concatemers. In some embodiments, the sequencing reaction uses detectably labeled nucleotide analogs. In some embodiments, the sequencing reaction uses a two-step sequencing reaction, which includes binding to a detectably labeled multivalent molecule and incorporating a nucleotide analog. The terms concatemer molecule and template molecule are used interchangeably herein.
[0241] In some embodiments, any of the rolling circle amplification reactions described herein (e.g., RCA performed on a support or in solution) can be used to generate immobilized concatemers, each containing a sequence of interest and tandem repeat units of any adapter molecules present in the covalently closed circular library molecule (600).
[0242] In some embodiments, an exemplary tandem repeat unit comprises (i) a universal adapter sequence having a binding sequence for a surface capture primer (subregion 413), (ii) a universal adapter sequence having a binding sequence for a reverse sequencing primer (130), (iii) a sequence of interest (110), (iv) a universal adapter sequence having a binding sequence for a forward sequencing primer (120), (v) a universal adapter sequence having a binding sequence for a surface pinning primer (subregion 411), and (vi) a sample index, a short random sequence (NNN), and / or a unique molecular index (UMI) (subregion 412).
[0243] In some embodiments, an exemplary tandem repeat unit comprises: (i) a universal adapter sequence having a binding sequence for a surface capture primer (subregion 412); (ii) a sample index, a short random sequence (NNN), and / or a unique molecular index (UMI) (subregion 413); (iii) a universal adapter sequence having a binding sequence for a reverse sequencing primer (130); (iv) a sequence of interest (110); (v) a universal adapter sequence having a binding sequence for a forward sequencing primer (120); and (vi) a universal adapter sequence having a binding sequence for a surface pinning primer (subregion 411).
[0244] In some embodiments, an exemplary tandem repeat unit comprises (i) a universal adapter sequence having a binding sequence for a surface capture primer (subregion 413), (ii) a universal adapter sequence having a binding sequence for a reverse sequencing primer (150), (iii) a sequence of interest (110), (iv) a universal adapter sequence having a binding sequence for a forward sequencing primer (140), (v) a sample index, a short random sequence (NNN) and / or a unique molecular index (UMI) (subregion 412), and (vi) a universal adapter sequence having a binding sequence for a surface pinning primer (subregion 411).
[0245] Immobilized concatemers can self-collapse into compact nucleic acid nanoballs. Including one or more compaction oligonucleotides in an RCA reaction can further compact the size and / or shape of the nanoballs. Increasing the number of tandem repeat units in a given concatemer increases the number of sites along the concatemer for hybridizing to multiple sequencing primers (e.g., sequencing primers with universal sequences) that serve as multiple initiation sites for polymerase-catalyzed sequencing reactions. When sequencing reactions use detectably labeled nucleotides and / or detectably labeled multivalent molecules (e.g., those with nucleotide units), signals emitted by nucleotides or nucleotide units participating in parallel sequencing reactions along the concatemer can result in increased signal intensity for each concatemer. Multiple portions of a given concatemer can be sequenced simultaneously. Furthermore, multiple binding complexes can be formed along a particular concatemeric molecule, each containing a sequencing polymerase bound to a multivalent molecule, and the multiple binding complexes remain stable without dissociation, resulting in increased duration, which increases signal intensity and reduces imaging time.
[0246] Sequencing methods using engineered polymerases In some aspects, the present disclosure provides concatemeric template molecules that can be sequenced using any nucleic acid sequencing method using labeled or unlabeled chain-terminating nucleotides, where the chain-terminating nucleotides contain a 3'-O-azido group (or a 3'-O-methyl azido group) or any other type of bulky blocking group at the 3' position of the sugar. In some embodiments, the concatemeric template molecules can be sequenced using a sequencing-by-avidity method (SBA) using labeled multivalent molecules and unlabeled chain-terminating nucleotides. In some embodiments, the concatemeric template molecules can be sequenced using a sequencing-by-synthesis method (SBS) using labeled chain-terminating nucleotides. In some embodiments, the concatemeric template molecules can be sequenced using a sequencing-by-binding method (SBB) using unlabeled chain-terminating nucleotides. In some embodiments, the concatemeric template molecules can be sequenced using phosphate-chain-labeled nucleotides.
[0247] Methods for sequencing using nucleotide analogs In some aspects, the present disclosure provides a method for sequencing, comprising: step (a): contacting a sequencing polymerase with (i) nucleic acid concatemer molecules and (ii) nucleic acid primers, wherein the contacting is performed under conditions suitable for binding of the sequencing polymerase to the nucleic acid concatemer molecules hybridized to the nucleic acid primers, such that the nucleic acid concatemer molecules hybridized to the nucleic acid primers form a nucleic acid duplex. In some embodiments, the sequencing polymerase comprises a recombinant mutant sequencing polymerase. In some embodiments, the primers comprise a 3' extendable end.
[0248] In some embodiments, the method for sequencing further comprises step (b): contacting a sequencing polymerase with a plurality of nucleotides under conditions suitable for binding at least one nucleotide to the sequencing polymerase bound to the nucleic acid duplex and suitable for polymerase-catalyzed nucleotide incorporation. In some embodiments, the sequencing polymerase is contacted with the plurality of nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises at least one nucleotide analog having a chain-terminating moiety at the sugar 2' or 3' position. In some embodiments, the plurality of nucleotides comprises at least one nucleotide lacking a chain-terminating moiety.
[0249] In some embodiments, the method for sequencing further comprises step (c): incorporating at least one nucleotide into the 3' end of the extendable primer under conditions suitable for incorporation of the at least one nucleotide. In some embodiments, the conditions suitable for nucleotide binding to the polymerase and nucleotide incorporation may be the same or different. In some embodiments, the conditions suitable for nucleotide incorporation comprise including at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, at least one nucleotide binds to the sequencing polymerase and is incorporated into the 3' end of the extendable primer. In some embodiments, incorporating a nucleotide into the 3' end of the primer in step (c) comprises a primer extension reaction.
[0250] In some embodiments, the method for sequencing further comprises step (d): repeating the incorporation of at least one nucleotide into the 3' end of the extendable primer of step (c) at least once. In some embodiments, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base by a linker, the linker being cleavable / removable from the base. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to enable detection and identification of the nucleotide base. In some embodiments, the method further comprises detecting at least one incorporated nucleotide in steps (c) and / or (d). In some embodiments, the method further comprises identifying at least one incorporated nucleotide in steps (c) and / or (d). In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined by detecting and identifying the nucleotide bound to the sequencing polymerase, thereby determining the sequence of the concatemer molecule. In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined by detecting and identifying the nucleotide incorporated within the 3' end of the primer, thereby determining the sequence of the concatemer molecule.
[0251] In some embodiments, the method for sequencing includes a plurality of sequencing polymerases bound to a nucleic acid duplex, the plurality of sequencing polymerases having at least a first and a second complexed polymerase, wherein (a) the first complexed polymerase comprises a first sequencing polymerase bound to a first nucleic acid duplex comprising a first nucleic acid template sequence hybridized to a first nucleic acid primer; (b) the second complexed polymerase comprises a second sequencing polymerase bound to a second nucleic acid duplex comprising a second nucleic acid template sequence hybridized to a second nucleic acid primer; (c) the first and second nucleic acid template sequences comprise the same or different sequences; (d) the first and second nucleic acid concatemers are clonally amplified; (e) the first and second primers comprise an extendable 3' end or a non-extendable 3' end; and (f) the plurality of complexed polymerases is immobilized on a support. In some embodiments, the density of the plurality of complexed polymerases is greater than 1 mm 2 Approximately 10 per 2 ~10 15 (e.g., 10 2 ~10 15 or more, e.g., 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 ) is a conjugated polymerase immobilized on a support.
[0252] A two-step method for nucleic acid sequencing In some aspects, the present disclosure provides a two-step method for sequencing nucleic acid molecules. In some embodiments, the first step generally comprises binding a multivalent molecule to a complexed polymerase to form a multivalent complexed polymerase, and detecting the multivalent complexed polymerase.
[0253] In some embodiments, the first stage comprises step (a): contacting a plurality of first sequencing polymerases with (i) a plurality of nucleic acid concatemer molecules and (ii) a plurality of nucleic acid primers, wherein the contacting is performed under conditions suitable for binding the plurality of first sequencing polymerases to the plurality of nucleic acid concatemer molecules and the plurality of nucleic acid primers, thereby forming a plurality of first complexed polymerases, each of the plurality of first complexed polymerases comprising a first sequencing polymerase bound to a nucleic acid duplex, and the nucleic acid duplex comprising the nucleic acid concatemer molecules hybridized to the nucleic acid primers. In some embodiments, the first polymerase comprises a recombinant mutant sequencing polymerase.
[0254] In some embodiments, in the method for sequencing concatemer molecules, the primer comprises a 3' extendable end or a 3' non-extendable end. In some embodiments, the plurality of nucleic acid concatemer molecules comprises amplified template molecules (e.g., clonally amplified template molecules). In some embodiments, the plurality of nucleic acid concatemer molecules comprises one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules comprises two or more tandem copies (e.g., concatemers) of a target sequence of interest. In some embodiments, the nucleic acid concatemer molecules in the plurality of nucleic acid concatemer molecules comprise the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid concatemer molecules and / or the plurality of nucleic acid primers are in solution or immobilized on a support. In some embodiments, when the plurality of nucleic acid concatemer molecules and / or the plurality of nucleic acid primers are immobilized on a support, binding with a first sequencing polymerase generates a plurality of immobilized first complexed polymerases. In some embodiments, the plurality of nucleic acid concatemer molecules and / or nucleic acid primers are 10 2 ~10 15In some embodiments, the binding of the plurality of concatemer molecules and nucleic acid primers to the plurality of first sequencing polymerases is performed at 10 different sites on the support. 2 ~10 15 different sites (e.g., 10 2 ~10 15 More than 10 sites, e.g., 10 2 Pieces, 10 3 Pieces, 10 4 Pieces, 10 5 Pieces, 10 6 Pieces, 10 7 Pieces, 10 8 Pieces, 10 9 Pieces, 10 10 Pieces, 10 11 Pieces, 10 12 Pieces, 10 13 Pieces, 10 14 Pieces, 10 15 In some embodiments, the plurality of immobilized first complexed polymerases on the support are immobilized at predetermined or random sites on the support. In some embodiments, the plurality of immobilized first complexed polymerases are in fluid communication with each other, allowing a solution of reagents (e.g., enzymes, including sequencing polymerases, multivalent molecules, nucleotides, and / or divalent cations) to flow over the support, whereby the plurality of immobilized complexed polymerases on the support react with the solution of reagents in a massively parallel manner.
[0255] In some embodiments, the method for sequencing further comprises step (b): contacting a plurality of first complexed polymerases with a plurality of multivalent molecules to form a plurality of multivalent complexed polymerases (e.g., bound complexes). In some embodiments, each multivalent molecule in the plurality of multivalent molecules comprises a core bound to a plurality of nucleotide arms, each nucleotide arm being bound to a nucleotide (e.g., a nucleotide unit) (e.g., Figures 22-25). In some embodiments, the contacting in step (b) is performed under conditions suitable for binding complementary nucleotide units of the multivalent molecule to at least two of the plurality of first complexed polymerases, thereby forming a plurality of multivalent complexed polymerases. In some embodiments, the conditions are suitable for inhibiting polymerase-catalyzed incorporation of complementary nucleotide units into primers of the plurality of multivalent complexed polymerases. In some embodiments, the plurality of multivalent molecules includes at least one multivalent molecule having multiple nucleotide arms (e.g., Figures 22-25), each of the multiple nucleotide arms being linked to a nucleotide analog (e.g., a nucleotide analog unit) that includes a chain-terminating moiety at the sugar 2' and / or 3' position. In some embodiments, the plurality of multivalent molecules includes at least one multivalent molecule including multiple nucleotide arms, each of the multiple nucleotide arms being linked to a nucleotide unit lacking a chain-terminating moiety. In some embodiments, at least one of the multivalent molecules in the plurality of multivalent molecules is labeled with a detectable reporter moiety. Any portion of the multivalent molecule can be labeled, including the core, the nucleotide arms, or the nucleobase. In some embodiments, the detectable reporter moiety includes a fluorophore. In some embodiments, the contacting in step (b) is performed in the presence of at least one non-catalytic cation, including strontium, barium, and / or calcium.
[0256] In some embodiments, the method for sequencing further comprises step (c): detecting the plurality of multivalent conjugated polymerases. In some embodiments, detecting comprises detecting multivalent molecules bound to the conjugated polymerases, wherein complementary nucleotide units of the multivalent molecules are bound to the primers but incorporation of the complementary nucleotide units is inhibited. In some embodiments, the multivalent molecules are labeled with a detectable reporter moiety to enable detection. In some embodiments, the labeled multivalent molecule comprises a fluorophore attached to the core, linker, and / or nucleotide units of the multivalent molecule.
[0257] In some embodiments, the method for sequencing further comprises step (d): identifying bases of complementary nucleotide units bound to the plurality of first complexed polymerases, thereby determining the sequence of the concatemeric molecule. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety corresponding to a particular nucleotide unit bound to the nucleotide arm, so as to enable identification of the complementary nucleotide units (e.g., the nucleotide bases adenine, guanine, cytosine, thymine, or uracil) bound to the plurality of first complexed polymerases.
[0258] In some embodiments, the second step of the two-step sequencing method generally comprises nucleotide incorporation. In some embodiments, the method for sequencing further comprises step (e): dissociating the plurality of multivalent complexed polymerases, removing the plurality of first sequencing polymerases and their bound multivalent molecules, and retaining the plurality of nucleic acid duplexes.
[0259] In some embodiments, the method for sequencing further comprises step (f): contacting the plurality of retained nucleic acid duplexes of step (e) with a plurality of second sequencing polymerases, wherein the contacting is performed under conditions suitable for binding the plurality of second sequencing polymerases to the plurality of retained nucleic acid duplexes, thereby forming a plurality of second complexed polymerases, each of the plurality of second complexed polymerases comprising a second sequencing polymerase bound to a nucleic acid duplex. In some embodiments, the second sequencing polymerase comprises a recombinant mutant sequencing polymerase.
[0260] In some embodiments, the plurality of first sequencing polymerases in step (a) have amino acid sequences that are 100% identical to the amino acid sequence of the plurality of second sequencing polymerases in step (f). In some embodiments, the plurality of first sequencing polymerases in step (a) have amino acid sequences that are different from the amino acid sequences of the plurality of second sequencing polymerases in step (f).
[0261] In some embodiments, the method for sequencing further includes step (g): contacting a plurality of second complexed polymerases with a plurality of nucleotides, wherein the contacting is performed under conditions suitable for binding complementary nucleotides from the plurality of nucleotides to at least two of the second complexed polymerases, thereby forming a plurality of nucleotide-complexed polymerases. In some embodiments, the contacting in step (g) is performed under conditions suitable for promoting polymerase-catalyzed incorporation of the bound complementary nucleotides into the primer by the nucleotide-complexed polymerase, thereby forming a plurality of nucleotide-complexed polymerases. In some embodiments, incorporating the nucleotide into the 3' end of the primer in step (g) comprises a primer extension reaction. In some embodiments, the contacting in step (g) is performed in the presence of at least one catalytic cation, including magnesium and / or manganese. In some embodiments, the contacting in step (g) is performed in the presence of magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises natural nucleotides (e.g., non-analog nucleotides) or nucleotide analogs. In some embodiments, the plurality of nucleotides comprises 2' and / or 3' chain terminating moieties, which may be removable or non-removable. In some embodiments, the plurality of nucleotides comprises nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety may comprise a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base by a linker, which may be cleavable / removable from the base or may not be removable from the base. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to the nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and identification of the nucleotide base.In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the plurality of nucleotides is labeled with a detectable reporter moiety.
[0262] In some embodiments, the method for sequencing further comprises step (h): if the nucleotide is labeled with a detectable reporter moiety, step (h) comprises detecting a complementary nucleotide incorporated into the primer of the nucleotide-conjugated polymerase. In some embodiments, the plurality of nucleotides are labeled with a detectable reporter moiety to enable detection. In some embodiments, in the method for sequencing concatemeric molecules, if the nucleotide is not labeled, the detection step is omitted.
[0263] In some embodiments, the method for sequencing further comprises step (i): if the nucleotide is labeled with a detectable reporter moiety, step (i) comprises identifying the base of the complementary nucleotide incorporated into the primer of the nucleotide-conjugated polymerase. In some embodiments, the identification of the incorporated complementary nucleotide in step (i) can be used to confirm the identity of the complementary nucleotide of the multivalent molecule bound to the plurality of first conjugated polymerases in step (d). In some embodiments, the identifying in step (i) can be used to determine the sequence of the nucleic acid concatemer molecule. In some embodiments, in the method for sequencing concatemer molecules, if the nucleotide is not labeled, the identifying step is omitted.
[0264] In some embodiments, the method for sequencing further comprises step (j): if step (g) is carried out by contacting a plurality of second complexed polymerases with a plurality of nucleotides comprising at least one nucleotide having a 2' and / or 3' chain-terminating moiety, removing the chain-terminating moiety from the incorporated nucleotides.
[0265] In some embodiments, the method for sequencing further comprises step (k): repeating steps (a) through (j) at least once. In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined in steps (c) and (d) by detecting and identifying multivalent molecules that bind to the sequencing polymerase but are not incorporated into the 3' end of the primer. In some embodiments, the sequence of the nucleic acid concatemer molecule can be determined (or confirmed) in steps (h) and (i) by detecting and identifying nucleotides that are incorporated into the 3' end of the primer.
[0266] In some embodiments, in any of the methods for sequencing nucleic acid molecules, binding of a plurality of first complexed polymerases to a plurality of multivalent molecules forms at least one avidity complex, the method comprising the steps of: (a) binding a first nucleic acid primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemeric template molecule, thereby forming a first binding complex, wherein a first nucleotide unit of the first multivalent molecule is bound to the first sequencing polymerase; and (b) binding a second nucleic acid primer, a second sequencing polymerase, and a first multivalent molecule to a second portion of the same concatemeric template molecule, thereby forming a second binding complex, wherein a second nucleotide unit of the first multivalent molecule is bound to the second sequencing polymerase; and the first and second binding complexes comprising the same multivalent molecule form an avidity complex. In some embodiments, the first sequencing polymerase comprises any wild-type or mutant polymerase described herein. In some embodiments, the second sequencing polymerase comprises any wild-type or mutant polymerase described herein. The concatemeric template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. First and second nucleic acid primers can bind to the sequencing primer binding sites along the concatemeric template molecule. Exemplary multivalent molecules are shown in Figures 22-25.
[0267] In some embodiments, in any of the methods for sequencing nucleic acid molecules, the method comprises combining a plurality of first complexed polymerases with a plurality of multivalent molecules to form at least one avidity complex, the method comprising the steps of: (a) contacting a plurality of sequencing polymerases and a plurality of nucleic acid primers with different portions of concatemeric nucleic acid concatemeric molecules to form at least first and second complexed polymerases on the same concatemeric template molecule; and (b) contacting the plurality of multivalent molecules with at least first and second complexed polymerases on the same concatemeric template molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first and second complexed polymerases, wherein at least a first nucleotide unit of the single multivalent molecule hybridizes to a first portion of the concatemeric template molecule, thereby forming a first binding complex (e.g., a first ternary complex). (c) contacting a first complexed polymerase comprising a primer, wherein at least a second nucleotide unit of the single multivalent molecule hybridizes to a second portion of the concatemeric template molecule, thereby forming a second binding complex (e.g., a second ternary complex), under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second binding complexes, such that the first and second binding complexes bound to the same multivalent molecule form an avidity complex; (d) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the concatemeric template molecule. In some embodiments, the plurality of sequencing polymerases comprises any wild-type or mutant sequencing polymerase described herein.The concatemer template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. Multiple nucleic acid primers can bind to the sequencing primer binding sites along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 22-25.
[0268] Sequencing by ligation In some aspects, the present disclosure provides a method for sequencing any of the immobilized concatemeric molecules described herein, wherein the sequencing method comprises a sequencing-by-binding (SBB) procedure using unlabeled chain-terminating nucleotides. In some embodiments, the SBB method includes: (a) sequentially contacting a primed template nucleic acid molecule with at least two separate mixtures under ternary complex-stabilizing conditions, each of the at least two separate mixtures comprising a polymerase and a nucleotide, thereby obtaining a primed template nucleic acid in contact with nucleotide analogs for the first, second, and third base types in the template under ternary complex-stabilizing conditions; (b) examining the at least two separate mixtures to determine whether a ternary complex has formed; and (c) examining the primed template nucleic acid with at least two separate mixtures to determine whether a ternary complex has formed. (d) adding the next correct nucleotide to the primer of the primed template nucleic acid molecule after step (b), thereby generating an extended primer; and (e) repeating steps (a)-(d) at least once on the primed template nucleic acid molecule containing the extended primer. Exemplary sequencing-by-binding methods are described in U.S. Patent Nos. 10,246,744 and 10,731,141 (the contents of both patents are incorporated herein by reference in their entireties).
[0269] Methods for sequencing using phosphate-chain labeled nucleotides The present disclosure provides a method for sequencing using an immobilized sequencing polymerase that binds a non-immobilized template molecule, wherein the sequencing reaction is performed with phosphate-chain-labeled nucleotides. In some embodiments, a covalently closed circular library molecule (600) can serve as the non-immobilized template molecule. In some embodiments, the sequencing method includes step (a): providing a support having a plurality of sequencing polymerases immobilized thereon. In some embodiments, the sequencing polymerase comprises a processive DNA polymerase. In some embodiments, the sequencing polymerase comprises a wild-type or mutant DNA polymerase, including, but not limited to, Phi29 DNA polymerase. In some embodiments, the support comprises a plurality of separate compartments, and the sequencing polymerase is immobilized at the bottom of the compartment. In some embodiments, the separate compartments comprise a silica bottom through which light can pass. In some embodiments, the separate compartments comprise a silica bottom configured with a nanophotonic confinement structure comprising holes in a metal cladding film (e.g., an aluminum cladding film). In some embodiments, the holes in the metal cladding have small openings, for example, on the order of 70 nm. In some embodiments, the height of the nanophotonic confinement structure is on the order of 100 nm. In some embodiments, the nanophotonic confinement structure comprises a zero-mode waveguide (ZMW). In some embodiments, the nanophotonic confinement structure contains a liquid.
[0270] In some embodiments, the sequencing method further includes step (b): contacting a plurality of immobilized sequencing polymerases with a plurality of single-stranded circular nucleic acid template molecules (e.g., covalently closed circular library molecules (600)) and a plurality of oligonucleotide sequencing primers under conditions suitable for each immobilized sequencing polymerase to bind to the single-stranded circular template molecule and for each sequencing primer to hybridize to each single-stranded circular template molecule, thereby generating a plurality of polymerase / template / primer complexes. In some embodiments, each sequencing primer hybridizes to a universal sequencing primer binding site on the single-stranded circular template molecule.
[0271] In some embodiments, the sequencing method further includes step (c): contacting a plurality of polymerase / template / primer complexes with a plurality of phosphate-chain-labeled nucleotides, each of which comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and a phosphate chain containing 3 to 20 phosphate groups (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 phosphate groups), wherein the terminal phosphate groups are attached to detectable reporter moieties (e.g., fluorophores). The first, second, and third phosphate groups may be referred to as alpha, beta, and gamma phosphate groups. In some embodiments, the specific detectable reporter moieties attached to the terminal phosphate groups correspond to nucleotide bases (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to enable detection and identification of the nucleobases. In some embodiments, a plurality of polymerase / template / primer complexes are contacted with a plurality of phosphate-chain-labeled nucleotides under conditions suitable for polymerase-catalyzed nucleotide incorporation. In some embodiments, a sequencing polymerase can bind to a complementary phosphate-chain-labeled nucleotide and incorporate a complementary nucleotide opposite the nucleotide in the template molecule. In some embodiments, the polymerase-catalyzed nucleotide incorporation reaction cleaves between the alpha phosphate group and the beta phosphate group, thereby releasing a plurality of phosphate chains attached to fluorophores.
[0272] In some embodiments, the sequencing method further comprises step (d): detecting a fluorescent signal emitted by a phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated onto the end of the sequencing primer. In some embodiments, step (d) further comprises identifying the phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated onto the end of the sequencing primer.
[0273] In some embodiments, the sequencing method further comprises step (d): repeating steps (c) through (d) at least once. In some embodiments, the sequencing method using phosphate chain-labeled nucleotides can be performed according to the methods described in U.S. Patent Nos. 7,170,050, 7,302,146, and / or 7,405,281.
[0274] Traditional pooling workflow for multiplexing In one embodiment, a method for preparing a multiplex mixture of sequences of interest isolated from multiple sample sources includes (a) providing two or more populations of single-stranded nucleic acid library molecules (100), each population of library molecules (100) contained in a separate compartment, and wherein the nucleic acid library molecules in a given population include (i) a sequence of interest (110), (ii) a universal adapter sequence (120) having a binding sequence for a forward sequencing primer, and (iii) a universal adapter sequence (130) having a binding sequence for a reverse sequencing primer (e.g., FIG. 1 ).
[0275] In some embodiments, the method for preparing a multiplex mixture of sequences of interest isolated from multiple sample sources further comprises step (b): providing a plurality of double-stranded splint adapters (200), each comprising a first splint strand (300) hybridized to a second splint strand (400) (e.g., FIG. 2). In some embodiments, the second splint strand (400) comprises at least one sample index sequence and optionally a short random sequence NNN.
[0276] In some embodiments, the method for preparing a multiplex mixture of sequences of interest isolated from multiple sample sources includes step (c): contacting, in separate compartments, a population of single-stranded nucleic acid library molecules (100) with an assignment of a plurality of double-stranded splint adapters (200), wherein the contacting is performed under conditions suitable for hybridizing a portion of the first splint strand (300) to a portion of the library molecule (100), thereby circularizing the library molecules to generate a population of library-sprint complexes (500), such that a region (320) of each first splint strand becomes a universal target for the forward sequencing primer binding site of each library molecule (100). The method further includes contacting a library molecule (100) with a first splint strand (300) that hybridizes to a monkey adapter sequence (120), a region (330) of each first splint strand hybridizes to a universal adapter sequence (130) for the reverse sequencing primer binding site of each library molecule (100), each library-sprint complex (500) comprising a first nick between the 5' end of the library molecule and the 3' end of the second splint strand (300), each library-sprint complex (500) comprising a second nick between the 5' end of the second splint strand (300) and the 3' end of the library molecule (100), the first and second nicks being enzymatically ligatable (see, for example, Figures 3 and 4).
[0277] In some embodiments, the method for preparing a multiplex mixture of sequences of interest isolated from multiple sample sources further comprises step (d): contacting, in separate compartments, the population of library-sprint complexes (500) with a ligase under conditions suitable for enzymatic ligation of the first and second nicks, thereby generating a population of covalently closed circular library molecules (600) each hybridized to a first splint strand (300) (see, e.g., Figure 4).
[0278] In some embodiments, the method for preparing a multiplex mixture of sequences of interest isolated from multiple sample sources further comprises step (e): pooling together the populations of covalently closed circular library molecules (600) from the separate compartments to produce a multiplex mixture of covalently closed circular library molecules (600) comprising a multiplex mixture of sequences of interest isolated from multiple sample sources.
[0279] In some embodiments, sequences of interest can be isolated from two or more different sample sources (e.g., 2-10, 10-50, 50-100, 100-250, any range therebetween, or more than 250 different sample sources). Exemplary nucleic acid sources include naturally occurring sources, recombinant sources, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or whole organisms. Exemplary nucleic acid sources include fresh sources, frozen sources, fresh frozen sources, or archived (e.g., formalin-fixed, paraffin-embedded; FFPE) sources. Those skilled in the art will recognize that nucleic acids can be isolated from many other sources. The sequences of interest in a given population may have the same or different sequences.
[0280] In some embodiments, the number of populations of single-stranded nucleic acid library molecules (100) in step (a) can be 2 to 10, 10 to 50, 50 to 100, 100 to 250, any range therebetween, or more than 250 different populations of single-stranded nucleic acid library molecules (100). In some embodiments, any number of different populations of covalently closed circular library molecules (600) can be pooled together in step (e), for example, 2 to 10, 10 to 50, 50 to 100, 100 to 200, any range therebetween, or more than 200 different populations of covalently closed circular library molecules (600) can be pooled together.
[0281] Those skilled in the art will recognize that any number of distinct compartments (e.g., 2-10, 10-50, 50-100, 100-250, any range therebetween, or more distinct compartments) (e.g., a multi-well plate such as a 96-well plate) can be used in step (a).
[0282] In some embodiments, the 3' end of the first splint strand (300) hybridized to the covalently closed circular library molecule (600) of step (d) or (e) comprises an extendable 3' OH terminus, which can serve as a starting point for a primer extension reaction (e.g., a rolling circle amplification reaction).
[0283] In some embodiments, in step (d) or (e), the population of covalently closed circular library molecules (600) hybridized to the first splint strands (300) can optionally be reacted with at least one exonuclease enzyme to remove a plurality of the first splint strands (300) and retain a plurality of covalently closed circular library molecules (600). In some embodiments, the at least one exonuclease enzyme comprises exonuclease I, thermolabile exonuclease I, and / or T7 exonuclease.
[0284] In some embodiments, the single-stranded nucleic acid library molecules (100) of step (a) further comprise any one or any combination of two or more of a universal binding sequence for a forward amplification primer, a universal binding sequence for a reverse amplification primer, and / or a universal binding sequence for a compaction oligonucleotide.
[0285] Sample indexing for improved base calling Generally, it is desirable to prepare a nucleic acid library distributed on a support (e.g., a coated flow cell), and the library molecules are converted into template molecules immobilized at high density on the support for massively parallel sequencing. With template molecules immobilized at high density at random locations on the support, the challenge of resolving high-density fluorescent images for accurate base calling during sequencing runs becomes difficult.
[0286] The nucleotide diversity of a population of immobilized template molecules refers to the relative proportions of the nucleotides A, G, C, and T present in each sequencing cycle. In some embodiments, optimal high-diversity template molecules contain a region of interest (insert) with approximately equal proportions of all four nucleotides represented in each cycle of a sequencing run. In some embodiments, low-diversity template molecules contain a region of interest (insert) with a high proportion of certain nucleotides and a low proportion of other nucleotides. To overcome the problem of low-diversity template molecules, small amounts of high-diversity molecules prepared from PhiX bacteriophage can be mixed with the template molecules of interest (e.g., a PhiX spike-in library) and sequenced together, for example, on the same flow cell. The PhiX spike-in library provides nucleotide diversity, but also occupies space on the flow cell, thereby displacing template molecules bearing the sequence of interest and reducing the amount of sequencing data that can be obtained from the template molecules (e.g., reducing sequencing throughput).
[0287] Another way to overcome the problem of low-diversity template molecules is to prepare template molecules with at least one sample index sequence designed to be color-balanced. However, it may be desirable to design single index sample sequences or sets of paired index sample sequences for multiple sample index sets, for example, 16-plex, 24-plex, 96-plex, or larger plex levels. It may be difficult to design single or paired sample index sequences for large sample index sets in which all sample index sequences are color-balanced (see, for example, Figures 19 and 20).
[0288] As described herein, an alternative method for overcoming the challenge of sequencing low-diversity template molecules (e.g., at high density on a support) is to prepare template molecules having at least one sample index sequence comprising a short random sequence (e.g., NNN) directly linked to the sample index sequence, where the short random sequence (e.g., NNN) provides nucleotide diversity and color balance. In some embodiments, the sample index sequence comprises a short random sequence (NNN) directly linked to the sample index sequence. In some embodiments, the short random sequence (NNN) is upstream or downstream of the sample index sequence. Some exemplary sample index sequences include, but are not limited to, NNNGTAGGAGCC, NNNCCGCTGCTA, NNNAACAACAAG, NNNGGTGGTCTA, NNNTTGGCCAAC, NNNCAGGAGTGC, and NNNATCACACTA (see, e.g., Table 3).
[0289] Those skilled in the art will recognize that the universal sample index may be of any length and have any sequence that can be used to distinguish sequences of interest obtained from different sample sources in a multiplex assay. For a given sample index, for example, but not limited to, a population of NNNGTAGGAGCC, the population contains a mixture of individual sample index molecules, each containing the same universal sample index sequence (e.g., GTAGGAGCC) and a different short random sequence (e.g., NNN), and up to 64 different short random sequences may be present within a given population of sample indexes.
[0290] In a population of sample-indexed template molecules, the short random sequence (NNN) of the sample index can provide high nucleotide diversity, with all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run being included in approximately equal proportions (see Figures 19 and 20). The high nucleotide diversity of the short random sequence can also provide color balance during each cycle of a sequencing run. An advantage of designing the sample index to include a short random sequence (e.g., NNN) is that in a low-plex population of template molecules (e.g., 2-plex or 4-plex), the universal sample index sequence identifying two or four different samples does not need to exhibit nucleotide diversity (see, e.g., Figures 19 and 20). Additionally, the nucleotide diversity of the short random sequence (e.g., NNN) can obviate the need to include a PhiX spike-in library or allow for the use of a smaller amount of PhiX spike-in library dispensed onto the flow cell and sequenced.
[0291] A template molecule may include a first sample index sequence that includes a short random sequence (NNN) and a second sample index sequence that lacks the short random sequence. In some embodiments, the short random sequence (e.g., NNN) provides sufficient nucleotide diversity and color balance, so sequencing data from only the sample index sequences with the short random sequence (NNN) are used for polony mapping and template registration. Both types of sample indexes can be used to distinguish sequences of interest obtained from different sample sources in a multiplex assay.
[0292] The order in which the region of the sequence of interest and the sample index region(s) are sequenced can also be used to improve the sequencing challenge of low-diversity template molecules. For example, but not limited to, the sample index region may be sequenced first, followed by sequencing the region of the sequence of interest, and the sample index sequence may be associated with the sequence region of interest. For example, but not limited to, the sample index region may be sequenced by first sequencing a short random sequence (e.g., NNN), optionally sequencing at least a portion of the universal sample index, and then sequencing the region of the sequence of interest. In a population of sample-indexed template molecules, the short random sequence (e.g., NNN) can provide nucleotide diversity that may not be provided by the region of the sequence of interest of the template molecule. The sequence of the sample index can provide improved nucleotide diversity and color balance for polony mapping and template registration.
[0293] In addition, when sequencing the sample index region first, the length of the sequenced sample index region may be relatively short (e.g., less than 30 nucleotides in length) to ensure more complete dehybridization of the sequenced sample index region products. Milder dehybridization conditions can be used to remove most or all of the sequenced sample index region products, reducing the level of residual signal from any sequencing products that remain hybridized to the template molecule. In contrast, the region of interest is typically much longer than the sample index region (e.g., more than 100 nucleotides in length). In some embodiments, if the region of interest is sequenced before the sample index region, the products of the sequenced region of interest must be subjected to more stringent dehybridization conditions, which may damage the template molecule, to remove any products that remain hybridized to the template molecule.
[0294] In some embodiments, the present disclosure provides template molecules each comprising at least one sample index sequence that can be used to distinguish sequences of interest obtained from different sample sources in a multiplex assay, wherein at least one sample index sequence comprises a short random sequence (e.g., NNN) linked to a universal sample index sequence. At least one sample index sequence can comprise sequence diversity for improved base calling. At least one sample index sequence can be used to improve base calling accuracy.
[0295] In some embodiments, the short random sequence (e.g., NNN) is located upstream of the universal sample index sequence, such that during a sequencing run, the random sequence portion is sequenced before the universal sample index sequence. In some embodiments, the short random sequence is located downstream of the universal sample index sequence, such that during a sequencing run, the random portion is sequenced after the universal sample index sequence.
[0296] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, for example, but not limited to, AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, in the population of template molecules, the universal sample index sequence includes short random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing attempt.
[0297] In some embodiments, the short random sequence (e.g., NNN) comprises 3 to 20 nucleotides, 3 to 10 nucleotides, 3 to 8 nucleotides, 3 to 6 nucleotides, 3 to 5 nucleotides, or 3 to 4 nucleotides, or any range therebetween.
[0298] In some embodiments, short random sequences (e.g., NNN) include, but are not limited to, AGC, AGT, GAC, GAT, CAT, CAG, TAG, TAC. One of skill in the art will recognize that many more random sequences (e.g., 64 possible combinations) may be prepared, with each base "N" at a given position within a random sequence being independently selected from A, G, C, T, or U.
[0299] In some embodiments, the universal sample index sequence comprises between 5 and 20 nucleotides, between 7 and 18 nucleotides, or between 9 and 16 nucleotides, or any range therebetween.
[0300] In some embodiments, individual sample index sequences in a population of sample indexes include a universal sample index sequence and a short random sequence (e.g., NNN). In some embodiments, the short random sequences in the population of sample index sequences have an overall base composition of about 25% or about 20-30% of all four nucleotide bases (e.g., A, G, C, and T / U), providing nucleotide diversity in each sequencing cycle during sequencing of the short random sequences (e.g., NNN).
[0301] In some embodiments, in a population of sample index sequences, the percentage of adenine (A) at any given position within the short random sequence is about 20-30%, about 15-35%, about 10-40%, or any range therebetween. In some embodiments, in a population of sample index sequences, the percentage of guanine (G) at any given position within the short random sequence is about 20-30%, about 15-35%, about 10-40%, or any range therebetween. In some embodiments, in a population of sample index sequences, the percentage of cytosine (C) at any given position within the short random sequence is about 20-30%, about 15-35%, about 10-40%, or any range therebetween. In some embodiments, in a population of sample index sequences, the percentage of thymine (T) or uracil (U) at any given position within the short random sequence is about 20-30%, about 15-35%, about 10-40%, or any range therebetween.
[0302] In some embodiments, the population of sample index sequences has a ratio of adenine (A) and thymine (T), or adenine (A) and uracil (U), at any given position within the short random sequence, of about 10-65%. In some embodiments, the population of sample index sequences has a ratio of guanine (G) and cytosine (C) at any given position within the short random sequence, of about 10-65%.
[0303] In some embodiments, in the population of sample index sequences, the sequence diversity of the short random sequences ensures that no sequencing cycle is represented with fewer than four different nucleotide bases during sequencing of at least the short random sequence (e.g., NNN).
[0304] In some embodiments, the random sequence (e.g., NNN) provides a balanced ratio of the nucleobases adenine, cytosine, guanine, thymine, and / or uracil (see, e.g., Figure 19). In some embodiments, in a population of sample-indexed template molecules, the random sequence (e.g., NNN), together with at least a portion of the universal sample index sequence, provides a balanced ratio of the nucleobases adenine, cytosine, guanine, thymine, and / or uracil represented in each cycle of a sequencing run.
[0305] In some embodiments, the sequencing reaction involves using a polymerase and nucleotides (e.g., nucleotide analogs) labeled with different fluorophores corresponding to the nucleobases. In some embodiments, sequencing a random sequence (e.g., NNN) using the labeled nucleotides provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil in each cycle of the sequencing run. In some embodiments, sequencing a random sequence (e.g., NNN) and at least a portion of the universal sample index sequence using the labeled nucleotides provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil (see, e.g., Figure 19). The labeled nucleotides emit a fluorescent signal during the sequencing reaction. In some embodiments, the sequencing reaction is performed on a sequencing instrument having a detector that captures fluorescent images from the sequencing reaction on the immobilized template molecules. The sequencing device can be configured to relay fluorescent imaging data captured by the detector to a computer system programmed to determine (e.g., map) the locations of the immobilized template molecules on the flow cell. The computer system can generate a map of the locations of the immobilized template molecules based on fluorescent imaging data of only the random sequence (e.g., NNN) or based on the random sequence (e.g., NNN) and at least a portion of the universal sample index sequence. Thus, a reduced number of sequencing cycles used to sequence the random sequence (e.g., NNN) and optionally a portion of the universal sample index sequence can be used to generate a map of the locations of the immobilized template molecules. The computer system can be configured to extract the fluorescent color and intensity of only the random sequence (e.g., NNN) or of at least a portion of the random sequence (e.g., NNN) and the universal sample index sequence.The computer system may be configured to use the location of a given immobilized template molecule and the fluorescent color and intensity associated with the given template molecule (established while sequencing the random sequence) for base calling while sequencing the insert region (110). The computer system may be configured to detect phasing and pre-phasing while sequencing the random sequence (e.g., NNN) and the universal sample index sequence, and the insert region (110). In some embodiments, a balanced ratio of fluorescent colors provided by the random sequence (e.g., NNN) in each sequencing cycle can improve the quality of data processed from the fluorescent images captured by the detector, which in turn can improve the computer system's ability to determine the location, color, and intensity of the immobilized template molecule on the flow cell, all of which can improve base calling accuracy and quality scores for the sequenced insert region (110).
[0306] In some embodiments, the sequencing reaction involves the use of a polymerase and multivalent molecules labeled with different fluorophores that correspond to the nucleic acid bases (e.g., adenine, guanine, cytosine, thymine, or uracil) of the nucleotide units attached to the nucleotide arms in a given multivalent molecule. In some embodiments, the core of each multivalent molecule is bound to a fluorophore that corresponds to the nucleotide unit (e.g., adenine, guanine, cytosine, thymine, or uracil) attached to the nucleotide arms in a given multivalent molecule (see, e.g., Figures 22-25). In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or nucleotide base attached to a fluorophore, and the fluorophore attached to a given linker or nucleotide base corresponds to the nucleotide base (e.g., adenine, guanine, cytosine, thymine, or uracil) of the nucleotide arm. In some embodiments, sequencing a random sequence (e.g., NNN) using labeled multivalent molecules provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil in each cycle of a sequencing run. In some embodiments, sequencing a random sequence (e.g., NNN) and at least a portion of a universal sample index sequence using labeled multivalent molecules provides a balanced ratio of fluorescent colors corresponding to the nucleobases adenine, cytosine, guanine, thymine, and / or uracil (see, e.g., Figure 19). The labeled multivalent molecules emit a fluorescent signal during the sequencing reaction. In some embodiments, the sequencing reaction is performed on a sequencing device having a detector that captures fluorescent images from the sequencing reaction on the immobilized template molecules. The sequencing device can be configured to relay fluorescent imaging data captured by the detector to a computer system programmed to determine (e.g., map) the location of the immobilized template molecules (polonies) on the flow cell.The computer system can generate a map of the locations of the immobilized template molecules based on fluorescent imaging data of only the random sequence (e.g., NNN) or based on the random sequence (e.g., NNN) and at least a portion of the universal sample index sequence. Thus, a reduced number of sequencing cycles used to sequence the random sequence (e.g., NNN) and, optionally, a portion of the universal sample index sequence, can be used to generate a map of the locations of the immobilized template molecules. The computer system may be configured to extract the fluorescent color and intensity of only the random sequence (e.g., NNN) or of only the random sequence (e.g., NNN) and the universal sample index sequence. The computer system may be configured to use the location of a given immobilized template molecule and the fluorescent color and intensity associated with the given template molecule (established while sequencing the random sequence) for base calling while sequencing the insert region (110). The computer system may be configured to detect phasing and pre-phasing while sequencing the random sequence (e.g., NNN) and the universal sample index sequence, and the insert region (110). In some embodiments, a balanced ratio of fluorescent colors provided by the random sequence (e.g., NNN) in each sequencing cycle can improve the quality of data processed from the fluorescent images captured by the detector, which in turn can improve the ability of a computer system to determine the location, as well as the color and intensity, of immobilized template molecules on the flow cell, all of which can improve base-calling accuracy and quality scores of the sequenced insert region (110).
[0307] First embodiment: Order of sequencing sample index sequences In some embodiments, the covalently closed circular molecule (600) comprises: (i) a first subregion (411) comprising a universal binding sequence for an immobilized surface pinning primer; (ii) a second subregion (412) comprising a short random sequence (NNN) and a sample index sequence; (iii) a third subregion (413) comprising a universal binding sequence for an immobilized surface capture primer; (iv) a universal binding sequence for a reverse sequencing primer (130); (v) an insert sequence (110); and (vi) a universal binding sequence for a forward sequencing primer (120) (see, e.g., Figure 4). In some embodiments, rolling circle amplification is performed on a plurality of covalently closed circular library molecules (600) by hybridizing the plurality of covalently closed circular library molecules (600) to a plurality of surface capture primers immobilized on a support, where the surface capture primers initiate amplification, thereby generating a plurality of immobilized concatemers, each having a tandem repeat sequence of its cognate covalently closed circular library molecule (600). In some embodiments, rolling circle amplification is performed in the presence of a plurality of compaction oligonucleotides that can hybridize to the concatemers at their universal binding sequences for the surface pinning primers, the universal binding sequence for the surface capture primers, the universal binding sequence for the reverse sequencing primer, and / or the universal binding sequence for the forward sequencing primer. In some embodiments, the plurality of immobilized capture primers lack uracil bases. In some embodiments, the rolling circle amplification reaction comprises a plurality of nucleotides, including dATP, dGTP, dCTP, dTTP, and dUTP, and produces a plurality of immobilized concatemers, each of which comprises randomly distributed uracil bases. In some embodiments, at least a portion of the immobilized concatemers are sequenced.
[0308] In some embodiments, the sequencing order includes (1) sequencing the short random sequence (NNN) and sample index sequence (412) of the concatemer, and (2) sequencing the insert region (110) of the concatemer. In some embodiments, the sequencing order further includes (3) performing a pairwise turn reaction such that the immobilized concatemer molecule is replaced with an immobilized second strand that is complementary to the concatemer molecule, and (4) sequencing the insert region (110) on the second strand. In some embodiments, sequencing the short random sequence (NNN) and sample index sequence (412) may provide sufficient nucleotide diversity and color balance for polony mapping and concatemer registration.
[0309] In some embodiments, a method for sequencing concatemeric molecules immobilized to a support comprises the step (a): hybridizing the concatemeric molecules with a first plurality of soluble sequencing primers that hybridize to at least a portion of the third subregion (413), and sequencing the short random sequence (e.g., NNN) and universal sample index sequence of the third subregion (413), thereby generating a plurality of sample index extension products hybridized to the immobilized concatemeric molecules, wherein the plurality of sample index extension products are complementary to the short random sequence (e.g., NNN) and universal sample index sequence.
[0310] In some embodiments, the method for sequencing further comprises step (b): removing the first plurality of sample index extension products and retaining the immobilized concatemeric molecules.
[0311] In some embodiments, the method for sequencing further comprises step (c): hybridizing the retained immobilized concatemer molecules with a second plurality of soluble sequencing primers that hybridize to the universal binding sequence (120) for the forward sequencing primer and sequencing the insert region (110), thereby generating a plurality of insert sequence extension products hybridized to the immobilized concatemer molecules.
[0312] In some embodiments, the method for sequencing further comprises step (d): displacing the plurality of insert sequence extension products hybridized to the immobilized concatemer molecules by performing a primer extension reaction using a strand-displacing polymerase and a plurality of nucleotides to generate second strand extension products hybridized to the immobilized concatemer molecules comprising the immobilized capture primer.
[0313] In some embodiments, the method for sequencing further comprises step (e): removing the immobilized concatemer molecules by generating abasic sites within the immobilized concatemer molecules at uracil sites and generating gaps at the abasic sites, thereby retaining the second-strand retention products generated in step (d) while generating concatemer molecules containing gaps, and each second-strand retention product is retained by hybridization to an immobilized capture primer. In some embodiments, a pairwise turn is achieved by performing steps (d) and (e).
[0314] In some embodiments, the method for sequencing further comprises step (f): hybridizing the retained second strand extension products with a third plurality of soluble sequencing primers that hybridize to universal binding sequences for reverse sequencing primers (130) and sequencing at least a portion of the insert region (110).
[0315] In some embodiments, the method for sequencing further comprises (i) assigning the sequence of the insert region (110) to (ii) a sample index sequence, thereby identifying the insert region as having been obtained from the first source. In some embodiments, the assigning may occur after steps (c) and / or (f).
[0316] In some embodiments, removing the plurality of sequencing extension products in step (b) can be performed using a denaturing reagent containing an SSC (e.g., saline-sodium citrate) buffer, with or without formamide, at a temperature that promotes nucleic acid denaturation, e.g., 50-90°C. In some embodiments, removing the plurality of sequencing extension products in step (b) can be performed using a dehybridization reagent at a temperature that promotes nucleic acid denaturation, e.g., 50-90°C. In some embodiments, the dehybridization reagent includes a pH buffer, a reducing agent, a monovalent salt, and a crowding agent. In some embodiments, the dehybridization reagent further includes a chaotropic agent.
[0317] In some embodiments, the sequencing of steps (a), (c), and (f) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (f) comprises performing any of the two-step sequencing methods described herein using a sequencing polymerase, a detectably labeled multivalent molecule, and labeled or unlabeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (f) comprises performing any of the binding sequencing methods or sequencing methods described herein using phosphate-chain labeled nucleotides.
[0318] In some embodiments, the density of the plurality of concatemeric molecules immobilized on the support is greater than 1 mm 2 Approximately 10 per 2 ~10 15 (e.g., 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 In some embodiments, the plurality of concatemer molecules are immobilized at random locations on the support. In some embodiments, the plurality of concatemer molecules are immobilized in a predetermined pattern on the support.
[0319] Second embodiment: Order of sequencing sample index sequences In some embodiments, the covalently closed circular molecule (600) comprises: (i) a first subregion (411) comprising a universal binding sequence for an immobilized surface pinning primer; (ii) a second subregion (412) comprising a short random sequence (NNN) and a sample index sequence; (iii) a third subregion (413) comprising a universal binding sequence for an immobilized surface capture primer; (iv) a universal binding sequence for a reverse sequencing primer (130); (v) an insert sequence (110); and (vi) a universal binding sequence for a forward sequencing primer (120) (see, e.g., Figure 4). In some embodiments, rolling circle amplification is performed on a plurality of covalently closed circular library molecules (600) by hybridizing the plurality of covalently closed circular library molecules (600) to a plurality of surface capture primers immobilized on a support, where the surface capture primers initiate amplification, thereby generating a plurality of immobilized concatemers, each having a tandem repeat sequence of its cognate covalently closed circular library molecule (600). In some embodiments, rolling circle amplification is performed in the presence of a plurality of compaction oligonucleotides that can hybridize to the concatemers at their universal binding sequences for the surface pinning primers, the universal binding sequence for the surface capture primers, the universal binding sequence for the reverse sequencing primer, and / or the universal binding sequence for the forward sequencing primer. In some embodiments, the plurality of immobilized capture primers lack uracil bases. In some embodiments, the rolling circle amplification reaction comprises a plurality of nucleotides, including dATP, dGTP, dCTP, dTTP, and dUTP, and produces a plurality of immobilized concatemers, each of which comprises randomly distributed uracil bases. In some embodiments, at least a portion of the immobilized concatemers are sequenced.
[0320] In some embodiments, the sequencing order includes (1) sequencing the insert region (110) of the concatemer, and (2) sequencing the short random sequence (NNN) and sample index sequence (412) of the concatemer. In some embodiments, the sequencing order further includes (3) performing a pairwise turn reaction such that the immobilized concatemer molecule is replaced with an immobilized second strand that is complementary to the concatemer molecule, and (4) sequencing the insert region (110) on the second strand. In some embodiments, sequencing the short random sequence (NNN) and sample index sequence (412) may provide sufficient nucleotide diversity and color balance for polony mapping and concatemer registration.
[0321] In some embodiments, a method for sequencing concatemeric molecules immobilized on a support comprises step (a): hybridizing the immobilized concatemeric molecules with a first plurality of soluble sequencing primers that hybridize to a universal binding sequence for the forward sequencing primer (120) and sequencing the insert region (110), thereby generating a plurality of insert sequence extension products hybridized to the immobilized concatemeric molecules.
[0322] In some embodiments, the method for sequencing further comprises step (b): removing the plurality of insert sequence extension products and retaining the immobilized concatemeric molecules.
[0323] In some embodiments, the method for sequencing further comprises step (c): hybridizing the retained concatemer molecules with a second plurality of soluble sequencing primers that hybridize to at least a portion of the third subregion (413) and sequencing the short random sequence (e.g., NNN) and universal sample index sequence of the third subregion (412), thereby generating a plurality of sample index extension products hybridized to the immobilized concatemer molecules, wherein the plurality of sample index extension products are complementary to the short random sequence (e.g., NNN) and universal sample index sequence.
[0324] In some embodiments, the method for sequencing further comprises step (d): displacing a plurality of the sample index extension products hybridized to the immobilized concatemer molecules by performing a primer extension reaction using a strand-displacing polymerase and a plurality of nucleotides to generate second strand extension products hybridized to the immobilized concatemer molecules comprising the immobilized capture primer.
[0325] In some embodiments, the method for sequencing further comprises step (e): removing the immobilized concatemer molecules by generating abasic sites within the immobilized concatemer molecules at uracil sites and generating gaps at the abasic sites, thereby retaining the second-strand retention products generated in step (d) while generating concatemer molecules containing gaps, and each second-strand retention product is retained by hybridization to an immobilized capture primer. In some embodiments, a pairwise turn is achieved by performing steps (d) and (e).
[0326] In some embodiments, the method for sequencing further comprises step (f): hybridizing the retained second strand extension products with a third plurality of soluble sequencing primers that hybridize to universal binding sequences for reverse sequencing primers (130) and sequencing at least a portion of the insert region (110).
[0327] In some embodiments, the method for sequencing further comprises (i) assigning the sequence of the insert region (110) to (ii) a sample index sequence, thereby identifying the insert region as having been obtained from the first source. In some embodiments, the assigning may occur after steps (c) and / or (f).
[0328] In some embodiments, removing the plurality of sequencing extension products in step (b) can be performed using a denaturing reagent containing an SSC (e.g., saline-sodium citrate) buffer, with or without formamide, at a temperature that promotes nucleic acid denaturation, e.g., 50-90°C. In some embodiments, removing the plurality of sequencing extension products in step (b) can be performed using a dehybridization reagent at a temperature that promotes nucleic acid denaturation, e.g., 50-90°C. In some embodiments, the dehybridization reagent includes a pH buffer, a reducing agent, a monovalent salt, and a crowding agent. In some embodiments, the dehybridization reagent further includes a chaotropic agent.
[0329] In some embodiments, the sequencing of steps (a), (c), and (f) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (f) comprises performing any of the two-step sequencing methods described herein using a sequencing polymerase, a detectably labeled multivalent molecule, and labeled or unlabeled nucleotide analogs. In some embodiments, the sequencing of steps (a), (c), and (f) comprises performing any of the binding sequencing methods or sequencing methods described herein using phosphate-chain labeled nucleotides.
[0330] In some embodiments, the density of the plurality of concatemeric molecules immobilized on the support is greater than 1 mm 2 Approximately 10 per 2 ~10 15 (e.g., 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , or 10 15 In some embodiments, the plurality of concatemer molecules are immobilized at random locations on the support. In some embodiments, the plurality of concatemer molecules are immobilized in a predetermined pattern on the support.
[0331] Third embodiment: Sequencing order In some embodiments, the covalently closed circular molecule (600) comprises: (i) a first subregion (411) comprising a universal binding sequence for an immobilized surface pinning primer; (ii) a second subregion (412) comprising a short random sequence (NNN) and a sample index sequence; (iii) a third subregion (413) comprising a universal binding sequence for an immobilized surface capture primer; (iv) a universal binding sequence for a reverse sequencing primer (130); (v) an insert sequence (110); and (vi) a universal binding sequence for a forward sequencing primer (120) (see, e.g., Figure 4). In some embodiments, rolling circle amplification is performed on a plurality of covalently closed circular library molecules (600) by hybridizing the plurality of covalently closed circular library molecules (600) to a plurality of surface capture primers immobilized on a support, where the surface capture primers initiate amplification, thereby generating a plurality of immobilized concatemers, each having a tandem repeat sequence of its cognate covalently closed circular library molecule (600). In some embodiments, rolling circle amplification is performed in the presence of a plurality of compaction oligonucleotides that can hybridize to the concatemers at their universal binding sequences for the surface pinning primers, the universal binding sequence for the surface capture primers, the universal binding sequence for the reverse sequencing primer, and / or the universal binding sequence for the forward sequencing primer. In some embodiments, the plurality of immobilized capture primers lack uracil bases. In some embodiments, the rolling circle amplification reaction comprises a plurality of nucleotides, including dATP, dGTP, dCTP, dTTP, and dUTP, and produces a plurality of immobilized concatemers, each of which comprises randomly distributed uracil bases. In some embodiments, at least a portion of the immobilized concatemers are sequenced.
[0332] In some embodiments, the sequencing order includes (1) sequencing the first 3-5 bases of the insert region (110) of the concatemer, (2) sequencing the short random sequence (NNN) and the sample index sequence (412) of the concatemer, and (3) sequencing the remainder of the insert region (110) of the concatemer. In some embodiments, the sequencing order further includes (4) performing a pairwise turn reaction such that the immobilized concatemer molecules are replaced with an immobilized second strand that is complementary to the concatemer molecules, and (5) sequencing the insert region (110) on the second strand. In some embodiments, sequencing the first 3-5 bases of the insert region (110) of the concatemer may provide sufficient nucleotide diversity and color balance for polony mapping and concatemer registration. In some embodiments, sequencing the short random sequence (NNN) and the sample index sequence (412) may provide sufficient nucleotide diversity and color balance for polony mapping and concatemer registration.
[0333] In some embodiments, a method for sequencing concatemer molecules immobilized on a support includes step (a): hybridizing the concatemer molecules with a first plurality of soluble sequencing primers that hybridize to a universal binding sequence (120) for the forward sequencing primer, thereby sequencing the first 3-5 bases of the insert region (110), thereby generating a plurality of insert extension products hybridized to the immobilized concatemer molecules, the plurality of insert extension products being complementary to the sequence of interest (110). The sequence of the first 3-5 bases of the insert region (110) may provide sufficient sequence diversity and color balance for polony mapping and concatemer registration.
[0334] In some embodiments, the method for sequencing further comprises step (b): removing the multiple insert extension products and retaining the immobilized concatemeric molecules.
[0335] In some embodiments, the method for sequencing further comprises step (c): hybridizing the retained concatemer molecules with a second plurality of soluble sequencing primers that hybridize to at least a portion of the third subregion (413) and sequencing the short random sequence (e.g., NNN) and universal sample index sequence of the third subregion (412), thereby generating a plurality of sample index extension products hybridized to the immobilized concatemer molecules, wherein the plurality of sample index extension products are complementary to the short random sequence (e.g., NNN) and universal sample index sequence.
[0336] In some embodiments, the method for sequencing further comprises step (d): removing the plurality of sample index extension products and retaining the immobilized concatemeric molecules.
[0337] In some embodiments, the method for sequencing further comprises step (e): hybridizing the retained concatemer molecules with a third plurality of soluble sequencing primers that hybridize to the universal binding sequence (120) for the forward sequencing primer, and sequencing the remainder of the insert region (110) or sequencing the entire length of the insert region (110), thereby generating a plurality of insert sequence extension products hybridized to the immobilized concatemer molecules, wherein the plurality of insert extension products are complementary to the sequence of the insert region (110). In some embodiments, in step (e), the first 3-5 bases of the insert region (110) can be sequenced using labeled nucleotides and / or labeled multivalent molecules. In some embodiments, in step (e), the first 3-5 bases of the insert region (110) can be sequenced using unlabeled nucleotides and / or unlabeled multivalent molecules (e.g., dark sequencing). In some embodiments, in step (e), the remainder of the insert region (110) may be sequenced using labeled nucleotides and / or labeled multivalent molecules. In some embodiments, in step (e), the full length of the insert region (110) may be sequenced using labeled nucleotides and / or labeled multivalent molecules.
[0338] In some embodiments, the method for sequencing further comprises step (f): displacing the plurality of insert sequence extension products hybridized to the immobilized concatemer molecules by performing a primer extension reaction using a strand-displacing polymerase and a plurality of nucleotides to generate second strand extension products hybridized to the immobilized concatemer molecules comprising the immobilized capture primer.
[0339] In some embodiments, the method for sequencing further comprises step (g): removing the immobilized concatemer molecules by generating abasic sites within the immobilized concatemer molecules at uracil sites and generating gaps at the abasic sites, thereby retaining the second-strand retention products generated in step (f) while generating concatemer molecules containing gaps, and each second-strand retention product is retained by hybridization to an immobilized capture primer. In some embodiments, a pairwise turn is achieved by performing steps (f) and (g).
[0340] In some embodiments, the method for sequencing further comprises step (h): hybridizing the retained second strand extension products with a third plurality of soluble sequencing primers that hybridize to universal binding sequences for reverse sequencing primers (130) and sequencing at least a portion of the insert region (110).
[0341] In some embodiments, the method for sequencing further comprises (i) assigning the sequence of the insert region (110) to (ii) a sample index sequence, thereby identifying the insert region as having been obtained from the first source. In some embodiments, the assigning may occur after steps (c), (e), and / or (h).
[0342] In some embodiments, removing the plurality of sequencing extension products in steps (b) and (d) can be performed using a denaturing reagent comprising an SSC (e.g., saline-sodium citrate) buffer, with or without formamide, at a temperature that promotes nucleic acid denaturation, e.g., 50-90°C. In some embodiments, removing the plurality of sequencing extension products in steps (b) and (d) can be performed using a dehybridization reagent at a temperature that promotes nucleic acid denaturation, e.g., 50-90°C. In some embodiments, the dehybridization reagent comprises a pH buffer, a reducing agent, a monovalent salt, and a crowding agent. In some embodiments, the dehybridization reagent further comprises a chaotropic agent.
[0343] In some embodiments, the sequencing in steps (a), (c), and (e) comprises performing any of the sequencing methods described herein using a sequencing polymerase and detectably labeled nucleotide analogs. In some embodiments, the sequencing in steps (a), (c), and (e) comprises performing any of the two-step sequencing methods described herein using a sequencing polymerase, detectably labeled nucleotide analogs, and labeled or unlabeled nucleotide analogs. In some embodiments, the sequencing in steps (a), (c), and (e) comprises performing any of the binding sequencing methods or sequencing methods described herein using phosphate-chain labeled nucleotides.
[0344] In some embodiments, the density of the plurality of concatemeric molecules immobilized on the support is greater than 1 mm 2 Approximately 10 per 2 ~10 15 (e.g., 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8, 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 In some embodiments, the plurality of concatemer molecules are immobilized at random locations on the support. In some embodiments, the plurality of concatemer molecules are immobilized in a predetermined pattern on the support.
[0345] Fourth embodiment: Sequencing order In some embodiments, the covalently closed circular molecule (600) comprises: (i) a first subregion (411) comprising a universal binding sequence for an immobilized surface pinning primer; (ii) a second subregion (412) comprising a short random sequence (NNN) and a sample index sequence; (iii) a third subregion (413) comprising a universal binding sequence for an immobilized surface capture primer; (iv) a universal binding sequence for a reverse sequencing primer (130); (v) an insert sequence (1...
Claims
1. A method for forming a plurality of library-sprint composites (500), comprising: a) providing a plurality of double-stranded splint adapters (200), each double-stranded splint adapter (200) in the plurality of double-stranded splint adapters (200) comprising a first splint strand (300) hybridized to a second splint strand (400), the double-stranded splint adapter comprising a double-stranded region and two adjacent single-stranded regions, the first splint strand comprising a first region (320), an internal region (310), and a second region (330), the internal region (310) of the first splint strand hybridizing to the second splint strand (400); b) hybridizing said plurality of double-stranded splint adapters to a plurality of single-stranded nucleic acid library molecules (100), each library molecule comprising a sequence of interest (110) flanked on a first side by a universal adapter sequence (120) for a forward sequencing primer binding site and on a second side by a universal adapter sequence (130) for a reverse sequencing primer binding site; hybridizing the plurality of library molecules, thereby circularizing the plurality of library molecules to form a plurality of library-splint complexes (500), each having two nicks.
2. 10. The method of claim 1, further comprising: (c) contacting the plurality of library-sprint complexes (500) with a ligase to generate a plurality of covalently closed circular library molecules (600).
3. The method of claim 1 or 2, wherein the hybridizing is performed under conditions suitable for hybridizing the first region (320) of the first splint strand to a universal adapter sequence (120) for the forward sequencing primer binding site of the library molecule.
4. The method of claim 3, wherein the conditions are suitable for hybridizing the second region (330) of the first splint strand to a universal adapter sequence (130) for the reverse sequencing primer binding site of the library molecule.
5. The method according to any one of claims 1 to 4, wherein the inner region (310) of the first splint chain (300) comprises at least three sub-regions.
6. The method of claim 5 , wherein the at least three sub-regions include sub-region (311), sub-region (312), and sub-region (313).
7. 7. The method of claim 6, wherein the subregion (311) comprises a universal adapter sequence for a surface capture primer binding site, a universal adapter sequence for a surface pinning primer binding site, a sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI).
8. 8. The method of claim 6 or 7, wherein the subregion (312) comprises a universal adapter sequence for a surface capture primer binding site, a universal adapter sequence for a surface pinning primer binding site, a sample index sequence, a short random sequence (NNN), and / or a unique molecular index (UMI).
9. 9. The method of any one of claims 6 to 8, wherein the subregion (313) comprises a universal adapter sequence for surface capture primer binding sites, a universal adapter sequence for surface pinning primer binding sites, a sample index sequence, a short random sequence (NNN) and / or a unique molecular index (UMI).
10. The sub-region (311), (312) or (313) comprises a sample index array, the sample index array comprising: A sample index sequence lacking short random sequences (NNN), Sample index sequences and short random sequences (NNN), a sample index flanked on both sides by nucleotide bases that can be converted to abasic bases; at least one nucleotide base that can be converted to an abasic base; at least one deoxyinosine, 18 carbon spacer, and / or - The method of any one of claims 6 to 9, comprising an 18-carbon spacer and at least one deoxyinosine.
11. 11. The method of any one of claims 6 to 10, further comprising: i) distributing the plurality of covalently closed circular library molecules (600) onto a support having a plurality of surface capture primers immobilized thereon under conditions suitable for hybridizing each covalently closed circular library molecule (600) to each immobilized surface capture primer, thereby immobilizing the plurality of covalently closed circular library molecules (600) to the support.
12. The method of claim 11 , wherein the support further comprises a plurality of surface pinning primers immobilized to the support.
13. ii) contacting a plurality of said immobilized covalently closed circular library molecules (600) with a plurality of strand displacement polymerases and a plurality of nucleotides under conditions suitable for performing a rolling circle amplification reaction on said support using said plurality of surface capture primers as immobilized amplification primers and said plurality of covalently closed circular library molecules (600) as template molecules; 12. The method of claim 11, further comprising thereby generating a plurality of nucleic acid concatemer molecules immobilized to said surface capture primers.
14. 14. The method of claim 13, further comprising: iii) sequencing the plurality of nucleic acid concatemer molecules immobilized on the surface capture primers, wherein the sequencing comprises: (i) sequencing the sample index; and (ii) sequencing the sequence of interest (110).
15. iv) further comprising sequencing the plurality of nucleic acid concatemer molecules immobilized on the surface capture primers, wherein the sequencing comprises: (A) sequencing one or more short random sequences NNN; (B) sequencing one or more sample indices; and (C) sequencing the sequence of interest (110).
16. The method of any one of claims 7 to 15, wherein the interior region (310) of the first splint strand (300) comprises one sample index.
17. The method of any one of claims 7 to 15, wherein the internal region (310) of the first splint strand (300) comprises one sample index and a short random sequence (NNN).
18. The method of any one of claims 1 to 17, wherein the library molecules (100) comprise one or more nucleotide sequences selected from Table 1.
19. The method of any one of claims 1 to 18, wherein the first splint strand (300) comprises one or more nucleotide sequences selected from Table 2.
20. 20. The method of any one of claims 1 to 19, wherein the second splint strand (400) comprises one or more nucleotide sequences selected from Table 3.