Dual-stranded splint adapter with universal long splint strand and method of use
The library-splint complex efficiently circularizes DNA molecules with unique index sequences, addressing inefficiencies in current methods and enhancing sequencing throughput and compatibility.
Patent Information
- Application Number
- JP2025514527
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-12
- Filing Date
- 2023-09-12
- Publication Date
- 2025-09-11
AI Technical Summary
Current methods for preparing covalently closed circular DNA libraries are inefficient and do not support the addition of unique index sequences, limiting the throughput and compatibility with next-generation sequencing technologies.
A library-splint complex is formed by hybridizing a single-stranded nucleic acid library molecule with a double-stranded splint adaptor, comprising a first and second splint strand, which circularizes the molecule and includes universal and index sequences, enabling efficient downstream amplification and sequencing.
The method allows for high-throughput sequencing by generating circular library molecules with unique index sequences, enhancing compatibility with next-generation sequencing technologies and improving sequencing efficiency.
Smart Images

Figure 2025530259000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 405,733, filed September 12, 2022, the contents of which are incorporated herein by reference in their entirety.
[0002] Electronic Sequence Listing Reference The contents of the Electronic Sequence Listing (ELEM-015_001WO_SeqList_ST26.xml, size: 88,981 bytes, and creation date: September 8, 2023) are incorporated herein by reference in their entirety.
[0003] The present disclosure is directed to methods of DNA sequencing and library preparation, including compositions comprising nucleic acid double-stranded splint adaptors, and methods for using double-stranded splint adaptors. The double-stranded splint adaptors can hybridize to a portion of a library molecule to form a library-splint complex having a nick, which can be ligated to form a covalently closed circular molecule that can be subjected to downstream amplification and sequencing workflows. [Background technology]
[0004] The present disclosure relates to preparing libraries of covalently closed circular molecules using double-stranded splint adapters and sequencing the libraries prepared using the compositions and methods described herein. Improvements in next-generation sequencing technologies have greatly increased sequencing speed and data output, resulting in high sample throughput for current sequencing platforms. Efficient preparation of closed circular library molecules containing target sequences is important for downstream amplification and sequencing workflows. Another aspect of increasing sequencing throughput is the addition of unique index sequences to DNA fragments during library preparation, which allows multiple libraries to be pooled and sequenced simultaneously during each sequencing run. Therefore, there is a need for alternative methods for generating and sequencing circular library molecules containing target sequences and unique index sequences that are compatible with downstream next-generation sequencing technologies. Compositions, methods, and kits that address this need are provided herein. Summary of the Invention
[0005] The present disclosure provides a library-sprint complex (500) comprising: (i) a single-stranded nucleic acid library molecule (100) comprising a sequence of interest (110) flanked on one side by at least a first left universal adaptor sequence (120) and on the other side by at least a first right universal adaptor sequence (130); and (ii) a double-stranded splint adaptor (200) comprising a first splint strand (300) and a second splint strand (400), wherein the double-stranded splint adaptor (200) comprises a double-stranded region and two single-stranded regions, one on each side of the double-stranded region, and the first splint strand comprises a first region (320), an internal region (310), and a second region (330), wherein the internal region (310) of the first splint strand hybridizes to the second splint strand (400), the first region (320) of the first splint strand hybridizes to at least a first left universal adapter sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least a first right universal sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500).
[0006] In some embodiments of the library-sprint complex (500) of the present disclosure, the nucleic acid library molecule (100) further comprises a second left universal adaptor sequence (140). In some embodiments, the second left universal adaptor sequence (140) is located at least between the first left universal adaptor sequence (120) and the sequence of interest (110). In some embodiments, the nucleic acid library molecule (100) further comprises a second right universal adaptor sequence (150). In some embodiments, the second right universal adaptor sequence (150) is located between the sequence of interest (110) and at least the first right universal adaptor sequence (130). In some embodiments, the nucleic acid library molecule (100) further comprises a first left index sequence (160). In some embodiments, the first left index sequence (160) is located at least between the first left universal adaptor sequence (120) and the sequence of interest (110). In some embodiments, the nucleic acid library molecule (100) further comprises a first right index sequence (170). In some embodiments, the first right index sequence (170) is located between the second right universal adaptor sequence (150) and at least the first right universal adaptor sequence (130). In some embodiments, the nucleic acid library molecule (100) further comprises a first left unique identifier sequence (180). In some embodiments, the first left unique identifier sequence (180) is located between at least the first left universal adaptor sequence (120) and the first left index sequence (160). In some embodiments, the nucleic acid library molecule (100) further comprises a first right unique identifier sequence (190). In some embodiments, the first right unique identifier sequence (190) is located between the first right index sequence (170) and at least the first right universal adaptor sequence (130).
[0007] In some embodiments of the library-sprint complex (500) of the present disclosure, the nucleic acid library molecule (100) further comprises any one or any combination of two or more of: (i) a second left universal adaptor sequence (140), (ii) a second right universal adaptor sequence (150), (iii) a first left index sequence (160), (iv) a first right index sequence (170), (v) a first left unique identifier sequence (180), and / or (vi) a first right unique identifier sequence (190).
[0008] In some embodiments of the library-sprint complex (500) of the present disclosure, the first left universal adaptor sequence (120) and / or the second left universal adaptor sequence (140) comprise (i) a universal binding sequence for the forward sequencing primer, (ii) a universal binding sequence for the reverse sequencing primer, (iii) a universal binding sequence for the first surface primer, (iv) a universal binding sequence for the second surface primer, (v) a universal binding sequence for the forward amplification primer, (vi) a universal binding sequence for the reverse amplification primer, and / or (vii) a universal binding sequence for the compaction oligonucleotide. In some embodiments, the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) comprise (i) a universal binding sequence for the forward sequencing primer, (ii) a universal binding sequence for the reverse sequencing primer, (iii) a universal binding sequence for the first surface primer, (iv) a universal binding sequence for the second surface primer, (v) a universal binding sequence for the forward amplification primer, (vi) a universal binding sequence for the reverse amplification primer, and / or (vii) a universal binding sequence for the compaction oligonucleotide.
[0009] In some embodiments of the library-splint complex (500) of the present disclosure, the second splint strand (400) comprises at least two subregions, the first subregion comprising a universal binding sequence for the third surface primer and the second subregion comprising a universal binding sequence for the fourth surface primer, wherein the first and second subregions do not hybridize to the first and second surface primers or exhibit very little hybridization to the first and second surface primers. In some embodiments, the second splint strand (400) comprises an optional third subregion, wherein the third subregion comprises a sample index sequence having 5-20 bases and / or a unique identification sequence having 2-10 or more bases. In some embodiments, the unique identification sequence comprises a random sequence.
[0010] In some embodiments of the library-splint complex (500) of the present disclosure, the first splint strand (300) comprises an internal region (310) comprising at least two subregions, wherein the fourth subregion comprises a universal binding sequence for a third surface primer, the fourth subregion hybridizes to the first subregion of the second splint strand (400), the fifth subregion comprises a universal binding sequence for the fourth surface primer, the fifth subregion hybridizes to the second subregion of the second splint strand (400), and the fourth and fifth subregions do not hybridize to the first and second surface primers or show very little hybridization to the first and second surface primers. In some embodiments, the first splint strand (300) comprises an internal region (310) further comprising a sixth subregion, the sixth subregion comprising a sample index sequence having 5 to 20 bases and / or a unique identification sequence having 2 to 10 or more bases, the sixth subregion hybridizing to the third subregion of the second splint strand (400). In some embodiments, the unique identification sequence comprises a random sequence.
[0011] The present disclosure provides a library-sprint complex (500) comprising: (a) components arranged in 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer (120); (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; (iii) a sequence of interest (110); and (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer. (b) a first splint strand (300) including components arranged in 5' to 3' order: a first region (320), an internal region (310), and a second region (330); and (c) a first subregion arranged in 3' to 5' order: a first subregion having a universal binding sequence for a third surface primer; and a fourth surface primer. and a second splint strand (400) comprising a second subregion having a universal binding sequence for the first surface primer (120), wherein the first splint strand (300) hybridizes to a portion of the library molecule (100), thereby circularizing the library molecule to produce a library-splint complex (500), whereby a first region (320) of the first splint strand hybridizes to a binding sequence for the first surface primer (120), and a third region (330) of the first splint strand hybridizes to a binding sequence for the second surface primer (120). The library-sprint complex (500) comprises a first nick between the 5' end of the library molecule and the 3' end of the second splint strand, and the second splint strand (400) hybridizes to the binding sequence for the primer (130), the second splint strand (400) hybridizes to an internal region (310) of the first splint strand (300), and the library-sprint complex (500) comprises a second nick between the 5' end of the second splint strand and the 3' end of the library molecule.
[0012] In some embodiments of the library-sprint complex (500) of the present disclosure, the first and second nicks are enzymatically ligatable.
[0013] The present disclosure provides a plurality of library-sprint complexes, including the library-sprint complex (500) of the present disclosure, wherein the target sequences (110) of each library-sprint complex in the plurality of library-sprint complexes include the same target sequence or different target sequences.
[0014] The present disclosure provides a method for generating a library-sprint complex of the present disclosure, the method comprising: (a) providing a plurality of single-stranded nucleic acid library molecules (100); (b) providing a plurality of double-stranded splint adapters (200), a first splint strand (300), and a second splint strand (400); and (c) contacting the plurality of single-stranded nucleic acid library molecules with a plurality of double-stranded splint adapters under conditions sufficient for the terminus of the first splint strand to hybridize to the terminus of the library molecule, thereby generating a plurality of library-sprint complexes.
[0015] The present disclosure provides a method for producing a library-sprint complex of the present disclosure, the method comprising: (a) providing a plurality of single-stranded nucleic acid library molecules, a plurality of first splint strands, and a plurality of second splint strands; and (b) contacting the plurality of single-stranded nucleic acid library molecules with the plurality of first splint strands and second splint strands under conditions sufficient for the second splint strand to hybridize to the first splint strand and sufficient for the end of the first splint strand to hybridize to the end of the library molecule, thereby producing a plurality of library-sprint complexes.
[0016] The present disclosure provides a method of sequencing a plurality of concatemeric template molecules, the method including: (a) providing a plurality of library-sprint complexes of the present disclosure; (b) performing rolling circle amplification on the plurality of library-sprint complexes to generate a plurality of concatemeric template molecules; and (c) sequencing the plurality of concatemeric template molecules.
[0017] The present disclosure provides a kit comprising a plurality of double-stranded splint adaptors of the present disclosure.
[0018] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief explanation of the drawings]
[0019] [Figure 1] Schematic diagram showing an exemplary linear, single-stranded library molecule (100) hybridized to a double-stranded splint molecule (200, also referred to as a "ds-sprint adaptor"), thereby circularizing the library molecule and forming a library-sprint complex (500) with two nicks. The library molecule (100) contains a sequence of interest (insert (110)) flanked on one side by a first left universal adaptor sequence (120) and on the other side by a first right universal adaptor sequence (130). The double-stranded splint molecule contains a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint strand hybridizes to the second splint strand (400). "P" indicates a 5'-terminal phosphate group. [Figure 2]1, with further details regarding the embodiment of the internal region (310) of the first splint strand (300) and the second splint strand (400). The second splint strand (400) may include two subregions, the first subregion containing a universal binding sequence for the third surface primer and the second subregion containing a universal binding sequence for the fourth surface primer. The internal region (310) of the first splint strand (300) may include two subregions, the fourth subregion hybridizing to the first subregion of the second splint strand (400) and the fifth subregion hybridizing to the second subregion of the second splint strand (400). [Figure 3] 1 with further details regarding an embodiment of the internal region (310) of the first splint strand (300) and the second splint strand (400). The second splint strand (400) may include three subregions, the first subregion including a universal binding sequence for the third surface primer, the second subregion including a universal binding sequence for the fourth surface primer, and the third subregion including a sample index sequence having 5-20 bases and / or a unique identifier sequence (e.g., NN) having 2-10 or more bases. The internal region (310) of the first splint strand (300) may include three subregions, the fourth subregion hybridizing to the first subregion of the second splint strand (400), the fifth subregion hybridizing to the second subregion of the second splint strand (400), and the sixth subregion hybridizing to the third subregion of the second splint strand (400). [Figure 4]1 is a schematic diagram showing an exemplary linear, single-stranded library molecule (100) hybridized to a double-stranded splint molecule (200), thereby circularizing the library molecule and forming a library-splint complex (500) with two nicks. The library molecule (100) comprises a sequence of interest (insert, 110) flanked on one side by a first left universal adaptor sequence (120) and a second left universal adaptor sequence (140), and on the other side by a second right universal adaptor sequence (150) and a first right universal adaptor sequence (130). The double-stranded splint molecule comprises a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. An internal region (310) of the first splint strand hybridizes to the second splint strand (400). [Figure 5]1 is a schematic diagram showing an exemplary linear, single-stranded library molecule (100) hybridizing to a double-stranded splint molecule (200), thereby circularizing the library molecule and forming a library-splint complex (500) with two nicks. The library molecule (100) comprises a first left universal adaptor sequence (120), a first left unique identifier sequence (180), a first left index sequence (160), a second left universal adaptor sequence (140), a sequence of interest (110), a second right universal adaptor sequence (150), a first right index sequence (170), and a first right universal adaptor sequence (130). The double-stranded splint molecule comprises a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. An internal region (310) of the first splint strand hybridizes to the second splint strand (400). [Figure 6]1 is a schematic diagram showing an exemplary linear, single-stranded library molecule (100) hybridizing to a double-stranded splint molecule (200), thereby circularizing the library molecule and forming a library-splint complex (500) with two nicks. The library molecule (100) comprises a first left universal adaptor sequence (120), a first left index sequence (160), a second left universal adaptor sequence (140), a sequence of interest (insert, 110), a second right universal adaptor sequence (150), a first right index sequence (170), a first right unique identifier sequence (190), and a first right universal adaptor sequence (130). The double-stranded splint molecule comprises a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. An internal region (310) of the first splint strand hybridizes to the second splint strand (400). [Figure 7A]1 is a schematic diagram showing an exemplary linear, single-stranded library molecule (100) hybridized to a double-stranded splint molecule (200), thereby circularizing the library molecule and forming a library-splint complex (500) with two nicks. The library molecule (100) includes a first left universal adaptor sequence (120), a first left index sequence (160), a second left universal adaptor sequence (140), a sequence of interest (also referred to as an insert, 110), a second right universal adaptor sequence (150), a first right index sequence (170), and a first right universal adaptor sequence (130). The double-stranded splint molecule includes a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint strand hybridizes to the second splint strand (400). The second splint strand (400) comprises two subregions: the first subregion comprises a universal binding sequence for a fourth surface primer (e.g., a surface pinning primer), and the second subregion comprises a universal binding sequence for a third surface primer (e.g., a surface capture primer). A random sequence (e.g., NNN) is inserted within the first subregion, or the random sequence replaces a region within the first subregion. The random sequence may comprise 3 to 20 bases. In some embodiments, the random sequence further comprises a sample index sequence. The random sequence may be sequenced, and the sequence information may be used for polony mapping and / or template registration. The random sequence in the second splint strand (400) is represented by a horizontal bar patterned region. The internal region (310) of the first splint strand (300) includes two subregions: the fourth subregion hybridizes to the first subregion of the second splint strand (400), and the fifth subregion hybridizes to the second subregion of the second splint strand (400).A predetermined sequence is inserted into the fourth subregion, or the predetermined sequence replaces a region within the fourth subregion. The predetermined sequence within the first splint strand (300) contains 3 to 20 bases and has a sequence that may or may not be complementary to the random sequence within the first subregion of the second splint strand (400). The predetermined sequence is represented by an unpatterned white region. [Figure 7B]1 is a schematic diagram showing an exemplary linear, single-stranded library molecule (100) hybridized to a double-stranded splint molecule (200), thereby circularizing the library molecule and forming a library-splint complex (500) with two nicks. The library molecule (100) comprises a first left universal adaptor sequence (120), a first left index sequence (160), a second left universal adaptor sequence (140), a sequence of interest (insert, 110), a second right universal adaptor sequence (150), a first right index sequence (170), and a first right universal adaptor sequence (130). The double-stranded splint molecule comprises a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint strand hybridizes to the second splint strand (400). The second splint strand (400) may comprise two subregions: the first subregion comprises a universal binding sequence for a fourth surface primer (e.g., a surface pinning primer), and the second subregion comprises a universal binding sequence for a third surface primer (e.g., a surface capture primer). A random sequence (e.g., NNN) is added to the 3' end of the first subregion. The random sequence may comprise 3 to 20 bases. In some embodiments, the random sequence further comprises a sample index sequence. The random sequence may be sequenced, and the sequence information may be used for polony mapping and / or template registration. The random sequence in the second splint strand (400) is represented by a horizontal bar patterned region. The interior region (310) of the first splint strand (300) may include two subregions, a fourth subregion that hybridizes to the first subregion of the second splint strand (400) and a fifth subregion that hybridizes to the second subregion of the second splint strand (400).A predetermined sequence may be added to the 5' end of the fourth subregion. The predetermined sequence in the first splint strand (300) may contain 3 to 20 bases and may or may not be complementary to the random sequence in the first subregion of the second splint strand (400). The predetermined sequence is represented by the unpatterned white area. [Figure 8]Schematic diagram showing an exemplary linear, single-stranded library molecule (100) hybridizing to a double-stranded splint molecule (200), thereby circularizing the library molecule to form a library-splint complex (500) with two nicks. The library molecule (100) comprises a first added left universal adaptor sequence (121), a first left universal adaptor sequence (120), a first left junction adaptor sequence (125), a first left index sequence (160), a second left junction adaptor sequence (165), a second left universal adaptor sequence (140), a third left junction adaptor sequence (145), a sequence of interest (insert, 110), a third right junction adaptor sequence (155), a second right universal adaptor sequence (150), a second right junction adaptor sequence (175), a first right index sequence (170), a first right junction adaptor sequence (135), a first right universal adaptor sequence (130), and a first added right universal adaptor sequence (131). The double-stranded splint molecule comprises a first splint strand (long strand (300)) hybridized to a second splint strand (short strand (400)). The first splint strand comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint strand hybridizes to the second splint strand (400). For simplicity, the library-sprint complex (500) does not show either the junction adaptor sequences or the added universal adaptor sequences. Those skilled in the art will recognize that the linear library molecule (100) can comprise any combination of any one or more of the junction adaptors, with or without one or both of the added universal adaptor sequences. Those skilled in the art will recognize that the library-sprint complex (500) can include any one or any combination of two or more of the junction adaptors, with or without one or both of the added universal adaptor sequences present in the library molecule (100). [Figure 9] Three schematic diagrams of exemplary covalently closed circular library molecules (600) are shown, each hybridized to a first splint strand (300). The top schematic diagram shows a covalently closed circular library molecule (600) having a sequence of interest (insert, 110), a first right universal adaptor sequence (130), a second splint strand sequence (400), and a first left universal adaptor sequence (120). The middle schematic diagram shows a covalently closed circular library molecule (600) having a sequence of interest (110), a second right universal adaptor sequence (150), a first right universal adaptor sequence (130), a second splint strand sequence (400), a first left universal adaptor sequence (120), and a second left universal adaptor sequence (140). The bottom schematic shows a covalently closed circular library molecule (600) having a sequence of interest (110), a second right universal adaptor sequence (150), a first right index sequence (170), a first right universal adaptor sequence (130), a second splint strand sequence (400), a first left universal adaptor sequence (120), a first left unique identifier sequence (180), a first left index sequence (160), and a second left universal adaptor sequence (140). [Figure 10] 1 is a schematic diagram showing an exemplary library-sprint complex (500) that is subjected to a ligation reaction to seal the nick and form a covalently closed circular library molecule (600) that is hybridized to a first splint strand (300), which is used as an amplification primer to perform a rolling circle amplification reaction. The dotted lines represent the nascent extension products. [Figure 11A]11A shows the nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (300) and a second splint strand (400). The exemplary first splint strand includes a first region (320; SEQ ID NO:4), a second region (330; SEQ ID NO:5), and an internal region (310) having a fourth subregion (SEQ ID NO:6) and a fifth subregion (SEQ ID NO:7). The exemplary second splint strand (400) includes a first subregion (SEQ ID NO:1) and a second subregion (SEQ ID NO:2). In FIG. 11A, the second splint strand (top strand, 400) has the sequence of SEQ ID NO:202, and the first splint strand (bottom strand, 300) has the sequence of SEQ ID NO:199. [Figure 11B] The nucleotide sequences of exemplary first splint strands (300), each having a cleavage sequence at the 5' end of the first region (320), are shown. The cleavage sequences of the first region (320) are different from SEQ ID NO:4 (see FIG. 11A). In the exemplary cleaved first splint strand shown in FIG. 11B, the fourth subregion comprises the sequence of SEQ ID NO:6, the fifth subregion comprises the sequence of SEQ ID NO:7, and the second region (330) comprises the sequence of SEQ ID NO:5. The cleaved first strand (300) can hybridize with the second splint strand (400) shown in FIG. 11A, which comprises the first subregion (SEQ ID NO:1) and the second subregion (SEQ ID NO:2). The full-length sequences in FIG. 11B, from top to bottom, are SEQ ID NO:217, SEQ ID NO:218, SEQ ID NO:219, SEQ ID NO:220, and SEQ ID NO:221. [Figure 11C]11C shows the nucleotide sequences of exemplary first splint strands (300), each having a mismatch sequence within the first region (320). The mismatch sequence is shown in lowercase and underlined. The mismatch sequence is different from SEQ ID NO: 4 (see FIG. 11A). The first region (320) can hybridize with the first left universal adaptor sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatch sequence in the first region (320). In the exemplary mismatched first splint strand shown in FIG. 11C, the fourth subregion comprises the sequence of SEQ ID NO: 6, the fifth subregion comprises the sequence of SEQ ID NO: 7, and the second region (330) comprises the sequence of SEQ ID NO: 5. The mismatched first strand (300) can hybridize to the second splint strand (400) shown in Figure 11A, which includes a first subregion (SEQ ID NO: 1) and a second subregion (SEQ ID NO: 2). The full-length sequences in Figure 11C, from top to bottom, are SEQ ID NO: 222, SEQ ID NO: 223, SEQ ID NO: 224, SEQ ID NO: 225, SEQ ID NO: 226, and SEQ ID NO: 227. [Figure 11D]The nucleotide sequence of an exemplary first splint strand (300) is shown, which has either an abasic site or a uracil. The first splint strand shown at the top contains abasic sites in the fourth and fifth subregions. The abasic sites are represented by solid black bars. The first region (320) of the top first splint strand can hybridize with the first left universal adaptor sequence (120) of the library molecule (100). The second region (330) of the top first splint strand can hybridize with the first right universal adaptor sequence (130) of the library molecule (100). The first splint strand shown at the bottom contains at least one uracil in the first region (320), the second region (330), and the internal region (310). The uracil is underlined. The first region (320) of the bottom first splint strand can hybridize with the first left universal adaptor sequence (120) of the library molecule (100). The second region (330) of the bottom first splint strand can hybridize with the first right universal adaptor sequence (130) of the library molecule (100). Top strand: SEQ ID NO:228-abasic site-SEQ ID NO:229-abasic site-SEQ ID NO:230, bottom strand: SEQ ID NO:231. [Figure 12A]1 shows the nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (300, bottom strand) and a second splint strand (400, top strand). The exemplary first splint strand includes a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint strand (400) includes a first subregion and a second subregion. The internal region (310) of the first splint strand (300) includes two subregions, the fourth subregion hybridizing to the first subregion of the second splint strand (400), and the fifth subregion hybridizing to the second subregion of the second splint strand (400). A 3-mer random sequence (e.g., NNN) is inserted into the sequence of the first subregion of the second splint strand (400). A three-base predetermined sequence (e.g., 5'-gcg-3') is inserted into the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule with a bubble at the position of the three-mer random sequence (e.g., NNN) within the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 232, bottom strand: SEQ ID NO: 233. [Figure 12B]The nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (300, bottom strand) and a second splint strand (400, top strand) is shown. The exemplary first splint strand includes a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint strand (400) includes a first subregion and a second subregion. The internal region (310) of the first splint strand (300) includes two subregions, the fourth subregion hybridizing to the first subregion of the second splint strand (400), and the fifth subregion hybridizing to the second subregion of the second splint strand (400). A 4-mer random sequence (e.g., NNNN) is inserted into the sequence of the first subregion of the second splint strand (400). A four-base predetermined sequence (e.g., 5'-gtcg-3') is inserted into the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule with a bubble at the position of the four-mer random sequence (e.g., NNNN) within the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 234, bottom strand: SEQ ID NO: 235. [Figure 13A]1 shows the nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (300, bottom strand) and a second splint strand (400, top strand). The exemplary first splint strand includes a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint strand (400) includes a first subregion and a second subregion. The internal region (310) of the first splint strand (300) includes two subregions, the fourth subregion hybridizing to the first subregion of the second splint strand (400), and the fifth subregion hybridizing to the second subregion of the second splint strand (400). A 3-mer random sequence (e.g., NNN) replaces a portion of the sequence of the first subregion of the second splint strand (400). A three-base predetermined sequence (e.g., 5'-tgc-3') replaces a portion of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule with a bubble at the position of the three-mer random sequence (e.g., NNN) within the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 236, bottom strand: SEQ ID NO: 237. [Figure 13B]1 shows the nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (bottom strand, 300) and a second splint strand (top strand, 400). The exemplary first splint strand includes a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint strand (400) includes a first subregion and a second subregion. The internal region (310) of the first splint strand (300) includes two subregions, the fourth subregion hybridizing to the first subregion of the second splint strand (400), and the fifth subregion hybridizing to the second subregion of the second splint strand (400). A 4-mer random sequence (e.g., NNNN) replaces a portion of the sequence of the first subregion of the second splint strand (400). A four-base predetermined sequence (e.g., 5'-gtgc-3') replaces a portion of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule with a bubble at the position of the four-mer random sequence (e.g., NNN) within the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 238, bottom strand: SEQ ID NO: 239. [Figure 14A]1 shows the nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (bottom strand, 300) and a second splint strand (top strand, 400). The exemplary first splint strand includes a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint strand (400) includes a first subregion and a second subregion. The internal region (310) of the first splint strand (300) includes two subregions: the fourth subregion hybridizes to the first subregion of the second splint strand (400), and the fifth subregion hybridizes to the second subregion of the second splint strand (400). A 3-mer random sequence (e.g., NNN) and an index sequence (e.g., 5'-cactattcc-3') are added to the 3' end of the first subregion of the second splint strand (400). A predetermined sequence equal in length to the 3-mer random sequence and the index sequence (e.g., 5'-ggaatagtgaca-3') is added to the 5' end of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule with a bubble or mismatched end at the position of the 3-mer random sequence (e.g., NNN) and the index sequence within the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 240, bottom strand: SEQ ID NO: 241. [Figure 14B]1 shows the nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (bottom strand, 300) and a second splint strand (top strand, 400). The exemplary first splint strand includes a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint strand (400) includes a first subregion and a second subregion. The internal region (310) of the first splint strand (300) includes two subregions: the fourth subregion hybridizes to the first subregion of the second splint strand (400), and the fifth subregion hybridizes to the second subregion of the second splint strand (400). A 4-mer random sequence (e.g., NNNN) and an index sequence (e.g., 5'-cactattcc-3') are added to the 3' end of the first subregion of the second splint strand (400). A predetermined sequence equal in length to the 4-mer random sequence and the index sequence (e.g., 5'-ggaatagtgacag-3') is added to the 5' end of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule with a bubble or mismatched end at the position of the 4-mer random sequence (e.g., NNNN) and the index sequence within the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 242, bottom strand: SEQ ID NO: 243. [Figure 15A] 1 is a graph showing sequencing quality scores for C base calls of first-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail mixture of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 15B] 1 is a graph showing sequencing quality scores for A base calls of first-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail mixture of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 15C] 1 is a graph showing sequencing quality scores for G base calls of first-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 15D]1 is a graph showing sequencing quality scores for T base calls of first-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail mixture of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 16A] 1 is a graph showing sequencing quality scores for C base calls of second-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail mixture of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 16B]1 is a graph showing sequencing quality scores for A base calls of second-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail mixture of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 16C] 1 is a graph showing sequencing quality scores for G base calls of second-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail mixture of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 16D]1 is a graph showing sequencing quality scores for T base calls of second-strand concatemeric template molecules generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see FIG. 5 or 6, but without the unique identifier sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of heat, NaOH, HEPES buffer, or a cocktail mixture of enzymes that can generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score. [Figure 17]
[0033] Figure 1 is a set of three graphs showing sequencing quality scores for A, G, C, and T base calls of first-strand concatemeric template molecules (Read 1) generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see Figures 5 or 6, but without the unique identifier sequences (180) and (190)) with ligase enzyme inactivation and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of a ligase high heat-kill control (left panel), a ligase low heat-kill control (middle panel), and a ligase NaOH inactivation (right panel). The x-axis is the number of sequencing cycles. The y-axis is the quality score. [Figure 18]
[0033] Figure 1 is a set of three graphs showing sequencing quality scores for A, G, C, and T base calls of second-strand concatemeric template molecules (Read 2) generated by a workflow that includes circularizing linear library molecules using double-stranded splint adapters (e.g., see Figures 5 or 6, but without the unique identifier sequences (180) and (190)) with ligase enzyme inactivation, and performing on-support rolling circle amplification to generate concatemeric template molecules immobilized on a coated support. The library preparation workflow compared the effects of a ligase high heat-kill control (left panel), a ligase low heat-kill control (center panel), and a ligase NaOH inactivation (right panel). The x-axis is the number of sequencing cycles. The y-axis is the quality score. [Figure 19] 1 is a schematic diagram of an exemplary low-binding support comprising a glass substrate and alternating layers of a hydrophilic coating covalently or non-covalently adhered to the glass, and further comprising chemically reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides and circularization oligonucleotides). In alternative embodiments, the support can be made from any material, such as glass, plastic, or a polymeric material. [Figure 20] Schematic diagrams of various exemplary configurations of multivalent molecules. Left: Schematic diagram of a multivalent molecule with a starburst or helter-skelter configuration. Center: Schematic diagram of a multivalent molecule with a dendrimer configuration. Right: Schematic diagram of multiple multivalent molecules formed by reacting streptavidin with a 4-arm or 8-arm PEG-NHS bearing biotin and dNTPs. Nucleotide units are represented as "N," biotin is represented as "B," and streptavidin is represented as "SA." [Figure 21] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a generic core attached to multiple nucleotide arms. [Figure 22] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms. [Figure 23]1 shows a schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, the nucleotide arms comprising biotin, spacers, linkers, and nucleotide units. [Figure 24] FIG. 1 is a schematic diagram of an exemplary nucleotide arm comprising a core attachment moiety, a spacer, a linker, and a nucleotide unit. [Figure 25] Shown are the chemical structures of an exemplary spacer (top) and various exemplary linkers (bottom), including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker. [Figure 26] 1 shows the chemical structures of various exemplary linkers, including linkers 1-9. [Figure 27] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 28] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 29] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 30] The chemical structure of an exemplary biotinylated nucleotide arm is shown, in which the nucleotide unit is connected to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base. [Figure 31] 1 is a schematic diagram of a G-quadruplex (e.g., a G-quadruplex). [Figure 32] FIG. 1 is a schematic diagram of an exemplary intramolecular G-quadruplex structure. [Figure 33-1] 1 is Table 1 (six sheets) listing exemplary first left index array (160) and first right index array (170) arrangements. [Figure 33-2] 1 is Table 1 (six sheets) listing exemplary first left index array (160) and first right index array (170) arrangements. [Figure 33-3] 1 is Table 1 (six sheets) listing exemplary first left index array (160) and first right index array (170) arrangements. [Figure 33-4] 1 is Table 1 (six sheets) listing exemplary first left index array (160) and first right index array (170) arrangements. [Figure 33-5] 1 is Table 1 (six sheets) listing exemplary first left index array (160) and first right index array (170) arrangements. [Figure 33-6] 1 is Table 1 (six sheets) listing exemplary first left index array (160) and first right index array (170) arrangements. [Figure 34] 1 is a bar graph showing the average percent recovery of covalently closed circular library molecules using input DNA from various species as determined by qPCR. Lane 1: Haemophilus influenzae (38% GC), Lane 2: E. coli (51% GC), Lane 3: Rhodopseudomonas palustris (65% GC), Lane 4: PhiX, Lane 5: human, Lane 6: human exome, Lane 7: human mRNA. See Examples 1-3. [Figure 35] 1 is a bar graph showing the average polony density obtained by distributing covalently closed circular library molecules onto a support and performing on-support rolling circle amplification. The covalently closed circular library molecules were prepared from input DNA from various species. Lane 1: cell-free DNA, Lane 2: E. coli, Lane 3: human, Lane 4: metagenomic DNA, and Lane 5: PhiX. See Example 4. [Figure 36] 1 is a graph showing the nucleotide base diversity of a right index sequence (170) containing a 3-mer random sequence (NNN). The graph shows the nucleotide diversity of the 3-mer random sequence (NNN) of approximately 30% for A and T base calls and approximately 20% for C and G base calls. [Figure 37]1 is a graph showing the nucleotide base diversity of the left index sequence (160) lacking the 3-mer random sequence (NNN). The graph shows a nucleotide diversity of approximately 40% for A and T base calls, approximately 15% for C base calls, and approximately 5% for G base calls. DETAILED DESCRIPTION OF THE INVENTION
[0020] definition Throughout this application, various publications, patents, and / or patent applications are referenced. The disclosures of these publications, patents, and / or patent applications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this disclosure pertains.
[0021] The headings provided herein are not limitations of various aspects of the disclosure, which aspects can be understood by reference to the specification as a whole.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art. Generally, terms relating to molecular biology, nucleic acid chemistry, protein chemistry, genetics, microbiology, transgenic cell production, and hybridization techniques described herein are well known and commonly used in the art. The techniques and procedures described herein are generally performed according to conventional methods well known in the art and as described in various general and more specific references cited and discussed throughout the specification. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY 2000). See also Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992). The nomenclature used in connection with, and the experimental procedures and techniques described herein are well known and commonly used in the art.
[0023] Unless otherwise required by context herein, singular terms include plurals and plural terms include the singular. The singular forms "a," "an," and "the," as well as use of the singular form of any word, include plural references unless expressly and unambiguously limited to a reference to one.
[0024] The use of alternative terms (eg, "or") is understood to mean either one or both of the alternatives, or any combination thereof.
[0025] The term "and / or" as used herein should be understood to mean a specific disclosure that each of the specified features or components may or may not have the other. For example, when used herein in phrases such as "A and / or B," the term "and / or" is intended to include "A and B," "A or B," "A" (A alone), and "B" (B alone). In a similar manner, when used in phrases such as "A, B, and / or C," the term "and / or" is intended to encompass each of the following embodiments: "A, B, and C," "A, B, or C," "A or C," "A or B," "B or C," "A and B," "B and C," "A and C," "A" (A alone), "B" (B alone), and "C" (C alone).
[0026] As used in this specification and the appended claims, the terms "comprising," "including," "having," and "containing," and grammatical variations thereof, as used herein, are intended to be open-ended so that one or more items in a list do not exclude other items that may be substituted for or added to the listed items. Wherever embodiments are described herein with the term "comprising," it is understood that alternative, similar embodiments described with the terms "consisting of" and / or "consisting essentially of" are also provided.
[0027] As used herein, the term "about" or "approximately" refers to a value or composition that is within an acceptable error range for a particular value or composition, as determined by one of ordinary skill in the art. The acceptable error range depends in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" or "approximately" can mean within one or more standard deviations per practice in the art. Alternatively, "about" or "approximately" can mean a range of up to 10% (i.e., ±10%) or more, depending on the limitations of the measurement system. For example, about 5 mg can include any number between 4.5 mg and 5.5 mg. Furthermore, particularly with respect to biological systems or processes, the term can mean up to one order of magnitude or up to five times the value. When a particular value or composition is provided in this disclosure, unless otherwise specified, the meaning of "about" or "approximately" should be considered to be within an acceptable error range for that particular value or composition. Additionally, when ranges and / or subranges of values are provided, the ranges and / or subranges can include the endpoints of the ranges and / or subranges.
[0028] As used herein, the terms "peptide," "polypeptide," and "protein," as well as other related terms, are used interchangeably and refer to polymers of amino acids and are not limited to any particular length. Polypeptides can contain natural and unnatural amino acids. Polypeptides include recombinant or chemically synthesized forms. Polypeptides also include precursor molecules that have not yet undergone post-translational modifications such as proteolytic cleavage, ribosomal skipping cleavage, hydroxylation, methylation, lipidation, acetylation, sumoylation, ubiquitination, glycosylation, phosphorylation, and / or disulfide bond formation. These terms encompass natural and artificial proteins, protein fragments, and polypeptide analogs of protein sequences (such as muteins, variants, chimeric proteins, and fusion proteins), as well as proteins that are post-translationally or otherwise covalently or non-covalently modified.
[0029] The term "cellular biological sample" refers to a single cell, multiple cells, tissue, organ, organism, or section of any of these cellular biological samples. Cellular biological samples may be extracted from an organism (e.g., biopsy) or obtained from cell cultures growing in liquid or in culture dishes. Cellular biological samples include fresh samples, frozen samples, fresh frozen samples, or archived (e.g., formalin-fixed paraffin-embedded; FFPE) samples. Cellular biological samples may be embedded in wax, resin, epoxy, or agar. Cellular biological samples may be fixed in, for example, any one or any combination of two or more of acetone, ethanol, methanol, formaldehyde, paraformaldehyde-Triton®, or glutaraldehyde. Cellular biological samples may or may not be sectioned. Cellular biological samples may be stained, destained, or unstained.
[0030] Nucleic acids of interest (sometimes referred to herein as sequences of interest) can be extracted from cells or cellular biological samples using any of several techniques known to those skilled in the art. For example, a typical DNA extraction procedure includes (i) collecting a cell or tissue sample from which DNA is to be extracted, (ii) disrupting cell membranes (i.e., lysing cells) to release DNA and other cytoplasmic components, (iii) treating the lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA, and (iv) purifying the DNA from the supernatant to remove detergents, proteins, salts, or other reagents used during cell membrane lysis. A variety of suitable commercially available nucleic acid extraction and purification kits are consistent with the disclosure herein. Examples include, but are not limited to, the QIAamp kit (for isolation of genomic DNA from human samples) and the DNAeasy kit (for isolation of genomic DNA from animal or plant samples) from Qiagen (Germantown, MD), or the Maxwell® and ReliaPrep™ series of kits from Promega (Madison, WI). The nucleic acid of interest can be ribonucleic acid (RNA) or deoxyribonucleic acid (DNA), such as genomic DNA or complementary DNA (cDNA) reverse transcribed from RNA.
[0031] As used herein, the term "polymerase" and variations thereof include enzymes that contain a nucleotide (or nucleoside)-binding domain, and the polymerase can form a complex with a template nucleic acid and a complementary nucleotide. A polymerase can have one or more activities, including, but not limited to, base analog detection activity, DNA polymerization activity, reverse transcriptase activity, DNA binding, strand displacement activity, and nucleotide binding and recognition. A polymerase can be any enzyme that can catalyze the polymerization of nucleotides (including their analogs) into a nucleic acid strand. Typically, although not necessarily, such nucleotide polymerization can occur in a template-dependent manner. Typically, a polymerase contains one or more active sites at which nucleotide binding and / or catalysis of nucleotide polymerization can occur. In some embodiments, a polymerase includes other enzymatic activities, such as 3' to 5' exonuclease activity or 5' to 3' exonuclease activity. In some embodiments, a polymerase has strand displacement activity. Polymerases can include, but are not limited to, naturally occurring polymerases and any subunits and truncations thereof, mutant polymerases, variant polymerases, recombinant, fused, or otherwise engineered polymerases, chemically modified polymerases, synthetic molecules or assemblies, and any analogs, derivatives, or fragments thereof (e.g., catalytically active fragments) that retain the ability to catalyze nucleotide polymerization. The term polymerase includes catalytically inactive polymerases, catalytically active polymerases, reverse transcriptases, and other enzymes that contain a nucleotide-binding domain. In some embodiments, polymerases can be isolated from cells or produced using recombinant DNA technology or chemical synthesis methods. In some embodiments, polymerases can be expressed in prokaryotic, eukaryotic, viral, or phage organisms. In some embodiments, polymerases can be post-translationally modified proteins or fragments thereof. Polymerases can be derived from prokaryotic, eukaryotic, viral, or phage organisms.Polymerases include DNA-directed DNA polymerases and RNA-directed DNA polymerases.
[0032] The term "strand displacement" refers to the ability of a polymerase to locally separate strands of double-stranded nucleic acid and synthesize a new strand in a template-based manner. Strand-displacing polymerases displace a complementary strand from the template strand and catalyze new strand synthesis. Strand-displacing polymerases include mesophilic and thermophilic polymerases. Strand-displacing polymerases include wild-type enzymes and variants, including exonuclease-minus mutants, mutated versions, chimeric enzymes, and truncated enzymes. Examples of strand-displacing polymerases include phi29 DNA polymerase, large fragment of Bst DNA polymerase, large fragment of Bsu DNA polymerase (exo-), Bca DNA polymerase (exo-), Klenow fragment of E. coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV viral reverse transcriptase, Deep Vent® DNA polymerase, and KOD DNA polymerase. The phi29 DNA polymerase can be a wild-type phi29 DNA polymerase (e.g., MagniPhi™ from Expedeon), or a variant EquiPhi29™ DNA polymerase (e.g., from Thermo Fisher Scientific), or a chimeric QualiPhi™ DNA polymerase (e.g., from 4basebio).
[0033] As used herein, the terms "nucleic acid," "polynucleotide," and "oligonucleotide," as well as other related terms, are used interchangeably and refer to a polymer of nucleotides and are not limited to any particular length. Nucleic acids include recombinant or chemically synthesized forms. Nucleic acids can be isolated. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of DNA or RNA produced using nucleotide analogs (e.g., peptide nucleic acids and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids comprise polymers of nucleotides, where the nucleotides comprise natural or non-natural bases and / or sugars. Nucleic acids comprise naturally occurring internucleoside linkages, such as phosphodiester linkages. Nucleic acids comprise non-natural internucleoside linkages, including phosphorothioate, phosphorothiolate, or peptide nucleic acid (PNA) linkages. Nucleic acids can comprise one type of polynucleotide or a mixture of two or more different types of polynucleotides.
[0034] As used herein, the terms "operably linked" and "operably joined," or related terms, refer to the juxtaposition of components. Juxtaposed components can be covalently linked together. For example, two nucleic acid components can be enzymatically ligated together, where the bond linking the two components together comprises a phosphodiester bond. A first and second nucleic acid component can be ligated together, where the first nucleic acid component can confer a function to the second nucleic acid component. For example, the linkage between a primer binding sequence and a sequence of interest forms a nucleic acid library molecule having a portion capable of binding to the primer. In another example, a transgene (e.g., a nucleic acid encoding a polypeptide or nucleic acid sequence of interest) can be ligated into a vector, where the linkage allows for expression or function of the transgene sequence contained within the vector. In yet a further example, the transgene is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects expression of the transgene. In exemplary vectors, the vector comprises at least one host cell regulatory sequence, including a promoter sequence, an enhancer, a transcriptional and / or translational initiation sequence, a transcriptional and / or translational termination sequence, a polypeptide secretion signal sequence, etc., which are said to be operably linked. In the above examples, the host cell regulatory sequence controls the level, timing, and / or location of expression of the transgene.
[0035] The terms "linked," "joined," "attached," and "appended," and variations thereof, include any type of fusion, bond, adhesion, or association between any combination of compounds or molecules that is stable enough to withstand use in a particular procedure. Procedures can include, but are not limited to, nucleotide binding, nucleotide incorporation, deblocking (e.g., removal of chain-terminating moieties), washing, removal, flow, detection, imaging, and / or identification. Such linkages can include, for example, covalent bonds, ionic, hydrogen, dipole-dipole, hydrophilic, hydrophobic, or affinity bonds, bonds or associations involving van der Waals forces, and mechanical bonds. Such linkages can occur intramolecularly, such as, for example, linking the ends of a single- or double-stranded linear nucleic acid molecule together to form a circular molecule. Alternatively, such linkages can occur between different molecular combinations or between molecules and non-molecules, including, but not limited to, bonds between nucleic acid molecules and solid surfaces, bonds between proteins and detectable reporter moieties, bonds between nucleotides and detectable reporter moieties, and the like. Some examples of linkages can be found, for example, in Hermanson, G., "Bioconjugate Techniques", Second Edition (2008), Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998), Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998).
[0036] As used herein, the term "primer" and related terms refer to an oligonucleotide capable of hybridizing to a DNA and / or RNA polynucleotide template to form a duplex molecule. Primers may be single-stranded along their entire length or may have single-stranded and double-stranded portions. Primers contain natural nucleotides and / or nucleotide analogs. Primers can be recombinant nucleic acid molecules. Primers can be of any length but typically range from 4 to 50 nucleotides. Typical primers contain a 5' end and a 3' end. The 3' end of a primer can contain a 3' OH moiety that functions as a nucleotide polymerization initiation site in a polymerase-catalyzed primer extension reaction. Alternatively, the 3' end of a primer can lack a 3' OH moiety or contain a terminal 3' blocking group that inhibits nucleotide polymerization in a polymerase-catalyzed reaction. Any one or more nucleotides along the length of a primer can be labeled with a detectable reporter moiety. Primers can be in solution (e.g., soluble primers) or immobilized on a support (e.g., capture primers).
[0037] The terms "template nucleic acid," "template polynucleotide," "target nucleic acid," "target polynucleotide," "template strand," and other variations thereof, refer to a nucleic acid strand that serves as the base nucleic acid molecule for any of the amplification and / or sequencing methods described herein. A template nucleic acid may be single-stranded or double-stranded, or may have single-stranded or double-stranded portions. A template nucleic acid may be obtained from a naturally occurring source or recombinant form, or may be chemically synthesized to contain any type of nucleic acid analog. A template nucleic acid may be linear, concatemeric, circular, or in other forms. A template nucleic acid may encode a sequence of interest.
[0038] The term "adapter" and related terms refer to an oligonucleotide that can be operably linked (appended) to a target polynucleotide, where the adapter confers a function to the co-linked adapter-target molecule. Adapters include DNA, RNA, chimeric DNA / RNA, or analogs thereof. Adapters can contain at least one ribonucleoside residue. Adapters can be single-stranded or double-stranded, or can have single-stranded and / or double-stranded portions. Adapters can be configured to be linear, stem-loop, hairpin, or Y-shaped. Adapters can be any length, from 4 to 100 nucleotides or more. Adapters can have blunt ends, overhanging ends, or a combination of both. Overhanging ends include 5' overhangs and 3' overhanging ends. The 5' end of a single-stranded adapter, or one strand of a double-stranded adapter, can have a 5' phosphate group or lack a 5' phosphate group. The adapter may include a 5' tail that does not hybridize to the target polynucleotide (e.g., a tailed adapter), or the adapter may be tailless. At least a portion of the adapter may include a known and predetermined sequence. The adapter may include a sequence complementary to at least a portion of a primer, such as an amplification primer, a sequencing primer, or a capture primer (e.g., a soluble or immobilized capture primer). The adapter may include a random or degenerate sequence. The adapter may include at least one inosine residue. The adapter may include at least one phosphorothioate, phosphorothiolate, and / or phosphoramidate linkage. The adapter may include at least one barcode / index sequence, which may be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. The adapter may include at least one unique identification sequence (e.g., a molecular tag), which may be used to uniquely identify the nucleic acid molecule to which the adapter is attached.Exemplary, but non-limiting, unique identification sequences include 2 to 12 or more nucleotides having a known sequence. For example, the unique identification sequence may include a known random sequence, where the nucleotides at each position in the known random sequence are randomly selected from nucleotides having the bases A, G, C, T, or U. The adapter may include at least one restriction enzyme recognition sequence, where the at least one restriction enzyme recognition sequence includes any one or any combination of two or more selected from the group consisting of type I, type II, type III, type IV, type Hs, and type IIB.
[0039] The term "universal sequence" and related terms refer to a sequence in a nucleic acid molecule that is common between two or more polynucleotide molecules. For example, an adapter having a universal sequence can be operably linked to multiple polynucleotides, such that a population of co-linked molecules possesses the same universal adapter sequence. Examples of universal adapter sequences include amplification primer sequences, sequencing primer sequences, e.g., those compatible with commercially available sequencing platforms, or capture primer sequences (e.g., soluble or immobilized capture primers).
[0040] When used in reference to nucleic acid molecules, the terms "hybridize" or "hybridizing" or "hybridization," or other related terms, refer to hydrogen bonding between two different nucleic acids to form a double-stranded nucleic acid. Hybridization also includes hydrogen bonding between two different regions of a single nucleic acid molecule to form a self-hybridizing molecule having a double-stranded region. Hybridization can involve Watson-Crick or Hoogstein binding to form a duplex double-stranded nucleic acid or a double-stranded region within a nucleic acid molecule. A double-stranded nucleic acid, or two different regions of a single nucleic acid, can be fully complementary or partially complementary. Complementary nucleic acid strands need not hybridize to each other throughout their entire length. Complementary base pairing can be standard AT or CG base pairing or other forms of base-pairing interactions. A double-stranded nucleic acid can contain mismatched base-pairing nucleotides ("bubbles") that can form single-stranded regions within the duplex.
[0041] When used with respect to nucleic acids, the terms "extend," "extending," "extension," and other variations refer to the incorporation of one or more nucleotides into a nucleic acid molecule. Nucleotide incorporation involves the polymerization of one or more nucleotides onto the terminal 3'OH terminus of a nucleic acid chain, resulting in the elongation of the nucleic acid chain. Nucleotide incorporation can be performed with natural nucleotides and / or nucleotide analogs. Typically, although not necessarily, nucleotide incorporation occurs in a template-dependent manner. Any suitable method for extending a nucleic acid molecule may be used, including primer extension catalyzed by DNA polymerase or RNA polymerase.
[0042] The term "nucleotide" and related terms refer to a molecule comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and at least one phosphate group. Standard or non-standard nucleotides are consistent with the use of this term. In some embodiments, a nucleotide comprises a monophosphate, diphosphate, or triphosphate, or the corresponding phosphate analog. The term "nucleoside" refers to a molecule comprising an aromatic base and a sugar. Nucleotides and nucleosides can be unlabeled or labeled with a detectable reporter moiety.
[0043] Nucleotides (and nucleosides) typically contain a heterocyclic base containing a substituted or unsubstituted nitrogen-containing parent heteroaromatic ring, which are commonly found in nucleic acids, including naturally occurring, substituted, modified, or engineered variants, or analogs thereof. The base of a nucleotide (or nucleoside) is capable of forming Watson-Crick and / or Hoogstein hydrogen bonds with an appropriate complementary base. Exemplary bases are purines and pyrimidines, such as 2-aminopurine, 2,6-diaminopurine, adenine (A), ethenoadenine, N, N-dimethylformamide ... 6 -Δ 2 -Isopentenyladenine (6iA), N 6 -Δ 2 -Isopentenyl-2-methylthioadenine (2ms6iA), N 6 -Methyladenine, guanine (G), isoguanine, N 2 -dimethylguanine (dmG), 7-methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine, and O 6 -methylguanine; 7-deaza-purines, such as 7-deazaadenine (7-deaza-A) and 7-deazaguanine (7-deaza-G); pyrimidines, such as cytosine (C), 5-propynylcytosine, isocytosine, thymine (T), 4-thiothymine (4sT), 5,6-dihydrothymine, O 4Examples of bases include, but are not limited to, methylthymine, uracil (U), 4-thiouracil (4sU), and 5,6-dihydrouracil (dihydrouracil; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; nebularine; inosine; hydroxymethylcytosine; 5-methycytosine; base (Y); and methylated, glycosylated, and acylated base moieties. Additional exemplary bases can be found in Fasman, 1989, "Practical Handbook of Biochemistry and Molecular Biology," pp. 385-394, CRC Press, Boca Raton, Fla.
[0044] Nucleotides (and nucleosides) typically include a sugar moiety, e.g., a carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100:4319-48), an acyclic moiety (Martinez, et al., 1999 Nucleic Acids Research 27:1271-1274; Martinez, et al., 1997 Bioorganic & Medicinal Chemistry Letters vol. 7:3013-3016), and another sugar moiety (Joeng, et al., 1993 J. Med. Chem. 36:2627-2638; Kim, et al., 1993 J. Med. Chem. 36:30-7; Eschenmosser 1999 Science 284:2118-2124; and U.S. Pat. No. 5,558,991). Sugar moieties include, but are not limited to, ribosyl; 2'-deoxyribosyl; 3'-deoxyribosyl; 2',3'-dideoxyribosyl; 2',3'-didehydrodideoxyribosyl; 2'-alkoxyribosyl; 2'-azidoribosyl; 2'-aminoribosyl; 2'-fluororibosyl; 2'-mercaptoriboxyl; 2'-alkylthioribosyl; 3'-alkoxyribosyl; 3'-azidoribosyl; 3'-aminoribosyl; 3'-fluororibosyl; 3'-mercaptoriboxyl; 3'-alkylthioribosyl carbocyclic; acyclic, or other modified sugars.
[0045] A nucleotide may contain a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramido linkage. In some embodiments, the nucleotide is an analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene. The phosphorus atoms in the chain may contain substituted side groups, including O, S, or BH3. Alternatively, or in addition, the chain may contain phosphate groups substituted with analogs, including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoramidite groups.
[0046] The term "rolling circle amplification" generally refers to an amplification method using a circularized nucleic acid template molecule containing a target sequence of interest, an amplification primer binding sequence, and, optionally, one or more adapter sequences, such as a sequencing primer binding sequence and / or a sample index sequence. A rolling circle amplification reaction may be performed under isothermal amplification conditions and includes a circularized nucleic acid template molecule, an amplification primer, a strand-displacing polymerase, and a plurality of nucleotides to generate concatemers containing tandem repeat sequences of any adapter sequences present in the circularized template molecule and the original circularized nucleic acid template molecule. The concatemers can self-collapse to form nucleic acid nanoballs. The shape and size of the nanoballs can be further compacted by including a pair of inverted repeat sequences within the circular template molecule or by performing the rolling circle amplification reaction with one or more compacting oligonucleotides. One advantage of using rolling circle amplification to generate clonal amplicons for sequencing workflows is that the repeat copies of the target sequence in the nanoballs can be simultaneously sequenced, increasing signal intensity. In some embodiments, a rolling circle amplification reaction can be performed in the presence of multiple compacted oligonucleotides with at least four consecutive guanines. The rolling circle amplification reaction generates concatemers containing repeated copies of the universal binding sequence for the compacted oligonucleotides. At least one compacted oligonucleotide can form a G-quadruplex (Figure 31) and hybridize to the universal binding sequence for the compacted oligonucleotide, and the resulting concatemer can fold and form an intramolecular G-quadruplex structure (Figure 32). The concatemer can self-collapse to form a compact nanoball.The formation of G-quadruplexes and G-quadruplexes in nanoballs can increase the stability of the nanoballs, allowing them to retain their compact size and shape, which can withstand repeated flows of reagents to perform any of the sequencing workflows described herein.
[0047] The terms "amplify," "amplifying," "amplification," and other related terms, when used with respect to nucleic acids, include producing multiple copies of an original polynucleotide template molecule, wherein the copies contain sequences that are complementary to the template sequence and / or the copies contain sequences that are identical to the template sequence. In some embodiments, the copies contain sequences that are substantially identical to the template sequence and / or sequences that are substantially identical to the sequence that is complementary to the template sequence.
[0048] The terms "reporter moiety," "reporter moieties," or related terms refer to a compound that produces or can be caused to produce a detectable signal. Reporter moieties are often referred to as "labels." Any suitable reporter moiety can be used, and suitable reporter moieties include luminescence, photoluminescence, electroluminescence, bioluminescence, chemiluminescence, fluorescence, phosphorescence, chromophores, radioisotopes, electrochemistry, mass spectrometry, Raman, haptens, affinity tags, atoms, or enzymes. A reporter moiety produces a detectable signal that results from a chemical or physical change (e.g., heat, light, electricity, pH, salt concentration, enzymatic activity, or a proximity event). A proximity event involves two reporter moieties coming into close proximity with, associating with, or binding to each other. It is well known to those skilled in the art to select reporter moieties so that each absorbs excitation radiation and / or emits fluorescence at a wavelength distinguishable from other reporter moieties, allowing for the monitoring of the presence of different reporter moieties in the same or different reactions. Two or more different reporter moieties can be selected that have spectrally distinct emission profiles or that have minimal overlapping spectral emission profiles. A reporter moiety can be bound (e.g., operably linked) to a nucleotide, a nucleoside, a nucleic acid, an enzyme (e.g., a polymerase or reverse transcriptase), or a support (e.g., a surface).
[0049] The reporter moiety (or label) may comprise a fluorescent label or fluorophore. Exemplary fluorescent moieties that can function as fluorescent labels or fluorophores include fluorescein and fluorescein derivatives, such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynapthofluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamidofluorescein, fluorescein maleimide, SAMSA-fluorescein, fluorescein thiosemicarbazide, carbohydrazinomethylthioacetyl-aminofluorescein, rhodamine and rhodamine derivatives, such as TRITC, TMR, lissamine rhodamine, Texas Red, rhodamine B, rhodamine 6G, rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, lissamine rhodamine B sulfonyl chloride, lissamine rhodamine B sulfonylhydrazine, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550 C3, BODIPY 530 / 550 C3-SE, BODIPY 530 / 550 C3 hydrazide, BODIPY 493 / 503 C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FL IA, BODIPY 530 / 551 IA, Br-BODIPY 493 / 503, Cascade Blue® and derivatives such as Cascade Blue acetyl azide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Yellow and derivatives such as Lucifer Yellow iodoacetamide, Lucifer Yellow CH, cyanines and derivatives such as indolium-based cyanine dyes, benzo-indolium-based cyanine dyes, pyridium-based cyanine dyes, thiozolium-based cyanine dyes, quinolinium-based cyanine dyes, imidazolium-based cyanine dyes,Cy3, Cy5, lanthanide chelates and derivatives such as BCPDA, TBP, TMT, BHHCT, BCOT, europium chelates, terbium chelates, Alexa Fluor® dyes, DyLight® dyes, Atto™ dyes, LightCycler® Red dyes, CAL Flour dyes, JOE and its derivatives, Oregon Green™ dyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, malachite green, stilbene, DEG dyes, NR dyes, near-infrared dyes, and others known in the art, e.g., Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th Edition; Lakowicz, Principles of Fluorescence Spectroscopy, 2nd Ed., Plenum Press New York (1999); or Hermanson, Bioconjugate Techniques, 2nd Ed. Edition, or derivatives thereof, or any combination thereof. The cyanine dyes may exist in either sulfonated or non-sulfonated form and consist of two indolenine, benzoindolium, pyridium, thiozolium, and / or quinolinium groups separated by a polymethine bridge between the two nitrogen atoms. Commercially available cyanine fluorophores include, for example, Cy3 (which is 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium; may include 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium-5-sulfonate), Cy5 (which may include1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-indolin-2-ylidene)penta-1,3-dien-1-yl)-3,3-dimethyl-3H-yne dol-1-ium, or 1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-sulfoindolin-2-ylidene)penta-1,3-dien-1-yl) Cy7 (which may include 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), and Cy8 (which may include 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), where "Cy" stands for "cyanine" and the first number identifies the number of carbon atoms between the two indolenine groups. Cy2, which is an oxazole derivative rather than an indolenine, and benzo-derivatized Cy3.5, Cy5.5, and Cy7.5 are exceptions to this rule.
[0050] In some embodiments, the reporter moieties can be FRET pairs, allowing multiple classifications to be performed under a single excitation and imaging step. As used herein, FRET can include excitation exchange (Förster) transfer or electron exchange (Dexter) transfer.
[0051] As used herein, the term "support" refers to a substrate designed for the deposition of biomolecules or biological samples for assay and / or analysis. Examples of biomolecules deposited on a support include nucleic acids (e.g., DNA, RNA), polypeptides, sugars, lipids, single cells, or multiple cells. Examples of biological samples include, but are not limited to, saliva, sputum, mucus, blood, plasma, serum, urine, feces, sweat, tears, and fluids from tissues or organs.
[0052] In some embodiments, the support is solid, semi-solid, or a combination of both. In some embodiments, the support is porous, semi-porous, non-porous, or any combination of porous. In some embodiments, the support is substantially planar, concave, convex, or any combination thereof. In some embodiments, the support is cylindrical, for example, comprising a capillary or the inner surface of a capillary.
[0053] In some embodiments, the surface of the support can be substantially smooth, hi some embodiments, the support can be regularly or irregularly textured, including ridges, etchings, pores, three-dimensional scaffolds, or any combination thereof.
[0054] In some embodiments, the support comprises beads having any shape, including spherical, hemispherical, cylindrical, barrel-shaped, toroidal, disk-shaped, rod-shaped, conical, triangular, cubic, polygonal, tubular, or wire-shaped.
[0055] The support can be made of any material, including, but not limited to, glass, fused silica, silicon, polymer (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic substrates are contemplated.
[0056] The present disclosure provides a plurality (e.g., two or more) of nucleic acid template molecules immobilized on a support. In some embodiments, the immobilized plurality of nucleic acid template molecules have the same sequence or different sequences. In some embodiments, individual nucleic acid template molecules in the plurality of nucleic acid template molecules are immobilized at different sites on the support. In some embodiments, two or more individual nucleic acid template molecules in the plurality of nucleic acid templates are immobilized at sites on the support.
[0057] The term "array" refers to a support comprising a plurality of sites located at predetermined locations on the support, forming an array of sites. The sites may be dispersed and separated by interstitial regions. In some embodiments, the predetermined sites on the support may be arranged in rows or columns in one dimension, or in rows or columns in two dimensions. In some embodiments, the plurality of predetermined sites are arranged in an organized manner on the support. In some embodiments, the plurality of predetermined sites are arranged in any organized pattern, including linear, hexagonal, lattice, patterns with reflection symmetry, patterns with rotational symmetry, etc. The pitch between different pairs of sites may be the same or may vary. In some embodiments, the support has at least 10 2 at least 10 sites 3 at least 10 sites 4at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites 9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13 at least 10 sites 14 sites, or at least 10 15 In some embodiments, the substrate comprises a plurality of predetermined sites (e.g., 10 or more sites), the sites being located at predetermined locations on the substrate. 2 ~10 15 The plurality of predetermined sites (one or more sites) comprise nucleic acid template molecules immobilized at the sites forming a nucleic acid template array. In some embodiments, the nucleic acid template molecules are immobilized at the plurality of predetermined sites by hybridization to immobilized surface capture primers, or the nucleic acid template molecules are covalently attached to the surface capture primers. In some embodiments, the plurality of predetermined sites are immobilized at, for example, 10 2 ~10 15 In some embodiments, the immobilized nucleic acid template molecule is clonally amplified to generate immobilized nucleic acid clusters at a plurality of predetermined sites. In some embodiments, the individual immobilized nucleic acid clusters comprise linear clusters or single- or double-stranded concatemers.
[0058] In some embodiments, a support comprising a plurality of sites located at random positions on the support is referred to herein as a support having randomly located sites thereon. The locations of the randomly located sites on the support are not predetermined locations. The plurality of randomly located sites are arranged on the support in a non-ordered and / or unpredictable manner. In some embodiments, the support has at least 10 2 at least 10 sites 3 at least 10 sites 4 at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites 9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13 at least 10 sites 14 sites, or at least 10 15 In some embodiments, the substrate comprises a plurality of randomly positioned sites (e.g., 10 sites, 100 sites, or more), where the sites are randomly located on the substrate. 2 ~10 15 Each of the plurality of randomly located sites (one or more sites) contains a nucleic acid template molecule immobilized at the site. In some embodiments, the nucleic acid template molecule is immobilized at the plurality of randomly located sites by hybridization to an immobilized surface capture primer, or the nucleic acid template molecule is covalently attached to a surface capture primer. In some embodiments, the nucleic acid template molecule is immobilized at the plurality of randomly located sites, e.g., 10 2 ~10 15In some embodiments, the immobilized nucleic acid template is clonally amplified to generate immobilized nucleic acid clusters at multiple randomly located sites. In some embodiments, the individual immobilized nucleic acid clusters comprise linear clusters or single- or double-stranded concatemers.
[0059] In some embodiments, multiple immobilized surface capture primers on a support (e.g., located at predetermined or random positions on the support) are in fluid communication with each other, allowing solutions of reagents (e.g., nucleic acid template molecules, soluble primers, enzymes, nucleotides, divalent cations, buffers, etc.) to flow over the support so that multiple immobilized surface capture primers on the support can react with reagents essentially simultaneously in a massively parallel manner. In some embodiments, fluid communication of multiple immobilized surface capture primers can be used to perform nucleic acid amplification reactions (e.g., RCA, MDA, PCR, and bridge amplification) essentially simultaneously on multiple immobilized surface capture primers. Exemplary supports that allow fluid communication to allow solution flow include, but are not limited to, the inner surface of a flow cell or a capillary tube.
[0060] In some embodiments, multiple immobilized nucleic acid clusters on a support are in fluid communication with each other, allowing solutions of reagents (e.g., enzymes, nucleotides, divalent cations, etc.) to flow over the support, thereby allowing the multiple immobilized nucleic acid clusters on the support to react with the reagents essentially simultaneously in a massively parallel manner. In some embodiments, the fluid communication of multiple immobilized nucleic acid clusters can be used to perform nucleotide binding assays and / or nucleotide polymerization reactions (e.g., primer extension or sequencing) substantially simultaneously on multiple immobilized nucleic acid clusters, and optionally, to perform detection and imaging for massively parallel sequencing.
[0061] The term "immobilized" and related terms refer to nucleic acid molecules that are bound to a support through covalent or non-covalent interactions, or that are attached to a coating on a support, or that are embedded within a matrix formed by a coating on a support, where the nucleic acid molecules include a surface capture primer, a nucleic acid template molecule, and an extension product of the capture primer. The extension product of the capture primer comprises a nucleic acid concatemer (e.g., a nucleic acid cluster). The nucleic acid molecules can be immobilized at predetermined or random positions on the support. The nucleic acid molecules can be immobilized at predetermined or random positions on or within a passivated coating on the support.
[0062] The term "immobilized" and related terms can also refer to an enzyme (e.g., a polymerase) that is bound to a support through covalent or non-covalent interactions, or that is attached to a coating on a support, or that is embedded within a matrix formed by a coating on a support. The enzyme can be immobilized at predetermined or random locations on the support. The enzyme can be immobilized at predetermined or random locations on or within a passivated coating on the support.
[0063] In some embodiments, one or more nucleic acid template molecules are immobilized on a support, e.g., at random or predetermined sites on the support. In some embodiments, one or more nucleic acid template molecules are clonally amplified. In some embodiments, one or more nucleic acid template molecules are clonally amplified off the support (e.g., in solution) and then deposited on the support and immobilized thereon. In some embodiments, a clonal amplification reaction of one or more nucleic acid template molecules is performed on the support, resulting in immobilization on the support. In some embodiments, one or more nucleic acid template molecules are clonally amplified (e.g., in solution or on the support) using a nucleic acid amplification reaction including any one or any combination of polymerase chain reaction (PCR), multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, bridge amplification, isothermal bridge amplification, rolling circle amplification (RCA), circle-circle amplification, helicase-dependent amplification, recombinase-dependent amplification, and / or single-strand binding (SSB) protein-dependent amplification.
[0064] The terms "surface primer," "capture primer," "surface capture primer," and related terms refer to a single-stranded oligonucleotide immobilized to a support and comprising a sequence capable of hybridizing to at least a portion of a nucleic acid template molecule. Surface capture primers can be used to immobilize template molecules to a support via hybridization. Surface capture primers can be immobilized to a support in a manner that resists primer removal during flow, washing, aspiration, and changes in temperature, pH, salt, chemical, and / or enzymatic conditions. Typically, although not necessarily, the 5' end of a surface capture primer can be immobilized to (or embedded within) a coating on) the support or a coating on the support. Alternatively, or in addition, an internal portion or 3' end of a surface capture primer can be immobilized to the support.
[0065] The sequences of the surface capture primers can be wholly or partially complementary to at least a portion of the nucleic acid template molecule along their length. The support can contain multiple immobilized surface capture primers with the same sequence or with two or more different sequences. Surface capture primers can be any length, for example, 4-50 nucleotides, 50-100 nucleotides, 100-150 nucleotides, or longer.
[0066] A surface capture primer can have a terminal 3' nucleotide with a sugar 3' OH moiety that is extendible for nucleotide polymerization (e.g., polymerase-catalyzed polymerization). A surface capture primer can have a terminal 3' nucleotide with a 3' sugar position linked to a chain-terminating moiety that inhibits nucleotide polymerization. The 3' chain-terminating moiety can be removed (e.g., deblocked) using a deblocking agent to convert the 3' end to an extendible 3' OH end. Examples of chain-terminating moieties include alkyl, alkenyl, alkynyl, allyl, aryl, benzyl, azide, amine, amide, keto, isocyanate, phosphate, thio, disulfide, carbonate, urea, or silyl groups. Azido-type chain-terminating moieties include azide, azido, and azidomethyl groups. Examples of deblocking agents include phosphine compounds, such as tris(2-carboxyethyl)phosphine (TCEP) and bis-sulfotriphenylphosphine (BS-TPP), for chain-terminating azide, azido, and azidomethyl groups. Examples of deblocking agents include tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ) for chain-terminating alkyl, alkenyl, alkynyl, and aryl groups. Examples of deblocking agents include Pd / C for chain-terminating aryl and benzyl groups. Examples of deblocking agents include phosphines, beta-mercaptoethanol, or dithiothreitol (DTT) for chain-terminating amine, amide, keto, isocyanate, phosphate, thio, and disulfide groups. Examples of deblocking agents include potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn in acetic acid (AcOH) for carbonate chain terminating groups. Examples of deblocking agents include tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, and triethylamine trihydrofluoride for urea and silyl chain terminating groups.
[0067] The term "sequencing" and related terms refer to methods for obtaining nucleotide sequence information from a nucleic acid molecule, typically by determining the identity of at least some nucleotides (including their nucleobase components) within the nucleic acid molecule. Sequence information for a given region of a nucleic acid molecule can include identifying each and every nucleotide within the sequenced region. Alternatively, the sequencing information determines only some of the region of nucleotides, with the identities of some nucleotides remaining undetermined or incorrectly determined. Any suitable sequencing method can be used. For example, sequencing can include label-free or ion-based sequencing methods. As a further example, sequencing can include labeled or dye-containing nucleotide or fluorescence-based nucleotide sequencing methods. Sequencing can include polony-based sequencing or bridge sequencing methods. Sequencing can use a polymerase and a multivalent molecule to generate at least one avidity complex, where each multivalent molecule comprises multiple nucleotide units tethered to a core (Figures 20-24). Sequencing can use a polymerase and free nucleotides to perform sequencing by synthesis, or a ligase enzyme and multiple sequence-specific oligonucleotides to perform sequencing by ligation.
[0068] Two-strand splint adapter The present disclosure provides compositions, including kits, that include nucleic acid double-stranded splint adaptors, and methods for using double-stranded splint adaptors.
[0069] The double-stranded splint adapter (200) can be used in a one-pot multi-enzyme reaction to introduce one or more new adapter sequences into the library molecule (100). The double-stranded splint adapter (200) comprises a first splint strand (long splint strand (300)) and a second splint strand (short splint strand (400)), which together hybridize to form the double-stranded splint adapter (200) having a double-stranded region and two adjacent single-stranded regions (see, e.g., Figures 1-8). The second splint strand (400) carries the new adapter sequence(s) to be introduced, e.g., a new universal binding sequence and / or a new index sequence. The first splint strand comprises a first region (320), an internal region (310), and a second region (330). The internal region (310) of the first splint strand hybridizes to the second splint strand (400). Two adjacent single-stranded regions (e.g., (320) and (330)) of the double-stranded splint adaptor are designed to hybridize to universal adaptor sequences at the ends of a single-stranded linear library molecule (100) having a sequence of interest (110). For example, the first region (320) of the first splint strand hybridizes to one end of the library molecule, and the second region (330) of the first splint strand hybridizes to the other end of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500) containing two nicks (see, e.g., Figures 1-8). The nicks can be enzymatically ligated to generate a covalently closed circular molecule (600) in which a second splint strand (400) is covalently linked to the library molecule at both ends, thereby introducing a new adapter sequence into the library molecule (see Figure 9).
[0070] Thus, the double-stranded splint adapters and methods described herein can be used to convert any linear library molecule into a covalently closed circular molecule, which can be used for different workflows, e.g., different massively parallel sequencing platforms. The double-stranded splint adapter offers flexibility because the two adjacent single-stranded regions (e.g., (320) and (330)) and the second splint strand (400) can be designed to contain any combination of universal adapter sequences. For example, the two adjacent single-stranded regions (e.g., (320) and (330)) can contain universal binding sequences (or their complementary sequences) for P5 and P7 sequences that bind to surface primers immobilized on a support (e.g., a flow cell), which are typically used to construct library molecules for Illumina sequencing platforms. The second splint strand (400) can include at least one new universal adapter sequence (e.g., a new surface primer sequence) not found in Illumina sequencing platforms, thereby enabling the use of the covalently closed circular molecule (600) in non-Illumina sequencing platforms.
[0071] The methods described herein also offer the advantage of introducing new adapter sequences using a ligation reaction rather than a gap-fill reaction, which results in highly efficient circularization with as few as 0.25 pmol of library molecules.
[0072] The methods described herein may be performed manually or adapted for automation, as annealing and multi-enzyme reactions can be performed in a single reaction vessel (one-pot) by combining several enzymatic reactions (e.g., phosphorylation and ligation) and adding subsequent enzymes (e.g., exonucleases) without alcohol precipitation or organic extraction.
[0073] The present disclosure provides a nucleic acid double-stranded splint adapter (200) comprising (i) a first splint strand (long splint strand (300)) that hybridizes to (ii) a second splint strand (short splint strand (400)) (see, e.g., Figures 1-8). The first splint strand comprises a first region (320), an internal region (310), and a second region (330). The internal region (310) of the first splint strand hybridizes to the second splint strand (400) to form a double-stranded splint adapter (200) having a double-stranded region and two adjacent single-stranded regions. The two adjacent single-stranded regions of the double-stranded splint adapter (200) are designed to hybridize to end sequences of a linear nucleic acid library molecule. The terminal sequences of the linear nucleic acid library molecule comprise first and second universal adaptor sequences, respectively. In some embodiments, the first and second universal adaptor sequences of the linear library molecule comprise binding sequences for first and second capture primers, respectively, immobilized on a support.
[0074] The first region (320) of the first splint strand comprises a first universal adapter sequence capable of hybridizing to a first universal binding sequence at one end of a linear nucleic acid library molecule (see, e.g., Figures 1-8). The second region (330) of the first splint strand comprises a second universal adapter sequence capable of hybridizing to a second universal binding sequence at the other end of the linear nucleic acid library molecule (see, e.g., Figures 1-8). In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0075] In some embodiments, the second splint strand (400) comprises at least two subregions, including a first and a second subregion (see, e.g., Figures 2 and 3). In some embodiments, the first subregion comprises a universal binding sequence for the third surface primer, and the second subregion comprises a universal binding sequence for the fourth surface primer, and the first and second subregions do not hybridize to (or at least exhibit very little hybridization to) the first and second surface primers. In some embodiments, the second splint strand (400) further comprises an optional third subregion, which comprises a sample index sequence having 5-20 bases and / or a unique identification sequence (e.g., NN) having 2-10 or more bases (see, e.g., Figure 3). In some embodiments, the second splint strand (400) includes only one subregion and lacks the second and third subregions, and the first subregion includes a sample index sequence having 5 to 20 bases. In some embodiments, the sample index sequence can be used in multiplex assays to distinguish sequences of interest obtained from different sample sources. In some embodiments, the unique identification sequence includes a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization to the first, second, third, and fourth surface primers. An exemplary arrangement of the subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[second subregion]-[first subregion]-3'. Another exemplary arrangement of subregions in the second splint strand (400) in the 5' to 3' direction includes 5'-[third subregion]-[second subregion]-[first subregion]-3' (see, e.g., Figures 2 and 3). In some embodiments, the second splint strand (400) can be 20 to 100 nucleotides in length, or 30 to 80 nucleotides in length, or 40 to 60 nucleotides in length. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or non-phosphorylated.In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.
[0076] In some embodiments, the first splint strand (300) comprises an internal region (310) comprising at least two subregions, including a fourth and fifth subregion (see, e.g., Figures 2 and 3). The fourth subregion hybridizes to the first subregion of the second splint strand (400). The fifth subregion hybridizes to the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize to (or at least show very little hybridization to) the first and second surface primers. In some embodiments, the internal region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes to the third subregion of the second splint strand (400) (see, e.g., Figure 3). An exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes 5'-[fourth subregion]-[fifth subregion]-3'. Another exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes 5'-[fourth subregion]-[fifth subregion]-[sixth subregion]-3' (see, e.g., Figures 2 and 3). In some embodiments, the first splint strand (300) can be 50 to 150 nucleotides in length, or 60 to 100 nucleotides in length, or 70 to 90 nucleotides in length. In some embodiments, the first splint strand (300) includes one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the first splint strand (300) contains one or more phosphorothioate linkages at internal positions to confer endonuclease resistance, hi some embodiments, the first splint strand (300) contains one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at internal positions.
[0077] Long splint chain variant: amputation In some embodiments, the first splint strand (300) comprises a cleavage strand having a first region (320) with a cleavage sequence at its 5' end (e.g., Figure 11B). In some embodiments, the 5' end of the first region can have a cleavage of any length, e.g., 1-10 nucleotides. In some embodiments, the cleaved first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which are uncleaved and do not possess any sequence variants, e.g., insertions, deletions, or base substitutions. In some embodiments, the cleaved first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first and second subregions of the second splint strand (400) to form the double-stranded splint adapter (200). In some embodiments, the cleaved first splint strand comprises (300), which comprises a cleaved first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the cleaved first splint strand, as part of a double-stranded splint adaptor (200), can hybridize to the library molecule (100) to form a library-sprint complex (500).
[0078] Long splint strand variants: mismatched sequences In some embodiments, the first splint strand (300) comprises a mismatched strand having a first region (320) with a mismatched sequence within the first region (320) (e.g., Figure 11C). In some embodiments, the mismatched sequence is internal to the first region. In some embodiments, the mismatched sequence can be of any length (e.g., 2-20 bases) and includes any sequence that is not perfectly complementary to the left universal adaptor sequence (120) of the library molecule (100). Some embodiments of the mismatched sequence in the first region (320) are shown in lowercase and underlined in Figure 11C. In some embodiments, the mismatched first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which do not possess any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first and second subregions of the second splint strand (400) to form the double-stranded splint adapter (200). In some embodiments, the mismatched first splint strand comprises (300), which comprises a mismatched first region (320) that hybridizes with a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes with a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the mismatched first region (320) can hybridize with the first left universal adaptor sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatched sequence in the first region (320). In some embodiments, the mismatched first splint strand can hybridize to the library molecule (100) as part of the double-stranded splint adaptor (200) to form a library-splint complex (500).
[0079] Long splint strand variants: abasic sites In some embodiments, the first splint strand (300) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site within the fourth subregion and / or at least one abasic site within the fifth subregion (e.g., in the top schematic diagram of Figure 11D, the abasic sites are shown as solid black bars). In some embodiments, the abasic sites each comprise 1',2'-dideoxyribose (e.g., dSpacer from Integrated DNA Technologies (IDT)).
[0080] In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO: 4) and a second region ((330); e.g., SEQ ID NO: 5), which do not possess any abasic sites and / or any sequence variants, such as, for example, insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint strand (400) to form the double-stranded splint adapter (200). In some embodiments, the abasic first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the abasic first splint strand can hybridize to the library molecule (100) as part of the double-stranded splint adaptor (200) to form a library-sprint complex (500).
[0081] Long splint chain variant: uracil In some embodiments, the first splint strand (300) comprises at least one uracil. In some embodiments, the first splint strand (300) comprises at least one uracil in any one or any combination of the regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base may be substituted with uracil. One embodiment of a first splint strand containing uracil is shown in Figure 11D (bottom schematic). Those skilled in the art will recognize that many other arrangements of the first splint strand (300) comprising one or more uracils are possible.
[0082] In some embodiments, the uracil-containing first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the uracil-containing first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the uracil-containing first splint strand, as part of the double-stranded splint adapter (200), can hybridize to the library molecule (100) to form a library-sprint complex (500).
[0083] Short splint strands (400) with inserted or replaced random sequences In some embodiments, the second splint strand (400) includes a random sequence inserted into a first subregion of the second splint strand (e.g., Figures 7A, 12A, and 12B). In some embodiments, the random sequence can replace a portion of the first subregion of the second splint strand (e.g., Figures 7A, 13A, and 13B).
[0084] In some embodiments, the second subregion of the second splint strand (400) does not have a random sequence inserted, and in some embodiments, a portion of the second subregion of the second splint strand (400) is not replaced with a random sequence.
[0085] In some embodiments, the random sequence can be any length, e.g., 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figures 12A and 13A) or 4 nucleotides in length (e.g., "NNNN" in Figures 12B and 13B). In some embodiments, the random sequence can be inserted at any position within the first subregion of the second splint strand (400).
[0086] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0087] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction, hi some embodiments, the random sequences provide high nucleotide diversity, with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run.
[0088] In some embodiments, the random sequence may be sequenced before sequencing the insertion region. In some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration. In some embodiments, the sequences of the left index (160), the right index (170), and / or any portion of the insertion region (110) do not provide sufficient nucleotide diversity to enable polony mapping and / or template registration. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to the left index (160), the right index (170), and / or any portion of the insertion region (110).
[0089] In some embodiments, a predetermined sequence is inserted into the sequence of the fourth subregion of the first splint strand (300). The length of the inserted predetermined sequence can be the same length as the random sequence inserted into the first subregion of the second splint strand (e.g., Figures 12A and 12B). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 12A and 12B.
[0090] In some embodiments, a portion of the fourth subregion of the first splint strand (300) is replaced with a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth subregion is the same as the random sequence replacing the portion of the first subregion of the second splint strand (e.g., Figures 13A and 13B). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 13A and 13B.
[0091] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth and fifth subregions of the first splint strand (300) to form the double-stranded splint adapter (200) (e.g., Figures 12A, 12B, 13A, and 13B). In some embodiments, the double-stranded splint adapter (200) forms a bubble at the location of the inserted or replacement random sequence.
[0092] In some embodiments, a second splint strand (400) bearing a random sequence can hybridize to a library molecule (100) as part of a double-stranded splint adaptor (200) to form a library-splint complex (500).
[0093] Short splint strands (400) with added random and index sequences In some embodiments, the second splint strand (400) comprises a random sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 7B, 14A, and 14B). In some embodiments, the random sequence comprises an index sequence.
[0094] In some embodiments, the second subregion of the second splint strand (400) does not have any random sequence added.
[0095] In some embodiments, the added random sequence can be any length, for example, 2 to 10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figure 14A) or 4 nucleotides in length (e.g., "NNNN" in Figure 14B).
[0096] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0097] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction, hi some embodiments, the random sequences provide high nucleotide diversity, with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run.
[0098] In some embodiments, the random sequence may be sequenced before sequencing the insertion region. In some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration. In some embodiments, the sequences of the left index (160), the right index (170), and / or any portion of the insertion region (110) do not provide sufficient nucleotide diversity to enable polony mapping and / or template registration. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to the left index (160), the right index (170), and / or any portion of the insertion region (110).
[0099] In some embodiments, index sequences can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay.
[0100] In some embodiments, the predetermined sequence is added to the 5' end of the fourth subregion of the first splint strand (300). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 14A and 14B). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence and index sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 14A and 14B). The added predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 14A and 14B.
[0101] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adapter (200) (e.g., Figures 14A and 14B). In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the location of the added random sequence. In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the location of the added random sequence and the index sequence.
[0102] In some embodiments, the second splint strand (400) with the added random sequence (and optionally an index sequence) can hybridize to the library molecule (100) as part of the double-stranded splint adapter (200) to form a library-sprint complex (500).
[0103] Short sprint chain arrangement (400) In some embodiments, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises first and second subregions comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A. In some embodiments, the 5' end of the second splint strand (400) may or may not be phosphorylated.
[0104] Long splint chain arrangement (300) In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence that includes a universal binding sequence for a first surface primer (or its complementary sequence), and the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complementary sequence of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence that includes a universal binding sequence for a second surface primer (or its complementary sequence), and the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize to a P7 surface primer or the complementary sequence of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fourth subregion having the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO: 198).In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) having fourth and fifth subregions, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAGGCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11A. In some embodiments, the 5' end of the first splint strand (300) can be phosphorylated or unphosphorylated. In some embodiments, the first subregion of the second splint strand (400) can hybridize to the fourth subregion of the first splint strand (300). In some embodiments, the second subregion of the second splint strand (400) can hybridize to the fifth subregion of the first splint strand (300).
[0105] Library-Sprint Complex The present disclosure provides a library-sprint complex (500) comprising: (i) a single-stranded nucleic acid library molecule (100) comprising a sequence of interest (110) flanked on one side by at least a first left universal adaptor sequence (120) and on the other side by at least a first right universal adaptor sequence (130); and (ii) a double-stranded splint adaptor (110) comprising a first splint strand (long splint strand (300)) and a second splint strand (short splint strand (400)). and a double-stranded splint adapter (200) comprising a first splint strand having a first region (320), an internal region (310), and a second region (330), wherein the internal region (310) of the first splint strand hybridizes to a second splint strand (400) to form a double-stranded splint adapter (200) having a double-stranded region adjacent to a single-stranded region on either side. In the library-sprint complex (500), the first region (320) of the first splint strand hybridizes to at least the first left universal adapter sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adapter sequence (130) of the library molecule, thereby circularizing the library molecule to generate the library-sprint complex (500) (see, for example, Figures 1-8).
[0106] In the library-splint complex (500), the first region (320) of the first splint strand comprises a first universal adaptor sequence capable of hybridizing to a first universal binding sequence at one end of a linear nucleic acid library molecule (see, e.g., Figures 1-8). In some embodiments, the first region (320) of the first splint strand comprises a first universal adaptor sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the first splint strand (300) can be 50-150 nucleotides in length, or 60-100 nucleotides in length, or 70-90 nucleotides in length. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0107] The second region (330) of the first splint strand comprises a second universal adapter sequence capable of hybridizing to a second universal binding sequence at the other end of the linear nucleic acid library molecule (see, e.g., Figures 1-8). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0108] In the library-sprint complex (500), the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adaptor sequence (130) of the library molecule, thereby circularizing the library molecule to generate the library-sprint complex (500). The library-sprint complex (500) comprises a first nick between the 5' end of the library molecule and the 3' end of the second splint strand. The library-sprint complex (500) also comprises a second nick between the 5' end of the second splint strand and the 3' end of the library molecule (see, e.g., Figures 1-8). In some embodiments, the first and second nicks are enzymatically ligatable.
[0109] In the library-sprint complex (500), the first region (320) of the first splint strand can hybridize to either the sense or antisense strand of the double-stranded nucleic acid library molecule. In the library-sprint complex (500), the second region (330) of the first splint strand can hybridize to either the sense or antisense strand of the double-stranded nucleic acid library molecule. The double-stranded nucleic acid library molecule can be denatured to generate single-stranded sense and antisense library strands.
[0110] In the library-sprint complex (500), the second splint strand (400) does not hybridize to the sequence of interest (110), and the internal region (310) of the first splint strand does not hybridize to the sequence of interest (110).
[0111] In the library-sprint complex (500), the first region (320) of the first splint strand does not hybridize to the sequence of interest (110), and the second region (330) of the first splint strand does not hybridize to the sequence of interest (110).
[0112] In some embodiments, the 5' ends of the single-stranded library molecules (100) in the library-sprint complex (500) are phosphorylated or lack a phosphate group. In some embodiments, the 3' ends of the single-stranded library molecules include a terminal 3' OH group or a terminal 3' blocking group.
[0113] In some embodiments, the nucleic acid library molecule (100) comprises a second left universal adaptor sequence (140). In some embodiments, the nucleic acid library molecule (100) comprises a second right universal adaptor sequence (150). Exemplary library molecules (100) are shown in Figures 4-8. In some embodiments, the nucleic acid library molecule (100) can further comprise additional left and / or right universal adaptor sequences.
[0114] In some embodiments, the nucleic acid library molecule (100) further comprises a first left index sequence (160). In some embodiments, the nucleic acid library molecule (100) further comprises a first right index sequence (170). In some embodiments, the first left index sequence (160) comprises a sample index sequence. In some embodiments, the first right index sequence (170) comprises another sample index sequence. In some embodiments, the first left index sequence (160) and the first right index sequence (170) are not the same sequence. In some embodiments, the nucleic acid library molecule (100) comprises a first left index sequence (160) and / or a first right index sequence (170). Sample index sequences can be used in multiplex assays to distinguish sequences of interest obtained from different sample sources. Exemplary library molecules (100) are shown in Figures 4-8. A list of exemplary first left index array (160) and first right index array (170) is provided in Table 1 in Figure 33. The first left index array (160) may include a random array (e.g., NNN) or may lack a random array. The first right index array (170) may include a random array (e.g., NNN) or may lack a random array.
[0115] In some embodiments, the nucleic acid library molecule (100) further comprises at least one junction adaptor sequence located between any of the universal adaptor sequences described herein (see, e.g., FIG. 8). For example, a first left junction adaptor sequence (125) can be located between the first left universal adaptor sequence (120) and the first left index sequence (160). A second left junction adaptor sequence (165) can be located between the first left index sequence (160) and the second left universal adaptor sequence (140). A third left junction adaptor sequence (145) can be located between the second left universal adaptor sequence (140) and the sequence of interest (110). A first right junction adaptor sequence (135) can be located between the first right universal adaptor sequence (130) and the first right index sequence (170). The second right junction adaptor sequence (175) can be located between the first right index sequence (170) and the second right universal adaptor sequence (150). The third right junction adaptor sequence (155) can be located between the second right universal adaptor sequence (150) and the sequence of interest (110). In some embodiments, the nucleic acid library molecule (100) further comprises at least 1 to up to 10 added universal adaptor sequences located 5' (upstream) of the first left universal adaptor sequence (120) (see, e.g., Figure 8). In some embodiments, the nucleic acid library molecule (100) further comprises at least 1 to up to 10 added universal adaptor sequences located 3' (downstream) of the first right universal adaptor sequence (130) (see, e.g., Figure 8). Any of the junction adaptor sequences can be 3 to 60 nucleotides in length and / or added universal adaptor sequences, including any sequence. Either the junction adapter sequence and / or the added universal adapter sequence comprises a universal sequence or a unique sequence.Any of the junction adapter sequences and / or the added universal adapter sequences include a binding sequence for an amplification primer, a sequencing primer, or a compaction oligonucleotide, or a combination thereof. Any of the junction adapter sequences and / or the added universal adapter sequences include a binding sequence for an immobilized surface primer (e.g., a capture primer). Any of the junction adapter sequences and / or the added universal adapter sequences include a sample index sequence. Any of the junction adapter sequences and / or the added universal adapter sequences include a unique identification sequence. Any of the junction adapter sequences and / or the added universal adapter sequences, particularly the junction adapter sequence (145) shown in FIG. 8, includes the Tn5 transposon end sequence 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 211). Any of the junction adapter sequences and / or the added universal adapter sequences, particularly the junction adapter sequence (155) shown in FIG. 8, includes the Tn5 transposon end sequence 5'-CTGTCTCTTATACACATCT-3' (SEQ ID NO: 212). The Tn5 transposon end sequence can be introduced into the library molecule (100) via a transposase-mediated reaction, which involves contacting double-stranded input DNA (e.g., genomic DNA) with a Tn5-type transposase enzyme and a double-stranded oligonucleotide comprising a Tn transposon end sequence (SEQ ID NO: 211) linked to a universal adapter sequence or a sample index sequence under conditions suitable for forming a transposon synaptic complex. In the double-stranded oligonucleotide, the Tn transposon end sequence (SEQ ID NO: 211) can be located 5' or 3' to the universal adapter sequence or sample index sequence.
[0116] Multiplex workflows are enabled by preparing sample-indexed libraries using one or both index sequences (e.g., left index sequence and / or right index sequence). The first left index sequence (160) and / or the first right index sequence (170) can be used to prepare separate sample-indexed libraries using input nucleic acids isolated from different sources. The sample-indexed libraries can be pooled together to generate a multiplexed library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insert region, along with the first left index sequence (160) and / or the first right index sequence (170), can be used to identify the source of the input nucleic acid. In some embodiments, any number of sample-indexed libraries can be pooled together, for example, 2 to 10, 10 to 50, 50 to 100, 100 to 200, or more than 200 sample-indexed libraries can be pooled. Exemplary nucleic acid sources include naturally occurring sources, recombinant sources, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or whole organisms. Exemplary nucleic acid sources include fresh sources, frozen sources, fresh frozen sources, or archived (e.g., formalin-fixed, paraffin-embedded; FFPE) sources. Those skilled in the art will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.
[0117] In some embodiments, the nucleic acid library molecule (100) further comprises an optional first left unique identification sequence (180) as shown in Figure 5. In some embodiments, the nucleic acid library molecule (100) further comprises an optional first right unique identification sequence (190) as shown in Figure 6. In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence used to uniquely identify an individual sequence of interest (e.g., an insert sequence) to which a unique adaptor has been added in a population of other sequences of molecules of interest. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular tagging. Exemplary library molecules (100) are shown in Figures 4-8.
[0118] In some embodiments, the nucleic acid library molecule (100) comprises any one or any combination of two or more of a first left universal adaptor sequence (120), a second left universal adaptor sequence (140), a first left index sequence (160), a first left unique identifier sequence (180), a first right universal adaptor sequence (130), a second right universal adaptor sequence (150), a first right index sequence (170), and / or a first right unique identifier sequence (190). Exemplary library molecules (100) are shown in Figures 4-8.
[0119] In some embodiments, the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) comprise a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises additional left universal adapter sequences.
[0120] In some embodiments, the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) comprise a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises additional right universal adapter sequences.
[0121] In some embodiments, the second splint strand (400) comprises at least two subregions, including a first and a second subregion (see, e.g., Figures 2 and 3). In some embodiments, the first subregion comprises a universal binding sequence for the third surface primer, and the second subregion comprises a universal binding sequence for the fourth surface primer, and the first and second subregions do not hybridize to (or at least exhibit very little hybridization to) the first and second surface primers. In some embodiments, the second splint strand (400) further comprises an optional third subregion, which comprises a sample index sequence having 5-20 bases and / or a unique identification sequence (e.g., NN) having 2-10 or more bases (see, e.g., Figures 2 and 3). In some embodiments, the second splint strand (400) includes only one subregion and lacks the second and third subregions, and the first subregion includes a sample index sequence having 5 to 20 bases. In some embodiments, the sample index sequence can be used in multiplex assays to distinguish sequences of interest obtained from different sample sources. In some embodiments, the unique identification sequence includes a random sequence. The unique identification sequence can be designed to exhibit reduced hybridization to, or no hybridization to, the first, second, third, and fourth surface primers. An exemplary arrangement of the subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[second subregion]-[first subregion]-3'. Another exemplary arrangement of the subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[third subregion]-[second subregion]-[first subregion]-3'. In some embodiments, the second splint strand (400) can be 20 to 100 nucleotides in length, or 30 to 80 nucleotides in length, or 40 to 60 nucleotides in length. In some embodiments, the second splint strand (400) includes one or more phosphorothioate linkages at the 5' and / or 3' ends to confer exonuclease resistance.In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0122] In some embodiments, the first splint strand (300) comprises an internal region (310) comprising at least two subregions, including a fourth and fifth subregion (see, e.g., Figures 2 and 3). The fourth subregion hybridizes to the first subregion of the second splint strand (400). The fifth subregion hybridizes to the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize to (or at least show very little hybridization to) the first and second surface primers. In some embodiments, the internal region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes to the third subregion of the second splint strand (400) (see, e.g., Figures 2 and 3). An exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes: 5'-[fourth subregion]-[fifth subregion]-3'. Another exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes: 5'-[fourth subregion]-[fifth subregion]-[sixth subregion]-3'.
[0123] In some embodiments, an exemplary library-sprint complex (500) includes (a) a single-stranded nucleic acid library molecule (100), (b) a first splint strand (300), and (c) a second splint strand (400).
[0124] In the exemplary library-sprint complex (500), the single-stranded nucleic acid library molecule (100) includes the following components arranged in 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer, (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer, (iii) a sequence of interest (110), (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer, and (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer.
[0125] In the exemplary library-sprint complex (500), the first splint strand (300) comprises components arranged in 5' to 3' order: a first region (320), an internal region (310), and a second region (330).
[0126] In the exemplary library-splint complex (500), the second splint strand (400) includes subregions arranged in 3' to 5' order: a first subregion having a universal binding sequence for a third surface primer, and a second subregion having a universal binding sequence for a fourth surface primer.
[0127] In the exemplary library-sprint complex (500), a portion of the first splint strand (300) hybridizes to a portion of the library molecule (100), thereby circularizing the library molecule to generate the library-sprint complex (500), whereby a first region (320) of the first splint strand hybridizes to the binding sequence for the first surface primer (120) and a third region (330) of the first splint strand hybridizes to the binding sequence for the second surface primer (130). Additionally, a second splint strand (400) hybridizes to an internal region (310) of the first splint strand (300). The library-sprint complex (500) comprises a first nick between the 5' end of the library molecule and the 3' end of the second splint strand, and a second nick between the 5' end of the second splint strand and the 3' end of the library molecule, wherein the first and second nicks are enzymatically ligatable.
[0128] In the exemplary library-sprint complex (500), the second splint strand (400) does not hybridize to the sequence of interest (110), and the internal region (310) of the first splint strand does not hybridize to the sequence of interest (110).
[0129] In the exemplary library-sprint complex (500), the first region (320) of the first splint strand does not hybridize to the sequence of interest (110), and the second region (330) of the first splint strand does not hybridize to the sequence of interest (110).
[0130] In some embodiments, any of the library-sprint complexes (500) described herein comprises a plurality of library-sprint complexes (500), wherein the target sequences (110) of the individual library-sprint complexes in the plurality of library-sprint complexes comprise the same target sequence or different target sequences.
[0131] Library splint complexes formed by double-stranded adapters with cleaved long splint strands In some embodiments, the library-sprint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adaptor (200). In some embodiments, a first region (320) of the first splint strand hybridizes to at least a first left universal adaptor sequence (120) of the library molecule, and a second region (330) of the first splint strand hybridizes to at least a first right universal sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500) having first and second nicks (see, e.g., Figures 1-6 and 8).
[0132] In some embodiments, the first splint strand (300) comprises a truncated strand having a first region (320) with a truncation sequence at its 5' end (e.g., FIG. 11B). In some embodiments, the first region has a truncation sequence at its 5' end when compared to SEQ ID NO: 199, as shown in FIG. 11A. In some embodiments, the 5' end of the first region can have a truncation of any length, e.g., 1-10 nucleotides. In some embodiments, the truncated first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which are untruncated and do not possess any sequence variants, e.g., insertions, deletions, or base substitutions. In some embodiments, the cleaved first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the cleaved first splint strand (300) comprises a cleaved first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the cleaved first splint strand, as part of the double-stranded splint adapter (200), can hybridize to the library molecule (100) to form a library-sprint complex (500) having first and second nicks.
[0133] Library splint complexes formed by double-stranded adapters with long splint strands containing mismatched sequences In some embodiments, the library-sprint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adapter (200). In some embodiments, a first region (320) of the first splint strand hybridizes to at least a first left universal adapter sequence (120) of the library molecule, and a second region (330) of the first splint strand hybridizes to at least a first right universal adapter sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500) having first and second nicks (see, e.g., Figures 1-8).
[0134] In some embodiments, the first splint strand (300) comprises a mismatched strand having a first region (320) with a mismatched sequence within the first region (320) (e.g., Figure 11C). In some embodiments, the mismatched sequence may be of any length (e.g., 2-20 bases) and includes any sequence that is not perfectly complementary to the left universal adaptor sequence (120) of the library molecule (100). Some embodiments of the mismatched sequence in the first region (320) are shown in lowercase and underlined in Figure 11C. In some embodiments, the mismatched first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which do not possess any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first and second subregions of the second splint strand (400) to form the double-stranded splint adapter (200). In some embodiments, the mismatched first splint strand comprises (300), which comprises a mismatched first region (320) that hybridizes with a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes with a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the mismatched first region (320) can hybridize with the first left universal adaptor sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatched sequence in the first region (320). In some embodiments, the mismatched first splint strand can hybridize to the library molecule (100) as part of the double-stranded splint adaptor (200) to form a library-splint complex (500) having a first and second nick.
[0135] Library splint complexes formed by double-stranded adapters with long splint strands containing abasic sites In some embodiments, the library-sprint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adapter (200). In some embodiments, a first region (320) of the first splint strand hybridizes to at least a first left universal adapter sequence (120) of the library molecule, and a second region (330) of the first splint strand hybridizes to at least a first right universal adapter sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500) having first and second nicks (see, e.g., Figures 1-8).
[0136] In some embodiments, the first splint strand (300) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site within the fourth subregion and / or at least one abasic site within the fifth subregion (e.g., in the top schematic diagram of Figure 11D, the abasic sites are shown as solid black bars). In some embodiments, the abasic sites each comprise 1',2'-dideoxyribose (e.g., dSpacer from Integrated DNA Technologies (IDT)).
[0137] In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO: 4) and a second region ((330); e.g., SEQ ID NO: 5), which do not possess any abasic sites and / or any sequence variants, such as, for example, insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint strand (400) to form the double-stranded splint adapter (200). In some embodiments, the abasic first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the abasic first splint strand, as part of a double-stranded splint adaptor (200), can hybridize to the library molecule (100) to form a library-splint complex (500) having first and second nicks.
[0138] Library splint complexes formed by double-stranded adapters with long splint strands containing uracil In some embodiments, the library-sprint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adapter (200). In some embodiments, a first region (320) of the first splint strand hybridizes to at least a first left universal adapter sequence (120) of the library molecule, and a second region (330) of the first splint strand hybridizes to at least a first right universal adapter sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500) having first and second nicks (see, e.g., Figures 1-8).
[0139] In some embodiments, the first splint strand (300) comprises at least one uracil in any one or any combination of the regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base may be substituted with uracil. One embodiment of a first splint strand containing uracil is shown in Figure 1 ID (bottom schematic). Those skilled in the art will recognize that many other arrangements of the first splint strand (300) comprising one or more uracils are possible.
[0140] In some embodiments, the uracil-containing first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the uracil-containing first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the uracil-containing first splint strand, as part of the double-stranded splint adapter (200), can hybridize to the library molecule (100) to form a library-sprint complex (500) having first and second nicks.
[0141] Library splint complexes formed by double-stranded adapters with short splint strands having random sequences In some embodiments, the library-sprint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adapter (200). In some embodiments, a first region (320) of the first splint strand hybridizes to at least a first left universal adapter sequence (120) of the library molecule, and a second region (330) of the first splint strand hybridizes to at least a first right universal adapter sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500) having first and second nicks (see, e.g., Figures 1-8).
[0142] In some embodiments, the second splint strand (400) includes a random sequence inserted into a first subregion of the second splint strand (e.g., Figures 7A, 12A, and 12B). In some embodiments, the random sequence can replace a portion of the first subregion of the second splint strand (e.g., Figures 7A, 13A, and 13B).
[0143] In some embodiments, the second subregion of the second splint strand (400) does not have a random sequence inserted, and in some embodiments, a portion of the second subregion of the second splint strand (400) is not replaced with a random sequence.
[0144] In some embodiments, the random sequence can be any length, e.g., 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figures 12A and 13A) or 4 nucleotides in length (e.g., "NNNN" in Figures 12B and 13B). In some embodiments, the random sequence can be inserted at any position within the first subregion of the second splint strand (400).
[0145] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0146] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction, hi some embodiments, the random sequences provide high nucleotide diversity, with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run.
[0147] In some embodiments, the random sequence may be sequenced before sequencing the insertion region. In some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration. In some embodiments, the sequences of the left index (160), the right index (170), and / or any portion of the insertion region (110) do not provide sufficient nucleotide diversity to enable polony mapping and / or template registration. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to the left index (160), the right index (170), and / or any portion of the insertion region (110).
[0148] In some embodiments, a predetermined sequence is inserted into the sequence of the fourth subregion of the first splint strand (300). The length of the inserted predetermined sequence can be the same length as the random sequence inserted into the first subregion of the second splint strand (e.g., Figures 12A and 12B). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 12A and 12B.
[0149] In some embodiments, a portion of the fourth subregion of the first splint strand (300) is replaced with a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth subregion is the same as the random sequence replacing the portion of the first subregion of the second splint strand (e.g., Figures 13A and 13B). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 13A and 13B.
[0150] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth and fifth subregions of the first splint strand (300) to form the double-stranded splint adapter (200) (e.g., Figures 12A, 12B, 13A, and 13B). In some embodiments, the double-stranded splint adapter (200) forms a bubble at the location of the inserted or replacement random sequence.
[0151] In some embodiments, a second splint strand (400) bearing a random sequence can hybridize to a library molecule (100) as part of a double-stranded splint adapter (200) to form a library-splint complex (500) having a first and second nick.
[0152] Library splint complexes formed by double-stranded adapters with short splint strands having appended random sequences and index sequences. In some embodiments, the library-sprint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adapter (200). In some embodiments, a first region (320) of the first splint strand hybridizes to at least a first left universal adapter sequence (120) of the library molecule, and a second region (330) of the first splint strand hybridizes to at least a first right universal adapter sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500) having first and second nicks (see, e.g., Figures 1-8).
[0153] In some embodiments, the second splint strand (400) comprises a random sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 7B, 14A, and 14B). In some embodiments, the random sequence further comprises an index sequence.
[0154] In some embodiments, the second subregion of the second splint strand (400) does not have any random sequence added.
[0155] In some embodiments, the added random sequence can be any length, for example, 2 to 10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figure 14A) or 4 nucleotides in length (e.g., "NNNN" in Figure 14B).
[0156] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0157] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction, hi some embodiments, the random sequences provide high nucleotide diversity, with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run.
[0158] In some embodiments, the random sequence may be sequenced before sequencing the insertion region. In some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration. In some embodiments, the sequences of the left index (160), the right index (170), and / or any portion of the insertion region (110) do not provide sufficient nucleotide diversity to enable polony mapping and / or template registration. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to the left index (160), the right index (170), and / or any portion of the insertion region (110).
[0159] In some embodiments, index sequences can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay.
[0160] In some embodiments, the predetermined sequence is added to the 5' end of the fourth subregion of the first splint strand (300). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 14A and 14B). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence and index sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 14A and 14B). The added predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 14A and 14B.
[0161] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adapter (200) (e.g., Figures 14A and 14B). In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the location of the added random sequence. In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the location of the added random sequence and the index sequence.
[0162] In some embodiments, a second splint strand (400) with added random sequence (and optionally index sequence) can hybridize to a library molecule (100) as part of a double-stranded splint adapter (200) to form a library-splint complex (500) with first and second nicks.
[0163] Short sprint chain arrangement (400) In some embodiments of the library-splint complex (500) described herein, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises first and second subregions comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A. In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or unphosphorylated.
[0164] In some embodiments, the second splint strand (400) comprises only one subregion, lacking the second and third subregions, and the first subregion comprises a sample index sequence having between 5 and 20 bases.
[0165] Long splint chain arrangement (300) In some embodiments of the library-splint complex (500) described herein, the first region (320) of the first splint strand comprises a first universal adapter sequence that includes a universal binding sequence for a first surface primer (or its complementary sequence), and the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complementary sequence of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence that includes a universal binding sequence for a second surface primer (or its complementary sequence), and the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize to a P7 surface primer or the complementary sequence of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fourth subregion having the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198).In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) having fourth and fifth subregions, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAGGCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11A. In some embodiments, the 5' end of the first splint strand (300) can be phosphorylated or unphosphorylated. In some embodiments, the first subregion of the second splint strand (400) can hybridize to the fourth subregion of the first splint strand (300). In some embodiments, the second subregion of the second splint strand (400) can hybridize to the fifth subregion of the first splint strand (300).
[0166] Library-Sprint Complex Sequence In some embodiments of the library-sprint complex (500) described herein, the first region (320) of the first splint strand comprises a sequence capable of binding to the first left universal adapter sequence (120) of the library molecule, and the first region (320) of the first splint strand comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 215) or a complementary sequence thereof.
[0167] In some embodiments of the library-sprint complex (500) described herein, the second region (330) of the first splint strand comprises a sequence capable of binding to the first right universal adapter sequence (130) of the library molecule, and the second region (330) of the first splint strand comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO: 216) or a complementary sequence thereof.
[0168] In some embodiments, in any of the library-splint complexes (500) described herein, the library molecule comprises a first left universal adaptor sequence (120) that binds to a first region (320) of a first splint strand, and the left universal binding sequence (120) comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (sequence number 203).
[0169] In some embodiments, in any of the library-splint complexes (500) described herein, the library molecule comprises a first left universal adaptor sequence (120) that binds to a first region (320) of the first splint strand, and the first left universal adaptor sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 213) or a complementary sequence thereof.
[0170] In some embodiments of the library-sprint complex (500) described herein, the library molecule comprises a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 204).
[0171] In some embodiments of the library-sprint complex (500) described herein, the library molecule comprises a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 207).
[0172] In some embodiments of the library-sprint complex (500) described herein, the library molecule comprises a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO: 208).
[0173] In some embodiments of the library-sprint complex (500) described herein, the library molecule comprises a second right universal adapter sequence comprising a sequence (150) for binding of a sequencing primer, wherein the second right universal adapter sequence comprises the sequence 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO: 205).
[0174] In some embodiments of the library-sprint complex (500) described herein, the library molecule comprises a second right universal adapter sequence comprising a sequence for binding of a sequencing primer (150), wherein the second right universal adapter sequence comprises the sequence 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO: 209).
[0175] In some embodiments of the library-sprint complex (500) described herein, the library molecule comprises a second right universal adapter sequence comprising a sequence for binding of a sequencing primer (150), wherein the second right universal adapter sequence comprises the sequence 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO: 210).
[0176] In some embodiments of the library-splint complex (500) described herein, the library molecule comprises a first right universal adaptor sequence (130) that binds to a first region of the first splint strand (330), and the first right universal adaptor sequence (130) comprises the sequence 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 206).
[0177] In some embodiments of the library-splint complex (500) described herein, the library molecule comprises a first right universal adaptor sequence (130) that binds to a first region of the first splint strand (330), and the first right universal adaptor sequence (130) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 214) or a complementary sequence thereof.
[0178] The present disclosure provides a reaction mixture comprising a plurality of any of the library-sprint complexes (500) described herein. In some embodiments, the reaction mixture comprises a plurality of any of the library-sprint complexes (500) described herein and T4 polynucleotide kinase. In some embodiments, the reaction mixture comprises a plurality of any of the library-sprint complexes (500) described herein and a ligase enzyme. In some embodiments, the reaction mixture comprises a plurality of any of the library-sprint complexes (500) described herein, T4 polynucleotide kinase, and a ligase enzyme. In some embodiments, the ligase enzyme comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase.
[0179] Covalent closed ring molecule The present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), at least a first left universal adaptor sequence (120), at least a first right universal adaptor sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9. In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adaptor sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adaptor sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adaptor sequences.
[0180] In some embodiments, the covalently closed circular molecule (600) further comprises a first left-indexing sequence (160) and / or a first right-indexing sequence (170). The indexing sequences can be used in multiplex assays to distinguish sequences of interest obtained from different sample sources. A list of exemplary first left-indexing sequences (160) and first right-indexing sequences (170) is provided in Table 1 in FIG. 33. The first left-indexing sequence (160) may include a random sequence (e.g., NNN) or may lack a random sequence. The first right-indexing sequence (170) may include a random sequence (e.g., NNN) or may lack a random sequence.
[0181] Multiplex workflows are enabled by preparing sample-indexed libraries using one or both index sequences (e.g., left index sequence and / or right index sequence). The first left index sequence (160) and / or the first right index sequence (170) can be used to prepare separate sample-indexed libraries using input nucleic acids isolated from different sources. The sample-indexed libraries can be pooled together to generate a multiplexed library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insert region, along with the first left index sequence (160) and / or the first right index sequence (170), can be used to identify the source of the input nucleic acid. In some embodiments, any number of sample-indexed libraries can be pooled together, for example, 2 to 10, 10 to 50, 50 to 100, 100 to 200, or more than 200 sample-indexed libraries can be pooled. Exemplary nucleic acid sources include naturally occurring sources, recombinant sources, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or whole organisms. Exemplary nucleic acid sources include fresh sources, frozen sources, fresh frozen sources, or archived (e.g., formalin-fixed, paraffin-embedded; FFPE) sources. Those skilled in the art will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.
[0182] In some embodiments, the covalently closed circular molecule (600) further comprises an optional first left unique identification sequence (180) and / or an optional first right unique identification sequence (190), as shown in Figures 5 and 6. In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence used to uniquely identify an individual sequence of interest (e.g., an insert sequence) to which a unique adaptor has been added, in a population of other sequences of molecules of interest. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) may be used for molecular tagging.
[0183] In some embodiments, the covalently closed circular molecule (600) comprises any one or any combination of two or more of a first left universal adaptor sequence (120), a second left universal adaptor sequence (140), a first left index sequence (160), a first left unique identifier sequence (180), a first right universal adaptor sequence (130), a second right universal adaptor sequence (150), a first right index sequence (170), and / or a first right unique identifier sequence (190). In some embodiments, the first left index sequence (160) comprises a sample index sequence. In some embodiments, the first right index sequence (170) comprises another sample index sequence. The sample index sequence may be used in a multiplex assay to distinguish sequences of interest obtained from different sample sources. In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence used to uniquely identify an individual sequence of interest (e.g., an insert sequence) to which a unique adaptor has been added in a population of other sequences of molecules of interest. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) may be used for molecular tagging.
[0184] In some embodiments, in the covalently closed circular molecule (600), the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) comprise a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the covalently closed circular molecule (600) can further comprise an additional left universal adapter sequence.
[0185] In some embodiments, in the covalently closed circular molecule (600), the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) comprise a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the covalently closed circular molecule (600) can further comprise an additional right universal adapter sequence.
[0186] In some embodiments, the covalently closed circular molecule (600) further comprises at least one junction adaptor sequence located between any of the universal adaptor sequences described herein (see, e.g., FIG. 8). For example, the first left junction adaptor sequence (125) can be located between the first left universal adaptor sequence (120) and the first left index sequence (160). The second left junction adaptor sequence (165) can be located between the first left index sequence (160) and the second left universal adaptor sequence (140). The third junction adaptor sequence (145) can be located between the second left universal adaptor sequence (140) and the sequence of interest (110). The first right junction adaptor sequence (135) can be located between the first right universal adaptor sequence (130) and the first right index sequence (170). The second right junction adaptor sequence (175) can be located between the first right index sequence (170) and the second right universal adaptor sequence (150). The third right junction adaptor sequence (155) can be located between the second right universal adaptor sequence (150) and the sequence of interest (110). In some embodiments, the covalently closed circular molecule (600) further comprises at least 1 to up to 10 additional universal adaptor sequences located 5' (upstream) of the first left universal adaptor sequence (120) (see, e.g., Figure 8). In some embodiments, the covalently closed circular molecule (600) further comprises at least 1 to up to 10 additional universal adaptor sequences located 3' (downstream) of the first right universal adaptor sequence (130) (see, e.g., Figure 8). Any of the junction adaptor sequences and / or the additional universal adaptor sequences can comprise any sequence and can be 3 to 60 nucleotides in length. Either the junction adapter sequence and / or the added universal adapter sequence comprises a universal sequence or a unique sequence.Any of the junction adapter sequences and / or the added universal adapter sequences include a binding sequence for an amplification primer, a sequencing primer, or a compaction oligonucleotide. Any of the junction adapter sequences and / or the added universal adapter sequences include a binding sequence for an immobilized surface primer (e.g., a capture primer). Any of the junction adapter sequences and / or the added universal adapter sequences include a sample index sequence. Any of the junction adapter sequences and / or the added universal adapter sequences include a unique identification sequence. Any of the junction adapter sequences and / or the added universal adapter sequences, particularly the junction adapter sequence (145) shown in FIG. 8, includes the Tn5 transposon end sequence 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 211). Any of the junction adapter sequences and / or the added universal adapter sequences, particularly the junction adapter sequence (155) shown in FIG. 8, includes the Tn5 transposon end sequence 5'-CTGTCTCTTATACACATCT-3' (SEQ ID NO: 212). The Tn5 transposon end sequence can be introduced into the library molecule (100) via a transposase-mediated reaction, which involves contacting double-stranded input DNA (e.g., genomic DNA) with a Tn5-type transposase enzyme and a double-stranded oligonucleotide comprising a Tn transposon end sequence (SEQ ID NO: 211) linked to a universal adapter sequence or a sample index sequence under conditions suitable for forming a transposon synaptic complex. In the double-stranded oligonucleotide, the Tn transposon end sequence (SEQ ID NO: 211) can be located 5' or 3' to the universal adapter sequence or sample index sequence.
[0187] In some embodiments, the second splint strand sequence (400) of the covalently closed circular molecule comprises at least two subregions, including a first and a second subregion. In some embodiments, the first subregion comprises a universal binding sequence for the third surface primer, and the second subregion comprises a universal binding sequence for the fourth surface primer, and the first and second subregions do not hybridize to (or at least exhibit very little hybridization to) the first and second surface primers. In some embodiments, the second splint strand (400) further comprises an optional third subregion, which comprises a sample index sequence having 5-20 bases and / or a unique identification sequence (e.g., NN) having 2-10 or more bases. In some embodiments, the second splint strand (400) comprises only one subregion, lacking the second and third subregions, and the first subregion comprises an index sequence (e.g., a sample index sequence) having 5-20 bases. In some embodiments, the index sequence can be used in a multiplex assay to distinguish sequences of interest obtained from different sample sources. In some embodiments, the unique identification sequence comprises a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization to the first, second, third, and fourth surface primers. An exemplary arrangement of subregions in the second splint strand (400) in the 5' to 3' direction includes 5'-[second subregion]-[first subregion]-3' (e.g., FIG. 2). Another exemplary arrangement of subregions in the second splint strand (400) in the 5' to 3' direction includes 5'-[third subregion]-[second subregion]-[first subregion]-3' (e.g., FIG. 3).
[0188] In some embodiments, the second splint strand sequence (400) of the covalently closed circular library molecule (600) can hybridize to the first splint strand (300). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising at least two subregions, including a fourth and fifth subregion. The fourth subregion hybridizes to the first subregion of the second splint strand (400). The fifth subregion hybridizes to the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize to (or at least exhibit very little hybridization to) the first and second surface primers. In some embodiments, the internal region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes to the third subregion of the second splint strand (400). An exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes 5'-[fourth subregion]-[fifth subregion]-3' (e.g., Figure 2). Another exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes 5'-[fourth subregion]-[fifth subregion]-[sixth subregion]-3' (e.g., Figure 3).
[0189] In some embodiments, an exemplary covalently closed circular molecule (600) comprises (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer, (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer, (iii) a sequence of interest (110), (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer, (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer, and (vi) a second splint strand sequence (400), wherein the covalently closed circular molecule (600) is optionally hybridized to the first splint strand (300).
[0190] In the exemplary covalently closed circular molecule (600), the second splint strand region (400) comprises at least two subregions, including a first and a second subregion. The first subregion comprises a universal binding sequence for the third surface primer, and the second subregion comprises a universal binding sequence for the fourth surface primer, with the first and second subregions not hybridizing to (or at least exhibiting very little hybridization to) the first and second surface primers. In some embodiments, the second splint strand (400) further comprises an optional third subregion, which comprises a sample index sequence having 5-20 bases and / or a unique identification sequence (e.g., NN) having 2-10 or more bases. In some embodiments, the second splint strand (400) comprises only one subregion, lacking the second and third subregions, and the first subregion comprises a sample index sequence having 5-20 bases. In some embodiments, the sample index sequence can be used in multiplex assays to distinguish sequences of interest obtained from different sample sources. In some embodiments, the unique identification sequence comprises a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization to the first, second, third, and fourth surface primers. An exemplary arrangement of subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[second subregion]-[first subregion]-3'. Another exemplary arrangement of subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[third subregion]-[second subregion]-[first subregion]-3'.
[0191] The present disclosure provides a plurality of covalently closed circular library molecules as described herein. In some embodiments of the plurality of covalently closed circular library molecules (600), the target sequences (110) of the individual covalently closed circular molecules (600) in the plurality of covalently closed circular molecules (600) comprise the same target sequence or different target sequences.
[0192] Covalently closed circular molecules formed by double-stranded adaptors with truncated long splint strands The present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), at least a first left universal adaptor sequence (120), at least a first right universal adaptor sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9. In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adaptor sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adaptor sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adaptor sequences.
[0193] In some embodiments, the covalently closed circular library molecule (600) is hybridized to a first splint strand. In some embodiments, the first splint strand (300) comprises a truncated strand having a first region (320) with a truncation sequence at its 5' end (e.g., as shown in Figure 11B, Figure 11A, showing an exemplary truncation compared to SEQ ID NO: 199). In some embodiments, the 5' end of the first region can have a truncation of any length, e.g., a truncation of 1 to 10 nucleotides. In some embodiments, the truncated first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which are untruncated and do not possess any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the cleaved first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) as part of the covalently closed circular library molecule (600) (e.g., FIG. 9). In some embodiments, the cleaved first splint strand comprises (300), which comprises a cleaved first region (320) that hybridizes to a sequence (e.g., 120) as part of the covalently closed circular library molecule (600), and a second region (330) that hybridizes to a sequence (e.g., 130) as part of the covalently closed circular library molecule (600).
[0194] Covalently closed circular molecules formed by double-stranded adapters with long splint strands containing mismatched sequences The present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), at least a first left universal adaptor sequence (120), at least a first right universal adaptor sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9. In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adaptor sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adaptor sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adaptor sequences.
[0195] In some embodiments, the covalently closed circular library molecule (600) hybridizes to a first splint strand. In some embodiments, the first splint strand (300) comprises a mismatched strand having a first region (320) with a mismatched sequence within the first region (320) (e.g., Figure 11C). In some embodiments, the mismatched sequence can be of any length (e.g., 2-20 bases) and includes any sequence that is not perfectly complementary to the left universal adaptor sequence (120) of the library molecule (100). Some embodiments of the mismatched sequence in the first region (320) are shown in lowercase and underlined in Figure 11C. In some embodiments, the mismatched first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which do not possess any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) as part of the covalently closed circular library molecule (600) (e.g., FIG. 9). In some embodiments, the mismatched first splint strand comprises (300), which comprises a mismatched first region (320) that hybridizes to a sequence (e.g., 120) as part of the covalently closed circular library molecule (600), and a second region (330) that hybridizes to a sequence (e.g., 130) as part of the covalently closed circular library molecule (600). In some embodiments, the first region of mismatch (320) can hybridize with the first left universal adaptor sequence (120) of the covalently closed circular library molecule (600) to form a double-stranded portion having a bubble at the position of the mismatch sequence in the first region (320).
[0196] Covalently closed circular molecules formed by double-stranded adapters with long splint strands containing abasic sites The present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), at least a first left universal adaptor sequence (120), at least a first right universal adaptor sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9. In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adaptor sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adaptor sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adaptor sequences.
[0197] In some embodiments, the covalently closed circular library molecule (600) hybridizes to a first splint strand. In some embodiments, the first splint strand (300) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site within the fourth subregion and / or at least one abasic site within the fifth subregion (e.g., in the top schematic diagram of Figure 11D, the abasic sites are shown as solid black bars). In some embodiments, the abasic sites each comprise 1',2'-dideoxyribose (e.g., dSpacer from Integrated DNA Technologies (IDT)). In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO: 4) and a second region ((330); e.g., SEQ ID NO: 5), which do not possess any abasic sites and / or any sequence variants, such as, for example, insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) as part of a covalently closed circular library molecule (600) (e.g., Figure 9). In some embodiments, the abasic first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) as part of a covalently closed circular library molecule (600), and a second region (330) that hybridizes to a sequence (e.g., 130) as part of a covalently closed circular library molecule (600).
[0198] Covalently closed circular molecules formed by double-stranded adaptors with long splint strands containing uracil The present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), at least a first left universal adaptor sequence (120), at least a first right universal adaptor sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9. In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adaptor sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adaptor sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adaptor sequences.
[0199] In some embodiments, the covalently closed circular library molecule (600) hybridizes to a first splint strand. In some embodiments, the first splint strand (300) comprises at least one uracil. In some embodiments, the first splint strand (300) comprises at least one uracil in any one or any combination of the regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base may be substituted with uracil. One embodiment of a first splint strand containing uracil is shown in Figure 11D (bottom schematic). One skilled in the art will recognize that many other arrangements of the first splint strand (300) comprising one or more uracils are possible. In some embodiments, the uracil-containing first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) as part of the covalently closed circular library molecule (600) (e.g., FIG. 9). In some embodiments, the uracil-containing first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) as part of the covalently closed circular library molecule (600), and a second region (330) that hybridizes to a sequence (e.g., 130) as part of the covalently closed circular library molecule (600).
[0200] Covalently closed circular molecules formed by double-stranded adapters with short splint strands having random and index sequences The present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), at least a first left universal adaptor sequence (120), at least a first right universal adaptor sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9. In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adaptor sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adaptor sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adaptor sequences.
[0201] In some embodiments, the covalently closed circular library molecule (600) comprises a second splint strand sequence (400) covalently linked to a first left universal adaptor sequence (120) and a first right universal adaptor sequence (130) (e.g., Figures 7A and 9). In some embodiments, the covalently closed circular library molecule (600) is hybridized to a first splint strand (300). In some embodiments, the second splint strand (400) comprises a random sequence inserted into a first subregion of the second splint strand (e.g., Figures 7A, 12A, and 12B). In some embodiments, the random sequence can replace a portion of the first subregion of the second splint strand (e.g., Figures 7A, 13A, and 13B). In some embodiments, the second subregion of the second splint strand (400) does not have a random sequence inserted. In some embodiments, a portion of the second subregion of the second splint strand (400) is not replaced with a random sequence.
[0202] In some embodiments, the random sequence can be any length, e.g., 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figures 12A and 13A) or 4 nucleotides in length (e.g., "NNNN" in Figures 12B and 13B). In some embodiments, the random sequence can be inserted at any position within the first subregion of the second splint strand (400).
[0203] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0204] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction. In some embodiments, the random sequences provide high nucleotide diversity, with all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run being present in approximately equal proportions. In some embodiments, the sequences of the left index (160), the right index (170), and / or any portion of the insertion region (110) do not provide sufficient nucleotide diversity to enable polony mapping and / or template registration. In some embodiments, the random sequences provide a higher level of nucleotide diversity than the left index (160), the right index (170), and / or any portion of the insertion region (110). In some embodiments, the covalently closed circular library molecules can be subjected to a rolling circle amplification reaction to generate concatemers immobilized on a support. The concatemers contain tandem repeat sequences of the circular library molecules, including any insert sequences and adapter sequences (e.g., random sequences) present in the original circularized nucleic acid template molecule. In some embodiments, the random index sequence can be sequenced prior to sequencing the insertion region, and the random index is located in the first subregion of the second splint strand (400).
[0205] In some embodiments, the random sequence may be sequenced prior to sequencing the insertion region, and in some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration.
[0206] In some embodiments, a predetermined sequence is inserted into the sequence of the fourth subregion of the first splint strand (300) (e.g., Figure 7A). The length of the inserted predetermined sequence can be the same length as the random sequence inserted into the first subregion of the second splint strand (e.g., Figures 12A and 12B). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 12A and 12B.
[0207] In some embodiments, a portion of the fourth subregion of the first splint strand (300) is replaced with a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth subregion is the same as the random sequence replacing the portion of the first subregion of the second splint strand (e.g., Figures 7A, 13A, and 13B). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 13A and 13B.
[0208] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth and fifth subregions of the first splint strand (300) as part of the covalently closed circular library molecule (600) (e.g., Figures 7A, 12A, 12B, 13A, and 13B) and may contain bubbles at the location of the inserted or replacement random sequence.
[0209] Covalently closed circular molecules formed by double-stranded adapters with short splint strands containing appended random and index sequences The present disclosure provides a covalently closed circular library molecule (600) comprising a sequence of interest (110), at least a first left universal adaptor sequence (120), at least a first right universal adaptor sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9. In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adaptor sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adaptor sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adaptor sequences.
[0210] In some embodiments, the covalently closed circular library molecule (600) comprises a second splint strand sequence (400) covalently linked to a first left universal adaptor sequence (120) and a first right universal adaptor sequence (130) (e.g., Figures 7B and 9). In some embodiments, the covalently closed circular library molecule (600) is hybridized to a first splint strand (300). In some embodiments, the second splint strand (400) comprises a random sequence added to the 3' end of a first subregion of the second splint strand (e.g., Figures 7B, 14A, and 14B). In some embodiments, the random sequence further comprises an index sequence. In some embodiments, the second subregion of the second splint strand (400) does not have an added random sequence.
[0211] In some embodiments, the added random sequence can be any length, for example, 2 to 10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figure 14A) or 4 nucleotides in length (e.g., "NNNN" in Figure 14B).
[0212] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0213] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction, hi some embodiments, the random sequences provide high nucleotide diversity, with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run.
[0214] In some embodiments, the random sequence can be sequenced before sequencing the insert region. In some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration. In some embodiments, the sequences of the left index (160), the right index (170), and / or any portion of the insert region (110) do not provide sufficient nucleotide diversity to enable polony mapping and / or template registration. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to the left index (160), the right index (170), and / or any portion of the insert region (110). In some embodiments, the covalently closed circular library molecule can be subjected to a rolling circle amplification reaction to generate concatemers immobilized on a support. The concatemers contain tandem repeat sequences of the circular library molecule, including any insert sequence and adapter sequence (e.g., random sequence) present in the original circularized nucleic acid template molecule. In some embodiments, the random index sequence can be sequenced prior to sequencing the insert region, with the random index located at the 3' end of the first subregion of the second splint strand (400).
[0215] In some embodiments, index sequences can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay.
[0216] In some embodiments, the predetermined sequence is added to the 5' end of the fourth subregion of the first splint strand (300) (e.g., Figure 7B). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 7B, 14A, and 14B). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence and index sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 7B, 14A, and 14B). The added predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 14A and 14B.
[0217] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth and fifth subregions of the first splint strand (300) as part of the covalently closed circular library molecule (600) (e.g., Figures 7B, 14A, and 14B) and may contain bubbles at the location of the added random sequence.
[0218] Short sprint chain arrangement In some embodiments of the covalently closed circular molecule (600) described herein, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises first and second subregions comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A. In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or non-phosphorylated.
[0219] Long sprint chain arrangementIn some embodiments of the covalently closed circular molecule (600) described herein, the first region (320) of the first splint strand comprises a first universal adapter sequence that includes a universal binding sequence for a first surface primer (or its complementary sequence), and the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complementary sequence of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence that includes a universal binding sequence for a second surface primer (or its complementary sequence), and the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize to a P7 surface primer or the complementary sequence of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fourth subregion having the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198).In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) having fourth and fifth subregions, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAGGCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11A. In some embodiments, the 5' end of the first splint strand (300) can be phosphorylated or unphosphorylated. In some embodiments, the first subregion of the second splint strand (400) can hybridize to the fourth subregion of the first splint strand (300). In some embodiments, the second subregion of the second splint strand (400) can hybridize to the fifth subregion of the first splint strand (300).
[0220] Covalently closed circular molecular arrangements In some embodiments of the covalently closed circular molecule (600) described herein, the first region (320) of the first splint strand comprises a sequence capable of binding to the first left universal adaptor sequence (120) of the library molecule, and the first region (320) of the first splint strand comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 215) or a complementary sequence thereof.
[0221] In some embodiments of the covalently closed circular molecule (600) described herein, the second region (330) of the first splint strand comprises a sequence capable of binding to the first right universal adaptor sequence (130) of the library molecule, and the second region (330) of the first splint strand comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO: 216) or a complementary sequence thereof.
[0222] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first left universal adaptor sequence (120) comprising the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203).
[0223] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first left universal adaptor sequence (120) that binds to a first region (320) of a first splint strand, wherein the first left universal adaptor sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 213) or a complementary sequence thereof.
[0224] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 204).
[0225] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal binding sequence comprises the sequence 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 207).
[0226] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO: 208).
[0227] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer, wherein the first right universal adapter sequence comprises the sequence 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO: 205).
[0228] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer, wherein the first right universal adapter sequence comprises the sequence 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO: 209).
[0229] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer, wherein the first right universal adapter sequence comprises the sequence 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO: 210).
[0230] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adaptor sequence (130) comprising the sequence 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 206).
[0231] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adaptor sequence (130) that binds to a first region of the first splint strand (330), and the first left universal adaptor sequence (130) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 214) or a complementary sequence thereof.
[0232] The present disclosure provides a reaction mixture comprising a plurality of any of the covalently closed circular molecules (600) described herein and at least one exonuclease enzyme, in some embodiments, the exonuclease enzyme comprises any one or any combination of two or more of Exonuclease I, Thermolabile Exonuclease I, and / or T7 Exonuclease.
[0233] Kit containing a double-stranded splint adapter The present disclosure provides kits for use in introducing one or more new adapter sequences into linear nucleic acid library molecules. In some embodiments, the kits can be used to circularize single-stranded nucleic acid library molecules having a sequence of interest (110) flanked on one side by at least a first left universal adapter sequence (120) and on the other side by at least a first right universal adapter sequence (130). In some embodiments, the circularized library molecules can be converted into covalently closed circular molecules, which can be subjected to a rolling circle amplification (RCA) reaction to generate nucleic acid concatemers. The concatemers can be immobilized on a support for massively parallel sequencing.
[0234] The present disclosure provides kits including a nucleic acid double-stranded splint adapter (200), comprising (i) a first splint strand (long splint strand (300)) hybridizing to (ii) a second splint strand (short splint strand (400)). In some embodiments, the first splint strand comprises a first region (320), an internal region (310), and a second region (330). The internal region (310) of the first splint strand hybridizes to the second splint strand (400) to form a double-stranded splint adapter (200) having a double-stranded region and two adjacent single-stranded regions. The second splint strand (400) comprises a new adapter sequence that can be introduced into a linear nucleic acid library molecule. Exemplary double-stranded splint adapters are shown in Figures 1-8. The kit may include a container containing a first splint strand (300) hybridized to a second splint strand (400). The kit may include a first container containing the first splint strand (300) and a second container containing the second splint strand (400).
[0235] In some embodiments of the kits of the present disclosure, the second splint strand (400) comprises at least two subregions, including a first and a second subregion (see, e.g., Figures 2 and 3). The first subregion comprises a universal binding sequence for the third surface primer, and the second subregion comprises a universal binding sequence for the fourth surface primer, and the first and second subregions do not hybridize to (or at least exhibit very little hybridization to) the first and second surface primers. In some embodiments, the second splint strand (400) further comprises an optional third subregion, which comprises a sample index sequence having 5-20 bases and / or a unique identification sequence (e.g., NN) having 2-10 or more bases (see, e.g., Figure 3). In some embodiments, the second splint strand (400) includes only one subregion and lacks the second and third subregions, and the first subregion includes an index sequence (e.g., a sample index) having 5 to 20 bases. In some embodiments, the sample index sequence can be used in multiplex assays to distinguish sequences of interest obtained from different sample sources. In some embodiments, the unique identification sequence includes a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization to the first, second, third, and fourth surface primers. An exemplary arrangement of the subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[second subregion]-[first subregion]-3'. Another exemplary arrangement of subregions in the second splint strand (400) in the 5' to 3' direction includes 5'-[third subregion]-[second subregion]-[first subregion]-3'. Exemplary first splint strands (300) and second splint strands (400) are shown in Figures 2 and 3. In some embodiments, the second splint strand (400) can be 20 to 100 nucleotides in length, or 30 to 80 nucleotides in length, or 40 to 60 nucleotides in length.In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0236] In some embodiments of the kits of the present disclosure, the first splint strand (300) comprises a first region (320), a second region (330), and an internal region (310). The first region (320) comprises a first universal adapter sequence capable of hybridizing to a first universal binding sequence at one end of a linear nucleic acid library molecule. The second region (330) comprises a second universal adapter sequence capable of hybridizing to a second universal binding sequence at the other end of the linear nucleic acid library molecule. In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the first splint strand (300) can be 50 to 150 nucleotides in length, or 60 to 100 nucleotides in length, or 70 to 90 nucleotides in length. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at the 5' and / or 3' ends to confer exonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more phosphorothioate linkages at internal positions to confer endonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' ends or at an internal position.In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or non-phosphorylated, and in some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0237] In some embodiments of the kits of the present disclosure, the first splint strand (300) comprises an internal region (310) comprising at least two subregions, including a fourth and fifth subregion (e.g., Figures 2 and 3). The fourth subregion hybridizes to the first subregion of the second splint strand (400). The fifth subregion hybridizes to the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize to (or at least show very little hybridization to) the first and second surface primers. In some embodiments, the internal region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes to the third subregion of the second splint strand (400). An exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes: 5'-[fourth subregion]-[fifth subregion]-3'. Another exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes: 5'-[fourth subregion]-[fifth subregion]-[sixth subregion]-3'. An exemplary first splint strand (300) is shown in Figures 2 and 3.
[0238] In some embodiments of the kits described herein, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises first and second subregions comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A. In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or non-phosphorylated.
[0239] In some embodiments of the kits described herein, the first region (320) of the first splint strand comprises a first universal adapter sequence that includes a universal binding sequence for a first surface primer (or its complementary sequence), and the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complementary sequence of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence that includes a universal binding sequence for a second surface primer (or its complementary sequence), and the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize to a P7 surface primer or the complementary sequence of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fourth subregion having the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198).In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) having fourth and fifth subregions, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAGGCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11A. In some embodiments, the 5' end of the first splint strand (300) can be phosphorylated or unphosphorylated. In some embodiments, the first subregion of the second splint strand (400) can hybridize to the fourth subregion of the first splint strand (300). In some embodiments, the second subregion of the second splint strand (400) can hybridize to the fifth subregion of the first splint strand (300).
[0240] In some embodiments of the kits described herein, the first region of the first splint strand (320) comprises a sequence capable of binding to the first left universal adapter sequence (120) of the library molecule, and the first region of the first splint strand (320) comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 215) or a complementary sequence thereof.
[0241] In some embodiments of the kits described herein, the second region of the first splint strand (330) comprises a sequence capable of binding to the first right universal adapter sequence (130) of the library molecule, and the second region of the first splint strand (330) comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO: 216) or a complementary sequence thereof.
[0242] In some embodiments, the kit includes an adapter having a first left universal adapter sequence (120) that binds to a first region (320) of a first splint strand for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0243] In some embodiments of the kits described herein, the library molecule comprises a first left universal adaptor sequence (120) that binds to a first region (320) of the first splint strand, and the first left universal adaptor sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 213) or a complementary sequence thereof.
[0244] In some embodiments, the kit includes an adapter having a second left universal adapter sequence (140) for a sequencing primer for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 204). In some embodiments, the adapter having the second left universal adapter sequence (140) for a sequencing primer also comprises a first left index sequence (160). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0245] In some embodiments, the kit includes an adapter having a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 207). In some embodiments, the adapter having the second left universal adapter sequence (140) for the sequencing primer also comprises a first left index sequence (160). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0246] In some embodiments, the kit includes an adapter having a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO: 208). In some embodiments, the adapter having the second left universal adapter sequence (140) for the sequencing primer also comprises a first left index sequence (160). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0247] In some embodiments, the kit includes an adapter having a second right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO: 205). In some embodiments, the adapter having the second right universal adapter sequence (150) for the sequencing primer also comprises a first right index sequence (170). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0248] In some embodiments, the kit includes an adapter having a second right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO: 209). In some embodiments, the adapter having the second right universal adapter sequence (150) for the sequencing primer also comprises a first right index sequence (170). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0249] In some embodiments, the kit includes an adapter having a second right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO: 210). In some embodiments, the adapter having the second right universal adapter sequence (150) for the sequencing primer also comprises a first right index sequence (170). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0250] In some embodiments, the kit includes an adapter having a first right universal adapter sequence (130) that binds to a first region of a first splint strand (330) for use in preparing a plurality of library molecules, the library molecules comprising the sequence 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 206). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.
[0251] In some embodiments of the kits described herein, the library molecule comprises a first right universal adaptor sequence (130) that binds to a first region of the first splint strand (330), and the first right universal adaptor sequence (130) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 214) or a complementary sequence thereof.
[0252] In some embodiments, the kit includes a plurality of polynucleotides comprising a first left index sequence (160) and / or a plurality of first right index sequences (170). In some embodiments, the kit can include separate containers holding polynucleotides comprising individual first left index (160) or individual first right index (170) sequences. In some embodiments, the kit can include separate containers holding pairs of polynucleotides comprising individual first left index (160) and individual first right index (170) sequences. In some embodiments, the kit contains polynucleotides comprising a first left index (160) and / or a plurality of first right index sequences (170) in a multiwell plate (e.g., a 96-well plate). A list of exemplary first left index sequences (160) and first right index sequences (170) is provided in Table 1 in Figure 33. The first left index sequence (160) may include a random sequence (e.g., NNN) or may lack a random sequence. The first right index sequence (170) may be random or may contain a sequence (eg, NNN), or may lack a random sequence.
[0253] In some embodiments, the kit comprises a nucleic acid double-stranded splint adaptor (200) and T4 polynucleotide kinase. In some embodiments, the kit comprises a ligase enzyme, wherein the ligase enzyme comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase. In some embodiments, the kit comprises at least one endonuclease, wherein the endonuclease comprises any one or any combination of two or more of exonuclease I, thermolabile exonuclease I, and / or T7 exonuclease.
[0254] In some embodiments, the kit includes at least one buffer for hybridizing a plurality of double-stranded splint adaptors (200) and a plurality of nucleic acid library molecules (100). In some embodiments, the kit includes one buffer for performing multiple enzymatic reactions in a single reaction vessel, the multiple enzymatic reactions including any combination of (i) phosphorylating the 5' ends of the first and / or second splint strands (e.g., (300) and / or (400)), (ii) ligating nicks in the library-splint complex (500), and / or (iii) exonuclease digestion of the first splint strand (300) from the covalently closed circular molecule (600). Alternatively, the kit includes two or more separate buffers, where a first buffer can be used to perform the phosphorylation reaction, a second buffer can be used to perform the ligation reaction, and a third buffer can be used to perform the exonuclease digestion reaction.
[0255] In some embodiments, the kit includes one or more containers containing any of the double-stranded splint adapters (200) described herein, or any of the first and second splint strands (300) and (400) described herein. The kit may further include one or more containers containing T4 polynucleotide kinase, at least one ligase, and / or at least one exonuclease. The kit can include any of these components in any combination, and may be contained in a single container, or in separate containers, or any combination thereof.
[0256] The kit can include instructions for using the kit to perform the reaction to introduce one or more new adapter sequences into the linear nucleic acid library molecule.
[0257] The kit may include polynucleotides encoding one or more exemplary sequences of interest for use as positive controls.
[0258] Methods for forming multiple library-splint complexes The present disclosure provides a method for forming a plurality of library-sprint complexes (500), the method comprising: (a) providing a plurality of double-stranded splint adapters (200), each double-stranded splint adapter (200) in the plurality of double-stranded splint adapters (200) comprising a first splint strand (300) hybridized to a second splint strand (400), the double-stranded splint adapter comprising a double-stranded region and two adjacent single-stranded regions, the first splint strand comprising a first region (320), an internal region (310), and a second region (330), the internal region (310) of the first splint strand hybridizing to the second splint strand (400). Exemplary double-stranded splint adapters (200) are shown in Figures 1-8.
[0259] In some embodiments, the method for forming a plurality of library-sprint complexes (500) includes step (b): hybridizing a plurality of double-stranded splint adapters to a plurality of single-stranded nucleic acid library molecules (100), each library molecule comprising a sequence of interest (110) flanked on one side by at least a first left universal adapter sequence (120) and on the other side by at least a first right universal adapter sequence (130) (e.g., Figures 1-8). The hybridizing is performed under conditions suitable for hybridizing a first region (320) of the first splint strand to at least the first left universal adapter sequence (120) of the library molecule and a second region (330) of the first splint strand to at least the first right universal sequence (130) of the library molecule, thereby circularizing the plurality of library molecules to form a plurality of library-sprint complexes (500).
[0260] In some embodiments of the method for forming a plurality of library-splint complexes (500), the first region (320) of the first splint strand comprises a first universal adaptor sequence capable of hybridizing to a first universal binding sequence at one end of a linear nucleic acid library molecule. In some embodiments, the first region (320) of the first splint strand comprises a first universal adaptor sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0261] In some embodiments of the method for forming a plurality of library-splint complexes (500), the second region (330) of the first splint strand comprises a second universal adaptor sequence capable of hybridizing to a second universal binding sequence at the other end of the linear nucleic acid library molecule. In some embodiments, the second region (330) of the first splint strand comprises a second universal adaptor sequence comprising a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0262] In some embodiments of the method for forming a plurality of library-sprint complexes (500), the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal sequence (130) of the library molecule, thereby circularizing the library molecule to generate the library-sprint complex (500). The library-sprint complex (500) comprises a first nick between the 5' end of the library molecule and the 3' end of the second splint strand (e.g., Figures 1-8). The library-sprint complex (500) also comprises a second nick between the 5' end of the second splint strand and the 3' end of the library molecule (e.g., Figures 1-8). In some embodiments, the first and second nicks are enzymatically ligatable.
[0263] In some embodiments of the method for forming a plurality of library-sprint complexes (500), the first region (320) of the first splint strand can hybridize to the sense strand or the antisense strand of a double-stranded nucleic acid library molecule. In the library-sprint complex (500), the second region (330) of the first splint strand can hybridize to the sense strand or the antisense strand of a double-stranded nucleic acid library molecule. The double-stranded nucleic acid library molecule can be denatured to generate single-stranded sense and antisense library strands.
[0264] In some embodiments of the method for forming multiple library-sprint complexes (500), the second splint strand (400) does not hybridize to the sequence of interest (110), and the internal region (310) of the first splint strand does not hybridize to the sequence of interest (110).
[0265] In some embodiments of the method for forming a plurality of library-sprint complexes (500), the first region (320) of the first splint strand does not hybridize to the sequence of interest (110), and the second region (330) of the first splint strand does not hybridize to the sequence of interest (110).
[0266] In some embodiments of the method for forming a plurality of library-splint complexes (500), the 5' ends of the single-stranded library molecules (100) are phosphorylated or lack a phosphate group. In some embodiments, the 3' ends of the single-stranded library molecules include a terminal 3' OH group or a terminal 3' blocking group.
[0267] In some embodiments of the method for forming a plurality of library-splint complexes (500), each individual nucleic acid library molecule (100) comprises a second left universal adaptor sequence (140). In some embodiments, each individual nucleic acid library molecule (100) comprises a second right universal adaptor sequence (150). In some embodiments, each nucleic acid library molecule (100) comprises an additional left universal adaptor sequence and / or a right universal adaptor sequence.
[0268] In some embodiments of the method for forming a plurality of library-splint complexes (500), the nucleic acid library molecule (100) comprises a first left index sequence (160). In some embodiments, the nucleic acid library molecule (100) comprises a first right index sequence (170). In some embodiments, the first left index sequence (160) comprises a sample index sequence. In some embodiments, the first right index sequence (170) comprises another sample index sequence. In some embodiments, the sequence of the first left index sequence is not the same as the sequence of the first right index sequence. The sample index sequence can be used in a multiplex assay to distinguish sequences of interest obtained from different sample sources. A list of exemplary first left index sequences (160) and first right index sequences (170) is provided in Table 1 in Figure 33. The first left index sequence (160) may comprise a random sequence (e.g., NNN) or may lack a random sequence. The first right index sequence (170) may include a random sequence (eg, NNN) or may lack a random sequence.
[0269] Multiplex workflows are enabled by preparing sample-indexed libraries using one or both index sequences (e.g., left index sequence and / or right index sequence). The first left index sequence (160) and / or the first right index sequence (170) can be used to prepare separate sample-indexed libraries using input nucleic acids isolated from different sources. The sample-indexed libraries can be pooled together to generate a multiplexed library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insert region, along with the first left index sequence (160) and / or the first right index sequence (170), can be used to identify the source of the input nucleic acid. In some embodiments, any number of sample-indexed libraries can be pooled together, for example, 2 to 10, 10 to 50, 50 to 100, 100 to 200, or more than 200 sample-indexed libraries can be pooled. Exemplary nucleic acid sources include naturally occurring sources, recombinant sources, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or whole organisms. Exemplary nucleic acid sources include fresh sources, frozen sources, fresh frozen sources, or archived (e.g., formalin-fixed, paraffin-embedded; FFPE) sources. Those skilled in the art will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.
[0270] In some embodiments of the method for forming a plurality of library-splint complexes (500), the nucleic acid library molecule (100) comprises a first left unique identification sequence (180). In some embodiments, the nucleic acid library molecule (100) comprises a first right unique identification sequence (190). In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence used to uniquely identify an individual sequence of interest (e.g., an insert sequence) to which a unique adaptor has been added in a population of other sequences of molecules of interest. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular tagging.
[0271] In some embodiments of the method for forming a plurality of library-sprint complexes (500), the nucleic acid library molecule (100) comprises any one or any combination of two or more of: a first left universal adaptor sequence (120), a second left universal adaptor sequence (140), a first left index sequence (160), a first left unique identifier sequence (180), a first right universal adaptor sequence (130), a second right universal adaptor sequence (150), a first right index sequence (170), and / or a first right unique identifier sequence (190).
[0272] In some embodiments of the method for forming a plurality of library-splint complexes (500), the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) comprise a universal binding sequence for a forward amplification primer or a reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward or reverse amplification primer, and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises an additional left universal adapter sequence.
[0273] In some embodiments of the method for forming a plurality of library-splint complexes (500), the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) comprises a universal binding sequence for a forward or reverse sequencing primer, a universal binding sequence for a first or second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises an additional right universal adapter sequence.
[0274] In some embodiments of the method for forming a plurality of library-splint complexes (500), the nucleic acid library molecule (100) comprises at least one junction adaptor sequence located between any of the universal adaptor sequences described herein (see, e.g., FIG. 8). For example, a first left junction adaptor sequence (125) can be located between the first left universal adaptor sequence (120) and the first left index sequence (160). A second left junction adaptor sequence (165) can be located between the first left index sequence (160) and the second left universal adaptor sequence (140). A third left junction adaptor sequence (145) can be located between the second left universal adaptor sequence (140) and the sequence of interest (110). A first right junction adaptor sequence (135) can be located between the first right universal sequence (130) and the first right index sequence (170). The second right junction adaptor sequence (175) can be located between the first right index sequence (170) and the second right universal adaptor sequence (150). The third right junction adaptor sequence (155) can be located between the second right universal adaptor sequence (150) and the sequence of interest (110). In some embodiments, the nucleic acid library molecule (100) further comprises at least 1 to up to 10 added universal adaptor sequences located 5' (upstream) of the first left universal adaptor sequence (120) (see, e.g., Figure 8). In some embodiments, the nucleic acid library molecule (100) further comprises at least 1 to up to 10 added universal adaptor sequences located 3' (downstream) of the first right universal sequence (130) (see, e.g., Figure 8). Any of the junction adaptor sequences and / or added universal adaptor sequences can comprise any sequence and can be 3 to 60 nucleotides in length. Either the junction adapter sequence and / or the added universal adapter sequence comprises a universal sequence or a unique sequence.Any of the junction adapter sequences and / or the added universal adapter sequences include a binding sequence for an amplification primer, a sequencing primer, or a compaction oligonucleotide, or a combination thereof. Any of the junction adapter sequences and / or the added universal adapter sequences include a binding sequence for an immobilized surface primer (e.g., a capture primer). Any of the junction adapter sequences and / or the added universal adapter sequences include a sample index sequence. Any of the junction adapter sequences and / or the added universal adapter sequences include a unique identification sequence. Any of the junction adapter sequences and / or the added universal adapter sequences, particularly junction adapter sequence (145), include the Tn5 transposon end sequence 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 211). Any of the junction adapter sequences and / or the added universal adapter sequences, particularly junction adapter sequence (155), include the Tn5 transposon end sequence 5'-CTGTCTCTTATACACATCT-3' (SEQ ID NO: 212). The Tn5 transposon end sequence can be introduced into the library molecule (100) via a transposase-mediated reaction, which involves contacting double-stranded input DNA (e.g., genomic DNA) with a Tn5-type transposase enzyme and a double-stranded oligonucleotide comprising a Tn transposon end sequence (SEQ ID NO: 211) linked to a universal adapter sequence or a sample index sequence under conditions suitable for forming a transposon synaptic complex. In the double-stranded oligonucleotide, the Tn transposon end sequence (SEQ ID NO: 211) can be located 5' or 3' to the universal adapter sequence or sample index sequence.
[0275] In some embodiments of the method for forming a plurality of library-splint complexes (500), the second splint strand (400) comprises at least two subregions, including a first and a second subregion (e.g., Figures 2 and 3). The first subregion comprises a universal binding sequence for the third surface primer, and the second subregion comprises a universal binding sequence for the fourth surface primer, and the first and second subregions do not hybridize to (or at least exhibit very little hybridization to) the first and second surface primers. In some embodiments, the second splint strand (400) further comprises an optional third subregion, which comprises a sample index sequence having 5-20 bases and / or a unique identification sequence (e.g., NN) having 2-10 or more bases (see, e.g., Figure 3). In some embodiments, the second splint strand (400) includes only one subregion and lacks the second and third subregions, and the first subregion includes a sample index sequence having 5 to 20 bases. In some embodiments, the sample index sequence can be used in multiplex assays to distinguish sequences of interest obtained from different sample sources. In some embodiments, the unique identification sequence includes a random sequence. The unique identification sequence can be designed to exhibit reduced hybridization to, or no hybridization to, the first, second, third, and fourth surface primers. An exemplary arrangement of the subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[second subregion]-[first subregion]-3'. Another exemplary arrangement of the subregions in the second splint strand (400) in the 5' to 3' direction includes: 5'-[third subregion]-[second subregion]-[first subregion]-3'. In some embodiments, the second splint strand (400) can be 20 to 100 nucleotides in length, or 30 to 80 nucleotides in length, or 40 to 60 nucleotides in length.In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at the 5' and / or 3' end to confer exonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more phosphorothioate linkages at an internal position to confer endonuclease resistance. In some embodiments, the second splint strand (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.
[0276] In some embodiments of the method for forming a plurality of library-sprint complexes (500), the first splint strand (300) comprises an internal region (310) comprising at least two subregions, including a fourth and fifth subregion (e.g., Figures 2 and 3). The fourth subregion hybridizes to the first subregion of the second splint strand (400). The fifth subregion hybridizes to the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize to (or at least exhibit very little hybridization to) the first and second surface primers. In some embodiments, the internal region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes to the third subregion of the second splint strand (400) (see, e.g., Figure 3). An exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes 5'-[fourth subregion]-[fifth subregion]-5'. Another exemplary arrangement of the subregions of the first splint strand (300) in the 5' to 3' direction includes 5'-[fourth subregion]-[fifth subregion]-[sixth subregion]-3'. In some embodiments, the first splint strand (300) can be 50 to 150 nucleotides in length, or 60 to 100 nucleotides in length, or 70 to 90 nucleotides in length. In some embodiments, the first splint strand (300) includes one or more phosphorothioate linkages at the 5' and / or 3' ends to confer exonuclease resistance. In some embodiments, the first splint strand (300) includes one or more phosphorothioate linkages at internal positions to confer endonuclease resistance. In some embodiments, the first splint strand (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' ends or at an internal position.
[0277] The present disclosure provides a method for forming a plurality of library-sprint complexes (500), the method comprising: (a) providing a plurality of double-stranded splint adapters (200), each double-stranded splint adapter (200) comprising a first splint strand (300) hybridized to a second splint strand (400), the first splint strand (300) comprising a first region (320), an internal region (310), and a second region (330) arranged in 5' to 3' order, the internal region (310) of the first splint strand hybridizing to the second splint strand (400), the second splint strand comprising, in 5' to 3' order, (i) a second subregion having a universal binding sequence for a fourth surface primer, and (ii) a first subregion having a universal binding sequence for a third surface primer.In some embodiments, the method for forming a plurality of library-splint complexes (500) comprises step (b): hybridizing a plurality of double-stranded splint adapters to a plurality of single-stranded nucleic acid library molecules (100), each library molecule comprising regions arranged in 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer; (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; (iii) a sequence of interest (110); (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer; and (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer (130), wherein hybridizing comprises hybridizing a first splint adapter to a plurality of single-stranded nucleic acid library molecules (100), each library molecule comprising regions arranged in 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer; (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; (iii) a sequence of interest (110); (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer; and (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer (130), wherein the hybridizing comprises hybridizing a first splint adapter to a plurality of double-stranded splint adapters to a plurality of single-stranded nucleic acid library molecules (100), each library molecule comprising regions arranged in 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer; (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; The method further includes hybridizing the first splint strand (300) to the library molecule (100) under conditions suitable for hybridizing the first splint strand to the library molecule (100), thereby circularizing the library molecule and producing a library-sprint complex (500), whereby a first region (320) of the first splint strand hybridizes to the binding sequence for the first surface primer (120), a third region (330) of the first splint strand hybridizes to the binding sequence for the second surface primer (130), the library-sprint complex (500) comprising a first nick between the 5' end of the library molecule and the 3' end of the second splint strand (300), and the library-sprint complex (500) comprising a second nick between the 5' end of the second splint strand (300) and the 3' end of the library molecule (100), wherein the first and second nicks are enzymatically ligatable. In some embodiments, the plurality of single-stranded nucleic acid library molecules (100) comprises a first left indexing sequence (160) and / or a first right indexing sequence (170) (see, e.g., Figure 5). A list of exemplary first left indexing sequences (160) and first right indexing sequences (170) is provided in Table 1 in Figure 33.In some embodiments, the first left index sequence (160) comprises or lacks a short random sequence (e.g., NNN). In some embodiments, the first right index sequence (170) comprises or lacks a short random sequence (e.g., NNN). In some embodiments, the plurality of single-stranded nucleic acid library molecules (100) comprises a first left unique identification sequence (180) and / or a first right unique identification sequence (190), each of which comprises a sequence used to uniquely identify an individual sequence of interest (e.g., an insert sequence) to which a unique adaptor has been added in a population of other sequences of molecules of interest. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular tagging. (See, e.g., Figure 6 ).
[0278] Multiplex workflows are enabled by preparing sample-indexed libraries using one or both index sequences (e.g., left index sequence and / or right index sequence). The first left index sequence (160) and / or the first right index sequence (170) can be used to prepare separate sample-indexed libraries using input nucleic acids isolated from different sources. The sample-indexed libraries can be pooled together to generate a multiplexed library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insert region, along with the first left index sequence (160) and / or the first right index sequence (170), can be used to identify the source of the input nucleic acid. In some embodiments, any number of sample-indexed libraries can be pooled together, for example, 2 to 10, 10 to 50, 50 to 100, 100 to 200, or more than 200 sample-indexed libraries can be pooled. Exemplary nucleic acid sources include naturally occurring sources, recombinant sources, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or whole organisms. Exemplary nucleic acid sources include fresh sources, frozen sources, fresh frozen sources, or archived (e.g., formalin-fixed, paraffin-embedded; FFPE) sources. Those skilled in the art will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.
[0279] In some embodiments, the plurality of single-stranded nucleic acid library molecules (100) further comprises a first left unique identification sequence (180) and / or a first right unique identification sequence (190) (see, e.g., FIG. 6). In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence used to uniquely identify an individual sequence of interest (e.g., an insert sequence) to which a unique adaptor has been added in a population of other sequences of molecules of interest. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) may be used for molecular tagging.
[0280] Methods for forming library-splint complexes using double-stranded adapters with cleaved long splints In some embodiments of the method for forming a library-splint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the first splint strand (300) of each double-stranded splint adaptor (200) comprises a truncated strand having a first region (320) with a truncation sequence at its 5' end (e.g., Figure 11B; compare, e.g., SEQ ID NO: 199 shown in Figure 11A). In some embodiments, the 5' end of the first region can have a truncation of any length, e.g., 1-10 nucleotides. In some embodiments, the truncated first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which are untruncated and do not possess any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the cleaved first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the cleaved first splint strand comprises (300), which comprises a cleaved first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the cleaved first splint strand, as part of the double-stranded splint adapter (200), can hybridize to the library molecule (100) to form a library-sprint complex (500).
[0281] Methods for forming library-splint complexes using double-stranded adapters with long splint strands having mismatched sequences In some embodiments of the method for forming a library-sprint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the first splint strand (300) of each double-stranded splint adaptor (200) comprises a first region (320) having a mismatch sequence within the first region (320) (e.g., Figure 11C). In some embodiments, the mismatch sequence can be of any length (e.g., 2-20 bases) and includes any sequence that is not perfectly complementary to the left universal adaptor sequence (120) of the library molecule (100). Some embodiments of the mismatch sequence in the first region (320) are shown in lowercase and underlined in Figure 11C. In some embodiments, the mismatched first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7), which do not possess any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint strand (400) to form the double-stranded splint adapter (200). In some embodiments, the mismatched first splint strand comprises (300), which comprises a first region of mismatch (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the first region of mismatch (320) can hybridize to the first left universal adaptor sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatched sequence in the first region (320).In some embodiments, the mismatched first splint strand can hybridize to the library molecule (100) as part of the double-stranded splint adaptor (200) to form a library-splint complex (500).
[0282] Methods for forming library-splint complexes using double-stranded adapters with long splint strands having abasic sites In some embodiments of the method for forming a library-sprint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adapters (200). In some embodiments, the first splint strand (300) of each double-stranded splint adapter (200) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site within the fourth subregion and / or at least one abasic site within the fifth subregion (e.g., in the top schematic diagram of Figure 11D, the abasic sites are shown as solid black bars). In some embodiments, the abasic sites each comprise 1',2'-dideoxyribose (e.g., dSpacer from Integrated DNA Technologies (IDT)). In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO: 4) and a second region ((330); e.g., SEQ ID NO: 5), which do not possess any abasic sites and / or any sequence variants, such as, for example, insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint strand (400) to form the double-stranded splint adapter (200). In some embodiments, the abasic first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the abasic first splint strand can hybridize to the library molecule (100) as part of the double-stranded splint adaptor (200) to form a library-sprint complex (500).
[0283] Methods for forming library-splint complexes using double-stranded adapters with long splint strands containing uracil In some embodiments of the method for forming a library-sprint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adapters (200). In some embodiments, the first splint strand (300) of each double-stranded splint adapter (200) comprises at least one uracil. In some embodiments, the first splint strand (300) of each double-stranded splint adapter (200) comprises at least one uracil in any one or any combination of the regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base can be substituted with uracil. One embodiment of a first splint strand containing uracil is shown in FIG. 11D (bottom schematic). Those skilled in the art will recognize that many other arrangements of first splint strands (300) comprising one or more uracils are possible. In some embodiments, the uracil-containing first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first and second subregions of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the uracil-containing first splint strand comprises (300), which comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the uracil-containing first splint strand, as part of the double-stranded splint adapter (200), can hybridize to the library molecule (100) to form a library-sprint complex (500).
[0284] Method for forming library-splint complexes using double-stranded adapters with short splint strands inserted with random and index sequences In some embodiments of the method for forming a library-sprint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the second splint strand (400) of each double-stranded splint adaptor (200) includes a random sequence inserted into a first subregion of the second splint strand (e.g., Figures 12A and 12B). In some embodiments, the random sequence can replace a portion of the first subregion of the second splint strand (e.g., Figures 13A and 13B). In some embodiments, the second subregion of the second splint strand (400) does not have a random sequence inserted. In some embodiments, a portion of the second subregion of the second splint strand (400) is not replaced with a random sequence.
[0285] In some embodiments, the random sequence can be any length, e.g., 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figures 12A and 13A) or 4 nucleotides in length (e.g., "NNNN" in Figures 12B and 13B). In some embodiments, the random sequence can be inserted at any position within the first subregion of the second splint strand (400).
[0286] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0287] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction, hi some embodiments, the random sequences provide high nucleotide diversity, with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run.
[0288] In some embodiments, the random sequence may be sequenced prior to sequencing the insertion region, and in some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration.
[0289] In some embodiments, a predetermined sequence is inserted into the sequence of the fourth subregion of the first splint strand (300). The length of the inserted predetermined sequence can be the same length as the random sequence inserted into the first subregion of the second splint strand (e.g., Figures 12A and 12B). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 12A and 12B.
[0290] In some embodiments, a portion of the fourth subregion of the first splint strand (300) is replaced with a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth subregion is the same as the random sequence replacing the portion of the first subregion of the second splint strand (e.g., Figures 13A and 13B). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 13A and 13B.
[0291] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth and fifth subregions of the first splint strand (300) to form the double-stranded splint adapter (200) (e.g., Figures 12A, 12B, 13A, and 13B). In some embodiments, the double-stranded splint adapter (200) forms a bubble at the location of the inserted or replacement random sequence.
[0292] In some embodiments, a second splint strand (400) bearing a random sequence can hybridize to a library molecule (100) as part of a double-stranded splint adaptor (200) to form a library-splint complex (500).
[0293] Method for forming library-splint complexes using double-stranded adapters with short splint strands appended with random and index sequences In some of the methods for forming a library-sprint complex (500) as described herein, a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the second splint strand (400) of each double-stranded splint adaptor (200) comprises a random sequence added to the 3' end of a first subregion of the second splint strand (e.g., Figures 14A and 14B). In some embodiments, the random sequence further comprises an index sequence. In some embodiments, the second subregion of the second splint strand (400) does not have a random sequence added.
[0294] In some embodiments, the added random sequence can be any length, for example, 2 to 10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., "NNN" in Figure 14A) or 4 nucleotides in length (e.g., "NNNN" in Figure 14B).
[0295] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeat sequences having two or three of the same nucleobases, e.g., AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of second splint strands (400) includes random sequences with high diversity sequences containing approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of the sequencing attempt.
[0296] In some embodiments, the random sequences provide nucleotide diversity and color balance for the sequencing reaction, hi some embodiments, the random sequences provide high nucleotide diversity, with approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) represented in each cycle of a sequencing run.
[0297] In some embodiments, the random sequence may be sequenced prior to sequencing the insertion region, and in some embodiments, the random sequence provides sufficient nucleotide diversity and color balance so that sequencing data from the random sequence can be used for polony mapping and / or template registration.
[0298] In some embodiments, index sequences can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay.
[0299] In some embodiments, the predetermined sequence is added to the 5' end of the fourth subregion of the first splint strand (300). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 14A and 14B). In some embodiments, the length of the added predetermined sequence can be the same length as the random sequence and index sequence added to the 3' end of the first subregion of the second splint strand (e.g., Figures 14A and 14B). The added predetermined sequence can have any sequence. Exemplary predetermined sequences are shown in Figures 14A and 14B.
[0300] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion that can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adapter (200) (e.g., Figures 14A and 14B). In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the location of the added random sequence. In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the location of the added random sequence and the index sequence.
[0301] In some embodiments, the second splint strand (400) with the added random sequence (and optionally an index sequence) can hybridize to the library molecule (100) as part of the double-stranded splint adapter (200) to form a library-sprint complex (500).
[0302] Short sprint chain arrangement In some embodiments of the methods for forming a plurality of library-splint complexes (500) described herein, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises first and second subregions comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A. In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or unphosphorylated.
[0303] Long sprint chain arrangement In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the first region (320) of the first splint strand comprises a first universal adapter sequence that includes a universal binding sequence for a first surface primer (or its complementary sequence), and the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complementary sequence of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence that includes a universal binding sequence for a second surface primer (or its complementary sequence), and the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize to a P7 surface primer or the complementary sequence of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fourth subregion having the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) that includes a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198).In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) having fourth and fifth subregions, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAGGCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11A. In some embodiments, the 5' end of the first splint strand (300) can be phosphorylated or unphosphorylated. In some embodiments, the first subregion of the second splint strand (400) can hybridize to the fourth subregion of the first splint strand (300). In some embodiments, the second subregion of the second splint strand (400) can hybridize to the fifth subregion of the first splint strand (300).
[0304] In some embodiments of the methods for forming multiple library-sprint complexes (500) described herein, the first region (320) of the first splint strand comprises a sequence capable of binding to the first left universal adaptor sequence (120) of the library molecule, and the first region (320) of the first splint strand comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 215) or a complementary sequence thereof.
[0305] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the second region (330) of the first splint strand comprises a sequence capable of binding to the first right universal adapter sequence (130) of the library molecule, and the second region (330) of the first splint strand comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (sequence number 216) or a complementary sequence thereof.
[0306] Library-Sprint Complex Sequence In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule comprises a first left universal adaptor sequence (120) that binds to a first region (320) of a first splint strand, and the first left universal adaptor sequence (120) comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203).
[0307] In some embodiments of the methods for forming a plurality of library-splint complexes (500) described herein, the library molecule comprises a first left universal adaptor sequence (120) that binds to a first region (320) of a first splint strand, and the first left universal adaptor sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 213) or a complementary sequence thereof.
[0308] In some embodiments of the methods for forming a plurality of library-sprint complexes (500) described herein, the library molecules comprise a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 204).
[0309] In some embodiments of the methods for forming a plurality of library-sprint complexes (500) described herein, the library molecules comprise a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 207).
[0310] In some embodiments of the methods for forming a plurality of library-sprint complexes (500) described herein, the library molecules comprise a second left universal adapter sequence (140) comprising a sequence for binding of a sequencing primer, wherein the second left universal adapter sequence comprises the sequence 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO: 208).
[0311] In some embodiments, in any of the methods for forming a plurality of library-sprint complexes (500) described herein, the library molecules comprise a second right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer, wherein the second right universal adapter sequence comprises the sequence 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO: 205).
[0312] In some embodiments of the methods for forming a plurality of library-sprint complexes (500) described herein, the library molecules comprise a second right universal adapter sequence (150) comprising a sequence for binding of a sequencing primer, wherein the second right universal adapter sequence comprises the sequence 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO: 209).
[0313] In some embodiments of the methods for forming a plurality of library-sprint complexes (500) described herein, the library molecules include a second right universal adapter sequence (150) that includes a sequence for binding of a sequencing primer, and the second right universal adapter sequence includes the sequence 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO: 210).
[0314] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule comprises a first right universal adaptor sequence (130) that binds to a first region (330) of a first splint strand, and the right universal binding sequence (130) comprises the...
Claims
1. A library-sprint complex (500), comprising: (i) a single-stranded nucleic acid library molecule (100) comprising a sequence of interest (110) flanked on one side by at least a first left universal adaptor sequence (120) and on the other side by at least a first right universal adaptor sequence (130); (ii) a double-stranded splint adapter (200) comprising a first splint strand (300) and a second splint strand (400), wherein the double-stranded splint adapter (200) comprises a double-stranded region and two single-stranded regions, one on each side of the double-stranded region, and the first splint strand comprises a first region (320), an internal region (310), and a second region (330); A library-sprint complex (500) in which the internal region (310) of the first splint strand hybridizes to the second splint strand (400), the first region (320) of the first splint strand hybridizes to the at least first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to the at least first right universal sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-sprint complex (500).
2. The library-sprint complex (500) of claim 1, wherein the nucleic acid library molecule (100) further comprises a second left universal adaptor sequence (140).
3. 3. The library-sprint complex (500) of claim 2, wherein the second left universal adaptor sequence (140) is between the at least first left universal adaptor sequence (120) and the sequence of interest (110).
4. The library-sprint complex (500) of any one of claims 1 to 3, wherein the nucleic acid library molecule (100) further comprises a second right universal adaptor sequence (150).
5. 5. The library-sprint complex (500) of claim 4, wherein the second right universal adaptor sequence (150) is between the sequence of interest (110) and the at least first right universal adaptor sequence (130).
6. The library-sprint complex (500) of any one of claims 1 to 5, wherein the nucleic acid library molecule (100) further comprises a first left index sequence (160).
7. 7. The library-sprint complex (500) of claim 6, wherein the first left index sequence (160) is between the at least first left universal adaptor sequence (120) and the sequence of interest (110).
8. The library-sprint complex (500) of any one of claims 1 to 7, wherein the nucleic acid library molecule (100) further comprises a first right index sequence (170).
9. 9. The library-sprint complex (500) of claim 8, wherein the first right index sequence (170) is between the second right universal adaptor sequence (150) and the at least first right universal adaptor sequence (130).
10. The library-sprint complex (500) of any one of claims 1 to 9, wherein the nucleic acid library molecule (100) further comprises a first left unique identifier sequence (180).
11. 11. The library-sprint complex (500) of claim 10, wherein the first left unique identification sequence (180) is between the at least first left universal adapter sequence (120) and the first left index sequence (160).
12. The library-sprint complex (500) of any one of claims 1 to 11, wherein the nucleic acid library molecule (100) further comprises a first right unique identification sequence (190).
13. 13. The library-sprint complex (500) of claim 12, wherein a first right unique identification sequence (190) is located between the first right index sequence (170) and the at least a first right universal adaptor sequence (130).
14. The nucleic acid library molecule (100) (i) a second left universal adaptor sequence (140); (ii) a second right universal adapter sequence (150); (iii) a first left index array (160); (iv) a first right index array (170); (v) a first left unique identification sequence (180), and / or (vi) A library-sprint complex (500) according to any one of claims 1 to 13, further comprising any one or any combination of two or more of the first right unique identification sequences (190).
15. The first left universal adapter array (120) and / or the second left universal adapter array (140) are (i) a universal binding sequence for the forward sequencing primer; (ii) a universal binding sequence for the reverse sequencing primer; (iii) a universal binding sequence for the first surface primer; (iv) a universal binding sequence for the second surface primer; (v) a universal binding sequence for the forward amplification primer; (vi) a universal binding sequence for the reverse amplification primer, and / or (vii) A library-sprint complex (500) according to any one of claims 1 to 14, comprising a universal binding sequence for compacted oligonucleotides.
16. The first right universal adapter array (130) and / or the second right universal adapter array (150) are (i) a universal binding sequence for the forward sequencing primer; (ii) a universal binding sequence for the reverse sequencing primer; (iii) a universal binding sequence for the first surface primer; (iv) a universal binding sequence for the second surface primer; (v) a universal binding sequence for the forward amplification primer; (vi) a universal binding sequence for the reverse amplification primer, and / or (vii) A library-sprint complex (500) according to any one of claims 1 to 15, comprising a universal binding sequence for compacted oligonucleotides.
17. The library-sprint complex (500) of claim 15 or 16, wherein the second splint strand (400) comprises at least two subregions, a first subregion comprising a universal binding sequence for a third surface primer and a second subregion comprising a universal binding sequence for a fourth surface primer, and wherein the first and second subregions do not hybridize to the first and second surface primers or exhibit very little hybridization to the first and second surface primers.
18. 18. The library-sprint complex (500) of claim 17, wherein the second splint strand (400) comprises an optional third subregion, the third subregion comprising a sample index sequence having 5 to 20 bases and / or a unique identification sequence having 2 to 10 or more bases.
19. 20. The library-sprint complex (500) of claim 18, wherein the unique identification sequence comprises a random sequence.
20. 17. The library-sprint complex of claim 15 or 16, wherein the first splint strand comprises an internal region comprising at least two subregions, a fourth subregion comprising a universal binding sequence for a third surface primer, the fourth subregion hybridizing to the first subregion of the second splint strand, a fifth subregion comprising a universal binding sequence for a fourth surface primer, the fifth subregion hybridizing to the second subregion of the second splint strand, and the fourth and fifth subregions do not hybridize to the first and second surface primers or show very little hybridization to the first and second surface primers.
21. The library-sprint complex (500) of claim 20, wherein the first splint strand (300) comprises an internal region (310) further comprising a sixth subregion, the sixth subregion comprising a sample index sequence having 5 to 20 bases and / or a unique identification sequence having 2 to 10 or more bases, and the sixth subregion hybridizes to the third subregion of the second splint strand (400).
22. 22. The library-sprint complex (500) of claim 21, wherein the unique identification sequence comprises a random sequence.
23. A library-sprint complex (500), comprising: a) a single-stranded nucleic acid library molecule (100) comprising, in 5' to 3' order, the following components: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer (120), (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer, (iii) a sequence of interest (110), (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer, and (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer (130); b) a first splint strand (300) comprising the following components arranged in 5' to 3' order: a first region (320), an inner region (310), and a second region (330); and c) a second splint strand (400) comprising subregions arranged in 3' to 5' order: a first subregion having a universal binding sequence for a third surface primer; and a second subregion having a universal binding sequence for a fourth surface primer, wherein the first splint strand (300) hybridizes to a portion of the library molecule (100), thereby circularizing the library molecule to form a library-sprint complex (500), whereby the first region (320) of the first splint strand hybridizes to the binding sequence for the first surface primer (120). a third region (330) of the first splint strand hybridizes to the binding sequence for the second surface primer (130); the second splint strand (400) hybridizes to the internal region (310) of the first splint strand (300); the library-sprint complex (500) comprises a first nick between the 5' end of the library molecule and the 3' end of the second splint strand; and the library-sprint complex (500) comprises a second nick between the 5' end of the second splint strand and the 3' end of the library molecule.
24. 24. The library splint complex (500) of claim 23, wherein the first and second nicks are enzymatically ligatable.
25. A plurality of library-sprint complexes, comprising the library-sprint complex (500) of any one of the preceding claims, wherein the sequences of interest (110) of individual library-sprint complexes in the plurality of library-sprint complexes comprise the same sequence of interest or different sequences of interest.
26. A method for producing a library-splint complex according to any one of claims 1 to 22, comprising: a. providing a plurality of single-stranded nucleic acid library molecules (100); b. Providing a plurality of double-stranded splint adapters (200), first splint strands (300) and second splint strands (400); c) contacting the plurality of single-stranded nucleic acid library molecules with the plurality of double-stranded splint adaptors under conditions sufficient for the termini of the first splint strands to hybridize to the termini of the library molecules, thereby generating a plurality of library-splint complexes.
27. 25. A method for generating a library-splint complex according to claim 23 or 24, comprising: a. providing a plurality of single-stranded nucleic acid library molecules, a plurality of first splint strands, and a plurality of second splint strands; b) contacting the plurality of single-stranded nucleic acid library molecules with the plurality of first splint strands and second splint strands under conditions sufficient for the second splint strand to hybridize to the first splint strand and sufficient for the terminus of the first splint strand to hybridize to the terminus of the library molecule, thereby generating a plurality of library-splint complexes.
28. 1. A method of sequencing a plurality of concatemeric template molecules, comprising: a. providing a plurality of library-sprint complexes according to any one of claims 1 to 24; b. performing rolling circle amplification on the plurality of library-splint complexes to generate a plurality of concatemeric template molecules; c. sequencing the plurality of concatemeric template molecules.
29. A kit comprising a plurality of double-stranded splint adaptors according to any one of claims 1 to 22.