Double chain cleat adapter with universal long cleat chain and method of use thereof

By hybridizing with single-stranded nucleic acid library molecules to form covalent closed loop molecules, the problem of insufficient sequencing throughput in the prior art is solved, and efficient library preparation and sequencing is achieved to meet the needs of high sample throughput.

CN120380144APending Publication Date: 2025-07-25ELEMENT BIOSCIENCES INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380077893.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-12
Filing Date
2023-09-12
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing second-generation sequencing technology is difficult to effectively add target sequences and unique index sequences during library preparation, resulting in insufficient sequencing throughput and unable to meet the needs of high sample throughput.

Method used

Double-stranded splint adaptors are used to hybridize with single-stranded nucleic acid library molecules to form a library-splint complex, and covalently closed circular molecules are formed through enzymatic ligation, and rolling ring amplification and sequencing are performed.

Benefits of technology

The sequencing throughput is improved, and a large number of libraries can be efficiently collected on the second-generation sequencing platform for sequencing, meeting the needs of high sample throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380144A_ABST
    Figure CN120380144A_ABST
Patent Text Reader

Abstract

The present disclosure provides compositions comprising nucleic acid double stranded splint adapters, including kits, as well as methods of employing the double stranded splint adapters. The double-stranded splint adaptor (200) can be used in a one-pot multienzyme reaction to introduce one or more new adaptor sequences into a library molecule. The double-stranded splint adaptor (200) comprises a first splint chain (long splint chain (300)) and a second splint chain (short splint chain (400)), where the first splint chain and the second splint chain hybridize together to form the double-stranded splint adaptor (200) having a double-stranded region and two flanked single-stranded regions. The second splint chain (400) carries the new adapter sequence to be introduced, such as, for example, a universal binding sequence, an index sequence and / or a random sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 405,733, filed September 12, 2022, the contents of which are incorporated herein by reference in their entirety.

[0003] References to electronic sequence listings

[0004] The contents of the Electronic Sequence Listing (ELEM-015_001WO_SeqList_ST26.xml; size: 88,981 bytes; and creation date: September 8, 2023) are incorporated herein by reference in their entirety. Technical Field

[0005] The present disclosure relates to methods for DNA sequencing and library preparation, including compositions comprising nucleic acid double-stranded splint adapters, and methods of using the double-stranded splint adapters. The double-stranded splint adapters can hybridize to portions of library molecules to form a library-splint complex with gaps, wherein the gaps can be ligated to form covalently closed circular molecules that can undergo downstream amplification and sequencing workflows. Background Art

[0006] The present disclosure relates to a library of covalently closed circular molecules prepared using double-stranded splint adapters, and a method for sequencing the prepared library using the compositions and methods described herein. Improvements in second-generation sequencing technology have greatly improved sequencing speed and data output, resulting in high sample throughput for current sequencing platforms. Effectively preparing closed circular library molecules with target sequences is very important for downstream amplification and sequencing workflows. Another aspect of improving sequencing throughput is to add unique index sequences to DNA fragments during library preparation, which allows a large number of libraries to be assembled and sequenced simultaneously during each sequencing run. Therefore, there is a need for an alternative method for generating and sequencing circular library molecules containing target sequences and unique index sequences that is compatible with downstream second-generation sequencing technology. Compositions, methods and kits that meet this need are provided herein. Summary of the Invention

[0007] The present disclosure provides a library-splint complex (500) comprising: (i) a single-stranded nucleic acid library molecule (100) comprising a sequence of interest (110) flanked on one side by at least a first left universal adapter sequence (120) and on the other side by at least a first right universal adapter sequence (130); and (ii) a double-stranded splint adapter (200) comprising a first splint strand (300) and a second splint strand (400), wherein the double-stranded splint adapter (200) comprises a double-stranded region and two single-stranded regions, the two single-stranded regions being Each is located on both sides of the double-stranded region, wherein the first splint chain comprises a first region (320), an internal region (310) and a second region (330); wherein the internal region (310) of the first splint chain hybridizes with the second splint chain (400), wherein the first region (320) of the first splint chain hybridizes with at least the first left universal adapter sequence (120) of the library molecule, and wherein the second region (330) of the first splint chain hybridizes with at least the first right universal sequence (130) of the library molecule, thereby circularizing the library molecule to produce a library-splint complex (500).

[0008] In some embodiments of the library-splint complex (500) of the present disclosure, the nucleic acid library molecule (100) further comprises: a second left universal adapter sequence (140). In some embodiments, the second left universal adapter sequence (140) is located between at least the first left universal adapter sequence (120) and the target sequence (110). In some embodiments, the nucleic acid library molecule (100) further comprises: a second right universal adapter sequence (150). In some embodiments, the second right universal adapter sequence (150) is located between the target sequence (110) and at least the first right universal adapter sequence (130). In some embodiments, the nucleic acid library molecule (100) further comprises: a first left index sequence (160). In some embodiments, the first left index sequence (160) is located between at least the first left universal adapter sequence (120) and the target sequence (110). In some embodiments, the nucleic acid library molecule (100) further comprises: a first right index sequence (170). In some embodiments, the first right index sequence (170) is located between the second right universal adapter sequence (150) and at least the first right universal adapter sequence (130). In some embodiments, the nucleic acid library molecule (100) further comprises: a first left unique identification sequence (180). In some embodiments, the first left unique identification sequence (180) is located between at least the first left universal adapter sequence (120) and the first left index sequence (160). In some embodiments, the nucleic acid library molecule (100) further comprises: a first right unique identification sequence (190). In some embodiments, the first right unique identification sequence (190) is located between the first right index sequence (170) and at least the first right universal adapter sequence (130).

[0009] In some embodiments of the library-splint complex (500) of the present disclosure, the nucleic acid library molecule (100) further comprises any one or any combination of two or more of the following: (i) a second left universal adapter sequence (140); (ii) a second right universal adapter sequence (150); (iii) a first left index sequence (160); (iv) a first right index sequence (170); (v) a first left unique identification sequence (180); and / or (vi) a first right unique identification sequence (190).

[0010] In some embodiments of the library-splint complex (500) of the present disclosure, the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) include: (i) a universal binding sequence for a forward sequencing primer; (ii) a universal binding sequence for a reverse sequencing primer; (iii) a universal binding sequence for a first surface primer; (iv) a universal binding sequence for a second surface primer; (v) a universal binding sequence for a forward amplification primer; (vi) a universal binding sequence for a reverse amplification primer; and / or (vii) a universal binding sequence for a compacted oligonucleotide. In some embodiments, the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) include: (i) a universal binding sequence for a forward sequencing primer; (ii) a universal binding sequence for a reverse sequencing primer; (iii) a universal binding sequence for a first surface primer; (iv) a universal binding sequence for a second surface primer; (v) a universal binding sequence for a forward amplification primer; (vi) a universal binding sequence for a reverse amplification primer; and / or (vii) a universal binding sequence for a compacted oligonucleotide.

[0011] In some embodiments of the library-splint complex (500) of the present disclosure, the second splint chain (400) includes at least two sub-regions: a first sub-region comprising a universal binding sequence for a third surface primer; and a second sub-region comprising a universal binding sequence for a fourth surface primer, wherein the first sub-region and the second sub-region do not hybridize with the first surface primer and the second surface primer or exhibit very little hybridization. In some embodiments, the second splint chain (400) includes an optional third sub-region, wherein the third sub-region comprises a sample index sequence having 5-20 bases and / or a unique identification sequence having 2-10 or more bases. In some embodiments, the unique identification sequence comprises a random sequence.

[0012] In some embodiments of the library-splint complex (500) of the present disclosure, the first splint chain (300) includes an internal region (310) comprising at least two subregions: a fourth subregion comprising a universal binding sequence for a third surface primer, and the fourth subregion hybridizes with the first subregion of the second splint chain (400); and a fifth subregion comprising a universal binding sequence for a fourth surface primer, and the fifth subregion hybridizes with the second subregion of the second splint chain (400); wherein the fourth subregion and the fifth subregion do not hybridize with the first surface primer and the second surface primer or at least exhibit very little hybridization. In some embodiments, the first splint chain (300) includes an internal region (310), the internal region further comprising a sixth subregion comprising a sample index sequence having 5-20 bases and / or a unique identification sequence having 2-10 or more bases, wherein the sixth subregion hybridizes with the third subregion of the second splint chain (400). In some embodiments, the unique identification sequence comprises a random sequence.

[0013] The present disclosure provides a library-splint complex (500) comprising: (a) a single-stranded nucleic acid library molecule (100) comprising components arranged in 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer (120); (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; (iii) a target sequence (110); (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer; and (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer (130); (b) a first splint chain (300) comprising components arranged in 5' to 3' order: a first region (320); an internal region (310); and a second region (330); and (c) a second splint chain (400) comprising a first region (320); an internal region (310); and a second region (330). to 5' sequentially arranged subregions: a first subregion having a universal binding sequence for a third surface primer; and a second subregion having a universal binding sequence for a fourth surface primer; wherein the first splint chain (300) hybridizes with a portion of the library molecule (100), thereby looping the library molecule to produce a library-splint complex (500), such that the first region (320) of the first splint chain hybridizes with the binding sequence (120) for the first surface primer, and the third region (330) of the first splint chain hybridizes with the binding sequence (130) for the second surface primer, wherein the second splint chain (400) hybridizes with the internal region (310) of the first splint chain (300), wherein the library-splint complex (500) comprises a first gap between the 5' end of the library molecule and the 3' end of the second splint chain, and wherein the library-splint complex (500) comprises a second gap between the 5' end of the second splint chain and the 3' end of the library molecule.

[0014] In some embodiments of the disclosed library-splint complex (500), the first gap and the second gap are enzymatically ligatable.

[0015] The present disclosure provides a plurality of library-splint complexes comprising the library-splint complex (500) of the present disclosure, wherein the sequence of interest (110) of each library-splint complex in the plurality of library-splint complexes comprises the same sequence of interest or a different sequence of interest.

[0016] The present disclosure provides a method for producing the library-splint complex of the present disclosure, which comprises: (a) providing a plurality of single-stranded nucleic acid library molecules (100); (b) providing a plurality of double-stranded splint adaptors (200), a first splint chain (300) and a second splint chain (400); and (c) contacting the plurality of single-stranded nucleic acid library molecules with the plurality of double-stranded splint adaptors under conditions sufficient to hybridize the ends of the first splint chain with the ends of the library molecules, thereby producing a plurality of library-splint complexes.

[0017] The present disclosure provides a method for producing the library-splint complex of the present disclosure, which comprises: (a) providing a plurality of single-stranded nucleic acid library molecules, a plurality of first splint chains, and a plurality of second splint chains; and (b) contacting the plurality of single-stranded nucleic acid library molecules with the plurality of first splint chains and the plurality of second splint chains under conditions sufficient to hybridize the second splint chains with the first splint chains and the ends of the first splint chains with the ends of the library molecules, thereby producing a plurality of library-splint complexes.

[0018] The present invention provides a method for sequencing a plurality of concatemer template molecules, comprising: (a) providing a plurality of library-splint complexes disclosed herein; (b) performing rolling circle amplification on the plurality of library-splint complexes to generate a plurality of concatemer template molecules; and (c) sequencing the plurality of concatemer template molecules.

[0019] The present disclosure provides kits comprising a plurality of double-stranded splint adaptors of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The novel features of the present invention are set forth with particularity in the appended claims. The features and advantages of the present invention will be better understood by reference to the following detailed description which sets forth exemplary embodiments in which the principles of the invention are utilized, and in the accompanying drawings which illustrate:

[0021] Figure 1Schematic diagram showing hybridization of an exemplary linear single-stranded library molecule (100) with a double-stranded splint molecule (200, also referred to as a "ds-splint adaptor"), thereby looping the library molecule to form a library-splint complex (500) with two gaps. The library molecule (100) comprises a sequence of interest (insert (110)) flanked on one side by a first left universal adaptor sequence (120) and on the other side by a first right universal adaptor sequence (130). The double-stranded splint molecule comprises a first splint chain (long chain (300)) that hybridizes to a second splint chain (short chain (400)). The first splint chain comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint chain hybridizes to the second splint chain (400). "P" represents a 5' terminal phosphate group.

[0022] Figure 2 is with Figure 1 The same schematic diagram as shown, which relates in more detail to an embodiment of the inner region (310) of the first splint chain (300) and the second splint chain (400). The second splint chain (400) may include two sub-regions, wherein the first sub-region comprises a universal binding sequence for the third surface primer and the second sub-region comprises a universal binding sequence for the fourth surface primer. The inner region (310) of the first splint chain (300) may include two sub-regions, wherein the fourth sub-region hybridizes with the first sub-region of the second splint chain (400) and the fifth sub-region hybridizes with the second sub-region of the second splint chain (400).

[0023] Figure 3 is with Figure 1 The same schematic diagram as shown in , which relates in more detail to an embodiment of the inner region (310) of the first splint chain (300) and the second splint chain (400). The second splint chain (400) may include three sub-regions, wherein the first sub-region comprises a universal binding sequence for a third surface primer, the second sub-region comprises a universal binding sequence for a fourth surface primer, and the third sub-region comprises a sample index sequence having 5-20 bases and / or a unique identification sequence having 2-10 or more bases (e.g., NN). The inner region (310) of the first splint chain (300) may include three sub-regions, wherein the fourth sub-region hybridizes with the first sub-region of the second splint chain (400), the fifth sub-region hybridizes with the second sub-region of the second splint chain (400), and the sixth sub-region hybridizes with the third sub-region of the second splint chain (400).

[0024] Figure 4Schematic diagram showing an exemplary linear single-stranded library molecule (100) hybridized with a double-stranded splint molecule (200), thereby looping the library molecule to form a library-splint complex (500) with two gaps. The library molecule (100) includes a target sequence (insert, 110), which is flanked on one side by a first left universal adapter sequence (120) and a second left universal adapter sequence (140), and on the other side by a second right universal adapter sequence (150) and a first right universal adapter sequence (130). The double-stranded splint molecule includes a first splint chain (long chain (300)), which is hybridized with a second splint chain (short chain (400)). The first splint chain includes a first region (320) hybridized with a sequence on one end of the linear single-stranded library molecule and a second region (330) hybridized with a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint chain hybridizes with the second splint chain (400).

[0025] Figure 5 Schematic diagram showing an exemplary linear single-stranded library molecule (100) hybridized with a double-stranded splint molecule (200), thereby looping the library molecule to form a library-splint complex (500) with two gaps. The library molecule (100) comprises: a first left universal adapter sequence (120); a first left unique identification sequence (180); a first left index sequence (160); a second left universal adapter sequence (140); a target sequence (110); a second right universal adapter sequence (150); a first right index sequence (170); and a first right universal adapter sequence (130). The double-stranded splint molecule comprises a first splint chain (long chain (300)) that hybridizes to a second splint chain (short chain (400)). The first splint chain comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The interior region (310) of the first splint strand is hybridized with the second splint strand (400).

[0026] Figure 6Schematic diagram showing an exemplary linear single-stranded library molecule (100) hybridized with a double-stranded splint molecule (200), thereby looping the library molecule to form a library-splint complex (500) with two gaps. The library molecule (100) comprises: a first left universal adapter sequence (120); a first left index sequence (160); a second left universal adapter sequence (140); a target sequence (insert, 110); a second right universal adapter sequence (150); a first right index sequence (170); a first right unique identification sequence (190); and a first right universal adapter sequence (130). The double-stranded splint molecule comprises a first splint chain (long chain (300)) that hybridizes to a second splint chain (short chain (400)). The first splint chain comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The interior region (310) of the first splint strand is hybridized with the second splint strand (400).

[0027] Figure 7ASchematic diagram showing hybridization of an exemplary linear single-stranded library molecule (100) with a double-stranded splint molecule (200), thereby looping the library molecule to form a library-splint complex (500) with two gaps. The library molecule (100) comprises: a first left universal adapter sequence (120); a first left index sequence (160); a second left universal adapter sequence (140); a target sequence (also referred to as an insert, 110); a second right universal adapter sequence (150); a first right index sequence (170); and a first right universal adapter sequence (130). The double-stranded splint molecule comprises a first splint chain (long chain (300)) that hybridizes to a second splint chain (short chain (400)). The first splint chain comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The inner region (310) of the first splint chain is hybridized with the second splint chain (400). The second splint chain (400) includes two sub-regions, wherein the first sub-region includes a universal binding sequence for a fourth surface primer (e.g., a surface pinning primer), and the second sub-region includes a universal binding sequence for a third surface primer (e.g., a surface capture primer). A random sequence (e.g., NNN) is inserted into the first sub-region, or the random sequence replaces the region in the first sub-region. The random sequence may comprise 3-20 bases. In some embodiments, the random sequence further comprises a sample index sequence. The random sequence may be sequenced, and the sequence information may be used for polymerase cloning mapping and / or template alignment. The random sequence in the second splint chain (400) is represented by a horizontal bar pattern region. The inner region (310) of the first splint chain (300) includes two sub-regions, wherein the fourth sub-region hybridizes with the first sub-region of the second splint chain (400), and the fifth sub-region hybridizes with the second sub-region of the second splint chain (400). The predetermined sequence is inserted into the fourth subregion, or the predetermined sequence replaces a region in the fourth subregion. The predetermined sequence in the first splint strand (300) comprises 3-20 bases and has a sequence that may or may not be complementary to the random sequence in the first subregion of the second splint strand (400). The predetermined sequence is represented by a white area without a pattern.

[0028] Figure 7BSchematic diagram showing hybridization of an exemplary linear single-stranded library molecule (100) with a double-stranded splint molecule (200), thereby looping the library molecule to form a library-splint complex (500) with two gaps. The library molecule (100) comprises: a first left universal adapter sequence (120); a first left index sequence (160); a second left universal adapter sequence (140); a target sequence (insert, 110); a second right universal adapter sequence (150); a first right index sequence (170); and a first right universal adapter sequence (130). The double-stranded splint molecule comprises a first splint chain (long chain (300)) that hybridizes to a second splint chain (short chain (400)). The first splint chain comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The inner region (310) of the first splint chain is hybridized with the second splint chain (400). The second splint chain (400) may include two sub-regions, wherein the first sub-region comprises a universal binding sequence for a fourth surface primer (e.g., a surface pinning primer), and the second sub-region comprises a universal binding sequence for a third surface primer (e.g., a surface capture primer). A random sequence (e.g., NNN) is attached to the 3' end of the first sub-region. The random sequence may comprise 3-20 bases. In some embodiments, the random sequence further comprises a sample index sequence. The random sequence may be sequenced, and the sequence information may be used for polymerase cloning mapping and / or template alignment. The random sequence in the second splint chain (400) is represented by a horizontal bar pattern area. The inner region (310) of the first splint chain (300) may include two sub-regions, wherein the fourth sub-region hybridizes with the first sub-region of the second splint chain (400), and the fifth sub-region hybridizes with the second sub-region of the second splint chain (400). A predetermined sequence may be attached to the 5' end of the fourth sub-region. The predetermined sequence in the first splint strand (300) may contain 3-20 bases and may have a sequence that is complementary or non-complementary to the random sequence in the first subregion of the second splint strand (400). The predetermined sequence is represented by a white area without a pattern.

[0029] Figure 8Schematic diagram showing hybridization of an exemplary linear single-stranded library molecule (100) to a double-stranded splint molecule (200), thereby circularizing the library molecule to form a library-splint complex (500) having two gaps. The library molecule (100) comprises: a first appended left universal adapter sequence (121); a first left universal adapter sequence (120); a first left connecting adapter sequence (125); a first left index sequence (160); a second left connecting adapter sequence (165); a second left universal adapter sequence (140); a third left connecting adapter sequence (145); a target sequence (insert, 110); a third right connecting adapter sequence (155); a second right universal adapter sequence (150); a second right connecting adapter sequence (175); a first right index sequence (170); a first right connecting adapter sequence (135); a first right universal adapter sequence (130); and a first appended right universal adapter sequence (131). The double-stranded splint molecule comprises a first splint chain (long chain (300)) that is hybridized to a second splint chain (short chain (400)). The first splint chain comprises a first region (320) that hybridizes to a sequence on one end of the linear single-stranded library molecule and a second region (330) that hybridizes to a sequence on the other end of the linear single-stranded library molecule. The internal region (310) of the first splint chain is hybridized to the second splint chain (400). For simplicity, the library-splint complex (500) does not show any connecting adapter sequence or additional universal adapter sequence. The skilled person will recognize that the linear library molecule (100) may include any one or any combination of two or more connecting adapters, with or without one or both of the additional universal adapter sequences present in the library molecule (100). The skilled person will recognize that the library-splint complex (500) may include any one or any combination of two or more connecting adapters, with or without one or both of the additional universal adapter sequences present in the library molecule (100).

[0030] Figure 9Three schematic diagrams of exemplary covalently closed circular library molecules (600) are shown, each of which is hybridized to a first splint strand (300). The top schematic diagram shows a covalently closed circular library molecule (600) having a target sequence (insert, 110), a first right universal adapter sequence (130), a second splint strand sequence (400), and a first left universal adapter sequence (120). The middle schematic diagram shows a covalently closed circular library molecule (600) having a target sequence (110), a second right universal adapter sequence (150), a first right universal adapter sequence (130), a second splint strand sequence (400), a first left universal adapter sequence (120), and a second left universal adapter sequence (140). The bottom schematic shows a covalently closed circular library molecule (600) having a target sequence (110), a second right universal adapter sequence (150), a first right index sequence (170), a first right universal adapter sequence (130), a second splint strand sequence (400), a first left universal adapter sequence (120), a first left unique identification sequence (180), a first left index sequence (160), and a second left universal adapter sequence (140).

[0031] Figure 10 Schematic diagram showing an exemplary library-splint complex (500) undergoing a ligation reaction to close the gap to form a covalently closed circular library molecule (600), which is hybridized to the first splint strand (300), where the first splint strand (300) is used as an amplification primer to perform a rolling circle amplification reaction. The dotted line represents the nascent extension product.

[0032] Figure 11A The nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (300) and a second splint strand (400) is shown. The exemplary first splint strand comprises a first region (320; SEQ ID NO:4), a second region (330; SEQ ID NO:5), and an inner region (310) having a fourth subregion (SEQ ID NO:6) and a fifth subregion (SEQ ID NO:7). The exemplary second splint strand (400) comprises a first subregion (SEQ ID NO:1) and a second subregion (SEQ ID NO:2). In Figure 11A, the second splint strand (top strand, 400) has a sequence of SEQ ID NO:202, and the first splint strand (bottom strand, 300) has a sequence of SEQ ID NO:199.

[0033] Figure 11B The nucleotide sequence of an exemplary first splint strand (300) is shown, each of which has a truncated sequence at the 5' end of the first region (320). The truncated sequence of the first region (320) is different from SEQ ID NO: 4 (see Figure 11A ).exist Figure 11B In the exemplary truncated first splint chain shown, the fourth subregion comprises the sequence of SEQ ID NO: 6, the fifth subregion comprises the sequence of SEQ ID NO: 7, and the second region (330) comprises the sequence of SEQ ID NO: 5. The truncated first chain (300) can be Figure 11A The second splint strand (400) is shown hybridized, wherein the second splint strand comprises a first sub-region (SEQ ID NO: 1) and a second sub-region (SEQ ID NO: 2). Figure 11B The full-length sequences in the sequence from top to bottom are: SEQ ID NO: 217, SEQ ID NO: 218, SEQ ID NO: 219, SEQ ID NO: 220, SEQ ID NO: 221.

[0034] Figure 11C The nucleotide sequence of an exemplary first splint strand (300) is shown, each of which has a mismatch sequence within the first region (320). The mismatch sequence is represented by lowercase letters and is underlined. The mismatch sequence is different from SEQ ID NO: 4 (see Figure 11A ). The first region (320) can hybridize with the first left universal adapter sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatch sequence in the first region (320). Figure 11C In the exemplary mismatched first splint strand shown, the fourth subregion comprises the sequence of SEQ ID NO: 6, the fifth subregion comprises the sequence of SEQ ID NO: 7, and the second region (330) comprises the sequence of SEQ ID NO: 5. The mismatched first strand (300) can be Figure 11A The second splint strand (400) is shown hybridized, wherein the second splint strand comprises a first sub-region (SEQ ID NO: 1) and a second sub-region (SEQ ID NO: 2). Figure 11C The full-length sequences from top to bottom are: SEQ ID NO: 222, SEQ ID NO: 223, SEQ ID NO: 224, SEQ ID NO: 225, SEQ ID NO: 226, SEQ ID NO: 227.

[0035] Figure 11DThe nucleotide sequence of an exemplary first splint chain (300) having an abasic site or uracil is shown. The first splint chain shown at the top contains abasic sites in the fourth subregion and the fifth subregion. The abasic sites are represented by solid black bars. The first region (320) of the top first splint chain can hybridize with the first left universal adapter sequence (120) of the library molecule (100). The second region (330) of the top first splint chain can hybridize with the first right universal adapter sequence (130) of the library molecule (100). The first splint chain shown at the bottom contains at least one uracil in the first region (320), the second region (330) and the internal region (310). The uracil is underlined. The first region (320) of the bottom first splint chain can hybridize with the first left universal adapter sequence (120) of the library molecule (100). The second region (330) of the bottom first splint strand can hybridize to the first right universal adapter sequence (130) of the library molecule (100). Top strand: SEQ ID NO: 228-basic site-SEQ ID NO: 229-basic site-SEQ ID NO: 230; Bottom strand: SEQ ID NO: 231.

[0036] Figure 12A The nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint chain (300, lower chain) and a second splint chain (400, upper chain) is shown. The exemplary first splint chain comprises a first region (320), a second region (330) and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint chain (400) comprises a first subregion and a second subregion. The internal region (310) of the first splint chain (300) comprises two subregions, wherein the fourth subregion hybridizes with the first subregion of the second splint chain (400) and the fifth subregion hybridizes with the second subregion of the second splint chain (400). A 3-mer random sequence (e.g., NNN) is inserted into the sequence of the first subregion of the second splint chain (400). A 3-base predetermined sequence (e.g., 5'-gcg-3') is inserted into the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule having a bubble at the position of the 3-mer random sequence (e.g., NNN) in the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 232; bottom strand: SEQ ID NO: 233.

[0037] Figure 12BThe nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint chain (300, lower chain) and a second splint chain (400, upper chain) is shown. The exemplary first splint chain comprises a first region (320), a second region (330) and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint chain (400) comprises a first subregion and a second subregion. The internal region (310) of the first splint chain (300) comprises two subregions, wherein the fourth subregion hybridizes with the first subregion of the second splint chain (400) and the fifth subregion hybridizes with the second subregion of the second splint chain (400). A 4-mer random sequence (e.g., NNNN) is inserted into the sequence of the first subregion of the second splint chain (400). A 4-base predetermined sequence (e.g., 5'-gtcg-3') is inserted into the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule having a bubble at the position of the 4-mer random sequence (e.g., NNNN) in the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 234; bottom strand: SEQ ID NO: 235.

[0038] Figure 13A The nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint chain (300, lower chain) and a second splint chain (400, upper chain) is shown. The exemplary first splint chain comprises a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint chain (400) comprises a first subregion and a second subregion. The internal region (310) of the first splint chain (300) comprises two subregions, wherein the fourth subregion hybridizes with the first subregion of the second splint chain (400) and the fifth subregion hybridizes with the second subregion of the second splint chain (400). A 3-mer random sequence (e.g., NNN) replaces a portion of the sequence of the first subregion of the second splint chain (400). A 3-base predetermined sequence (e.g., 5'-tgc-3') replaces a portion of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule having a bubble at the position of the 3-mer random sequence (e.g., NNN) in the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 236; bottom strand: SEQ ID NO: 237.

[0039] Figure 13BThe nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint strand (lower strand, 300) and a second splint strand (upper strand, 400) is shown. The exemplary first splint strand comprises a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint strand (400) comprises a first subregion and a second subregion. The internal region (310) of the first splint strand (300) comprises two subregions, wherein the fourth subregion hybridizes with the first subregion of the second splint strand (400) and the fifth subregion hybridizes with the second subregion of the second splint strand (400). A 4-mer random sequence (e.g., NNNN) replaces a portion of the sequence of the first subregion of the second splint strand (400). A 4-base predetermined sequence (e.g., 5'-tgc-3') replaces a portion of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule having a bubble at the position of the 4-mer random sequence (e.g., NNN) in the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 238; bottom strand: SEQ ID NO: 239.

[0040] Figure 14A The nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint chain (lower chain, 300) and a second splint chain (upper chain, 400) is shown. The exemplary first splint chain comprises a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint chain (400) comprises a first subregion and a second subregion. The internal region (310) of the first splint chain (300) comprises two subregions, wherein the fourth subregion hybridizes with the first subregion of the second splint chain (400) and the fifth subregion hybridizes with the second subregion of the second splint chain (400). A 3-mer random sequence (e.g., NNN) and an index sequence (e.g., 5'-cactattcc-3') are attached to the 3' end of the first subregion of the second splint chain (400). A predetermined sequence of equal length to the 3-mer random sequence and the index sequence (e.g., 5'-ggaatagtgaca-3') is appended to the 5' end of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule having bubbles or mismatched ends at the locations of the 3-mer random sequence (e.g., NNN) and the index sequence in the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 240; bottom strand: SEQ ID NO: 241.

[0041] Figure 14BThe nucleotide sequence of an exemplary double-stranded splint molecule (200) having a first splint chain (lower chain, 300) and a second splint chain (upper chain, 400) is shown. The exemplary first splint chain comprises a first region (320), a second region (330), and an internal region (310) having a fourth subregion and a fifth subregion. The exemplary second splint chain (400) comprises a first subregion and a second subregion. The internal region (310) of the first splint chain (300) comprises two subregions, wherein the fourth subregion hybridizes with the first subregion of the second splint chain (400) and the fifth subregion hybridizes with the second subregion of the second splint chain (400). A 4-mer random sequence (e.g., NNNN) and an index sequence (e.g., 5'-cactattcc-3') are attached to the 3' end of the first subregion of the second splint chain (400). A predetermined sequence of equal length to the 4-mer random sequence and the index sequence (e.g., 5'-ggaatagtgacag-3') is appended to the 5' end of the sequence of the fourth subregion. The second splint strand (400) can hybridize with the first splint strand (300) to form a double-stranded molecule having bubbles or mismatched ends at the locations of the 4-mer random sequence (e.g., NNNN) and the index sequence in the first subregion of the second splint strand (400). Top strand: SEQ ID NO: 242; bottom strand: SEQ ID NO: 243.

[0042] Figure 15A is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6 Sequencing quality scores for C base recognition of first-strand concatemer template molecules generated by workflows that circularize linear library molecules (without unique recognition sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0043] Figure 15B is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6 Sequencing quality scores for A base recognition of first-strand concatemer template molecules generated by workflows for circularizing linear library molecules (without unique recognition sequences (180) and (190)), heating or not heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0044] Figure 15C is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6 Sequencing quality scores for G base recognition of first-strand concatemer template molecules generated by workflows that circularize linear library molecules, but do not have unique recognition sequences (180) and (190), with or without heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0045] Figure 15D is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6 Sequencing quality scores for T base recognition of first-strand concatemer template molecules generated by workflows that circularize linear library molecules (without unique recognition sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0046] Figure 16A is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6 Sequencing quality scores for C base recognition of second-strand concatemer template molecules generated by workflows that circularize linear library molecules, but do not have unique recognition sequences (180) and (190), with or without heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0047] Figure 16B is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6Sequencing quality scores for A base recognition of second-strand concatemer template molecules generated by workflows that circularize linear library molecules (without unique recognition sequences (180) and (190)), with or without heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0048] Figure 16C is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6 Sequencing quality scores for G base recognition of second-strand concatemer template molecules generated by workflows that circularize linear library molecules, but do not have unique recognition sequences (180) and (190), with or without heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0049] Figure 16D is a diagram illustrating a method comprising using a double-stranded splint adapter (e.g., see Figure 5 or Figure 6 Sequencing quality scores for T base recognition of second-strand concatemer template molecules generated by workflows that circularize linear library molecules, but do not have unique recognition sequences (180) and (190), with or without heating, and performing on-support rolling circle amplification to generate concatemer template molecules immobilized to a coated support. Library preparation workflows compared the effects of heat, NaOH, HEPES buffer, or enzyme cocktails that generate and remove abasic sites. The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0050] Figure 17 is a series of 3 figures illustrating the use of double-stranded splint adapters (e.g., see Figure 5 or Figure 6 Sequencing quality scores for A, G, C, and T base recognition of the first-strand concatemer template molecule (read 1) generated by a workflow that circularizes the linear library molecule without unique recognition sequences (180) and (190), accompanied by ligase inactivation and on-support rolling circle amplification to generate concatemer template molecules immobilized to the coated support. Library preparation workflow comparing the effects of a high-heat inactivation control for ligase (left panel), a low-heat inactivation for ligase (center panel), and a NaOH inactivation for ligase (right panel). The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0051] Figure 18 is a series of 3 figures illustrating the use of double-stranded splint adapters (e.g., see Figure 5 or Figure 6 Sequencing quality scores for A, G, C, and T base recognition of the second-strand concatemer template molecule (read 2) generated by a workflow that circularizes the linear library molecule without unique recognition sequences (180) and (190), accompanied by ligase inactivation and on-support rolling circle amplification to generate concatemer template molecules immobilized to the coated support. Library preparation workflow comparing the effects of a high-heat inactivation control for ligase (left panel), a low-heat inactivation for ligase (center panel), and a NaOH inactivation for ligase (right panel). The X-axis is the number of sequencing cycles. The Y-axis is the quality score.

[0052] Figure 19 is a schematic diagram of an exemplary low binding support comprising a glass substrate and layers of alternating hydrophilic coatings covalently or non-covalently adhered to the glass, and which further comprises chemically reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides and circularizing oligonucleotides). In alternative embodiments, the support can be made of any material, such as glass, plastic, or polymeric materials.

[0053] Figure 20 Schematic diagrams of various exemplary configurations of multivalent molecules. Left: Schematic diagram of a multivalent molecule having a starburst or helical configuration. Center: Schematic diagram of a multivalent molecule having a dendrimer configuration. Right: Schematic diagram of various multivalent molecules formed by reacting streptavidin with 4-arm or 8-arm PEG-NHS, biotin, and dNTPs. The nucleotide unit is designated as 'N', biotin as 'B', and streptavidin as 'SA'.

[0054] Figure 21 is a schematic diagram of an exemplary multivalent molecule comprising a universal core attached to multiple nucleotide arms.

[0055] Figure 22 is a schematic diagram of an exemplary multivalent molecule comprising a dendritic core attached to multiple nucleotide arms.

[0056] Figure 23 A schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, wherein the nucleotide arms comprise biotin, a spacer, a linker, and nucleotide units is shown.

[0057] Figure 24 is a schematic diagram of an exemplary nucleotide arm comprising a core attachment portion, a spacer, a linker, and a nucleotide unit.

[0058] Figure 25Shown are the chemical structures of exemplary spacers (top) and various exemplary linkers, including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker (bottom).

[0059] Figure 26 The chemical structures of various exemplary linkers are shown, including Linkers 1-9.

[0060] Figure 27 The chemical structures of various exemplary linkers joined / attached to the nucleotide units are shown.

[0061] Figure 28 The chemical structures of various exemplary linkers joined / attached to the nucleotide units are shown.

[0062] Figure 29 The chemical structures of various exemplary linkers joined / attached to the nucleotide units are shown.

[0063] Figure 30 The chemical structure of an exemplary biotinylated nucleotide arm is shown. In this example, the nucleotide unit is linked to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base.

[0064] Figure 31 is a schematic diagram of a guanine tetrad (eg, a G-tetrad).

[0065] Figure 32 is a schematic diagram of an exemplary intramolecular G-quadruplex structure.

[0066] FIG33 is Table 1 (of 6) which lists the sequences of exemplary first left index sequences (160) and first right index sequences (170).

[0067] Figure 34 Bar graph showing the average percent recovery of covalently closed circular library molecules using input DNA from different species, as determined by qPCR. Lane 1: Haemophilus influenzae (38% GC); Lane 2: Escherichia coli (51% GC); Lane 3: Rhodopseudomonas palustris (65% GC); Lane 4: PhiX; Lane 5: Human; Lane 6: Human exome; Lane 7: Human mRNA. See Examples 1-3.

[0068] Figure 35 This is a bar graph showing the average polymerase clone density obtained by distributing covalently closed circular library molecules onto a support and performing on-support rolling circle amplification. Covalently closed circular library molecules were prepared from input DNA from various species. Lane 1: cell-free DNA; Lane 2: E. coli; Lane 3: human; Lane 4: metagenomic DNA; Lane 5: PhiX. See Example 4.

[0069] Figure 36 1 is a graph showing the nucleotide base diversity of the right index sequence (170) including the 3-mer random sequence (NNN). The graph shows that the nucleotide diversity of the 3-mer random sequence (NNN) recognized by A and T bases is approximately 30%, and the nucleotide diversity of the 3-mer random sequence (NNN) recognized by C and G bases is approximately 20%.

[0070] Figure 37 is a graph showing the nucleotide base diversity of the left index sequence (160) lacking a 3-mer random sequence (NNN). The graph shows that the nucleotide diversity for A and T base recognition is approximately 40%, the nucleotide diversity for C base recognition is approximately 15%, and the nucleotide diversity for G base recognition is approximately 5%. DETAILED DESCRIPTION

[0071] definition

[0072] Throughout this application, various publications, patents and / or patent applications are cited. The disclosures of the publications, patents and / or patent applications are hereby incorporated by reference into this application in their entirety in order to more fully describe the state of the art to which this disclosure pertains.

[0073] The headings provided herein are not limitations of the various aspects of the disclosure, which can be understood by reference to the specification as a whole.

[0074] Unless otherwise defined, the technology and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. Generally speaking, the terms related to the technology of molecular biology as described herein, nucleic acid chemistry, protein chemistry, genetics, microbiology, transgenic cell production and hybridization are well-known in the art and commonly used those. Technology and procedures as described herein are generally performed according to conventional methods well known in the art and as described in the various general and more specific references cited and discussed throughout this specification. For example, referring to Sambrook et al., Molecular Cloning:A Laboratory Manual (3rd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY2000). See also Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992). The nomenclature utilized in conjunction with laboratory procedures as described herein and technology is well-known in the art and commonly used those.

[0075] Unless the context otherwise requires, singular terms shall include pluralities and plural terms shall include the singular. The singular forms "a," "an," and "the," as well as any words used in the singular, include plural referents unless expressly and unequivocally limited to one referent.

[0076] It will be understood that use of alternative terms (eg, "or") is taken to mean one or both or any combination of the alternatives.

[0077] As used herein, the term "and / or" should be taken to mean a specific disclosure of each of the specified features or components with or without the other. For example, the term "and / or" as used in phrases such as "A and / or B" herein is intended to include: "A and B"; "A or B"; "A" (only A); and "B" (only B). Similarly, the term "and / or" as used in phrases such as "A, B, and / or C" herein is intended to cover each of the following aspects: "A, B, and C"; "A, B, or C"; "A or C"; "A or B"; "B or C"; "A and B"; "B and C"; "A and C"; "A" (only A); "B" (only B); and "C" (only C).

[0078] As used herein and in the appended claims, the terms "comprises," "including," "having," and "containing," and grammatical variations thereof, as used herein, are intended to be non-limiting, such that one or more items in a list are not exclusive of other items that can be substituted or added to the listed items. It should be understood that whenever aspects are described herein with the language "comprising," similar aspects described with "consisting of" and / or "consisting essentially of" are also otherwise provided.

[0079] As used herein, the terms "about" and "approximately" refer to values ​​or compositions within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, "about" or "approximately" may mean within one or more standard deviations. Alternatively, "about" or "approximately" may mean a range of up to 10% (i.e., ±10%) or greater, depending on the limitations of the measurement system. For example, about 5 mg may include any amount between 4.5 mg and 5.5 mg. In addition, particularly for biological systems or processes, these terms may mean a value of up to an order of magnitude or up to 5 times. When a specific value or composition is provided in the present disclosure, unless otherwise stated, the meaning of "about" or "approximately" should be assumed to be within an acceptable error range for that specific value or composition. In addition, where a range and / or subrange of values ​​is provided, the range and / or subrange may include the endpoints of the range and / or subrange.

[0080] As used herein, the terms "peptide," "polypeptide," and "protein," and other related terms, are used interchangeably and refer to polymers of amino acids and are not limited to any particular length. Polypeptides may include natural and non-natural amino acids. Polypeptides include recombinant or chemically synthesized forms. Polypeptides also include precursor molecules that have not undergone post-translational modifications, such as proteolytic cleavage, cleavage due to ribosome skipping, hydroxylation, methylation, lipidation, acetylation, sumoylation, ubiquitination, glycosylation, phosphorylation, and / or disulfide bond formation. These terms encompass natural and artificial proteins, protein fragments, and polypeptide analogs (such as mutants, variants, chimeric proteins, and fusion proteins) of protein sequences, as well as proteins that have been post-translationally or otherwise covalently or non-covalently modified.

[0081] The term "cell biological sample" refers to a section of a single cell, a plurality of cells, a tissue, an organ, an organism or any of these cell biological samples. A cell biological sample can be extracted from an organism (e.g., a biopsy) or obtained from a cell culture grown in a liquid or a culture dish. A cell biological sample comprises a fresh, frozen, freshly frozen or archived sample (e.g., formalin-fixed paraffin-embedded; FFPE). A cell biological sample can be embedded in wax, a resin, an epoxy resin or agar. A cell biological sample can be fixed in, for example, any one of acetone, ethanol, methanol, formaldehyde, paraformaldehyde-Triton or glutaraldehyde or any combination of two or more thereof. A cell biological sample can be sectioned or unsectioned. A cell biological sample can be stained, decolorized or unstained.

[0082] Nucleic acids of interest, sometimes referred to herein as sequences of interest, can be extracted from cells or cellular biological samples using a variety of techniques known to those skilled in the art. For example, a typical DNA extraction procedure involves: (i) collecting a cell sample or tissue sample from which DNA is to be extracted, (ii) disrupting the cell membrane (i.e., cell lysis) to release DNA and other cytoplasmic components, (iii) treating the lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA, and (iv) purifying the DNA from the supernatant to remove detergents, proteins, salts, or other reagents used during cell membrane lysis. A variety of suitable commercial nucleic acid extraction and purification kits are consistent with the disclosure herein. Examples include, but are not limited to, the QIAamp kit (for isolating genomic DNA from human samples) and the DNAeasy kit (for isolating genomic DNA from animal or plant samples) from Qiagen (Germantown, MD), or the PERIK™ ELISA Kit (for isolating genomic DNA from animal or plant samples) from Promega (Madison, WI). and ReliaPrep TMThe target nucleic acid can be ribonucleic acid (RNA) or deoxyribonucleic acid (DNA), such as genomic DNA or complementary DNA (cDNA) reverse transcribed from RNA.

[0083] As used herein, the term "polymerase" and its variants include enzymes comprising a domain that binds nucleotides (or nucleosides), wherein the polymerase can form a complex with a template nucleic acid and a complementary nucleotide. The polymerase can have one or more activities, including but not limited to: base analog detection activity, DNA polymerization activity, reverse transcriptase activity, DNA binding, strand displacement activity, and nucleotide binding and recognition. The polymerase can be any enzyme that can catalyze the polymerization of nucleotides (including analogs thereof) into nucleic acid chains. Typically, but not necessarily, such nucleotide polymerization can occur in a template-dependent manner. Typically, the polymerase is included in one or more active sites at which nucleotide binding and / or nucleotide polymerization catalysis can occur. In some embodiments, the polymerase includes other enzymatic activities, such as, for example, 3' to 5' exonuclease activity or 5' to 3' exonuclease activity. In some embodiments, the polymerase has strand displacement activity. Polymerases may include, but are not limited to, naturally occurring polymerases and any of their subunits and truncations, mutant polymerases, variant polymerases, recombinant, fused or otherwise engineered polymerases, chemically modified polymerases, synthetic molecules or assemblies, and any analogs, derivatives or fragments thereof (e.g., catalytically active fragments) that retain the ability to catalyze nucleotide polymerization. The term polymerase includes catalytically inactive polymerases, catalytically active polymerases, reverse transcriptases and other enzymes comprising a nucleotide binding domain. In certain embodiments, the polymerase can be isolated from a cell or produced using recombinant DNA technology or chemical synthesis methods. In certain embodiments, the polymerase can be expressed in a prokaryotic organism, a eukaryotic organism, a virus or a bacteriophage organism. In certain embodiments, the polymerase can be a post-translationally modified protein or fragment thereof. The polymerase can be derived from a prokaryotic organism, a eukaryotic organism, a virus or a bacteriophage. The polymerase includes DNA-guided DNA polymerases and RNA-guided DNA polymerases.

[0084] The term "strand displacement" refers to the ability of a polymerase to locally separate double-stranded nucleic acid chains and synthesize new chains in a template-based manner. A strand-displacing polymerase displaces a complementary strand from a template strand and catalyzes the synthesis of a new strand. Strand-displacing polymerases include mesophilic and thermophilic polymerases. Strand-displacing polymerases include wild-type enzymes and variants (including exonuclease-minus mutants, mutant versions, chimeric enzymes, and truncated enzymes). Examples of strand-displacing polymerases include: phi29 DNA polymerase, large fragment of Bst DNA polymerase, large fragment (exo-) of Bsu DNA polymerase, Bca DNA polymerase (exo-), Klenow fragment of Escherichia coli (E. coli) DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV viral reverse transcriptase, Deep DNA polymerase and KOD DNA polymerase. The phi29 DNA polymerase can be a wild-type phi29 DNA polymerase (e.g., MagniPhi from Expedeon). TM ), or variant EquiPhi29 TM DNA polymerase (e.g., from Thermo Fisher Scientific), or chimeric QualiPhi TM DNA polymerase (eg, from 4basebio).

[0085] As used herein, the terms "nucleic acid," "polynucleotide," and "oligonucleotide," as well as other related terms, are used interchangeably and refer to polymers of nucleotides and are not limited to any specific length. Nucleic acids include recombinant and chemically synthesized forms. Nucleic acids can be isolated. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of DNA or RNA generated using nucleotide analogs (e.g., peptide nucleic acids and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids comprise polymers of nucleotides, wherein the nucleotides comprise natural or non-natural bases and / or sugars. Nucleic acids comprise naturally occurring internucleoside bonds, such as phosphodiester bonds. Nucleic acids comprise non-natural internucleoside bonds, including phosphorothioate, phosphorothiol, or peptide nucleic acid (PNA) bonds. Nucleic acids can comprise one type of polynucleotide, or a mixture of two or more different types of polynucleotides.

[0086] As used herein, the terms "operably linked" and "operably connected" or related terms refer to the juxtaposition of components. The juxtaposed components can be covalently linked together. For example, two nucleic acid components can be enzymatically linked together, wherein the bond joining the two components together comprises a phosphodiester bond. A first nucleic acid component and a second nucleic acid component can be linked together, wherein the first nucleic acid component can confer a function to the second nucleic acid component. For example, the bond between the primer binding sequence and the target sequence forms a nucleic acid library molecule having a portion that can bind to a primer. In another example, a transgene (e.g., a nucleic acid encoding a polypeptide or a target sequence nucleic acid) can be linked to a vector, wherein the bond allows the transgene sequence contained in the vector to be expressed or function. In yet another example, the transgene is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects transgene expression. In an exemplary vector, the vector comprises at least one host cell regulatory sequence comprising a promoter sequence, an enhancer, a transcription and / or translation initiation sequence, a transcription and / or translation termination sequence, a polypeptide secretion signal sequence, etc., which are said to be operably linked. In the foregoing examples, host cell regulatory sequences control the level, timing, and / or location of expression of the transgene.

[0087] The terms "linked," "connected," "attached," "appended," and variations thereof include any type of fusion, binding, attachment, or association between any combination of compounds or molecules that is sufficiently stable to withstand use in a particular procedure. The procedure may include, but is not limited to, nucleotide binding; nucleotide incorporation; deblocking (e.g., removal of a chain terminating moiety); washing; removal; flow; detection; imaging and / or identification. For example, such bonds may include, for example, covalent bonds, ionic bonds, hydrogen bonds, dipole-dipole bonds, hydrophilic bonds, hydrophobic bonds, or affinity bonds, bonds or associations involving van der Waals forces, mechanical bonds, and the like. Such connections may occur within a molecule, such as joining the ends of a single-stranded or double-stranded linear nucleic acid molecule together to form a circular molecule. Alternatively, such bonds may occur between a combination of different molecules, or between a molecule and a non-molecule, including, but not limited to, bonds between a nucleic acid molecule and a solid surface; bonds between a protein and a detectable reporter moiety; bonds between a nucleotide and a detectable reporter moiety; and the like. Examples of some bonds can be found, for example, in Hermanson, G., “Bioconjugate Techniques”, 2nd ed. (2008); Aslam, M., Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences”, London: Macmillan (1998); Aslam, M., Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences”, London: Macmillan (1998).

[0088] The term "primer" and related terms used herein refer to oligonucleotides that can hybridize with DNA and / or RNA polynucleotide templates to form double-stranded molecules. Primers can be single-stranded or have single-stranded and double-stranded portions along their entire length. Primers can include natural nucleotides and / or nucleotide analogs. Primers can be recombinant nucleic acid molecules. Primers can have any length, but generally range from 4 to 50 nucleotides. Typical primers include 5' ends and 3' ends. The 3' end of a primer can include a 3'OH portion that serves as a nucleotide polymerization initiation site in a primer extension reaction catalyzed by a polymerase. Alternatively, the 3' end of a primer can lack a 3'OH portion or can include an end 3' blocking group that inhibits nucleotide polymerization in a polymerase-catalyzed reaction. Any one nucleotide along the length of the primer or more than one nucleotide can be labeled with a detectable reporter portion. Primers can be in solution (e.g., soluble primers) or can be fixed to a support (e.g., capture primers).

[0089] The terms "template nucleic acid," "template polynucleotide," "target nucleic acid," "target polynucleotide," "template strand," and other variations refer to a nucleic acid strand used as a base nucleic acid molecule in any amplification and / or sequencing method described herein. A template nucleic acid can be single-stranded or double-stranded, or a template nucleic acid can have single-stranded or double-stranded portions. A template nucleic acid can be obtained from naturally occurring sources, recombinant forms, or chemically synthesized to include any type of nucleic acid analog. A template nucleic acid can be linear, concatemeric, circular, or in other forms. A template nucleic acid can encode a sequence of interest.

[0090] The term "adapter" and related terms refer to oligonucleotides that can be operably linked (appended) to a target polynucleotide, wherein the adaptor pair conferring function to a co-linked adaptor-target molecule. Adaptors include DNA, RNA, chimeric DNA / RNA or their analogs. The adaptor may include at least one ribonucleoside residue. The adaptor may be single-stranded, double-stranded or have a single-stranded and / or double-stranded portion. The adaptor may be configured as a linear, stem-loop, hairpin or Y-shaped form. The adaptor may be of any length, including 4 to 100 nucleotides or longer. The adaptor may have a blunt end, an overhanging end or a combination of the two. The overhanging end includes a 5' overhang and a 3' overhanging end. The 5' end of a single-stranded adaptor or one chain of a double-stranded adaptor may have a 5' phosphate group or lack a 5' phosphate group. The adaptor may include a 5' tail (e.g., a tailed adaptor) that is not hybridized with the target polynucleotide, or the adaptor may be tailless. At least a portion of the adaptor may include a known and predetermined sequence. The adapter may comprise a sequence that is complementary to at least a portion of a primer, such as an amplification primer, a sequencing primer, or a capture primer (e.g., a soluble or immobilized capture primer). The adapter may comprise a random sequence or a degenerate sequence. The adapter may comprise at least one inosine residue. The adapter may comprise at least one phosphorothioate, phosphorothiol, and / or phosphoramidite bond. The adapter may comprise at least one barcode / index sequence that can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. The adapter may comprise at least one unique identification sequence (e.g., a molecular identifier) ​​that can be used to uniquely identify the nucleic acid molecule to which the adapter is attached. Exemplary, but non-limiting, unique identification sequences comprise 2-12 or more nucleotides of known sequence. For example, the unique identification sequence comprises a known random sequence in which the nucleotides at each position are randomly selected from nucleotides having bases A, G, C, T, or U. The adapter may comprise at least one restriction enzyme recognition sequence, including any one or any combination of two or more selected from the group consisting of Type I, Type II, Type III, Type IV, Type Hs, or Type IIB.

[0091] The term "universal sequence" and related terms refer to a sequence common to two or more polynucleotide molecules in a nucleic acid molecule. For example, an adapter with a universal sequence can be operably joined to multiple polynucleotides so that the population of co-joined molecules carries the same universal adapter sequence. Examples of universal adapter sequences include amplification primer sequences, sequencing primer sequences (e.g., sequences compatible with commercial sequencing platforms), or capture primer sequences (e.g., soluble or immobilized capture primers).

[0092] When used to refer to nucleic acid molecules, the term "hybridize" or "hybridizing" or "hybridization" or other related terms refer to hydrogen bonding between two different nucleic acids to form a duplex nucleic acid. Hybridization also includes hydrogen bonding between two different regions of a single nucleic acid molecule to form a self-hybridizing molecule with a duplex region. Hybridization can include Watson-Crick or Hoogstein bonding to form a duplex double-stranded nucleic acid or a double-stranded region within a nucleic acid molecule. The two different regions of a double-stranded nucleic acid or a single nucleic acid can be completely complementary or partially complementary. The complementary nucleic acid chains do not need to hybridize to each other across their entire length. Complementary base pairing can be standard AT or CG base pairing or can be other forms of base pairing interactions. The duplex nucleic acid can contain mismatched base pairing nucleotides, which can form single-stranded regions ("bubbles") in the duplex.

[0093] When used in reference to nucleic acids, the terms "extend," "extending," "extension," and other variants refer to the incorporation of one or more nucleotides into a nucleic acid molecule. Nucleotide incorporation comprises the polymerization of one or more nucleotides into the terminal 3' OH end of a nucleic acid chain, resulting in extension of the nucleic acid chain. Nucleotide incorporation can be performed using natural nucleotides and / or nucleotide analogs. Typically, but not necessarily, nucleotide incorporation occurs in a template-dependent manner. Any suitable method for extending nucleic acid molecules can be used, including primer extension catalyzed by DNA polymerase or RNA polymerase.

[0094] The term "nucleotide" and related terms refer to molecules comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose) and at least one phosphate group. Regular or non-regular nucleotides are consistent with the use of the term. In certain embodiments, the nucleotide comprises a monophosphate, a diphosphate, or a triphosphate, or a corresponding phosphate analog. The term "nucleoside" refers to a molecule comprising an aromatic base and a sugar. Nucleotides and nucleosides can be unlabeled or labeled with a detectable reporter moiety.

[0095] Nucleotides (and nucleosides) typically contain heterocyclic bases, including substituted or unsubstituted nitrogen-containing parent heteroaromatic rings, which are commonly found in nucleic acids, including naturally occurring, substituted, modified or engineered variants or analogs thereof. The base of a nucleotide (or nucleoside) is capable of forming Watson-Crick and / or Hoogstein hydrogen bonds with an appropriate complementary base. Exemplary bases include, but are not limited to, purines and pyrimidines, such as 2-aminopurine, 2,6-diaminopurine, adenine (A), ethyleneadenine, N 6 -Δ 2 -Isopentenyl adenine (6iA), N 6 -Δ 2 -Isopentenyl-2-methylthioadenine (2ms6iA), N 6 -methyladenine, guanine (G), isoguanine, N 2 -dimethylguanine (dmG), 7-methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine and O 6 -methylguanine; 7-deazapurines such as 7-deazaadenine (7-deaza-A) and 7-deazaguanine (7-deaza-G); pyrimidines such as cytosine (C), 5-propynylcytosine, isocytosine, thymine (T), 4-thiothymine (4sT), 5,6-dihydrothymine, O 4 -methylthymine, uracil (U), 4-thiouracil (4sU) and 5,6-dihydrouracil (dihydrouracil; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; muscimol; inosines; hydroxymethylcytosines; 5-methylcytosines; bases (Y); and methylated, glycosylated and acylated base moieties; etc. Additional exemplary bases can be found in Fasman, 1989, in "Practical Handbook of Biochemistry and Molecular Biology", pp. 385-394, CRC Press, Boca Raton, Fla.

[0096] Nucleotides (and nucleosides) generally contain a sugar moiety, such as a carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100:4319-48), an acyclic moiety (Martinez et al., 1999 Nucleic Acids Research 27:1271-1274; Martinez et al., 1997 Bioorganic & Medicinal Chemistry Letters Vol. 7:3013-3016), and other sugar moieties (Joeng et al., 1993 J. Med. Chem. 36:2627-2638; Kim et al., 1993 J. Med. Chem. 36:30-7; Eschenmosser 1999 Science 284:2118-2124; and U.S. Pat. No. 5,558,991). Sugar moieties include, but are not limited to, ribosyl; 2'-deoxyribosyl; 3'-deoxyribosyl; 2',3'-dideoxyribosyl; 2',3'-didehydrodideoxyribosyl; 2'-alkoxyribosyl; 2'-azidoribosyl; 2'-aminoribosyl; 2'-fluororibosyl; 2'-mercaptocarboxyl; 2'-alkylthioribosyl; 3'-alkoxyribosyl; 3'-azidoribosyl; 3'-aminoribosyl; 3'-fluororibosyl; 3'-mercaptocarboxyl; 3'-alkylthioribosyl carbocyclic; acyclic or other modified sugars.

[0097] Nucleotide can comprise a chain of one, two or three phosphorus atoms, wherein the chain is generally attached to the 5' carbon of the sugar moiety via an ester or phosphoramide bond. In certain embodiments, nucleotide is an analog with a phosphorus chain, wherein the phosphorus atom is linked together with an O, S, NH, methylene or ethylene group in the middle. The phosphorus atom in the chain can include a substituted side group (including O, S or BH3). Alternatively or in addition, the chain can include a phosphate group substituted with an analog, and the analog includes phosphoramide, phosphorothioate, phosphorodithioate and O-methylphosphoramidite groups.

[0098] The term "rolling circle amplification" generally refers to an amplification method that utilizes a circularized nucleic acid template molecule containing a target sequence of interest, an amplification primer binding sequence, and optionally one or more adapter sequences, such as a sequencing primer binding sequence and / or a sample index sequence. The rolling circle amplification reaction can be performed under isothermal amplification conditions and includes the circularized nucleic acid template molecule, an amplification primer, a strand-displacing polymerase, and a plurality of nucleotides to generate concatemers containing tandem repeats of the circular template molecule and any adapter sequences present in the original circularized nucleic acid template molecule. The concatemers can self-collapse to form nucleic acid nanospheres. The shape and size of the nanospheres can be further compressed by including a pair of inverted repeats in the circular template molecule or by performing the rolling circle amplification reaction with one or more compacting oligonucleotides. One advantage of using rolling circle amplification to generate clonal amplicons for sequencing workflows is that duplicate copies of the target sequence in the nanosphere can be sequenced simultaneously to increase signal strength. In some embodiments, the rolling circle amplification reaction can be performed in the presence of multiple compactions having at least four consecutive guanines. The rolling circle amplification reaction produces concatemers containing duplicate copies of the universal binding sequence for the compacting oligonucleotides. At least one compacted oligonucleotide can form a guanine tetrad ( Figure 31 ) and hybridizes to the universal binding sequence of the compacted oligonucleotide, and the resulting concatemer can fold to form an intramolecular G-quadruplex structure ( Figure 32 The concatemer can self-collapse to form a compact nanosphere. The formation of guanine quadruplexes and G-quadruplexes in the nanosphere can increase the stability of the nanosphere, maintaining its compact size and shape, so that it can withstand the repeated flow of reagents used to perform any of the sequencing workflows described herein.

[0099] The terms "amplify," "amplifying," "amplification," and other related terms when used in reference to nucleic acids include generating multiple copies of an original polynucleotide template molecule, wherein the copies comprise a sequence complementary to the template sequence, and / or the copies comprise a sequence identical to the template sequence. In some embodiments, the copies comprise a sequence substantially identical to the template sequence, and / or a sequence substantially identical to a sequence complementary to the template sequence.

[0100] The term "reporter moiety," "reporter moieties," or related terms refers to a compound that produces or causes a detectable signal. A reporter moiety is sometimes referred to as a "label." Any suitable reporter moiety can be used, including luminescence, photoluminescence, electroluminescence, bioluminescence, chemiluminescence, fluorescence, phosphorescence, chromophores, radioisotopes, electrochemistry, mass spectrometry, Raman, haptens, affinity tags, atoms, or enzymes. The reporter moiety generates a detectable signal caused by a chemical or physical change (e.g., heat, light, electricity, pH, salt concentration, enzyme activity, or a neighboring event). A neighboring event includes two reporter moieties that are close to each other, or associate with each other, or are bound to each other. It is well known to those skilled in the art that the reporter moiety is selected so that each absorbs excitation radiation and / or fluoresces at a wavelength that can be distinguished from other reporter moieties, to allow monitoring of the presence of different reporter moieties in the same reaction or in different reactions. Two or more different reporter moieties can be selected that have spectrally distinct emission curves or have minimal overlapping spectral emission curves. The reporter moiety can be linked (eg, operably linked) to a nucleotide, nucleoside, nucleic acid, enzyme (eg, a polymerase or reverse transcriptase), or support (eg, a surface).

[0101] The reporter moiety (or label) may comprise a fluorescent label or fluorophore. Exemplary fluorescent moieties that can be used as fluorescent labels or fluorophores include, but are not limited to, fluorescein and fluorescein derivatives such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynaphthol fluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamido fluorescein, maleimide fluorescein, SAMSA-fluorescein, thiosemicarbazide fluorescein, hydrazine carbonate methylthioacetamido fluorescein, rhodamine and rhodamine derivatives such as TRITC, TMR, lissamine rhodamine, Texas Red, Rhodamine B, Rhodamine 6G, Rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, lissamine rhodamine B sulfonyl chloride, lissamine rhodamine B sulfonyl hydrazide, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550 C3, BODIPY530 / 550 C3-SE, BODIPY 530 / 550 C3 hydrazide, BODIPY 493 / 503 C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FL IA, BODIPY 530 / 551 IA, Br-BODIPY 493 / 503,Cascade and derivatives such as Cascade Blue acetyl triazoide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Yellow and derivatives such as Lucifer Yellow iodoacetamide, Lucifer Yellow CH, cyanines and derivatives such as indolium cyanine dyes, benzindolium cyanine dyes, pyridinium cyanine dyes, thiazolium cyanine dyes, quinolinium cyanine dyes, imidazolium cyanine dyes, Cy 3, Cy 5, lanthanide chelates and derivatives such as BCPDA, TBP, TMT, BHHCT, BCOT, europium chelates, terbium chelates, Alexa Fluor 500, 480, 1467-1496, 1474 ... dye, Dye, Atto TM dye, Red dye, CAL Flour dye, JOE and its derivatives, Oregon Green TMDyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, malachite green, diphenylethylene, DEG dyes, NR dyes, near infrared dyes and other dyes known in the art, such as those described in Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th edition; Lakowicz, Principles of Fluorescence Spectroscopy, 2nd edition, Plenum Press New York (1999) or Hermanson, Bioconjugate Techniques, 2nd edition, or derivatives thereof; or any combination thereof. Cyanine dyes can exist in sulfonated or non-sulfonated form and consist of two indolenine, benzindolium, pyridinium, thiazolium and / or quinolinium groups separated by a polymethine bridge between the two nitrogen atoms.Commercially available cyanine fluorophores include, for example, Cy3 (which may comprise 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium or 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium oxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium-5-sulfonate), Cy5 (which may include 1-(6-((2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl)-3,3-dimethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium-5-sulfonate). )penta-1,3-dien-1-yl)-3,3-dimethyl-3H-indol-1-ium or 1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-sulfonindol-2-ylidene)penta-1,3-dien-1-yl)-3,3-dimethyl-3H-indol-1-ium-5-sulfonate) and Cy7 (which may contain 1-(5-carboxypentyl) )-2-[(1E,3E,5E,7Z)-7-(1-ethyl-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium or 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), where "Cy" stands for 'cyanine' and the first digit identifies the number of carbon atoms between the two indolenine groups. Cy2 is an oxazole derivative rather than an indolenine, and the benzo-derivatized Cy3.5, Cy5.5, and Cy7.5 are exceptions to this rule.

[0102] In some embodiments, the reporter moiety can be a FRET pair, so that multiple classifications can be performed in a single excitation and imaging step. As used herein, FRET can include excitation exchange (Forster) transfer or electron exchange (Dexter) transfer.

[0103] As used herein, the term "support" refers to a substrate designed for the deposition of biological molecules or biological samples for measurement and / or analysis. Examples of biological molecules to be deposited on a support include nucleic acids (e.g., DNA, RNA), polypeptides, carbohydrates, lipids, single cells or multiple cells. Examples of biological samples include, but are not limited to, saliva, sputum, mucus, blood, plasma, serum, urine, feces, sweat, tears, and fluids from tissues or organs.

[0104] In some embodiments, the support is solid, semi-solid, or a combination thereof. In some embodiments, the support is porous, semi-porous, non-porous, or any combination of porosities. In some embodiments, the support is substantially planar, concave, convex, or any combination thereof. In some embodiments, the support is cylindrical, for example, comprising a capillary or the inner surface of a capillary.

[0105] In some embodiments, the surface of the support can be substantially smooth. In some embodiments, the support can have a regular or irregular texture, including bumps, etchings, holes, three-dimensional scaffolds, or any combination thereof.

[0106] In some embodiments, the support comprises microbeads having any shape, including spheres, hemispheres, cylinders, barrels, rings, disks, rods, cones, triangles, cubes, polygons, tubes, or threads.

[0107] The support can be made of any material, including but not limited to: glass, fused silica, silicon, polymers (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high-density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic substrates are contemplated.

[0108] The present disclosure provides multiple (for example, two or more) nucleic acid template molecules that are fixed to a support. In certain embodiments, the multiple nucleic acid template molecules fixed have the same sequence or have different sequences. In certain embodiments, each nucleic acid template molecule in a plurality of nucleic acid template molecules is fixed to different sites on the support. In certain embodiments, two or more independent nucleic acid template molecules in the multiple nucleic acid templates are fixed to sites on the support.

[0109] The term "array" refers to a support comprising a plurality of sites located at predetermined positions on the support to form an array of sites. The sites may be discrete and separated by gap regions. In some embodiments, the predetermined sites on the support may be arranged in one dimension into rows or columns, or in two dimensions into rows and columns. In some embodiments, the plurality of predetermined sites are arranged in an organized manner on the support. In some embodiments, the plurality of predetermined sites are arranged in any organized pattern, including straight lines, hexagonal patterns, grid patterns, patterns with reflection symmetry, patterns with rotational symmetry, and the like. The spacing between different pairs of sites may be the same or may be different. In some embodiments, the support comprises at least 10 2 sites, at least 10 3 sites, at least 10 4 sites, at least 10 5 sites, at least 10 6 sites, at least 10 7 sites, at least 10 8 sites, at least 10 9 sites, at least 10 10 sites, at least 10 11 sites, at least 10 12 sites, at least 10 13 sites, at least 10 14 sites, at least 10 15 In some embodiments, the plurality of predetermined sites (e.g., 10 2 -10 15 In some embodiments, the nucleic acid template molecules are immobilized at a plurality of predetermined sites by hybridizing with an immobilized surface capture primer, or the nucleic acid template molecules are covalently attached to the surface capture primer. In some embodiments, the nucleic acid template molecules are immobilized at a plurality of predetermined sites, for example, at 10 2 –10 15 In certain embodiments, the fixed nucleic acid template molecules are cloned and amplified to generate fixed nucleic acid clusters at multiple predetermined sites. In certain embodiments, each fixed nucleic acid cluster comprises a linear cluster, or comprises a single-stranded or double-stranded concatemer.

[0110] In some embodiments, a support comprising a plurality of sites located at random positions on the support is referred to herein as a support having randomly positioned sites thereon. The positions of the randomly positioned sites on the support are not predetermined. The plurality of randomly positioned sites are arranged in a disordered and / or unpredictable manner on the support. In some embodiments, the support comprises at least 10 2sites, at least 10 3 sites, at least 10 4 sites, at least 10 5 sites, at least 10 6 sites, at least 10 7 sites, at least 10 8 sites, at least 10 9 sites, at least 10 10 sites, at least 10 11 sites, at least 10 12 sites, at least 10 13 sites, at least 10 14 sites, at least 10 15 sites or more sites, wherein the sites are randomly positioned on the support. In some embodiments, the plurality of randomly positioned sites (e.g., 10 2 –10 15 In some embodiments, the nucleic acid template molecules are immobilized at a plurality of randomly positioned sites by hybridization with an immobilized surface capture primer, or the nucleic acid template molecules are covalently attached to a surface capture primer. In some embodiments, the nucleic acid template is immobilized at a plurality of randomly positioned sites, for example, at 10 2 to 10 15 In certain embodiments, the nucleic acid templates are cloned and amplified to generate fixed nucleic acid clusters at multiple randomly positioned sites. In certain embodiments, each fixed nucleic acid cluster comprises a linear cluster, or comprises a single-stranded or double-stranded concatemer.

[0111] In some embodiments, a plurality of fixed surface capture primers on a support (e.g., located at a predetermined or random position on the support) are fluidically connected to each other to allow a solution of a reagent (e.g., a nucleic acid template molecule, a soluble primer, an enzyme, a nucleotide, a divalent cation, a buffer, etc.) to flow onto the support so that a plurality of fixed surface capture primers on the support can react with the reagent in a large-scale parallel manner substantially at the same time. In some embodiments, the fluid communication of a plurality of fixed surface capture primers can be used to carry out nucleic acid amplification reactions (e.g., RCA, MDA, PCR, and bridge amplification) substantially simultaneously on the plurality of fixed surface capture primers. Exemplary supports allowing fluid communication to allow solution flow include, but are not limited to, the inner surface of a circulation pool or a capillary.

[0112] In some embodiments, the multiple fixed nucleic acid clusters on the support are in fluid communication with each other to allow a solution of a reagent (e.g., an enzyme, a nucleotide, a divalent cation, etc.) to flow onto the support so that the multiple fixed nucleic acid clusters on the support can react with the reagent in a large-scale parallel manner substantially simultaneously. In some embodiments, the fluid communication of the multiple fixed nucleic acid clusters can be used to perform nucleotide binding assays and / or perform nucleotide polymerization reactions (e.g., primer extension or sequencing) on ​​the multiple fixed nucleic acid clusters substantially simultaneously, and optionally perform detection and imaging of large-scale parallel sequencing.

[0113] The term "fixed" and related terms refer to nucleic acid molecules that are attached to a support or the coating on the support or are embedded in the matrix formed by the coating on the support by covalent bonds or non-covalent interactions, wherein nucleic acid molecules include the extension products of surface capture primers, nucleic acid template molecules and capture primers. The extension products of capture primers include nucleic acid concatemers (for example, nucleic acid clusters). Nucleic acid molecules can be fixed on a predetermined or random position on the support. Nucleic acid molecules can be fixed on a predetermined or random position on or within a passivated coating on the support.

[0114] The term "immobilized" and related terms can also refer to an enzyme (e.g., a polymerase) that is attached to a support, or to a coating on a support, or embedded in a substrate formed by a coating on a support, by covalent bonds or non-covalent interactions. The enzyme can be immobilized at a predetermined or random position on the support. The enzyme can be immobilized at a predetermined or random position on or within a passivated coating on a support.

[0115] In some embodiments, one or more nucleic acid template molecules are fixed on a support, for example, fixed at a random or predetermined position on a support. In some embodiments, one or more nucleic acid template molecules are clonally amplified. In some embodiments, one or more nucleic acid template molecules are clonally amplified from a support (for example, in a solution), then deposited onto a support and fixed on a support. In some embodiments, the clonal amplification reaction of one or more nucleic acid template molecules is carried out on a support, resulting in immobilization on a support. In some embodiments, one or more nucleic acid template molecules are clonally amplified (for example, in a solution or on a support) using a nucleic acid amplification reaction, including any one or any combination of the following: polymerase chain reaction (PCR), multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, bridge amplification, isothermal bridge amplification, rolling circle amplification (RCA), loop-to-loop amplification, helicase-dependent amplification, recombinase-dependent amplification and / or single-stranded binding (SSB) protein-dependent amplification.

[0116] The terms "surface primer," "capture primer," "surface capture primer" and related terms refer to single-stranded oligonucleotides that are fixed to a support and comprise a sequence that can hybridize with at least a portion of a nucleic acid template molecule. Surface capture primers can be used to fix the template molecule to a support via hybridization. The surface capture primer can be fixed to the support in a manner that resists primer removal during flow, washing, suction, and changes in temperature, pH, salt, chemicals, and / or enzyme conditions. Typically, but not necessarily, the 5' end of the surface capture primer can be fixed to a support or a coating on a support (or embedded in a coating on a support). Alternatively or in addition, the inner portion or 3' end of the surface capture primer can be fixed to a carrier support.

[0117] The sequence of the surface capture primer can be fully or partially complementary to at least a portion of the nucleic acid template molecule along its length. The support can include a plurality of fixed surface capture primers with the same sequence or with two or more different sequences. The surface capture primer can be any length, for example 4-50 nucleotides, or 50-100 nucleotides, or 100-150 nucleotides, or longer lengths.

[0118] The surface capture primer can have terminal 3' nucleotides, and this terminal 3' nucleotides has sugar 3' OH part, and it can be extended for nucleotide polymerization (for example, polymerization catalyzed by polymerase).Surface capture primer can have terminal 3' nucleotides, and this terminal 3' nucleotides has 3' sugar position linked with the chain termination part that inhibits nucleotide polymerization.Can use deblocking agent to remove (for example, to block) 3' chain termination part, so that 3' end is converted into extendable 3' OH end.The example of chain termination part includes alkyl group, alkenyl group, alkynyl group, allyl group, aryl group, benzyl group, azido group, amine group, amide group, keto group, isocyanate group, phosphate group, thio group, disulfide group, carbonate group, urea group or silyl group.Azide type chain termination part includes azido, azido and azidomethyl group. Examples of deblocking agents include phosphine compounds such as tris(2-carboxyethyl)phosphine (TCEP) and bissulfotriphenylphosphine (BS-TPP) for azide, azido and azidomethyl chain terminating groups. Examples of deblocking agents include tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine, or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ) for alkyl, alkenyl, alkynyl and allyl chain terminating groups. Examples of deblocking agents include Pd / C for aryl and benzyl chain terminating groups. Examples of deblocking agents include phosphine, β-mercaptoethanol or dithiothreitol (DTT) for amine, amide, ketone, isocyanate, phosphate, thio and disulfide chain terminating groups. Examples of deblocking agents include potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine and Zn in acetic acid (AcOH) for carbonate chain terminating groups. Examples of deblocking agents include tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, and triethylamine trihydrofluoride for the chain-terminating groups urea and silyl.

[0119] The term "sequencing" and related terms refer to methods for obtaining nucleotide sequence information from nucleic acid molecules, typically by determining the identity of at least some nucleotides (including their nucleobase components) within the nucleic acid molecule. The sequence information for a given region of a nucleic acid molecule can include identifying each nucleotide in the region being sequenced. Alternatively, the sequencing information only determines some nucleotides in the region, while the identities of some nucleotides remain undetermined or incorrectly determined. Any suitable method for sequencing can be used. For example, sequencing can include label-free or ion-based sequencing methods. As a further example, sequencing can include labeled or dye-based nucleotide or fluorescent nucleotide sequencing methods. Sequencing can include polymerase cloning-based sequencing or bridge sequencing methods. Sequencing can employ a polymerase and a multivalent molecule to generate at least one affinity complex, wherein each multivalent molecule comprises a plurality of nucleotide units ( Figures 20 to 24Sequencing can be performed by sequencing-by-synthesis using a polymerase and free nucleotides. Sequencing can also be performed by sequencing-by-ligation using a ligase and multiple sequence-specific oligonucleotides.

[0120] Double-stranded splint adapter

[0121] The present disclosure provides compositions comprising nucleic acid double-stranded splint adaptors, including kits, and methods employing double-stranded splint adaptors.

[0122] The double-stranded splint adapter (200) can be used in a one-pot multi-enzyme reaction to introduce one or more new adapter sequences into the library molecule (100). The double-stranded splint adapter (200) comprises a first splint strand (long splint strand (300)) and a second splint strand (short splint strand (400)), wherein the first splint strand and the second splint strand are hybridized together to form a double-stranded splint adapter (200) having a double-stranded region and two flanking single-stranded regions (see, for example, Figures 1 to 8 ). The second splint chain (400) carries a new adapter sequence to be introduced, such as, for example, a new universal binding sequence and / or a new index sequence. The first splint chain comprises a first region (320), an internal region (310) and a second region (330). The internal region of the first splint chain (310) hybridizes with the second splint chain (400). The two flanking single-stranded regions (e.g., (320) and (330)) of the double-stranded splint adapter are designed to hybridize with the universal adapter sequence at the end of a single-stranded linear library molecule (100) having a target sequence (110). For example, the first region (320) of the first splint chain hybridizes with one end of the library molecule, and the second region (330) of the first splint chain hybridizes with the other end of the library molecule, thereby looping the library molecule to produce a library-splint complex (500) comprising two gaps (e.g., see Figures 1 to 8 The gap can be enzymatically ligatable to produce a covalently closed circular molecule (600) in which the second splint strand (400) is covalently linked to the library molecule at both ends, thereby introducing a new adapter sequence into the library molecule (see Figure 9 ).

[0123] Thus, the double-stranded splint adapters and methods described herein can be used to convert any linear library molecule into a covalently closed circular molecule that can be used in different workflows, such as different massively parallel sequencing platforms. The double-stranded splint adapter provides flexibility because the two flanking single-stranded regions (e.g., (320) and (330)) and the second splint chain (400) can be designed to include any combination of universal adapter sequences. For example, the two flanking single-stranded regions (e.g., (320) and (330)) can include universal binding sequences (or their complements) of P5 and P7 sequences that bind to surface primers immobilized on a support (e.g., a flow cell), where the P5 and P7 sequences are typically used to construct library molecules for the Illumina sequencing platform. The second splint chain (400) can include at least one new universal adapter sequence (e.g., a new surface primer sequence) that is not found on the Illumina sequencing platform, thereby allowing the covalently closed circular molecule (600) to be used on non-Illumina sequencing platforms.

[0124] The methods described herein also offer the advantage of using a ligation reaction rather than a gap-filling reaction to introduce new adapter sequences. The ligation reaction can achieve efficient circularization using as little as 0.25 pmol of library molecules.

[0125] The methods described herein can be performed manually or for automation because annealing and multi-enzyme reactions can be performed in a single reaction vessel (one pot) by combining some enzymatic reactions (e.g., phosphorylation and ligation) and by adding subsequent enzymes (e.g., exonucleases), without the need for intermediate alcohol precipitation or organic extraction.

[0126] The present disclosure provides a nucleic acid double-stranded splint adaptor (200) comprising: (i) a first splint strand (long splint strand (300)) hybridized to (ii) a second splint strand (short splint strand (400)) (e.g., see Figures 1 to 8 ). The first splint chain comprises a first region (320), an internal region (310) and a second region (330). The internal region (310) of the first splint chain hybridizes with the second splint chain (400) to form a double-stranded splint adapter (200) having a double-stranded region and two flanking single-stranded regions. The two flanking single-stranded regions of the double-stranded splint adapter (200) are designed to hybridize with the end sequences of the linear nucleic acid library molecules. The end sequences of the linear nucleic acid library molecules respectively comprise the first and second universal adapter sequences. In some embodiments, the first universal adapter sequence and the second universal adapter sequence of the linear library molecule respectively comprise binding sequences for the first capture primer and the second capture primer fixed on the support.

[0127] The first region (320) of the first splint strand comprises a first universal adaptor sequence that can hybridize to a first universal binding sequence at one end of a linear nucleic acid library molecule (e.g., see Figures 1 to 8The second region (330) of the first splint strand comprises a second universal adapter sequence that can hybridize to a second universal binding sequence at the other end of the linear nucleic acid library molecule (e.g., see Figures 1 to 8 ). In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence comprising a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3'OH group or a terminal 3' blocking group.

[0128] In some embodiments, the second cleat chain (400) includes at least two sub-regions, including first and second sub-regions (see, for example, Figure 2 and Figure 3 In some embodiments, the first subregion comprises a universal binding sequence for a third surface primer, and the second subregion comprises a universal binding sequence for a fourth surface primer, wherein the first subregion and the second subregion do not hybridize (or at least exhibit very little hybridization) with the first surface primer and the second surface primer. In some embodiments, the second splint strand (400) further comprises an optional third subregion comprising a sample index sequence of 5-20 bases and / or a unique identification sequence (e.g., NN) of 2-10 or more bases (e.g., see Figure 3 ). In some embodiments, the second splint chain (400) comprises only one subregion and lacks a second subregion and a third subregion, wherein the first subregion comprises a sample index sequence having 5-20 bases. In some embodiments, the sample index sequence can be used to distinguish target sequences obtained from different sample sources in multiplexed assays. In some embodiments, the unique identification sequence comprises a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization with the first, second, third, and fourth surface primers. An exemplary arrangement of the subregions in the second splint chain (400) in the 5' to 3' direction comprises: 5'-[second subregion]–[first subregion]–3'. Another exemplary arrangement of the subregions in the second splint chain (400) in the 5' to 3' direction comprises: 5'-[third subregion]–[second subregion]–[first subregion]–3' (see, for example, Figure 2 and Figure 3 ). In some embodiments, the length of the second splint chain (400) may be 20-100 nucleotides, or 30-80 nucleotides, or 40-60 nucleotides in length. In some embodiments, the 5' end of the second splint chain (400) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the second splint chain (400) comprises a terminal 3'OH group or a terminal 3' blocking group. In some embodiments, the second splint chain (400) comprises one or more thiophosphate bonds at the 5' and / or 3' end to impart exonuclease resistance. In some embodiments, the second splint chain (400) comprises one or more thiophosphate bonds at an internal position to impart endonuclease resistance. In some embodiments, the second splint chain (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.

[0129] In some embodiments, the first cleat chain (300) includes an inner region (310) comprising at least two sub-regions, including a fourth sub-region and a fifth sub-region (e.g., see Figure 2 and Figure 3 ). The fourth subregion hybridizes to the first subregion of the second splint strand (400). The fifth subregion hybridizes to the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize to (or at least exhibit very little hybridization with) the first and second surface primers. In some embodiments, the interior region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes to the third subregion of the second splint strand (400) (e.g., see Figure 3 ). An exemplary arrangement of the sub-regions of the first splint chain (300) in the 5' to 3' direction includes: 5'-[fourth sub-region]-[fifth sub-region]-3'. Another exemplary arrangement of the sub-regions of the first splint chain (300) in the 5' to 3' direction includes: 5'-[fourth sub-region]-[fifth sub-region]-[sixth sub-region]-3' (see, for example, Figure 2 and Figure 3 ). In some embodiments, the length of the first splint chain (300) may be 50-150 nucleotides, or 60-100 nucleotides, or 70-90 nucleotides in length. In some embodiments, the first splint chain (300) comprises one or more phosphorothioate bonds at the 5' and / or 3' end to confer resistance to exonucleases. In some embodiments, the first splint chain (300) comprises one or more phosphorothioate bonds at an internal position to confer resistance to endonucleases. In some embodiments, the first splint chain (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.

[0130] Long splint chain variant: shortened

[0131] In some embodiments, the first splint strand (300) comprises a truncated strand having a first region (320) with a truncated sequence at the 5' end (e.g., Figure 11B ). In some embodiments, the 5' end of the first region can have a truncation of any length, for example, a truncation of 1-10 nucleotides. In some embodiments, the truncated first splint chain (300) comprises a second region (330; for example, SEQ ID NO: 5), a fourth subregion (for example, SEQ ID NO: 6), and a fifth subregion (for example, SEQ ID NO: 7) that are not truncated and do not carry any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the truncated first splint chain comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint chain (400) to form a double-stranded splint adapter (200). In some embodiments, the truncated first splint chain (300) comprises a truncated first region (320) that hybridizes to a sequence (for example, 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (for example, 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, a truncated first splint strand that is part of a double-stranded splint adaptor (200) can hybridize to a library molecule (100) to form a library-splint complex (500).

[0132] Variants of long splint chains: mismatched sequences

[0133] In some embodiments, the first splint strand (300) comprises a mismatch strand having a first region (320) with a mismatch sequence (e.g., Figure 11C ). In some embodiments, the mismatch sequence is located within the first region. In some embodiments, the mismatch sequence can be of any length (e.g., 2-20 bases) and comprises any sequence that is not completely complementary to the left universal adapter sequence (120) of the library molecule (100). Some embodiments of the mismatch sequence in the first region (320) are Figure 11CIn some embodiments, the mismatched first splint chain (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7) that do not carry any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint chain (300) comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint chain (400) to form a double-stranded splint adapter (200). In some embodiments, the mismatched first splint chain (300) comprises a mismatched first region (320) hybridized to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) hybridized to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the mismatched first region (320) can hybridize to the first left universal adapter sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatched sequence in the first region (320). In some embodiments, the mismatched first splint strand that is part of the double-stranded splint adapter (200) can hybridize to the library molecule (100) to form a library-splint complex (500).

[0134] Variants of long splint chains: abasic sites

[0135] In some embodiments, the first splint strand (300) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site in the fourth subregion and / or at least one abasic site in the fifth subregion (e.g., Figure 11D ). In some embodiments, the abasic sites each comprise a 1',2'-dideoxyribose sugar (e.g., dSpacer from Integrated DNA Technologies (IDT)).

[0136] In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO:4), and a second region ((330); e.g., SEQ ID NO:5), which does not carry any abasic sites and / or any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the abasic first splint strand (300) comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the abasic first splint strand that is part of the double-stranded splint adaptor (200) can hybridize to a library molecule (100) to form a library-splint complex (500).

[0137] Long splint chain variant: uracil

[0138] In some embodiments, the first splint strand (300) comprises at least one uracil. In some embodiments, the first splint strand (300) comprises at least one uracil in any one or any combination of the regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base may be substituted with a uracil. Examples of first splint strands comprising uracil are shown in Figure 11D The skilled person will recognize that many other sequences of the first splint strand (300) comprising one or more uracils are possible.

[0139] In some embodiments, the uracil-containing first splint strand comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint strand (400) to form a double-stranded splint adaptor (200). In some embodiments, the uracil-containing first splint strand (300) comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of a linear single-stranded library molecule (100), and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the uracil-containing first splint strand as part of the double-stranded splint adaptor (200) can hybridize with a library molecule (100) to form a library-splint complex (500).

[0140] Insert or replace short splint chains in random sequence (400)

[0141] In some embodiments, the second cleat chain (400) comprises a random sequence (eg, Figure 7A 、 Figure 12A and Figure 12B In some embodiments, the random sequence may replace a portion of the first sub-region of the second splint chain (e.g., Figure 7A 、 Figure 13A and Figure 13B ).

[0142] In some embodiments, the second sub-region of the second cleat chain (400) does not have an inserted random sequence. In some embodiments, a portion of the second sub-region of the second cleat chain (400) is not replaced by a random sequence.

[0143] In some embodiments, the random sequence can be of any length, for example, 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 12A and Figure 13A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 12B and Figure 13B In some embodiments, the random sequence may be inserted at any position in the first sub-region of the second splint chain (400).

[0144] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0145] In some embodiments, random sequences provide nucleotide diversity and color balance for sequencing reactions. In some embodiments, random sequences provide high nucleotide diversity, comprising approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U), which will be represented in each cycle of a sequencing run.

[0146] In some embodiments, the random sequence can be sequenced before the insertion zone is sequenced. In some embodiments, the sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides enough nucleotide diversity and color balance. In some embodiments, the sequence of any portion of the left index (160), right index (170) and / or insertion zone (110) does not provide enough nucleotide diversity to achieve polymerase clone mapping and / or template alignment. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to any portion of the left index (160), right index (170) and / or insertion zone (110).

[0147] In some embodiments, a predetermined sequence is inserted into the sequence of the fourth sub-region of the first splint chain (300). The length of the inserted predetermined sequence may be the same as the length of the random sequence inserted into the first sub-region of the second splint chain (e.g., Figure 12A and Figure 12B ). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 12A and Figure 12B shown.

[0148] In some embodiments, a portion of the fourth sub-region of the first cleat chain (300) is replaced by a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth sub-region is the same as the length of the random sequence replacing the portion of the first sub-region of the second cleat chain (e.g., Figure 13A and Figure 13B ). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 13A and Figure 13B shown.

[0149] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion, which can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adaptor (200) (e.g., Figure 12A 、 Figure 12B 、 Figure 13A and Figure 13B In some embodiments, the double-stranded splint adaptor (200) forms a bubble at the position where the random sequence is inserted or replaced.

[0150] In some embodiments, the second splint strand (400) as part of the double-stranded splint adaptor (200) carries a random sequence, which second splint strand can hybridize with the library molecule (100) to form a library-splint complex (500).

[0151] Short splint chain with additional random sequence and index sequence (400)

[0152] In some embodiments, the second splint chain (400) comprises a random sequence (eg, Figure 7B 、 Figure 14A and Figure 14B ). In some embodiments, the random sequence includes an index sequence.

[0153] In some embodiments, the second sub-section of the second cleat chain (400) does not have an additional random sequence.

[0154] In some embodiments, the additional random sequence can be of any length, such as 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 14A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 14B NNNN in the .

[0155] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0156] In some embodiments, random sequences provide nucleotide diversity and color balance for sequencing reactions. In some embodiments, random sequences provide high nucleotide diversity, comprising approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U), which will be represented in each cycle of a sequencing run.

[0157] In some embodiments, the random sequence can be sequenced before the insertion zone is sequenced. In some embodiments, the sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides enough nucleotide diversity and color balance. In some embodiments, the sequence of any portion of the left index (160), right index (170) and / or insertion zone (110) does not provide enough nucleotide diversity to achieve polymerase clone mapping and / or template alignment. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to any portion of the left index (160), right index (170) and / or insertion zone (110).

[0158] In some embodiments, index sequences can be used to distinguish polynucleotides (eg, insert sequences) from different sample sources in a multiplex assay.

[0159] In some embodiments, a predetermined sequence is attached to the 5' end of the fourth sub-region of the first splint chain (300). In some embodiments, the length of the attached predetermined sequence can be the same as the length of the random sequence attached to the 3' end of the first sub-region of the second splint chain (e.g., Figure 14A and Figure 14B In some embodiments, the length of the additional predetermined sequence may be the same as the length of the random sequence and index sequence attached to the 3' end of the first subregion of the second splint chain (e.g., Figure 14A and Figure 14B ). The additional predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 14A and Figure 14B shown.

[0160] In some embodiments, the second splint strand (400) includes a first subregion and a second subregion, which can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adaptor (200) (e.g., Figure 14A and Figure 14B In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the position of the additional random sequence. In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the position of the additional random sequence and the index sequence.

[0161] In some embodiments, a second splint strand (400) that is part of a double-stranded splint adaptor (200) is appended with a random sequence (and optionally an index sequence) that can hybridize with a library molecule (100) to form a library-splint complex (500).

[0162] Sequence of short splint chain (400)

[0163] In some embodiments, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence

[0164] 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises first and second sub-regions comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). Figure 11AIn some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or non-phosphorylated.

[0165] Sequence of long splint chain (300)

[0166] In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence comprising a universal binding sequence (or its complement) for a first surface primer, wherein the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complement of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a second surface primer (or its complement), wherein the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize with a P7 surface primer or a complement of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fourth subregion having the sequence

[0167] 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fifth subregion having the following sequence: 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO: 198). In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) comprising a fourth and a fifth subregion, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAG GCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11A In some embodiments, the 5' end of the first splint chain (300) can be phosphorylated or non-phosphorylated. In some embodiments, the first subregion of the second splint chain (400) can hybridize with the fourth subregion of the first splint chain (300). In some embodiments, the second subregion of the second splint chain (400) can hybridize with the fifth subregion of the first splint chain (300).

[0168] Library-splint complex

[0169] The present disclosure provides a library-splint complex (500) comprising: (i) a single-stranded nucleic acid library molecule (100) comprising a target sequence (110) flanked on one side by at least a first left universal adapter sequence (120) and on the other side by at least a first right universal adapter sequence (130); and (ii) a double-stranded splint adapter (200) comprising a first splint chain (a long splint chain (300)) and a second splint chain (a short splint chain (400)), wherein the first splint chain comprises a first region (320), an internal region (310) and a second region (330), wherein the internal region (310) of the first splint chain hybridizes with the second splint chain (400) to form a double-stranded splint adapter (200) having a double-stranded region flanked on both sides by single-stranded regions. In the library-splint complex (500), the first region (320) of the first splint strand hybridizes to at least the first left universal adapter sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adapter sequence (130) of the library molecule, thereby circularizing the library molecule to produce the library-splint complex (500) (see, e.g., Figures 1 to 8 ).

[0170] In the library-splint complex (500), the first region (320) of the first splint strand comprises a first universal adaptor sequence that can hybridize to a first universal binding sequence at one end of a linear nucleic acid library molecule (e.g., see Figures 1 to 8 ). In some embodiments, the first region (320) of the first splint chain includes a first universal adapter sequence comprising a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the first splint chain (300) may be 50-150 nucleotides in length, or 60-100 nucleotides in length, or 70-90 nucleotides in length. In some embodiments, the first splint chain (300) comprises one or more phosphorothioate bonds at the 5' and / or 3' end to confer nuclease resistance. In some embodiments, the first splint chain (300) comprises one or more phosphorothioate bonds at an internal position to confer nuclease resistance. In some embodiments, the first splint chain (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the first splint chain (300) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the first splint chain (300) comprises a terminal 3' OH group or a terminal 3' blocking group.

[0171] The second region (330) of the first splint strand comprises a second universal adaptor sequence that can hybridize to a second universal binding sequence at the other end of the linear nucleic acid library molecule (e.g., see Figures 1 to 8 ). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.

[0172] In the library-splint complex (500), the first region (320) of the first splint strand hybridizes to at least the first left universal adapter sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adapter sequence (130) of the library molecule, thereby looping the library molecule to produce the library-splint complex (500). The library-splint complex (500) comprises a first gap between the 5' end of the library molecule and the 3' end of the second splint strand. The library-splint complex (500) also comprises a second gap between the 5' end of the second splint strand and the 3' end of the library molecule (e.g., see Figures 1 to 8 ). In some embodiments, the first gap and the second gap are enzymatically ligatable.

[0173] In the library-splint complex (500), the first region (320) of the first splint strand can hybridize with the sense strand or the antisense strand of a double-stranded nucleic acid library molecule. In the library-splint complex (500), the second region (330) of the first splint strand can hybridize with the sense strand or the antisense strand of a double-stranded nucleic acid library molecule. The double-stranded nucleic acid library molecule can be denatured to generate single-stranded sense and antisense library strands.

[0174] In the library-splint complex (500), the second splint strand (400) does not hybridize to the target sequence (110), and the internal region (310) of the first splint strand does not hybridize to the target sequence (110).

[0175] In the library-splint complex (500), the first region (320) of the first splint strand does not hybridize to the target sequence (110), and the second region (330) of the first splint strand does not hybridize to the target sequence (110).

[0176] In some embodiments, in the library-splint complex (500), the 5' end of the single-stranded library molecule (100) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the single-stranded library molecule includes a terminal 3' OH group or a terminal 3' blocking group.

[0177] In some embodiments, the nucleic acid library molecule (100) comprises a second left universal adapter sequence (140). In some embodiments, the nucleic acid library molecule (100) comprises a second right universal adapter sequence (150). Exemplary library molecules (100) are shown in Figures 4 to 8 In some embodiments, the nucleic acid library molecule (100) may further comprise an additional left universal adapter sequence and / or a right universal adapter sequence.

[0178] In some embodiments, the nucleic acid library molecule (100) further comprises a first left index sequence (160). In some embodiments, the nucleic acid library molecule (100) further comprises a first right index sequence (170). In some embodiments, the first left index sequence (160) comprises a sample index sequence. In some embodiments, the first right index sequence (170) comprises another sample index sequence. In some embodiments, the first left index sequence (160) and the first right index sequence (170) are not the same sequence. In some embodiments, the nucleic acid library molecule (100) comprises a first left index sequence (160) and / or a first right index sequence (170). The sample index sequence can be used to distinguish target sequences obtained from different sample sources in multiplexed assays. An exemplary library molecule (100) is shown in Figures 4 to 8 A list of exemplary first left index sequences (160) and first right index sequences (170) is provided in Table 1 of FIG33. The first left index sequence (160) may include a random sequence (e.g., NNN) or lack a random sequence. The first right index sequence (170) may include a random sequence (e.g., NNN) or lack a random sequence.

[0179] In some embodiments, the nucleic acid library molecule (100) further comprises at least one ligation adapter sequence positioned between any of the universal adapter sequences described herein (e.g., see Figure 8 ). For example, the first left connecting adapter sequence (125) may be located between the first left universal adapter sequence (120) and the first left index sequence (160). The second left connecting adapter sequence (165) may be located between the first left index sequence (160) and the second left universal adapter sequence (140). The third left connecting adapter sequence (145) may be located between the second left universal adapter sequence (140) and the target sequence (110). The first right connecting adapter sequence (135) may be located between the first right universal adapter sequence (130) and the first right index sequence (170). The second right connecting adapter sequence (175) may be located between the first right index sequence (170) and the second right universal adapter sequence (150). The third right connecting adapter sequence (155) may be located between the second right universal adapter sequence (150) and the target sequence (110). In some embodiments, the nucleic acid library molecule (100) further comprises at least one, and up to ten, additional universal adapter sequences located 5' (upstream) of the first left universal adapter sequence (120) (e.g., see Figure 8 In some embodiments, the nucleic acid library molecule (100) further comprises at least one and up to ten additional universal adapter sequences located 3' (downstream) of the first right universal adapter sequence (130) (e.g., see Figure 8). Any connecting adapter sequence comprises any sequence and may be 3-60 nucleotides in length and / or an additional universal adapter sequence. Any connecting adapter sequence and / or an additional universal adapter sequence comprises a universal sequence or a unique sequence. Any connecting adapter sequence and / or an additional universal adapter sequence comprises a binding sequence for an amplification primer, a sequencing primer, a compaction oligonucleotide or a combination thereof. Any connecting adapter sequence and / or an additional universal adapter sequence comprises a binding sequence for an immobilized surface primer (e.g., a capture primer). Any connecting adapter sequence and / or an additional universal adapter sequence comprises a sample index sequence. Any connecting adapter sequence and / or an additional universal adapter sequence comprises a unique identification sequence. Any connecting adapter sequence and / or an additional universal adapter sequence, in particular as Figure 8 The ligation adapter sequence (145) shown comprises the Tn5 transposon end sequence 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 211). Any ligation adapter sequence and / or additional universal adapter sequence, particularly as Figure 8 The ligation adapter sequence (155) shown comprises the Tn5 transposon end sequence 5'-CTGTCTCTTATACACATCT-3' (SEQ ID NO: 212). The Tn5 transposon end sequence can be introduced into the library molecule (100) by a transposase-mediated reaction under conditions suitable for forming a transposon synaptic complex, the reaction comprising contacting double-stranded input DNA (e.g., genomic DNA) with a Tn-5 type transposase and a double-stranded oligonucleotide comprising the Tn transposon end sequence (SEQ ID NO: 211) linked to a universal adapter sequence or a sample index sequence. In the double-stranded oligonucleotide, the Tn transposon end sequence (SEQ ID NO: 211) can be located 5' or 3' relative to the universal adapter sequence or the sample index sequence.

[0180] A multiplex workflow is initiated by preparing a library with a sample index using one or two index sequences (e.g., a left index sequence and / or a right index sequence). A separate library with a sample index can be prepared using a first left index sequence (160) and / or a first right index sequence (170) using input nucleic acids isolated from different sources. Libraries with sample indexes can be pooled together to generate a multiple library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insertion region together with the first left index sequence (160) and / or the first right index sequence (170) can be used to identify the source of the input nucleic acid. In some embodiments, any number of libraries with sample indexes can be pooled together, for example, 2-10, or 10-50, or 50-100, or 100-200, or more than 200 libraries with sample indexes can be pooled. Exemplary nucleic acid sources include naturally occurring, recombinant, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or entire organisms. Exemplary nucleic acid sources include fresh, frozen, fresh frozen or archived sources (e.g., formalin fixed paraffin embedded; FFPE). The skilled artisan will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.

[0181] In some embodiments, the nucleic acid library molecule (100) further comprises Figure 5 In some embodiments, the nucleic acid library molecule (100) further comprises the following: Figure 6 Optional first right unique identification sequence (190) shown. In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence that uniquely identifies each target sequence (e.g., an insert sequence) to which a unique adapter is attached, among a population of other target molecules. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular labeling. Exemplary library molecules (100) are shown in Figures 4 to 8 middle.

[0182] In some embodiments, the nucleic acid library molecule (100) comprises any one or any combination of two or more of the following: a first left universal adapter sequence (120); a second left universal adapter sequence (140); a first left index sequence (160); a first left unique identification sequence (180); a first right universal adapter sequence (130); a second right universal adapter sequence (150); a first right index sequence (170); and / or a first right unique identification sequence (190). An exemplary library molecule (100) is shown in Figures 4 to 8 middle.

[0183] In some embodiments, the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) comprise: a universal binding sequence for a forward sequencing primer or a reverse sequencing primer; a universal binding sequence for a first surface primer or a second surface primer; a universal binding sequence for a forward amplification primer or a reverse amplification primer; and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises an additional left universal adapter sequence.

[0184] In some embodiments, the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) comprise: a universal binding sequence for a forward sequencing primer or a reverse sequencing primer; a universal binding sequence for a first surface primer or a second surface primer; a universal binding sequence for a forward amplification primer or a reverse amplification primer; and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises an additional right universal adapter sequence.

[0185] In some embodiments, the second cleat chain (400) includes at least two sub-regions, including first and second sub-regions (see, for example, Figure 2 and Figure 3 In some embodiments, the first subregion comprises a universal binding sequence for a third surface primer, and the second subregion comprises a universal binding sequence for a fourth surface primer, wherein the first subregion and the second subregion do not hybridize (or at least exhibit very little hybridization) with the first surface primer and the second surface primer. In some embodiments, the second splint strand (400) further comprises an optional third subregion comprising a sample index sequence of 5-20 bases and / or a unique identification sequence (e.g., NN) of 2-10 or more bases (e.g., see Figure 2 and 3). In some embodiments, the second splint chain (400) comprises only one subregion and lacks a second subregion and a third subregion, wherein the first subregion comprises a sample index sequence having 5-20 bases. In some embodiments, the sample index sequence can be used to distinguish target sequences obtained from different sample sources in multiple assays. In some embodiments, the unique identification sequence comprises a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization with the first, second, third, and fourth surface primers. An exemplary arrangement of the subregions in the second splint chain (400) in the 5' to 3' direction comprises: 5'-[second subregion]–[first subregion]–3'. Another exemplary arrangement of the subregions in the second splint chain (400) in the 5' to 3' direction comprises: 5'-[third subregion]–[second subregion]–[first subregion]–3'. In some embodiments, the second splint chain (400) may be 20-100 nucleotides in length, or 30-80 nucleotides in length, or 40-60 nucleotides in length. In some embodiments, the second splint chain (400) comprises one or more phosphorothioate bonds at the 5' and / or 3' ends to confer resistance to nucleic acid exonucleases. In some embodiments, the second splint chain (400) comprises one or more phosphorothioate bonds at an internal position to confer resistance to nucleic acid endonucleases. In some embodiments, the second splint chain (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' ends or at an internal position. In some embodiments, the 5' end of the second splint chain (400) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the second splint chain (400) comprises a terminal 3'OH group or a terminal 3' blocking group.

[0186] In some embodiments, the first cleat chain (300) includes an inner region (310) comprising at least two sub-regions, including a fourth and a fifth sub-region (see, e.g., Figure 2 and Figure 3 ). The fourth subregion hybridizes to the first subregion of the second splint strand (400). The fifth subregion hybridizes to the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize to (or at least exhibit very little hybridization with) the first and second surface primers. In some embodiments, the interior region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes to the third subregion of the second splint strand (400) (e.g., see Figure 2 and Figure 3 ). An exemplary arrangement of the sub-regions of the first splint chain (300) in the 5' to 3' direction includes: 5'-[fourth sub-region]-[fifth sub-region]-3'. Another exemplary arrangement of the sub-regions of the first splint chain (300) in the 5' to 3' direction includes: 5'-[fourth sub-region]-[fifth sub-region]-[sixth sub-region]-3'.

[0187] In some embodiments, an exemplary library-splint complex (500) comprises: (a) a single-stranded nucleic acid library molecule (100); (b) a first splint strand (300); and (c) a second splint strand (400).

[0188] In an exemplary library-splint complex (500), a single-stranded nucleic acid library molecule (100) comprises components arranged in 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer; (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; (iii) a target sequence (110); (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer; and (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer.

[0189] In the exemplary library-splint complex (500), the first splint strand (300) comprises components arranged in 5' to 3' order: a first region (320); an internal region (310); and a second region (330).

[0190] In the exemplary library-splint complex (500), the second splint strand (400) comprises subregions arranged in 3' to 5' order: a first subregion having a universal binding sequence for a third surface primer; and a second subregion having a universal binding sequence for a fourth surface primer.

[0191] In an exemplary library-splint complex (500), a portion of a first splint strand (300) hybridizes with a portion of a library molecule (100), thereby looping the library molecule to produce a library-splint complex (500) such that a first region (320) of the first splint strand hybridizes with a binding sequence (120) for a first surface primer, and a third region (330) of the first splint strand hybridizes with a binding sequence (130) for a second surface primer. Additionally, a second splint strand (400) hybridizes with an internal region (310) of the first splint strand (300). The library-splint complex (500) comprises a first gap between the 5' end of the library molecule and the 3' end of the second splint strand, and a second gap between the 5' end of the second splint strand and the 3' end of the library molecule, and the first gap and the second gap are enzymatically ligatable.

[0192] In the exemplary library-splint complex (500), the second splint strand (400) does not hybridize to the sequence of interest (110), and the internal region (310) of the first splint strand does not hybridize to the sequence of interest (110).

[0193] In the exemplary library-splint complex (500), the first region (320) of the first splint strand does not hybridize to the sequence of interest (110), and the second region (330) of the first splint strand does not hybridize to the sequence of interest (110).

[0194] In some embodiments, any of the library-splint complexes (500) described herein comprises a plurality of library-splint complexes (500), wherein the sequence of interest (110) of each library-splint complex in the plurality of library-splint complexes comprises the same sequence of interest or a different sequence of interest.

[0195] Library splint complexes formed with double-stranded adapters with truncated long splint strands

[0196] In some embodiments, the library-splint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adaptor (200). In some embodiments, the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-splint complex (500) having a first gap and a second gap (e.g., Figures 1 to 6 and Figure 8 ).

[0197] In some embodiments, the first splint strand (300) comprises a truncated strand having a first region (320) with a truncated sequence at the 5' end (e.g., Figure 11B ). In some embodiments, the first region has a truncated sequence at the 5' end when compared to SEQ ID NO: 199, such as Figure 11AAs shown. In some embodiments, the 5' end of the first region can have a truncation of any length, such as a truncation of 1-10 nucleotides. In some embodiments, the truncated first splint chain (300) comprises a second region (330; for example, SEQ ID NO: 5), a fourth subregion (for example, SEQ ID NO: 6) and a fifth subregion (for example, SEQ ID NO: 7) that are not truncated and do not carry any sequence variants, such as insertions, deletions or base substitutions. In some embodiments, the truncated first splint chain comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint chain (400) to form a double-stranded splint adapter (200). In some embodiments, the truncated first splint chain (300) comprises a truncated first region (320) hybridized with a sequence (for example, 120) on one end of the linear single-stranded library molecule (100) and a second region (330) hybridized with a sequence (for example, 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, a truncated first splint strand that is part of a double-stranded splint adaptor (200) can hybridize to a library molecule (100) to form a library-splint complex (500) having a first gap and a second gap.

[0198] Library splint complexes formed with double-stranded adapters containing long splint strands with mismatched sequences

[0199] In some embodiments, the library-splint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adaptor (200). In some embodiments, the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adaptor sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-splint complex (500) having a first gap and a second gap (e.g., Figures 1 to 8 ).

[0200] In some embodiments, the first splint strand (300) comprises a mismatch strand having a first region (320) with a mismatch sequence (e.g., Figure 11C ). In some embodiments, the mismatch sequence can be of any length (e.g., 2-20 bases) and comprises any sequence that is not completely complementary to the left universal adapter sequence (120) of the library molecule (100). Some embodiments of the mismatch sequence in the first region (320) are Figure 11CIn some embodiments, the mismatched first splint chain (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7) that do not carry any sequence variants, such as, for example, insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint chain comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint chain (400) to form a double-stranded splint adapter (200). In some embodiments, the mismatched first splint chain (300) comprises a mismatched first region (320) hybridized to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) hybridized to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the mismatched first region (320) can hybridize to the first left universal adapter sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatched sequence in the first region (320). In some embodiments, the mismatched first splint strand that is part of the double-stranded splint adapter (200) can hybridize to the library molecule (100) to form a library-splint complex (500) having a first gap and a second gap.

[0201] Library splint complexes formed using double-stranded adapters with long splint strands containing abasic sites

[0202] In some embodiments, the library-splint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adaptor (200). In some embodiments, the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adaptor sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-splint complex (500) having a first gap and a second gap (e.g., Figures 1 to 8 ).

[0203] In some embodiments, the first splint strand (300) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site in the fourth subregion and / or at least one abasic site in the fifth subregion (e.g., Figure 11D ). In some embodiments, the abasic sites each comprise a 1',2'-dideoxyribose sugar (e.g., dSpacer from Integrated DNA Technologies (IDT)).

[0204] In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO:4), and a second region ((330); e.g., SEQ ID NO:5), which does not carry any abasic sites and / or any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the abasic first splint strand (300) comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, an abasic first splint strand that is part of a double-stranded splint adaptor (200) can hybridize to a library molecule (100) to form a library-splint complex (500) having a first gap and a second gap.

[0205] Library splint complexes formed with double-stranded adapters having long splint strands of uracil

[0206] In some embodiments, the library-splint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adaptor (200). In some embodiments, the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adaptor sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-splint complex (500) having a first gap and a second gap (e.g., Figures 1 to 8 ).

[0207] In some embodiments, the first splint strand (300) comprises at least one uracil in any one or any combination of the regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base may be substituted with uracil. Examples of first splint strands comprising uracil are shown in Figure 11D The skilled person will recognize that many other sequences of the first splint strand (300) comprising one or more uracils are possible.

[0208] In some embodiments, the first uracil-containing splint strand comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint strand (400) to form a double-stranded splint adaptor (200). In some embodiments, the first uracil-containing splint strand (300) comprises a first region (320) that hybridizes with a sequence (e.g., 120) at one end of a linear single-stranded library molecule (100), and a second region (330) that hybridizes with a sequence (e.g., 130) at the other end of the linear single-stranded library molecule (100). In some embodiments, the first uracil-containing splint strand as part of the double-stranded splint adaptor (200) can hybridize with a library molecule (100) to form a library-splint complex (500) having a first gap and a second gap.

[0209] Library splint complexes formed with double-stranded adapters with short splint strands of random sequence

[0210] In some embodiments, the library-splint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adaptor (200). In some embodiments, the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adaptor sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-splint complex (500) having a first gap and a second gap (e.g., Figures 1 to 8 ).

[0211] In some embodiments, the second cleat chain (400) comprises a random sequence (eg, Figure 7A 、 Figure 12A and Figure 12B In some embodiments, the random sequence may replace a portion of the first sub-region of the second splint chain (e.g., Figure 7A 、 Figure 13A and Figure 13B ).

[0212] In some embodiments, the second sub-region of the second cleat chain (400) does not have an inserted random sequence. In some embodiments, a portion of the second sub-region of the second cleat chain (400) is not replaced by a random sequence.

[0213] In some embodiments, the random sequence can be of any length, for example, 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 12A and Figure 13A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 12B and Figure 13BIn some embodiments, the random sequence may be inserted at any position in the first sub-region of the second splint chain (400).

[0214] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0215] In some embodiments, random sequences provide nucleotide diversity and color balance for sequencing reactions. In some embodiments, random sequences provide high nucleotide diversity, comprising approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U), which will be represented in each cycle of a sequencing run.

[0216] In some embodiments, the random sequence can be sequenced before the insertion zone is sequenced. In some embodiments, the sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides enough nucleotide diversity and color balance. In some embodiments, the sequence of any portion of the left index (160), right index (170) and / or insertion zone (110) does not provide enough nucleotide diversity to achieve polymerase clone mapping and / or template alignment. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to any portion of the left index (160), right index (170) and / or insertion zone (110).

[0217] In some embodiments, a predetermined sequence is inserted into the sequence of the fourth sub-region of the first splint chain (300). The length of the inserted predetermined sequence may be the same as the length of the random sequence inserted into the first sub-region of the second splint chain (e.g., Figure 12A and Figure 12B ). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 12A and Figure 12B shown.

[0218] In some embodiments, a portion of the fourth sub-region of the first cleat chain (300) is replaced by a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth sub-region is the same as the length of the random sequence replacing the portion of the first sub-region of the second cleat chain (e.g., Figure 13A and Figure 13B). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 13A and Figure 13B shown.

[0219] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion, which can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adaptor (200) (e.g., Figure 12A 、 Figure 12B 、 Figure 13A and Figure 13B In some embodiments, the double-stranded splint adaptor (200) forms a bubble at the position where the random sequence is inserted or replaced.

[0220] In some embodiments, a second splint strand (400) that is part of a double-stranded splint adaptor (200) carries a random sequence and can hybridize with a library molecule (100) to form a library-splint complex (500) having a first gap and a second gap.

[0221] Library splints formed with double-stranded adapters with short splint strands of attached random and index sequences Compound

[0222] In some embodiments, the library-splint complex (500) comprises a library molecule (100) hybridized to a first splint strand (300) of a double-stranded splint adaptor (200). In some embodiments, the first region (320) of the first splint strand hybridizes to at least the first left universal adaptor sequence (120) of the library molecule, and the second region (330) of the first splint strand hybridizes to at least the first right universal adaptor sequence (130) of the library molecule, thereby circularizing the library molecule to generate a library-splint complex (500) having a first gap and a second gap (e.g., Figures 1 to 8 ).

[0223] In some embodiments, the second splint chain (400) comprises a random sequence (eg, Figure 7B 、 Figure 14A and Figure 14B ). In some embodiments, the random sequence further comprises an index sequence.

[0224] In some embodiments, the second sub-section of the second cleat chain (400) does not have an additional random sequence.

[0225] In some embodiments, the additional random sequence can be of any length, such as 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 14A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 14B NNNN in the .

[0226] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0227] In some embodiments, random sequences provide nucleotide diversity and color balance for sequencing reactions. In some embodiments, random sequences provide high nucleotide diversity, comprising approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U), which will be represented in each cycle of a sequencing run.

[0228] In some embodiments, the random sequence can be sequenced before the insertion zone is sequenced. In some embodiments, the sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides enough nucleotide diversity and color balance. In some embodiments, the sequence of any portion of the left index (160), right index (170) and / or insertion zone (110) does not provide enough nucleotide diversity to achieve polymerase clone mapping and / or template alignment. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to any portion of the left index (160), right index (170) and / or insertion zone (110).

[0229] In some embodiments, index sequences can be used to distinguish polynucleotides (eg, insert sequences) from different sample sources in a multiplex assay.

[0230] In some embodiments, a predetermined sequence is attached to the 5' end of the fourth sub-region of the first splint chain (300). In some embodiments, the length of the attached predetermined sequence can be the same as the length of the random sequence attached to the 3' end of the first sub-region of the second splint chain (e.g., Figure 14A and Figure 14B In some embodiments, the length of the additional predetermined sequence may be the same as the length of the random sequence and index sequence attached to the 3' end of the first subregion of the second splint chain (e.g., Figure 14A and Figure 14B ). The additional predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 14A and Figure 14B shown.

[0231] In some embodiments, the second splint strand (400) includes a first subregion and a second subregion, which can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adaptor (200) (e.g., Figure 14A and Figure 14B In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the position of the additional random sequence. In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the position of the additional random sequence and the index sequence.

[0232] In some embodiments, a second splint strand (400) that is part of a double-stranded splint adaptor (200) is appended with a random sequence (and an optional index sequence), which second splint strand can hybridize with a library molecule (100) to form a library-splint complex (500) having a first gap and a second gap.

[0233] Sequence of short splint chain (400)

[0234] In some embodiments of the library-splint complex (500) described herein, the first subregion of the second splint strand (400) comprises the sequence

[0235] 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence

[0236] 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises a first sub-region and a second sub-region comprising the sequence

[0237] 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or non-phosphorylated.

[0238] In some embodiments, the second splint strand (400) comprises only one subregion and lacks the second subregion and the third subregion, wherein the first subregion comprises a sample index sequence having 5-20 bases.

[0239] Sequence of long splint chain (300)

[0240] In some embodiments of the library-splint complex (500) described herein, the first region of the first splint strand (320) comprises a first universal adapter sequence comprising a universal binding sequence for a first surface primer (or its complement), wherein the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complement of a P5 surface primer. For example, a P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or a P5 surface primer comprises the sequence

[0241] 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; length P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence (or its complement) for a second surface primer, wherein the second region (330) comprises the sequence

[0242] 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize with a P7 surface primer or the complement of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fourth subregion having the following sequence: 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198). In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) comprising a fourth and a fifth subregion, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAGGCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11AIn some embodiments, the 5' end of the first splint chain (300) can be phosphorylated or non-phosphorylated. In some embodiments, the first subregion of the second splint chain (400) can hybridize with the fourth subregion of the first splint chain (300). In some embodiments, the second subregion of the second splint chain (400) can hybridize with the fifth subregion of the first splint chain (300).

[0243] Sequence of the library-splint complex

[0244] In some embodiments of the library-splint complex (500) described herein, the first region of the first splint chain (320) comprises a sequence that can bind to the first left universal adapter sequence (120) of the library molecule, wherein the first region of the first splint chain (320) comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO:215) or its complementary sequence.

[0245] In some embodiments of the library-splint complex (500) described herein, the second region of the first splint strand (330) comprises a sequence that can bind to the first right universal adapter sequence (130) of the library molecule, wherein the second region of the first splint strand (330) comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO:216) or its complement.

[0246] In some embodiments, in any library-splint complex (500) described herein, the library molecule comprises a first left universal adapter sequence (120) that binds to a first region of a first splint strand (320), wherein the left universal binding sequence (120) comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203).

[0247] In some embodiments, in any library-splint complex (500) described herein, the library molecule includes a first left universal adapter sequence (120) that binds to a first region of the first splint strand (320), wherein the first left universal adapter sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO:213) or its complement.

[0248] In some embodiments of the library-splint complex (500) described herein, the library molecule includes a second left universal adapter sequence that includes a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence includes the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO:204).

[0249] In some embodiments of the library-splint complex (500) described herein, the library molecule includes a second left universal adapter sequence that includes a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence includes the sequence 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO:207).

[0250] In some embodiments of the library-splint complex (500) described herein, the library molecule includes a second left universal adapter sequence that includes a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence includes the sequence 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO:208).

[0251] In some embodiments of the library-splint complex (500) described herein, the library molecule includes a second right universal adapter sequence that includes a sequence for binding to a sequencing primer (150), wherein the right universal adapter sequence includes the sequence 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO:205).

[0252] In some embodiments of the library-splint complex (500) described herein, the library molecule includes a second right universal adapter sequence that includes a sequence for binding to a sequencing primer (150), wherein the second right universal adapter sequence includes the sequence 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO:209).

[0253] In some embodiments of the library-splint complex (500) described herein, the library molecule includes a second right universal adapter sequence that includes a sequence for binding to a sequencing primer (150), wherein the second right universal adapter sequence includes the sequence 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO:210).

[0254] In some embodiments of the library-splint complex (500) described herein, the library molecule comprises a first right universal adapter sequence (130) that binds to the first region (330) of the first splint strand, wherein the first right universal adapter sequence (130) comprises the sequence 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO:206).

[0255] In some embodiments of the library-splint complex (500) described herein, the library molecule comprises a first right universal adapter sequence (130) that binds to a first region (330) of a first splint strand, wherein the first right universal adapter sequence (130) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO:214) or its complement.

[0256] The present disclosure provides a reaction mixture comprising a plurality of any library-splint complexes (500) described herein. In some embodiments, the reaction mixture comprises a plurality of any library-splint complexes (500) described herein and T4 polynucleotide kinase. In some embodiments, the reaction mixture comprises a plurality of any library-splint complexes (500) described herein and a ligase. In some embodiments, the reaction mixture comprises a plurality of any library-splint complexes (500) described herein and a T4 polynucleotide kinase and a ligase. In some embodiments, the ligase comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase.

[0257] Covalently closed ring molecules

[0258] The present disclosure provides a covalently closed circular library molecule (600) comprising: a target sequence (110), at least a first left universal adapter sequence (120), at least a first right universal adapter sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9 In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adapter sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adapter sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adapter sequences.

[0259] In some embodiments, the covalently closed circular molecule (600) further comprises a first left index sequence (160) and / or a first right index sequence (170). The index sequences can be used to distinguish target sequences obtained from different sample sources in a multiplex assay. A list of exemplary first left index sequences (160) and first right index sequences (170) is provided in Table 1 of Figure 33. The first left index sequence (160) can include a random sequence (e.g., NNN) or lack a random sequence. The first right index sequence (170) can include a random sequence (e.g., NNN) or lack a random sequence.

[0260] A multiplex workflow is initiated by preparing a library with a sample index using one or two index sequences (e.g., a left index sequence and / or a right index sequence). A separate library with a sample index can be prepared using a first left index sequence (160) and / or a first right index sequence (170) using input nucleic acids isolated from different sources. Libraries with sample indexes can be pooled together to generate a multiple library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insertion region together with the first left index sequence (160) and / or the first right index sequence (170) can be used to identify the source of the input nucleic acid. In some embodiments, any number of libraries with sample indexes can be pooled together, for example, 2-10, or 10-50, or 50-100, or 100-200, or more than 200 libraries with sample indexes can be pooled. Exemplary nucleic acid sources include naturally occurring, recombinant, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or entire organisms. Exemplary nucleic acid sources include fresh, frozen, fresh frozen or archived sources (e.g., formalin fixed paraffin embedded; FFPE). The skilled artisan will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.

[0261] In some embodiments, the covalently closed circular molecule (600) further comprises an optional first left unique recognition sequence (180) and / or an optional first right unique recognition sequence (190), such as Figure 5 and Figure 6 In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence that uniquely identifies each target sequence (e.g., an insert sequence) to which a unique adaptor is attached, among a population of other target molecules. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular labeling.

[0262] In some embodiments, the covalently closed circular molecule (600) includes any one or any combination of two or more of: a first left universal adapter sequence (120); a second left universal adapter sequence (140); a first left index sequence (160); a first left unique identification sequence (180); a first right universal adapter sequence (130); a second right universal adapter sequence (150); a first right index sequence (170); and / or a first right unique identification sequence (190). In some embodiments, the first left index sequence (160) includes a sample index sequence. In some embodiments, the first right index sequence (170) includes another sample index sequence. The sample index sequence can be used to distinguish target sequences obtained from different sample sources in a multiplexed assay. In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each include a sequence that is used to uniquely identify each target sequence (e.g., an insert sequence) to which a unique adapter is attached, among a population of other target molecules. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) may be used for molecular labeling.

[0263] In some embodiments, in the covalently closed circular molecule (600), the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) comprise: a universal binding sequence for a forward sequencing primer or a reverse sequencing primer; a universal binding sequence for a first surface primer or a second surface primer; a universal binding sequence for a forward amplification primer or a reverse amplification primer; and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the covalently closed circular molecule (600) may further comprise additional left universal adapter sequences.

[0264] In some embodiments, in the covalently closed circular molecule (600), the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) include: a universal binding sequence for a forward sequencing primer or a reverse sequencing primer; a universal binding sequence for a first surface primer or a second surface primer; a universal binding sequence for a forward amplification primer or a reverse amplification primer; a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the covalently closed circular molecule (600) may further include additional right universal adapter sequences.

[0265] In some embodiments, the covalently closed circular molecule (600) further comprises at least one ligating adapter sequence positioned between any of the universal adapter sequences described herein (e.g., see Figure 8). For example, the first left connecting adapter sequence (125) may be located between the first left universal adapter sequence (120) and the first left index sequence (160). The second left connecting adapter sequence (165) may be located between the first left index sequence (160) and the second left universal adapter sequence (140). The third connecting adapter sequence (145) may be located between the second left universal adapter sequence (140) and the target sequence (110). The first right connecting adapter sequence (135) may be located between the first right universal adapter sequence (130) and the first right index sequence (170). The second right connecting adapter sequence (175) may be located between the first right index sequence (170) and the second right universal adapter sequence (150). The third right connecting adapter sequence (155) may be located between the second right universal adapter sequence (150) and the target sequence (110). In some embodiments, the covalently closed circular molecule (600) further comprises at least one and up to ten additional universal adaptor sequences located 5' (upstream) of the first left universal adaptor sequence (120) (e.g., see Figure 8 In some embodiments, the covalently closed circular molecule (600) further comprises at least one and up to ten additional universal adapter sequences located 3' (downstream) of the first right universal adapter sequence (130) (e.g., see Figure 8 ). Any connecting adapter sequence and / or additional universal adapter sequence comprises any sequence and may be 3-60 nucleotides in length. Any connecting adapter sequence and / or additional universal adapter sequence comprises a universal sequence or a unique sequence. Any connecting adapter sequence and / or additional universal adapter sequence comprises a binding sequence for an amplification primer, a sequencing primer or a compaction oligonucleotide. Any connecting adapter sequence and / or additional universal adapter sequence comprises a binding sequence for an immobilized surface primer (e.g., a capture primer). Any connecting adapter sequence and / or additional universal adapter sequence comprises a sample index sequence. Any connecting adapter sequence and / or additional universal adapter sequence comprises a unique identification sequence. Any connecting adapter sequence and / or additional universal adapter sequence, in particular as Figure 8 The ligation adapter sequence (145) shown comprises the Tn5 transposon end sequence 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 211). Any ligation adapter sequence and / or additional universal adapter sequence, particularly as Figure 8The ligation adapter sequence (155) shown comprises the Tn5 transposon end sequence 5'-CTGTCTCTTATACACATCT-3' (SEQ ID NO: 212). The Tn5 transposon end sequence can be introduced into the library molecule (100) by a transposase-mediated reaction under conditions suitable for forming a transposon synaptic complex, the reaction comprising contacting double-stranded input DNA (e.g., genomic DNA) with a Tn-5 type transposase and a double-stranded oligonucleotide comprising the Tn transposon end sequence (SEQ ID NO: 211) linked to a universal adapter sequence or a sample index sequence. In the double-stranded oligonucleotide, the Tn transposon end sequence (SEQ ID NO: 211) can be located 5' or 3' relative to the universal adapter sequence or the sample index sequence.

[0266] In some embodiments, the second splint chain sequence (400) of the covalently closed circular molecule includes at least two sub-regions, the at least two sub-regions including a first sub-region and a second sub-region. In some embodiments, the first sub-region includes a universal binding sequence for a third surface primer, and the second sub-region includes a universal binding sequence for a fourth surface primer, wherein the first sub-region and the second sub-region do not hybridize with the first surface primer and the second surface primer (or at least exhibit very little hybridization). In some embodiments, the second splint chain (400) further includes an optional third sub-region, which includes a sample index sequence with 5-20 bases and / or a unique identification sequence with 2-10 or more bases (e.g., NN). In some embodiments, the second splint chain (400) includes only one sub-region and lacks a second sub-region and a third sub-region, wherein the first sub-region includes an index sequence with 5-20 bases (e.g., a sample index sequence). In some embodiments, the index sequence can be used to distinguish between target sequences obtained from different sample sources in multiple determinations. In some embodiments, the unique identification sequence includes a random sequence. The unique recognition sequence can be designed to exhibit reduced or no hybridization with the first, second, third, and fourth surface primers. Exemplary arrangements of subregions in the second splint strand (400) in the 5' to 3' direction include: 5'-[second subregion]-[first subregion]-3' (e.g., Figure 2 ). Another exemplary arrangement of sub-regions in the second splint chain (400) in the 5' to 3' direction includes: 5'-[third sub-region] - [second sub-region] - [first sub-region] - 3' (e.g., Figure 3 ).

[0267] In some embodiments, the second splint chain sequence (400) of the covalently closed circular library molecule (600) can hybridize with the first splint chain (300). In some embodiments, the first splint chain (300) includes an internal region (310) comprising at least two subregions, including a fourth subregion and a fifth subregion. The fourth subregion hybridizes with the first subregion of the second splint chain (400). The fifth subregion hybridizes with the second subregion of the second splint chain (400). The fourth and fifth subregions do not hybridize with the first and second surface primers (or at least exhibit very little hybridization). In some embodiments, the internal region (310) of the first splint chain further includes an optional sixth subregion that hybridizes with the third subregion of the second splint chain (400). An exemplary arrangement of the subregions of the first splint chain (300) in the 5' to 3' direction includes: 5'–[fourth subregion]–[fifth subregion]–3' (e.g., Figure 2 ). Another exemplary arrangement of the sub-regions of the first splint chain (300) in the 5' to 3' direction includes: 5'-[fourth sub-region]-[fifth sub-region]-[sixth sub-region]-3' (e.g., Figure 3 ).

[0268] In some embodiments, an exemplary covalently closed circular molecule (600) includes: (i) a first left universal adapter sequence (120) having a binding sequence for a first surface primer; (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; (iii) a target sequence (110); (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer; (v) a first right universal adapter sequence (130) having a binding sequence for a second surface primer; and (vi) a second splint chain sequence (400), wherein the covalently closed circular molecule (600) is optionally hybridized to the first splint chain (300).

[0269] In the exemplary covalently closed circular molecule (600), the second splint chain region (400) includes at least two subregions, the at least two subregions including the first and second subregions. The first subregion includes a universal binding sequence for a third surface primer, and the second subregion includes a universal binding sequence for a fourth surface primer, wherein the first subregion and the second subregion do not hybridize with the first surface primer and the second surface primer (or at least exhibit very little hybridization). In some embodiments, the second splint chain (400) further includes an optional third subregion, which includes a sample index sequence with 5-20 bases and / or a unique identification sequence with 2-10 or more bases (e.g., NN). In some embodiments, the second splint chain (400) includes only one subregion and lacks the second and third subregions, wherein the first subregion includes a sample index sequence with 5-20 bases. In some embodiments, the sample index sequence can be used to distinguish target sequences obtained from different sample sources in multiple determinations. In some embodiments, the unique identification sequence includes a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization with the first, second, third, and fourth surface primers. An exemplary arrangement of the sub-regions in the second splint chain (400) in the 5' to 3' direction includes: 5'-[second sub-region] - [first sub-region] - 3'. Another exemplary arrangement of the sub-regions in the second splint chain (400) in the 5' to 3' direction includes: 5'-[third sub-region] - [second sub-region] - [first sub-region] - 3'.

[0270] The present disclosure provides a plurality of covalently closed circular library molecules as described herein. In some embodiments of the plurality of covalently closed circular molecules (600), the target sequence (110) of each covalently closed circular molecule (600) in the plurality of covalently closed circular molecules comprises the same target sequence or a different target sequence.

[0271] Covalently closed circular molecules formed using double-stranded adaptors with truncated long splint strands

[0272] The present disclosure provides a covalently closed circular library molecule (600) comprising: a target sequence (110), at least a first left universal adapter sequence (120), at least a first right universal adapter sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9 In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adapter sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adapter sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adapter sequences.

[0273] In some embodiments, the covalently closed circular library molecule (600) is hybridized to the first splint strand. In some embodiments, the first splint strand (300) comprises a strand having a first region (320) having a truncated sequence at the 5' end (e.g., Figure 11B , showing exemplary truncations compared to SEQ ID NO: 199, such as Figure 11A As shown). In some embodiments, the 5' end of the first region can have a truncation of any length, for example, a truncation of 1-10 nucleotides. In some embodiments, the truncated first splint strand (300) comprises a second region (330; for example, SEQ ID NO: 5), a fourth subregion (for example, SEQ ID NO: 6), and a fifth subregion (for example, SEQ ID NO: 7) that is not truncated and does not carry any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the truncated first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint strand (400) that is part of a covalently closed circular library molecule (600) (for example, Figure 9 In some embodiments, the truncated first splint strand (300) comprises a truncated first region (320) that hybridizes to a sequence (e.g., 120) that is part of the covalently closed circular library molecule (600) and a second region (330) that hybridizes to a sequence (e.g., 130) that is part of the covalently closed circular library molecule (600).

[0274] Covalently closed circular molecules formed using double-stranded adapters with long splint chains having mismatched sequences

[0275] The present disclosure provides a covalently closed circular library molecule (600) comprising: a target sequence (110), at least a first left universal adapter sequence (120), at least a first right universal adapter sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9 In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adapter sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adapter sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adapter sequences.

[0276] In some embodiments, the covalently closed circular library molecule (600) is hybridized to the first splint strand. In some embodiments, the first splint strand (300) comprises a mismatch strand having a first region (320) having a mismatch sequence (e.g., Figure 11C). In some embodiments, the mismatch sequence can be of any length (e.g., 2-20 bases) and comprises any sequence that is not completely complementary to the left universal adapter sequence (120) of the library molecule (100). Some embodiments of the mismatch sequence in the first region (320) are Figure 11C In some embodiments, the mismatched first splint strand (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7) that do not carry any sequence variants, such as, for example, insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint strand (400) that is part of a covalently closed circular library molecule (600) (e.g., Figure 9 In some embodiments, the mismatched first splint strand (300) comprises a mismatched first region (320) that hybridizes to a sequence (e.g., 120) that is part of the covalently closed circular library molecule (600) and a second region (330) that hybridizes to a sequence (e.g., 130) that is part of the covalently closed circular library molecule (600). In some embodiments, the mismatched first region (320) can hybridize to the first left universal adapter sequence (120) of the covalently closed circular library molecule (600) to form a double-stranded portion having a bubble at the position of the mismatched sequence in the first region (320).

[0277] Covalently closed circular molecules formed using double-stranded adaptors with long splint chains having abasic sites

[0278] The present disclosure provides a covalently closed circular library molecule (600) comprising: a target sequence (110), at least a first left universal adapter sequence (120), at least a first right universal adapter sequence (130), and a second splint strand sequence (400). An exemplary covalently closed circular library molecule is shown in Figure 9 In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adapter sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adapter sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adapter sequences.

[0279] In some embodiments, the covalently closed circular library molecule (600) hybridizes to the first splint strand. In some embodiments, the first splint strand (300) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site in the fourth subregion and / or at least one abasic site in the fifth subregion (e.g., Figure 11D, with the abasic sites shown as solid black bars). In some embodiments, the abasic sites each comprise a 1',2'-dideoxyribose sugar (e.g., dSpacer from Integrated DNA Technologies (IDT)). In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO:4), and a second region ((330); e.g., SEQ ID NO:5) that does not carry any abasic sites and / or any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first subregion and the second subregion of the second splint strand (400) that is part of a covalently closed circular library molecule (600) (e.g., Figure 9 In some embodiments, the abasic first splint strand (300) comprises a first region (320) that hybridizes to a sequence (e.g., 120) that is part of the covalently closed circular library molecule (600) and a second region (330) that hybridizes to a sequence (e.g., 130) that is part of the covalently closed circular library molecule (600).

[0280] Covalently closed circular molecules formed using double-stranded adapters with long splint chains of uracil

[0281] The present disclosure provides a covalently closed circular library molecule (600) comprising: a sequence of interest (110), at least a first left universal adapter sequence (120), at least a first right universal adapter sequence (130), and a second splint strand sequence (400). Exemplary covalently closed circular library molecules are as follows: Figure 9 In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adapter sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adapter sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adapter sequences.

[0282] In some embodiments, the covalently closed circular library molecule (600) is hybridized to the first splint strand. In some embodiments, the first splint strand (300) comprises at least one uracil. In some embodiments, the first splint strand (300) comprises at least one uracil in any one or any combination of the regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base can be substituted with uracil. Examples of first splint strands comprising uracil are shown in Figure 11D(bottom schematic). The skilled artisan will recognize that many other sequences of the first splint strand (300) comprising one or more uracils are possible. In some embodiments, the uracil-containing first splint strand comprises a fourth subregion and a fifth subregion that can hybridize to the first subregion and the second subregion of the second splint strand (400) that is part of a covalently closed circular library molecule (600) (e.g., Figure 9 In some embodiments, the uracil-containing first splint strand (300) comprises a first region (320) that hybridizes to a sequence (e.g., 120) that is part of the covalently closed circular library molecule (600) and a second region (330) that hybridizes to a sequence (e.g., 130) that is part of the covalently closed circular library molecule (600).

[0283] Covalently closed circular molecules formed using double-stranded adaptors with short splint chains having random and index sequences

[0284] The present disclosure provides a covalently closed circular library molecule (600) comprising: a sequence of interest (110), at least a first left universal adapter sequence (120), at least a first right universal adapter sequence (130), and a second splint strand sequence (400). Exemplary covalently closed circular library molecules are as follows: Figure 9 In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adapter sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adapter sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adapter sequences.

[0285] In some embodiments, the covalently closed circular library molecule (600) comprises a second splint strand sequence (400) covalently linked to a first left universal adapter sequence (120) and a first right universal adapter sequence (130) (e.g., Figure 7A and Figure 9 In some embodiments, the covalently closed circular library molecule (600) hybridizes to the first splint strand (300). In some embodiments, the second splint strand (400) comprises a random sequence (e.g., Figure 7A 、 Figure 12A and Figure 12B In some embodiments, the random sequence may replace a portion of the first sub-region of the second splint chain (e.g., Figure 7A 、 Figure 13A and Figure 13B In some embodiments, the second sub-region of the second splint chain (400) does not have an inserted random sequence. In some embodiments, a portion of the second sub-region of the second splint chain (400) is not replaced by a random sequence.

[0286] In some embodiments, the random sequence can be of any length, for example, 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 12A and Figure 13A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 12B and Figure 13B In some embodiments, the random sequence may be inserted at any position in the first sub-region of the second splint chain (400).

[0287] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0288] In some embodiments, random sequence provides nucleotide diversity and color balance for sequencing reaction. In some embodiments, random sequence provides high nucleotide diversity, and the sequence includes all four nucleotides (for example, A, G, C, T and / or U) of roughly equal proportion. This will be represented in each cycle of sequencing operation. In some embodiments, the sequence of any part of left index (160), right index (170) and / or insertion zone (110) does not provide enough nucleotide diversity to realize polymerase clone mapping and / or template alignment. In some embodiments, compared with any part of left index (160), right index (170) and / or insertion zone (110), random sequence provides higher level of nucleotide diversity. In some embodiments, covalently closed circular library molecule can undergo rolling circle amplification reaction to produce concatemers fixed to support. Concatemers include tandem repeats of circular library molecules, which include any insertion sequence and adapter sequence (for example, random sequence) present in the original cyclized nucleic acid template molecule. In some embodiments, random index sequence can be a sequence before sequencing the insertion zone, wherein random index is located in the first sub-region of the second splint chain (400).

[0289] In some embodiments, the random sequence can be sequenced before sequencing the insert region. In some embodiments, sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides sufficient nucleotide diversity and color balance.

[0290] In some embodiments, the predetermined sequence is inserted into the sequence of the fourth sub-region of the first splint chain (300) (eg, Figure 7A The length of the inserted predetermined sequence may be the same as the length of the random sequence inserted into the first sub-region of the second splint chain (e.g., Figure 12A and Figure 12B ). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 12A and Figure 12B shown.

[0291] In some embodiments, a portion of the fourth sub-region of the first cleat chain (300) is replaced by a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth sub-region is the same as the length of the random sequence replacing the portion of the first sub-region of the second cleat chain (e.g., Figure 7A 、 Figure 13A and Figure 13B ). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 13A and Figure 13B shown.

[0292] In some embodiments, the second splint strand (400) includes a first subregion and a second subregion that can hybridize to the fourth subregion and the fifth subregion of the first splint strand (300) as part of a covalently closed circular library molecule (600) (e.g., Figure 7A 、 Figure 12A 、 Figure 12B 、 Figure 13A and Figure 13B ), the covalently closed circular library molecules can include bubbles at the positions where random sequences are inserted or replaced.

[0293] Covalently closed circles formed using double-stranded adapters with short splint strands having attached random and index sequences Molecules

[0294] The present disclosure provides a covalently closed circular library molecule (600) comprising: a sequence of interest (110), at least a first left universal adapter sequence (120), at least a first right universal adapter sequence (130), and a second splint strand sequence (400). Exemplary covalently closed circular library molecules are as follows: Figure 9 In some embodiments, the covalently closed circular library molecule (600) further comprises a second left universal adapter sequence (140). In some embodiments, the covalently closed circular molecule (600) further comprises a second right universal adapter sequence (150). In some embodiments, the covalently closed circular molecule (600) further comprises additional left and / or right universal adapter sequences.

[0295] In some embodiments, the covalently closed circular library molecule (600) comprises a second splint strand sequence (400) covalently linked to a first left universal adapter sequence (120) and a first right universal adapter sequence (130) (e.g., Figure 7B and Figure 9 In some embodiments, the covalently closed circular library molecule (600) hybridizes to the first splint strand (300). In some embodiments, the second splint strand (400) comprises a random sequence (e.g., Figure 7B 、 Figure 14A and Figure 14B In some embodiments, the random sequence further comprises an index sequence. In some embodiments, the second sub-region of the second splint chain (400) does not have an additional random sequence.

[0296] In some embodiments, the additional random sequence can be of any length, such as 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 14A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 14B NNNN in the .

[0297] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0298] In some embodiments, random sequences provide nucleotide diversity and color balance for sequencing reactions. In some embodiments, random sequences provide high nucleotide diversity, comprising approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U), which will be represented in each cycle of a sequencing run.

[0299] In some embodiments, the random sequence can be sequenced before the insertion zone is sequenced. In some embodiments, the sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides enough nucleotide diversity and color balance. In some embodiments, the sequence of any part of the left index (160), the right index (170) and / or the insertion zone (110) does not provide enough nucleotide diversity to achieve polymerase clone mapping and / or template alignment. In some embodiments, the random sequence provides a higher level of nucleotide diversity compared to any part of the left index (160), the right index (170) and / or the insertion zone (110). In some embodiments, the covalently closed circular library molecule can undergo a rolling circle amplification reaction to produce a concatemer fixed to a support. The concatemer comprises a tandem repeat sequence of the circular library molecule, which includes any insertion sequence and adapter sequence (e.g., random sequence) present in the original cyclized nucleic acid template molecule. In some embodiments, the random index sequence can be a sequence before the insertion zone is sequenced, wherein the random index is located at the 3' end of the first sub-region of the second splint chain (400).

[0300] In some embodiments, index sequences can be used to distinguish polynucleotides (eg, insert sequences) from different sample sources in a multiplex assay.

[0301] In some embodiments, a predetermined sequence is appended to the 5' end of the fourth sub-region of the first splint strand (300) (e.g., Figure 7B In some embodiments, the length of the additional predetermined sequence may be the same as the length of the random sequence attached to the 3' end of the first sub-region of the second splint chain (e.g., Figure 7B 、 Figure 14A and Figure 14B In some embodiments, the length of the additional predetermined sequence may be the same as the length of the random sequence and index sequence attached to the 3' end of the first subregion of the second splint chain (e.g., Figure 7B 、 Figure 14A and Figure 14B ). The additional predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 14A and Figure 14B shown.

[0302] In some embodiments, the second splint strand (400) includes a first subregion and a second subregion that can hybridize to the fourth subregion and the fifth subregion of the first splint strand (300) as part of a covalently closed circular library molecule (600) (e.g., Figure 7B 、 Figure 14A and Figure 14B ), the covalently closed circular library molecule may include a bubble at the location of the appended random sequence.

[0303] Sequence of short splint chains

[0304] In some embodiments of the covalently closed circular molecules (600) described herein, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence

[0305] 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises a first sub-region and a second sub-region comprising the sequence

[0306] 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or non-phosphorylated.

[0307] Sequence of long splint chains

[0308] In some embodiments of the covalently closed circular molecule (600) described herein, the first region of the first splint strand (320) comprises a first universal adaptor sequence comprising a universal binding sequence (or its complement) for a first surface primer, wherein the first region (320) comprises the sequence

[0309] 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize with a P5 surface primer or the complement of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand includes a second universal adapter sequence comprising a universal binding sequence (or its complement) for a second surface primer, wherein the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize with a P7 surface primer or the complement of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fourth subregion having the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198). In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) having fourth and fifth subregions, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAG GCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). Figure 11A In some embodiments, the 5' end of the first splint chain (300) can be phosphorylated or non-phosphorylated. In some embodiments, the first subregion of the second splint chain (400) can hybridize with the fourth subregion of the first splint chain (300). In some embodiments, the second subregion of the second splint chain (400) can hybridize with the fifth subregion of the first splint chain (300).

[0310] Sequence of covalently closed circular molecules

[0311] In some embodiments of the covalently closed circular molecules (600) described herein, the first region of the first splint chain (320) comprises a sequence that can bind to the first left universal adapter sequence (120) of the library molecule, wherein the first region of the first splint chain (320) comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO:215) or its complementary sequence.

[0312] In some embodiments of the covalently closed circular molecules (600) described herein, the second region of the first splint strand (330) comprises a sequence that can bind to the first right universal adapter sequence (130) of the library molecule, wherein the second region of the first splint strand (330) comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO:216) or its complementary sequence.

[0313] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecules include a first left universal adaptor sequence (120) comprising the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203).

[0314] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule includes a first left universal adapter sequence (120) that binds to a first region of the first splint strand (320), wherein the first left universal adapter sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 213) or its complement.

[0315] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a second left universal adapter sequence comprising a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence comprises the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 204).

[0316] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a second left universal adapter sequence comprising a sequence for binding to a sequencing primer (140), wherein the second left universal binding sequence comprises the sequence 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 207).

[0317] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a second left universal adapter sequence comprising a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence comprises the sequence 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO: 208).

[0318] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adapter sequence comprising a sequence for binding to a sequencing primer (150), wherein the first right universal adapter sequence comprises the sequence 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO:205).

[0319] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adapter sequence comprising a sequence for binding to a sequencing primer (150), wherein the first right universal adapter sequence comprises the sequence 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO: 209).

[0320] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule comprises a first right universal adapter sequence comprising a sequence for binding a sequencing primer (150), wherein the first right universal adapter sequence comprises the sequence 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO:210).

[0321] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecules include a first right universal adaptor sequence (130) comprising the sequence 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 206).

[0322] In some embodiments of the covalently closed circular molecules (600) described herein, the library molecule includes a first right universal adapter sequence (130) that binds to a first region of the first splint strand (330), wherein the first right universal adapter sequence (130) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO:214) or its complement.

[0323] The present disclosure provides a reaction mixture comprising a plurality of any covalently closed circular molecules (600) described herein and at least one exonuclease. In some embodiments, the exonuclease comprises any one of Exonuclease I, Thermostable Exonuclease I, and / or T7 Exonuclease, or any combination of two or more thereof.

[0324] Kits containing double-stranded splint adapters

[0325] The present disclosure provides a kit for introducing one or more new adapter sequences into linear nucleic acid library molecules. In some embodiments, the kit can be used to circularize single-stranded nucleic acid library molecules having a target sequence (110) flanked on one side by at least a first left universal adapter sequence (120) and on the other side by at least a first right universal adapter sequence (130). In some embodiments, the circularized library molecules can be converted into covalently closed circular molecules, which can undergo rolling circle amplification (RCA) reactions to generate nucleic acid concatemers. The concatemers can be fixed to a support for large-scale parallel sequencing.

[0326] The present disclosure provides a kit comprising a nucleic acid double-stranded splint adapter (200), comprising: (i) a first splint chain (long splint chain (300)) hybridized with (ii) a second splint chain (short splint chain (400)). In some embodiments, the first splint chain comprises a first region (320), an internal region (310) and a second region (330). The internal region (310) of the first splint chain hybridizes with the second splint chain (400) to form a double-stranded splint adapter (200) having a double-stranded region and two flanking single-stranded regions. The second splint chain (400) includes a new adapter sequence that can be introduced into a linear nucleic acid library molecule. Exemplary double-stranded splint adapters are shown in Figure 1-8 The kit may include a container containing a first splint strand (300) hybridized with a second splint strand (400). The kit may include a first container containing the first splint strand (300) and a second container containing the second splint strand (400).

[0327] In some embodiments of the kits of the present disclosure, the second splint chain (400) includes at least two sub-regions, the at least two sub-regions including a first sub-region and a second sub-region (e.g., see Figure 2 and Figure 3). The first subregion comprises a universal binding sequence for a third surface primer, and the second subregion comprises a universal binding sequence for a fourth surface primer, wherein the first subregion and the second subregion do not hybridize (or at least exhibit little hybridization) with the first surface primer and the second surface primer. In some embodiments, the second splint strand (400) further comprises an optional third subregion comprising a sample index sequence having 5-20 bases and / or a unique identification sequence (e.g., NN) having 2-10 or more bases (e.g., see Figure 3 ). In some embodiments, the second splint chain (400) includes only one subregion and lacks a second subregion and a third subregion, wherein the first subregion includes an index sequence (e.g., a sample index) having 5-20 bases. In some embodiments, the sample index sequence can be used to distinguish target sequences obtained from different sample sources in multiple assays. In some embodiments, the unique identification sequence comprises a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization with the first, second, third, and fourth surface primers. An exemplary arrangement of subregions in the second splint chain (400) in the 5' to 3' direction includes: 5'-[second subregion]–[first subregion]–3'. Another exemplary arrangement of subregions in the second splint chain (400) in the 5' to 3' direction includes: 5'-[third subregion]–[second subregion]–[first subregion]–3'. Exemplary first splint chain (300) and second splint chain (400) are shown in Figure 2 and 3 In some embodiments, the second splint chain (400) may be 20-100 nucleotides in length, or 30-80 nucleotides in length, or 40-60 nucleotides in length. In some embodiments, the second splint chain (400) comprises one or more phosphorothioate bonds at the 5' and / or 3' ends to confer resistance to nucleic acid exonucleases. In some embodiments, the second splint chain (400) comprises one or more phosphorothioate bonds at an internal position to confer resistance to nucleic acid endonucleases. In some embodiments, the second splint chain (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' ends or at an internal position. In some embodiments, the 5' end of the second splint chain (400) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the second splint chain (400) comprises a terminal 3'OH group or a terminal 3' blocking group.

[0328] In some embodiments of the kits of the present disclosure, the first splint strand (300) comprises a first region (320), a second region (330) and an internal region (310). The first region (320) comprises a first universal adapter sequence that can hybridize with a first universal binding sequence at one end of a linear nucleic acid library molecule. The second region (330) comprises a second universal adapter sequence that can hybridize with a second universal binding sequence at the other end of a linear nucleic acid library molecule. In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence that comprises a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compacted oligonucleotide, or a combination thereof. In some embodiments, the second region (330) of the first splint chain includes a second universal adapter sequence comprising a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compaction oligonucleotide, or a combination thereof. In some embodiments, the first splint chain (300) may be 50-150 nucleotides in length, or 60-100 nucleotides in length, or 70-90 nucleotides in length. In some embodiments, the first splint chain (300) comprises one or more thiophosphate bonds at the 5' and / or 3' end to confer nuclease resistance. In some embodiments, the first splint chain (300) comprises one or more thiophosphate bonds at an internal position to confer nuclease resistance. In some embodiments, the first splint chain (300) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the first splint chain (300) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the first splint chain (300) comprises a terminal 3' OH group or a terminal 3' blocking group.

[0329] In some embodiments of the kits of the present disclosure, the first cleat chain (300) includes an inner region (310) comprising at least two sub-regions, the at least two sub-regions comprising a fourth sub-region and a fifth sub-region (e.g., Figure 2 and Figure 3). The fourth subregion hybridizes with the first subregion of the second splint chain (400). The fifth subregion hybridizes with the second subregion of the second splint chain (400). The fourth and fifth subregions do not hybridize with the first and second surface primers (or at least exhibit very little hybridization). In some embodiments, the interior region (310) of the first splint chain further comprises an optional sixth subregion that hybridizes with the third subregion of the second splint chain (400). An exemplary arrangement of the subregions of the first splint chain (300) in the 5' to 3' direction comprises: 5'–[fourth subregion]–[fifth subregion]–3'. Another exemplary arrangement of the subregions of the first splint chain (300) in the 5' to 3' direction comprises: 5-[fourth subregion]–[fifth subregion]–[sixth subregion]–3'. An exemplary first splint chain (300) is shown in Figure 2 and Figure 3 middle.

[0330] In some embodiments of the kits described herein, the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or non-phosphorylated.

[0331] In some embodiments of the kits described herein, the first region of the first splint strand (320) comprises a first universal adapter sequence comprising a universal binding sequence (or its complement) for a first surface primer, wherein the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or the complement of a P5 surface primer. For example, the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence for a second surface primer (or its complement), wherein the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize with a P7 surface primer or a complement of a P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fourth subregion having the sequence

[0332] 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198). In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) comprising a fourth and a fifth subregion, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAG GCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). See Figure 11AIn some embodiments, the 5' end of the first splint chain (300) can be phosphorylated or non-phosphorylated. In some embodiments, the first subregion of the second splint chain (400) can hybridize with the fourth subregion of the first splint chain (300). In some embodiments, the second subregion of the second splint chain (400) can hybridize with the fifth subregion of the first splint chain (300).

[0333] In some embodiments of the kits described herein, the first region (320) of the first splint strand comprises a sequence that can bind to the first left universal adapter sequence (120) of the library molecule, wherein the first region (320) of the first splint strand comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 215) or its complementary sequence.

[0334] In some embodiments of the kits described herein, the second region of the first splint strand (330) comprises a sequence that can bind to the first right universal adapter sequence (130) of the library molecule, wherein the second region of the first splint strand (330) comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO:216) or its complementary sequence.

[0335] In some embodiments, the kit comprises an adapter having a first left universal adapter sequence (120) that binds to a first region (320) of a first splint strand, the adapter being used to prepare a plurality of library molecules, wherein the library molecules comprise the sequence 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.

[0336] In some embodiments of the kits described herein, the library molecule comprises a first left universal adapter sequence (120) that binds to the first region (320) of the first splint strand, wherein the first left universal adapter sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 213) or its complement.

[0337] In some embodiments, the kit comprises an adaptor having a second left universal adaptor sequence for a sequencing primer (140) for preparing a plurality of library molecules, wherein the library molecules comprise the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 204). In some embodiments, the adaptor having a second left universal adaptor sequence for a sequencing primer (140) further comprises a first left index sequence (160). The adaptor can be a single-stranded adaptor (e.g., a PCR primer), a double-stranded adaptor, a bubble adaptor, or a Y-shaped adaptor.

[0338] In some embodiments, the kit comprises an adaptor having a second left universal adaptor sequence comprising a sequence for binding to a sequencing primer (140), the adaptor being used to prepare a plurality of library molecules, wherein the library molecules comprise the sequence

[0339] 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 207). In some embodiments, the adapter having the second left universal adapter sequence for the sequencing primer (140) further comprises a first left index sequence (160). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.

[0340] In some embodiments, the kit comprises an adaptor having a second left universal adaptor sequence comprising a sequence for binding to a sequencing primer (140), the adaptor being used to prepare a plurality of library molecules, wherein the library molecules comprise the sequence

[0341] 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO: 208). In some embodiments, the adapter having a second left universal adapter sequence for the sequencing primer (140) further comprises a first left index sequence (160). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.

[0342] In some embodiments, the kit comprises an adaptor having a second right universal adaptor sequence comprising a sequence for binding to a sequencing primer (150), the adaptor being used to prepare a plurality of library molecules, wherein the library molecules comprise the sequence

[0343] 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO: 205). In some embodiments, the adapter having the second right universal adapter sequence for the sequencing primer (150) further comprises a first right index sequence (170). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.

[0344] In some embodiments, the kit comprises an adaptor having a second right universal adaptor sequence comprising a sequence for binding to a sequencing primer (150), the adaptor being used to prepare a plurality of library molecules, wherein the library molecules comprise the sequence

[0345] 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO: 209). In some embodiments, the adapter having a second right universal adapter sequence for the sequencing primer (150) further comprises a first right index sequence (170). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.

[0346] In some embodiments, the kit comprises an adaptor having a second right universal adaptor sequence comprising a sequence for binding to a sequencing primer (150), the adaptor being used to prepare a plurality of library molecules, wherein the library molecules comprise the sequence

[0347] 5'-ATGTCGGAAGGTGTGCAGGCTACCGCTTGTCAACT-3' (SEQ ID NO: 210). In some embodiments, the adapter having the second right universal adapter sequence for the sequencing primer (150) further comprises a first right index sequence (170). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.

[0348] In some embodiments, the kit comprises an adapter having a first right universal adapter sequence (130) that binds to a first region of a first splint strand (330), the adapter being used to prepare a plurality of library molecules, wherein the library molecules comprise the sequence 5'-TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 206). The adapter can be a single-stranded adapter (e.g., a PCR primer), a double-stranded adapter, a bubble adapter, or a Y-shaped adapter.

[0349] In some embodiments of the kits described herein, the library molecule comprises a first right universal adapter sequence (130) that binds to a first region of the first splint strand (330), wherein the first right universal adapter sequence (130) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 214) or its complement.

[0350] In some embodiments, the test kit comprises a plurality of polynucleotides comprising a first left index sequence (160) and / or a plurality of first right index sequences (170). In some embodiments, the test kit may comprise a separate container for accommodating polynucleotides comprising each first left index (160) or each first right index (170) sequence. In some embodiments, the test kit may comprise a separate container for accommodating a pair of polynucleotides comprising each first left index (160) and each first right index (170) sequence. In some embodiments, the test kit contains polynucleotides in a multi-well plate (e.g., a 96-well plate) comprising a first left index (160) and / or a plurality of first right index sequences (170). A list of exemplary first left index sequences (160) and first right index sequences (170) is provided in Table 1 of Figure 33. The first left index sequence (160) may comprise a random sequence (e.g., NNN) or lack a random sequence. The first right index sequence (170) may comprise a random or sequence (e.g., NNN) or lack a random sequence.

[0351] In some embodiments, the kit comprises a nucleic acid double-stranded splint adapter (200) and T4 polynucleotide kinase. In some embodiments, the kit comprises a ligase, wherein the ligase comprises T7 DNA ligase, T3 ligase, T4 ligase, or Taq ligase. In some embodiments, the kit comprises at least one endonuclease, which comprises any one or any combination of two or more of exonuclease I, thermolabile exonuclease I, and / or T7 exonuclease.

[0352] In some embodiments, the kit comprises at least one buffer for hybridizing a plurality of double-stranded splint adapters (200) and a plurality of nucleic acid library molecules (100). In some embodiments, the kit comprises a buffer for performing multiple enzymatic reactions in a single reaction vessel, including (i) phosphorylating the 5' end of the first and / or second splint strands (e.g., (300) and / or (400)), (ii) ligating a gap in the library-splint complex (500), and / or (iii) exonuclease digestion of the first splint strand (300) from the covalently closed circular molecule (600). Alternatively, the kit comprises two or more separate buffers, wherein a first buffer can be used to perform the phosphorylation reaction, a second buffer can be used to perform the ligation reaction, and a third buffer can be used to perform the exonuclease digestion reaction.

[0353] In some embodiments, the kit comprises one or more containers containing any double-stranded splint adapter (200) described herein or any of the first splint strand (300) and the second splint strand (400) described herein. The kit may further comprise one or more containers containing T4 polynucleotide kinase, at least one ligase, and / or at least one exonuclease. The kit may comprise any of these components in any combination and may be contained in a single container, or may be contained in separate containers, or any combination thereof.

[0354] The kit can include instructions for using the kit to perform a reaction to introduce one or more new adaptor sequences into a linear nucleic acid library molecule.

[0355] The kit may include a polynucleotide encoding one or more exemplary sequences of interest for use as a positive control. Methods for forming multiple library-splint complexes

[0356] The present disclosure provides a method for forming a plurality of library-splint complexes (500), comprising: (a) providing a plurality of double-stranded splint adaptors (200), wherein each double-stranded splint adaptor (200) in the plurality of double-stranded splint adaptors comprises a first splint strand (300) hybridized to a second splint strand (400), wherein the double-stranded splint adaptor comprises a double-stranded region and two flanking single-stranded regions, wherein the first splint strand comprises a first region (320), an internal region (310), and a second region (330), and wherein the internal region (310) of the first splint strand hybridizes to the second splint strand (400). Exemplary double-stranded splint adaptors (200) are shown in Figures 1 to 8 middle.

[0357] In some embodiments, a method for forming a plurality of library-splint complexes (500) comprises step (b): hybridizing a plurality of double-stranded splint adaptors to a plurality of single-stranded nucleic acid library molecules (100), wherein each library molecule comprises a sequence of interest (110) flanked on one side by at least a first left universal adaptor sequence (120) and on another side by at least a first right universal adaptor sequence (130) (e.g., Figures 1 to 8 The hybridization is performed under conditions suitable for hybridizing the first region (320) of the first splint strand to at least the first left universal adaptor sequence (120) of the library molecule, and hybridizing the second region (330) of the first splint strand to at least the first right universal sequence (130) of the library molecule, thereby circularizing the plurality of library molecules to form a plurality of library-splint complexes (500).

[0358] In some embodiments of the method for forming multiple library-splint complexes (500), the first region (320) of the first splint strand comprises a first universal adapter sequence that can hybridize to the first universal binding sequence at one end of the linear nucleic acid library molecule. In some embodiments, the first region (320) of the first splint strand comprises a first universal adapter sequence that comprises a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compacted oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the first splint strand (300) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the first splint strand (300) comprises a terminal 3' OH group or a terminal 3' blocking group.

[0359] In some embodiments of the method for forming multiple library-splint complexes (500), the second region (330) of the first splint strand comprises a second universal adapter sequence that can hybridize to the second universal binding sequence at the other end of the linear nucleic acid library molecule. In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence that comprises a universal binding sequence for a forward sequencing primer or a reverse sequencing primer, a universal binding sequence for a first surface primer or a second surface primer, a universal binding sequence for a forward amplification primer or a reverse amplification primer, a universal binding sequence for a compacted oligonucleotide, or a combination thereof. In some embodiments, the 5' end of the second splint strand (400) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the second splint strand (400) comprises a terminal 3' OH group or a terminal 3' blocking group.

[0360] In some embodiments of the method for forming a plurality of library-splint complexes (500), a first region (320) of a first splint strand hybridizes to at least a first left universal adaptor sequence (120) of a library molecule, and a second region (330) of the first splint strand hybridizes to at least a first right universal sequence (130) of a library molecule, thereby circularizing the library molecule to produce the library-splint complex (500). The library-splint complex (500) comprises a first gap (e.g., Figures 1 to 8 The library-splint complex (500) further comprises a second gap (e.g., Figures 1 to 8 ). In some embodiments, the first gap and the second gap are enzymatically ligatable.

[0361] In some embodiments of the method for forming a plurality of library-splint complexes (500), the first region of the first splint strand (320) can hybridize to the sense strand or the antisense strand of a double-stranded nucleic acid library molecule. In the library-splint complex (500), the second region (330) of the first splint strand can hybridize to the sense strand or the antisense strand of a double-stranded nucleic acid library molecule. The double-stranded nucleic acid library molecule can be denatured to produce single-stranded sense and antisense library strands.

[0362] In some embodiments of the method for forming a plurality of library-splint complexes (500), the second splint strand (400) does not hybridize to the sequence of interest (110), and the internal region (310) of the first splint strand does not hybridize to the sequence of interest (110).

[0363] In some embodiments of the method for forming multiple library-splint complexes (500), the first region of the first splint strand (320) does not hybridize to the target sequence (110), and the second region (330) of the first splint strand does not hybridize to the target sequence (110).

[0364] In some embodiments of the method for forming a plurality of library-splint complexes (500), the 5' end of the single-stranded library molecule (100) is phosphorylated or lacks a phosphate group. In some embodiments, the 3' end of the single-stranded library molecule comprises a terminal 3' OH group or a terminal 3' blocking group.

[0365] In some embodiments of the method for forming a plurality of library-splint complexes (500), each nucleic acid library molecule (100) comprises a second left universal adapter sequence (140). In some embodiments, each nucleic acid library molecule (100) comprises a second right universal adapter sequence (150). In some embodiments, the nucleic acid library molecule (100) comprises additional left universal adapter sequences and / or right universal adapter sequences.

[0366] In some embodiments of the method for forming multiple library-splint complexes (500), the nucleic acid library molecule (100) comprises a first left index sequence (160). In some embodiments, the nucleic acid library molecule (100) comprises a first right index sequence (170). In some embodiments, the first left index sequence (160) comprises a sample index sequence. In some embodiments, the first right index sequence (170) comprises another sample index sequence. In some embodiments, the sequence of the first left index sequence is different from the sequence of the first right index sequence. The sample index sequence can be used to distinguish target sequences obtained from different sample sources in multiple assays. A list of exemplary first left index sequences (160) and first right index sequences (170) is provided in Table 1 of Figure 33. The first left index sequence (160) can include a random sequence (e.g., NNN) or lack a random sequence. The first right index sequence (170) can include a random sequence (e.g., NNN) or lack a random sequence.

[0367] A multiplex workflow is initiated by preparing a library with a sample index using one or two index sequences (e.g., a left index sequence and / or a right index sequence). A separate library with a sample index can be prepared using a first left index sequence (160) and / or a first right index sequence (170) using input nucleic acids isolated from different sources. Libraries with sample indexes can be pooled together to generate a multiple library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insertion region together with the first left index sequence (160) and / or the first right index sequence (170) can be used to identify the source of the input nucleic acid. In some embodiments, any number of libraries with sample indexes can be pooled together, for example, 2-10, or 10-50, or 50-100, or 100-200, or more than 200 libraries with sample indexes can be pooled. Exemplary nucleic acid sources include naturally occurring, recombinant, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or entire organisms. Exemplary nucleic acid sources include fresh, frozen, fresh frozen or archived sources (e.g., formalin fixed paraffin embedded; FFPE). The skilled artisan will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.

[0368] In some embodiments of the method for forming a plurality of library-splint complexes (500), the nucleic acid library molecule (100) comprises a first left unique identification sequence (180). In some embodiments, the nucleic acid library molecule (100) comprises a first right unique identification sequence (190). In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence that uniquely identifies each target sequence (e.g., an insert sequence) to which a unique adaptor is attached, among a population of other target molecules. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular labeling.

[0369] In some embodiments of the method for forming multiple library-splint complexes (500), the nucleic acid library molecule (100) comprises any one or any combination of two or more of the following: a first left universal adapter sequence (120); a second left universal adapter sequence (140); a first left index sequence (160); a first left unique identification sequence (180); a first right universal adapter sequence (130); a second right universal adapter sequence (150); a first right index sequence (170); and / or a first right unique identification sequence (190).

[0370] In some embodiments of the method for forming a plurality of library-splint complexes (500), the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) comprise a universal binding sequence for a forward sequencing primer or a reverse sequencing primer; a universal binding sequence for a first surface primer or a second surface primer; a universal binding sequence for a forward amplification primer or a reverse amplification primer; and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises an additional left universal adapter sequence.

[0371] In some embodiments of the method for forming a plurality of library-splint complexes (500), the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) comprise a universal binding sequence for a forward sequencing primer or a reverse sequencing primer; a universal binding sequence for a first surface primer or a second surface primer; a universal binding sequence for a forward amplification primer or a reverse amplification primer; and / or a universal binding sequence for a compaction oligonucleotide. In some embodiments, the nucleic acid library molecule (100) comprises an additional right universal adapter sequence.

[0372] In some embodiments of the method for forming a plurality of library-splint complexes (500), the nucleic acid library molecule (100) comprises at least one ligation adapter sequence positioned between any of the universal adapter sequences described herein (e.g., see Figure 8). For example, the first left-joining adapter sequence (125) may be located between the first left universal adapter sequence (120) and the first left index sequence (160). The second left-joining adapter sequence (165) may be located between the first left index sequence (160) and the second left universal adapter sequence (140). The third left-joining adapter sequence (145) may be located between the second left universal adapter sequence (140) and the target sequence (110). The first right-joining adapter sequence (135) may be located between the first right universal sequence (130) and the first right index sequence (170). The second right-joining adapter sequence (175) may be located between the first right index sequence (170) and the second right universal adapter sequence (150). The third right-joining adapter sequence (155) may be located between the second right universal adapter sequence (150) and the target sequence (110). In some embodiments, the nucleic acid library molecule (100) further comprises at least one and up to ten additional universal adapter sequences located 5' (upstream) of the first left universal adapter sequence (120) (e.g., see Figure 8 In some embodiments, the nucleic acid library molecule (100) further comprises at least one and up to ten additional universal adapter sequences located 3' (downstream) of the first right universal sequence (130) (e.g., see Figure 8). Any connection adapter sequence and / or additional universal adapter sequence comprises any sequence and can be 3-60 nucleotides in length. Any connection adapter sequence and / or additional universal adapter sequence comprises a universal sequence or a unique sequence. Any connection adapter sequence and / or additional universal adapter sequence comprises a binding sequence for an amplification primer, a sequencing primer, a compaction oligonucleotide or a combination thereof. Any connection adapter sequence and / or additional universal adapter sequence comprises a binding sequence for a fixed surface primer (e.g., a capture primer). Any connection adapter sequence and / or additional universal adapter sequence comprises a sample index sequence. Any connection adapter sequence and / or additional universal adapter sequence comprises a unique identification sequence. Any connection adapter sequence and / or additional universal adapter sequence, particularly connection adapter sequence (145) comprises a Tn5 transposon end sequence 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 211). Any ligation adapter sequence and / or additional universal adapter sequence, in particular the ligation adapter sequence (155) comprises the Tn5 transposon end sequence 5'-CTGTCTCTTATACACATCT-3' (SEQ ID NO: 212). The Tn5 transposon end sequence can be introduced into the library molecule (100) by a transposase-mediated reaction under conditions suitable for forming a transposon synapse complex, the reaction comprising contacting double-stranded input DNA (e.g., genomic DNA) with a Tn-5 type transposase and a double-stranded oligonucleotide comprising the Tn transposon end sequence (SEQ ID NO: 211) linked to a universal adapter sequence or a sample index sequence. In the double-stranded oligonucleotide, the Tn transposon end sequence (SEQ ID NO: 211) can be located 5' or 3' relative to the universal adapter sequence or the sample index sequence.

[0373] In some embodiments of the method for forming a plurality of library-splint complexes (500), the second splint strand (400) comprises at least two sub-regions, the at least two sub-regions comprising a first sub-region and a second sub-region (e.g., Figure 2 and Figure 3 ). The first subregion comprises a universal binding sequence for a third surface primer, and the second subregion comprises a universal binding sequence for a fourth surface primer, wherein the first subregion and the second subregion do not hybridize (or at least exhibit little hybridization) with the first surface primer and the second surface primer. In some embodiments, the second splint strand (400) further comprises a third subregion comprising a sample index sequence of 5-20 bases and / or a unique identification sequence (e.g., NN) of 2-10 or more bases (e.g., Figure 3). In some embodiments, the second splint chain (400) comprises only one subregion and lacks a second subregion and a third subregion, wherein the first subregion comprises a sample index sequence having 5-20 bases. In some embodiments, the sample index sequence can be used to distinguish target sequences obtained from different sample sources in multiple assays. In some embodiments, the unique identification sequence comprises a random sequence. The unique identification sequence can be designed to exhibit reduced or no hybridization with the first, second, third, and fourth surface primers. An exemplary arrangement of the subregions in the second splint chain (400) in the 5' to 3' direction comprises: 5'-[second subregion]–[first subregion]–3'. Another exemplary arrangement of the subregions in the second splint chain (400) in the 5' to 3' direction comprises: 5'-[third subregion]–[second subregion]–[first subregion]–3'. In some embodiments, the second splint chain (400) may be 20-100 nucleotides in length, or 30-80 nucleotides in length, or 40-60 nucleotides in length. In some embodiments, the second splint chain (400) comprises one or more phosphorothioate bonds at the 5' and / or 3' end to confer resistance to nucleic acid exonucleases. In some embodiments, the second splint chain (400) comprises one or more phosphorothioate bonds at an internal position to confer resistance to nucleic acid endonucleases. In some embodiments, the second splint chain (400) comprises one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position. In some embodiments, the 5' end of the second splint chain (400) is phosphorylated or non-phosphorylated. In some embodiments, the 3' end of the second splint chain (400) comprises a terminal 3'OH group or a terminal 3' blocking group.

[0374] In some embodiments of the method for forming a plurality of library-splint complexes (500), the first splint strand (300) comprises an inner region (310) comprising at least two subregions, the at least two subregions comprising a fourth subregion and a fifth subregion (e.g., Figure 2 and Figure 3 ). The fourth subregion hybridizes with the first subregion of the second splint strand (400). The fifth subregion hybridizes with the second subregion of the second splint strand (400). The fourth and fifth subregions do not hybridize (or at least exhibit very little hybridization) with the first and second surface primers. In some embodiments, the interior region (310) of the first splint strand further comprises an optional sixth subregion that hybridizes with the third subregion of the second splint strand (400) (e.g., Figure 3). An exemplary arrangement of the subregions of the first splint chain (300) in the 5' to 3' direction includes: 5'-[fourth subregion]-[fifth subregion]-5'. Another exemplary arrangement of the subregions of the first splint chain (300) in the 5' to 3' direction includes: 5'-[fourth subregion]-[fifth subregion]-[sixth subregion]-3'. In some embodiments, the length of the first splint chain (300) may be 50-150 nucleotides, or 60-100 nucleotides, or 70-90 nucleotides in length. In some embodiments, the first splint chain (300) includes one or more phosphorothioate bonds at the 5' and / or 3' end to confer resistance to nucleic acid exonucleases. In some embodiments, the first splint chain (300) includes one or more phosphorothioate bonds at an internal position to confer resistance to nucleic acid endonucleases. In some embodiments, the first splint chain (300) includes one or more 2'-O-methylcytosine bases at the 5' and / or 3' end or at an internal position.

[0375] The present disclosure provides a method for forming multiple library-splint complexes (500), which includes: (a) providing multiple double-stranded splint adapters (200), wherein each double-stranded splint adapter (200) comprises a first splint chain (300) hybridized to a second splint chain (400), wherein the first splint chain (300) comprises regions arranged in 5' to 3' order: a first region (320), an internal region (310) and a second region (330), and wherein the internal region (310) of the first splint chain hybridizes to the second splint chain (400), wherein the second splint chain comprises regions arranged in 5' to 3' order: (i) a second sub-region having a universal binding sequence for a fourth surface primer, and (ii) a first sub-region having a universal binding sequence for a third surface primer. In some embodiments, the method for forming a plurality of library-splint complexes (500) further comprises step (b): hybridizing a plurality of double-stranded splint adaptors to a plurality of single-stranded nucleic acid library molecules (100), wherein each library molecule comprises regions arranged in 5' to 3' order: (i) a first left universal adaptor sequence (120) having a binding sequence for a first surface primer; (ii) a second left universal adaptor sequence (140) having a binding sequence for a first sequencing primer; (iii) a target sequence (110); (iv) a second right universal adaptor sequence (150) having a binding sequence for a second sequencing primer; and (v) a first right universal adaptor sequence (130) having a binding sequence for a second surface primer, wherein the hybridization The method is performed under conditions suitable for hybridizing the first splint strand (300) to the library molecule (100), thereby circularizing the library molecule to produce a library-splint complex (500), such that the first region (320) of the first splint strand hybridizes to the binding sequence for the first surface primer (120), and the third region (330) of the first splint strand hybridizes to the binding sequence (130) for the second surface primer, wherein the library-splint complex (500) comprises a first gap between the 5' end of the library molecule and the 3' end of the second splint strand (300), wherein the library-splint complex (500) comprises a second gap between the 5' end of the second splint strand (300) and the 3' end of the library molecule (100), and wherein the first gap and the second gap are enzymatically ligatable. In some embodiments, the plurality of single-stranded nucleic acid library molecules (100) comprise a first left index sequence (160) and / or a first right index sequence (170) (e.g., see Figure 5). A list of exemplary first left index sequences (160) and first right index sequences (170) is provided in Table 1 of Figure 33. In some embodiments, the first left index sequence (160) includes or lacks a short random sequence (e.g., NNN). In some embodiments, the first right index sequence (170) includes or lacks a short random sequence (e.g., NNN). In some embodiments, a plurality of single-stranded nucleic acid library molecules (100) comprise a first left unique identification sequence (180) and / or a first right unique identification sequence (190), each of which is contained in a population of other target sequence molecules for uniquely identifying the sequence of each target sequence (e.g., an insertion sequence) to which a unique adapter is attached. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular labeling. (See, for example, Figure 6 ).

[0376] A multiplex workflow is initiated by preparing a library with a sample index using one or two index sequences (e.g., a left index sequence and / or a right index sequence). A separate library with a sample index can be prepared using a first left index sequence (160) and / or a first right index sequence (170) using input nucleic acids isolated from different sources. Libraries with sample indexes can be pooled together to generate a multiple library mixture, and the pooled libraries can be amplified and / or sequenced. The sequence of the insertion region together with the first left index sequence (160) and / or the first right index sequence (170) can be used to identify the source of the input nucleic acid. In some embodiments, any number of libraries with sample indexes can be pooled together, for example, 2-10, or 10-50, or 50-100, or 100-200, or more than 200 libraries with sample indexes can be pooled. Exemplary nucleic acid sources include naturally occurring, recombinant, or chemically synthesized sources. Exemplary nucleic acid sources include single cells, multiple cells, tissues, biological fluids, environmental samples, or entire organisms. Exemplary nucleic acid sources include fresh, frozen, fresh frozen or archived sources (e.g., formalin fixed paraffin embedded; FFPE). The skilled artisan will recognize that nucleic acids can be isolated from many other sources. Nucleic acid library molecules can be prepared in single-stranded or double-stranded form.

[0377] In some embodiments, the plurality of single-stranded nucleic acid library molecules (100) further comprises a first left unique identification sequence (180) and / or a first right unique identification sequence (190) (see, e.g., Figure 6In some embodiments, the first left unique identification sequence (180) and the first right unique identification sequence (190) each comprise a sequence that uniquely identifies each target sequence (e.g., an insert sequence) to which a unique adaptor is attached, among a population of other target molecules. In some embodiments, the first left unique identification sequence (180) and / or the first right unique identification sequence (190) can be used for molecular labeling.

[0378] Method for forming library-splint complexes using double-stranded adapters with truncated long splints

[0379] In some embodiments of the method for forming a library-splint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the first splint strand (300) of each double-stranded splint adaptor (200) comprises a truncated strand having a first region (320) having a truncated sequence at the 5' end (e.g., Figure 11B , for example with Figure 11A 199 shown in ). In some embodiments, the 5' end of the first region can have a truncation of any length, for example, a truncation of 1-10 nucleotides. In some embodiments, the truncated first splint chain (300) comprises a second region (330; for example, SEQ ID NO: 5), a fourth subregion (for example, SEQ ID NO: 6), and a fifth subregion (for example, SEQ ID NO: 7) that are not truncated and do not carry any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the truncated first splint chain comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint chain (400) to form a double-stranded splint adapter (200). In some embodiments, the truncated first splint chain (300) comprises a truncated first region (320) that hybridizes to a sequence (for example, 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (for example, 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, a truncated first splint strand that is part of a double-stranded splint adaptor (200) can hybridize to a library molecule (100) to form a library-splint complex (500).

[0380] Method for forming library-splint complexes using double-stranded adapters with long splint strands having mismatched sequences

[0381] In some embodiments of the method for forming a library-splint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the first splint strand (300) of each double-stranded splint adaptor (200) comprises a first region (320) having a mismatch sequence (e.g., Figure 11C). In some embodiments, the mismatch sequence can be of any length (e.g., 2-20 bases) and comprises any sequence that is not completely complementary to the left universal adapter sequence (120) of the library molecule (100). Some embodiments of the mismatch sequence in the first region (320) are Figure 11C In some embodiments, the mismatched first splint chain (300) comprises a second region (330; e.g., SEQ ID NO: 5), a fourth subregion (e.g., SEQ ID NO: 6), and a fifth subregion (e.g., SEQ ID NO: 7) that do not carry any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the mismatched first splint chain (300) comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint chain (400) to form a double-stranded splint adapter (200). In some embodiments, the mismatched first splint chain (300) comprises a mismatched first region (320) hybridized to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) hybridized to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the mismatched first region (320) can hybridize to the first left universal adapter sequence (120) of the library molecule (100) to form a double-stranded portion having a bubble at the position of the mismatched sequence in the first region (320). In some embodiments, the mismatched first splint strand that is part of the double-stranded splint adapter (200) can hybridize to the library molecule (100) to form a library-splint complex (500).

[0382] Method for forming library-splint complexes using double-stranded adapters with long splint strands having abasic sites

[0383] In some embodiments of the method for forming a library-splint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the first splint strand (300) of each double-stranded splint adaptor (200) comprises at least one abasic site lacking a nitrogenous base. In some embodiments, the first splint strand (300) comprises at least one abasic site in the fourth subregion and / or at least one abasic site in the fifth subregion (e.g., Figure 11D, with the abasic sites shown as solid black bars). In some embodiments, the abasic sites each comprise a 1',2'-dideoxyribose (e.g., dSpacer from Integrated DNA Technologies (IDT)). In some embodiments, the abasic first splint strand (300) comprises a first region ((320); e.g., SEQ ID NO:4), and a second region ((330); e.g., SEQ ID NO:5) that does not carry any abasic sites and / or any sequence variants, such as insertions, deletions, or base substitutions. In some embodiments, the abasic first splint strand comprises a fourth subregion and a fifth subregion that can hybridize with the first subregion and the second subregion of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the abasic first splint strand (300) comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100) and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the abasic first splint strand as part of the double-stranded splint adaptor (200) can hybridize to the library molecule (100) to form a library-splint complex (500).

[0384] Method for forming library-splint complexes using double-stranded adapters having long splint strands of uracil

[0385] In some embodiments of the method for forming a library-splint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adapters (200). In some embodiments, the first splint strand (300) of each double-stranded splint adapter (200) comprises at least one uracil. In some embodiments, the first splint strand (300) of each double-stranded splint adapter (200) comprises at least one uracil in any one or any combination of regions comprising the first region (320), the second region (330), the fourth subregion, and / or the fifth subregion. In some embodiments, at least one thymine base can be substituted with uracil. An embodiment of a first splint strand containing uracil is shown in Figure 11D(bottom schematic). The skilled person will recognize that many other sequences of the first splint strand (300) comprising one or more uracils are possible. In some embodiments, the uracil-containing first splint strand comprises a fourth subregion and a fifth subregion, which can hybridize with the first subregion and the second subregion of the second splint strand (400) to form a double-stranded splint adapter (200). In some embodiments, the uracil-containing first splint strand (300) comprises a first region (320) that hybridizes to a sequence (e.g., 120) on one end of the linear single-stranded library molecule (100), and a second region (330) that hybridizes to a sequence (e.g., 130) on the other end of the linear single-stranded library molecule (100). In some embodiments, the uracil-containing first splint strand that is part of the double-stranded splint adapter (200) can hybridize with the library molecule (100) to form a library-splint complex (500).

[0386] Formation of library-splint complexes using double-stranded adapters with short splint strands containing inserted random and index sequences Method of things

[0387] In some embodiments of the method for forming a library-splint complex (500), a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the second splint strand (400) of each double-stranded splint adaptor (200) comprises a random sequence (e.g., Figure 12A and Figure 12B In some embodiments, the random sequence may replace a portion of the first sub-region of the second splint chain (e.g., Figure 13A and Figure 13B In some embodiments, the second sub-region of the second splint chain (400) does not have an inserted random sequence. In some embodiments, a portion of the second sub-region of the second splint chain (400) is not replaced by a random sequence.

[0388] In some embodiments, the random sequence can be of any length, for example, 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 12A and Figure 13A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 12B and Figure 13B In some embodiments, the random sequence may be inserted at any position in the first sub-region of the second splint chain (400).

[0389] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0390] In some embodiments, random sequences provide nucleotide diversity and color balance for sequencing reactions. In some embodiments, random sequences provide high nucleotide diversity, comprising approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U), which will be represented in each cycle of a sequencing run.

[0391] In some embodiments, the random sequence can be sequenced before sequencing the insert region. In some embodiments, sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides sufficient nucleotide diversity and color balance.

[0392] In some embodiments, a predetermined sequence is inserted into the sequence of the fourth sub-region of the first splint chain (300). The length of the inserted predetermined sequence may be the same as the length of the random sequence inserted into the first sub-region of the second splint chain (e.g., Figure 12A and Figure 12B ). The inserted predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 12A and Figure 12B shown.

[0393] In some embodiments, a portion of the fourth sub-region of the first cleat chain (300) is replaced by a predetermined sequence. The length of the predetermined sequence replacing the portion of the fourth sub-region is the same as the length of the random sequence replacing the portion of the first sub-region of the second cleat chain (e.g., Figure 13A and Figure 13B ). The replacement predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 13A and Figure 13B shown.

[0394] In some embodiments, the second splint strand (400) comprises a first subregion and a second subregion, which can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adaptor (200) (e.g., Figure 12A 、 Figure 12B 、 Figure 13A and Figure 13BIn some embodiments, the double-stranded splint adaptor (200) forms a bubble at the position where the random sequence is inserted or replaced.

[0395] In some embodiments, the second splint strand (400) as part of the double-stranded splint adaptor (200) carries a random sequence, which second splint strand can hybridize with the library molecule (100) to form a library-splint complex (500).

[0396] Formation of library-splint complexes using double-stranded adapters with short splint strands having additional random and index sequences Method of things

[0397] In any of the methods for forming a library-splint complex (500) as described herein, a single-stranded library molecule (100) can be hybridized to a plurality of double-stranded splint adaptors (200). In some embodiments, the second splint strand (400) of each double-stranded splint adaptor (200) comprises a random sequence (e.g., Figure 14A and Figure 14B In some embodiments, the random sequence further comprises an index sequence. In some embodiments, the second sub-region of the second splint chain (400) does not have an additional random sequence.

[0398] In some embodiments, the additional random sequence can be of any length, such as 2-10 bases in length. For example, the random sequence can be 3 nucleotides in length (e.g., Figure 14A 'NNN' in ) or 4 nucleotides in length (e.g., Figure 14B NNNN in the .

[0399] In some embodiments, in the random sequence, each base "N" at a given position is independently selected from A, G, C, T, or U. In some embodiments, the random sequence lacks consecutive repeats of 2 or 3 identical nucleobases, such as AA, TT, CC, GG, UU, AAA, TTT, CCC, GGG, or UUU. In some embodiments, the population of the second splint strand (400) comprises a random sequence having a high diversity sequence that includes approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U) that will be represented in each cycle of the sequencing run.

[0400] In some embodiments, random sequences provide nucleotide diversity and color balance for sequencing reactions. In some embodiments, random sequences provide high nucleotide diversity, comprising approximately equal proportions of all four nucleotides (e.g., A, G, C, T, and / or U), which will be represented in each cycle of a sequencing run.

[0401] In some embodiments, the random sequence can be sequenced before sequencing the insert region. In some embodiments, sequencing data from the random sequence can be used for polymerase clone mapping and / or template alignment because the random sequence provides sufficient nucleotide diversity and color balance.

[0402] In some embodiments, index sequences can be used to distinguish polynucleotides (eg, insert sequences) from different sample sources in a multiplex assay.

[0403] In some embodiments, a predetermined sequence is attached to the 5' end of the fourth sub-region of the first splint chain (300). In some embodiments, the length of the attached predetermined sequence can be the same as the length of the random sequence attached to the 3' end of the first sub-region of the second splint chain (e.g., Figure 14A and Figure 14B In some embodiments, the length of the additional predetermined sequence may be the same as the length of the random sequence and index sequence attached to the 3' end of the first subregion of the second splint chain (e.g., Figure 14A and Figure 14B ). The additional predetermined sequence can have any sequence. Exemplary predetermined sequences are as follows Figure 14A and Figure 14B shown.

[0404] In some embodiments, the second splint strand (400) includes a first subregion and a second subregion, which can hybridize with the fourth subregion and the fifth subregion of the first splint strand (300) to form a double-stranded splint adaptor (200) (e.g., Figure 14A and Figure 14B In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the position of the additional random sequence. In some embodiments, the double-stranded splint adapter (200) forms a bubble or mismatched end at the position of the additional random sequence and the index sequence.

[0405] In some embodiments, a second splint strand (400) that is part of a double-stranded splint adaptor (200) is appended with a random sequence (and optionally an index sequence) that can hybridize with a library molecule (100) to form a library-splint complex (500).

[0406] Sequence of short splint chains

[0407] In some embodiments of the methods described herein for forming a plurality of library-splint complexes (500), the first subregion of the second splint strand (400) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 200). In some embodiments, the second subregion of the second splint strand (400) comprises the sequence 5'-AGTCGTCGCAGCCTCACCTGATC-3' (SEQ ID NO: 201). In some embodiments, the second splint strand (400) comprises a first and a second subregion comprising the sequence 5'-AGTCGTCGCAGCCTCACCTGATCCATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 202). See Figure 11A In some embodiments, the 5' end of the second splint strand (400) can be phosphorylated or non-phosphorylated.

[0408] Sequence of long splint chains

[0409] In some embodiments of the methods for forming the plurality of library-splint complexes (500) described herein, the first region (320) of the first splint strand comprises a first universal adapter sequence comprising a universal binding sequence for a first surface primer (or its complement), wherein the first region (320) comprises the sequence 5'-TCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 193). For example, the first region (320) of the first splint strand can hybridize to a P5 surface primer or a complement of a P5 surface primer. For example, the P5 surface primer comprises the sequence

[0410] 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO: 203; short P5), or the P5 surface primer comprises the sequence 5'-AATGATACGGCGACCACCGAGATC-3' (SEQ ID NO: 194; long P5). In some embodiments, the second region (330) of the first splint strand comprises a second universal adapter sequence comprising a universal binding sequence (or its complement) for a second surface primer, wherein the second region (330) comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195). For example, the second region (330) of the first splint strand can hybridize to the P7 surface primer or the complement of the P7 surface primer. For example, the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 195; short P7), or the P7 surface primer comprises the sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 196; long P7). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fourth subregion having the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO: 197). In some embodiments, the first splint strand (300) comprises an internal region (310) comprising a fifth subregion having the sequence 5'-GATCAGGTGAGGCTGCGACGACT'3' (SEQ ID NO: 198). In some embodiments, the first splint strand (300) comprises a first region (320), an internal region (310) having fourth and fifth subregions, and a second region (330) having the sequence 5'-TCGGTGGTCGCCGTATCATTACCCTGAAAGTACGTGCATTACATGGATCAGGTGAG GCTGCGACGACTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 199). Figure 11A In some embodiments, the 5' end of the first splint chain (300) can be phosphorylated or non-phosphorylated. In some embodiments, the first subregion of the second splint chain (400) can hybridize with the fourth subregion of the first splint chain (300). In some embodiments, the second subregion of the second splint chain (400) can hybridize with the fifth subregion of the first splint chain (300).

[0411] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the first region (320) of the first splint chain comprises a sequence that can bind to the first left universal adapter sequence (120) of the library molecule, wherein the first region (320) of the first splint chain comprises the sequence 5'-ACCCTGAAAGTACGTGCATTACATG-3' (SEQ ID NO:215) or its complement.

[0412] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the second region (330) of the first splint chain comprises a sequence that can bind to the first right universal adapter sequence (130) of the library molecule, wherein the second region (330) of the first splint chain comprises the sequence 5'-GATCAGGTGAGGCTGCGACGACT-3' (SEQ ID NO:216) or its complement.

[0413] Sequence of the library-splint complex

[0414] In some embodiments of the methods for forming a plurality of library-splint complexes (500) described herein, the library molecule comprises a first left universal adapter sequence (120) that binds to a first region (320) of a first splint strand, wherein the first left universal adapter sequence (120) comprises the sequence

[0415] 5'-AATGATACGGCGACCACCGA-3' (SEQ ID NO:203).

[0416] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule includes a first left universal adapter sequence (120) that binds to the first region (320) of the first splint chain, wherein the first left universal adapter sequence (120) comprises the sequence 5'-CATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO:213) or its complement.

[0417] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule includes a second left universal adapter sequence that includes a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence comprises the sequence 5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO:204).

[0418] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule includes a second left universal adapter sequence that includes a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence includes the sequence 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO:207).

[0419] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule includes a second left universal adapter sequence that includes a sequence for binding to a sequencing primer (140), wherein the second left universal adapter sequence includes the sequence 5'-CGTGCTGGATTGGCTCACCAGACACCTTCCGACAT-3' (SEQ ID NO:208).

[0420] In some embodiments, in any of the methods for forming multiple library-splint complexes (500) described herein, the library molecule includes a second right universal adapter sequence that includes a sequence for binding to a sequencing primer (150), wherein the second right universal adapter sequence includes the sequence 5'-AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC-3' (SEQ ID NO:205).

[0421] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule includes a second right universal adapter sequence that includes a sequence for binding to a sequencing primer (150), wherein the second right universal adapter sequence includes the sequence 5'-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3' (SEQ ID NO:209).

[0422] In some embodiments of the methods for forming multiple library-splint complexes (500) described herein, the library molecule includes a second right universal adapter sequence that includes a sequence for binding to a sequencing primer (150), wherein the second right universal adapter sequence inc...

Claims

1. A library-clamp complex (500) comprising: (i) A single-stranded nucleic acid library molecule (100) comprising a target sequence (110) flanked on one side by at least a first left universal adapter sequence (120) and on the other side by at least a first right universal adapter sequence (130); and (ii) A double-stranded clamp adapter (200) comprising a first clamp strand (300) and a second clamp strand (400), wherein the double-stranded clamp adapter (200) comprises a double-stranded region and two single-stranded regions, each of the two single-stranded regions being located on both sides of the double-stranded region, wherein the first clamp strand comprises a first region (320), an internal region (310), and a second region (330); wherein the internal region (310) of the first clamp strand hybridizes with the second clamp strand (400), wherein the first region (320) of the first clamp strand hybridizes with at least the first left universal adapter sequence (120) of the library molecule, and wherein the second region (330) of the first clamp strand hybridizes with at least the first right universal sequence (130) of the library molecule, thereby looping the library molecule to produce the library-clamp complex (500).

2. The library-clamp complex (500) according to claim 1, wherein the nucleic acid library molecule (100) further comprises: a second left universal adapter sequence (140).

3. The library-clamp complex (500) according to claim 2, wherein the second left universal adapter sequence (140) is located between the at least first left universal adapter sequence (120) and the target sequence (110).

4. The library-clamp complex (500) according to any one of claims 1 to 3, wherein the nucleic acid library molecule (100) further comprises: a second right universal adapter sequence (150).

5. The library-clamp complex (500) according to claim 4, wherein the second right universal adapter sequence (150) is located between the target sequence (110) and at least the first right universal adapter sequence (130).

6. The library-clamp complex (500) according to any one of claims 1 to 5, wherein the nucleic acid library molecule (100) further comprises: a first left index sequence (160).

7. The library-clamp complex (500) according to claim 6, wherein the first left index sequence (160) is located between the at least first left universal adapter sequence (120) and the target sequence (110).

8. The library-clamp complex (500) according to any one of claims 1 to 7, wherein the nucleic acid library molecule (100) further comprises: a first right index sequence (170).

9. The library-clamp complex (500) according to claim 8, wherein the first right index sequence (170) is located between the second right universal adapter sequence (150) and at least the first right universal adapter sequence (130).

10. The library-clamp complex (500) according to any one of claims 1 to 9, wherein the nucleic acid library molecule (100) further comprises: a first left unique identification sequence (180).

11. The library-clamp complex (500) according to claim 10, wherein the first left unique identification sequence (180) is located between the at least first left universal adapter sequence (120) and the first left index sequence (160).

12. The library-clamp complex (500) according to any one of claims 1 to 11, wherein the nucleic acid library molecule (100) further comprises: a first right unique identification sequence (190).

13. The library-clamp complex (500) according to claim 12, wherein the first right unique identification sequence (190) is located between the first right index sequence (170) and the at least first right universal adapter sequence (130).

14. The library-clamp complex (500) according to any one of claims 1 to 13, wherein the nucleic acid library molecule (100) further comprises any one or both or any combination of two or more of the following: (i) A second left universal adapter sequence (140); (ii) A second right universal adapter sequence (150); (iii) A first left index sequence (160); (iv) A first right index sequence (170); (v) A first left unique identification sequence (180); and / or (vi) A first right unique identification sequence (190).

15. The library-clamp complex (500) according to any one of claims 1 to 14, wherein the first left universal adapter sequence (120) and / or the second left universal adapter sequence (140) comprises: (i) A universal binding sequence for a forward sequencing primer; (ii) A universal binding sequence for a reverse sequencing primer; (iii) A universal binding sequence for a first surface primer; (iv) A universal binding sequence for a second surface primer; (v) A universal binding sequence for a forward amplification primer; (vi) A universal binding sequence for a reverse amplification primer; and / or (vii) A universal binding sequence for a compaction oligonucleotide.

16. The library-clamp complex (500) according to any one of claims 1 to 15, wherein the first right universal adapter sequence (130) and / or the second right universal adapter sequence (150) comprises: (i) A universal binding sequence for a forward sequencing primer; (ii) A universal binding sequence for a reverse sequencing primer; (iii) A universal binding sequence for a first surface primer; (iv) A universal binding sequence for a second surface primer; (v) A universal binding sequence for a forward amplification primer; (vi) A universal binding sequence for a reverse amplification primer; and / or (vii) A universal binding sequence for a compaction oligonucleotide.

17. The library-clamp complex (500) according to claim 15 or 16, wherein the second clamp strand (400) comprises at least two sub-regions: a first sub-region comprising a universal binding sequence for a third surface primer; and a second sub-region comprising a universal binding sequence for a fourth surface primer, wherein the first sub-region and the second sub-region do not hybridize or exhibit very little hybridization with the first surface primer and the second surface primer.

18. The library-clamp complex (500) according to claim 17, wherein the second clamp strand (400) comprises an optional third sub-region, wherein the third sub-region comprises a sample index sequence having 5-20 bases and / or a unique identification sequence having 2-10 or more bases.

19. The library-clamp complex (500) according to claim 18, wherein the unique identification sequence comprises a random sequence.

20. The library-clamp complex (500) according to claim 15 or 16, wherein the first clamp strand (300) comprises an internal region (310), the internal region comprising at least two sub-regions: a fourth sub-region, which comprises a universal binding sequence for a third surface primer, and the fourth sub-region hybridizes to a first sub-region of the second clamp strand (400); and a fifth sub-region comprising a universal binding sequence for a fourth surface primer, and the fifth sub-region hybridizes with the second sub-region of the second clamp strand (400); wherein the fourth sub-region and the fifth sub-region do not hybridize or at least exhibit very little hybridization with the first surface primer and the second surface primer.

21. The library-clamp complex (500) according to claim 20, wherein the first clamp strand (300) comprises an internal region (310), and the internal region further comprises a sixth sub-region, and the sixth sub-region comprises a sample index sequence having 5-20 bases and / or a unique identification sequence having 2-10 or more bases, wherein the sixth sub-region hybridizes with the third sub-region of the second clamp strand (400).

22. The library-clamp complex (500) according to claim 21, wherein the unique identification sequence comprises a random sequence.

23. A library-clamp complex (500) comprising: a) a single-stranded nucleic acid library molecule (100) comprising components arranged in a 5' to 3' order: (i) a first left universal adapter sequence (120) having a binding sequence (120) for a first surface primer; (ii) a second left universal adapter sequence (140) having a binding sequence for a first sequencing primer; (iii) a target sequence (110); (iv) a second right universal adapter sequence (150) having a binding sequence for a second sequencing primer; and (v) a first right universal adapter sequence (130) having a binding sequence (130) for a second surface primer; b) a first clamp strand (300) comprising components arranged in a 5' to 3' order: a first region (320); an internal region (310); and a second region (330); and c) A second splint strand (400), which comprises sub-regions arranged in a 3'-to-5' order: a first sub-region having a universal binding sequence for a third surface primer; and a second sub-region having a universal binding sequence for a fourth surface primer; wherein said first splint strand (300) hybridizes to a portion of said library molecule (100), thereby looping said library molecule to produce a library-splint complex (500), such that said first region (320) of said first splint strand hybridizes to said binding sequence (120) for said first surface primer, and said third region (330) of said first splint strand hybridizes to said binding sequence (130) for said second surface primer, wherein said second splint strand (400) hybridizes to said internal region (310) of said first splint strand (300), wherein said library-splint complex (500) comprises a first gap between the 5'-end of said library molecule and the 3'-end of said second splint strand, and wherein said library-splint complex (500) comprises a second gap between the 5'-end of said second splint strand and the 3'-end of said library molecule.

24. The library-splint complex (500) according to claim 23, wherein said first gap and said second gap are enzymatically ligatable.

25. A plurality of library-splint complexes, which comprises the library-splint complex (500) according to any one of the preceding claims, wherein the target sequence (110) of each library-splint complex in said plurality of library-splint complexes comprises the same target sequence or different target sequences.

26. A method for generating a library-splint complex according to any one of claims 1 to 22, which comprises: a. Providing a plurality of single-stranded nucleic acid library molecules (100); b. Providing a plurality of double-stranded splint adaptors (200), a first splint strand (300) and a second splint strand (400); and c. Contacting said plurality of single-stranded nucleic acid library molecules with said plurality of double-stranded splint adaptors under conditions sufficient to hybridize the ends of said first splint strand to the ends of said library molecule, thereby generating a plurality of library-splint complexes.

27. A method for generating a library-splint complex according to claim 23 or 24, which comprises: a. Providing a plurality of single-stranded nucleic acid library molecules, a plurality of first splint strands and a plurality of second splint strands; and b. Contacting said plurality of single-stranded nucleic acid library molecules with said plurality of first splint strands and said plurality of second splint strands under conditions sufficient to hybridize said second splint strand to said first splint strand and to hybridize the ends of said first splint strand to the ends of said library molecule, thereby generating a plurality of library-splint complexes.

28. A method for sequencing a plurality of concatemeric template molecules, which comprises: a. Providing a plurality of library-splint complexes according to any one of claims 1 to 24; b. Performing rolling circle amplification on said plurality of library-splint complexes to generate a plurality of concatemeric template molecules; and c. Sequencing said plurality of concatemeric template molecules.

29. A kit, which comprises a plurality of double-stranded splint adaptors according to any one of claims 1 to 22.

Citation Information

Patent Citations

  • Method and system for sequencing nucleic acids

    US10246744B2

  • Engineered polymerases for improved sequencing

    US10731141B2

  • DNA sequencing method using acyclonucleoside triphosphates

    US5558991A