Sequence determination template containing multiple inserts, and composition and method for improving sequence determination throughput.

JP2026104930A5Pending Publication Date: 2026-07-21ILLUMINA INC +1

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2026-04-13
Publication Date
2026-07-21

Smart Images

  • Figure 00000137_0000
    Figure 00000137_0000
  • Figure 00000137_0001
    Figure 00000137_0001
  • Figure 00000138_0000
    Figure 00000138_0000
Patent Text Reader

Abstract

This application relates to a polynucleotide comprising a lead primer binding sequence, an insert sequence derived from a target nucleic acid, a ligation sequence, and an adhesion sequence. [Solution] Polynucleotides for use as sequencing templates containing multiple inserts are described herein. Methods for generating and using these polynucleotides, as well as methods for using such templates, including analysis of continuity information, are also described herein. Furthermore, sequencing templates containing insert sequences and copies of insert sequences can be used to correct random errors occurring during sequencing or amplification, or to identify nucleic acid base damage or other mutations resulting in non-standard base pairing in double-stranded nucleic acids. Methods for performing methylation analysis are also described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 094,422, filed on 21 October 2020, and U.S. Provisional Patent Application No. 63 / 256,040, filed on 15 October 2021, the contents of which are incorporated herein by reference in their entirety for all purposes. explanation

[0002] This application relates to polynucleotides comprising a lead-primer-binding sequence, an insert sequence derived from a target nucleic acid, a ligation sequence, and an adherent sequence. Compositions comprising these polynucleotides, as well as methods for generating and sequencing ligated nucleic acid sequencing templates, are also described. In addition, this disclosure relates to methods for preparing sequencing templates comprising multiple inserts. This disclosure also relates to methods for using such templates, including the analysis of continuity information. Furthermore, sequencing templates comprising two copies of the same insert sequence (i.e., the insert sequence and a copy of the insert sequence) can be used to correct random errors occurring during sequencing or amplification, or to identify damage to nucleic acid bases or other mutations resulting in non-standard base pairing in double-stranded nucleic acids. These sequencing templates comprising the insert sequence and a copy of the insert sequence can also be used for methylation analysis. [Background technology]

[0003] Typically, read lengths in synthetic sequencing (SBS) platforms are limited to 250–300 base pairs due to phasing / prephasing. This read length limits the throughput of the SBS platform.

[0004] Previously, methods for improving SBS throughput using polynucleotides containing multiple inserts have been described. In many cases, these methods relied on orthogonal SBS reactions using, for example, combinations of different polymerases or substrates, or primer blocking (see International Publication 2015 / 0002789 and U.S. Patent Application Publication 20180312917). However, there is a need for a simpler, more cost-effective, and user-friendly method to increase sequencing output from flow cells without requiring non-standard reagents.

[0005] This disclosure describes polynucleotides comprising multiple insert sequences derived from one or more target nucleic acids. These polynucleotides can be generated from multiple DNA libraries. They are then made extendable by annealing a hybridization sequence in one library product to the complement of a hybridization sequence in another library product, thereby forming a hybridized adduct and creating a polynucleotide comprising multiple insert sequences. Sequencing of these multiple insert sequences can be performed by a sequential SBS extension reaction based on multiple different read-primer binding sequences contained within the polynucleotide.

[0006] In addition, conventional short-read sequencing methods involve the initial generation of short, distinct fragments derived from intact genomic DNA or RNA. These fragments are generated by several methods, such as physical shearing, enzymatic digestion, or polymerase extension from one or more primers. Template preparations then modify these fragments and add synthetic adapters to enable sequencing. These sequencing templates are often ordered and juxtaposed in the same order as the intact genome. The template always contains a single fragment derived from the original sample containing the base sequence. When the template is double-stranded, the sequence complements are associated by hybridization of the two strands. However, when the double-stranded template is denatured, the two complementary strands separate, and the template becomes single-stranded, containing a single sequence fragment derived from the original sample. In this process, any association between the two complementary strands is lost. Furthermore, in this fragmentation and template preparation process, any association between two or more adjacent fragments in the original unfragmented genome is also lost.

[0007] An exception to this rule, which involves the loss of continuity information, is found in template preparation methods that use ligation to ligate two or more distal fragments together before adding a sequencing adapter. One example is the “mate pair” library, where the ends of a large DNA fragment are ligated together to form a ring, which is then further fragmented, and subsequently, smaller fragments spanning the ligated ends are recovered. The resulting template contains two sequences derived from the original large fragment, tandem-ligated. Another example is chromatin-based structural capture, where distal fragments of DNA in a genome are spatially organized in close proximity by the structural arrangement of the DNA complexed with chromatin in vivo. The ligation and subsequent processing of the closely spaced fragments generate a sequencing template with tandem inserts that provide information about the spatial relevance and, by inference, the functional relevance of the individual inserts.

[0008] Many different methods exist, such as Duplex Sequencing (Schmitt, et al. Proc. Natl. Acad. Sci. USA 109:14508-14513 (2012), Duplex Proximity Sequencing (Pro-Seq, described in Pel et al. PLoS One 13:1-19 (2018)), CypherSeq (Gregory et al. Nucleic Acids Res. 44:e22 (2016)), o2n-seq (Wang et al. Nat. Commun. 8, 15335 (2017)), Circle Sequencing (Lou et al., Proc. Natl. Acad. Sci. USA 110:19872-19877 (2013)), and Bot Sequencing (Hoang et al. al. Proc. Natl. Acad. Sci. USA 113:9846-9851 (2016) and Abascal et al. Nature 593, 405-410 (2021) have been developed as potentially improved methods for preparing sequencing templates with multiple inserts. However, all of these methods have drawbacks and none are universally applicable.

[0009] Recently described in Bae et al., bioRxiv, 10.1101 / 2021.06.11.448110, posted on June 12, 2021, Concatenating The Original Duplex for Error Correction (CODEC) method involves physically ligating both strands of a double-stranded DNA molecule for sequencing a single double-stranded molecule with a single read pair using a special CODEC adapter complex. Using the CODEC method, it is possible to identify non-standard base pairings that may result from nucleic acid base damage or changes present only on one strand of a double-stranded nucleic acid, as well as errors that may have been introduced during PCR amplification or sequencing. However, the CODEC method requires two consecutive ligations, which can limit conversion efficiency, and byproducts may also be formed by undesirable ligations.

[0010] If there are no inherent structural relationships between sequences in the genome, a barcode-like surrogate "association marker" may be used. For example, large fragments of DNA, such as those exceeding 1000 base pairs, or even exceeding 5000 base pairs, can be isolated by dilution, compartmentalization, or immobilization on a surface, and further fragmented, where each small fragment is then given a common barcode sequence. Thus, if many fragments are processed in parallel and each isolated fragment receives a unique barcode sequence attached to its subsequent subsequence, then a pool of all small fragments derived from all fragments can be formed. This allows sequencing in a single experiment, and the small fragments can be clearly identified by identifying and matching their barcode sequences. This approach enables the association of adjacent sequences within the genome, allowing for the in silico assembly of numerous small fragments into much larger in silico fragments, which can be useful for phasing variants in the genome.

[0011] Another type of barcoding uses unique molecular identifiers (UMIs) to preserve the relationships between sequences within the genome that are physically separated during template preparation and sequencing. UMIs consist of short barcode sequences appended to DNA or RNA fragments during template preparation, so that each individual single molecule receives its own unique barcode. Reading UMIs through sequencing allows for the identification of individual molecules (such as fragments in the template preparation) even if the original sample contained two or more identical fragments in terms of length and sequence. UMIs also help identify errors (e.g., modifications to the natural genome sequence) generated and propagated during PCR or other such methods for creating copies of the original template. This is useful in experiments sequencing samples containing low-frequency natural variants that would otherwise be difficult to identify in the background of artificial variants produced by PCR. Another application of UMI is that double-stranded fragments can be ligated by attaching them to a double-stranded adapter containing double-stranded UMI (i.e., a UMI barcode hybridized to its complement in the double-stranded adapter so that the first and second strands of the genome fragment each bear a common UMI barcode). In this way, after separation by denaturation, the first and second strands can be identified and reassigned by UMI. The use of such UMI can help improve sequencing accuracy by providing two types of “reads” of sequences in the genome, in other words, by identifying and using “sense” and “antisense” pairs of templates derived from the fragment, and by inferring the validity of base calls in sequencing reads of either template.

[0012] The use of barcodes to associate either distal or complementary sequences within a genome is actually complex due to constraints regarding the design and incorporation of barcodes during adapter and sequencing reactions. For example, there are a finite number of permutations for a given length of barcode. In one example, a 4-base barcode has only 256 permutations, and not all are actually functional considering self-complementarity and other sequences. Similar problems arise when the barcode is longer but with the added penalty of requiring more sequencing cycles to read the barcode.

[0013] Adding a barcode to an adapter adds complexity to the adapter itself. For example, adding a change in performance from one adapter to another brings issues regarding normalization in a library pool. Complex barcodes also require complex manufacturing, especially when the barcode and its complement hybridize in a double-stranded adapter.

[0014] The use of in vivo structural associations such as mate pairs or chromatin conformation capture also requires a complex workflow and is limited in the associations it can identify. For example, the issues with mate pairs are extremely large fragments, while the issues with chromatin conformation capture are chromatin-induced associations.

[0015] Methods that do not use barcodes to provide relevant information about adjacent and complementary sequences within a genome are disclosed herein. These methods can utilize surfaces for concatenating sequences tandemly within a single template. The methods can also use partitioning to generate templates for proximity or haplotype data. Once sequenced, the resulting templates can provide information for correcting errors in sequencing or identifying non-standard base pair formation, and can also provide continuous information for assembly and phasing of genomic information. and can do so.

[0016] Methods for detecting methylation status are also disclosed herein. Conventional methods for detecting methylation status in genomic DNA generally use chemical or biochemical reactions to convert a base of interest to a different base. Detection of this conversion is used to infer whether a base is methylated or not. These methods require splitting a sample into two aliquots. One aliquot is chemically / biochemically treated, and the other is left untreated. Both are then sequenced and compared to each other to estimate the methylation status. An example of such a chemical method is bisulfite sequencing, which uses the conversion of unmethylated C bases to U bases with sodium bisulfite. The uracil nucleotides are then converted to thymine nucleotides during an amplification step such as PCR. After sequencing both the treated and untreated samples, the reads are compared, and if a C base in the untreated sample is read as a T base in the treated sample, it indicates that this C base was not methylated in the original sample. However, if a C base in the untreated sample is still read as a C base in the treated sample, theoretically, the C base was methylated in the original sample.

[0017] As described in Vaisvilas et al., Genome Res. 31(7):1280-1289 (2021), a similar strategy is used in EM-Seq assays, with the exception that an enzymatic reaction, rather than a chemical reaction, is used to convert unmethylated C. A recent publication (Liu et al., Nature Biotechnology 37(4):424-429 (2019)) introduced an alternative borane-based chemical method that converts methylated C nucleotides but not unmethylated C nucleotides. This has been reported to be advantageous over conventional C conversion chemical methods such as bisulfite sequencing because, as opposed to bisulfite chemistry where the converted genome is mostly A, G, and T, only a small percentage of the genome is methylated, so the converted genome is mostly still a 4-nucleotide genome containing A, C, G, and T.

[0018] A common feature of current methods for methylation analysis is the need to split the sample into two aliquots, which are then processed and sequenced in parallel. Techniques exist that directly detect the methylation status of bases without the need to split the sample. These methods rely on single-molecule sequencing techniques that employ sequencing strategies capable of distinguishing between methylated and unmethylated bases in the original sample. An example of such a technique is nanopore sequencing (e.g., "Epigenetics and methylation analysis," Oxford Nanopore Technologies, downloaded). Examples include (see nanoporetech.com / applications / investigation / epigenetics-and-methylation-analysis on October 7, 2021) and SMRT sequencing (described in Flusberg et al., Nat Methods. 7(6):461-465 (2010)). However, these strategies require high-throughput sequencing or are unsuitable for methods where the target genome is small in fragment size, such as cell-free DNA. A method for processing and sequencing single aliquot methylated samples to identify the methylation status of a genome is described herein. This method includes a method for distinguishing hydroxymethylated cytosine from methylated cytosine. The method of the present invention can reduce the effort involved in sample preparation and sequencing and can potentially reduce the amount of starting material required for methylation analysis. [Prior art documents] [Non-patent literature]

[0019] [Non-Patent Document 1] Schmitt,et al.Proc.Natl.Acad.Sci.USA109:14508-14513(2012) [Non-Patent Document 2] Pel et al.PLoS One 13:1-19(2018) [Non-Patent Document 3] Gregory et al., Nucleic Acids Res. 44: e22 (2016) Non-Patent Document 4 Wang et al., Nat. Commun. 8, 15335 (2017) Non-Patent Document 5 Lou et al., Proc. Natl. Acad. Sci. U.S.A. 110: 19872 - 19877 (2013) Non-Patent Document 6 Hoang et al., Proc. Natl. Acad. Sci. U.S.A. 113: 9846 - 9851 (2016) Non-Patent Document 7 Abascal et al., Nature 593, 405 - 410 (2021) Non-Patent Document 8 Bae et al., bioRxiv, 10.1101 / 2021.06.11.448110 Non-Patent Document 9 Vaisvilas et al., Genome Res. 31(7): 1280 - 1289 (2021) Non-Patent Document 10 Liu et al., Nature Biotechnology 37(4): 424 - 429 (2019) Non-Patent Document 11 Flusberg et al., Nat Methods. 7(6): 461 - 465 (2010) Summary of the Invention Means for Solving the Problems

[0020] Polynucleotides containing multiple insert sequences are described herein. These polynucleotides may be used in methods that enable sequencing of multiple insert sequences derived from a target nucleic acid. Also described herein are polynucleotides containing multiple inserts for use as sequencing templates in methods for error correction and identification of non-standard base pairings, determination of continuity data, and methylation analysis.

[0021] Embodiment 1 is a polynucleotide comprising: (a) a 5'-terminal polynucleotide including a first read primer binding sequence; (b) a first insert sequence located on the 3' side of the 5'-terminal polynucleotide, which is derived from a target nucleic acid; (c) a ligation sequence located on the 3' side of the first insert sequence, which includes a second read primer binding sequence and a hybridization sequence; (d) a second insert sequence located on the 3' side of the ligation sequence, which is derived from a non-adjacent sequence of the target nucleic acid or from a target nucleic acid different from the first insert sequence; and (e) a 3'-terminal polynucleotide sequence.

[0022] Embodiment 2 is a polynucleotide comprising a 3'-terminal polynucleotide containing a first read primer binding sequence, a first insert sequence on the 5' side of the 3'-terminal polynucleotide derived from a target nucleic acid, and a ligation sequence containing a second read primer binding sequence orthogonal to the first read primer binding sequence, wherein the second read primer binding sequence contains a hybridization sequence, a second insert sequence located on the 5' side of the ligation sequence and derived from a non-adjacent sequence of the target nucleic acid or from a target nucleic acid different from the first insert sequence, and an attached polynucleotide located at the 5' end of the polynucleotide and containing an attachment sequence, wherein the 3'-terminal polynucleotide, ligation sequence, and attached polynucleotide do not originate from a target nucleic acid.

[0023] Embodiment 3 is an embodiment of Embodiment 1 or 2 in which two insert sequences originate from different target nucleic acids. These are polynucleotides as described in [reference].

[0024] Embodiment 4 is a polynucleotide according to any one of Embodiments 1 to 3, wherein the first insert sequence and the second insert sequence each independently contain 40 to 400 nucleotides, 100 to 200 nucleotides, or 150 nucleotides.

[0025] Embodiment 5 is a polynucleotide according to any one of Embodiments 1 to 4, wherein the first read primer binding sequence includes the first adapter sequence.

[0026] Embodiment 6 is a polynucleotide according to any one of Embodiments 1 to 5, wherein the first read primer binding sequence further comprises a complement of the transposon terminal sequence.

[0027] Embodiment 7 is a polynucleotide according to Embodiment 5 or 6, wherein the first adapter sequence is a complement of the A14 primer sequence (A14') or a complement of the B15 primer sequence (B15').

[0028] Embodiment 8 is a polynucleotide according to any one of Embodiments 3 to 7, wherein the 3' terminal polynucleotide is a complement of the P7 primer sequence (P7') or a complement of the P5 primer sequence (P5').

[0029] Embodiment 9 is a polynucleotide according to any one of Embodiments 2 to 7, wherein the 3' terminal polynucleotide contains a complement of the P5 primer sequence (P5') and the attached polynucleotide contains a P7 primer sequence (P7), or the 3' terminal polynucleotide contains a complement of the P7 primer sequence (P7') and the attached polynucleotide contains a P5 primer sequence (P5).

[0030] Embodiment 10 is a polynucleotide according to any one of Embodiments 2 to 9, wherein the ligation sequence includes (a) a hybridization sequence and optionally includes (b) a complement of the 3' transposon terminal sequence of the hybridization unit and the 5' transposon terminal sequence of the hybridization unit.

[0031] Embodiment 11 is the polynucleotide described in Embodiment 10, wherein the second read primer binding sequence includes a complement of the hybridization sequence and the transposon terminal sequence.

[0032] Embodiment 12 is a polynucleotide according to any one of Embodiments 2 to 11, wherein the attached polynucleotide comprises a second adapter sequence and optionally a transposon terminal sequence.

[0033] Embodiment 13 is the polynucleotide according to Embodiment 12, wherein the second adapter sequence is the A14 sequence or the B15 sequence.

[0034] Embodiment 14 is the polynucleotide according to Embodiment 13, wherein the first adapter sequence is the complement (A14') of the A14 sequence and the second adapter sequence is the B15 sequence, or the first adapter sequence is the complement (B15') of the B15 sequence and the second adapter sequence is the A14 sequence.

[0035] Embodiment 15 is one of the embodiments described in any one of Embodiments 2-7 or 9-14, wherein the 3' terminal polynucleotide and / or attached polynucleotide each independently includes at least one of the following: a barcode sequence, a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence. It is a polynucleotide.

[0036] Embodiment 16 is a polynucleotide according to any one of Embodiments 2-7 and 9-14, wherein the polynucleotide is immobilized on a solid support.

[0037] Embodiment 17 is the polynucleotide described in Embodiment 16, wherein the polynucleotide is immobilized on a solid support via an attached polynucleotide.

[0038] Embodiment 18 is the polynucleotide according to Embodiment 17, wherein the polynucleotide is immobilized on a solid support via hybridization of the attached polynucleotide to an attached polynucleotide complement on the surface of the solid support.

[0039] Embodiment 19 is the polynucleotide according to Embodiment 17, wherein the polynucleotide is immobilized on a solid support via the binding of affinity moieties on the attached polynucleotide to binding moieties on the surface of the solid support.

[0040] Embodiment 20 is a polynucleotide according to any one of Embodiments 16 to 19, wherein the solid support is a flow cell or beads.

[0041] Embodiment 21 is a polynucleotide according to any one of Embodiments 2 to 7 or 9 to 20, wherein the polynucleotide comprises at least one insertion unit between a second insert sequence and an attached polynucleotide, the insert sequence comprising a 5' terminal insert sequence derived from a non-adjacent sequence of the target nucleic acid or derived from a target nucleic acid different from other insert sequences, and a 3' terminal ligation sequence comprising a read primer binding sequence, the read primer binding sequence being orthogonal to other read primer binding sequences.

[0042] Embodiment 22 is the polynucleotide described in Embodiment 21, wherein the polynucleotide is hybridized with its complement.

[0043] Embodiment 23 is a composition comprising a polynucleotide described in any one of Embodiments 1, 3 to 8, or 22, and its complement, wherein the complement comprises (a) a 5'-terminal complement comprising a first complementary read-primer binding sequence, (b) a complementary sequence of a second insert sequence located on the 3' side of the 5'-terminal complement, (c) a complementary linking sequence located on the 3' side of the complementary sequence of the second insert sequence, wherein the second insert sequence comprises (i) a second complementary read-primer binding sequence and (ii) a complementary hybridization sequence, (d) a complementary sequence of the first insert sequence located on the 3' side of the complementary linking sequence, and (e) a 3'-terminal complement.

[0044] Embodiment 24 is a composition comprising a polynucleotide described in any one of Embodiments 2 to 7 or 9 to 22 and its complement, wherein the complement is a 3'-terminal complement comprising a first complementary read-primer binding sequence, the first complementary read-primer binding sequence comprising a 3'-terminal complement orthogonal to the first and second read-primer binding sequences, a complement of a second insert sequence located on the 5' side of the 3'-terminal complement, and a complementary linking sequence located on the 5' side of the complement of the second insert sequence and comprising a second complementary read-primer binding sequence from 3' to 5', the second complementary read-primer binding sequence comprising a complementary linking sequence orthogonal to the first and second read-primer binding sequences and the first complementary read-primer binding sequence, a complement of a first insert sequence located on the 5' side of the complementary linking sequence, and a complementary attached polynucleotide comprising a complementary attached sequence at its 5' end.

[0045] Embodiment 25 is a configuration in which the first complementary read-primer binding sequence is a second adapter sequence, and The composition according to Embodiment 24, wherein, if present, the attached polynucleotide is complementary to the transposon terminal sequence, the complementary linking sequence is complementary to the linking sequence, and the complementary attached polynucleotide is complementary to the first adapter sequence and, if present, to the complement of the transposon terminal sequence.

[0046] Embodiment 26 is the composition according to Embodiment 24 or 25, wherein the polynucleotide is immobilized on a solid support via a first attached polynucleotide.

[0047] Embodiment 27 is the composition according to Embodiment 24 or 25, wherein the complement is immobilized on a solid support via complementary attached polynucleotides.

[0048] Embodiment 28 is a polynucleotide according to any one of Embodiments 2 to 7 or 9 to 22, or a composition according to any one of Embodiments 24 to 27, wherein the polynucleotide has the structure 3'-P7'-B15'-ME'-insert1-ME-HYB-ME'-insert2-ME-A14-P5-5', and ME' is a complement to a mosaic terminal sequence (e.g., SEQ ID NO: 3).

[0049] Embodiment 29 is the polynucleotide or composition described in Embodiment 28, wherein the complement of the polynucleotide has the structure 3'-P5'-A14'-ME'-insert2-ME-HYB'-ME'-insert1-ME-B15-P7-5'.

[0050] Embodiment 30 is a transposome complex comprising a transposase, a first transposon including a complement of a first read-primer binding sequence, the complement of the first read-primer binding sequence including a 3' portion including a transposon terminal sequence, a second transposon including a complement of a first adapter sequence and a 5' portion including a complement of a transposon terminal sequence, and a complement of a hybridization sequence.

[0051] Embodiment 31 is the transposome complex according to Embodiment 30, wherein the complement of the first adapter sequence is the B15 sequence.

[0052] Embodiment 32 is a transposome complex according to Embodiment 30 or 31, wherein the second transposon includes a complementary attachment sequence on the 5' side of the first read primer binding sequence, and optionally the complementary attachment sequence includes a P7 sequence.

[0053] Embodiment 33 is a transposome complex,

[0054] [ka] The transposome complex according to Embodiment 30 has the structure and ME is a mosaic terminal sequence such as SEQ ID NO: 6.

[0055] Embodiment 34 is a transposome complex according to any one of embodiments 30 to 33, wherein the transposome complex is immobilized on a bead via a first or second transposon.

[0056] Embodiment 35 is a first transposon comprising a transposase and an attached polynucleotide, wherein the attached polynucleotide comprises a 5' portion including an attached sequence, and the first transposon A transposome complex comprising a sposon, a 3' portion containing a second read-primer binding sequence, a 3' portion containing a transposon terminal sequence, an adapter, a second transposon containing a 5' portion containing a complement to the transposon terminal sequence, and a hybridization sequence.

[0057] Embodiment 36 is the transposome complex according to Embodiment 35, wherein the adapter is an A14 sequence.

[0058] Embodiment 37 is the transposome complex according to Embodiment 35 or 36, wherein the attachment sequence includes a P5 sequence.

[0059] Embodiment 38 is a transposome complex,

[0060] [ka] This is the transposome complex according to Embodiment 35, having the structure shown.

[0061] Embodiment 39 is a transposome complex according to any one of embodiments 35 to 38, wherein the transposome complex is immobilized on a solid support via a first or second transposon.

[0062] Embodiment 40 is a transposome complex according to any one of Embodiments 35 to 38, wherein the transposome complex is immobilized on a bead.

[0063] Embodiment 41 is a transposome complex according to any one of Embodiments 30 to 40, wherein the transposome complex is immobilized on an affinity binding partner on a solid support or beads via affinity elements connected to a linker attached to a first or second transposon.

[0064] Embodiment 42 is a composition or kit comprising two or more transposome complexes, such as the transposome complex described in any one of Embodiments 30 to 41.

[0065] Embodiment 43 is a composition or kit comprising a solid support, optionally the support being beads, a solid support, a component for generating a transposome complex, a transposase, and oligonucleotides for generating an oligonucleotide double helix, wherein the first oligonucleotide comprises a 3' transposon terminal sequence and a 5' first adapter sequence, and the second oligonucleotide comprises a 5' transposon terminal sequence and a 3' second adapter sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence. A composition or kit comprising: a component comprising oligonucleotides, the first and second adapter sequences of which are not the same; and first and second primer sets for adding an attachment sequence and a hybridization sequence to a fragment by PCR, wherein the first primer set comprises a primer for adding a hybridization sequence and a first attachment sequence to the fragment, and the second primer set comprises a primer for adding a complementary hybridization sequence and a second attachment sequence to the fragment, the first and second attachment sequences of which are not the same; and first and second primer sets.

[0066] Embodiment 44 comprises a first fork-type adapter complex and a second fork-type adapter An adapter composition or kit comprising a complex, wherein the first fork-type adapter complex is a complementary attachment polynucleotide comprising a 5' portion containing a complementary attachment sequence and a 3' portion containing an adapter, and a hybridization polynucleotide comprising (a) a 5' portion that hybridizes with a complement of a part of the adapter and (b) a complement of the hybridization sequence, wherein the complement of the hybridization sequence is not complementary to the complementary attachment polynucleotide, and The adapter composition or kit comprises a hybridization polynucleotide and a second fork-type adapter complex, the second fork-type adapter complex comprising an attached polynucleotide comprising a 5' portion containing an attached sequence and a 3' portion containing an adapter, and a hybridization polynucleotide comprising (a) a 5' portion that hybridizes with a complement of a part of the adapter and (b) a hybridization sequence that is not complementary to the attached polynucleotide.

[0067] Embodiment 45 is the adapter composition or kit described in Embodiment 44, wherein the attachment sequence includes a P5 primer sequence and the complementary attachment sequence includes a P7 primer sequence.

[0068] Embodiment 46 is the adapter composition or kit according to Embodiment 44 or 45, wherein the complementary attached polynucleotide comprises the B15 sequence and the hybridization polynucleotide comprises the A14 sequence.

[0069] Embodiment 47 is a first fork-type adapter complex,

[0070] [ka] The structure is such that the second fork-type adapter complex is

[0071] [ka] The adapter composition or kit described in Embodiment 46 has the structure of the above.

[0072] Embodiment 48 is an adapter composition or kit according to any one of Embodiments 44 to 47, wherein the adapter complex contains a methylated nucleotide (for example, a methylated cytosine).

[0073] Embodiment 49 is a method for generating a linked nucleic acid sequencing template, comprising: attaching a first read-primer binding sequence to the 3' end of a first insert sequence derived from a first target nucleic acid; attaching a hybridization sequence to the 5' end of the first insert sequence; attaching a complement of the hybridization sequence to the 3' end of a second insert sequence derived from a discontinuous region of the first target nucleic acid or a second target nucleic acid; annealing the hybridization sequence to its complement to form a hybridized adduct; and synthesizing a fully double-stranded linked nucleic acid sequencing template from the hybridized adduct, wherein the region between the first insert sequence and the second insert sequence is a second read-primer binding sequence comprising a hybridization sequence and orthogonal to the first read-primer binding sequence, the second read-primer binding sequence This method includes a lymer-binding sequence, thereby generating a concatenated nucleic acid sequencing template.

[0074] Embodiment 50 is the method of Embodiment 48, wherein the attachment of a first lead primer binding sequence and the attachment of a hybridization sequence are performed by contacting one or more target nucleic acids with a transposome complex under conditions suitable for tagmentation.

[0075] Embodiment 51 is the method of Embodiment 49 or 50, which includes contacting one or more target nucleic acids with a transposome complex under conditions suitable for tagmentation, by attaching a complementary hybridization sequence to the 3' end of a discontinuous region of a first target nucleic acid or a second insert sequence derived from a second target nucleic acid.

[0076] Embodiment 52 is the method of Embodiment 49, comprising attaching a first read primer binding sequence to the 3' end of a first insert sequence and attaching a hybridization sequence to the 5' end of a first insert sequence, and contacting one or more target nucleic acids with a first fork-type adapter complex described in any one of Embodiments 44 to 48 under conditions suitable for ligation of the ends of an adapter complex fragment, thereby forming a fragment ligated with the first adapter complex at both ends and a fragment ligated with the second adapter complex at both ends, and denaturing the ligated fragment.

[0077] Embodiment 53 is the method of Embodiment 49 or 50, which includes attaching the complement of the hybridization sequence to the 3' end of the second insert sequence, and contacting one or more target nucleic acids with the second fork-type adapter complex under conditions suitable for ligation to the ends of the adapter complex fragments to form a fragment ligated with the first adapter complex at both ends and a fragment ligated with the second adapter complex at both ends, and denaturing the ligated fragments.

[0078] Embodiment 54 is a method for generating a linked nucleic acid sequencing template, comprising contacting a first sample containing a first target nucleic acid with a first transposome complex and a second transposome complex, each transposome complex comprising a transposase, a first transposon comprising a 3' portion containing a transposon terminal sequence and a 5' portion containing an adapter sequence, and a second transposon comprising a complement to the transposon terminal sequence and a 5' portion that hybridizes thereto, wherein the adapter sequence in the first transposome complex is the complement to the first adapter sequence, and the adapter sequence in the second transposome complex is the second adapter sequence, and the first target nucleic acid is fragmented under conditions sufficient to generate a first tagged product comprising an insert sequence derived from the first target nucleic acid, with one end tagged by the transposon of the first transposome complex and the other end tagged by the transposon of the second transposome complex. The process involves: optionally adding a complementary attachment sequence to the 3' end of the first tagged product and adding a complementary hybridization sequence to the 5' end of the first tagged product by polymerase chain reaction to form a first modified tagged product; and contacting a second sample containing a second target nucleic acid with a transposome complex, under conditions sufficient to fragment the second target nucleic acid and generate a second tagged product containing an insert sequence derived from the second target nucleic acid, with one end tagged by a transposon of the first transposome complex and the other end tagged by a transposon of the second transposome complex; optionally adding an attachment sequence to the 3' end of the second tagged product and adding a hybridization sequence to the 5' end of the second tagged product by polymerase chain reaction to form a second modified tagged product; and connecting the hybridization sequence of the first modified tagged product to the second modified tagged A method comprising: annealing to the complement of a hybridization sequence in the product to form a hybridized adduct; and synthesizing a fully double-stranded linked nucleic acid sequencing template from the hybridized adduct, wherein the linked nucleic acid sequencing template comprises: (a) a first read-primer-binding sequence located at the 3' end of an insert sequence derived from a second target nucleic acid, comprising a first adapter sequence and a complement of a transposon terminal sequence; and (b) a second read-primer-binding sequence located between the two insert sequences, comprising a transposon terminal sequence and a hybridization sequence, wherein the first read-primer-binding sequence is orthogonal to the second read-primer-binding sequence.

[0079] Embodiment 55 is a method for generating a linked nucleic acid sequencing template, comprising contacting a first sample containing a first target nucleic acid with a first transposome complex, the first transposome complex comprising a transposase, a first transposon comprising a 3' portion containing a transposon terminal sequence, and a 5' portion containing a complement of an attachment sequence and a first adapter sequence, and a second transposon comprising a complement of the transposon terminal sequence and a 5' portion that hybridizes thereto, thereby fragmenting the first target nucleic acid so that each end is the first transpo The process involves carrying out under conditions sufficient to generate a first tagged product containing an insert sequence derived from a first target nucleic acid, tagged by a transposon of the somosome complex; optionally adding a hybridization sequence complement to the 5' end of the first tagged product by polymerase chain reaction to form a first modified tagged product; and contacting a second sample containing a second target nucleic acid with a second transpososome complex, wherein the second transpososome complex comprises a transposase, a 3' portion containing the transposon terminal sequence, and The first transposon comprises a 5' portion containing two adapter sequences and a complementary attachment sequence, and the second transposon comprises a 5' portion containing a complement to the transposon terminal sequence and a hybridizing portion thereof, and the second target nucleic acid is fragmented under conditions sufficient to produce a second tagged product containing an insert sequence derived from the second target nucleic acid, with each end tagged by the transposon of the second transpososome complex, and optionally a polymerase chain reaction is used to add a hybridization sequence to the 5' end of the second tagged product. The method involves adding a complement to form a second modified tagged product, annealing the hybridization sequence of the first modified tagged product to the complement of the hybridization sequence in the second modified tagged product to form a hybridized adduct, and synthesizing a fully double-stranded linked nucleic acid sequencing template from the hybridized adduct, wherein the linked nucleic acid sequencing template comprises (a) a first read-primer binding sequence located at the 3' end of an insert sequence derived from a second target nucleic acid, and a first adapter sequence,A method comprising: (b) a first read-primer binding sequence containing a complement to the transposon terminal sequence; and (b) a second read-primer binding sequence located between two insert sequences, the second read-primer binding sequence containing the transposon terminal sequence and a hybridization sequence, wherein the first read-primer binding sequence is orthogonal to the second read-primer binding sequence.

[0080] Embodiment 56 is the method of Embodiment 54 or 55, wherein the transposome complex is immobilized on a solid support.

[0081] Embodiment 57 is a method for generating a linked nucleic acid sequencing template, comprising: (a) contacting a first double-stranded polynucleotide containing a first target nucleic acid with a first restriction enzyme, and (ii) contacting a second double-stranded polynucleotide containing a second target nucleic acid with a second restriction enzyme to produce first and second polynucleotides having a compatibility overhang, wherein the restriction enzyme is selected from type II, type IIS, type IIP, and type IIT restriction enzymes; and (b) using a ligase to provide the compatibility overhang of the first and second polynucleotides. This method includes attaching a coating.

[0082] Embodiment 58 is the method of Embodiment 57, wherein, before the contact step, (a) a first restriction enzyme cleavage site is optionally attached to a first target nucleic acid using an adapter, and a first double-stranded polynucleotide is generated by primer extension, and (b) a second restriction enzyme cleavage site is optionally attached to a second target nucleic acid using an adapter, and a second double-stranded polynucleotide is generated by primer extension.

[0083] Embodiment 59 is a method for generating a linked nucleic acid sequencing template, comprising: (a) shearing or digesting a first nucleic acid source and a second nucleic acid source to generate a first library of nucleic acid fragments and a second library of nucleic acid fragments, respectively; (b) attaching a first adapter to each nucleic acid fragment derived from the first nucleic acid source and attaching a second adapter to each nucleic acid fragment derived from the second nucleic acid source, comprising: (i) contacting the nucleic acid fragments with a first polymerase to generate nucleic acid fragments having blunt ends; (ii) phosphorylating the 5'-hydroxyl of the nucleic acid fragments using a kinase; and (iii) the second A method comprising: (iv) adding 3' adenine to a nucleic acid fragment using polymerase; (iv) ligating a first adapter to each nucleic acid fragment of the first library and a second adapter to each nucleic acid fragment of the second library; and (c) optionally mixing and annealing the first and second nucleic acid libraries by PCR, wherein (i) the nucleic acid denatures at high temperatures, (ii) the A and A' sequences hybridize with each other at lower temperatures, and (d) optionally synthesizing a fully double-stranded linked nucleic acid sequencing template by PCR.

[0084] Embodiment 60 is a method according to any one of Embodiments 54 to 59, wherein the method comprises sequencing a concatenated nucleic acid sequencing template.

[0085] Embodiment 61 is a method for sequencing a linked nucleic acid sequencing template, comprising: sequencing a first insert sequence of a polynucleotide described in any one of Embodiments 1 to 22 by initiating sequencing using a first read sequencing primer complementary to a first read primer binding sequence; and sequencing a second insert sequence by initiating sequencing using a second read sequencing primer complementary to a second read primer binding sequence.

[0086] Embodiment 62 is the method of Embodiment 61, further comprising: sequencing a complement of a second insert sequence by initiating sequencing using a first complementary read sequencing primer complementary to a first complementary read primer binding sequence; and sequencing a complement of a first insert sequence by initiating sequencing using a second complementary read sequencing primer complementary to a second complementary read primer binding sequence.

[0087] Embodiment 63 is a method according to any one of Embodiments 49 to 59, wherein a sample containing a target double-stranded nucleic acid is partitioned into several different compartments, and the generation of a concatenated nucleic acid sequencing template is performed within each of the different compartments.

[0088] Embodiment 64 is a polynucleotide comprising: (a) a 5'-terminal polynucleotide containing a first read sequencing primer sequence; (b) an insert sequence derived from a target nucleic acid, located at the 3' end of the 5'-terminal polynucleotide; (c) a hybridization sequence located at the 3' end of the insert sequence; (d) a copy of the insert sequence located at the 3' end of the hybridization sequence; and (e) a second read sequencing primer sequence. It is a polynucleotide containing a 3' terminal polynucleotide that includes the complementary of .

[0089] Embodiment 65 is a polynucleotide comprising: (a) a 5'-terminal polynucleotide containing a first read sequencing primer sequence; (b) an insert sequence derived from a target nucleic acid, located on the 3' side of the 5'-terminal polynucleotide, which is a first insert sequence; (c) a hybridization sequence located on the 3' side of the insert sequence; (d) a second insert sequence located on the 3' side of the hybridization sequence; and (e) a 3'-terminal polynucleotide containing a complement of the second read sequencing primer sequence.

[0090] Embodiment 66 is a polynucleotide according to Embodiment 64 or 65, wherein the insert sequence contains 40 to 400 nucleotides, and optionally, the insert sequence contains 1000 or fewer nucleotides.

[0091] Embodiment 67 is a polynucleotide according to any one of Embodiments 64 to 66, wherein the hybridization sequence comprises 10 to 30 nucleotides, and optionally, one or more nucleotides in the hybridization sequence are lock nucleic acids.

[0092] Embodiment 68 is a polynucleotide according to any one of Embodiments 64 to 67, wherein the first read sequencing primer sequence and the second read sequencing primer sequence are different.

[0093] Embodiment 69 is a polynucleotide according to any one of Embodiments 64 to 68, wherein the first read sequencing primer sequence and the second read sequencing primer sequence each include the A14 sequence or the B15 sequence, or their complements.

[0094] Embodiment 70 is a polynucleotide according to any one of Embodiments 64 to 69, wherein the 3' terminal polynucleotide contains a complement of the P5 primer sequence (P5') and the 5' terminal polynucleotide contains a P7 primer sequence (P7(SEQ ID NO: 8)), or the 3' terminal polynucleotide contains a complement of the P7 primer sequence (P7') and the 5' terminal polynucleotide contains a P5 primer sequence (P5(SEQ ID NO: 7)).

[0095] Embodiment 71 is a polynucleotide according to any one of Embodiments 64 to 70, wherein the 3'-terminal polynucleotide and / or the 5'-terminal polynucleotide each independently comprises at least one of the following: an adapter, a barcode sequence, a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence.

[0096] Embodiment 72 is a polynucleotide according to any one of Embodiments 64 to 71, wherein the polynucleotide is immobilized on a solid support.

[0097] Embodiment 73 is the polynucleotide according to Embodiment 72, wherein the polynucleotide is immobilized on a solid support via a 5' terminal polynucleotide.

[0098] Embodiment 74 is the polynucleotide according to Embodiment 73, wherein the polynucleotide is immobilized on a solid support via the binding of the affinity moiety on the 5' terminal polynucleotide to a binding moiety on the surface of the solid support.

[0099] Embodiment 75 is a polynucleotide according to any one of Embodiments 64 to 74, wherein the affinity moiety is attached to the 5' terminal polynucleotide via a linker.

[0100] Embodiment 76 has an affinity portion that is biotin, desthiobiotin, or dualbiotin This is a polynucleotide according to any one of embodiments 64 to 75.

[0101] Embodiment 77 is a polynucleotide according to any one of Embodiments 64 or 66-76, wherein the polynucleotide has the structure 5'-P5-A14-insert-HYB-insert-B15'-P7'-3' or 5'-P7-B15-insert-HYB'-insert-A14'-P5'-3', where HYB is a hybridization sequence and HYB' is a complement of the hybridization sequence.

[0102] Embodiment 78 is a polynucleotide according to any one of Embodiments 65 to 77, wherein the polynucleotide has the structure 5'-P5-A14-insert1-HYB-insert2-B15'-P7'-3' or 5'-P7-B15-insert1-HYB'-insert2-A14'-P5'-3', where HYB is a hybridization sequence and HYB' is a complement of the hybridization sequence.

[0103] Embodiment 79 is a composition comprising a polynucleotide according to any one of Embodiments 64 to 78, which is hybridized to its complement.

[0104] Embodiment 80 is a composition comprising any one of Embodiments 64 to 78 or the composition according to Embodiment 79, immobilized on the surface of a solid support, wherein the affinity moiety is biotin, desthiobiotin, or dualbiotin, and the binding moiety is avidin or streptavidin.

[0105] Embodiment 81 is the composition according to Embodiment 80, wherein the solid support is a bead, a slide, a container wall, a flow cell, or a nanowell contained in a flow cell.

[0106] Embodiment 82 is a fork-type adapter comprising two polynucleotide chains, (a) a first chain containing a sequencing primer sequence, and (b) a second chain containing a 3' hybridization sequence or its complement, wherein the 3' end of the first chain is fully or partially complementary to the 5' end of the second chain.

[0107] Embodiment 83 is a fork-type adapter according to Embodiment 82, wherein a hybridization sequence or its complement is bound to a blocking oligonucleotide that is completely or partially complementary to the hybridization sequence or its complement.

[0108] Embodiment 84 is a fork-type adapter according to Embodiment 83, wherein a hybridization sequence or its complement is bound to a blocking oligonucleotide that is perfectly complementary to the hybridization sequence or its complement.

[0109] Embodiment 85 is a fork-type adapter according to any one of Embodiments 82 to 84, wherein the first and / or second strands further comprises at least one of the following: an adapter, a barcode sequence, a unique molecular identifier (UMI) sequence, a sample index sequence, a capture sequence, or a cleavage sequence.

[0110] Embodiment 86 is a fork-type adapter according to any one of embodiments 82 to 85, wherein the first and / or second strands further comprise a P7 or P5 primer sequence, or a complement thereof.

[0111] Embodiment 87 is a fork-type adapter according to any one of Embodiments 82 to 86, wherein the sequencing primer sequence includes the B15 sequence (SEQ ID NO: 6) or the A14 sequence (SEQ ID NO: 4), or their complements.

[0112] Embodiment 88 is a fork-type adapter according to any one of embodiments 82 to 87, wherein the first chain includes a 5' affinity element that can be bonded to an affinity bonding partner on a solid support or beads.

[0113] Embodiment 89 is a fork-type adapter according to Embodiment 88, wherein the affinity elements are connected via linkers attached to the first chain.

[0114] Embodiment 90 is a composition or kit comprising two fork adapters as described in any one of Embodiments 82 to 89, wherein (a) the first fork adapter comprises a first strand containing a first read sequencing primer sequence and a second strand containing a complementary hybridization sequence, and (b) the second fork adapter comprises a first strand containing a second read sequencing primer sequence and a second strand containing a hybridization sequence.

[0115] Embodiment 91 is a composition or kit according to Embodiments 44-48 or 90, wherein one or both fork-type adapters contain a blocking oligonucleotide.

[0116] Embodiment 92 is a method for generating one or more linked nucleic acid sequencing templates, comprising: (a) contacting a sample comprising double-stranded nucleic acid fragments, each comprising an insert prepared from a target nucleic acid, with a composition or kit comprising two fork-type adapters, one or both of which comprise a blocking oligonucleotide, and optionally, the first read sequencing adapter sequence comprises a first read primer binding sequence; (b) ligating the fork-type adapters to the double-stranded fragments to prepare tagged double-stranded fragments; (c) immobilizing the tagged double-stranded fragments on a solid support; and (d) immobilization The method comprises (2) denaturing a tagged double-stranded fragment fixed to generate a single-stranded fragment, and (3) denaturing a blocking oligonucleotide to deblock a hybridization sequence and a complement of the hybridization sequence, (4) hybridizing the two fixed single-stranded fragments to form a crosslink by linking the hybridization sequence in the first fragment to the complement of the hybridization sequence in the second fragment, and (5) extending from the 3' end of each single-stranded fragment to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both fixed single-stranded fragments.

[0117] Embodiment 93 is a method for generating one or more linked nucleic acid sequencing templates, comprising (a) contacting a sample containing a double-stranded target nucleic acid with two pools of transposome complexes in solution, wherein the first pool of transposome complexes comprises (i) a transposase, (ii) a first transposon comprising a 3' transposon terminal sequence and a first read sequencing adapter sequence, and (iii) a second transposon comprising a 3' complement of a 5' sequence and a hybridization sequence that is completely or partially complementary to the 3' transposon terminal sequence, and the second pool of transposome complexes comprises (i) a transposase, and (ii) a 3' transposon terminal sequence and a second read sequencing adapter sequence (iii) comprising a first transposon containing a 5' sequence and a 3' hybridization sequence that are completely or partially complementary to the 3' transposon terminal sequence, wherein one or both of the second transposons contain a blocking oligonucleotide; (b) tagging a double-stranded nucleic acid to produce a tagged double-stranded fragment; (c) releasing the transpososome complex from the double-stranded fragment; (d) extending and ligating the double-stranded fragment; (e) immobilizing the tagged double-stranded fragment onto a solid support; and (f) (1) the immobilized tagged double-stranded fragment to produce an immobilized single-stranded fragment, and (2) the hybridization sequence A method comprising: (g) denaturing a blocking oligonucleotide to unblock the complements of the sequence and hybridization sequence; (h) hybridizing two fixed single-stranded fragments to form a crosslink by linking the hybridization sequence in the first fragment to the complement of the hybridization sequence in the second fragment; and (h) extending from the 3' end of each single-stranded fragment to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both fixed single-stranded fragments.

[0118] Embodiment 94 is the method according to Embodiment 92 or 93, wherein the denaturation is carried out by increasing the temperature, changing the pH, and / or adding one or more chaotropic agents.

[0119] Embodiment 95 is the method of Embodiment 94, wherein the temperature increase is from 45°C to 55°C to 85°C to 95°C, and optionally, the temperature increase is from 50°C to 90°C.

[0120] Embodiment 96 is a method according to any one of Embodiments 92 to 95, wherein one or more chaotropic agents include formamide and / or NaOH.

[0121] Embodiment 97 is a method according to any one of Embodiments 92 to 96, wherein fixation is by binding an affinity portion, which is (1) included in a first and / or second fork-shaped adapter, or (2) included in a tag derived from a second transposome, to one or more binding portions on the surface of a solid support.

[0122] Embodiment 98 is the method according to any one of Embodiments 92 to 97, wherein the affinity moiety is biotin, desthiobiotin, or dualbiotin, and the binding moiety is avidin or streptavidin.

[0123] Embodiment 99 is a method according to any one of Embodiments 92 to 98, wherein one or more additional modification, hybridization, and extension steps are performed.

[0124] Embodiment 100 is a method according to any one of Embodiments 92 to 99, wherein the first single-stranded fragment includes an insert, and the second single-stranded fragment includes an insert that is a complement of the insert contained in the first fragment.

[0125] Embodiment 101 is a method according to any one of Embodiments 92 to 100, wherein the first single-stranded fragment includes an insert, and the second fragment includes an insert that is not a complement of the insert contained in the first fragment.

[0126] Embodiment 102 is the method according to any one of Embodiments 92 to 101, wherein hybridization occurs between single-stranded fragments prepared from double-stranded fragments, comprising (1) a first fork-shaped adapter ligated to one end of each fragment and a second fork-shaped adapter ligated to the other end of each fragment, or (2) a tag derived from a second transposon of the first transpososome complex at one end of each fragment and a tag derived from a second transposon of the second transpososome at the other end of each fragment.

[0127] Embodiment 103 is a method according to any one of Embodiments 92 to 102, wherein two fixed single-stranded fragments hybridize to each other and do not form crosslinks, without the hybridization sequence in the first fragment binding to the complement of the hybridization sequence in the second fragment.

[0128] Embodiment 104 hybridizes two fixed single-stranded fragments to form a crosslink. The method according to Embodiment 103 is such that what happens does not occur between single-stranded fragments prepared from double-stranded fragments, which include (1) the same fork-shaped adapters ligated to both ends of each fragment, or (2) tags derived from the same transposome complex at both ends of each fragment.

[0129] Embodiment 105 is a method for generating one or more linked nucleic acid sequencing templates, comprising: (a) partitioning a sample containing a target double-stranded nucleic acid into a plurality of different compartments; (b) preparing fragments each containing an insert derived from the double-stranded nucleic acid in the plurality of different compartments; (c) contacting the plurality of different compartments with a composition or kit comprising two fork-type adapters as described in Embodiment 91, wherein one or both fork-type adapters contain a blocking oligonucleotide; (d) ligating the fork-type adapters to the double-stranded fragments to prepare tagged double-stranded fragments in the plurality of different compartments; and (e) (1) single-stranded fragments The method comprises (2) denaturing a tagged double-stranded fragment fixed for generation, and (3) denaturing a blocking oligonucleotide to unblock hybridization sequences and complementary hybridization sequences in a plurality of different compartments, (4) hybridizing two single-stranded fragments in the same compartment to form a crosslink by linking the hybridization sequence in the first fragment to the complementary hybridization sequence in the second fragment, and (5) extending from the 3' end of each single-stranded fragment to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both single-stranded fragments in the same compartment.

[0130] Embodiment 106 is the method of Embodiment 105, wherein the target double-stranded nucleic acid includes a double-stranded DNA fragment, and the preparation of the fragment is the preparation of a small fragment of the double-stranded DNA fragment.

[0131] Embodiment 107 is the method according to Embodiments 63, 105, or 107, wherein the compartment is a well, a tube, or a droplet.

[0132] Embodiment 108 is a method according to any one of Embodiments 105 to 107, wherein the denaturation is carried out by increasing the temperature, changing the pH, and / or adding one or more chaotropic agents.

[0133] Embodiment 109 is the method of Embodiment 108, wherein the temperature increase is from 45°C to 55°C to 85°C to 95°C, and optionally, the temperature increase is from 50°C to 90°C.

[0134] Embodiment 110 is the method according to Embodiment 108 or 109, wherein one or more chaotropic agents include formamide and / or NaOH.

[0135] Embodiment 111 is a method according to any one of Embodiments 105 to 110, wherein one or more additional modification, hybridization, and extension steps are performed.

[0136] Embodiment 112 is the method according to any one of Embodiments 105 to 111, wherein the first single-stranded fragment includes an insert, and the second single-stranded fragment includes an insert that is a complement of the insert contained in the first fragment.

[0137] Embodiment 113 is a method according to any one of Embodiments 105 to 111, wherein the first single-stranded fragment includes an insert, and the second fragment includes an insert that is not a complement of the insert contained in the first fragment.

[0138] Embodiment 114 hybridizes a first fork-shaped adapter ligated to one end of each segment and a second fork-shaped adapter ligated to the other end of each segment. The method according to any one of embodiments 105 to 113, which occurs between single-stranded fragments prepared from double-stranded fragments, including a fork-type adapter.

[0139] Embodiment 115 is a method according to any one of Embodiments 105 to 114, wherein the single-stranded fragments do not hybridize with each other, in a manner in which the hybridization sequence in the first fragment does not bind to the complementary hybridization sequence in the second fragment.

[0140] Embodiment 116 is the method of Embodiment 115, wherein hybridization of two single-stranded fragments does not occur between single-stranded fragments prepared from double-stranded fragments, which include the same fork-shaped adapters ligated to both ends of each fragment.

[0141] Embodiment 117 is a method according to any one of Embodiments 63 or 105-116, wherein compartmentalization involves diluting the sample such that most compartments either contain one target double-stranded nucleic acid or do not contain a target double-stranded nucleic acid.

[0142] Embodiment 118 is the method of Embodiment 117, wherein the inserts contained in the same linked sequencing template are prepared from the same target nucleic acid.

[0143] Embodiment 119 is a method according to any one of Embodiments 63 or 105-118, wherein partitioning separates the most different haplotypes into different compartments, and the method is used for haplotype phasing.

[0144] Embodiment 120 is the method of Embodiment 119, wherein haplotype fading does not require a barcode.

[0145] Embodiment 121 is a solid support comprising two pools of immobilized transposome complexes, (a) a first pool of transposome complexes comprising: (i) a transposase; (ii) a first transposon comprising a 3' transposon terminal sequence, a first read sequencing adapter sequence, and a 5' affinity moiety; (iii) a second transposon comprising a 5' sequence that is completely or partially complementary to the 3' transposon terminal sequence, and a 3' complement of the hybridization sequence; and (b) The second pool of the nsposome complex comprises a solid support, each first transposon being fixed by the binding of its 5' affinity moiety to a binding site on the surface of the solid support.

[0146] Embodiment 122 is a solid support according to Embodiment 121, wherein the first or second pool of transposome complexes comprises the transposome complex described in any one of Embodiments 30 to 42, and the first read sequencing adapter sequence comprises the first read primer binding sequence.

[0147] Embodiment 123 is a solid support according to Embodiment 121 or 122, wherein the first and / or second pool of transposome complexes comprises homodimers and / or heterodimers.

[0148] Embodiment 124 is the solid support according to Embodiment 122 or 123, wherein the solid support is a bead, a slide, a container wall, a flow cell, or a nanowell contained in a flow cell.

[0149] Embodiment 125 is a solid support according to any one of Embodiments 121 to 124, wherein one or more transposons include an index sequence and / or UMI.

[0150] Embodiment 126 is a solid support according to Embodiment 125, wherein the first transposon contained in the first pool of the transposome complex and / or the first transposon contained in the second pool of the transposome complex include a sample index.

[0151] Embodiment 127 is the solid support described in Embodiment 126, wherein both the first transposon contained in the first pool of the transposome complex and the first transposon contained in the second pool of the transposome complex contain a sample index.

[0152] Embodiment 128 is a solid support according to any one of Embodiments 121 to 127, wherein the second transposon contained in the first pool of the transposome complex and / or the second transposon contained in the second pool of the transposome complex includes a sample index and / or a unique molecular identifier (UMI).

[0153] Embodiment 129 is a solid support according to Embodiment 128, wherein both the second transposon contained in the first pool of the transposome complex and the second transposon contained in the second pool of the transposome complex contain a sample index.

[0154] Embodiment 130 is a solid support according to Embodiment 128 or Embodiment 129, wherein both the second transposon in the first pool of the transposome complex and the second transposon in the second pool of the transposome complex contain UMI.

[0155] Embodiment 131 is a method for generating one or more double-stranded linked nucleic acid sequencing templates, comprising: (a) applying a sample containing double-stranded nucleic acids immobilized on a solid support; (b) tagging the double-stranded nucleic acids to generate tagged double-stranded fragments containing inserts derived from the double-stranded nucleic acids, wherein the double-stranded fragments are immobilized on the solid support by binding of 5' affinity moieties to binding moieties on the surface of the solid support; (c) releasing transposome complexes from the double-stranded fragments; and (d) two The method comprises (e) extending and ligating a double-stranded fragment, (f) denaturing a double-stranded fragment into a single-stranded fragment, wherein the single-stranded fragment containing the 5' affinity moiety remains fixed on a solid support, (g) hybridizing a hybridization sequence contained in a first fixed single-stranded fragment to a complement of a hybridization sequence contained in a second fixed single-stranded fragment, thereby forming a crosslink, and (g) extending and generating a double-stranded linked nucleic acid sequencing template.

[0156] Embodiment 132 is the same method as in Embodiment 131, wherein the transposome complex is released from the double-stranded fragment using SDS.

[0157] Embodiment 133 is the method according to Embodiment 131 or 132, wherein hybridization includes cooling a solid support and / or applying a hybridization buffer.

[0158] Embodiment 134 is the method of Embodiment 133, wherein cooling includes lowering the temperature of the solid support to 60°C or below.

[0159] Embodiment 135 is the method according to Embodiment 133 or 134, wherein the hybridization buffer contains a high salt concentration, and optionally the high salt concentration is 750 mM NaCl. .

[0160] Embodiment 136 is a method according to any one of Embodiments 131 to 135, wherein the modification includes heating the solid support or applying a chemical modifier.

[0161] Embodiment 137 is the method of Embodiment 136, wherein the modification includes raising the temperature of the solid support to 90°C or higher.

[0162] Embodiment 138 is a method according to any one of Embodiments 131 to 137, wherein the extension process includes providing polymerase, dNTPs, and an extension buffer.

[0163] Embodiment 139 is the method according to any one of Embodiments 131 to 138, further comprising the additional step of hybridizing, extending, and generating a double-stranded linked nucleic acid sequencing template.

[0164] Embodiment 140 is the method according to Embodiments 131-139, wherein hybridization of a hybridization sequence contained in a first fixed single-stranded fragment to a complement of a hybridization sequence contained in a second fixed single-stranded fragment occurs only when the first and second fragments are in close proximity to each other on the surface of a solid support that is closer than the longer length of the first or second fragment.

[0165] Embodiment 141 is a method according to Embodiments 131 to 140, wherein a first fixed fragment and a second fixed fragment are fixed in close proximity on a solid support, and by being in close proximity, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more nucleotides contained in the hybridization sequence contained in the first fixed fragment can bind to nucleotides contained in the complement of the hybridization sequence contained in the second fixed fragment.

[0166] Embodiment 142 is a method according to any one of Embodiments 131 to 141, wherein the first fixed fragment and the second fixed fragment are fixed to each other within 20 to 500 nanometers on the surface of a solid support.

[0167] Embodiment 143 is a method according to any one of Embodiments 93-121 or 131-142, wherein the sample comprises multiple double-stranded nucleic acids.

[0168] Embodiment 144 is the method of Embodiment 143, wherein both the first and second immobilized fragments are prepared from the same double-stranded nucleic acid, and the double-stranded linked nucleic acid sequencing template includes two inserts derived from the same double-stranded nucleic acid.

[0169] Embodiment 145 is the method according to Embodiment 144, wherein the two inserts are derived from two adjacent sequences contained in the same double-stranded nucleic acid.

[0170] Embodiment 146 is the method according to Embodiment 144, wherein the two inserts are derived from two adjacent sequences contained in the same double-stranded nucleic acid, and the adjacent sequences are separated by 100 or fewer nucleotides, 200 or fewer nucleotides, 300 or fewer nucleotides, 400 or fewer nucleotides, 500 or fewer nucleotides, 700 or fewer nucleotides, or 1,000 or fewer nucleotides in the double-stranded nucleic acid.

[0171] Embodiment 147 includes a solid support region comprising a plurality of double-stranded linked nucleic acid sequencing templates that share a common insert sequence derived from adjacent sequences contained in the same double-stranded nucleic acid. This is the method of form 146.

[0172] Embodiment 148 is a double-stranded linked nucleic acid sequencing template prepared by the method of any one of Embodiments 131 to 147, wherein the template structure comprises (a) 5'-P5-i5-A14-ME-insert1-ME'-HYB-ME-insert2-ME'-B15'-i7'-P7'-3', (b) 5'-P5-A14-ME-insert1-ME'-i6-HYB-i8'-ME-insert2-ME'-B15'-P7'-3', or (c) 5'-P5-i5-A14-ME-insert1-ME'-i6-HYB-i8'-ME-insert2-ME'-B15'-i7'-P7'-3', or their complements.

[0173] Embodiment 149 is a method according to any one of Embodiments 131 to 148, further comprising (a) releasing a double-stranded linked nucleic acid sequencing template from a solid support, and (b) sequencing the template to determine the insert sequence contained in the template.

[0174] Embodiment 150 is the method of Embodiment 149, wherein the release includes enzymatic digestion or chemical cleavage.

[0175] Embodiment 151 is the method of Embodiment 149 or 150, further comprising amplifying the template after liberation and before sequencing.

[0176] Embodiment 152 is a method for generating one or more linked nucleic acid sequencing templates, comprising (a) partitioning a sample containing a target double-stranded nucleic acid into a plurality of different compartments, and (b) tagmenting the double-stranded nucleic acid to generate tagged double-stranded fragments containing inserts derived from the double-stranded nucleic acid in the plurality of different compartments, wherein the tagmentation is carried out using two pools of transposome complexes, the first pool of transposome complexes comprising (i) a transposase, (ii) a first transposon containing a 3' transposon terminal sequence and a first read sequencing adapter sequence, and (iii) a second transposon containing a 5' sequence and a 3' complement of a hybridization sequence that is completely or partially complementary to the 3' transposon terminal sequence, and the transposome complex The second pool is a method comprising (i) a transposase, (ii) a first transposon comprising a 3' transposon terminal sequence and a second read sequencing adapter sequence, and (iii) a second transposon comprising a 5' sequence and a 3' hybridization sequence that are fully or partially complementary to the 3' transposon terminal sequence, (b) denaturing the tagged double-stranded fragment to generate a single-stranded fragment, (c) hybridizing two single-stranded fragments within the same compartment by linking the hybridization sequence in the first fragment to the complement of the hybridization sequence in the second fragment, and (d) extending from the 3' end of each single-stranded fragment to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both single-stranded fragments.

[0177] Embodiment 153 is the method of Embodiment 152, wherein the double-stranded linked nucleic acid sequencing template is generated only from the hybridization of two single-stranded fragments located in the same compartment.

[0178] Embodiment 154 is the method according to Embodiment 152 or 153, wherein a hybridization sequence and / or its complement is bound to a blocking oligonucleotide that is completely or partially complementary to the hybridization sequence or its complement, and denaturation includes denaturing the blocking oligonucleotide to deblock the hybridization sequence and / or its complement.

[0179] Embodiment 155 is a method according to any one of Embodiments 152 to 154, wherein the transposome complex is present in a solution.

[0180] Embodiment 156 is a method according to any one of Embodiments 152 to 155, wherein the compartment is a well, a tube, or a droplet.

[0181] Embodiment 157 is a method according to any one of Embodiments 152 to 156, wherein the denaturation is carried out by increasing the temperature, changing the pH, and / or adding one or more chaotropic agents.

[0182] Embodiment 158 ​​is the method of Embodiment 157, wherein the temperature increase is from 45°C to 55°C to 85°C to 95°C, and optionally, the temperature increase is from 50°C to 90°C.

[0183] Embodiment 159 is the method according to Embodiment 157 or 158, wherein one or more chaotropic agents include formamide and / or NaOH.

[0184] Embodiment 160 is a method according to any one of Embodiments 152 to 159, wherein one or more additional modification, hybridization, and extension steps are performed.

[0185] Embodiment 161 is a method according to any one of Embodiments 152 to 160, wherein the first single-stranded fragment includes an insert, and the second fragment includes an insert that is not a complement of the insert contained in the first fragment.

[0186] Embodiment 162 is a method according to any one of Embodiments 105 to 161, wherein compartmentalization involves diluting the sample such that most compartments either contain one target double-stranded nucleic acid or do not contain a target double-stranded nucleic acid.

[0187] Embodiment 163 is the method of Embodiment 162, wherein the inserts contained in the same linked sequencing template are prepared from the same target nucleic acid.

[0188] Embodiment 164 is a method according to any one of Embodiments 63 or 152-163, wherein partitioning separates the most different haplotypes into different compartments, and the method is used for haplotype phasing.

[0189] Embodiment 165 is the method of Embodiment 164, wherein haplotype fading does not require a barcode.

[0190] Embodiment 166 is a method according to any one of embodiments 93-121 or 131-165, further comprising amplifying the mold.

[0191] Embodiment 167 is a method according to any one of Embodiments 49-55, 57-59, 93-121, or 131-166, further comprising determining the arrangement of the mold.

[0192] Embodiment 168 is the method of Embodiment 167, wherein sequencing is performed using sequencing primers that bind to A14, B15, and / or hybridization sequences (HYB).

[0193] Embodiment 169 is the method of Embodiment 167 or 168, wherein the sequencing includes a dark cycle and data is not recorded for part of the sequencing.

[0194] Embodiment 170 is the method of Embodiment 169, wherein the unrecorded data is sequence data related to the 3' transposon terminal sequence or its complement.

[0195] Embodiment 171 is a method according to any one of Embodiments 167 to 170, further comprising (a) evaluating the sequences of inserts contained in the same template, and (b) determining proximity data for sequences contained in double-stranded nucleic acids based on the inserts contained in the same template.

[0196] Embodiment 172 is the method of Embodiment 171, wherein proximity data determines that the insert sequence (or its complement) was contained in the same target nucleic acid.

[0197] Embodiment 173 is a method according to any one of Embodiments 167 to 172, further comprising (a) evaluating sequencing results derived from multiple sequences of a given insert prepared from different templates, and (b) determining non-standard base pairing events based on sequencing data derived from (i) inserts and their complements contained in the same linked sequencing template, and / or (ii) inserts contained in multiple linked sequencing templates.

[0198] Embodiment 174 is a method according to any one of Embodiments 167 to 173, further comprising: evaluating sequencing results derived from multiple sequences of a given insert prepared from different templates; and correcting errors in the sequencing results for this insert based on sequencing data derived from (i) inserts and their complements contained in the same linked sequencing template, and / or (ii) inserts contained in multiple linked sequencing templates.

[0199] Embodiment 175 is a method for identifying modified cytosines contained in an insert sequence contained in a linked sequencing template, comprising: (a) preparing a double-stranded linked sequencing template, wherein each strand contains an insert sequence and a copy of the insert sequence, and the two strands are complementary to each other; (b) subjecting the double-stranded linked sequencing template to conditions for modifying and / or unmodified cytosines; (c) preparing amplicons for each strand of the double-stranded linked sequencing template; (d) sequencing the amplicons and evaluating the sequencing results for the insert sequence and the copy of the insert sequence in the amplicons generated from each strand; and (e) determining the location of modified cytosines contained in the insert sequence based on the sequences of each strand of the double-stranded linked sequencing template.

[0200] Embodiment 176 is the method of Embodiment 175, wherein the modified cytosine is methylated or hydroxymethylated cytosine.

[0201] Embodiment 177 is the method of Embodiment 175 or 176, wherein the linked sequence determination template is prepared by the method described in any one of Embodiments 93-121 or 131-165.

[0202] Embodiment 178 is the method of Embodiment 177, wherein the extension for generating a double-stranded linked sequencing template is carried out using a reaction solution containing methylated dCTP.

[0203] Embodiment 179 is a method according to any one of Embodiments 175 to 178, wherein uracil contained in the linked sequencing template is converted to thymine when preparing the amplicon.

[0204] Embodiment 180 is a modified cytosine or unmodified cytosine, which is optionally modified. The method according to any one of Embodiments 175 to 179, wherein cytosine is modified by TET-assisted pyridineborane sequencing (TAPS) treatment, or unmodified cytosine is modified by sodium bisulfite or enzymatic treatment.

[0205] Embodiment 181 is the method of Embodiment 180, wherein the modified cytosine is modified, the position of the modified cytosine is determined by the presence of (T,C) in the insert sequence and the copy of the insert sequence, respectively, the position of the unmodified cytosine is determined by the presence of (C,C) in the insert sequence and the copy of the insert sequence, respectively, and the modified and unmodified cytosine are paired with G' in the complementary chain.

[0206] Embodiment 182 is the method of Embodiment 180, wherein the unmodified cytosine is modified, the position of the modified cytosine is determined by the presence of (C,T) in the insert sequence and the copy of the insert sequence, respectively, the position of the unmodified cytosine is determined by the presence of (T,T) in the insert sequence and the copy of the insert sequence, respectively, and the modified and unmodified cytosine are paired with G' in the complementary chain.

[0207] Embodiment 183 is the method of Embodiment 180, wherein the method distinguishes the position of methylated cytosine from that of hydroxymethylated cytosine.

[0208] Embodiment 184 is the method of Embodiment 183, wherein subjecting each strand to conditions for modifying and / or unmodified cytosine comprises (a) reacting each strand with β-glycosyltransferase, (b) reacting each strand with DNA methyltransferase (DNMT), and (c) reacting each strand under conditions for converting unmodified cytosine to uracil.

[0209] Embodiment 185 is the method of Embodiment 184, wherein (a) the position of methylated cytosine is determined by the presence of (C,C) in the insert sequence and the copy of the insert sequence, respectively; (b) the position of hydroxymethylated cytosine is determined by the presence of (C,T) in the insert sequence and the copy of the insert sequence, respectively; and (c) the position of unmodified cytosine is determined by the presence of (T,T) in the insert sequence and the copy of the insert sequence, respectively, and the methylated, hydroxymethylated, and unmodified cytosines are paired with G' in the complementary chain.

[0210] Embodiment 186 involves subjecting each chain to conditions for modifying modified cytosine and / or unmodified cytosine, (a) reacting each chain with DNMT, and (b) methylating each chain with dihydroxyuracil. DH The method according to Embodiment 183, which includes reacting under conditions that convert to U).

[0211] Embodiment 187 is the method of Embodiment 186, wherein (a) the position of methylated cytosine is determined by the presence of (T,T) in the insert sequence and the copy of the insert sequence, respectively; (b) the position of hydroxymethylated cytosine is determined by the presence of (T,C) in the insert sequence and the copy of the insert sequence, respectively; and (c) the position of unmodified cytosine is determined by the presence of (C,C) in the insert sequence and the copy of the insert sequence, respectively, and the methylated, hydroxymethylated, and unmodified cytosines are paired with G' in the complementary chain.

[0212] Additional objectives and benefits are partially described below, some of which are evident from the description or can be learned through practice. These objectives and benefits will be realized and achieved by the elements and combinations specifically indicated in the attached claims.

[0213] Please understand that the general explanation above and the detailed explanation below are illustrative and explanatory only, and do not limit the scope of the claims.

[0214] The accompanying drawings incorporated herein and constituting part of this specification illustrate one (or several) embodiments and serve to illustrate the principles described herein together with the specification. [Brief explanation of the drawing]

[0215] [Figure 1] This paper provides an overview of a method for increasing sequencing throughput in a flow cell using polynucleotides containing two insert sequences. Sequencing is performed using lead 1 (R1) sequencing primer, followed by lead 2 (R2) sequencing primer. A turnaround is then performed, followed by sequencing using lead 3 (R3) sequencing primer, followed by lead 4 (R4) sequencing primer. [Figure 2] This describes the sequencing of a representative polynucleotide having two insert sequences, including P5' and P7 sequences, and a hybridization (HYB) sequence. The polynucleotide is first sequenced using a read 1 sequencing primer that hybridizes to the 3' polynucleotide (containing the P5' sequence), followed by a read 2 sequencing primer that hybridizes to the HYB sequence. A turnaround is then performed. The polynucleotide is then sequenced using a read 3 sequencing primer that hybridizes to the 3 polynucleotide (containing the P7' sequence) and a read 4 sequencing primer that hybridizes to the complementary sequence (HYB') of the hybridization sequence. [Figure 3]The sequencing of a representative polynucleotide with two insert sequences, generated from Library A or Library B, is shown. The polynucleotide is sequenced using a Read 1 sequencing primer that first hybridizes to the 3' polynucleotide (containing the P5' sequence), followed by a Read 2 sequencing primer that hybridizes to the HYB sequence and the SBS sequence. The SBS sequence assists in the binding of the sequencing primers; for example, the SBS sequence may contain ME or ME'. A turnaround is then performed. The polynucleotide is then sequenced using a Read 3 sequencing primer that hybridizes to the 3' polynucleotide (containing the P7' sequence), followed by a Read 4 sequencing primer that hybridizes to the hybridization sequence complement (HYB') and the SBS sequence. The representative polynucleotide is also shown to originate from two separate libraries (Library A and Library B). [Figure 4A]This outlines the sequencing of a standard Illumina paired-end library containing one insert, compared to the sequencing of a polynucleotide containing two insert sequences. (A) Using a standard Illumina paired-end library, 150 cycles of sequential synthesis sequencing (SBS) are performed forward using a read 1 sequencing (seq) primer (SEQ ID NO: 22) that hybridizes to A14' and ME'. Then, a paired-end turnaround is performed, and 150 cycles of SBS sequencing are performed on the reverse strand using a read 2 seq primer (SEQ ID NO: 23) that hybridizes to B15' and ME'. (B) Using a paired-end library of polynucleotides containing two insert sequences, 150 cycles of SBS sequencing are performed forward using a read 1 sequencing (first read) primer that hybridizes to A14' and ME'. Then, the SBS synthesis strand can be denatured, and the read 1-B seq primer (second read) hybridizes to HYB and ME'. Paired-end turnaround is performed, and 150 cycles of sequencing by SBS sequencing are carried out on each reverse strand of the two insert sequences of the polynucleotide, using lead 2-A sequencing (third read) primers that hybridize to B15' and ME', followed by lead 2-B sequencing (fourth read) primers that hybridize to HYB' and ME'. In this way, the sequences of the two insert sequences derived from the target nucleic acid are obtained using the same flow cell region as in standard methods. [Figure 4B]This outlines the sequencing of a standard Illumina paired-end library containing one insert, compared to the sequencing of a polynucleotide containing two insert sequences. (A) Using a standard Illumina paired-end library, 150 cycles of sequential synthesis sequencing (SBS) are performed forward using a read 1 sequencing (seq) primer (SEQ ID NO: 22) that hybridizes to A14' and ME'. Then, a paired-end turnaround is performed, and 150 cycles of SBS sequencing are performed on the reverse strand using a read 2 seq primer (SEQ ID NO: 23) that hybridizes to B15' and ME'. (B) Using a paired-end library of polynucleotides containing two insert sequences, 150 cycles of SBS sequencing are performed forward using a read 1 sequencing (first read) primer that hybridizes to A14' and ME'. Then, the SBS synthesis strand can be denatured, and the read 1-B seq primer (second read) hybridizes to HYB and ME'. Paired-end turnaround is performed, and 150 cycles of sequencing by SBS sequencing are carried out on each reverse strand of the two insert sequences of the polynucleotide, using lead 2-A sequencing (third read) primers that hybridize to B15' and ME', followed by lead 2-B sequencing (fourth read) primers that hybridize to HYB' and ME'. In this way, the sequences of the two insert sequences derived from the target nucleic acid are obtained using the same flow cell region as in standard methods. [Figure 5A] This illustrates the steps in a standard Nextera Flex workflow that yield a sequence-ready fragment containing a single insert sequence derived from a target nucleic acid (genomic DNA or gDNA). [Figure 5B] This illustrates the steps in a standard Nextera Flex workflow that yield a sequence-ready fragment containing a single insert sequence derived from a target nucleic acid (genomic DNA or gDNA). [Figure 5C]This illustrates the steps in a standard Nextera Flex workflow that yield a sequence-ready fragment containing a single insert sequence derived from a target nucleic acid (genomic DNA or gDNA). [Figure 6A] A general overview of the preparation of a tandem-read library using transpososomes to incorporate the A14 and B15 sequences (A), followed by PCR to add either the P5 and HYB(H) sequences (B) or HYB'(H') and P7' (C). Library products enclosed in the box in (D) can form hybridization adducts with other library products (via HYB / HYB' hybridization) to enable extension. At least 1 / 9 of the extended products are expected to be sequenceable (E). [Figure 6B] A general overview of the preparation of a tandem-read library using transpososomes to incorporate the A14 and B15 sequences (A), followed by PCR to add either the P5 and HYB(H) sequences (B) or HYB'(H') and P7' (C). Library products enclosed in the box in (D) can form hybridization adducts with other library products (via HYB / HYB' hybridization) to enable extension. At least 1 / 9 of the extended products are expected to be sequenceable (E). [Figure 6C] A general overview of the preparation of a tandem-read library using transpososomes to incorporate the A14 and B15 sequences (A), followed by PCR to add either the P5 and HYB(H) sequences (B) or HYB'(H') and P7' (C). Library products enclosed in the box in (D) can form hybridization adducts with other library products (via HYB / HYB' hybridization) to enable extension. At least 1 / 9 of the extended products are expected to be sequenceable (E). [Figure 6D]A general overview of the preparation of a tandem-read library using transpososomes to incorporate the A14 and B15 sequences (A), followed by PCR to add either the P5 and HYB(H) sequences (B) or HYB'(H') and P7' (C). Library products enclosed in the box in (D) can form hybridization adducts with other library products (via HYB / HYB' hybridization) to enable extension. At least 1 / 9 of the extended products are expected to be sequenceable (E). [Figure 6E] A general overview of the preparation of a tandem-read library using transpososomes to incorporate the A14 and B15 sequences (A), followed by PCR to add either the P5 and HYB(H) sequences (B) or HYB'(H') and P7' (C). Library products enclosed in the box in (D) can form hybridization adducts with other library products (via HYB / HYB' hybridization) to enable extension. At least 1 / 9 of the extended products are expected to be sequenceable (E). [Figure 7A] The method for forming a P5-HYB' fork library in one tube using bead-based tagmentation and a P7-HYB fork library in another tube using solution-based tagmentation is shown (A). The library products can form hybridized adducts based on the hybridization of HYB and HYB', and can generate polynucleotides via extension (B). [Figure 7B] The method for forming a P5-HYB' fork library in one tube using bead-based tagmentation and a P7-HYB fork library in another tube using solution-based tagmentation is shown (A). The library products can form hybridized adducts based on the hybridization of HYB and HYB', and can generate polynucleotides via extension (B). [Figure 8A] The preparation of library products using bead-linked transposomes (BLTs) is shown in Tube 1 (Type 1 BLT immobilized on beads by P5) and Tube 2 (Type 2 BLT immobilized on beads by P7). P7 can be immobilized on beads using a single desthiobiotin, which can be readily removed from streptavidin-coated beads using a free buffer (A). Thus, the P7-HYB library can be selectively released from the beads and hybridized to a P5-HYB' library on bead type 1 (B). After extension, a linked nucleic acid sequencing template is generated. [Figure 8B] The preparation of library products using bead-linked transposomes (BLTs) is shown in Tube 1 (Type 1 BLT immobilized on beads by P5) and Tube 2 (Type 2 BLT immobilized on beads by P7). P7 can be immobilized on beads using a single desthiobiotin, which can be readily removed from streptavidin-coated beads using a free buffer (A). Thus, the P7-HYB library can be selectively released from the beads and hybridized to a P5-HYB' library on bead type 1 (B). After extension, a linked nucleic acid sequencing template is generated. [Figure 9A] A simple single-tube workflow based on bead-linked transposons is shown, enabling the creation of two libraries, one containing HYB' and the other containing HYB (A). The linked nucleic acid sequencing templates are prepared through the processes of denaturation, hybridization, and extension (B). [Figure 9B] A simple single-tube workflow based on bead-linked transposons is shown, enabling the creation of two libraries, one containing HYB' and the other containing HYB (A). The linked nucleic acid sequencing templates are prepared through the processes of denaturation, hybridization, and extension (B). [Figure 10]This diagram shows a typical Truseq method for generating two library products that can be used to produce polynucleotides containing two inserts that can be used for sequencing. An SBS sequence is a sequence that can bind to a sequencing primer; for example, an SBS sequence may contain a sequence complementary to a known sequencing primer. In this diagram, "SBS" refers collectively to either an SBS sequence or a sequence that is completely or partially complementary to an SBS sequence (e.g., SBS or SBS'). [Figure 11] The results of a bioanalyzer test on the size of the tandem library (i.e., polynucleotides containing two insert sequences) generated by the Trueq method are shown, compared to the two library products (P5-HYB' and P7-HYB) used to generate the tandem library. [Figure 12] Two libraries generated by the Trueq method are shown, in which the attached polynucleotides and hybridization polynucleotides of each fork-type adapter contain SBS sequences. As shown in Figure 12, "SBS" can refer collectively to either SBS or SBS' sequences (i.e., the tandem SBS sequences in Figure 12 may include fully or partially complementary SBS / SBS' sequences). [Figure 13] Two libraries generated by the Trueq method are shown, in which the attached polynucleotides of each fork-type adapter contain either A14 and ME or B15 and ME. [Figure 14A] This shows thumbnail images of data from the sequencing of a polynucleotide containing two insert sequences, each having a lead 1-A seq primer (first lead primer 1, (A)) and a lead 1-B seq primer (second lead primer, (B)). [Figure 14B] This shows thumbnail images of data from the sequencing of a polynucleotide containing two insert sequences, each having a lead 1-A seq primer (first lead primer 1, (A)) and a lead 1-B seq primer (second lead primer, (B)). [Figure 15A]An exemplary method for preparing a tandem insert library using ligation is shown. Figure 15A shows an exemplary first starting library with a BtgZI cleavage site. Figure 15B shows an exemplary second starting library with a BglII cleavage site. Each of the two starting libraries is digested with its respective restriction enzyme to generate compatibility overhangs (Figures 15C-D). The digested DNA is purified using streptavidin magnetic beads, and the digested DNA is ligated together (Figure 15E). Each new fragment of DNA has a unique adapter that mitigates the complementarity issue of the fork handle. The new libraries are sequenced using primer reads 1, 2, 3, and 4 (Figure 15F). Exemplary P5 and P7 sequences are shown in black highlight and white text. [Figure 15B] An exemplary method for preparing a tandem insert library using ligation is shown. Figure 15A shows an exemplary first starting library with a BtgZI cleavage site. Figure 15B shows an exemplary second starting library with a BglII cleavage site. Each of the two starting libraries is digested with its respective restriction enzyme to generate compatibility overhangs (Figures 15C-D). The digested DNA is purified using streptavidin magnetic beads, and the digested DNA is ligated together (Figure 15E). Each new fragment of DNA has a unique adapter that mitigates the complementarity issue of the fork handle. The new libraries are sequenced using primer reads 1, 2, 3, and 4 (Figure 15F). Exemplary P5 and P7 sequences are shown in black highlight and white text. [Figure 15C]An exemplary method for preparing a tandem insert library using ligation is shown. Figure 15A shows an exemplary first starting library with a BtgZI cleavage site. Figure 15B shows an exemplary second starting library with a BglII cleavage site. Each of the two starting libraries is digested with its respective restriction enzyme to generate compatibility overhangs (Figures 15C-D). The digested DNA is purified using streptavidin magnetic beads, and the digested DNA is ligated together (Figure 15E). Each new fragment of DNA has a unique adapter that mitigates the complementarity issue of the fork handle. The new libraries are sequenced using primer reads 1, 2, 3, and 4 (Figure 15F). Exemplary P5 and P7 sequences are shown in black highlight and white text. [Figure 15D] An exemplary method for preparing a tandem insert library using ligation is shown. Figure 15A shows an exemplary first starting library with a BtgZI cleavage site. Figure 15B shows an exemplary second starting library with a BglII cleavage site. Each of the two starting libraries is digested with its respective restriction enzyme to generate compatibility overhangs (Figures 15C-D). The digested DNA is purified using streptavidin magnetic beads, and the digested DNA is ligated together (Figure 15E). Each new fragment of DNA has a unique adapter that mitigates the complementarity issue of the fork handle. The new libraries are sequenced using primer reads 1, 2, 3, and 4 (Figure 15F). Exemplary P5 and P7 sequences are shown in black highlight and white text. [Figure 15E]An exemplary method for preparing a tandem insert library using ligation is shown. Figure 15A shows an exemplary first starting library with a BtgZI cleavage site. Figure 15B shows an exemplary second starting library with a BglII cleavage site. Each of the two starting libraries is digested with its respective restriction enzyme to generate compatibility overhangs (Figures 15C-D). The digested DNA is purified using streptavidin magnetic beads, and the digested DNA is ligated together (Figure 15E). Each new fragment of DNA has a unique adapter that mitigates the complementarity issue of the fork handle. The new libraries are sequenced using primer reads 1, 2, 3, and 4 (Figure 15F). Exemplary P5 and P7 sequences are shown in black highlight and white text. [Figure 15F] An exemplary method for preparing a tandem insert library using ligation is shown. Figure 15A shows an exemplary first starting library with a BtgZI cleavage site. Figure 15B shows an exemplary second starting library with a BglII cleavage site. Each of the two starting libraries is digested with its respective restriction enzyme to generate compatibility overhangs (Figures 15C-D). The digested DNA is purified using streptavidin magnetic beads, and the digested DNA is ligated together (Figure 15E). Each new fragment of DNA has a unique adapter that mitigates the complementarity issue of the fork handle. The new libraries are sequenced using primer reads 1, 2, 3, and 4 (Figure 15F). Exemplary P5 and P7 sequences are shown in black highlight and white text. [Figure 16A]An exemplary method for preparing tandem insert libraries with two different ends is shown. Figure 16A shows an exemplary workflow for preparing a first library using an adapter with a BtgZI cleavage site and a P5-read 1 site. Figure 16B shows an exemplary workflow for preparing a second library using an adapter with a BglII cleavage site and a P7-read 2 site. Both libraries are double-stranded by primer extension using a single primer. [Figure 16B] An exemplary method for preparing tandem insert libraries with two different ends is shown. Figure 16A shows an exemplary workflow for preparing a first library using an adapter with a BtgZI cleavage site and a P5-read 1 site. Figure 16B shows an exemplary workflow for preparing a second library using an adapter with a BglII cleavage site and a P7-read 2 site. Both libraries are double-stranded by primer extension using a single primer. [Figure 17]An exemplary method for preparing tandem insert libraries using the Strand Overlap Extension (SOE) method is described. DNA1 and DNA2 represent the inputs to exemplary first and second libraries. DNA1 and DNA2 are prepared separately, and as a result, each obtained tandem insert library has DNA attached to a unique adapter. Each library is sheared to generate DNA fragments and treated with polymerase to remove damaged DNA ends resulting from the shearing process. The DNA fragments are treated with polymerase to generate blunt-ended DNA double helixes and treated with kinase to phosphorylate the 5'OH of the DNA fragments. Then, using polymerase, adenine is added to the 3' end of each double helix to ligate the DNA fragments to the adapters. The first library is ligated with a P5-read 1 / A adapter (adapter 1). The second library is ligated with a P7-index-read 2 / A' adapter (adapter 2 or 3). The libraries are purified and 150-200 base pair fragments are selected. The library is mixed and added to the PCR reaction. The DNA fragments are denatured at high temperature and re-annealed at low temperature. This causes the A and A' complementary sequences to hybridize with each other. The chain is extended by polymerase to form a tandem insert polynucleotide. ER = end repair. A-tail = adenine tail. Tag = exemplary index in the barcode sequence. P5 = P5 primer sequence. P7 = P7 primer sequence. In some embodiments, the tag is added adjacent to P7. In some embodiments, the tag is added adjacent to P5. [Figure 18] An exemplary library fragment is shown with two inserts separated by adapter sequences. As shown, four sequencing reads are possible. Reads 1 and 4 give paired-end data from the first insert. Reads 2 and 3 give paired-end data from the second insert. P5 = P5 primer sequence. P7 = P7 primer sequence. [Figure 19]An exemplary tandem insert library fragment is shown, containing inserts from two distinct genomes (E. coli and human), or two distinct amplicons from the same genome. The two inserts are separated by an adapter sequence. As shown, four sequencing reads are possible. For example, reads 1 and 4 give paired-end data from the E. coli insert. Reads 2 and 3 give paired-end data from the human insert. P5 = P5 primer sequence. P7 = P7 primer sequence. [Figure 20A] Figures 15A to 15F show the sequencing data of the tandem insert library generated using the ligation method shown. Figure 20A = Read 1. Figure 20B = Read 2. Figure 20C = Read 3. Figure 20D = Read 4. [Figure 20B] Figures 15A to 15F show the sequencing data of the tandem insert library generated using the ligation method shown. Figure 20A = Read 1. Figure 20B = Read 2. Figure 20C = Read 3. Figure 20D = Read 4. [Figure 20C] Figures 15A to 15F show the sequencing data of the tandem insert library generated using the ligation method shown. Figure 20A = Read 1. Figure 20B = Read 2. Figure 20C = Read 3. Figure 20D = Read 4. [Figure 20D] Figures 15A to 15F show the sequencing data of the tandem insert library generated using the ligation method shown. Figure 20A = Read 1. Figure 20B = Read 2. Figure 20C = Read 3. Figure 20D = Read 4. [Figure 21A] Figures 15A to 15F show sequencing data for tandem insert libraries generated using the ligation method shown. The number of cycles or base call percentage for each read is indicated. Each insert shows the accurate base composition for the corresponding genome. Figure 21A = Reads 1 and 4 for the E. coli insert. Figure 21B = Reads 2 and 3 for the human insert. [Figure 21B]Figures 15A to 15F show sequencing data for tandem insert libraries generated using the ligation method shown. The number of cycles or base call percentage for each read is indicated. Each insert shows the accurate base composition for the corresponding genome. Figure 21A = Reads 1 and 4 for the E. coli insert. Figure 21B = Reads 2 and 3 for the human insert. [Figure 22] Figure 17 shows fragments of a tandem insert library prepared using the SOE method. Instead of using sheared genomic DNA fragments, this experiment used monotemplates, with a PhiX amplicon used as insert 1 and an E. coli amplicon used as insert 2. The adapter was ligated to the monotemplate, and the tandem insert library was prepared using the SOE method as shown in Figure 17. Reads 1 and 4 provide paired-end data from the PhiX amplicon. Reads 2 and 3 provide paired-end data from the E. coli amplicon. P5 = P5 primer sequence. P7 = P7 primer sequence. [Figure 23A] Figure 17 shows the sequencing data of a tandem insert library generated using the SOE method shown. Figure 23A = Read 1. Figure 23B = Read 2. Figure 23C = Read 3. Figure 23D = Read 4. [Figure 23B] Figure 17 shows the sequencing data of a tandem insert library generated using the SOE method shown. Figure 23A = Read 1. Figure 23B = Read 2. Figure 23C = Read 3. Figure 23D = Read 4. [Figure 23C] Figure 17 shows the sequencing data of a tandem insert library generated using the SOE method shown. Figure 23A = Read 1. Figure 23B = Read 2. Figure 23C = Read 3. Figure 23D = Read 4. [Figure 23D] Figure 17 shows the sequencing data of a tandem insert library generated using the SOE method shown. Figure 23A = Read 1. Figure 23B = Read 2. Figure 23C = Read 3. Figure 23D = Read 4. [Figure 24A]Figure 17 shows sequencing data for a tandem insert library generated using the SOE method. Figure 24A shows the predicted sequences of reads 1, 2, 3, and 4 derived from the tandem insert library polynucleotides. The double slash mark " / / " indicates that the shown DNA sequence belongs to a single polynucleotide template. Figures 24B-C show the observed read 1 (Figure 24B) and read 2 (Figure 24C) sequences. [Figure 24B] Figure 17 shows sequencing data for a tandem insert library generated using the SOE method. Figure 24A shows the predicted sequences of reads 1, 2, 3, and 4 derived from the tandem insert library polynucleotides. The double slash mark " / / " indicates that the shown DNA sequence belongs to a single polynucleotide template. Figures 24B-C show the observed read 1 (Figure 24B) and read 2 (Figure 24C) sequences. [Figure 24C] Figure 17 shows sequencing data for a tandem insert library generated using the SOE method. Figure 24A shows the predicted sequences of reads 1, 2, 3, and 4 derived from the tandem insert library polynucleotides. The double slash mark " / / " indicates that the shown DNA sequence belongs to a single polynucleotide template. Figures 24B-C show the observed read 1 (Figure 24B) and read 2 (Figure 24C) sequences. [Figure 25]This document provides an overview of fork-type adapters that can be used to prepare sequencing templates containing multiple inserts derived from target nucleic acids. The first oligonucleotide of a first fork-type adapter ("first adapter") may include a 3' end containing a transposon terminal sequence and a 5' end containing an adapter such as a first read sequencing adapter sequence (P5.R1). The first adapter may also include a second oligonucleotide that includes a 5' end containing a complement to the transposon terminal sequence contained in the first oligonucleotide and a 3' end containing a complement to a hybridization sequence (X'). The first adapter may also include a third oligonucleotide that is a blocking oligonucleotide (X'B) that can bind to X'. In parallel, the first oligonucleotide of a second fork-type adapter ("second adapter") may include a first oligonucleotide that includes a 3' end containing a transposon terminal sequence and a 5' end containing an adapter such as a second read sequencing adapter sequence (P7.R2). The second adapter may also include a second oligonucleotide comprising a 5' end containing a complement to the transposon terminal sequence contained in the first oligonucleotide and a 3' end containing a hybridization sequence (X). The second adapter may also include a third oligonucleotide, which is a blocking oligonucleotide (X'B') capable of binding to X. The blocking oligonucleotide acts to block the hybridization of X' in the first fork-type adapter to X in the second fork-type adapter until the blocking oligonucleotide is removed. Both the first and second adapters may be used in a method for preparing a sequencing template comprising two inserts, as described herein. Bio = biotin, which can be used as the affinity moiety. [Figure 26A]Different combinations of first and second fork-type adapters that may be used in the method of the present invention are shown, along with a presentation of a method for preparing similar fragments using transposomes in solution. (A) The second oligonucleotides of both the first and second fork-type adapters are bound to a blocking oligonucleotide. (B) The second oligonucleotide of the first fork-type adapter is bound to a blocking oligonucleotide. (C) The second oligonucleotide of the second fork-type adapter is bound to a blocking oligonucleotide. (D) Using two pools of transposomes in solution, target nucleic acids can be tagged to fragments in solution. After inactivation (e.g., by SDS) and extension and ligation using an extension-ligation mix (ELM), similar tagged fragments can be prepared as shown in A-C for ligation of fork-type adapters. [Figure 26B] Different combinations of first and second fork-type adapters that may be used in the method of the present invention are shown, along with a presentation of a method for preparing similar fragments using transposomes in solution. (A) The second oligonucleotides of both the first and second fork-type adapters are bound to a blocking oligonucleotide. (B) The second oligonucleotide of the first fork-type adapter is bound to a blocking oligonucleotide. (C) The second oligonucleotide of the second fork-type adapter is bound to a blocking oligonucleotide. (D) Using two pools of transposomes in solution, target nucleic acids can be tagged to fragments in solution. After inactivation (e.g., by SDS) and extension and ligation using an extension-ligation mix (ELM), similar tagged fragments can be prepared as shown in A-C for ligation of fork-type adapters. [Figure 26C]Different combinations of first and second fork-type adapters that may be used in the method of the present invention are shown, along with a presentation of a method for preparing similar fragments using transposomes in solution. (A) The second oligonucleotides of both the first and second fork-type adapters are bound to a blocking oligonucleotide. (B) The second oligonucleotide of the first fork-type adapter is bound to a blocking oligonucleotide. (C) The second oligonucleotide of the second fork-type adapter is bound to a blocking oligonucleotide. (D) Using two pools of transposomes in solution, target nucleic acids can be tagged to fragments in solution. After inactivation (e.g., by SDS) and extension and ligation using an extension-ligation mix (ELM), similar tagged fragments can be prepared as shown in A-C for ligation of fork-type adapters. [Figure 26D] Different combinations of first and second fork-type adapters that may be used in the method of the present invention are shown, along with a presentation of a method for preparing similar fragments using transposomes in solution. (A) The second oligonucleotides of both the first and second fork-type adapters are bound to a blocking oligonucleotide. (B) The second oligonucleotide of the first fork-type adapter is bound to a blocking oligonucleotide. (C) The second oligonucleotide of the second fork-type adapter is bound to a blocking oligonucleotide. (D) Using two pools of transposomes in solution, target nucleic acids can be tagged to fragments in solution. After inactivation (e.g., by SDS) and extension and ligation using an extension-ligation mix (ELM), similar tagged fragments can be prepared as shown in A-C for ligation of fork-type adapters. [Figure 27A]Figures 26A-26D show different tagged fragments that can be produced by ligation or tagmentation in a solution containing a mixture of the first and second fork adapters shown. (A) A fragment tagged with one end of the first fork adapter and the other end of the ligated second fork adapter. (B) A fragment tagged with both ends of the first fork adapter. (C) A fragment tagged with the ligated second fork adapter at both ends. The expected ratio of tagged fragments would be 50%(A):25%(B):25%(C). [Figure 27B] Figures 26A-26D show different tagged fragments that can be produced by ligation or tagmentation in a solution containing a mixture of the first and second fork adapters shown. (A) A fragment tagged with one end of the first fork adapter and the other end of the ligated second fork adapter. (B) A fragment tagged with both ends of the first fork adapter. (C) A fragment tagged with the ligated second fork adapter at both ends. The expected ratio of tagged fragments would be 50%(A):25%(B):25%(C). [Figure 27C] Figures 26A-26D show different tagged fragments that can be produced by ligation or tagmentation in a solution containing a mixture of the first and second fork adapters shown. (A) A fragment tagged with one end of the first fork adapter and the other end of the ligated second fork adapter. (B) A fragment tagged with both ends of the first fork adapter. (C) A fragment tagged with the ligated second fork adapter at both ends. The expected ratio of tagged fragments would be 50%(A):25%(B):25%(C). [Figure 28A]This describes how different types of tagged fragments (using the representative first and second adapters shown in Figure 25 or the method in Figure 26D) hybridize or do not hybridize after being immobilized on the surface of a solid support. For ease of explanation, the solid supports shown on the left and right present two different figures of the same surface on the solid support, all nucleic acid fragments extend upward from the same surface on the solid support, and hybridized fragments form a cross-linked structure. (A) Two single-stranded fragments are generated by immobilizing a double-stranded fragment containing an insert on the surface of a solid support and denaturing it. A first single-stranded fragment containing a first oligonucleotide ligated to a first fork adapter (P5.R1) ​​at one end and a second oligonucleotide ligated to a second fork adapter at the other end (X) can hybridize to a second single-stranded fragment containing a second strand ligated to a first fork adapter (X') at one end and a first oligonucleotide ligated to a second fork adapter at the other end (P7.R2). Since both strands derived from the double-stranded fragment are likely to be fixed in close proximity to each other after the double-stranded fragment has denatured (as shown), these two fragments are likely to be complements of each other (i.e., they were two single strands contained in the same double-stranded fragment). The two fragments may also be sequences that are not complements of each other (not shown). This hybridization of the two single-stranded fragments occurs via the binding of the hybridization sequence (X) to the hybridization sequence's complement (X'). After hybridization of two fragments by X / X', extension can be performed from the 3' end of the ligated sequence. (B) Single-stranded tagged fragments having ligated first / second oligonucleotides from the first fork adapter at both ends cannot hybridize with each other (because both contain an X' sequence at one end). (C) Single-stranded tagged fragments having ligated first / second oligonucleotides from the second fork adapter at both ends cannot hybridize with each other in the hybridization sequence (because both contain an X sequence at one end).Therefore, 100% of single-stranded fragments having the same inserts that can hybridize with each other in the hybridization sequence are prepared from double-stranded fragments having one fork-shaped adapter at the first end and a second fork-shaped adapter at the second end. [Figure 28B]This describes how different types of tagged fragments (using the representative first and second adapters shown in Figure 25 or the method in Figure 26D) hybridize or do not hybridize after being immobilized on the surface of a solid support. For ease of explanation, the solid supports shown on the left and right present two different figures of the same surface on the solid support, all nucleic acid fragments extend upward from the same surface on the solid support, and hybridized fragments form a cross-linked structure. (A) Two single-stranded fragments are generated by immobilizing a double-stranded fragment containing an insert on the surface of a solid support and denaturing it. A first single-stranded fragment containing a first oligonucleotide ligated to a first fork adapter (P5.R1) ​​at one end and a second oligonucleotide ligated to a second fork adapter at the other end (X) can hybridize to a second single-stranded fragment containing a second strand ligated to a first fork adapter (X') at one end and a first oligonucleotide ligated to a second fork adapter at the other end (P7.R2). Since both strands derived from the double-stranded fragment are likely to be fixed in close proximity to each other after the double-stranded fragment has denatured (as shown), these two fragments are likely to be complements of each other (i.e., they were two single strands contained in the same double-stranded fragment). The two fragments may also be sequences that are not complements of each other (not shown). This hybridization of the two single-stranded fragments occurs via the binding of the hybridization sequence (X) to the hybridization sequence's complement (X'). After hybridization of two fragments by X / X', extension can be performed from the 3' end of the ligated sequence. (B) Single-stranded tagged fragments having ligated first / second oligonucleotides from the first fork adapter at both ends cannot hybridize with each other (because both contain an X' sequence at one end). (C) Single-stranded tagged fragments having ligated first / second oligonucleotides from the second fork adapter at both ends cannot hybridize with each other in the hybridization sequence (because both contain an X sequence at one end).Therefore, 100% of single-stranded fragments having the same inserts that can hybridize with each other in the hybridization sequence are prepared from double-stranded fragments having one fork-shaped adapter at the first end and a second fork-shaped adapter at the second end. [Figure 28C]This describes how different types of tagged fragments (using the representative first and second adapters shown in Figure 25 or the method in Figure 26D) hybridize or do not hybridize after being immobilized on the surface of a solid support. For ease of explanation, the solid supports shown on the left and right present two different figures of the same surface on the solid support, all nucleic acid fragments extend upward from the same surface on the solid support, and hybridized fragments form a cross-linked structure. (A) Two single-stranded fragments are generated by immobilizing a double-stranded fragment containing an insert on the surface of a solid support and denaturing it. A first single-stranded fragment containing a first oligonucleotide ligated to a first fork adapter (P5.R1) ​​at one end and a second oligonucleotide ligated to a second fork adapter at the other end (X) can hybridize to a second single-stranded fragment containing a second strand ligated to a first fork adapter (X') at one end and a first oligonucleotide ligated to a second fork adapter at the other end (P7.R2). Since both strands derived from the double-stranded fragment are likely to be fixed in close proximity to each other after the double-stranded fragment has denatured (as shown), these two fragments are likely to be complements of each other (i.e., they were two single strands contained in the same double-stranded fragment). The two fragments may also be sequences that are not complements of each other (not shown). This hybridization of the two single-stranded fragments occurs via the binding of the hybridization sequence (X) to the hybridization sequence's complement (X'). After hybridization of two fragments by X / X', extension can be performed from the 3' end of the ligated sequence. (B) Single-stranded tagged fragments having ligated first / second oligonucleotides from the first fork adapter at both ends cannot hybridize with each other (because both contain an X' sequence at one end). (C) Single-stranded tagged fragments having ligated first / second oligonucleotides from the second fork adapter at both ends cannot hybridize with each other in the hybridization sequence (because both contain an X sequence at one end).Therefore, 100% of single-stranded fragments having the same inserts that can hybridize with each other in the hybridization sequence are prepared from double-stranded fragments having one fork-shaped adapter at the first end and a second fork-shaped adapter at the second end. [Figure 29] A double-stranded linked sequencing template is shown, with two inserts in each strand prepared using a fork-type adapter. In this representative example, both inserts are copies of the same insert sequence in strand A or strand A' (illustrated). In other examples, the two insert sequences in each strand of the double-stranded linked sequencing template may be different from each other (not shown). [Figure 30] Methods for denaturation (to separate the strands of a double-stranded fragment and remove blocking oligonucleotides) and annealing of fixed single-stranded fragments are described. When used after ligation of a fork-type adapter, these methods can prepare a linked sequencing template containing two inserts on each strand. Because both strands of the double-stranded fragment are constrained and likely to bind within the same region of the solid support, this method often produces a linked sequencing template containing two copies of the same insert sequence (e.g., A' / A' and A / A). A linked sequencing template can be prepared from single-stranded fragments containing different adapters (e.g., A / A', B / B', and D / D'), but a linked sequencing template (produced from two single-stranded fragments generated from one double-stranded fragment) cannot be prepared from single-stranded fragments containing the same adapters at both ends (e.g., C / C' and E / E'). [Figure 31] This describes a method for preparing a sequencing template using tubes or wells as compartments. f1, f2, and f3 refer to different relatively large fragments that can later be converted into smaller fragments. [Figure 32] This document describes a method for preparing a linked arrangement determination template using droplets as compartments. [Figure 33]This document describes a method for preparing concatenated sequencing templates for haplotype phasing using compartments. By limiting dilution of the sample within compartments, the likelihood of two chromosomes from different haplotypes being contained within the same compartment is made extremely low. In this example, Chr1-Hap1 and Chr2-Hap1 are contained within one compartment, while Chr1-Hap2 and Chr2-Hap2 are contained within different compartments. Boxes indicated by checked arrows contain concatenated sequencing templates that may be generated after the processes of degeneration, re-annealing, and extension. Boxes indicated by "X" arrows indicate concatenated sequencing templates that cannot be generated (because these chromosomes were contained within different compartments). Concatenated sequencing templates can only contain insert sequences derived from chromosomes contained within the same compartment, and these templates are contained within the boxes indicated by checked arrows. Dashed ellipses within the boxes indicated by checked arrows represent concatenated sequencing templates that constitute the original haplotype. Other concatenated sequencing templates within the boxes indicated by checked arrows (i.e., those not within the dashed ellipses) contain inserts derived from different chromosomes. [Figure 34]Shows a transposon that can be used to prepare an array determination template containing two or more inserts. The first and second transposons each contain a fork-type adapter. As used herein, "first oligo" or "first strand" may refer to the first transposon contained in the fork-type adapter, and "second oligo" or "second strand" may refer to the second transposon contained in the fork-type adapter. The fork-type adapter of the first transposon includes a first strand containing a 3' transposon end sequence (ME, such as SEQ ID NO: 6) and a 5' first read sequencing adapter sequence (P5.R1), and a second strand containing a 5' complement of the transposon end sequence (ME', such as SEQ ID NO: 3) and a 3' complement of the hybridization sequence (X'). The fork-type adapter of the second transposon includes a first strand containing a 3' transposon end sequence and a 5' second read sequencing adapter sequence (P7.R2), and a second strand containing a 5' complement of the transposon end sequence and a 3' hybridization sequence (X). This representative example shows two pools of transposons, each pool being a homodimer (shown by two checkered transposons or two striped transposons). As described herein, the transposon may include a heterodimer. [Figure 35] Shows a solid support having a fixed transposon (shown in more detail in Figure 34) fixed on its surface. B = biotin, which is used as an affinity moiety for binding the transposon to the surface of the solid support. [Figure 36] Shows the process of tagmentation using the solid support shown in Figure 35. Double-stranded nucleic acid is added to the solid support. Next, fragments are prepared by tagmentation. SDS and washing are used to remove the transposase. Finally, extension and ligation are performed using an extension ligation mix (ELM) buffer. This example shows tagmentation by only one pair of transposons. [Figure 37]Shows the cross-linking of fragments generated by a transposon. The double-stranded DNA may contain sequence A in the sense strand and A' in the antisense strand. The cross-linking may be between the first transposon and the second transposon, or between the first transposon and the first transposon, or between the second transposon and the second transposon. Such permutations occur in a ratio of 50:25:25, respectively. [Figure 38] Shows the released transposon and the immobilized fragments after denaturation of the fragments. The single-stranded fragments may be prepared from the first transposon and the second transposon (50%), or from the first transposon and the first transposon (25%), or from the second transposon and the second transposon (25%). Thus, the fragments have either X or X' at their free ends based on which transposon prepared each fragment. [Figure 39] Shows representative single-stranded fragments and whether they can hybridize with each other to form cross-links. The X / X' set of sequences in two different single-stranded fragments can hybridize (resulting in 100% hybridization), the X' / X' set of sequences cannot hybridize (0%), and the X / X set of sequences cannot hybridize (0%). Thus, 100% cross-linked single-stranded fragments are prepared from the binding of the X sequence in one fragment to X' in another fragment (i.e., the binding of the hybridization sequence to its complement). [Figure 40] Shows the formation (or non-formation) of a ligated sequencing template containing two copies of an insert sequence. After hybridization of the X / X' sequences (100%), a double-stranded ligated sequencing template containing two copies of the tandem A strands in the sense strand and two copies of the tandem A' strands in the antisense strand is formed, but no ligated sequencing template is formed between single-stranded fragments both containing X' (0%) or both containing the X sequence (0%). The resulting double-stranded ligated sequencing template may contain P5 or P5' at one end and P7 or P7' at the other end. [Figure 41] This shows the crosslinks that can be formed when a double-stranded nucleic acid is tagged by a transposome to prepare two crosslinking inserts. In this representative example, the double-stranded nucleic acid contains sequences A and B in the sense strand and sequences A' and B' in the antisense strand. Exemplary options for tagging the two crosslinking fragments with adapter sequences different from the first and / or second fork-type adapter contained in the transposome are shown. [Figure 42] Exemplary hybridizations between single-stranded fragments for generating a concatenated sequencing template are shown. These hybridizations can occur between fragments containing an insert and its complementary sequence (e.g., A / A' or B / B'), or between fragments containing two different inserts (e.g., A / B, A' / B, A / B', and A' / B'). All of the hybridizations generate a sequenceable concatenated sequencing template (after extension) with P5 / P5' at one end and P7 / P7' at the other. Other hybridizations generate some unsequential concatenated sequencing templates (after extension). Unsequential concatenated sequencing templates can include those with P5 / P5' at both ends or P7 / P7' at both ends, and these representative templates are depicted in dashed boxes. [Figure 43] Two cross-linking inserts are shown, prepared either from transpososomes containing a second fork-shaped adapter or from transpososomes containing a first fork-shaped adapter. [Figure 44] This shows that single-stranded fragments having adapters derived from a second fork-type adapter at both ends cannot hybridize together, and that single-stranded fragments having adapters derived from a first fork-type adapter at both ends cannot hybridize together. This lack of hybridization is because X sequences cannot hybridize with other X sequences, and similarly, X' sequences cannot hybridize with other X' sequences. [Figure 45]The group of five bridging inserts illustrates a representative example of how diverse hybridization can occur between fragments containing different insert sequences. Although not shown in the figure, fragments with the same sequence sense and antisense (e.g., A and A') can also hybridize. Not all pair formations result in a sequenceable, concatenated sequencing template (after extension) with different adapters at the template's ends, but many combinations are possible. An exemplary concatenated sequencing template generated from hybridized single-stranded fragments is shown in the box. [Figure 46A] The following are examples of sequencing templates that include sample indices: (A) A transposome complex containing sample index i5 on the first strand of a fork-type adapter contained in a first transposome complex, and sample index i7 on the first strand of a fork-type adapter contained in a second transposome complex, along with a sequencing template that can be prepared using these transpososomes. (B) A transposome complex containing sample index i8 on the second strand of a fork-type adapter contained in a first transposome complex, and sample index i6 on the second strand of a fork-type adapter contained in a second transposome complex, along with a sequencing template that can be prepared using these transpososomes. (C) A representative sequencing template that can be prepared when the first and second strands of the first and second transpososomes contain sample indices. [Figure 46B]The following are examples of sequencing templates that include sample indices: (A) A transposome complex containing sample index i5 on the first strand of a fork-type adapter contained in a first transposome complex, and sample index i7 on the first strand of a fork-type adapter contained in a second transposome complex, along with a sequencing template that can be prepared using these transpososomes. (B) A transposome complex containing sample index i8 on the second strand of a fork-type adapter contained in a first transposome complex, and sample index i6 on the second strand of a fork-type adapter contained in a second transposome complex, along with a sequencing template that can be prepared using these transpososomes. (C) A representative sequencing template that can be prepared when the first and second strands of the first and second transpososomes contain sample indices. [Figure 46C] The following are examples of sequencing templates that include sample indices: (A) A transposome complex containing sample index i5 on the first strand of a fork-type adapter contained in a first transposome complex, and sample index i7 on the first strand of a fork-type adapter contained in a second transposome complex, along with a sequencing template that can be prepared using these transpososomes. (B) A transposome complex containing sample index i8 on the second strand of a fork-type adapter contained in a first transposome complex, and sample index i6 on the second strand of a fork-type adapter contained in a second transposome complex, along with a sequencing template that can be prepared using these transpososomes. (C) A representative sequencing template that can be prepared when the first and second strands of the first and second transpososomes contain sample indices. [Figure 47] This describes a method that may use dark cycling to avoid sequencing of the ME sequence after primer binding to the A14, B15', or X sequence used as a primer binding site for a concatenated sequencing template. Primer binding is indicated by an arrow showing the direction of the sequencing read. [Figure 48]The following shows a typical double-stranded linked sequencing template, each strand containing an insert and a copy of the insert, where the insert sequence contains methylated cytosine (mC) and hydroxymethylated cytosine (hmC), which may be referred to herein as modified cytosine. One single-strand template contains a sense insert (S) and a copy thereof (S-copy), and the other single-strand template contains an antisense insert (S') and a copy thereof (S'-copy). The S-copy and S'-copy do not contain modified cytosine. The underlined positions A, T, and G indicate non-cytosine nucleotides. [Figure 49] Figure 48 shows the results of treating the template by converting unmethylated cytosine to uracil (using sodium bisulfite, for example). [Figure 50A] Figure 25 shows the upper (A) and lower (B) strands of the double-stranded sequencing template shown in Figure 25 before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 50B] Figure 25 shows the upper (A) and lower (B) strands of the double-stranded sequencing template shown in Figure 25 before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 50C] Figure 25 shows the upper (A) and lower (B) strands of the double-stranded sequencing template shown in Figure 25 before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 51] Figure 48 shows the results of processing the template by converting modified cytosine (methylated cytosine and hydroxymethylated cytosine) to dihydroxyuracil (DHU, e.g., by the TAPS method). [Figure 52A] Figure 51 shows the upper (A) and lower (B) strands of the double-stranded sequencing template shown before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 52B] Figure 51 shows the upper (A) and lower (B) strands of the double-stranded sequencing template shown before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 52C]Figure 51 shows the upper (A) and lower (B) strands of the double-stranded sequencing template shown before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 53] The sequencing template prepared by extension in the presence of methylated dCTP is shown. The S-copy and S'-copy may contain methylated cytosine when prepared by this method. [Figure 54] Figure 53 shows the results after processing the sequencing template by converting unmethylated cytosine to uracil. [Figure 55A] Figure 54 shows the upper (A) and lower (B) strands of the double-stranded sequencing template before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 55B] Figure 54 shows the upper (A) and lower (B) strands of the double-stranded sequencing template before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 55C] Figure 54 shows the upper (A) and lower (B) strands of the double-stranded sequencing template before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 56] Figure 53 shows the results after processing the sequencing template by converting unmethylated cytosine to uracil. [Figure 57A] Figure 54 shows the upper (A) and lower (B) strands of the double-stranded sequencing template before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 57B] Figure 54 shows the upper (A) and lower (B) strands of the double-stranded sequencing template before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 57C] Figure 54 shows the upper (A) and lower (B) strands of the double-stranded sequencing template before and after PCR for preparing the amplicon, as well as the analysis of the sequencing results (C). [Figure 58]A representative process included in a method for methylation analysis to distinguish unmodified cytosine, methylated cytosine, and hydroxymethylated cytosine, using β-glucosyltransferase treatment followed by DNA methyltransferase 1 (DNMT1) treatment, is shown. [Figure 59] A method for converting unmethylated cytosine in the sequencing template prepared in FIG. 58 to uracil is shown. [Figure 60A] The upper strand (A) and lower strand (B) of the double-stranded ligated sequencing template shown in FIG. 59 before and after PCR for preparing an amplicon, and the analysis (C) of the sequencing result are shown. [Figure 60B] The upper strand (A) and lower strand (B) of the double-stranded ligated sequencing template shown in FIG. 59 before and after PCR for preparing an amplicon, and the analysis (C) of the sequencing result are shown. [Figure 60C] The upper strand (A) and lower strand (B) of the double-stranded ligated sequencing template shown in FIG. 59 before and after PCR for preparing an amplicon, and the analysis (C) of the sequencing result are shown. [Figure 61] A representative process included in a method for methylation analysis to distinguish cytosine, methylated cytosine, and hydroxymethylated cytosine, using DNA methyltransferase 1 (DNMT1) and conversion of methylated cytosine to DHU, is shown. [Figure 62A] The upper strand (A) and lower strand (B) of the double-stranded ligated sequencing template shown in FIG. 61 before and after PCR for preparing an amplicon, and the analysis (C) of the sequencing result are shown. [Figure 62B] The upper strand (A) and lower strand (B) of the double-stranded ligated sequencing template shown in FIG. 61 before and after PCR for preparing an amplicon, and the analysis (C) of the sequencing result are shown. [Figure 62C] The upper strand (A) and lower strand (B) of the double-stranded ligated sequencing template shown in FIG. 61 before and after PCR for preparing an amplicon, and the analysis (C) of the sequencing result are shown.

[0216] Sequence description Table 1 provides a list of specific sequences referenced herein.

[0217] [Table 1-1]

[0218] [Table 1-2] [Modes for carrying out the invention]

[0219] This specification describes polynucleotides comprising multiple insert sequences, where the insert sequences are derived from one or more target nucleic acids. These polynucleotides may include ligature sequences and multiple primer sequences. This application also describes methods for generating these polynucleotides and the uses of these polynucleotides. The presence of multiple insert sequences in a given polynucleotide can increase the output of a sequencing platform by increasing the number of reads generated per flow cell.

[0220] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. All patents, applications, published applications, and other publications referenced herein are incorporated by reference in their entirety unless otherwise stated. Where multiple definitions exist for a term herein, the one in the definitions section takes precedence unless otherwise stated. Where used herein and in the appended claims, the singular forms "a," "an," and "the" refer to multiple subjects unless the context explicitly indicates otherwise. Unless otherwise stated, conventional methods of mass spectrometry, NMR, HPLC, protein chemistry, biochemistry, recombinant DNA technology, and pharmacology are employed. The use of "or" or "and" means "and / or" unless otherwise stated. Furthermore, the use of the term "including," and other forms such as "include," "includes," and "included," is not limited to these. As used herein, the terms “comprise” and “comprising,” whether in a transitional clause or in the text of a claim, are used in the open clause. These terms should be interpreted as having the meaning of "comprising." That is, these terms should be interpreted as synonymous with the phrases "at least having" or "at least including." When used in the context of a process, the term "comprising" means that the process includes at least the listed steps, but may include additional steps. When used in the context of a compound, component, or device, the term "comprising" means that the compound, component, or device includes at least the listed features or components, but may also include additional features or components.

[0221] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described herein.

[0222] I. Definition As used herein, "hybridization sequence" or "HYB" refers to a sequence that can hybridize to a complementary hybridization sequence. Hybridization of a HYB in one library product to a HYB' in another library product can result in a hybridization adduct in which the two library products anneal to each other via HYB / HYB' hybridization.

[0223] As used herein, “concatenated nucleic acid sequencing template” means a double-stranded composition of polynucleotides and their complements. A concatenated nucleic acid sequencing template may be generated by the association of two library products by HYB / HYB' hybridization, followed by extension to generate a double-stranded template.

[0224] As used herein, "insert sequence" refers to a region of a target nucleic acid contained within a polynucleotide. A polynucleotide may contain multiple insert sequences.

[0225] As used herein, “stacked reads” or “tandem reads” relate to sequencing reads of multiple insert sequences generated from a single polynucleotide. These sequencing reads may be sequential. For example, tandem reads can be generated using a polynucleotide containing two or more insert sequences and two or more primer sequences. As used herein, “tandem read library” refers to a library of polynucleotides containing multiple insert sequences that can be used to generate tandem reads.

[0226] As used herein, "SBS" refers to a sequence incorporated into a polynucleotide to improve lead primer binding. In embodiments where the polynucleotide is prepared from a library product produced by tagmentation, the SBS may be a mosaic terminal sequence, and the SBS' may be a complement to the mosaic terminal sequence, such as ME and ME'. If the library product is produced using the Trueq method (Illumina), the SBS and SBS' sequences may also be included in the adapter.

[0227] II. Polynucleotides containing multiple insert sequences This specification describes a polynucleotide comprising multiple insert sequences, where each insert contains a portion of one or more target nucleic acids. A single polynucleotide comprising multiple insert sequences enables sequencing of multiple regions of one or more target nucleic acids in the same region of a flow cell. In this way, more regions of one or more target nucleic acids can be sequenced without requiring a larger flow cell.

[0228] In some embodiments, polynucleotides are hybridized by hybridizing the HYB in one library product to the HYB' sequence in the other library product. It is generated from two separate library products, based on the process of forming additives and then extending and ligating them to produce a nucleic acid sequencing template.

[0229] These polynucleotides may also include additional sequences, such as one or more primer sequences, ligation sequences, or attached polynucleotides.

[0230] In some embodiments, the polynucleotide comprises a 3'-terminal polynucleotide containing a first read-primer binding sequence, a ligation sequence containing a first insert sequence on the 5' side of the 3'-terminal polynucleotide derived from a target nucleic acid, and a second read-primer binding sequence orthogonal to the first read-primer binding sequence, wherein the second read-primer binding sequence contains a hybridization sequence, a second insert sequence located on the 5' side of the ligation sequence and derived from a non-adjacent sequence of the target nucleic acid or from a target nucleic acid different from the first insert sequence, and an attached polynucleotide located at the 5' end of the polynucleotide and containing an attached sequence, wherein the 3'-terminal polynucleotide, the ligation sequence, and the attached polynucleotide do not originate from the target nucleic acid.

[0231] Figure 1 presents an overview of these polynucleotides, illustrating how sequencing of two distinct insert sequences is possible by sequencing an exemplary polynucleotide with four primer sequences.

[0232] Figure 2 shows an exemplary polynucleotide structure in which the ligation sequence includes a second lead-primer-binding sequence (read 2) containing a hybridization sequence (HYB), a first lead-primer-binding sequence (read 1) that binds to a 3' polynucleotide containing a P5' sequence, and an adhesion sequence containing a P7 sequence. As shown in Figure 3, different inserts in the polynucleotide can be generated from different libraries.

[0233] Polynucleotides with multiple insert sequences can allow for the generation of a larger quantity of sequences from a flow cell compared to a standard Illumina paired-end library, as shown in Figure 4A compared to Figure 4B. In Figures 4A and 4B, since the same amount of flow cell surface was used in both cases, twice as much sequence was generated for the same area of ​​flow cell surface using polynucleotides with two insert sequences compared to polynucleotides with a single insert.

[0234] Polynucleotides that can be used as sequencing templates are also described herein. These sequencing templates can be used with any standard sequencing method known in the art.

[0235] In some embodiments, the polynucleotide includes two or more insert sequences. As used herein, “insert sequence” or “insert” refers to a region of a target nucleic acid, such as a double-stranded nucleic acid, contained within the polynucleotide. The polynucleotide may include multiple insert sequences. In some embodiments, the polynucleotide includes two insert sequences. In some embodiments, the polynucleotide includes three, four, or five insert sequences. A polynucleotide containing two or more inserts that can be used as a sequencing template may be referred herein to as a “concatenated nucleic acid sequencing template” or “concatenated sequencing template.”

[0236] In some embodiments, the polynucleotide includes a hybridization sequence or a complement of a hybridization sequence. As used herein, “hybridization sequence” or “HYB” refers to a sequence that can hybridize to a complementary hybridization sequence. For example, in one fragment (e.g., a library product) a HYB in another fragment Hybridization to HYB' (complementary to the hybridization sequence) can result in a hybridization adduct or crosslink in which the two fragments anneal to each other via HYB / HYB' hybridization. In some embodiments, HYB contains enough nucleotides to attach the two single-stranded fragments together when HYB hybridizes to HYB'. In some embodiments, the HYB sequence contained in the concatenated sequencing template can be used as a primer binding site, as shown in Figure 47.

[0237] In some embodiments, HYB or HYB' contains 10 to 30 nucleotides. In some embodiments, the binding of HYB in the first single-stranded nucleic acid fragment to HYB' in the second single-stranded nucleic acid fragment is sufficient to "crosslink" the two fragments (as described in the methods herein using examples shown in Figures 28A and 39). The nucleotides contained in HYB or HYB' may be natural, artificial, or modified nucleotides. In some embodiments, HYB or HYB' containing artificial or modified nucleotides may require fewer nucleotides in their sequences to enable crosslinking between the two single-stranded fragments.

[0238] In some embodiments, one or more nucleotides in HYB or HYB' are lock nucleic acids or crosslink nucleic acids. As used herein, “lock nucleic acid” or “LNA” refers to a modified RNA nucleotide in which the ribose portion is modified with an extra crosslink connecting the 2' oxygen and 4' carbon. In some embodiments, the LNA confers enhanced structural stability in the HYB or HYB' sequence and thus increases the hybridization melting temperature (Tm) of the HYB / HYB' interaction. For example, a HYB or HYB' sequence containing one or more LNAs may contain only a relatively short sequence (e.g., 10-20 nucleotides), but may still confer a bond strong enough to allow the formation of a crosslink between a first single-stranded fragment containing HYB and a second single-stranded fragment containing HYB'.

[0239] In some embodiments, the polynucleotide includes two or more inserts. As described herein, these inserts may be copies of the same sequence derived from the target nucleic acid, or they may be distinct sequences derived from the target nucleic acid. As used herein, “chimeric template” refers to a template containing different inserts.

[0240] A wide variety of different polynucleotides containing two inserts, such as those shown in Figures 29 and 40, are described herein. In addition to two or more inserts and hybridization sequences (or their complements), the polynucleotides of the present invention may also contain various other types of inserts.

[0241] For example, a polynucleotide may comprise one or more sequencing primer sequences. Such sequencing primer sequences may be used for primer binding to initiate sequencing when the polynucleotide is used as a sequencing template. In some embodiments, the polynucleotide comprises a first read sequencing primer sequence and / or a second read sequencing primer sequence. As used herein, “first read sequencing primer sequence” and “second read sequencing primer sequence” refer to sequences that can bind to primers that may be used with different sequencing reads. These terms are not limited to any particular sequence; for example, a first read sequencing primer sequence may be used to initiate a second sequencing read in a given experiment, and a second read sequencing primer may be used to initiate a first sequencing read in a given experiment. Such primer sequences may vary based on the sequencing platform that the user plans to use, and such primer sequences are well known in the art, such as sequences A14 (SEQ ID NO: 4) and B15 (SEQ ID NO: 5).

[0242] In some embodiments, the first read sequencing primer sequence and the second read sequencing primer sequence are different. In some embodiments, the first read sequencing primer sequence and the second read sequencing primer sequence each include the A14 sequence or the B15 sequence, or their complements. In some embodiments, the 3'-terminal polynucleotide includes the complement of the P5 primer sequence (P5') and the 5'-terminal polynucleotide includes the P7 primer sequence (P7, SEQ ID NO: 48), or the 3'-terminal polynucleotide includes the complement of the P7 primer sequence (P7') and the 5'-terminal polynucleotide includes the P5 primer sequence (P5, SEQ ID NO: 7).

[0243] In some embodiments, the 3'-terminal polynucleotide and / or the 5'-terminal polynucleotide each independently includes at least one of the following: an adapter, a barcode sequence, a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence. In other words, the polynucleotide may include additional sequences that are useful in any way the user wishes, such as sequencing.

[0244] Using the methods described herein, one insert in a polynucleotide may be prepared from a fragment containing a portion of the sense strand of the target nucleic acid, and the other insert may be prepared by extension from a fragment containing a portion of the antisense strand of the target nucleic acid. Using the methods described herein, one insert may be prepared from a fragment containing a portion of the antisense strand of the target nucleic acid, and the other insert may be prepared by extension from a fragment containing a portion of the strand of the target nucleic acid.

[0245] In some embodiments, the polynucleotide includes two insert sequences that are copies of each other. In some embodiments, the polynucleotide includes a 5'-terminal polynucleotide comprising (a) a first read sequencing primer sequence, (b) an insert sequence derived from a target nucleic acid and located at the 3' end of the 5'-terminal polynucleotide, (c) a hybridization sequence located at the 3' end of the insert sequence, (d) a copy of the insert sequence located at the 3' end of the hybridization sequence, and (e) a 3'-terminal polynucleotide containing a complement of the second read sequencing primer sequence. In some embodiments, this polynucleotide can be a sequencing template. While the two copies of the insert (i.e., the insert sequence and the copy of the insert sequence) may be expected to be identical, sequencing results may indicate that they are not identical. For example, the two copies of the insert may differ based on mismatch mutations in the target nucleic acid or on the introduction of errors during PCR amplification.

[0246] In some embodiments, the polynucleotide includes two insert sequences that are not copies of each other. In some embodiments, the two insert sequences may be different. In some embodiments, the two insert sequences included in the polynucleotide were prepared from different regions of the target nucleic acid. In some embodiments, the polynucleotide includes (a) a 5'-terminal polynucleotide containing a first read sequencing primer sequence, (b) an insert sequence derived from the target nucleic acid and located at the 3' end of the 5'-terminal polynucleotide, a first insert sequence, (c) a hybridization sequence located at the 3' end of the insert sequence, (d) a second insert sequence located at the 3' end of the hybridization sequence, and (e) a 3'-terminal polynucleotide containing a complement of the second read sequencing primer sequence. Continuity data can be determined using such a template having two different insert sequences, as described herein with respect to a method using immobilized transposomes.

[0247] The two inserts contained in the polynucleotide may be the same or different in size. In some embodiments, the inserts, which are copies, have the same number of nucleotides. The insert sequence contains 40 to 400 nucleotides, and optionally, the insert sequence contains 1000 or fewer nucleotides. In some embodiments, the pair sequencing read protocol can be performed on larger inserts, such as those containing more than 500 nucleotides.

[0248] In some embodiments, the polynucleotide is immobilized on a solid support. In some embodiments, the polynucleotide is immobilized on the solid support via a 5'-terminal polynucleotide (such as the embodiment shown in Figure 29). In some embodiments, the polynucleotide is immobilized on the solid support via the binding of an affinity moiety on the 5'-terminal polynucleotide to a binding moiety on the surface of the solid support. In some embodiments, the affinity moiety is attached to the 5'-terminal polynucleotide via a linker. In some embodiments, the affinity moiety is biotin, desthiobiotin, or dualbiotin.

[0249] In some embodiments, polynucleotides are 5'-P5-A14-insert-HYB-insert-B15'-P7'-3', or

[0250] The insert sequence has the structure 5'-P7-B15-insert-HYB'-insert-A14'-P5'-3', where HYB is the hybridization sequence and HYB' is the complement of the hybridization sequence. In some embodiments, the two insert sequences are identical copies of the same sequence, or two sequences with more than 95% sequence homology. Potential reasons for differences between the two copies of the insert sequence, such as non-standard base pairing or random errors introduced during sequencing, are described herein. Figure 40 shows a typical double-stranded polynucleotide containing two complementary linked sequencing templates. One template contains two A inserts, and the complementary strand contains two A' inserts.

[0251] In some embodiments, polynucleotides are 5'-P5-A14-insert1-HYB-insert2-B15'-P7'-3', or

[0252] The structure is 5'-P7-B15-insert1-HYB'-insert2-A14'-P5'-3', where HYB is the hybridization sequence and HYB' is the complement of the hybridization sequence. In some embodiments, insert1 and insert2 contain different sequences with little or no sequence homology. Figure 45 shows typical crosslinking methods that may be used to generate two complementary polynucleotides, each containing two different sequences.

[0253] In some embodiments, the composition comprises a polynucleotide hybridized to its complement. In some embodiments, the polynucleotide hybridized to its complement may be referred to as a double-stranded linked sequencing template. In some embodiments, the double-stranded linked sequencing template is fixed to the surface of a solid support by both of its 5' ends.

[0254] In some embodiments, a polynucleotide or a composition comprising a polynucleotide and its complement is immobilized on the surface of a solid support, where the affinity moiety is biotin, desthiobiotin, or dualbiotin, and the binding moiety is avidin or streptavidin.

[0255] A wide range of different solid supports can be used for immobilization. In some embodiments, the solid support may be beads, slides, container walls, flow cells, or nano-containers within flow cells. It is well.

[0256] In some embodiments, the linker for attaching the affinity moiety to the polynucleotide is a cleavable linker. In some embodiments, the user can release the polynucleotide from the solid support at a desired time by cleaving this cleavable linker.

[0257] A. Target nucleic acid The target nucleic acids used herein may consist of DNA, RNA, or analogues thereof. The source of the target nucleic acid may be genomic DNA, messenger RNA, or other nucleic acids derived from natural sources. In some cases, target nucleic acids derived from such sources may be amplified before use in the methods or compositions described herein.

[0258] Examples of biological samples from which target nucleic acids may be derived include mammals such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cattle, cats, dogs, primates, humans, or non-human primates; plants such as Arabidopsis thaliana, maize, sorghum, oats, wheat, rice, canola, or soybeans; algae such as Chlamydomonas reinhardtii; nematodes such as Caenorhabditis elegans; insects such as Drosophila melanogaster, mosquitoes, fruit flies, bees, or spiders; fish such as zebrafish; reptiles; amphibians such as frogs or Xenopus laevis; fungi such as dictyostelium discoideum, Pneumocystis carinii, Takifugu rubripes, yeast, Saccharamoyces cerevisiae, or Schizosaccharomyces pombe; or Plasmodium Examples include those derived from falciparum. Target nucleic acids may also be derived from prokaryotes, bacteria such as Escherichia coli, staphylococci, or mycoplasma pneumoniae; archaea; viruses such as hepatitis C virus or human immunodeficiency virus; or viloids. Target nucleic acids may be derived from homogeneous cultures or populations of the above organisms, or alternatively, for example, from a collection of several different organisms in a population or ecosystem. Nucleic acids can be isolated using methods known in the art, for example, those described in Sambrook et al, Molecular Cloning: A Laboratory Manual, 3rd edition, Cold Spring Harbor Laboratory, New York (2001) or Ausubel et al, Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1998) (each of which is incorporated herein by reference).

[0259] In some embodiments, the target nucleic acid can be obtained as fragments of one or more larger nucleic acids. Fragmentation can be carried out using any of the various techniques known in the art, including, for example, atomization, sonication, chemical cleavage, enzymatic cleavage, or physical shearing. Fragmentation can also result from the use of certain amplification techniques that produce amplicons by copying only a portion of the larger nucleic acid. For example, PCR amplification produces fragments whose size is determined by the length of the fragments between adjacent primers used for amplification.

[0260] The target nucleic acid or the population of its amplicons may have an average chain length desirable or appropriate for a particular use of the method or composition described herein. For example, the average chain length may be 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or less than 50 nucleotides. Alternatively or additionally, the average chain length may be 10 nucleotides, The chain length may be 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or greater than 100,000 nucleotides. The mean chain length of the target nucleic acid or the population of its amplicons may be within the range between the above maximum and minimum values. It will be understood that the amplicons produced at the amplification site (or otherwise prepared or used herein) may have a mean chain length within the range between the upper and lower limits selected from those exemplified above.

[0261] In some embodiments, the target nucleic acid has a relatively short mean chain length, such as less than 200 nucleotides, less than 150 nucleotides, less than 100 nucleotides, less than 75 nucleotides, less than 50 nucleotides, or less than 36 nucleotides. Sequencing of target nucleic acids with a relatively short mean chain length is not limited by read length, and increasing the number of reads can significantly increase the sequencing output. Examples of sample types with a relatively short mean chain length are cell-free DNA (cfDNA) and exome sequencing samples.

[0262] In some embodiments, the target nucleic acid is cell-free DNA (cfDNA) derived from a maternal blood sample. In some embodiments, the cfDNA is extracted from a maternal plasma sample. In some embodiments, the cfDNA is for non-invasive prenatal testing (NIPT).

[0263] In some embodiments, the target nucleic acid is an exome. In some embodiments, the exome is prepared by targeted rearrangement. In some embodiments, the exome is prepared by whole-genome enrichment. In some embodiments, the exome is prepared by hybridization-based enrichment.

[0264] In some embodiments, the target nucleic acids are DNA and RNA. Separate libraries of RNA and DNA can be prepared to produce hybrid DNA / RNA polynucleotides. In some embodiments, the polynucleotide comprises one or more inserts containing RNA and one or more inserts containing DNA. Such polynucleotides containing RNA and DNA inserts can be called “hybrid polynucleotides” and allow for the generation of multiple readouts from a single sequencing run. In some embodiments, the polynucleotides containing RNA and DNA inserts have a dual sample index that allows for self-normalization. In some embodiments, the minimum amount of DNA or RNA in the starting library determines the amount of hybrid polynucleotide to be produced.

[0265] Any known amplification techniques may be used to increase the amount of template sequence available for use in the methods described herein. Exemplary techniques include, but are not limited to, polymerase chain reactions (PCR), rolling circle amplification (RCA), multiple substitution amplification (MDA), or random primer amplification (RPA) of nucleic acid molecules having a template sequence. It will be understood that pre-use amplification of the target nucleic acid in the methods or compositions described herein is optional. Therefore, the target nucleic acid is not amplified before use in some embodiments of the methods and compositions described herein. The target nucleic acid may optionally be derived from a synthetic library. The synthetic nucleic acid may have a natural DNA or RNA composition, or an analog thereof. Solid-phase amplification methods may also be used, including, for example, cluster amplification, crosslinking amplification, or other methods described below in the context of array-based methods.

[0266] In some embodiments, the polynucleotides disclosed herein can be sequenced using any suitable nucleic acid sequencing platform. In some embodiments, the target sequence correlates with or is associated with one or more congenital or genetic disorders, pathogenicity, antibiotic resistance, or gene modifications. Sequencing is performed using short-term sequencing. These methods and compositions can be used to determine nucleic acid sequences of ndem repeats, single nucleotide polymorphisms, genes, exons, coding regions, exomes, or parts thereof. Accordingly, the methods and compositions described herein are, in no particular way, useful in cancer and disease diagnosis, prognosis and treatment, DNA fingerprinting applications (e.g., DNA data banking, crime casework), metagenomic research and discovery, Aglai genome applications, and pathogen identification and monitoring.

[0267] In some embodiments, the sample used to prepare the sequencing comprises a double-stranded nucleic acid. This double-stranded nucleic acid may be referred to as the target nucleic acid. In some embodiments, the double-stranded nucleic acid may be added to a solid support containing immobilized transposomes. In some embodiments, the double-stranded nucleic acid may be fragmented and combined with a mixture of fork-shaped adapters.

[0268] In some embodiments, the sample comprises multiple double-stranded nucleic acids.

[0269] The biological sample used in accordance with this disclosure may be of any type, including the target nucleic acid. However, the sample does not need to be completely purified and may include, for example, nucleic acids mixed with proteins, other nucleic acid species, other cellular components, and / or any other impurities. In some embodiments, the biological sample comprises a mixture of nucleic acids, proteins, other nucleic acid species, other cellular components, and / or any other impurities present in proportions similar to those found in vivo. For example, in some embodiments, the components are found in proportions similar to those found in intact cells. In some embodiments, the biological sample has a 260 / 280 absorbance ratio of 2.0, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, 0.7, or 0.60 or less. In some embodiments, the biological sample has a 260 / 280 absorbance ratio of at least 2.0, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, 0.7, or 0.60. The method provided herein allows nucleic acids to be bound to a solid support so that other contaminants can be removed simply by washing the solid support after surface binding tagmentation has occurred. The biological sample may include, for example, crude cell lysates or whole cells. For example, crude cell lysates applied to a solid support in the method herein do not require one or more of the conventional separation steps used to isolate nucleic acids from other cellular components. Exemplary separation steps are described in Maniatis et al., Molecular Cloning: A Laboratory Manual, 2nd Edition, 1989, and Short Protocols in Molecular This is described in Biology, ed. Ausubel, et al., and is incorporated herein by reference.

[0270] In some embodiments, the sample applied to the solid support has an absorbance ratio of 260 / 280 which is 1.7 or less.

[0271] Therefore, in some embodiments, the biological sample may include, for example, blood, plasma, serum, lymph, mucus, sputum, urine, semen, cerebrospinal fluid, bronchial aspirate, feces, and their macerated tissue or lysates, or any other biological specimen containing nucleic acids.

[0272] In some embodiments, the sample is blood. In some embodiments, the sample is cell lysate. In some embodiments, the cell lysate is crude cell lysate. In some embodiments, the method further includes the step of applying the sample to a solid support and then lysing the cells in the sample to produce cell lysate.

[0273] In some embodiments, the sample is a biopsy sample. The sample is either liquid or solid. In some embodiments, a biopsy sample from a cancer patient is used to evaluate the target sequence to determine whether the subject has a specific mutation or variant in the predicted gene.

[0274] In some embodiments, the sample contains target double-stranded DNA. In some embodiments, the DNA is genomic DNA. In some embodiments, the DNA is cell-free DNA (cfDNA). In some embodiments, the DNA is circulating tumor DNA (ctDNA).

[0275] In some embodiments, the DNA is double-stranded cDNA prepared from RNA. In some embodiments, the RNA is mRNA. In some embodiments, the RNA includes coding, untranslated region (UTR) sequences, intron sequences, and / or intergenetic sequences.

[0276] B. 3'-terminal polynucleotide In some embodiments, the 3' terminal polynucleotide includes a first read primer binding sequence.

[0277] In some embodiments, the 3'-terminal polynucleotide includes at least one of a barcode sequence, a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence. In some embodiments, the 3'-terminal polynucleotide and / or attached polynucleotide each independently include at least one of a barcode sequence, a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence.

[0278] In some embodiments, the 3'-terminal polynucleotide comprises the ME', B15', and / or P7' sequences.

[0279] In some embodiments, the 3'-terminal polynucleotide comprises a complement (P5') of the P5 primer sequence, and the attached polynucleotide comprises a P7 primer sequence (P7). In some embodiments, the 3'-terminal polynucleotide comprises a complement (P7') of the P7 primer sequence, and the attached polynucleotide comprises a P5 primer sequence (P5).

[0280] In some embodiments, the 3' terminal polynucleotide contains the ME'-B15'-P7' sequence.

[0281] C. Insert Array The insert sequences contained within the polynucleotides include sequences derived from the target nucleic acid. Therefore, the polynucleotides described herein can be used for several purposes, such as generating tandem reads during sequencing.

[0282] The polynucleotides described herein include two or more insert sequences. In some embodiments, the polynucleotide includes 2, 3, 4, 5, 6, 7, 8, 9, 10, or more insert sequences. In some embodiments, the polynucleotide includes two insert sequences. In some embodiments, the polynucleotide includes three insert sequences.

[0283] The insert sequence may originate from one or more target nucleic acids.

[0284] In some embodiments, polynucleotides are multiple ingredients derived from multiple target nucleic acids. Includes a sart sequence.

[0285] In some embodiments, a polynucleotide may contain multiple insert sequences, all derived from the same target nucleic acid. In some embodiments, the multiple insert sequences are derived from non-adjacent sequences of the target nucleic acid. Non-adjacent sequences mean that the multiple insert sequences in the polynucleotide are not adjacent to each other in the original target nucleic acid. In some embodiments, the multiple insert sequences are derived from random regions of the target nucleic acid. In some embodiments, the method for generating the polynucleotide of the present invention does not select specific insert sequences.

[0286] In some embodiments, the multiple insert sequences each contain 40 to 400 nucleotides, or 100 to 200 nucleotides, or 150 nucleotides. In some embodiments, the first insert sequence and the second insert sequence each contain 40 to 400 nucleotides, or 100 to 200 nucleotides, or 150 nucleotides.

[0287] In some embodiments, the polynucleotide comprises three or more insert sequences. In some embodiments, the polynucleotide comprises at least one insertion unit between a second insert sequence and an attached polynucleotide, the insertion unit comprising a 5'-terminal insert sequence derived from a non-adjacent sequence of the target nucleic acid or from a target nucleic acid different from the other insert sequences, and a 3'-terminal ligation sequence comprising a read-primer binding sequence, wherein the read-primer binding sequence is orthogonal to the other read-primer binding sequences.

[0288] In embodiments where the polynucleotide includes three or more insert sequences, the polynucleotide may include multiple different linking sequences, each linking sequence including a primer sequence, and the primer sequences included in different linking sequences are different. In some embodiments, one or more primer sequences include a hybridization sequence, and the hybridization sequences are different in different primer sequences.

[0289] For example, two different HYB / HYB' sequence pairs, such as HYB1 / HYB1' and HYB2 / HYB2', can be used to generate a polynucleotide containing three insert sequences. Insert 1 and insert 2 can be linked using HYB1 / HYB1', and insert 2 and insert 3 can be linked using HYB2 / HYB2'. The fork-type adapter of insert 1 may contain P5 and HYB1, the adapter of insert 2 may contain HYB1' and HYB2, and the adapter of insert 3 may contain HYB2' and P7'.

[0290] Insert sequences can be generated by several methods for producing nucleic acid fragments, such as tagmentation or fragmentation.

[0291] D. Adapter Array In some embodiments, the polynucleotide may include one or more adapter sequences.

[0292] The adapter sequence may include one or more functional sequences or components selected from the group consisting of primer sequences, anchor sequences, universal sequences, spacer regions, index sequences, capture sequences, barcode sequences, cleavage sequences, sequencing-related sequences, and combinations thereof. In some embodiments, the adapter sequence includes a primer sequence. In other embodiments, the adapter sequence includes a primer sequence and an index or barcode sequence. The primer sequence may also be a universal sequence. This disclosure is not limited to the types of adapter sequences that can be used, and those skilled in the art will recognize additional sequences that can be used for library preparation and next-generation sequencing. A universal sequence is a region of nucleotide sequence common to two or more nucleic acid fragments. Optionally, two or more nucleic acid fragments may also have regions of sequence difference. A universal sequence that may be present in different members of multiple nucleic acid fragments can enable the replication or amplification of multiple different sequences using a single universal primer complementary to the universal sequence.

[0293] In some embodiments, the first lead-primer binding sequence includes a first adapter sequence. In some embodiments, the first adapter sequence is a complement of the A14 primer sequence (A14') or a complement of the B15 primer sequence (B15').

[0294] In some embodiments, the adapter sequence includes an SBS or SBS' sequence. In some embodiments, the SBS or SBS' sequence may include all or part of a standard sequence contained in an oligonucleotide used in the Trueq workflow, so that a standard sequence primer can be used. In some embodiments, the SBS may be a mosaic-ended sequence, and the SBS' may be a complement of a mosaic-ended sequence such as ME and ME'.

[0295] In some embodiments, the SBS or SBS' sequence may include A14-ME or B15-ME, or their complements. Sequence IDs 15-21 show some exemplary SBS or SBS' sequences, or adapters containing SBS or SBS' sequences.

[0296] In some embodiments, SBS and SBS' are all or partially complementary sequences that can form an adapter duplex. In some embodiments, SBS and SBS' are partially complementary. In some embodiments, SBS and SBS' are fully complementary. In some embodiments, SBS and / or SBS' include a 13-base pair sequence. In some embodiments, the adapter duplex includes P5-HYB' and P7-HYB in addition to SBS or SBS'. Thus, for example, if two library fragments are stacked together (i.e., in tandem) to produce a polynucleotide with two inserts, the resulting polynucleotide can be sequenced using standard sequencing primers.

[0297] In some embodiments, the adapter sequence has a melting temperature of 65°C or higher for binding to the sequencing primer. In some embodiments, the adapter sequence binds to the sequencing primer in such a way that the binding is not lost at the temperature used for sequencing. In some embodiments, the adapter sequence contains a substantial amount (more than 10%) of each of A, T, C, and G. In some embodiments, the G / C content of the adapter sequence is 40% to 60%. In some embodiments, the G / C content of the adapter sequence is 30% or more and 70% or less. In some embodiments, the G / C content of the adapter sequence is 40% or more and 50% or less, or 50% or more and 60% or less.

[0298] In some embodiments, the attached polynucleotide includes a second adapter sequence. In some embodiments, the second adapter sequence is either the A14 sequence or the B15 sequence.

[0299] In some embodiments, the first adapter sequence is the complement (A14') of the A14 sequence, and the second adapter sequence is the B15 sequence. In some embodiments, the first adapter sequence is the complement (B15') of the B15 sequence, and the second adapter sequence is the A14 sequence.

[0300] In some embodiments, the adapter sequence is transcribed to the 5' end of the nucleic acid fragment by a tagmentation reaction.

[0301] E. Linked Arrays In some embodiments, the linking sequence includes a second read-primer binding sequence orthogonal to the first read-primer binding sequence, and the second read-primer binding sequence includes a hybridization sequence. In some embodiments, the hybridization sequence is HYB'. In some embodiments, the second read-primer binding sequence includes a hybridization sequence (HYB) and a complement (ME') of the SBS' sequence, as shown in Figure 4B. In some embodiments, the fourth read-primer binding sequence includes a complement (HYB') of the hybridization sequence and a complement (SBS') of the SBS sequence, as shown in Figure 4B.

[0302] In some embodiments, the linked sequence includes a complement of the 3' transposon terminal sequence of the hybridization sequence and the 5' transposon terminal sequence of the hybridization sequence.

[0303] In some embodiments, the concatenation sequence includes ME', HYB', and / or ME. In some embodiments, the concatenation sequence includes ME', HYB', and ME. In some embodiments, the concatenation sequence is ME'-HYB'-ME.

[0304] In some embodiments, the second read-primer binding sequence includes a complement to the hybridization sequence and a complement to the transposon terminal sequence. In some embodiments, the second read-primer binding sequence includes HYB' or ME'. In some embodiments, the second read-primer binding sequence includes HYB' and ME'. In some embodiments, the second read-primer binding sequence is HYB'-ME'.

[0305] F. Immobilization and attachment of polynucleotides In some embodiments, the polynucleotides are immobilized on a solid support.

[0306] In some embodiments, the polynucleotides are immobilized on a solid support via attached polynucleotides. In some embodiments, the attached polynucleotides include an attached sequence.

[0307] In some embodiments, the attached polynucleotide includes an attached sequence. In some embodiments, the attached sequence is a nucleic acid sequence that hybridizes to a transposon in a transposome complex and is immobilized on a solid support such as a slide, flow cell, or bead. In some embodiments, the attached sequence functions to attach the transposome complex to the solid support. In some embodiments, the attached sequence functions to attach the polynucleotide to the solid support. In some embodiments, the attached sequence is P5.

[0308] In some embodiments, polynucleotides are immobilized on a solid support via hybridization of the attached polynucleotide to an attached polynucleotide complement on the surface of the solid support. In some embodiments, polynucleotides are immobilized on a solid support via binding of affinity moieties on the attached polynucleotide to binding moieties on the surface of the solid support.

[0309] In some embodiments, the solid support includes a flow cell or beads.

[0310] In some embodiments, the attached polynucleotide includes at least one of a barcode sequence, a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence.

[0311] In some embodiments, the attached polynucleotide includes a second adapter sequence. In some embodiments, the second adapter sequence is A14 or B15.

[0312] In some embodiments, the attached polynucleotide includes a transposon terminal sequence. In some embodiments, the transposon terminal sequence is ME.

[0313] In some embodiments, the attached sequence is P5, the second adapter sequence is A14, and / or the transposon terminal sequence is ME. In some embodiments, the attached polynucleotide comprises P5, A14, and / or ME. In some embodiments, the attached polynucleotide comprises P5, A14, and ME. In some embodiments, the attached polynucleotide comprises P5-A14-ME.

[0314] G. Sample Index and UMI In some embodiments, the polynucleotide includes a hybridization sequence (or its complement) and at least two inserts, in addition to a primer sequence, an index sequence, a barcode sequence, a purified tag, or any combination thereof. In some embodiments, the polynucleotide includes a sample index and / or a unique molecular identifier (UMI). In some embodiments, one or more of these sequences are incorporated into the polynucleotide using a fork-shaped adapter ligated to the double-stranded fragment, or using a fork-shaped adapter contained within a transposome that is incorporated into the double-stranded fragment during tagmentation. Alternatively, additional sequences may be added to the polynucleotide (e.g., a ligated sequencing template) after the polynucleotide has been generated using PCR or the like.

[0315] A unique molecular identifier (UMI) is a nucleotide sequence applied to or used to identify a nucleic acid molecule, which can be used to distinguish individual nucleic acid molecules from one another. UMIs may be sequenced together with the relevant nucleic acid molecules to determine whether a read sequence belongs to one source nucleic acid molecule or another. The term "UMI" may be used herein to refer to both the sequence information of a polynucleotide and the physical polynucleotide itself. While a UMI is similar to a barcode commonly used to distinguish a read from one sample from a read from another, a UMI is instead used to distinguish a nucleic acid template fragment from another fragment when many fragments from an individual sample are sequenced together. UMIs can be defined in many ways, as described in International Publication 2019 / 108972 and International Publication 2018 / 136248, which are incorporated herein by reference.

[0316] In some embodiments, a unique dual index (UDI) is prepared using two sample indices. In some embodiments, the sample indices are the i5-i8 sequences. Alternatively, the i6 and i8 sequences may be used as the UMI.

[0317] UMIs are useful for removing PCR replicas and detecting low-frequency variants in double-stranded nucleic acids, while UDIs are useful for mitigating inaccurate sample assignment due to index hopping in library sequencing and demultiplexing. UDIs, such as unique i5 and i7 index sequences, can be appended to the ends of target nucleic acids so that both ends contain the UDI. UDIs can be used with patterned flow cells such as Illumina's NovaSeq 6000 system (e.g., International Publication Nos. 2018 / 204423, 2018 / 208699, 2019 / 055715, and 2019 / 055715). See publication 2016 / 176091 (they are incorporated herein in their entirety by reference). In some embodiments described herein, such as those shown in Figures 46A and 46B, transposons contained in different pools of transposomal complexes are designed to prepare polynucleotides that incorporate UDI or UMI during tagmentation, eliminating the need for a separate PCR step for incorporating UDI or UMI. Exemplary polynucleotides containing UDI (e.g., i5 and i7) or UMI (e.g., i6 or i8) are shown in Figures 46A–46C.

[0318] Composition comprising H. polynucleotides and their complements In some embodiments, the composition comprises a polynucleotide and its complement. In some embodiments, the polynucleotide hybridizes with its complement. In some embodiments, the polynucleotide and its complement are included in a double-stranded composition.

[0319] In some embodiments, the composition comprises a polynucleotide and its complement, the complement being a 3'-terminal complement comprising a first complementary read-primer binding sequence, the first complementary read-primer binding sequence comprising a 3'-terminal complement orthogonal to the first and second read-primer binding sequences, a complement of a second insert sequence located on the 5' side of the 3'-terminal complement, and a complementary linking sequence located on the 5' side of the complement of the second insert sequence, comprising a second complementary read-primer binding sequence from 3' to 5', the second complementary read-primer binding sequence comprising a complementary linking sequence orthogonal to the first and second read-primer binding sequences and the first complementary read-primer binding sequence, a complement of a first insert sequence located on the 5' side of the complementary linking sequence, and a complementary attached polynucleotide comprising a complementary attached sequence at its 5' end.

[0320] In some embodiments, the composition comprises a polynucleotide and a complement, with either the polynucleotide or the complement immobilized on a solid support. In some embodiments, the composition comprises a polynucleotide immobilized on a solid support via a first attached polynucleotide. In some embodiments, the complement is immobilized on the solid support via a complementary attached polynucleotide.

[0321] In some embodiments, the complementary attached polynucleotide includes an attached sequence. In some embodiments, the complementary sequence included in the complementary attached polynucleotide is P7.

[0322] In some embodiments, the complementary adherent polynucleotide includes the ME-B15-P7 sequence. In some embodiments, the complementary adherent sequence includes P7. In some embodiments, the complementary ligation sequence includes ME-HYB-ME. In some embodiments, the second read complementary primer sequence includes HYB-ME'. In some embodiments, the 3'-terminal polynucleotide complement includes P5'-A14'-ME'. In some embodiments, the first read complementary read-primer binding sequence includes A14'-ME'. In some embodiments, the complementary hybridization sequence includes HYB.

[0323] I. Structure of polynucleotides or compositions Polynucleotides can have a variety of structures. In some embodiments, the composition comprises one polynucleotide or its complement from the following structures.

[0324] In some embodiments, polynucleotides are It has the structure 3'-P7'-B15'-ME'-insert1-ME-HYB-ME'-insert2-ME-A14-P5-5'.

[0325] In some embodiments, the polynucleotide complement is 3'-P5'-A14'-ME'-Insert2-ME-HYB'-ME'-Insert It has the structure of 1-ME-B15-P7-5'.

[0326] J. Polynucleotide-containing kit In some embodiments, the kit or composition comprises a first transposome complex and a second transposome complex, wherein the first transposome complex comprises a transposon containing a hybridization sequence complement, and the second transposome complex comprises a transposon containing a hybridization sequence.

[0327] In some embodiments, the composition or kit comprises a solid support, optionally the optional support being beads; a component for generating a transposome complex, comprising a transposase; oligonucleotides for generating oligonucleotide doubles, wherein the first oligonucleotide comprises a 3' transposon terminal sequence and a 5' first adapter sequence, and the second oligonucleotide comprises a 5' transposon terminal sequence and a 3' second adapter sequence, wherein the 5' transposon terminal sequence is complementary to the 3' transposon terminal sequence, and the first and second adapter sequences are not the same; and first and second primer sets for adding an attachment sequence and a hybridization sequence to a fragment by PCR, wherein the first primer set comprises a primer for adding a hybridization sequence and a first attachment sequence to the fragment, and the second primer set comprises a primer for adding a complementary hybridization sequence and a second attachment sequence to the fragment, and the first and second attachment sequences are not the same.

[0328] In some embodiments, the kit or composition includes one or more fork-type adapter complexes. In some embodiments, the kit or composition includes a first fork-type adapter complex and a second fork-type adapter complex.

[0329] In some embodiments, the kit or composition includes one or more assembled adapter twin strands. In some embodiments, the kit or composition includes an assembled adapter twin strand comprising a first adapter twin strand and a second adapter twin strand.

[0330] In some embodiments, the kit or composition includes a fork-type adapter complex and an assembled adapter twin.

[0331] In some embodiments, the kit or composition comprises an assembled enzyme and a transposon.

[0332] In some embodiments, the kit or composition includes a purified oligonucleotide.

[0333] III. Method for preparing polynucleotides containing multiple insert sequences The polynucleotides described herein can be produced using a variety of methods.

[0334] A. Methods involving rearrangement reactions In some embodiments, polynucleotides are prepared by a method involving a rearrangement reaction.

[0335] Transposition reactions are reactions in which one or more transposons are inserted into a target nucleic acid at random or nearly random sites. Components of a transposition reaction include transposases (or other enzymes capable of fragmenting and labeling nucleic acids as described herein (e.g., integrases)), and double-stranded transposons bound to transposases (or other enzymes as described herein). Examples include transposon elements containing transposon terminal sequences, and adapter sequences attached to one of the two transposon terminal sequences. One strand of the double-stranded transposon terminal sequence is transcribed to one strand of the target nucleic acid, and the complementary transposon terminal sequence strand is not a (non-transcribed transposon sequence). The adapter sequence may optionally or desiredly include one or more functional sequences or components (e.g., primer sequences, anchor sequences, universal sequences, spacer regions, or index tag sequences).

[0336] Transposon-based technologies can be used to fragment DNA, for example, as exemplified in the workflow of the NEXTERA® FLEX DNA Sample Preparation Kit (Illumina, Inc.), where target nucleic acids, such as genomic DNA, are processed with a transposomal complex that simultaneously fragments and tags the target ("tagmentation"), thereby generating a population of fragmented nucleic acid molecules tagged with a specific adapter sequence at the ends of the fragments.

[0337] Figures 6A to 9B illustrate various approaches for preparing library products containing HYB or HYB' sequences using rearrangement reactions. In some embodiments, bead-linked transpososomes (BLTs) are used. In some embodiments, transpososomes in solution are used.

[0338] A “transposomal complex” consists of at least one transposase (or other enzyme described herein) and a transposon recognition sequence. In some such systems, the transposase binds to the transposon recognition sequence to form a functional complex capable of catalyzing the transposition reaction. In some embodiments, the transposon recognition sequence is a double-stranded transposon terminal sequence. The transposase binds to a transposase recognition site in the target nucleic acid and inserts the sequence transposon recognition sequence into the target nucleic acid. In some such insertion events, one strand of the transposon recognition sequence (or terminal sequence) is transcribed into the target nucleic acid, resulting in a cleavage event. Exemplary transposition procedures and systems can be readily adapted for use with transposases.

[0339] Exemplary transposases that may be used in certain embodiments provided herein include (or are encoded by) Tn5 transposase, Sleeping Beauty (SB) transposase, Vibrio harveyi, MuA transposase and Mu transposase recognition sites including R1 and R2 terminal sequences, Staphylococcus aureus Tn552, Ty1, Tn7 transposase, Tn / O and IS10, Mariner transposase, Tc1, P Element, Tn3, bacterial insert sequences, retroviruses, and yeast retrotransposons. More examples include IS5, Tn10, Tn903, IS911, and designed versions of transposase family enzymes. The methods described herein also include combinations of transposases, not just single transposases.

[0340] In some embodiments, the transposase is Tn5, Tn7, MuA, or Vibrioharvey transposase, or active variants thereof. In other embodiments, the transposase is Tn5 transposase or a variant thereof. In other embodiments, the transposase is Tn5 transposase or a variant thereof. In other embodiments, the transposase is Tn5 transposase or an active variant thereof. In some embodiments, the Tn5 transposase is a highly active Tn5 transposase, or an active variant thereof. In some embodiments, the Tn5 transposase is a Tn5 transposase described in PCT International Publication 2015 / 160895, which is incorporated herein by reference. In some embodiments, the Tn5 transposase is wild-type Tn The Tn5 transposase is a highly active Tn5 having mutations at positions 54, 56, 372, 212, 214, 251, and 338 of the Tn5 transposase. In some embodiments, the Tn5 transposase is a highly active Tn5 having the following mutations from the wild-type Tn5 transposase: E54K, M56A, L372P, K212R, P214R, G251R, and A338V. In some embodiments, the Tn5 transposase is a fusion protein. In some embodiments, the Tn5 transposase fusion protein contains a fusion elongation factor Ts(Tsf) tag. In some embodiments, the Tn5 transposase is a highly active Tn5 transposase containing mutations at amino acids 54, 56, and 372 of the wild-type sequence. In some embodiments, the highly active Tn5 transposase is a fusion protein, and optionally, the fusion protein is the elongation factor Ts(Tsf). In some embodiments, the recognition site is a Tn5 transposase recognition site (Goryshin and Reznikoff, J. Biol. Chem., 273:7367, 1998). In one embodiment, a transposase recognition site that forms a complex with an overactive Tn5 transposase is used (e.g., EZ Tn5® Transposase, Epicentre Biotechnologies, Madison, Wis.). In some embodiments, the Tn5 transposase is a wild-type Tn5 transposase.

[0341] In some embodiments, the transposome complex comprises a dimer of two molecules of transposase. In some embodiments, the transposome complex is homodimer, and the two molecules of transposase are each bound to a first and second transposon of the same type (for example, the sequences of the two transposons bound to each monomer are the same and form a “homodimer”). In some embodiments, the compositions and methods described herein employ two populations of transposome complexes. In some embodiments, the transposons in each population are the same. In some embodiments, the transposome complexes in each population are homodimer, and the first population has a first adapter sequence in each monomer, while the second population has different adapter sequences in each monomer.

[0342] In some embodiments, the transposase complex comprises a transposase (e.g., Tn5 transposase) dimer containing first and second monomers. In some embodiments, each monomer comprises a first transposon, a second transposon, and an attached polynucleotide, wherein the first transposon comprises a transposon terminal sequence (also called the 3' transposon terminal sequence) at its 3' end and an adapter sequence (also called the 5' adapter sequence) at its 5' end, the second transposon comprises a transposon terminal sequence (also called the 5' transposon terminal sequence) at its 5' end and an adapter sequence (also called the 3' adapter sequence) at its 3' end, and the attached polynucleotide comprises an attached adapter sequence that hybridizes to the 5' adapter sequence of the first transposon, a primer sequence, and a linker. In some embodiments, the 5' transposon terminal sequence of the second transposon is at least partially complementary to the 3' transposon terminal sequence of the first transposon. In some embodiments, the attached adapter sequence of the attached polynucleotide is at least partially complementary to the 5' adapter sequence of the first transposon. In some embodiments, the linker of the attached polynucleotide includes a binding element.

[0343] 1. Transposome complex In some embodiments, the transposome complex comprises a first transposon including a complement of a first read-primer binding sequence, wherein the complement of the first read-primer binding sequence includes a 3' portion containing a transposon terminal sequence and a complement of a first adapter sequence, and a 5' portion containing the complement of the transposon terminal sequence. It includes a second transposon and a complementary hybridization sequence. In some embodiments, the first read primer binding sequence includes a first read sequencing adapter sequence.

[0344] In some embodiments, the 3' transposon terminal sequence includes a mosaic terminal (ME) sequence, and the 5' transposon terminal sequence includes an ME' sequence.

[0345] In some embodiments, the complement of the first adapter sequence is the B15 sequence.

[0346] In some embodiments, the first read primer binding sequence is ME'-B15'.

[0347] In some embodiments, the second transposon includes a complementary attachment sequence on the 5' side of the first read-primer binding sequence. In some embodiments, the complementary attachment sequence includes a P7 sequence.

[0348] In some embodiments, the transposome complex has the following structure:

[0349] [ka]

[0350] In some embodiments, the targeted transpososome complex comprises a transposase, a first transposon comprising an attached polynucleotide, the attached polynucleotide comprising a 5' portion comprising an attached sequence, a second 3' portion comprising a read-primer binding sequence, the 3' portion comprising a transposon terminal sequence, an adapter, a second transposon comprising a 5' portion comprising a complement of the transposon terminal sequence, and a hybridization sequence.

[0351] In some embodiments, the adapter is an A14 sequence. In some embodiments, the attachment sequence includes a P5 sequence.

[0352] In some embodiments, the transposome complex has the following structure:

[0353] [ka]

[0354] In some embodiments, the first and second transposons described herein are annealed to each other, and the first transposon is annealed to an attached polynucleotide. The annealed polynucleotide is then loaded onto a transposase such as Tn5 transposase, thereby forming a transposome complex, which is then brought into contact with and bound to a solid support such as a bead. In some embodiments, the annealed transposon is bound to a solid support such as a bead, and then the transposase complexes with the transposon, thereby producing a transposome that is bound to the solid support.

[0355] 2. Terminal sequence In some embodiments, the first transposon includes a 3' transposon terminal sequence, and the second transposon includes a 5' transposon terminal sequence. In some embodiments, the 5' transposon terminal sequence is at least partially complementary to the 3' transposon terminal sequence. In some embodiments, the complementary transposon terminal sequence hybridizes to form a double-stranded transposon terminal sequence that binds to a transposase (or other enzyme described herein). In some embodiments, the transposon terminal sequence is a mosaic terminal (ME) sequence. Thus, in some embodiments, the 3' transposon terminal sequence is an ME sequence, and the 5' transposon terminal sequence is an ME' sequence.

[0356] 3. Adapter Array As described above in Section II.D, in any embodiment of the method described herein, the first transposon comprises a 5' adapter sequence, and the second transposon comprises a 3' adapter sequence. In some embodiments, the attached polynucleotide comprises an attached adapter sequence that hybridizes to the 5' adapter sequence. In some embodiments, the attached adapter sequence is at least partially complementary to the 5' adapter sequence. In some embodiments, the adapter sequence is an A14 sequence or a B15 sequence. Thus, in some embodiments, the 5' adapter sequence is an A14 sequence, and the attached adapter sequence is an A14' sequence. In some embodiments, the 3' adapter sequence is a B15' sequence.

[0357] In any embodiment, adapter sequences or transposon terminal sequences including A14-ME, ME, B15-ME, ME', A14, B15, and ME are provided below: A14-ME:5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3'(Sequence ID 1) B15-ME:5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3'(Sequence ID 2) ME':5'-phoS-CTGTCTCTTATACACATCT-3'(Sequence ID 3) A14:5'-TCGTCGGCAGCGTC-3'(Sequence ID 4) B15:5'-GTCTCGTGGGCTCGG-3'(Sequence ID 5) ME:AGATGTGTATAAGAGACAG (Sequence ID 6)

[0358] 4. Immobilized transposomes and solid supports In some embodiments, the transposome complex is immobilized on a solid support via a first or second transposon. In some embodiments, the transposome complex is immobilized on a bead. In some embodiments, the transposome complex is immobilized on a bead via a first or second transposon.

[0359] The terms “solid surface,” “solid support,” and other grammatical equivalents refer to any material that is suitable for, or can be modified to be suitable for, the binding of transposome complexes. As understood in the art, there are many possible substrates. Possible substrates include glass and modified or functionalized glass, plastics (acrylic, polystyrene, and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon®, etc.), polysaccharides, polyhedral organic silsesquioxane (POSS) materials, nylon or nitrocellulose, ceramics, resins, silica, or silicon and modified silicon, carbon, metals, inorganic glass, plastics, fiber optic bundles, beads, paramagnetic beads, and silica-based materials including various other polymers.

[0360] In some embodiments, the transposomal complex comprises a binding element (and an optional phosphorus). It is immobilized on a solid support via a ker. In some embodiments, the solid support is beads, paramagnetic beads, flow cells, the surface of a microfluidic device, a tube, a plate well, a slide, a patterned surface, or microparticles. In some embodiments, the solid support contains or is beads. In one embodiment, the beads are paramagnetic beads. In some embodiments, the solid support comprises multiple solid supports. In some embodiments, the transposome complex is immobilized on multiple solid supports. In some embodiments, the multiple solid supports contain multiple beads. In some embodiments, the multiple transposome complex is 1 mm 2 At least 10 3 , 10 4 , 10 5 , 10 6 The complexes are immobilized on a solid support at a density of 10,000. In some embodiments, the solid support is a bead or a paramagnetic bead, and each bead contains 10,000, 20,000, 30,000, 40,000, 50,000, or more than 60,000 transposome complexes bound to it.

[0361] Suitable bead compositions include plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, triasol, and carbon black. Examples of materials for solid supports include, but are not limited to, lead, titanium dioxide, latex, or Sepharose, cellulose, nylon, crosslinked micelles, and crosslinked dextrans such as Teflon®, as well as any other materials outlined herein. In certain embodiments, microspheres are magnetic microspheres or beads, e.g., paramagnetic particles, spheres, or beads. The beads do not need to be spherical. Irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead size ranges from nanometers, e.g., 100 nm, to millimeters, e.g., 1 mm, with beads preferably 0.2 to 200 micrometers, and particularly preferably 0.5 to 5 micrometers, although in some embodiments smaller or larger beads may be used. The beads may be coated with a binding partner, for example, the beads may be coated with streptavidin. In some embodiments, the beads are streptavidin-coated paramagnetic beads, such as Dynabeads MyOne streptavidin C1 beads (Thermo Scientific catalog number 65601), Streptavidin MagneSphere Paramagnetic particles (Promega catalog number Z5481), Streptavidin Magnetic beads (NEB catalog number S1420S), and MaxBead Streptavidin (Abnova catalog number U0087). The solid support may also be a slide, such as a flow cell or other slide modified so that a transposome complex can be immobilized on it.

[0362] In some embodiments, the binding partners are present on a solid support or beads at densities of 1000-6000 pmol / mg, 2000-5000 pmol / mg, 3000-5000 pmol / mg, or 3500-4500 pmol / mg.

[0363] In some embodiments, the solid surface is the inner surface of the sample tube. In some embodiments, the solid surface is a capture membrane. In one embodiment, the capture membrane is a biotin capture membrane (e.g., available from Promega Corporation). In some embodiments, the capture membrane is filter paper. In some embodiments of this disclosure, a solid support consisting of an inert substrate or matrix (e.g., a glass slide, polymer beads, etc.) is functionalized by applying a layer or coating of an intermediate material containing reactive groups that enable covalent bonding to molecules such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass, particularly the polyacrylamide hydrogels described in International Publication No. 2005 / 065814 and U.S. Patent Application Publication No. 2008 / 0280773, the contents of which are incorporated herein by reference in their entirety. Methods for tagging (fragmenting and tagging) DNA on a solid surface for construction are described in International Publication 2016 / 189331 and U.S. Patent Application Publication 2014 / 0093916(A1), which are incorporated herein by reference in their entirety. In some embodiments, the transposome complex described herein is immobilized on a solid support via a binding element. In some such embodiments, the solid support contains streptavidin as the binding partner and biotin as the binding element.

[0364] In some embodiments, transposome complexes are immobilized on a solid support, such as beads, at a specific density or density range. In some embodiments, the density of the complex on the solid support refers to the concentration of the transposome complex in the solution during the immobilization reaction. The complex density assumes that the immobilization reaction is quantitative. Once the complex is formed at a specific density, the density remains constant in a batch of surface-bound transposome complexes. The resulting beads can be diluted, and the resulting concentration of the complex in the diluted solution is obtained by dividing the prepared density of the beads by the dilution factor. Diluted bead stocks retain the complex density from their preparations, but the complex is present at a lower concentration in the diluted solution. The dilution step does not change the density of the complex on the beads and therefore affects the library yield, but does not affect the size of the inserts (fragments). In some embodiments, the density is 5 nM to 1000 nM, or 5 to 150 nM, or 10 nM to 800 nM. In other embodiments, the density is 10 nM, or 25 nM, or 50 nM, or 100 nM, or 200 nM, or 300 nM, or 400 nM, or 500 nM, or 600 nM, or 700 nM, or 800 nM, or 900 nM, or 1000 nM. In some embodiments, the density is 100 nM. In some embodiments, the density is 300 nM. In some embodiments, the density is 600 nM. In some embodiments, the density is 800 nM. In some embodiments, the density is 100 nM. In some embodiments, the density is 1000 nM.

[0365] In some embodiments, the composition comprises a solid support and a transposome complex immobilized on the solid support. In some embodiments, the transposome complex comprises a transposon, a first transposon, an attached polynucleotide, and a second transposon. In some embodiments, the first transposon comprises a 3' transposon terminal sequence and a 5' adapter sequence. In some embodiments, the attached polynucleotide comprises a 5' adapter sequence and an attached adapter sequence that hybridizes to a binding element. In some embodiments, the second transposon comprises a 5' transposon terminal sequence and a 3' adapter sequence. In some embodiments, the transposome complex is immobilized on the solid support via the attached polynucleotide. In some embodiments, the attached polynucleotide further comprises a primer sequence.

[0366] In some embodiments, the binding element contains or is optionally substituted biotin. In some embodiments, the binding element is linked to the attached polynucleotide via a linker. In some embodiments, the binding element contains or is a biotin linker. In some embodiments, the binding element contains or is a 3', 5', or internal biotin.

[0367] Some embodiments of the transposome complex described herein include an attached polynucleotide. When used herein, the attached polynucleotide is a polynucleotide that hybridizes to the transposon at one end and binds to the surface at the second end. Thus, the transposome complex described herein is immobilized on a solid support via the attached polynucleotide. In some embodiments, the attached polynucleotide is an adapter sequence of the first transposon or an adapter sequence of the second transposon, ply It includes a mer array and an adhesive adapter array that hybridizes to the linker. In some embodiments, the linker includes a coupling element.

[0368] As described herein, the attachment adapter sequence may be at least partially complementary to the adapter sequence of the first or second transposon. In some embodiments, the attachment adapter sequence hybridizes to the 5' adapter sequence. In embodiments where the attachment adapter sequence hybridizes to the 5' adapter sequence, the 5' adapter sequence is the A14 sequence and the attachment adapter sequence is the A14' sequence. In some embodiments, the attachment adapter sequence hybridizes to the 3' adapter sequence. In embodiments where the attachment adapter sequence hybridizes to the 3' adapter sequence, the 3' adapter sequence is the B15' sequence and the attachment adapter sequence is the B15 sequence. In any of these embodiments, the attachment adapter sequence may be fully complementary or partially complementary to the adapter sequence of the first or second transposon.

[0369] In some embodiments, the attachment polynucleotide comprises a primer sequence. In some embodiments, the primer sequence is the P5 primer sequence or the P7 primer sequence or a complement thereof (e.g., P5' or P7'). The P5 and P7 primers are used on the surface of commercially available flow cells sold by Illumina, Inc. for sequencing on various Illumina platforms. The primer sequences are described in U.S. Patent Application Publication No. 2011 / 0059865, which is incorporated herein by reference in its entirety. Examples of P5 and P7 primers that can be alkyne termini at the 5' end include: P5: AATGATACGGCGACCACCGAGAUCTACAC (SEQ ID NO: 7) P7: CAAGCAGAAGACGGCATACGAG * AT (SEQ ID NO: 8) and derivatives thereof. In some examples, the P7 sequence contains a modified guanine at the G * position, e.g., 8-oxo-guanine. In other examples, * is G *This indicates that the bond between and the adjacent 3'A is a phosphorothioate bond. In some examples, the P5 and / or P7 primers contain a non-natural linker. Optionally, one or both of the P5 and P7 primers may contain a poly-T tail. The poly-T tail is generally located at the 5' end of the sequence described above, for example, between the 5' base and the terminal alkyne unit, but may also be located at the 3' end. The poly-T sequence may contain any number, e.g., 2 to 20 T nucleotides. The P5 and P7 primers are given as examples, but it should be understood that any suitable primers may be used in the examples presented herein. An index sequence having a primer sequence containing the P5 and P7 primer sequences serves to add P5 and P7 to activate the library for sequencing. While the P5 and P7 primers are illustrated, it should be understood that any suitable amplification primers may be used in the examples presented herein.

[0370] As used herein, an example of a linker is a binding element that covalently bonds to the end of a nucleotide portion of an attached polynucleotide and may be used to immobilize the attached polynucleotide to a solid support. The linker may be a cleavable linker, for example, a linker that can be cleaved to remove the attached polynucleotide, and therefore a transposomal complex or tagmentation product from a solid support. As used herein, a cleavable linker is a linker that can be cleaved by chemical or physical means such as photolysis, chemical cleavage, thermal cleavage, or enzymatic cleavage. In some embodiments, cleavage may be by biochemical, chemical, enzymatic, nucleophilic, reduction-sensitive agents, or other means. A cleavable linker may be a restriction endonuclease site; at least one ribonucleotide cleavable with an RNase; a nucleotide analog cleavable in the presence of a specific chemical; a photocleavable linker unit; or (e.g.) treatment with a periodate. The molecule may include a portion selected from the group consisting of a diol bond cleavable by a chemical reducing agent; a disulfide group cleavable by a chemical reducing agent; a cleavable portion that can be subjected to photochemical cleavage; and a peptide cleavable by a peptidase enzyme, or other suitable means. Cleavage may be enzymatically mediated by incorporating a cleavable nucleotide or nucleic acid base into a cleavable linker such as uracil or 8-oxo-guanine.

[0371] In some embodiments, the linkers described herein may be covalently and directly attached to the attached nucleotide, for example, to form an -O- bond, or covalently attached via another group such as a phosphate or ester. Alternatively, the linkers described herein may be covalently attached to the phosphate group of the attached polynucleotide, for example, to the 3' hydroxyl group via the phosphate group, and therefore -O- A P(O)3- bond may be formed.

[0372] As used herein, a binding element is a portion that can be used to covalently or non-covalently bond to a binding partner. In some embodiments, the binding element is located on a transposome complex, and the binding partner is located on a solid support. In some embodiments, the binding element can bind to or non-covalently to the binding partner on the solid support, thereby allowing the transposome complex to be non-covalently attached to the solid support. In some embodiments, the binding element can bind (covalently or non-covalently) to the binding partner on the solid support. In some embodiments, the binding element is bound (covalently or non-covalently) to the binding partner on the solid support, resulting in an immobilized transposome complex.

[0373] In such embodiments, the binding element may include, for example, biotin, or be biotin, and the binding partner may include, avidin or streptavidin, or be avidin or streptavidin. In other embodiments, the binding element / binding partner combination may include FITC / anti-FITC, dioxygenin / dioxygenin antibody, or hapten / antibody, or FITC / anti-FITC, dioxygenin / dioxygenin antibody, or hapten / antibody. Further preferred binding pairs include, but are not limited to, desthiobiotin-avidin, dithiobiotin-avidin, iminobiotin-avidin, biotin-avidin, dithiobiotin-succinylated avidin, iminobiotin-succinylated avidin, biotin-streptavidin, and biotin-succinylated avidin. In some embodiments, the binding element is biotin and the binding partner is streptavidin.

[0374] In some embodiments, the binding element can bind to a binding partner via a chemical reaction or covalently bond with a binding partner on a solid support, thereby covalently attaching the transposome complex to the solid support. In some embodiments, the binding element / binding partner combination includes or is an amine / carboxylic acid (e.g., bound by a standard peptide coupling reaction under conditions known to those skilled in the art, such as EDC or NHS-mediated coupling). The reaction of the two components links the binding element and the binding partner by an amide bond. Alternatively, the binding element and binding partner may be two click chemical partners (e.g., an azide / alkyne that react to form a triazole bond).

[0375] In some embodiments, the attached polynucleotide further includes additional sequences or components such as a universal sequence, a spacer region, an anchor sequence, or an index tag sequence, or a combination thereof. The universal sequence is a region of nucleotide sequence common to two or more nucleic acid fragments. Optionally, the two or more nucleic acid fragments also have a region of sequence difference. The universal sequence, which may be present in different members of multiple nucleic acid fragments, is universal Using a single universal primer complementary to the sequence, it is possible to enable the replication or amplification of multiple different sequences.

[0376] Mutations in the transposome complex, including transposases, transposons, and attached polynucleotides, can be achieved. For example, mutations in composition, design, hybridization, structural elements, and the overall arrangement of the transposome complex can be achieved. While the disclosures and drawings provided herein offer several variations, it is understood that further mutations within the scope of this disclosure can be readily achieved.

[0377] In some embodiments, one or more library products used to generate polynucleotides are produced by bead-based tagmentation. In some embodiments, one or more library products used to generate polynucleotides are produced by solution-based tagmentation.

[0378] B.Truseq method Techniques based on fork-type adapters can be used to generate polynucleotides, for example, as exemplified in the workflow of the Truseq sample preparation kit (Illumina, Inc.). Figures 10, 12, and 13 illustrate various approaches for preparing library products containing HYB or HYB' sequences using the Truseq method.

[0379] In some embodiments, the adapter composition or kit comprises a first fork-type adapter complex and a second fork-type adapter complex, wherein the first fork-type adapter complex is a complementary attachment polynucleotide comprising a 5' portion containing a complementary attachment sequence and a 3' portion containing an adapter, and a hybridization polynucleotide comprising (a) a complement of a portion of the adapter and a 5' portion that hybridizes thereto, and (b) a complement of the hybridization sequence, wherein the complement of the hybridization sequence is the complementary attachment polynucleotide The second fork-type adapter complex comprises a hybridization polynucleotide, which includes a hybridization polynucleotide containing a complementary hybridization sequence that is not complementary to the ocide, and the second fork-type adapter complex comprises an attached polynucleotide, which includes a 5' portion containing the attached sequence and a 3' portion containing the adapter, and a hybridization polynucleotide, which includes (a) a 5' portion that hybridizes to a complementary part of the adapter and (b) a hybridization sequence, which is not complementary to the attached polynucleotide.

[0380] In some embodiments, the attachment sequence includes a P5 primer sequence, and the complementary attachment sequence includes a P7 primer sequence.

[0381] In some embodiments, the complementary attached polynucleotide contains the B15 sequence, and the hybridization polynucleotide contains the A14 sequence.

[0382] In some embodiments, the first fork-type adapter complex has the following structure:

[0383] [ka]

[0384] In some embodiments, the second fork-type adapter complex has the following structure:

[0385] [ka]

[0386] In some embodiments, the adapter complex includes a methylated nucleotide (for example, a methylated cytosine).

[0387] Methods including C. ligation In some embodiments, a library of polynucleotides is prepared by a method including a ligation step (Figures 15A-15F) such that each polynucleotide contains two inserts separated by an adapter sequence (Figures 18-19). Each starting polynucleotide has one insert. Starting polynucleotides from two or more libraries can be treated with restriction enzymes to produce polynucleotides with compatibility overhangs, and as a result, the polynucleotides can be linked together in various desired configurations to produce a new library of polynucleotides. The overhangs avoid any problems that may arise due to the complementarity of the fork-type adapter handles. In some embodiments, the new library is prepared from two starting libraries.

[0388] In some embodiments, the overhang is generated using a restriction enzyme and a restriction enzyme recognition site. In some embodiments, the enzyme is a type II, type IIS, type IIP, or type IIT restriction enzyme. In some embodiments, the enzyme is BtgZI. In some embodiments, the enzyme is BgLII. In some embodiments, the overhang is ligated together using a ligase.

[0389] In some embodiments, the polynucleotide is bound to a binding element such as biotin. In some embodiments, the digested ends of the polynucleotide are removed by applying a binding partner such as streptavidin magnetic beads.

[0390] Figures 15A to 15F illustrate exemplary ligation methods for preparing a tandem insert library. In some embodiments, the tandem insert library is sequenced using multiple reads. In some embodiments, reads 1 and 4 provide paired-end data from a first insert. In some embodiments, reads 2 and 3 provide paired-end data from a second insert.

[0391] In some embodiments, a fork-type adapter is ligated to an insert used to generate polynucleotides having different ends (Figures 16A-16B). In some embodiments, the fork-type adapter for a first library contains (1) P5 and read 1 on its first strand, and (2) a BtzI restriction enzyme recognition site on its second strand. In some embodiments, the fork-type adapter for a second library contains (1) P7 and read 2 on its first strand, and (2) a BglII restriction enzyme recognition site on its second strand. In some embodiments, primer extension is used to generate polynucleotides that are double-stranded along the entire length of each polynucleotide, i.e., polynucleotides that do not have a fork-type configuration (Figures 16A-16B).

[0392] D. Methods including chain overlap elongation (SOE) In some embodiments, the polynucleotide library is configured such that each polynucleotide contains two inserts separated by an adapter sequence (Figures 17-18). The materials are prepared by methods including chain overlap extension (SOE) (Figures 17-18). In some embodiments, the adapter sequence is a ligation sequence defined herein as a hybridization sequence that may contain one or more primer-binding sequences. Each starting polynucleotide has one insert. Starting polynucleotides from two or more libraries are ligated to the adapter. In some embodiments, these adapters are fork-type adapters or Y-type adapters. Fork-type adapters are designed so that all starting libraries have a unique adapter sequence attached to their polynucleotides. These adapter sequences provide complementary sequences for annealing in various desired configurations to generate a new library of polynucleotides (Figure 17). In some embodiments, the new library is prepared from two starting libraries. In some embodiments, the new library is prepared from three or more starting libraries.

[0393] For example, the first library contains a polynucleotide having a first adapter sequence at one end and a second adapter sequence at the other end. In these embodiments, the first or second adapter sequence has a 3' sequence complementary to the 3' terminal sequence of the third adapter sequence in the second library. Mixing of the two libraries by denaturation and re-annealing allows for hybridization of complementary ends from both libraries. In these embodiments, a polymerase elongation reaction elongates the complementary region to its full length, thus producing a double-insert polynucleotide.

[0394] Figures 17-18 show an exemplary SOE method for preparing a tandem insert library. In some embodiments, the starting library DNA is sheared to generate DNA fragments. Using polymerase, damaged DNA ends are removed, and the DNA strands are extended to produce blunt-ended double helices. Using kinase, the 5'-hydroxyls of the DNA strands are phosphorylated. Then, using polymerase, a single adenine base is added to the 3' end of each double helice. This adenine overhang ("A tail" in Figure 17) allows each end of the DNA fragment to be ligated to a single thymine overhang of the adapter. After ligating the DNA fragments with the adapters, the library is purified, and fragments of 150-200 base pairs are selected, mixed, and prepared for PCR reaction. The DNA strands are denatured at high temperature and re-annealed at low temperature. This allows the A and A' complementary adapter sequences to hybridize with each other. The strands are then extended by polymerase in the PCR reaction to form tandem insert polynucleotides.

[0395] In many embodiments, the adapter may contain various sequences in various combinations. In some embodiments, the adapter is a fork-type adapter that may contain P5, read 1, tag, and / or A sequence. In some embodiments, the adapter is a fork-type adapter that may contain P7, index, read 2, tag, and / or A' sequence.

[0396] In some embodiments, the tandem insert library is sequenced using multiple reads. In some embodiments, reads 1 and 4 provide paired-end data from a first insert. In some embodiments, reads 2 and 3 provide paired-end data from a second insert.

[0397] IV. Method for generating concatenated nucleic acid sequencing templates This application also discloses a method for generating a concatenated nucleic acid sequencing template. Multiple insert sequences can be sequenced from the concatenated nucleic acid sequencing template. In other words, the concatenated nucleic acid sequencing template can be used to generate tandem reads.

[0398] In some embodiments, the concatenated nucleic acid sequencing template is a hybridized adduct. It is produced by the formation of the hybridized adduct. As used herein, “hybridized adduct” means a hybridization sequence annealed to the complement of the hybridization sequence. In some embodiments, a fully double-stranded linked nucleic acid sequencing template is produced after the formation of the hybridized adduct.

[0399] In some embodiments, a method for generating a linked nucleic acid sequencing template includes: attaching a first read-primer binding sequence to the 3' end of a first insert sequence derived from a first target nucleic acid; attaching a hybridization sequence to the 5' end of the first insert sequence; attaching a complement of the hybridization sequence to the 3' end of a second insert sequence derived from a discontinuous region of the first target nucleic acid or a second target nucleic acid; annealing the hybridization sequence to its complement to form a hybridized adduct; and synthesizing a fully double-stranded linked nucleic acid sequencing template from the hybridized adduct, wherein the region between the first and second insert sequences is a second read-primer binding sequence comprising a hybridization sequence and a second read-primer binding sequence orthogonal to the first read-primer binding sequence, thereby generating a linked nucleic acid sequencing template.

[0400] In some embodiments, attaching a first read primer binding sequence and attaching a hybridization sequence involves contacting one or more target nucleic acids with a transposomal complex under conditions suitable for tagmentation.

[0401] In some embodiments, attaching a hybridization sequence complement to a discontinuous region of a first target nucleic acid or to the 3' end of a second insert sequence derived from a second target nucleic acid involves contacting one or more target nucleic acids with a transposome complex under conditions suitable for tagmentation.

[0402] In some embodiments, attaching a first read-primer binding sequence to the 3' end of a first insert sequence and attaching a hybridization sequence to the 5' end of a first insert sequence involves contacting one or more target nucleic acids with a first fork-type adapter complex under conditions suitable for ligation of the ends of the adapter complex fragments to form fragments ligated with the first adapter complex at both ends and fragments ligated with the second adapter complex at both ends, and denaturing the ligated fragments.

[0403] In some embodiments, attaching the complementary of the hybridization sequence to the 3' end of the second insert sequence involves contacting one or more target nucleic acids with the second fork-type adapter complex under conditions suitable for ligation of the ends of the adapter complex fragments to form fragments ligated with the first adapter complex at both ends and fragments ligated with the second adapter complex at both ends, and denaturing the ligated fragments.

[0404] In some embodiments, a method for generating a linked nucleic acid sequencing template involves contacting a first sample containing a first target nucleic acid with a first transposome complex and a second transposome complex, where each transposome complex is Transposase and, A first transposon comprising a 3' portion containing the transposon terminal sequence and a 5' portion containing the adapter sequence, It includes a 5' portion containing the complement of the transposon terminal sequence, and a second transposon that hybridizes to it, The adapter sequence in the first transposome complex is complementary to the first adapter sequence, and the adapter sequence in the second transposome complex is complementary to the second adapter sequence. can be, This is done under conditions sufficient to produce a first tagged product containing an insert sequence derived from the first target nucleic acid, in which the first target nucleic acid is fragmented and one end is tagged with a transposon of the first transposome complex and the other end is tagged with a transposon of the second transposome complex. The process involves selectively adding a complementary attachment sequence to the 3' end of the first tagged product via polymerase chain reaction, and adding a complementary hybridization sequence to the 5' end of the first tagged product to form the first modified tagged product. The method involves contacting a second sample containing a second target nucleic acid with a transposome complex, under conditions sufficient to fragment the second target nucleic acid and generate a second tagged product containing an insert sequence derived from the second target nucleic acid, where one end is tagged with a transposon of the first transposome complex and the other end is tagged with a transposon of the second transposome complex. The process involves selectively adding an attachment sequence to the 3' end of the second tagged product and a hybridization sequence to the 5' end of the second tagged product via polymerase chain reaction, thereby forming a second modified tagged product. The hybridization sequence of the first modified tagged product is annealed to the complementary of the hybridization sequence in the second modified tagged product to form a hybridized adduct. The synthesis of a fully double-stranded, linked nucleic acid sequencing template from a hybridized adduct, wherein the linked nucleic acid sequencing template is (a) A first read-primer binding sequence located at the 3' end of an insert sequence derived from a second target nucleic acid, comprising a first adapter sequence and a complement of the transposon terminal sequence, (b) A second read-primer binding sequence located between two insert sequences, comprising a transposon terminal sequence and a hybridization sequence, The first read-primer binding sequence is orthogonal to the second read-primer binding sequence, and includes the following:

[0405] In some embodiments, a method for generating a concatenated nucleic acid sequencing template is: The first method involves contacting a first sample containing a first target nucleic acid with a first transposome complex, wherein the first transposome complex is Transposase and, A first transposon comprising a 3' portion containing the transposon terminal sequence and a 5' portion containing the attachment sequence and the complement of the first adapter sequence, It includes a 5' portion containing the complement of the transposon terminal sequence, and a second transposon that hybridizes to it, This is done under conditions sufficient to fragment the first target nucleic acid and generate a first tagged product containing an insert sequence derived from the first target nucleic acid, with each end tagged by a transposon of the first transposomal complex. The first modified tagged product is formed by selectively adding a hybridization sequence complement to the 5' end of the first tagged product via polymerase chain reaction, The method involves contacting a second sample containing a second target nucleic acid with a second transposome complex, wherein the second transposome complex is Transposase and, A first transposon comprising a 3' portion containing a transposon terminal sequence and a 5' portion containing a second adapter sequence and a complementary attachment sequence, It contains the 5' portion which includes the complement of the transposon terminal sequence, and hybridizes to it. Including two transposons, This is done under conditions sufficient to fragment the second target nucleic acid and generate a second tagged product containing an insert sequence derived from the second target nucleic acid, with each end tagged by a transposon of the second transposomal complex. The process involves selectively adding a hybridization sequence complement to the 5' end of the second tagged product via polymerase chain reaction to form a second modified tagged product, The hybridization sequence of the first modified tagged product is annealed to the complementary of the hybridization sequence in the second modified tagged product to form a hybridized adduct. The synthesis of a fully double-stranded, linked nucleic acid sequencing template from a hybridized adduct, wherein the linked nucleic acid sequencing template is (a) A first read-primer binding sequence located at the 3' end of an insert sequence derived from a second target nucleic acid, comprising a first adapter sequence and a complement of the transposon terminal sequence, (b) A second read-primer binding sequence located between two insert sequences, comprising a transposon terminal sequence and a hybridization sequence, The first read-primer binding sequence is orthogonal to the second read-primer binding sequence, and includes the following:

[0406] In some embodiments, the transposome complex is immobilized on a solid support.

[0407] V. Method for preparing a sequencing template using a fork-type adapter In some embodiments, a fork-type adapter can be used to prepare an array determination template containing two or more inserts.

[0408] In some embodiments, the adapter may be a fork-shaped adapter, also known as a Y-shaped adapter. Techniques based on the fork-shaped adapter can be used to generate polynucleotides, for example, as exemplified in the workflow of the Truseq® sample preparation kit (Illumina, Inc.). Alternatively, the fork-shaped adapter can be assembled using reagents from the workflow of the TruSight® Oncology kit (Illumina, Inc.). In some embodiments, the fork-shaped adapter contains a HYB or HYB' sequence.

[0409] As used herein, “fork adapter” refers to an adapter comprising two strands of nucleic acid, each containing a region complementary to the other strand and a region not complementary to the other strand. In some embodiments, the nucleic acids of the two strands in the fork adapter are annealed together prior to ligation, and the annealing is based on the complementary region. In some embodiments, the complementary region contains 12 nucleotides each. In some embodiments, the fork adapter is ligated to both strands at the ends of a double-stranded DNA fragment. In some embodiments, the fork adapter is ligated to one end of a double-stranded DNA fragment. In some embodiments, the fork adapter is ligated to both ends of a double-stranded DNA fragment. In some embodiments, the fork adapter on opposing ends of a fragment is different (shown in Figure 27A). In some embodiments, one strand of the fork adapter is phosphorylated at its 5' to facilitate ligation to the fragment. In some embodiments, one strand of the fork adapter has a phosphorothioate bond immediately before the 3'T. In some embodiments, 3'T is an overhang (i.e., the other end of the chain of the fork-type adapter) (Not paired with a creotide). In some embodiments, the 3'T overhang can base pair with the A tail present on the library fragment. In some embodiments, the phosphorothioate bond blocks exonuclease digestion of the 3'T overhang.

[0410] In some embodiments, each fork-type adapter includes a first oligonucleotide and a second oligonucleotide that partially hybridize with each other to form a double-stranded section and a single-stranded section.

[0411] Figure 25 shows a pair of fork-type adapters (i.e., a first adapter and a second adapter) that may be used to prepare a sequencing template. In some embodiments, the first strand of each fork-type adapter contains an adapter such as a sequencing primer sequence. In some embodiments, the second strand of each fork-type adapter contains either a hybridization sequence (X) or a complement (X') of the hybridization sequence.

[0412] Blocking oligonucleotides can be used to block the hybridization sequence (X) and its complement (X') from binding to each other at undesirable times. In some embodiments, the blocking oligonucleotides include one or more modifications so that they are not targeted for tagmentation. In other words, blocking oligonucleotides can be designed to be transposase-resistant and therefore to avoid cleavage of the double-stranded nucleic acid formed by hybridization of the blocking oligonucleotide to the hybridization sequence or its complement. In some embodiments, the blocking oligonucleotides include a phosphorothioate backbone.

[0413] In some embodiments, the blocking oligonucleotide comprises a complement to all or part of the sequence whose hybridization is to be blocked. Thus, in some embodiments, the blocking oligonucleotide may be all or part of the X or X' sequence. As used herein, “blocking oligonucleotide” means an oligonucleotide that can be used to inhibit the binding of two sequences to each other until the blocking oligonucleotide, which is bound to at least one of the two sequences, is removed. In some embodiments, the blocking oligonucleotide comprises a sequence that is completely or partially complementary to all or part of either the hybridization sequence (X or HYB) or its complement (X' or HYB'). For example, a blocking oligonucleotide (X'B') for blocking the HYB sequence (X in Figure 25) may comprise all or part of the HYB' sequence, and a blocking oligonucleotide (XB) for blocking the HYB' sequence (X' in Figure 25) may comprise all or part of the HYB sequence.

[0414] In the case of the fork-type adapter shown in Figure 26, one or more blocking oligonucleotides may act to block the binding of the X sequence in one fork-type adapter to the X' sequence in the other fork-type adapter.

[0415] In some embodiments, the blocking oligonucleotide (XB) binds to the X' sequence. In some embodiments, the blocking oligonucleotide (X'B') binds to the X sequence. In some embodiments, the blocking oligonucleotide binds to both the X and X' sequences. The blocking oligonucleotide may be completely or partially complementary to either the X or X' sequence. In some embodiments, the blocking oligonucleotide binds to the complete X or X' sequence. In some embodiments, the blocking oligonucleotide binds to a portion of the X or X' sequence.

[0416] One or both fork adapters may also include an affinity moiety on the 5' end of the first chain of the fork adapter. In some embodiments, such as those shown in Figure 26, both the first chain of the first fork adapter and the first chain of the second fork adapter include an affinity moiety at the 5' end of the chain. In some embodiments, the affinity moiety is biotin, desthiobiotin, or dualbiotin. In some embodiments, the affinity moiety is biotin (i.e., the first chain of one or both fork adapters is biotinized). In some embodiments, the affinity moiety binds to a binding moiety on the surface of a solid support. In some embodiments, the binding moiety is avidin or streptavidin that binds to avidin or streptavidin on the surface of a solid support. The range of affinity moieties that can bind to the binding moiety is known to those skilled in the art, and the user can select any pair of affinity / binding moieties of their choice.

[0417] In some embodiments, the binding portion serves to immobilize the tagged fragment (prepared by ligation to the fragment of the fork-type adapter) onto a solid support. In some embodiments, a single-stranded fragment ligated to at least one first strand of the fork-type adapter is immobilized on the solid support. In some embodiments, the immobilized fragment can be washed, and blocking oligonucleotides can be removed without the fragment being released from the surface of the solid support.

[0418] In some embodiments, the first chain of the fork-type adapter includes a 5' affinity element that can bind to an affinity binding partner on a solid support or beads. Such an affinity element may be biotin, as indicated by "Bio" in the first and second adapters shown in Figure 25.

[0419] In some embodiments, affinity elements are connected via linkers attached to the first chain. In some embodiments, these linkers are cleavable linkers.

[0420] In some embodiments, the affinity portion is linked to a first chain of a fork-type adapter by a linker. In some embodiments, the linker is a cleavable linker. In some embodiments, the user can detach the sequencing template prepared from the fixed fragments from the solid support at a desired time by detaching the cleavable linker between the affinity portion and the first chain of the fork-type adapter. In some embodiments, the amplicon of the sequencing template may be prepared on the surface of the solid support, in which case the amplicon can be sequenced without requiring the detachment of the sequencing template from the surface.

[0421] In some embodiments, hybridization sequences (HYBs) and their complements (HYB's) can hybridize with each other. However, in some cases, this can potentially lead to dimerization between different fork adapters based on the binding of a HYB in one fork adapter to a HYB' in another fork adapter. Such adapter dimerization can reduce the ability of the fork adapters to ligate to the ends of nucleic acid fragments.

[0422] In some embodiments, a blocking oligonucleotide is used to block the binding of HYB to HYB' between different fork-type adapters until the user desires this binding to occur. In some embodiments, the hybridization sequence or its complement is bound to a blocking oligonucleotide that is fully or partially complementary to the hybridization sequence or its complement.

[0423] Figures 26A–26C illustrate various different embodiments of fork-type adapters. A blocking oligonucleotide can bind to the second chain of both the first and second fork-type adapters (Figure 26A). Alternatively, the blocking oligonucleotide can bind to only the second chain of the first fork-type adapter (Figure 26B), or only the second chain of the second fork-type adapter. As long as either the hybridization sequence (X) or its complement (X') is bound by the blocking oligonucleotide, the blocking oligonucleotide blocks the annealing of the fork-type adapters to each other via the association of X to X'. A similar method can be carried out using transposome complexes in solution, as shown in Figure 26D.

[0424] In some embodiments, a fork-type adapter comprising two polynucleotide chains comprises (a) a first chain comprising a sequencing primer sequence, and (b) a second chain comprising a 3' hybridization sequence or its complement, wherein the 3' end of the first chain is fully or partially complementary to the 5' end of the second chain. In other words, the two chains of the fork-type adapter can hybridize together in a particular region, while the two chains are separated in another region. The sequences of the first and second chains may be different or entirely or partially complementary in the region where the two chains are separated, but the first and second chains may be the same and fully or partially complementary in the region where the two chains hybridize together.

[0425] As is well known in the art, further target sequences such as UMI and sample index may be included in the fork adapter. In other words, the fork adapter is not limited to the type of sequence shown in Figure 25, but may include one or more further types of sequences such as UMI or sample index.

[0426] In some embodiments, the first and / or second strands further include at least one of an adapter, a barcode sequence, a unique molecular identifier (UMI) sequence, a sample index sequence, a capture sequence, or a cleavage sequence.

[0427] In some embodiments, the sequencing primer sequence included in the first strand of the fork-type adapter includes a B15 sequence or an A14 sequence, or their complements. In some embodiments, the first strand of the fork-type adapter further includes a P7 or P5 primer sequence, or their complements. Such embodiments are shown in Figure 25, in which the first strand of the first adapter includes a P5 sequence and a first read sequencing adapter sequence (P5.R1), and the first strand of the second adapter includes a P7 sequence and a second read sequencing adapter sequence (P7.R2).

[0428] In some embodiments, the fork adapter is included in a mixture with another non-identical fork adapter. In some embodiments, the mixture includes a different first fork adapter and a second fork adapter.

[0429] In some embodiments, the composition or kit includes two fork-type adapters: (a) a first fork-type adapter comprising a first strand containing a first read sequencing primer sequence and a second strand containing a complement of a hybridization sequence; and (b) a second fork-type adapter comprising a first strand containing a second read sequencing primer sequence and a second strand containing a hybridization sequence. In some embodiments, one or both of the fork-type adapters included in the kit or composition include a blocking oligonucleotide.

[0430] A mixture of fork-type adapters can be ligated to double-stranded nucleic acid fragments. The fragments can be prepared from DNA (e.g., genomic DNA or cDNA prepared from RNA) using physical means, such as acoustic, spray, centrifugal force, needle, or hydrodynamic techniques, which are well known in the art. Enzymatic means for preparing the fragments, such as DNase treatment, are also well known.

[0431] When a mixture containing a first fork adapter and a second fork adapter is combined with a double-stranded nucleic acid fragment under conditions for ligation, the predicted ratios are that 50% of the fragment is tagged with the first fork adapter at one end and the second fork adapter at the second end (Figure 27A), 25% of the fragment is tagged with the first fork adapter at both ends (Figure 27B), and 25% of the fragment is tagged with the second fork adapter at both ends (Figure 27C). In some embodiments, the ligation products shown in Figures 27A-27C can be produced by a ligation reaction prepared in solution. In other words, the tagged fragments shown in Figures 27A-27C can be prepared in solution.

[0432] In some embodiments, tagged fragments prepared in solution by ligation of a fork-type adapter can then be immobilized on the surface of a solid support.

[0433] In some embodiments, a method for generating one or more linked nucleic acid sequencing templates includes contacting a sample containing double-stranded nucleic acid fragments, each containing an insert prepared from a target nucleic acid, with a composition or kit containing two fork-shaped adapters, one or both of which contain a blocking oligonucleotide. In some embodiments, after contacting the sample with the two fork-shaped adapters, the method includes ligating the fork-shaped adapters to the double-stranded fragments to prepare tagged double-stranded fragments and immobilizing the tagged double-stranded fragments on a solid support.

[0434] In some embodiments, the double-stranded fragment is applied to a solid support after ligation with a fork-type adapter. In some embodiments, both 5' ends of the tagged double-stranded fragment contain affinity moieties that can bind to a binding site on the surface of the solid support (based on ligation of the first strand of the fork-type adapter containing the affinity moieties). In some embodiments, binding of the affinity moieties to the binding site fixes the fragment on the solid support and, as a result, prevents it from being released from the support by temperature changes that could allow the release of a blocking oligonucleotide bound to the hybridization sequence or its complement.

[0435] After immobilizing a double-stranded fragment onto the surface of a solid support, the method comprises (1) denaturing the immobilized tagged double-stranded fragment to produce an immobilized single-stranded fragment, and (2) deblocking a blocking oligonucleotide to deblock the hybridization sequence and its complement. In some embodiments, denaturation is carried out by increasing the temperature, changing the pH, and / or adding one or more chaotropic agents. In some embodiments, for example, a single temperature change may mediate the denaturation of the two strands of the double-stranded fragment and the release of the blocking oligonucleotide. In some embodiments, the temperature increase associated with denaturation is from 45°C to 55°C to 85°C to 95°C, and optionally, the temperature increase is from 50°C to 90°C. In some embodiments, the one or more chaotropic agents include formamide and / or NaOH.

[0436] In some embodiments, the first single-stranded fragment includes an insert, and the second single-stranded fragment includes an insert that is a complement to the insert contained in the first fragment. In some embodiments, the first single-stranded fragment includes an insert, and the second fragment includes an insert that is not a complement to the insert contained in the first fragment. In some embodiments, hybridization involves a first fork-shaped adapter ligated to one end of each fragment and This occurs between single-stranded fragments prepared from double-stranded fragments, each containing a second fork-shaped adapter ligated to the other end of each fragment. In some embodiments, the two fixed single-stranded fragments do not hybridize to each other to form a bridge, without the binding of the hybridization sequence in the first fragment to the complement of the hybridization sequence in the second fragment. In some embodiments, the hybridization of two fixed single-stranded fragments to form a bridge does not occur between single-stranded fragments prepared from double-stranded fragments, each containing the same fork-shaped adapter ligated to both ends of each fragment.

[0437] In some embodiments, the surface of the solid support is washed after denaturation, and blocking oligonucleotides are removed by washing, but the single-stranded fragments remain immobilized due to interactions between 5' affinity moieties on the fragments having binding moieties on the surface of the solid support. In some embodiments, immobilization of double-stranded or single-stranded fragments is by binding of affinity moieties derived from a first and / or second fork-type adapter to one or more binding moieties on the surface of the solid support. In some embodiments, the affinity moieties are biotin, desthiobiotin, or dualbiotin, and the binding moieties are avidin or streptavidin.

[0438] Since the single-stranded fragments are prepared from double-stranded fragments already fixed on a single surface on a solid support, the complementary single-stranded fragments derived from the double-stranded fragments are likely to be very close in proximity (as shown in Figure 28A, where the left and right surfaces of the solid support are different diagrams of the same surface). The denaturation of the blocking oligonucleotide means that the hybridization sequence and its complement (X and X' in Figure 28A) are available to bind to each other at this stage.

[0439] Next, the method includes hybridizing two fixed single-stranded fragments to form a bridge by linking the hybridization sequence in the first fragment to the complement of the hybridization sequence in the second fragment, and extending from the 3' ends of both single-stranded fragments to generate a double-stranded linked nucleic acid sequencing template in which each strand of the template contains inserts (or their complements) derived from both fixed single-stranded fragments (as shown in Figure 29).

[0440] In some embodiments, a single-stranded fragment prepared from a double-stranded fragment ligated at the first end with the first strand of a first fork adapter (e.g., shown in Figure 25) and the second strand of a second fork adapter can be bound to another single-stranded fragment prepared from a double-stranded fragment ligated at the first end with the first strand of a first fork adapter and the second strand of a second fork adapter by association of the hybridization sequence (X) in the first fragment to the complementary hybridization sequence (X') in the second fragment (Figure 28A).

[0441] In some embodiments, one or more further steps of denaturation, hybridization, and extension are performed. In this way, the method can proceed with the preparation of the sequencing template until there are no other suitable single-stranded fragments to form a crosslink (and linked sequencing template) via HYB / HYB' bonds.

[0442] In some embodiments, both single-stranded fragments prepared from the double-stranded fragments are fixed to the surface of the same solid support. In some embodiments, the method is carried out using a single surface on the solid support, and as a result, all fragments are fixed to the same solid support. The left and right surfaces shown in Figures 28A–28C (shown with the first and second fragments attached) represent two different figures of the same surface on the solid support.

[0443] In some embodiments, the release of blocking oligonucleotides is accompanied by their complementary This generates “free” hybridization sequences that can be joined in a column. In some embodiments, a hybridization sequence contained in one single-stranded fragment may bind to the complement of a hybridization sequence in another single-stranded fragment. Such binding can create a “crosslink” as shown in Figure 28A.

[0444] After stretching, the linked sequencing template may contain two inserts that are copies of each other, as shown in Figure 29.

[0445] Single-stranded fragments with identical ligated adapters cannot hybridize with each other. For example, two fragments tagged with X' cannot pair with each other in the hybridization sequence (Figure 28B), and two fragments tagged with X cannot pair with each other in the hybridization sequence (Figure 28C). Therefore, a sequencing template containing two inserts cannot be prepared from fragments containing the same adapter (indicated by 0% shown in Figures 28B and 28C). Although two insert sequences can hybridize with each other (sequences of strands A and A' in Figures 28A-28C), direct hybridization between these sequences makes post-hybridization extension impossible because such pairing between strands A and A' is followed by a non-complementary 3' sequence (X / X').

[0446] In this way, 100% of the sequencing template, including the two copies of the insert, is prepared from fragments containing different adapters (Figure 28A). This aspect is important because the first fork adapter may contain a different sequence than the second fork adapter. For example, as shown in Figure 28A, the first fork adapter may contain the first read sequencing adapter sequence (P5.R1), and the second fork adapter may contain the second read sequencing adapter sequence (P7.R2).

[0447] Therefore, as shown in Figure 29, a full-length linked sequencing template can be prepared after stretching, including two copies of the same insert sequence and appropriate adapters that may be required for a desired sequencing platform. In other words, a person skilled in the art can design a fork-type adapter such that the resulting sequencing template includes the desired adapter sequence for a preferred sequencing platform.

[0448] Since the double-stranded fragments are first immobilized on a solid support and then denatured, two single-stranded fragments denatured from the same double-stranded fragment are likely to be immobilized in close proximity to each other on the surface. This sequence of steps means that two single-stranded fragments derived from the same double-stranded fragment (as shown in Figure 28A, one fragment containing the A-sequence and the other containing the A'-sequence) are likely to be able to interact with each other. This aspect increases the likelihood that the sequencing template prepared by the method of the present invention contains two copies of the same sequence derived from the target nucleic acid (one derived from the A-sequence and one from the complement of the A'-sequence prepared by extension). As described herein, this sequencing template having two copies of the same insert sequence (resulting from the complementary strand of the target nucleic acid) allows for error correction or identification of base pair mismatches between the strand and antisense strand of the target nucleic acid. Such base pair mismatches may be rare and otherwise difficult to resolve with standard sequencing.

[0449] Alternatively, single-stranded fragments containing unrelated insert sequences and complementary adapters may also hybridize to form bridges, thereby generating a concatenated sequencing template. A concatenated sequencing template with two different inserts may help increase sequencing depth by enabling additional sequence reads compared to sequencing using a standard sequencing template with a single insert.

[0450] A. Methods for partitioning data to evaluate proximity data Any method described herein may be used in conjunction with compartmentalization. In some embodiments, compartmentalization allows for the generation of proximity data, such as whether different inserts are contained in the same target nucleic acid. If the same target nucleic acid is a chromosome, compartmentalization may be used in the haplotype phasing methods described herein.

[0451] In some embodiments, compartmentalization is used in conjunction with the present invention's method of evaluating proximity data using a fork-type adapter or transposome. In some embodiments, compartmentalization may be used diluted to limit the number of available target nucleic acids. In some embodiments, each compartment generally contains one target nucleic acid or does not contain a target nucleic acid after dilution (as shown in Figure 31). Thus, fragments prepared in a given compartment are generally fragments prepared from the same target nucleic acid. In this way, it can be inferred that inserts contained in the same concatenated sequencing template prepared by these methods originated from the same target nucleic acid.

[0452] In some embodiments, the compartments are wells, tubes, or droplets. For example, Figure 31 shows a method using wells, and Figure 32 shows a method using droplets. A wide range of different wells, tubes, and droplets are known to those skilled in the art, and any type can be used in the method of the present invention.

[0453] A "droplet" refers to a volume of liquid on a droplet actuator. Typically, a droplet is at least partially bound by a filling fluid. For example, a droplet may be completely surrounded by the filling fluid, or surrounded by the filling fluid and one or more surfaces of the droplet actuator. In another example, a droplet may be surrounded by the filling fluid, one or more surfaces of the droplet actuator, and / or the atmosphere. In yet another example, a droplet may be surrounded by the filling fluid and the atmosphere. A droplet may be, for example, aqueous or non-aqueous, or a mixture or emulsion containing aqueous and non-aqueous components. Droplets can take on a wide variety of shapes, and non-limiting examples include substantially disc-shaped, slug-shaped, truncated sphere-shaped, elliptical, spherical, partially compressed sphere-shaped, hemispherical, oval, cylindrical, combinations of such shapes, and various shapes formed during droplet operation, e.g., merging or splitting, or various shapes formed as a result of contact between such shapes and one or more surfaces of the droplet actuator. For examples of droplet fluids that can be used for droplet manipulation using the approach of this disclosure, see Eckhardt et al., International Publication No. 2007 / 120241, “Droplet-Based Biochemistry,” published October 25, 2007, the entire disclosure of which is incorporated herein by reference. U.S. Patent No. 10,975,371, which teaches a wide variety of applications of droplets and droplet actuators, the entirety of which is incorporated herein by reference.

[0454] In some embodiments, the fragments can be prepared within a compartment using two pools of fork-type adapters, one pool containing a fork-type adapter containing a hybridization sequence (i.e., the second adapter in Figure 25), and the other pool containing a fork-type adapter containing a complement of the hybridization sequence (i.e., the first adapter in Figure 25).

[0455] In some embodiments, a method for generating one or more linked nucleic acid sequencing templates comprises partitioning a sample containing a target double-stranded nucleic acid into a plurality of different compartments, and preparing fragments each containing an insert derived from the double-stranded nucleic acid within the plurality of different compartments. The method then involves contacting the plurality of different compartments with a composition or kit comprising two fork-type adapters, one or both of which contain blocking oligonucleotides, and ligating the fork-type adapters to the double-stranded fragments, plurality This may include preparing double-stranded fragments tagged within different compartments.

[0456] In some embodiments, the method may then include denaturing (1) a tagged double-stranded fragment fixed to generate a single-stranded fragment, and (2) a blocking oligonucleotide to unblock hybridization sequences and complements of hybridization sequences in a plurality of different compartments, and hybridizing two single-stranded fragments in the same compartment to form a crosslink by linking the hybridization sequence in the first fragment to the complement of the hybridization sequence in the second fragment. In some embodiments, the method may include extending from the 3' end of each single-stranded fragment to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both single-stranded fragments in the same compartment.

[0457] In some embodiments, the target double-stranded nucleic acid comprises a double-stranded DNA fragment, and smaller fragments of the double-stranded DNA fragment are prepared by preparing the fragment. In other words, the target double-stranded nucleic acid may be fragmented into a relatively large fragment, which is then fragmented into smaller fragments within a compartment. This is shown in Figures 31 and 32, where the f1 fragment is fragmented into smaller fragments 1.1, 1.2, and 1.3.

[0458] In this method, the single-stranded fragments are not immobilized, so it is likely that a concatenated sequencing template containing two different insert sequences will be prepared. In some embodiments, the first single-stranded fragment contains an insert, and the second fragment contains an insert that is not a complement of the insert contained in the first fragment.

[0459] In some embodiments, hybridization occurs between single-stranded fragments prepared from double-stranded fragments, each fragment comprising a first fork-shaped adapter ligated to one end of the fragment and a second fork-shaped adapter ligated to the other end of the fragment.

[0460] In some embodiments, single-stranded fragments do not hybridize to each other to form a bridge without the binding of the hybridization sequence in the first fragment to the complementary hybridization sequence in the second fragment. In some embodiments, hybridization of two single-stranded fragments to form a bridge does not occur between single-stranded fragments prepared from double-stranded fragments, each containing the same fork-shaped adapters ligated at both ends of the fragment.

[0461] B. Haplotype fading As used herein, "haplotype phasing" refers to the identification of alleles coexisting on the same chromosome. Sequence data generally consists of unphasized genotypes, and such data cannot distinguish whether a particular allele belongs to two parental chromosomes or a specific haplotype.

[0462] Methods for compartmentalization (e.g., for use in preparation for whole-genome haplotyping) are well known in the art, for example, Amini et al., Nat Genet. 46(12):1343-9 (2014), Kaper F, et al. Proc. Natl. Acad. Sci. US A. 110(14):5552-5557 (2013), Kitzman JO, et al. Nat. Biotechnol. 29(1):59-63 (2011), Peters BA, et al. Nature. 487(7406):190-195 (2012), Fan HC, et al. Nat. Biotechnol. 29(1):51-57 (2011), Levy S, et al. PLoS Biol. 5(10):e254 (2007), Duitama J, et al. al.Nucleic Acids Res.40(5):2041-2053(2012), Suk This is taught in EK, et al. Genome Res. 21(10):1672-1685 (2011), and the full disclosures of each are incorporated herein by reference.

[0463] In some embodiments, compartmentalization separates different haplotypes into different compartments, and the method is used for haplotype phasing. In some embodiments, a target nucleic acid, such as double-stranded DNA, is divided equally into multiple compartments by limiting dilution, so that each compartment contains a limited number of DNA molecules, thereby making it likely that any location in the genome is represented by haploid DNA in the compartment.

[0464] In some embodiments, limiting dilution reduces the likelihood that both haplotypes (such as Chr1-Hap1 and Chr2-Hap2 in Figure 33) are in the same compartment, but the method does not require that only a single chromosome be contained within the compartment. In other words, dilution may reduce the likelihood (e.g., less than 5% or less than 1%) that two haploid copies of the same chromosome are in the same compartment, but the compartment may often contain two or more chromosomes (in which case the two or more chromosomes are generally not haploid copies of the same chromosome).

[0465] Such a method is shown in Figure 33, in which chromosomes are subjected to limiting dilution within compartments, followed by the preparation of single-stranded fragments, and then hybridization and extension to prepare linked sequencing templates within individual compartments.

[0466] In the example shown in Figure 33, Chr1-Hap1 is contained within the compartment containing Chr2-Hap1, while Chr1-Hap2 is contained within the compartment containing Chr2-Hap2. Since the concatenated sequencing templates are prepared using the compartments, these templates can only contain chromosomal inserts that were in the same compartment (indicated as boxes with checked arrows). Other combinations (indicated as boxes with "X" arrows) cannot be formed because these haplotypes were not contained within the same compartment in this example.

[0467] When this method is performed using samples from organisms with known genomes, the presence of inserts from different chromosomes in the same concatenated sequencing template (because these different chromosomes were in the same compartment during the method) can be elucidated from the sequencing data. Analysis to determine the chromosomes that were in the same compartment can determine information about the alleles contained in the haploid copy. In some embodiments, the method does not require barcodes. Instead, the use of concatenated sequencing templates prepared in compartments according to the present invention allows for analysis of which insert sequences were contained in the haploid copy without the need for barcodes.

[0468] VI. Method for preparing a sequencing template containing multiple inserts using transposomes in solution In some embodiments, tagmentation is performed in solution to prepare tagged double-stranded fragments. These tagged double-stranded fragments can be used to prepare a sequencing template containing multiple inserts, similar to the method described above for ligation of fork-type adapters. In some embodiments, the tagged double-stranded fragments are prepared in solution using two pools of transposomes, and then the tagged double-stranded fragments are immobilized on a solid support. In some embodiments, immobilization is performed by binding affinity moies incorporated into the tagged fragments during tagmentation to binding moies on the solid support. Figure 26D shows an embodiment of preparing tagged double-stranded fragments in solution using tagmentation, and these tagged double-stranded fragments can be used to prepare a linked sequencing template, similar to the method described above for using fork-type adapters.

[0469] In some embodiments, a method for generating one or more linked nucleic acid sequencing templates includes (a) contacting a sample containing a double-stranded target nucleic acid with two pools of transposome complexes in solution, wherein a first pool of transposome complexes comprises a transposase, a first transposon comprising a 3' transposon terminal sequence and a first read sequencing adapter sequence, and a second transposon comprising a 3' complement of a 5' sequence and a hybridization sequence that is fully or partially complementary to the 3' transposon terminal sequence, and a second pool of transposome complexes comprises a transposase, a first transposon comprising a 3' transposon terminal sequence and a second read sequencing adapter sequence, and a second transposon comprising a 5' sequence and a 3' hybridization sequence that is fully or partially complementary to the 3' transposon terminal sequence.

[0470] In some embodiments, one or both of the second transposons include a blocking oligonucleotide. Such blocking oligonucleotides have been described above in the method using a fork-type adapter, and the blocking oligonucleotide may be used to inhibit the binding of a hybridization sequence contained in one pool of the transposomal complex to its complement in the other pool of the transposomal complex.

[0471] In some embodiments, the method includes tagging a double-stranded nucleic acid to generate a tagged double-stranded fragment, releasing a transposomal complex from the double-stranded fragment, and extending and ligating the double-stranded fragment.

[0472] In some embodiments, the tagged double-stranded fragments are immobilized on a solid support. In some embodiments, this immobilization is achieved by binding the 5' affinity portion of the tag to the binding portion on the solid support.

[0473] In some embodiments, the method then comprises denaturing (1) an immobilized tagged double-stranded fragment to generate an immobilized single-stranded fragment, and (2) a blocking oligonucleotide to unblock a hybridization sequence and a complement of the hybridization sequence. In some embodiments, after denaturation, the method comprises hybridizing the two immobilized single-stranded fragments to form a crosslink by linking the hybridization sequence in the first fragment to the complement of the hybridization sequence in the second fragment, and extending from the 3' end of each single-stranded fragment to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both immobilized single-stranded fragments.

[0474] In some embodiments, the double-stranded linked nucleic acid sequencing template includes an insert sequence and a copy of the insert sequence. In some embodiments, the double-stranded linked nucleic acid sequencing template includes two insert sequences that are different from each other.

[0475] Hybridization of a hybridization sequence in one single-strand template to a complementary hybridization sequence in another single-strand template, and extension for preparing a linked sequencing template, can be carried out as described above in the fork adapter method. Essentially, once tagged double-stranded fragments are prepared in solution (either by ligation of a fork adapter or tagmentation in solution), a linked sequencing template can be carried out by a similar process after subsequent immobilization and crosslinking preparation steps.

[0476] In some embodiments, hybridization involves a tag derived from a second transposon of the first transposome complex at one end of each fragment, and a second transposon at the other end of each fragment. This occurs between single-stranded fragments prepared from double-stranded fragments, including a tag derived from the second transposon of the transposome.

[0477] In some embodiments, hybridization of two fixed single-stranded fragments to form crosslinks does not occur between single-stranded fragments prepared from double-stranded fragments containing tags derived from the same transposome complex at both ends of each fragment.

[0478] VII. Method for preparing a sequencing template containing multiple inserts using a solid support having immobilized transpososomes In some embodiments, sequencing templates containing multiple inserts are prepared using transposomes immobilized on a solid support. In some embodiments, the solid support is a bead, a slide, a container wall, a flow cell, or a nanowell contained within a flow cell.

[0479] When used herein, a “transposome complex” or “transposome” comprises at least one transposase (or other enzyme described herein) and a transposon recognition sequence. In some such systems, the transposase binds to the transposon recognition sequence to form a functional complex capable of catalyzing the transposition reaction. In some aspects, the transposon recognition sequence is a double-stranded transposon terminal sequence. The transposase binds to a transposase recognition site in the target nucleic acid and inserts the transposon recognition sequence into the target nucleic acid. In some such insertion events, one strand of the transposon recognition sequence (or terminal sequence) is transcribed into the target nucleic acid, resulting in a cleavage event. Exemplary transposition procedures and systems can be readily adapted for use with transposases.

[0480] "Transposase" means an enzyme that forms a functional complex having a transposon end-containing composition (e.g., a transposon, a transposon end, or a transposon end composition) and can catalyze the insertion or rearrangement of the transposon end-containing composition into a double-stranded target nucleic acid. The transposases presented herein may also include integrases from retrotransposons and retroviruses.

[0481] Transposon-based technologies can be used to fragment DNA, where target nucleic acids, such as genomic DNA, are treated with a transposomal complex that simultaneously fragments and tags the target ("tagmentation"), thereby generating a population of fragmented nucleic acid molecules tagged with adapter sequences specific to the ends of the fragments. Tagmentation involves the modification of DNA by a transposomal complex containing a transposase enzyme complexed with one or more tags (such as adapter sequences) containing transposon end sequences (hereinafter referred to as transposons). Tagmentation can result in simultaneous fragmentation of DNA and ligation of adapters to the 5' ends of both strands of the double-stranded fragments.

[0482] A transposition reaction is a reaction in which one or more transposons are inserted into a target nucleic acid at a random or nearly random site. Components of a transposition reaction may include a transposase (or other enzymes capable of fragmenting and tagging nucleic acids as described herein, e.g., integrase), a transposon element comprising a double-stranded transposon terminal sequence bound to the enzyme, and an adapter sequence bound to one of the two transposon terminal sequences. One strand of the double-stranded transposon terminal sequence is transposed to one strand of the target nucleic acid, while the complementary transposon terminal sequence strand is not transposed (i.e., a non-transposition transposon sequence). The adapter sequence may optionally or desiredly include one or more functional sequences (e.g., primer sequences).

[0483] The term “transposon end” refers to double-stranded nucleic acid DNA that exhibits only the nucleotide sequence (“transposon end sequence”) necessary to form a complex with a transposase or integrase enzyme that functions in an in vitro transposition reaction. In some embodiments, a transposon end can form a functional complex with a transposase in a transposition reaction. As a non-limiting example, as described in the disclosure of U.S. Patent Application Publication 2010 / 0120098, which is incorporated entirely herein by reference, a transposon end may be a 19-bp outer end (“outer end, OE”), an inner end (“inner end, IE”), or a “mosaic end” (“mosaic end, ME”) transposon recognized by wild-type or mutant Tn5 transposases. The transposon ends may include the terminals, or the R1 and R2 transposon ends. The transposon ends may include any nucleic acid or nucleic acid analog suitable for forming a functional complex with a transposase or integrase enzyme in an in vitro transposition reaction. For example, the transposon ends may include DNA, RNA, modified bases, non-native bases, and modified backbones, and may contain nicks in one or both strands. The term “DNA” is used throughout this disclosure in reference to transposon end compositions, but it should be understood that any suitable nucleic acid or nucleic acid analog may be utilized at the transposon ends.

[0484] Similarly, the term “transfer strand” refers to the transposed portion of both transposon ends. Similarly, the term “non-transfer strand” refers to the non-transfer portion of both “transposon ends.” The 3' end of the transfer strand is bound to or transferred to the target DNA in the in vitro transposition reaction. The non-transfer strand, which exhibits a transposon end sequence complementary to the transposed transposon end sequence, is not bound to or transferred to the target DNA in the in vitro transposition reaction.

[0485] In some embodiments, the transposon is a fork-type adapter transposon. The fork-type adapter transposon comprises two chains. In some embodiments, the second chain of the fork-type adapter transposon comprises an adapter array and an array that is fully or partially complementary to the first chain of the first fork-type adapter transposon. The arrays having full or partial complementarity in the first and second chains allow the two chains to hybridize together to form a fork-type structure.

[0486] In some embodiments, two or more transposome complexes are immobilized on the surface of a solid support. In some embodiments, fragments can be prepared with different tags based on the use of different transposomes.

[0487] In some embodiments, the solid support comprises two pools of immobilized transposome complexes. In some embodiments, the first pool of transposome complexes comprises (a) a transposase, (b) a first transposon comprising a 3' transposon terminal sequence, a first read sequencing adapter sequence, and a 5' affinity moiety, and (c) a second transposon comprising a 3' complement of a 5' sequence and a hybridization sequence that is fully or partially complementary to the 3' transposon terminal sequence. In some embodiments, the second pool of transposome complexes comprises (a) a transposase, (b) a first transposon comprising a 3' transposon terminal sequence, a second read sequencing adapter sequence, and a 5' affinity moiety, and (c) a second transposon comprising a 5' sequence and a 3' hybridization sequence that is fully or partially complementary to the 3' transposon terminal sequence. In some embodiments, each first transposon is immobilized by binding of its 5' affinity moiety to a binding portion on the surface of the solid support.

[0488] In some embodiments, the first pool of immobilized transposome complexes includes a first fork-type adapter comprising a first oligonucleotide containing P5.R1 and a second oligonucleotide containing X' (complementary to the hybridization sequence). In some embodiments, the second pool of immobilized transposome complexes includes a second fork-shaped adapter comprising a first oligonucleotide containing P7.R2 and a second oligonucleotide containing X (hybridization sequence). Such exemplary embodiments are shown in Figure 34.

[0489] In some embodiments, the transposomal complex comprises dimers of two molecules of the transposase. In some embodiments, the transposomal complex comprises a homodimer and / or a heterodimer.

[0490] In some embodiments, the transposome complex is a homodimer, where two molecules of the transposase each bind to a first and second transposon of the same type (for example, the sequences of the two transposons bound to each monomer are the same, forming a “homodimer”). In some embodiments, the compositions and methods described herein employ two populations of transposome complexes. In some embodiments, the transposons in each population are the same. As used herein, “homodimer” refers to a transposome dimer that has the same transposon sequence at both sites. In some embodiments, the compositions and methods described herein use a population of transposome complexes assembled by preparing a first transposome complex by contacting a first fork-shaped adapter with a transposase, assembling a second transposome complex by contacting a second fork-shaped adapter with a transposase, and then pooling the first and second transposome complexes together. In some embodiments, the pool of transposome complexes includes a homodimer containing a first fork-shaped adapter and a homodimer containing a second fork-shaped adapter.

[0491] In some embodiments, the transposome complex is a heterodimer, where two molecules of transposase are each bound to different fork-shaped adapters containing a first and a second transposon (for example, the sequences of the two transposons bound to each monomer of the transposome complex are different, forming a "heterodimer").

[0492] In some embodiments, the compositions and methods described herein utilize a population of transposome complexes assembled by pooling a first fork-type adapter and a second fork-type adapter together with a transposase and assembling the pool of transposome complexes. After this pooling, the expected proportions of the assembled transposome complexes are 25% homodimer transposome complexes containing the first fork-type adapter, 25% homodimer transposome complexes containing the second fork-type adapter, and 50% heterodimer transposome complexes containing the first and second fork-type adapters. In some embodiments, the first and / or second pools of transposome complexes are homodimers or heterodimers. In some embodiments, the first and second pools of transposome complexes are homodimers or heterodimers. Exemplary homodimers, heterodimers, and immobilized homodimers comprising solid supports, and methods of using them, are disclosed in U.S. Patent No. 9,683,230, which is incorporated herein by reference in whole. Figure 35 shows an exemplary solid support containing two pools of homodimers, where all homodimers are immobilized on the surface of the solid support. Using the two pools of homodimers or the pool containing heterodimers, tagged double-stranded fragments can be generated, in which at least some fragments contain a tag derived from a transposome complex contained in the first pool at one end and a tag derived from a transposome complex contained in the second pool at the other end.

[0493] In some embodiments, one or more transposons are adapters, barcode arrays. The transposon includes at least one of a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence. In other words, the transposon may include additional sequences that are useful in any way the user wishes to perform, such as sequencing. In some embodiments, one or more transposons include an index sequence and / or a UMI. Transposons containing UMIs and methods of using them are described in International Publications 2019 / 108972, 2018 / 136248, 2016176091, and 202014437, which are each incorporated in their entirety herein.

[0494] In some embodiments, the first transposon contained in the first pool of the transposome complex and / or the first transposon contained in the second pool of the transposome complex include a sample index. In some embodiments, both the first transposon contained in the first pool of the transposome complex and the first transposon contained in the second pool of the transposome complex include a sample index. In a typical example, the embodiment may include a first transposon containing i5 contained in the first pool of the transposome complex and a first transposon containing i7 contained in the second pool of the transposome complex, as shown in Figure 46A.

[0495] In some embodiments, the second transposon contained in the first pool of the transposome complex and / or the second transposon contained in the second pool of the transposome complex includes a sample index and / or UMI. In some embodiments, both the second transposon contained in the first pool of the transposome complex and the second transposon contained in the second pool of the transposome complex include a sample index.

[0496] In some embodiments, both the second transposon contained in the first pool of the transposome complex and the second transposon contained in the second pool of the transposome complex contain a UMI. In a typical example, the embodiment may include a second transposon containing i8 contained in the first pool of the transposome complex and a second transposon containing i6 contained in the second pool of the transposome complex, as shown in Figure 46B, where i6 and i8 function as a UMI.

[0497] In some embodiments, the first and second transposons contained in both the first and second pools of the transposome may contain either a sample index sequence or a UMI. When such transpososomes are used in the method of the present invention, polynucleotides as shown in Figure 46C may be produced.

[0498] In some embodiments, a method for generating one or more double-stranded linked nucleic acid sequencing templates (as shown in Figure 37) comprises applying a sample containing double-stranded nucleic acids immobilized on a solid support, and tagging the double-stranded nucleic acids to generate tagged double-stranded fragments containing inserts derived from the double-stranded nucleic acids, the double-stranded fragments being immobilized on the solid support by a 5' affinity moiety binding to a binding moiety on the surface of the solid support. In some embodiments, the 5' affinity moiety is contained in a first transposon (i.e., a first strand of a fork-shaped adapter contained in a transposome complex).

[0499] In some embodiments, the transposome complex is then released from the double-stranded fragment. In some embodiments, the release of the transposome complex from the double-stranded fragment is carried out using SDS and washing.

[0500] In some embodiments, the method involves extending two transposomal complexes after their release. This includes ligating chain fragments. In some embodiments, extension and ligation include providing polymerase, dNTPs, and extension buffer (ELMT).

[0501] In some embodiments, the method involves denaturing an extended and ligated double-stranded fragment into a single-stranded fragment, the single-stranded fragment containing the 5' affinity moiety remaining immobilized on a solid support, as shown in Figure 38. In some embodiments, the denaturation includes heating the solid support or applying a chemical denaturant. In some embodiments, the denaturation includes raising the temperature of the solid support to 90°C or higher.

[0502] In some embodiments, the method includes hybridizing a hybridization sequence contained in a first immobilized single-stranded fragment to a complement of a hybridization sequence contained in a second immobilized single-stranded fragment, thereby forming a crosslink. In some embodiments, hybridization includes cooling a solid support and / or applying a hybridization buffer. In some embodiments, cooling includes lowering the temperature of the solid support to 60°C or below. In some embodiments, the hybridization buffer contains a high salt concentration, optionally, the high salt concentration being 750 mM NaCl.

[0503] In some embodiments, a hybridization sequence (X or HYB) contained in a first single-stranded fragment can hybridize to a complement (X' or HYB') of a hybridization sequence contained in a second single-stranded fragment. In some embodiments, the hybridization sequence and / or its complement are bound to a blocking oligonucleotide that is fully or partially complementary to the hybridization sequence or its complement, and denaturation involves denaturing the blocking oligonucleotide to deblock the hybridization seq...

Claims

1. It is a polynucleotide, a. A 5'-terminal polynucleotide containing the first lead primer binding sequence, b. A first insert sequence located on the 3' side of the 5' terminal polynucleotide, wherein the first insert sequence is derived from the target nucleic acid, c. A linking sequence located at the 3' side of the first insert sequence, which includes a second lead primer binding sequence and a hybridization sequence, d. A second insert sequence located on the 3' side of the linked sequence, which is derived from a non-adjacent sequence of the target nucleic acid or from a target nucleic acid different from the first insert sequence, e. A polynucleotide containing a 3' terminal polynucleotide sequence.

2. It is a polynucleotide, a. A 3'-terminal polynucleotide containing the first read primer binding sequence, b. The first insert sequence at the 5' end of the 3' terminal polynucleotide derived from the target nucleic acid, c. A linking sequence comprising a second read primer binding sequence orthogonal to the first read primer binding sequence, wherein the second read primer binding sequence comprises a hybridization sequence, d. A second insert sequence located at the 5' end of the linked sequence and derived from a non-adjacent sequence of the target nucleic acid, or derived from a target nucleic acid different from the first insert sequence, e. A polynucleotide located at the 5' end of the polynucleotide and containing an attachment sequence, comprising: The 3' terminal polynucleotide, the ligation sequence, and the attached polynucleotide are polynucleotides that do not originate from the target nucleic acid.

3. The polynucleotide according to claim 1 or claim 2, wherein the two insert sequences are derived from different target nucleic acids.

4. The polynucleotide according to any one of claims 1 to 3, wherein the first insert sequence and the second insert sequence each independently contain 40 to 400 nucleotides, 100 to 200 nucleotides, or 150 nucleotides.

5. The polynucleotide according to any one of claims 1 to 4, wherein the first read primer binding sequence includes a first adapter sequence.

6. The polynucleotide according to any one of claims 1 to 5, wherein the first read primer binding sequence further comprises a complement of the transposon terminal sequence.

7. The polynucleotide according to any one of claims 2 to 6, wherein the linked sequence comprises (a) the hybridization sequence and optionally (b) a complement of the 3' transposon terminal sequence of the hybridization unit and the 5' transposon terminal sequence of the hybridization unit.

8. The polynucleotide according to any one of claims 2 to 7, wherein the attached polynucleotide comprises a second adapter sequence and optionally a transposon terminal sequence.

9. The 3'-terminal polynucleotide and / or the attached polynucleotide are, independently of each other, a barcode sequence, a unique molecular identifier (UMI) sequence, a capture sequence, or a cleavage sequence. A polynucleotide according to any one of claims 2 to 8, comprising at least one of the above.

10. The polynucleotide according to any one of claims 2 to 8, wherein the polynucleotide is immobilized on a solid support.

11. The polynucleotide according to any one of claims 2 to 12, comprising at least one insertion unit between the second insert sequence and the attached polynucleotide, the insertion unit comprising the 5' terminal insert sequence and the 3' terminal ligation sequence including a read primer binding sequence, wherein the read primer binding sequence is orthogonal to the other read primer binding sequences.

12. A composition comprising a polynucleotide according to any one of claims 1, 3 to 6, and 11, and its complement, wherein the complement is a. A 5'-terminal complement containing the first complementary read-primer binding sequence, b. The complementary sequence of the second insert sequence located on the 3' side of the 5' terminal complement, c. A complementary linking array located on the 3' side of the complementary array of the second insert array, wherein the second insert array is i. The second complementary read-primer binding sequence, ii. A complementary hybridization sequence and a complementary linking sequence including, d. The complementary sequence of the first insert sequence located on the 3' side of the complementary linking sequence, e. A composition comprising a 3'-terminal complement.

13. A composition comprising a polynucleotide according to any one of claims 2 to 11 and its complement, wherein the complement is a. A 3'-terminal complement comprising a first complementary read-primer binding sequence, wherein the first complementary read-primer binding sequence is orthogonal to the first read-primer binding sequence and the second read-primer binding sequence, b. The complement of the second insert sequence located on the 5' side of the 3' terminal complement, c. A complementary linking sequence located on the 5' side of the complement of the second insert sequence, including a second complementary read-primer binding sequence from 3' to 5', wherein the second complementary read-primer binding sequence is orthogonal to the first read-primer binding sequence, the second read-primer binding sequence, and the first complementary read-primer binding sequence. d. The complementary of the first insert array on the 5' side of the complementary link array, e. A composition comprising a complementary attached polynucleotide at the 5' end, which includes a complementary attached sequence.

14. A transposome complex, a. Transposase and, b. A first transposon comprising a complement of a first read-primer binding sequence, wherein the complement of the first read-primer binding sequence is i. The 3' portion containing the transposon terminal sequence, ii. A first transposon comprising a complement of the first adapter sequence, c. A second transposon, i. The 5' portion containing the complement of the transposon terminal sequence, ii. A transposomal complex comprising a second transposon containing a complementary hybridization sequence.

15. A transposome complex, a. Transposase and, b. A first transposon comprising an attached polynucleotide, wherein the attached polynucleotide is i. The 5' portion containing the attachment sequence, ii. The 3' portion containing the second lead primer binding sequence, iii. The 3' portion containing the transposon terminal sequence, iv. The first transposon, including the adapter and the 3' portion, c. A second transposon, i. The 5' portion containing the complement of the transposon terminal sequence, ii. A transposomal complex comprising a hybridization sequence and a second transposon.

16. A composition or kit comprising the transposome complex described in claim 14 or 15.

17. A composition or kit, a. A solid support, wherein optionally the optional support is a bead, b. Components for generating a transposome complex, i. Transposase and, ii. An oligonucleotide for generating a double-stranded oligonucleotide, comprising: a first oligonucleotide comprising a 3' transposon terminal sequence and a 5' first adapter sequence; a second oligonucleotide comprising a 5' transposon terminal sequence and a 3' second adapter sequence, wherein the 5' transposon terminal sequence is complementary to the 3' transposon terminal sequence, and the first and second adapter sequences are not the same; and a component thereof. c. A composition or kit comprising a first primer set and a second primer set for adding an attachment sequence and a hybridization sequence to a fragment by PCR, wherein the first primer set comprises a primer for adding a hybridization sequence and a first attachment sequence to the fragment, and the second primer set comprises a primer for adding a complementary hybridization sequence and a second attachment sequence to the fragment, and the first attachment sequence and the second attachment sequence are not the same.

18. An adapter composition or kit comprising a first fork-type adapter complex and a second fork-type adapter complex, The first fork-type adapter complex is a. Complementary attached polynucleotides, i. The 5' portion containing the complementary attachment sequence, ii. A 3' portion including an adapter, and a complementary attached polynucleotide, b. Hybridized polynucleotides, i. A 5' portion that includes a complementary part of the adapter and hybridizes thereto, ii. A hybridization polynucleotide comprising a complementary hybridization sequence, wherein the complementary hybridization sequence is not complementary to the complementary attached polynucleotide, The second fork-type adapter complex is, a. Adhering polynucleotides, i. The 5' portion containing the attachment sequence, ii. A 3' portion including the adapter, and an attached polynucleotide, b. Hybridized polynucleotides, i. A 5' portion that includes a complementary part of the adapter and hybridizes thereto, ii. An adapter composition or kit comprising a hybridization sequence and a hybridization polynucleotide, wherein the hybridization sequence is not complementary to the attached polynucleotide.

19. A method for generating a concatenated nucleic acid sequencing template, a. Attaching the first read primer binding sequence to the 3' end of the first insert sequence derived from the first target nucleic acid, b. Attaching a hybridization sequence to the 5' end of the first insert sequence, c. Attaching the complementary of the hybridization sequence to the 3' end of a second insert sequence derived from the discontinuous region of the first target nucleic acid or the second target nucleic acid, d. Annealing the hybridization sequence to its complement to form a hybridized adduct, e. The process includes synthesizing a fully double-stranded nucleic acid sequencing template from the hybridized adduct, The region between the first insert sequence and the second insert sequence is the second read A primer binding sequence comprising the hybridization sequence and a second read primer binding sequence orthogonal to the first read primer binding sequence, A method for generating a concatenated nucleic acid sequencing template.

20. A method for generating a concatenated nucleic acid sequencing template, a. Contacting a first sample containing a first target nucleic acid with a first transposome complex and a second transposome complex, Each transposome complex, i. Transposase and, ii. A first transposon comprising a 3' portion containing a transposon terminal sequence and a 5' portion containing an adapter sequence, iii. A second transposon comprising a 5' portion containing the complement of the transposon terminal sequence and hybridizing thereto, The adapter sequence in the first transposome complex is a complement of the first adapter sequence, and the adapter sequence in the second transposome complex is a second adapter sequence. The first target nucleic acid is fragmented under conditions sufficient to produce a first tagged product containing an insert sequence derived from the first target nucleic acid, with one end tagged by the transposon of the first transposome complex and the other end tagged by the transposon of the second transposome complex. b. To optionally add a complementary attachment sequence to the 3' end of the first tagged product and a complementary hybridization sequence to the 5' end of the first tagged product by polymerase chain reaction to form a first modified tagged product, c. Contacting a second sample containing a second target nucleic acid with the transposome complex, under conditions sufficient to fragment the second target nucleic acid and generate a second tagged product containing an insert sequence derived from the second target nucleic acid, where one end is tagged with the transposon of the first transposome complex and the other end is tagged with the transposon of the second transposome complex. d. To optionally add an attachment sequence to the 3' end of the second tagged product and a hybridization sequence to the 5' end of the second tagged product by polymerase chain reaction to form a second modified tagged product, e. The hybridization sequence of the first modified tagged product is annie to the complementary of the hybridization sequence in the second modified tagged product. Ringing and forming hybridized adducts, f. Synthesizing a fully double-stranded nucleic acid sequencing template from the hybridized adduct, wherein the linked nucleic acid sequencing template is i. A first read primer binding sequence located at the 3' end of the insert sequence derived from the second target nucleic acid, comprising the first adapter sequence and a complement of the transposon terminal sequence, ii. A second read-primer binding sequence located between the two insert sequences, comprising the transposon terminal sequence and the hybridization sequence, A method comprising the following: the first read-primer binding sequence is orthogonal to the second read-primer binding sequence.

21. A method for generating a concatenated nucleic acid sequencing template, a. Contacting a first sample containing a first target nucleic acid with a first transposome complex, wherein the first transposome complex is i. Transposase and, ii. A first transposon comprising a 3' portion containing a transposon terminal sequence and a 5' portion containing an attachment sequence and a complement of the first adapter sequence, iii. A second transposon comprising a 5' portion containing the complement of the transposon terminal sequence and hybridizing thereto, The first target nucleic acid is fragmented under conditions sufficient to produce a first tagged product containing an insert sequence derived from the first target nucleic acid, with each end tagged by the transposon of the first transpososome complex. b. To optionally add a hybridization sequence complement to the 5' end of the first tagged product by polymerase chain reaction to form a first modified tagged product, c. Contacting a second sample containing a second target nucleic acid with a second transposome complex, wherein the second transposome complex is i. Transposase and, ii. A first transposon comprising a 3' portion containing a transposon terminal sequence and a 5' portion containing a second adapter sequence and a complementary attachment sequence, iii. A second transposon comprising a 5' portion containing the complement of the transposon terminal sequence and hybridizing thereto, The second target nucleic acid is fragmented under conditions sufficient to produce a second tagged product containing an insert sequence derived from the second target nucleic acid, with each end tagged by the transposon of the second transposomal complex. d. To optionally add a complementary hybridization sequence to the 5' end of the second tagged product by polymerase chain reaction to form a second modified tagged product, e. Annealing the hybridization sequence of the first modified tagged product to the complement of the hybridization sequence in the second modified tagged product to form a hybridized adduct, f. Synthesizing a fully double-stranded nucleic acid sequencing template from the hybridized adduct, wherein the linked nucleic acid sequencing template is i. A first read primer binding sequence located at the 3' end of the insert sequence derived from the second target nucleic acid, comprising the first adapter sequence and a complement of the transposon terminal sequence, ii. A second read-primer binding sequence located between the two insert sequences, comprising the transposon terminal sequence and the hybridization sequence, A method comprising the following: the first read-primer binding sequence is orthogonal to the second read-primer binding sequence.

22. A method for generating a concatenated nucleic acid sequencing template, a. i. A first double-stranded polynucleotide containing the first target nucleic acid is used as the first restriction enzyme, and ii. The second double-stranded polynucleotide containing the second target nucleic acid is brought into contact with the second restriction enzyme. The method involves producing a first polynucleotide and a second polynucleotide having a compatibility overhang, wherein the restriction enzyme is selected from type II, type IIS, type IIP, and type IIT restriction enzymes. b. A method comprising using a ligase to attach the compatibility overhangs of the first polynucleotide and the second polynucleotide.

23. Before the aforementioned contact step, a. The first restriction enzyme cleavage site is optionally attached to the first target nucleic acid by using an adapter, and the first double-stranded polynucleotide is generated by primer extension. b. The method according to claim 22, wherein the second restriction enzyme cleavage site is optionally attached to a second target nucleic acid by using an adapter, and the second double-stranded polynucleotide is generated by primer extension.

24. A method for generating a concatenated nucleic acid sequencing template, a. The first nucleic acid source and the second nucleic acid source are sheared or digested to produce a first library of nucleic acid fragments and a second library of nucleic acid fragments, respectively. b. Attaching a first adapter to each nucleic acid fragment derived from the first nucleic acid source, and attaching a second adapter to each nucleic acid fragment derived from the second nucleic acid source, i. Contacting the nucleic acid fragment with a first polymerase to produce a nucleic acid fragment having a blunt end, ii. Phosphorylation of the 5'-hydroxyl group of the nucleic acid fragment using kinase, iii. Adding 3'-adenine to the nucleic acid fragment using a second polymerase, iv. The method includes ligating the first adapter to each nucleic acid fragment of the first library, and ligating the second adapter to each nucleic acid fragment of the second library. c. Selectively mixing and annealing the first nucleic acid library and the second nucleic acid library by PCR, i. The nucleic acid is denatured at high temperature, ii. The A and A' sequences hybridize with each other at lower temperatures, d. A method comprising optionally synthesizing a fully double-stranded nucleic acid sequencing template by PCR.

25. A method for sequencing a concatenated nucleic acid sequencing template, a. Sequence the first insert sequence of the polynucleotide according to any one of claims 1 to 11 by initiating sequencing using a first read sequencing primer complementary to the first read primer binding sequence, b. A method comprising sequencing a second insert sequence by initiating sequencing using a second read sequencing primer complementary to a second read primer binding sequence.

26. The method according to any one of claims 20 to 24, comprising partitioning a sample containing one or more target double-stranded nucleic acids into a plurality of different compartments, wherein the generation of a concatenated nucleic acid sequencing template is performed within the different compartments.

27. It is a polynucleotide, a. A 5'-terminal polynucleotide containing the first read sequencing primer sequence, b. An insert sequence derived from a target nucleic acid, the insert sequence located on the 3' side of the 5' terminal polynucleotide, c. A hybridization sequence located on the 3' side of the insert sequence, d. A copy of the insert sequence on the 3' side of the hybridization sequence or a second insert sequence on the 3' side of the hybridization sequence, e. A polynucleotide comprising a 3'-terminal polynucleotide containing a complement of the second read sequencing primer sequence.

28. The aforementioned polynucleotide is a. 5'-P5-A14-insert-HYB-insert-B15'-P7'-3', b. 5'-P7-B15-insert-HYB'-insert-A14'-P5'-3', c. 5'-P5-A14-insert1-HYB-insert2-B15'-P7'-3', or d. The polynucleotide according to claim 27, having the structure 5'-P7-B15-insert1-HYB'-insert2-A14'-P5'-3', where HYB is a hybridization sequence and HYB' is a complement of the hybridization sequence.

29. A fork-shaped adapter containing two polynucleotide chains, a. A first strand containing the sequencing primer sequence, b. A fork-type adapter comprising a second chain containing a 3' hybridization sequence or its complement, wherein the 3' end of the first chain is fully or partially complementary to the 5' end of the second chain, and the hybridization sequence or its complement is bound to a blocking oligonucleotide that is fully or partially complementary to the hybridization sequence or its complement.

30. The fork adapter according to claim 29, wherein the first chain includes 5' affinity elements that can be bonded to affinity bonding partners on a solid support or beads, and optionally the affinity elements are connected via linkers attached to the first chain.

31. A composition or kit comprising two fork-type adapters according to any one of claim 29 or 30, a. The first fork-type adapter comprises a first strand containing a first read sequencing primer sequence and a second strand containing a complementary hybridization sequence. b. The second fork-type adapter comprises a first strand containing a second read sequencing primer sequence and a second strand containing a hybridization sequence, A composition or kit comprising one or both fork-type adapters containing a blocking oligonucleotide.

32. A method for generating one or more concatenated nucleic acid sequencing templates, a. A sample containing double-stranded nucleic acid fragments, each containing an insert prepared from a target nucleic acid, Contacting a composition or kit according to any one of claims 18 or 29-31, which includes two fork-shaped adapters, wherein one or both fork-shaped adapters contain a blocking oligonucleotide, b. Ligating the fork-type adapter to the double-stranded fragment to prepare a tagged double-stranded fragment, c. Fixing the tagged double-stranded fragment onto a solid support, d. (1) Denaturing the fixed tagged double-stranded fragment to generate a fixed single-stranded fragment, and (2) denaturing the blocking oligonucleotide to deblock the hybridization sequence and its complement, e. Hybridizing two fixed single-stranded fragments to form a bridge by linking the hybridization sequence in the first fragment to the complementary hybridization sequence in the second fragment, f. A method comprising extending each single-stranded fragment from its 3' end to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both fixed single-stranded fragments.

33. A method for generating one or more concatenated nucleic acid sequencing templates, a. Contacting a sample containing a double-stranded target nucleic acid with two pools of transposome complexes in solution, wherein the first pool of transposome complexes is i. Transposase and, ii. A first transposon comprising a 3' transposon terminal sequence and a first read sequence determination adapter sequence, iii. A second transposon comprising a 5' sequence and a 3' complement of a hybridization sequence that are completely or partially complementary to the 3' transposon terminal sequence, wherein the second pool of the transposome complex is i. Transposase and, ii. A first transposon comprising a 3' transposon terminal sequence and a second read sequence determination adapter sequence, iii. A second transposon comprising a 5' sequence and a 3' hybridization sequence that are completely or partially complementary to the 3' transposon terminal sequence, wherein one or both of the second transposons contain a blocking oligonucleotide. b. Tag the double-stranded nucleic acid to generate tagged double-stranded fragments, c. Releasing the transposomal complex from the double-stranded fragment, d. Extending and ligating the double-stranded fragments, e. Fixing the tagged double-stranded fragment onto a solid support, f. (1) Denaturing the fixed tagged double-stranded fragment to generate a fixed single-stranded fragment, and (2) denaturing the blocking oligonucleotide to deblock the hybridization sequence and its complement, g. Hybridizing two fixed single-stranded fragments to form a bridge by linking the hybridization sequence in the first fragment to the complementary hybridization sequence in the second fragment, h. A method comprising extending each single-stranded fragment from its 3' end to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both fixed single-stranded fragments.

34. The method according to claim 33, wherein the first pool or the second pool of the transposome complex comprises the transposome complex described in claim 15, and the first read sequencing adapter sequence comprises the first read primer binding sequence.

35. A method for generating one or more concatenated nucleic acid sequencing templates, a. Compartmenting a sample containing the target double-stranded nucleic acid into multiple different compartments, b. Preparing fragments each containing the double-stranded nucleic acid-derived insert in a plurality of different compartments, c. The plurality of different compartments are brought into contact with the composition or kit according to any one of claims 18 or 29-31, which includes two fork-type adapters, wherein one or both of the fork-type adapters include a blocking oligonucleotide. d. Ligating the fork-type adapter to the double-stranded fragments to prepare tagged double-stranded fragments in the plurality of different compartments, e. (1) Denaturing the fixed tagged double-stranded fragment to generate a single-stranded fragment, and (2) denaturing the blocking oligonucleotide to deblock the hybridization sequences and complementary hybridization sequences within the plurality of different compartments, f. Hybridizing two single-stranded fragments within the same section to form a bridge by linking the hybridization sequence in the first fragment to the complementary of the hybridization sequence in the second fragment, g. A method comprising extending each single-stranded fragment from its 3' end to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both single-stranded fragments within the same compartment.

36. The method according to claim 26 or 35, wherein the compartment is a well, a tube, or a droplet, and / or the insert contained in the same linked sequencing template is prepared from the same target nucleic acid.

37. The method according to any one of claims 26, 35, or 36, wherein the partitioning separates the most different haplotypes into different partitions, and the method is used for haplotype phasing.

38. A solid support comprising two pools of immobilized transposome complexes, a. The first pool of the transposome complex is i. Transposase and, ii. A first transposon comprising a 3' transposon terminal sequence, a first read sequencing adapter sequence, and a 5' affinity moiety, iii. A second transposon comprising a 5' sequence and a 3' complement of a hybridization sequence that are completely or partially complementary to the 3' transposon terminal sequence, wherein the second pool of the transposome complex is i. Transposase and, ii. A first transposon comprising a 3' transposon terminal sequence, a second read sequencing adapter sequence, and a 5' affinity moiety, iii. A solid support comprising a second transposon having a 5' sequence and a 3' hybridization sequence that are completely or partially complementary to the 3' transposon terminal sequence, wherein each first transposon is fixed by the binding of its 5' affinity portion to a binding portion on the surface of the solid support.

39. The method according to claim 38, wherein the first pool or the second pool of the transposome complex comprises the transposome complex according to any one of claim 15, and the first read sequencing adapter sequence comprises the first read primer binding sequence.

40. A method for generating one or more double-stranded linked nucleic acid sequencing templates, a. A sample comprising double-stranded nucleic acid immobilized on a solid support according to claim 38 or 39 To apply, b. Tagmenting the double-stranded nucleic acid to generate a tagged double-stranded fragment containing an insert derived from the double-stranded nucleic acid, wherein the double-stranded fragment is fixed to the solid support by binding of its 5' affinity moiety to a binding portion on the surface of the solid support, c. Releasing the transposomal complex from the double-stranded fragment, d. Extending and ligating the double-stranded fragments, e. The double-stranded fragment is modified into a single-stranded fragment, wherein the single-stranded fragment containing the 5' affinity moiety remains fixed on the solid support. f. Hybridizing the hybridization sequence contained in the first fixed single-stranded fragment to the complementary hybridization sequence contained in the second fixed single-stranded fragment, thereby forming a crosslink. g. A method comprising extending and generating a double-stranded linked nucleic acid sequencing template.

41. The method according to claim 40, wherein the first fixed fragment and the second fixed fragment are fixed in close proximity on the solid support, and by being in close proximity, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more nucleotides contained in the hybridization sequence contained in the first fixed fragment can bind to nucleotides contained in the complement of the hybridization sequence contained in the second fixed fragment, or the first fixed fragment and the second fixed fragment are fixed to each other within 20 to 500 nanometers on the surface of the solid support.

42. The method according to claim 40 or 41, wherein both the first immobilized fragment and the second immobilized fragment are prepared from the same double-stranded nucleic acid, and the double-stranded linked nucleic acid sequencing template includes two inserts derived from the same double-stranded nucleic acid.

43. The method according to claim 42, wherein the two inserts are derived from two adjacent sequences contained in the same double-stranded nucleic acid, and the adjacent sequences are separated by 100 or fewer nucleotides, 200 or fewer nucleotides, 300 or fewer nucleotides, 400 or fewer nucleotides, 500 or fewer nucleotides, 700 or fewer nucleotides, or 1,000 or fewer nucleotides in the double-stranded nucleic acid.

44. a. Releasing a double-stranded nucleic acid sequencing template from the solid support, b. The method according to any one of claims 40 to 43, further comprising determining the arrangement of the mold and determining the arrangement of inserts contained in the mold.

45. A method for generating one or more concatenated nucleic acid sequencing templates, a. Compartmenting a sample containing the target double-stranded nucleic acid into multiple different compartments, b. Tagmenting the double-stranded nucleic acid to generate tagged double-stranded fragments containing inserts derived from the double-stranded nucleic acid in a plurality of different compartments, wherein the tagmentation is performed using two pools of transposome complexes, the first pool of the transposome complexes being i. Transposase and, ii. A first transposon comprising a 3' transposon terminal sequence and a first read sequence determination adapter sequence, iii. A second transposon comprising a 5' sequence and a 3' complement of a hybridization sequence that are completely or partially complementary to the 3' transposon terminal sequence, The second pool of the transposome complex is i. Transposase and, ii. A first transposon comprising a 3' transposon terminal sequence and a second read sequence determination adapter sequence, iii. A second transposon comprising a 5' sequence and a 3' hybridization sequence that are completely or partially complementary to the 3' transposon terminal sequence, c. Denature the tagged double-stranded fragment to generate a single-stranded fragment, d. Hybridizing two single-stranded fragments within the same region by linking the hybridization sequence in the first fragment to the complementary hybridization sequence in the second fragment, e. A method comprising extending each single-stranded fragment from its 3' end to generate a double-stranded linked nucleic acid sequencing template containing inserts derived from both single-stranded fragments.

46. The method according to claim 45, wherein the double-stranded linked nucleic acid sequencing template is generated solely from the hybridization of two single-stranded fragments located in the same compartment.

47. The method according to claim 45 or 46, wherein the hybridization sequence and / or its complement is bound to a blocking oligonucleotide that is completely or partially complementary to the hybridization sequence or its complement, and the denaturation comprises denaturing the blocking oligonucleotide to unblock the hybridization sequence and / or its complement.

48. The method according to any one of claims 45 to 47, wherein the partitioning separates the most different haplotypes into different partitions, and the method is used for haplotype phasing.

49. The method according to any one of claims 19 to 24, 26, 32 to 37, or 39 to 48, further comprising determining the arrangement of the mold.

50. The method according to claim 49, wherein sequencing is performed using sequencing primers that bind to A14, B15, and / or hybridization sequences (HYB), and optionally, the sequencing includes a dark cycle and data is not recorded for any portion of the sequencing.

51. a. Evaluating the arrangement of inserts contained in the same mold, b. The method according to claim 49 or 50, further comprising determining proximity data for sequences contained in the double-stranded nucleic acid based on inserts contained in the same template.

52. The method according to claim 51, wherein the proximity data determines that the insert sequence (or its complement) was included in the same target nucleic acid.

53. a. Evaluating the sequencing results of multiple sequences of a given insert prepared from different templates, b. Determining the event of non-standard base pairing based on the sequencing data, wherein the sequencing data is i. The insert and its complement included in the same linked sequence determination template, and / or ii. The method according to any one of claims 49 to 52, further comprising the fact that the insert is contained in a plurality of linked sequence determination templates.

54. a. Evaluating the sequencing results of multiple sequences of a given insert prepared from different templates, b. Based on the sequence determination data, the sequence determination result for this insert is This involves correcting Lar, and the sequence determination data is i. The insert and its complement included in the same linked sequence determination template, and / or ii. The method according to any one of claims 49 to 52, further comprising the fact that the insert is contained in a plurality of linked sequence determination templates.

55. A method for identifying modified cytosines contained in insert sequences included in a concatenated sequencing template, a. To prepare a double-stranded linked sequencing template, wherein each strand contains an insert sequence and a copy of the insert sequence, and the two strands are complementary to each other. b. The double-stranded linked sequencing template is subjected to conditions for modifying and / or unmodified cytosine, c. Preparing the amplicons of each strand of the double-stranded linked sequence determination template, d. Sequence the amplicon and evaluate the sequencing results for the insert sequence in the amplicon generated from each strand and the copy of the insert sequence, e. A method comprising determining the position of modified cytosines contained in the insert sequence based on the sequence of each strand of the double-stranded linked sequencing template.

56. The method according to claim 55, wherein the modified cytosine is methylated or hydroxymethylated cytosine.

57. The method according to claim 55 or 56, wherein the linked sequence determination template is prepared by the method according to any one of claims 19 to 24, 26, 32 to 37, or 39 to 48.

58. The method according to claim 57, wherein the extension for generating the double-stranded sequence determination template is carried out using a reaction solution containing methylated dCTP.

59. The method according to any one of claims 55 to 58, wherein modified cytosine or unmodified cytosine is modified, and optionally, the modified cytosine is modified by TET-assisted pyridineborane sequencing (TAPS) treatment, or the unmodified cytosine is modified by sodium bisulfite or enzymatic treatment.

60. The method according to claim 59, wherein the modified cytosine is modified, the position of the modified cytosine is determined by the presence of (T, C) in the insert sequence and the copy of the insert sequence, respectively, the position of the unmodified cytosine is determined by the presence of (C, C) in the insert sequence and the copy of the insert sequence, respectively, and the modified cytosine and the unmodified cytosine are paired with G' in the complementary chain.

61. The method according to claim 59, wherein the unmodified cytosine is modified, the position of the modified cytosine is determined by the presence of (C, T) in the insert sequence and the copy of the insert sequence, respectively, the position of the unmodified cytosine is determined by the presence of (T, T) in the insert sequence and the copy of the insert sequence, respectively, and the modified cytosine and the unmodified cytosine are paired with G' in the complementary chain.

62. The method according to claim 59, wherein the method distinguishes the position of methylated cytosine from that of hydroxymethylated cytosine.

63. To provide each chain with conditions for modifying modified cytosine and / or unmodified cytosine, a. Reacting each chain with β-glycosyltransferase, b. Reacting each strand with DNA methyltransferase (DNMT), c. The method according to claim 62, comprising reacting each chain under conditions that convert unmodified cytosine to uracil.

64. a. The position of the methylated cytosine is determined by the presence of (C,C) in the insert sequence and the copy of the insert sequence, respectively. b. The position of hydroxymethylated cytosine is determined by the presence of (C, T) in the insert sequence and the copy of the insert sequence, respectively. c. The position of unmodified cytosine is determined by the presence of (T,T) in the insert sequence and the copy of the insert sequence, respectively. The method according to claim 63, wherein the methylated cytosine, the hydroxymethylated cytosine, and the unmodified cytosine are paired with G' in the complementary chain.

65. The conditions for modifying modified cytosine and / or unmodified cytosine are to provide each chain as described above. a. Reacting each chain with DNMT, b. Methylated cytosine in each chain is replaced with dihydroxyuracil ( DH The method according to claim 62, comprising reacting under conditions that convert to U.

66. a. The position of the methylated cytosine is determined by the presence of (T,T) in the insert sequence and the copy of the insert sequence, respectively. b. The position of hydroxymethylated cytosine is determined by the presence of (T,C) in the insert sequence and the copy of the insert sequence, respectively. c. The position of unmodified cytosine is determined by the presence of (C,C) in the insert sequence and the copy of the insert sequence, respectively. The method according to claim 65, wherein the methylated cytosine, the hydroxymethylated cytosine, and the unmodified cytosine are paired with G' in the complementary chain.