Dislocations that maintain continuity

By transposing nucleic acids into transposomes with barcode sequences and immobilizing them on solid supports, the method maintains continuity and phasing information, addressing the challenges of sequencing library preparation and enhancing the accuracy of methylation and genomic variation detection.

JP7860177B2Active Publication Date: 2026-05-15ILLUMINA CAMBRIDGE LTD
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ILLUMINA CAMBRIDGE LTD
Filing Date
2024-07-31
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing methods face challenges in efficiently maintaining the continuity and phasing information of nucleic acid sequences, particularly in the presence of harsh conditions like bisulfite treatment, which disrupts the continuity and phasing information, making it difficult to determine the methylation state and genomic variations accurately.

Method used

The method involves transposing target nucleic acids into multiple transposomes, each containing a transposon and transposase, which hybridize with complementary capture sequences, fragmenting the nucleic acids while maintaining continuity, and attaching barcode sequences to each fragment, allowing for immobilization on solid supports for sequencing and analysis.

Benefits of technology

This approach preserves the continuity and phasing information of nucleic acids, enabling accurate determination of methylation states and genomic variations without the need for additional purification steps, thereby improving sequencing library preparation and data accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007860177000007
    Figure 0007860177000007
  • Figure 0007860177000008
    Figure 0007860177000008
  • Figure 0007860177000009
    Figure 0007860177000009
Patent Text Reader

Abstract

To provide a method for preparing a library of barcoded DNA fragments of a target nucleic acid.SOLUTION: Provided is a method for preparing a library of barcoded DNA fragments of a target nucleic acid comprising: (a) contacting a target nucleic acid with a plurality of transposome complexes; (b) fragmenting the target nucleic acid into a plurality of fragments and inserting a plurality of transfer strands at the 5' end of at least one strand of the fragments; (c) contacting a plurality of fragments of the target nucleic acid with a plurality of solid supports, each of the plurality of solid supports comprising a plurality of immobilized oligonucleotides, each oligonucleotide comprising a complementary capture sequence and a first barcode sequence; and (d) transferring information from the barcode sequence to a fragment of the target nucleic acid, thereby creating a library of double-stranded fragments.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 065,544, filed Oct. 17, 2014, and U.S. Provisional Patent Application No. 62 / 157,396, filed May 5, 2015, which are hereby incorporated by reference in their entirety.

[0002] Embodiments of the invention relate to nucleic acid sequencing. Specifically, embodiments of the methods and compositions provided herein relate to the preparation of nucleic acid templates and the acquisition of sequence data from nucleic acid templates.

Background Art

[0003] The detection of specific nucleic acid sequences present in a biological sample has been used, for example, as a method for the identification and classification of microorganisms, the diagnosis of infectious diseases, the detection and characterization of genetic abnormalities, the identification of gene changes associated with cancer, the study of genetic susceptibility to disease, and the measurement of responses to various types of treatment. A common technique for detecting specific nucleic acid sequences in a biological sample is nucleic acid sequencing.

[0004] Nucleic acid sequencing methods have evolved significantly from the chemical degradation method used by Maxam and Gilbert and the chain extension method used by Sanger. Today, several sequencing methods are used that enable the parallel processing of all nucleic acids in a single sequencing run. Thus, the information generated from a single sequencing run can be enormous.

Summary of the Invention

Means for Solving the Problems

[0005] ​​​​​​​​​​​​ In one embodiment, this specification describes the lyricity of barcoded DNA fragments of target nucleic acids. This document describes a method for preparing a brally. This method involves transposing a target nucleic acid into multiple transposables. This involves contacting the transpososome complex, and each transpososome complex is trans It contains a transposon and a transposase, and the transposon includes a transfer chain and a non-transfer chain. At least one transposon of the nsposome complex hybridizes with a complementary capture sequence. Includes an adapter sequence that can be soyed. Flags the target nucleic acid into multiple fragments. Mentating the fragments while maintaining the contiguity of the target nucleic acid. Insert multiple transfer strands into the 5' end of at least one strand. Multiple fragments of the target nucleic acid. The device is brought into contact with multiple solid supports. Each of the multiple solid supports is connected to multiple immobilized oligonucleotides. It contains a rheotide, and each oligonucleotide has a complementary capture sequence and a first barcode sequence. Including, the first barcode sequence from each solid support in a plurality of solid supports The first barcode sequence is different from that of other solid supports within the body. Barcode sequence information It is transferred to a fragment of the target nucleic acid, thereby creating at least two fragments of the same target nucleic acid. So that the ragments receive the same barcode information, at least one chain is the first bar We will create a library of immobilized double-stranded fragments tagged with 5' code sequences.

[0006] In one embodiment, this specification relates to a method for determining the continuity information of a target nucleic acid sequence. The method includes contacting a target nucleic acid with multiple transposomal complexes. Each transpososome complex contains a transposon and a transposase. Furthermore, a transposon includes a transfer chain and a non-transfer chain, and the transposome complex At least one of the zones is an adapter capable of hybridizing into a complementary capture sequence. - Includes sequences. Fragmenting the target nucleic acid into multiple fragments and maintaining the continuity of the target nucleic acid. Multiple transfer chains are inserted into multiple fragments while maintaining their position. The nucleotide is brought into contact with multiple solid supports. Each of the multiple solid supports contains multiple immobilized oligonucleotides. It contains a creotide, and each oligonucleotide has a complementary capture sequence and a first barcode sequence. The first barcode sequence from each solid support in a plurality of solid supports is included in the plurality of solid supports. The first barcode sequence is different from that of other solid supports within the support. Barcode sequence information This means that at least two fragments of the same target nucleic acid receive the same barcode information. The sea urchin is transferred to the target nucleic acid fragment. The sequence and barcode of the target nucleic acid fragment are also provided. The sequence is determined. The continuity information of the target nucleic acid is determined by identifying the barcode sequence. In some embodiments, the transposase of the transpososome complex is rearranged. After transposition, remove the transposon's adapter sequence. Hybridize to a complementary capture sequence. In some embodiments, the transposer The transposase is removed by SDS treatment. In some embodiments, the transposase is treated Removed by protein-degrading enzyme treatment.

[0007] In one embodiment, this specification provides information on the phasing and methylation status of a target nucleic acid sequence. This document describes a method for simultaneously measuring [the target nucleic acid]. This method involves transposing the target nucleic acid into multiple transposables. This involves contacting the transposomal complex, and each transposomal complex is transpo comprising a zone and a transposase, the transposase comprising a transfer strand and a non-transfer strand, at least one of the transposons of the transpososome complex comprising an adapter sequence capable of hybridizing to a complementary capture sequence. Fragmenting a target nucleic acid into a plurality of fragments, and inserting a plurality of transfer strands into the target nucleic acid fragments while maintaining the continuity of the target nucleic acid. Contacting a plurality of fragments of the target nucleic acid with a plurality of solid supports, each of the plurality of solid supports comprising a plurality of immobilized oligonucleotides, each oligonucleotide comprising a complementary capture sequence and a first barcode sequence, the first barcode sequences from each of the plurality of solid supports in the plurality of solid supports being different from the first barcode sequences from other solid supports in the plurality of solid supports. Transferring barcode sequence information to the target nucleic acid fragments such that at least two fragments of the same target nucleic acid receive the same barcode information. Subjecting the target nucleic acid fragments comprising a barcode to bisulfite treatment, thereby generating bisulfite-treated target nucleic acid fragments comprising a barcode. Determining the sequences of the bisulfite-treated target nucleic acid fragments and the barcode sequences. Determining the continuity information of the target nucleic acid by identifying the barcode sequences. In one aspect, the present disclosure describes a method of preparing an immobilized library of tagged DNA fragments. The method includes providing a plurality of solid supports having transpososome complexes immobilized thereon, the transpososome complexes being multimeric, the transposon monomer units of the same transpososome complex being bound to each other, the transposon monomer units of the same transpososome complex being bound to each other, and the transposon monomer units of the same transpososome complex being bound to each other.

[0008] In one embodiment, the present disclosure describes a method of preparing an immobilized library of tagged DNA fragments. The method includes providing a plurality of solid supports having transpososome complexes immobilized thereon, the transpososome complexes being multimeric, the transposon monomer units of the same transpososome complex being bound to each other, and the transposon monomer units of the same transpososome complex being bound to each other. are bound to each other, and the transposon monomer units of the same transpososome complex are bound to each other. The nsposome monomer unit contains a transposase bound to the first polynucleotide. The first polynucleotide consists of (i) a 3' portion containing the transposon terminal sequence, and (ii) ) Includes a first adapter containing a first barcode. The target DNA is transposed to the target DNA. Fragmented by the sporosome complex, the 3' transpo of the first polynucleotide Under the condition that the zoon terminal sequence is transferred to the 5' end of at least one strand of the fragment, multiple It is applied to a solid support. Thereafter, at least one chain is 5' in the first barcode. We will create an immobilized library of ligated double-stranded fragments.

[0009] In one embodiment, this specification provides a sequence for determining the methylation state of a target nucleic acid. This document describes a method for preparing a sampling library. This method involves selecting two target nucleic acids or This includes fragmenting into further fragments. First common adapter array The adapter sequence is incorporated into the 5' end of the target nucleic acid fragment, and the first primer It includes a binding sequence and an affinity moiety, the affinity moiety being present in one member of the binding pair. The target nucleic acid fragment is denatured. The target nucleic acid fragment is immobilized on a solid support. The solid support contains the other members of the binding pair, and the immobilization of the target nucleic acid is necessary for the binding of the binding pair. Further steps are taken. The immobilized target nucleic acid fragment is subjected to bisulfite treatment. Second common adapter - The sequence is incorporated into an immobilized target nucleic acid fragment treated with bisulfite, and a second common adapter is formed. The pterosulfite treatment immobilized on a solid support is provided. The immobilized target nucleic acid fragment is amplified, thereby determining the methylation state of the target nucleic acid. Create a sequencing library for this purpose.

[0010] In one embodiment, this specification provides a sequence for determining the methylation state of a target nucleic acid. This document describes a method for preparing a sampling library. This method involves immobilizing the library. The invention includes providing a plurality of solid supports containing immobilized transposome complexes. The sposome complex contains a transposon and a transposase, and the transposon is It includes a transfer chain and a non-transfer chain. The transfer chain has (i) a 3' end containing a transposase recognition sequence. (ii) the first part of and (ii) the first adapter array and the first member of the bonding pair 'Includes the second part located in the first part. The first member of the bond pair is a solid support. It binds to the second member of the binding pair on the body, thereby allowing the transposon to attach to a solid support. To immobilize. The first adapter also contains the first primer-binding sequence. The non-transfer chain is (i) the first portion of the 5' end containing the transposase recognition sequence and (ii) the end of the 3' end The 3' to 1st portion contains the second adapter sequence in which the terminal nucleotides are blocked. It includes a second portion located therein. The second adapter also includes a second primer binding sequence. The target nucleic acid is brought into contact with multiple solid supports containing an immobilized transposome complex. The target nucleic acid is fragmented into multiple fragments, and multiple transfer chains are formed into fewer fragments. Both are inserted into the 5' end of a single strand, thereby fixing the target nucleic acid fragment to a solid support. Optimize. The 3' end of the fragmented target nucleic acid is extended with DNA polymerase. The non-transfer strand is ligated to the 3' end of the fragmented target nucleic acid. Target nucleic acid fragments are subjected to bisulfite treatment. Fixed fragments damaged during bisulfite treatment are then treated. The 3' end of the immobilized target nucleic acid fragment is homopo The DNA polymerase is used to extend the rimtail to include it. Second adapter The sequence is directed to the 3' end of the immobilized target nucleic acid fragment that was damaged during bisulfite treatment. Insert. A bisulfite-treated target nucleic acid fragment immobilized on a solid support is subjected to the first and The second primer is used for amplification, thereby determining the methylation state of the target nucleic acid. Create a sequencing library.

[0011] In one embodiment, this specification provides a sequence for determining the methylation state of a target nucleic acid. This document describes a method for preparing a transcoding library. This method involves transcoding target nucleic acids. This involves contacting the sporosome complex, and the transposome complex is trans It contains transposons and transposases. Transposons include a transfer chain and a non-transfer chain. The transfer chain consists of (i) the first portion at the 3' end containing the transposase recognition sequence and (ii) Located in the 5' to first portion including the first adapter array and the first member of the bonding pair The second part includes the first member of the bond pair, and the first member of the bond pair is bonded to the second member of the bond pair. The non-transfer chain consists of (i) the first portion at the 5' end containing the transposase recognition sequence and (i i) 3' end containing a second adapter sequence with a blocked terminal nucleotide at the 3' end The second adapter includes a second part located in the first part, and the second primer bond Includes combined sequences. Fragmentation of the target nucleic acid into multiple fragments and multiple transfer strands. Insert into the 5' end of at least one strand of the fragment, thereby targeting the nucleic acid fragment. The target nucleic acid fragment, including the transposon end, is immobilized on a solid support. It is brought into contact with multiple solid supports, including the second member, and bonded with the first member of the bonding pair. The target nucleic acid is immobilized on a solid support by binding with the second member of the pair. The 3' end of the translocated target nucleic acid is extended with DNA polymerase. The non-transfer strand is then flattened. Ligate to the 3' end of the immobilized target nucleic acid fragment. The immobilized target nucleic acid fragments are subjected to bisulfite treatment. The 3' end of the immobilized target nucleic acid fragment contains a homopolymer tail. The sea urchin is extended using DNA polymerase. The second adapter sequence is connected to bisulfite. It is introduced into the 3' end of an immobilized target nucleic acid fragment that has been damaged during processing. Immobilized bisulfite-treated target nucleic acid fragments are used with first and second primers. It amplifies the signal, thereby determining the methylation state of the target nucleic acid using sequencing. Prepare an ibrary.

[0012] In some embodiments, the terminal nucleotide at the 3' end of the second adapter is Selected from the group consisting of deoxynucleotides, phosphate groups, thiophosphate groups, and azide groups. Blocked by one member.

[0013] In some embodiments, the affinity portion can be a member of the binding pair. In some cases, the modified nucleic acid contains the first member of the binding pair. Also, the capture probe may include the second member of the binding pair. In this case, the capture probe may be immobilized on a solid surface, and the modified nucleic acid will bind to the solid surface. It may include the first member of A, and the capture probe is the second member of the binding pair It may include. In such cases, the bond between the first and second members of the bond pair is This further immobilizes modified target nucleic acids onto solid surfaces. Examples of binding pairs are limited to this. Although not, biotin-avidin, biotin-streptavidin, biotin-neutral Raavidin, ligand-receptor, hormone-receptor, lectin-glycoprotein, oligonucleotide Examples include creotide-complementary oligonucleotides and antigen-antibodies.

[0014] In some embodiments, the first common adapter array is one-side d) Incorporate into the 5' terminal fragment of the target nucleic acid by rearrangement. In some embodiments, Then, the first common adapter sequence is ligated to the 5' end fragment of the target nucleic acid. It is incorporated into the to. In some embodiments, the second common adapter array is bisulfite The step of incorporating the processed target nucleic acid fragment into the immobilized target nucleic acid fragment is (i) immobilized target nucleic acid fragment The 3' end of the terminal is elongated using terminal transferase, forming a homopolymer tail. (ii) an oligonucleotide containing a single-chain homopolymer portion, and A step of hybridizing a double-stranded portion containing a common adapter sequence of 2, The single-chain homopolymer portion is complementary to the homopolymer tail, and (iii) ) Ligate the second common adapter sequence to the immobilized target nucleic acid fragment, thereby The second common adapter sequence is incorporated into the bisulfite-treated, immobilized target nucleic acid fragment. Includes steps.

[0015] In some embodiments, the target nucleic acid is derived from a single cell. In this state, the target nucleic acid originates from a single organelle. In some embodiments, The target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is other nucleic acids. It is crosslinked. In some embodiments, the target nucleic acid is formalin-fixed and paraffin-embedded. (FFPE:formarin fixed paraffin embedded) It originates from a sample. In some embodiments, the target nucleic acid crosslinks with a protein. In some embodiments, the target nucleic acid crosslinks with DNA. In some embodiments, the target nucleic acid is histone-protected DNA. Histones are removed from the target nucleic acid. In some embodiments, the target nucleic acid is an acellular tumor. This is tumor DNA. In some embodiments, cell-free tumor DNA is obtained from placental fluid. In some embodiments, cell-free tumor DNA is obtained from plasma. In this process, plasma is collected from whole blood using a membrane separator equipped with a plasma collection zone. In some embodiments, the plasma collection zone is a transporter immobilized on a solid support. It includes a some complex. In some embodiments, the target nucleic acid is cDNA. In some embodiments, the solid support is a bead. In some embodiments, Multiple solid supports are multiple beads, and the multiple beads are of various sizes.

[0016] In some embodiments, a single barcode sequence is placed on multiple individual solid supports. It is present in the immobilized oligonucleotide. In some embodiments, different barcodes The sequence is present in multiple immobilized oligonucleotides on each individual solid support. In one embodiment, the transfer of barcode sequence information to the target nucleic acid fragment is performed by Lige By means of. In some embodiments, barcode distribution to target nucleic acid fragments The transfer of column information is mediated by polymerase elongation. In some embodiments, the target nucleic acid f The transfer of barcode sequence information to the ligation is performed by both ligation and polymerase elongation. By means of. In some embodiments, polymerase elongation is performed on ligated immobilized oligonucleotides. Using a nucleotide as a template, the 3' end of the non-ligated transposon chain is converted to DNA. This is done by elongation with rimelase. In some embodiments, a small number of adapter sequences At least some of them include a second barcode sequence.

[0017] In some embodiments, the transposome complex is a multimer, and each monomer is a single monomer The adapter sequence of a transposon at a given position is a unit of other monomeric units of the same transposomal complex. This is different. In some embodiments, the adapter array is a first primer binding array Further including columns. In some embodiments, the first primer binding site is a capture sequence Alternatively, it does not have sequence homology with respect to the complementary of the capture sequence. In some embodiments, The immobilized oligonucleotide on the solid support further comprises a second primer-binding sequence.

[0018] In some embodiments, the transposome complex is a multimer, and Somal monomer units bind to each other within the same transposome complex. Several implementations Morphologically, the transposase of the transposomal monomer unit is the same transposase It binds to the transposase of another transposomal monomer unit of the complex. In this embodiment, the transposons of the transposomal monomer unit are the same transposons It binds to a transposon, another transposomal monomer unit of the complex. In one embodiment, the transposase of the transposomal monomer unit is the same transpo Covalently bound to the transposase of another transposomal monomer unit of the chromosome complex. Combine. In some embodiments, the transposase of one monomer unit is the same Disulfide transposase of another transposomal monomer unit of the transposome complex They are bound by id bonds. In some embodiments, transposomal monomer units A transposon is a transposon monomer unit of another transposomal complex. It binds to lansposons by covalent bonds.

[0019] In some embodiments, the continuity information of the target nucleic acid sequence indicates haplotype information. In some embodiments, the continuity information of the target nucleic acid sequence indicates genomic variation. In several embodiments, genomic mutations include deletions and translocations. The group consists of interchromosomal gene fusions, duplications, and paralogs, which are selected from several real In the application form, the oligonucleotide immobilized on the solid support has a partially double-stranded region and It includes a partial single-stranded region of the oligonucleotide. In some embodiments, a partial oligonucleotide The single-stranded region includes a second barcode sequence and a second primer-binding sequence. In the embodiment, a target nucleic acid fragment including a barcode is used to obtain the target nucleic acid fragment Amplify before sequencing. In some embodiments, subsequent amplification is performed on the target nucleic acid. Before determining the arrangement of the fragments, the reaction is carried out in a single reaction compartment. In some embodiments, A third barcode sequence is introduced into the target nucleic acid fragment during amplification.

[0020] In some embodiments, the above method involves a target nucleic acid fragment including a barcode. From multiple first sets of reaction compartments, a pool of target nucleic acid fragments containing barcodes The step of combining the target nucleic acid fragments, including barcodes, into a pool of second The steps of redistributing the set into the reaction compartment, and the target nucleic acid fragment into the second set By amplifying the third barcode within the reaction compartment before sequencing, the target nucleic acid The procedure may further include a step of introducing the lagment.

[0021] In some embodiments, the above method involves contacting the target nucleic acid with the transposomal complex. The process may further include a step of pre-fragmenting before proceeding. Morphologically, pre-fragmentation of target nucleic acids is a group process consisting of sonication and restriction digestion. It is done by the method selected from. [Brief explanation of the drawing]

[0022] [Figure 1] This flowchart shows an example of a method for attaching transposomes to the surface of beads. [Figure 2] This diagram illustrates the steps of the method shown in Figure 1. [Figure 3] This is a schematic diagram illustrating an example of the tagmentation process on the surface of a bead. [Figure 4] This data table shows an example of DNA yield from the perspective of the number of clusters from the bead-based tagging process shown in Figure 3. [Figure 5]Figure 3 shows another example of the reproducibility of the bead-based tagging process, specifically from the perspective of uniform size. [Figure 6] Figures 6A and 6B show plots of the insertion sizes for pool 1 and pool 2 of the indexed samples in Figure 5, respectively. [Figure 7] This bar graph shows the reproducibility of the total number of reads and the percentage of reads aligned for the experiment described in Figure 5. [Figure 8] Figures 8A, 8B, and 8C show the insertion size plots in the control library, the insertion size plots in the bead-based tagmented library, and the summary data table, respectively, for the exome enrichment assay. [Figure 9] Figures 9A, 9B, and 9C show bar graphs for the dups PF fraction, the selected bases fraction, and PCT usable bases on target in the exome enrichment assay, respectively. [Figure 10] This flowchart shows an example of a method for forming transposome complexes on the surface of beads. [Figure 11] This diagram illustrates the steps of the method shown in Figure 10. [Figure 12] This diagram illustrates the steps of the method shown in Figure 10. [Figure 13] This diagram illustrates the steps of the method shown in Figure 10. [Figure 14] Figure 13 is a schematic diagram of the tagging process using transposome-coated beads. [Figure 15] This figure shows an exemplary scheme for transposome formation on a solid support. [Figure 16] This figure shows an example scheme for creating a contiguously linked library with unique indexes. [Figure 17] This figure shows an example scheme for creating a linked library with unique indexes. [Figure 18]This diagram shows the capture of a single CPT-DNA on a single clone-indexed bead, with the CPT-DNA wrapped around the bead. [Figure 19] This diagram shows the capture of a single CPT-DNA on a single clone-indexed bead, with the CPT-DNA wrapped around the bead. [Figure 20] This figure shows an exemplary scheme for binding a Y-adapter immobilized on a solid surface to target DNA by ligation and gap filling. [Figure 21] This figure shows an exemplary scheme for fabricating the Y-adapter during ligation between CPT-DNA and immobilized oligonucleotides on a solid support. [Figure 22] This figure shows agarose gel electrophoresis demonstrating the removal of free transposomes from a linked library by size exclusion chromatography. [Figure 23] This figure shows an exemplary scheme for generating a shotgun sequence library of specific DNA fragments. [Figure 24] This figure shows an exemplary scheme for assembling sequence information from a sequenced library with clone indexes. [Figure 25] This figure shows the results of optimizing the capture probe density on the beads. [Figure 26] This figure shows the results of a feasibility study on preparing indexed sequencing libraries of CPT-DNA on beads using intramolecular hybridization. [Figure 27] This figure shows the results of a test to determine the feasibility of clone indexing. [Figure 28] This graph shows the frequency of sequencing reads relative to specific distances within and between neighboring reads aligned to the tagged template nucleic acid. [Figure 29] Figures 29A and 29B illustrate exemplary approaches for extracting continuity information on a solid support. [Figure 30] This figure shows a schematic diagram and the rearrangement results of indexed clonal bead rearrangement in a single reaction vessel (one-pot). [Figure 31] This figure shows a schematic diagram and the rearrangement results of indexed clonal bead rearrangement in a single reaction vessel (one-pot). [Figure 32] This is a schematic diagram showing the construction of clonal transposomes on beads using 5' or 3' biotinylated oligonucleotides. [Figure 33] This figure shows the library size relative to transposomes on beads. [Figure 34] This figure shows the effect of transposome surface density on insertion size. [Figure 35] This figure shows the effect of input DNA on the size distribution. [Figure 36] This figure shows the size and distribution of islands using bead-based and solution-based tagmentation reactions. [Figure 37] This figure shows the clonal indexing of several individual DNA molecules, each receiving its own unique index. [Figure 38] This is a schematic diagram of a device for separating plasma from whole blood. [Figure 39] This is a schematic diagram showing a device for separating plasma and the subsequent use of the separated plasma. [Figure 40] This is a schematic diagram showing a device for separating plasma and the subsequent use of the separated plasma. [Figure 41] This figure shows an exemplary scheme of targeted phasing by enriching specific regions of the genome. [Figure 42] This figure shows an exemplary scheme for exome phasing using inter-exon SNPs. [Figure 43] This figure shows an exemplary scheme for the simultaneous detection of phasing and methylation. [Figure 44] This figure shows another exemplary scheme for the simultaneous detection of phasing and methylation. [Figure 45] This figure shows an exemplary scheme for generating libraries of various sizes using clone-indexed beads of various sizes in a single assay. [Figure 46] This figure shows an exemplary scheme for determining genetic variation in libraries of different length scales. [Figure 47A] This figure shows the detection results for a 60kb heterozygous deletion in chromosome 1. [Figure 47B] This figure shows the detection results for a 60kb heterozygous deletion in chromosome 1. [Figure 48] This figure shows the results of gene fusion detection using the method of the present invention. [Figure 49] This figure shows the results of gene deletion detection using the method of the present invention. [Figure 50] This figure shows the ME sequence before and after bisulfite conversion. [Figure 51] This figure shows the results of optimizing the bisulfite conversion efficiency. [Figure 52] This figure shows the results after bisulfite conversion as an IVC plot (intensity per base versus cycle). [Figure 53] This figure shows images of agarose gel electrophoresis of an indexed conjugated library after PCR following BSC. [Figure 54] This figure shows the bioanalyzer trace of a bound CPT-seq library with a whole genome index before enrichment and without size selection. [Figure 55] This figure shows the agarose gel analysis of the concentrated library. [Figure 56] This figure shows the results of applying targeted haplotyping to the HLA region of a chromosome. [Figure 57] This diagram shows several possible mechanisms for ME swapping. [Figure 58] This diagram shows several possible mechanisms for ME swapping. [Figure 59]This figure shows a portion of a Tn5 transposase having exemplary amino acid residues that can be substituted with Cys: Asp468, Tyr407, Asp461, Lys459, Ser458, Gly462, Ala466, Met470. [Figure 60] This figure shows a portion of a Tn5 transposase having amino acid substitutions S458C, K459C, and A466C that enable the cysteine ​​residue to form a disulfide bond between two monomeric units. [Figure 61] This figure shows an exemplary scheme for the preparation and use of dimer transposase (dTnp) nanoparticle (NP) bioconjugates (dTnp-NP) using amine-coated nanoparticles. [Figure 62] This figure shows an exemplary scheme of conjugation between a transposome dimer and an amine-coated solid support. [Figure 63] This figure shows the Mu transposome complex with the transposon terminus bound to it. [Figure 64] This figure shows a schematic diagram of indexed binding reads for pseudogene assembly / phasing, and the advantages of pseudogene mutation identification using shorter fragments. [Figure 65] This figure shows plots of index exchanges from four separate experiments, expressed as the percentage of swapped indices. [Figure 66] This figure shows the analysis of fragment size in Ts-Tn5 titration using Agilent BioAnalyzer. [Figure 67] This figure shows an exemplary scheme for improving DNA yield in the Epi-CPTSeq protocol using an enzymatic method to recover damaged library elements after bisulfite treatment. [Figure 68] Figures 68A–68C show several exemplary schemes for improving DNA yield in the Epi-CPTSeq protocol using an enzymatic method to recover damaged library elements after bisulfite treatment. [Figure 69]This figure shows an exemplary scheme for mold rescue using random primer extension. [Figure 70] This figure shows the fragmentation of a DNA library during sodium bisulfite conversion. The left panel shows the fragmentation of a portion of tagged DNA on magnetic beads during sodium bisulfite conversion. The right panel shows the bioanalyzer traces of CPT-seq and Epi-CPT-seq (Me-CPT-seq) libraries. [Figure 71] This figure shows an exemplary scheme and results of a TdT-mediated ssDNA ligation reaction. [Figure 72] This figure shows the scheme and results of TdT-mediated recovery of sodium bisulfite conversion beads bound to a library. The left panel shows the operational flow for rescuing a damaged bisulfite conversion DNA library using a TdT-mediated ligation reaction. The results of the DNA library rescue experiment are shown in the right panel. [Figure 73] This figure shows the results of the methyl-CPT-seq assay. [Figure 74] This figure shows an illustrative scheme for bead-based bisulfite conversion of DNA. [Figure 75] Figures 75A and 75B show the results of optimizing the bisulfite conversion efficiency. [Modes for carrying out the invention]

[0023] Embodiments of the present invention relate to nucleic acid sequencing. Specifically, provided herein Embodiments of the method and composition include the preparation of nucleic acid templates and the acquisition of sequence data from nucleic acid templates. Regarding.

[0024] In one embodiment, the present invention relates to tagged (fragmented and tagged) targets To construct nucleic acid libraries, a method for tagging target nucleic acids on a solid support is needed. Regarding this, in one embodiment, the solid support is a bead. Therefore, the target nucleic acid is DNA.

[0025] In one embodiment, the present invention relates to a solid capable of extracting continuity information of a target nucleic acid. The present invention relates to a method, a method, and a composition based on a support, a transposase, and several embodiments. In this state, the composition and method can extract assembly / fading information. .

[0026] In one embodiment, the present invention involves capturing linked transposition target nucleic acids on a solid support. This relates to a method and composition for extracting continuity information.

[0027] In one embodiment, the methods and compositions disclosed herein relate to the analysis of genomic mutations. Examples of genomic mutations include, but are not limited to, deletions, interchromosomal transpositions, and duplications. Examples include paralogs and interchromosomal gene fusions. In some embodiments, as described herein The disclosed methods and compositions relate to the determination of phasing information of genomic mutations.

[0028] In one embodiment, the methods and compositions disclosed herein relate to a specific region of a target nucleic acid. Regarding phasing. In one embodiment, the target nucleic acid is DNA. In the application form, the target nucleic acid is genomic DNA. In some embodiments, the target Nucleic acids are RNA. In some embodiments, RNA is mRNA. In several embodiments, the target nucleic acid is complementary DNA (cDNA). (ary DNA). In some embodiments, the target nucleic acid is derived from a single cell. In some embodiments, the target nucleic acid is derived from circulating tumor cells. In some embodiments, the target nucleic acid is cell-free DNA. The target nucleic acid is cell-free tumor DNA. In some embodiments, the target nucleic acid is Forma Derived from a phosphate-fixed paraffin-embedded tissue sample. In some embodiments, the target nucleus The acid is the target nucleic acid for crosslinking. In some embodiments, the target nucleic acid is crosslinked to a protein. To bridge. In some embodiments, the target nucleic acid is cross-linked to the nucleic acid. In this embodiment, the target nucleic acid is histone-protected DNA. Then, the histone-protected DNA is precipitated from the cell lysate using an antibody against histones. It is suppressed and histones are removed.

[0029] In some aspects, an indexed library is a clone indexed library. It is prepared from target nucleic acids using beads. In some embodiments, tagmentation is performed. The target nucleic acid, with the transposase still bound to the target DNA, has a clonal index. It can be captured using beads. In some embodiments, specific capture points A lobe is used to capture a specific region of the target nucleic acid. The region is washed with various stringencies, selectively amplified, and then sequenced. It can be done. In some embodiments, the capture probe can be biotinylated. Good. The biotinylation capture probe hybridizes to a specific region of the indexed target nucleic acid. The fused complex can be separated using streptavidin beads. An exemplary scheme for aging is shown in Figure 41.

[0030] In some embodiments, the compositions and methods disclosed herein are for exome phase It can be used for zing. In some embodiments, the exons, promoters It can be enriched. Markers, for example, heterojunction SNPs between exon regions, in particular This may be useful for exon fading when the distance between exons is large. Typical exome phasing is shown in Figure 42. In some embodiments, index The binding read with the cap extends simultaneously to the heterojunction SNP of the adjacent exon (cover). It is not possible to phase two or more exons. It is difficult. The compositions and methods disclosed herein also address interexon heterojunction SNPs. This process involves concentrating the nucleotides, for example, phasing exon 1 into SNP1 and SNP2 into exon 2. Therefore, through the use of SNP1, exon 1 and exon 2 are converted as shown in Figure 42. It can be phased.

[0031] In one embodiment, the compositions and methods disclosed herein relate to phasing and simultaneous It can be used to detect methylation. (BSC: bisulfite) Methylation detection by conversion is difficult because the BSC reaction is harsh on DNA. It is a harsh process that fragments DNA, thereby disrupting continuity / phasing Removing the information is difficult. Furthermore, the method disclosed in this application is a conventional BSC approach. In contrast to what is required in the previous method, no further purification steps are needed, resulting in a higher yield. There are further advantages to improving it.

[0032] In one embodiment, the compositions and methods disclosed herein are used to produce different sized ra Ibrary can be prepared in a single assay. In some embodiments, different Using clone indexed beads of a certain size, libraries of different sizes can be created. It can be prepared. Figure 1 shows a method 100 for binding transposase to the surface of beads. An example flowchart is shown. Transposases are transposons oligonucleotides. Using a transposase and any other chemical that may be added to the solid phase, the bead surface They may be bound together. In one example, transpososomes are bound together with biotin-streptoavian The material is bonded to the bead surface via a din-bonding complex. Method 100 is not limited to this, however This includes the following steps:

[0033] In one embodiment, the transposon has a sequencing primer binding site. It may include, but is not limited to, exemplary sequences of sequence-binding sites. However, AATGATACGGCGACCACCGAGATCTACAC (P5 sequence) and C AAGCAGAAGACGGCATACGAGAT (P7 sequence) is one example. In this embodiment, the transposon may be biotinylated.

[0034] In step 110 of Figure 1, P5 and P7 biotinylated transposons are generated. Transposons also contain one or more index sequences (unique identifiers). It may be present. Exemplary index sequences include, but are not limited to, TAGATC. GC, CTCTCTAT, TATCCTCT, AGAGTAGA, GTAAGGAG, A Examples include CTGCATA, AAGGAGTA, and CTAAGCCT. Another example is P5. Biotinylation is performed on only transposons or only P7 transposons. In yet another example, Transposons consist only of mosaic end (ME) sequences, or ME sequences. In addition to the columns, it includes additional sequences that are not P5 and P7 sequences. In this example, P5 and P7 sequences This is added in the subsequent PCR amplification step.

[0035] In step 115 of Figure 1, the transposome is assembled. Transposomes are a mixture of P5 and P7 transposomes. The mixture of smposomes will be described in more detail in relation to Figures 11 and 12.

[0036] In step 120 of Figure 1, the P5 / P7 transposome mixture is placed on the bead surface. To bind. In this example, the beads are streptavidin-coated beads, and the transporter The soms are attached to the bead surface via a biotin-streptavidin binding complex. The beads may be of various sizes. In one example, the beads are 2.8 μm beads. It may also be 1 μm beads. In another example, it may be 1 μm beads. Suspension of 1 μm beads The liquid (e.g., 1 μL) provides a large surface area per unit volume for transposome binding. The surface area that can be used for transposomal binding allows for tagging per reaction. The number of finished products increases.

[0037] Figure 2 illustrates steps 110, 115, and 120 of Method 100 in Figure 1. Now, let's show the transposon as a double strand. In another example (not shown), we have a different structure such as a hairpin. A single oligonucleotide having a self-complementary region capable of forming a double helix. You may also use Ochido.

[0038] In step 110 of method 100, multiple biotinylated P5 transposons 210a and generates multiple P7 transposons 210b. P5 transposons 210a and P Biotinylation of transposon 210b.

[0039] In step 115 of method 100, the P5 transposon 210a and the P7 transposon Poson 210b was mixed with transposase Tn5 215, and multiple assembled transposases were used. It forms sphosome 220.

[0040] In step 120 of method 100, the transposome 220 is attached to the bead 225. Combine. Bead 225 is a streptavidin-coated bead. Transposome 2 20 is bound to bead 225 via a biotin-streptavidin binding complex.

[0041] In one embodiment, a mixture of transposomes is shown in Figures 10, 11, 12, and 1 As shown in 3, it may be formed on a solid support such as the surface of a bead. In this example, P5 and The P7 oligonucleotides are initially placed on the bead surface before the assembly of the transposome complex. To bond to a surface.

[0042] Figure 3 shows a schematic diagram of an example of the tagmentation process 300 on the bead surface. Process 30 In Figure 2, the bead 225 to which transposome 220 is bound is shown. DNA31 Add solution 0 to the suspension of beads 225. DNA 310 contacts transposome 220. When touched, the DNA is tagged (fragmented and tagged) and becomes a transposome. The bound and tagged DNA 310 is attached to bead 225 via 220. CR amplification may be performed to generate a pool of amplification product 315 in solution (without beads). The amplified product 315 may be transferred to the surface of the flow cell 320. Cluster generation process Tocol (for example, can be used for bridge amplification protocols or cluster generation) Using any other amplification protocol, multiple clusters 325 are connected to flow cell 320. It may be generated on the surface. Cluster 325 is a clone of tagged DNA 310. This is the amplification product. Now cluster 325 proceeds to the next step of the sequencing protocol. This means we're ready for the event.

[0043] In another embodiment, the transposome is an arbitrary solid such as the wall of a microcentrifuge tube. It may adhere to the body surface.

[0044] In another embodiment, in which a mixture of transposome complexes is formed on the surface of the beads, The ligonenucleotides are first bound to the bead surface before the transposome assembly. Figure 10 shows an example of a method 1000 for forming a transposome complex on a bead surface. A flowchart is shown. Method 1000 includes, but is not limited to, the following steps. .

[0045] In step 1010, P5 and P7 oligonucleotides are bound to the surface of the beads. In one example, P5 and P7 oligonucleotides are biotinylated, and the beads are These are streptavidin-coated beads. This step is also shown in schematic Figure 1100. This is illustrated in the diagram. Referring to Figure 11, P5 oligonucleotide 1110 and P7 oligonucleotide Ligonucleotide 1115 binds to the surface of bead 1120. In this example, 1 Two P5 oligonucleotides 1110 and one P7 oligonucleotide 1115 are used. Although bound to the surface of the P5 oligonucleotide 1120, any number of P5 oligonucleotides 1110 and Alternatively, P7 oligonucleotide 1115 may bind to the surface of multiple beads 1120. In one example, P5 oligonucleotide 1110 is a P5 primer sequence, DEX sequence (unique identifier), read 1 sequencing primer sequence, and mosaic It contains a C-terminal (ME) sequence. In this example, P7 oligonucleotide 1115 is P7 Primer sequence, index sequence (unique identifier), read 2 sequencing ply Includes the MAR sequence and the ME sequence. In another example (not shown), the index sequence is P It is present only in oligonucleotide 1110. In yet another example (not shown), The INDEX sequence is present only in P7 oligonucleotide 1115. Here's another example (Figure) (Not shown) The index sequences are P5 oligonucleotide 1110 and P7 oligonucleotide It is not present in any of the nucleotide 1115.

[0046] In step 1015, the complementary mosaic end (ME') oligonucleotide is used. Hybridize the P5 and P7 oligonucleotides bound to the oligonucleotides. This step also This is illustrated in schematic diagram 1200 of Figure 12. Referring to Figure 12, the complementary ME sequence ( ME')1125 contains P5 oligonucleotide 1110 and P7 oligonucleotide 11 Hybridizes to 15. Complementary ME sequence (ME') 1125 (for example, complementary ME The sequence (ME')1125a and complementary ME sequence (ME')1125b) are P5 oligonucleotides. Hybrids were added to the ME sequences of rheotide 1110 and P7 oligonucleotide 1115, respectively. The complementary ME sequence (ME')1125 is typically about 15 nucleotides long, and 5 It is phosphorylated at the terminal end.

[0047] In step 1020, the transposase enzyme is converted to a bead-bound oligonucleotide. Add to form a mixture of bead-bound transposome complexes. This step also A schematic diagram is shown in Figure 13, diagram 1300. Referring to Figure 13, the transposase enzyme The substance is added to form multiple transposome complexes 1310. In this example, Lansposome complex 1310 contains a transposase enzyme and two surface-bound oligonucleotides. Includes an Otid sequence and a complementary ME sequence (ME')1125 hybridized thereto. It has a double-stranded structure. For example, the transposome complex 1310a has a complementary ME sequence (M P5 oligonucleotide 1110 hybridized to E')1125 and complementary ME compound P7 oligonucleotide 1115 hybridized to column (ME')1125 (i.e., P 5:P7) is included, and the transpososome complex 1310b contains the complementary ME sequence (ME')1 Two P5 oligonucleotides 1110 hybridized to 125 (i.e., P5:P5) ) containing, the transpososome complex 1310c has complementary ME sequence (ME')1125 Contains two hybridized P7 oligonucleotides 1115 (i.e., P7:P7). The ratios of 5:P5, P7:P7, and P5:P7 transposome complexes are, for example, 25 :25:50 is also acceptable.

[0048] Figure 14 shows the tagging process using the transposome-coated beads 1120 shown in Figure 13. An exemplary schematic diagram 1400 is shown. In this example, the transposome complex 1310 is Add the beads 1120 to the DNA 1410 solution in the tagging buffer, and tag This induces mentation, and the DNA is transferred to the surface of the bead 1120 via the transposome 1310. They are then bound together. By sequential tagging of DNA1410, transposome 131 Multiple bridge molecules 1415 are formed between 0. The length of the bridge molecule 1415 is equal to the length of the bead 1 This may depend on the density of the transposome complex 1310 on the surface of 120. In one example, the density of the transposome complex 1310 on the surface of the bead 1120 is In step 1010 of method 100 in Figure 10, P5 is bonded to the surface of bead 1120. This may also be adjusted by changing the amount of P7 oligonucleotides. In another example, The density of the transposome complex 1310 on the surface of bead 1120 is determined by the method shown in Figure 10. In step 1015 of 1000, hybridized P5 and P7 oligonucleotides. The amount of complementary ME sequence (ME') that is used may be adjusted by changing it. In another example, the density of the transposome complex 1310 on the surface of bead 1120 is In step 1020 of Method 1000 in Figure 1, the amount of tramposase enzyme added is varied. You may adjust it by doing so.

[0049] The length of bridge molecule 1415 is the length of the transposomal complex used in the tagmentation reaction. The amount of combined 1310 is independent of the amount of bead 1120 to which it is bound. Similarly, in the tagmentation reaction Adding more or less DNA1410 in the final tagging production It does not change the size of the material, but it may affect the yield of the reaction.

[0050] In one example, bead 1120 is a paramagnetic bead. In this example, tagment The purification of the chemical reaction can be easily performed by fixing the beads 1120 with a magnet and washing them. Therefore, tagging and subsequent PCR amplification can be performed in a single reaction compartment ("One-Point"). It may also be carried out using a "one-pot" reaction.

[0051] In one embodiment, the present invention provides for extracting continuity information of a target nucleic acid on a solid support. The present invention relates to a transposase-based method, methods, and compositions that enable several implementations. In this state, the composition and method can extract assembly / phase information. In one embodiment, the solid support is a bead. In one embodiment, the target nucleus The acid is DNA. In one embodiment, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is RNA. In some embodiments, R NA is mRNA. In some embodiments, the target nucleic acid is complementary DNA (c It is DNA.

[0052] In some embodiments, transposons are dimerized on a solid support such as beads. It is then immobilized, and subsequently, the transposase binds to the transposon to form a transposome. It is permissible.

[0053] In some embodiments, by adding a solid-phase transposon and a transposase, In particular, in relation to the formation of transpososomes in the solid phase, two transposons support the solid The components may be fixed in close proximity to each other (preferably at a certain distance) within the body. The approach has several advantages. Firstly, preferably, two transformers Two linker lengths and orientations are optimal for the poson to efficiently form transpososomes. The transposons will always be fixed at the same time. Secondly, the transposons The efficiency of transposon formation will not be a function of transposon density. Transposons are always in the right orientation and distance between them to form transpososomes. This can be used for the following. Thirdly, random immobilized transposons on the surface Therefore, various distances are formed between transposons, and as a result, only one fraction is transposon It has the optimal orientation and distance for efficiently forming sposomes. Instead of lanceposons being converted into transpososomes, solid-phase uncomplexed transpososomes This means that transposons exist. These transposons have a double-stranded DNA ME portion. Therefore, it is easily targeted for translocation. This leads to a decrease in translocation efficiency and undesirable by-products. It has the potential to form objects. Therefore, it can be used in conjunction with tagging and sequencing. Transpososomes are prepared on a solid support that can derive continuity information through a process. This may be done. An exemplary scheme is shown in Figure 15. In some embodiments, transform Posons may be immobilized on a solid support by methods other than chemical bonding. Exemplary methods for immobilizing nsposons include, but are not limited to, streptavi Din-biotin, maltose-maltose-binding protein, antigen-antibody, DNA-DNA Alternatively, affinity bonding such as DNA-RNA hybridization can be mentioned.

[0054] In some embodiments, the transposomes are supported by a solid after pre-assembly. It can be immobilized on the body. In some embodiments, the transposon is unique Includes index, barcode, and amplified primer binding site. Transposase, A transposon that can be added to a solution containing transposons and immobilized on a solid support. Somotic dimers can be formed. In one embodiment, each set is immobilized They share the same index derived from the nsposon, thereby producing indexed beads. As shown in Figure 29A, multiple bead sets can be generated. This can be added to each set of indexed beads.

[0055] In some embodiments, the target nucleic acid is added to each set of indexed beads. This can be done, and tagging and subsequent PCR amplification may be performed separately.

[0056] In some embodiments, the target nucleic acid, indexed beads, and transposable The mixture consists of many droplets, one bead and one or more DNA molecules and sufficient transistors. They can be combined within a droplet to include sposomes.

[0057] In some embodiments, indexed beads can be pooled, The target nucleic acid can be added to the reaction zone, and tagging and subsequent PCR amplification can be performed in a single reaction zone. You can also do it with a drawing ("One Pot").

[0058] In one embodiment, the present invention involves capturing a linked transposition target nucleic acid on a solid support. The present invention relates to a method and composition for extracting continuity information. In some embodiments, Contiguity-preserving transpos The process (CP) is performed on the DNA, but the DNA remains intact. T-DNA), and therefore forms a linked library. Continuity information is transmitted via transposase. The association of template nucleic acid fragments adjacent to the target nucleic acid is determined using this method. It can be preserved by maintaining the following: CPT-DNA can be stored on a solid support, for example. , complementary oligonucleotides having a unique index or barcode fixed to the beads It can be captured by rheotide hybridization (Figure 29B). In this embodiment, oligonucleotides immobilized on a solid support are added to the barcode. Furthermore, the primer binding site and the unique molecular index (UMI) are important. It may also include ular indices.

[0059] Advantageously, using transposomes in this way allows for the fragmentation of nucleic acids. By maintaining rational proximity, molecules of the same origin, for example, fragments from chromosomes... Nucleic acids, with the same unique barcode and index information immobilized on a solid support, The possibility of receiving from ligonenucleotides increases. This results in a unique barcode. This will result in a concatenated sequencing library. The rally can be sequenced to extract sequential sequence information.

[0060] Figures 16 and 17 show how to create a linked library with a unique barcode or index. A schematic diagram of an exemplary embodiment of the above-described aspect of the present invention is shown. The exemplary method is CPT - DNA and immobilized oligonucleotides on a solid support having a unique index and barcode Using ligation with creotide and strand substitution PCR, sequencing live Generates a rally. In one embodiment, clone indexed beads run It may be generated using immobilized DNA sequences such as dams or specific primers and indices. The resulting library involves hybridization to immobilized oligonucleotides and subsequent labs. Molecules can be captured on clone-indexed beads by eregulation. Intramolecular hybridization capture was much faster than intermolecular hybridization. The continuous dislocation library then "wraps around" the bead. Figures 18 and 19 show the capture of CPT-DNA on clone-indexed beads. This also shows the preservation of continuity information. Strand substitution PCR preserves clone bead index information individually. It can be transferred to the molecule. Therefore, each linked library is specifically indexed It will be attached.

[0061] In some embodiments, oligonucleotides immobilized on a solid support are, on the other hand, One chain is fixed to a solid support, and the other chain is partially complementary to the fixed chain. This allows for the inclusion of a partially double-stranded structure that acts as a Y-adapter. In one embodiment, the Y-adapter fixed to a solid support is ligation It also binds to ligated tagged DNA by gap filling, as shown in Figure 20.

[0062] In some embodiments, the Y-adapter is a probe on a solid support such as a bead. Formation through hybridization capture of CPT-DNA using an index. Figure 21 shows an exemplary scheme for fabricating such a Y-adapter. By using their Y-adapter, each fragment can potentially be sequenced This ensures that it has the potential to become a library. The scope of application (coverage) increases.

[0063] In some embodiments, free transpososomes are separated from CPT-DNA. Alternatively, in some embodiments, the separation of free transpososomes is performed by size exclusion cross-section. By matrix. In one embodiment, separation is performed using MicroSpin S-4 00 HR Columns (Pittsburgh, Pennsylvania, GE Healthcare) This may be done by (Re Life Sciences Inc.). Figure 22 shows free transporters. This shows the agarose gel electrophoresis of CPT-DNA isolated from the chromosomes.

[0064] For capturing linked transposition target nucleic acids onto a solid support via hybridization, It has several unique advantages. Firstly, the method is based on hybridization. Therefore, it is not based on rearrangement. Intramolecular hybridization rate >> Intermolecular hybridization rate This is the bridging rate. Therefore, a continuous transposition library of a single target DNA molecule The possibility of wrapping around a uniquely indexed bead is two or more different single targets. This is far more difficult than the process of DNA molecules wrapping around beads with unique indices. Therefore, DNA transposition and the barcoding of transposed DNA occur in two separate steps. The third is the assembly of activated transposomes on beads and on solid surfaces. This can avoid the challenges associated with optimizing the surface density of transposons. Fourth One way is to remove the autotransformation product by column purification. The fifth way is to... Because the transposition DNA contains gaps, the DNA is more flexible, and therefore, transposable Compared to methods that fix the chromosome to beads, this method places less stress on dislocation density (insertion size). 6 The goal is to use a combined barcode scheme as the method. It can be used. The seventh method involves covalently bonding indexed oligos to beads. It is easy to do. Therefore, the possibility of index replacement is low. The eighth point is tags Mentation and subsequent PCR amplification may be multiplexed, and a single reaction compartment ("one-pot") may be used. Since it can be done in a reaction, there is no need to perform individual reactions for each index sequence. Yes.

[0065] In some embodiments, during the transposition, multiple unique barcodes are transposed across the target nucleic acid. They may be inserted across. In some embodiments, each barcode is separated by a fragment. It includes a first barcode sequence and a second barcode sequence in which the punctuation portion is arranged. The first barcode sequence and the second barcode sequence are identified or designed to pair with each other. It is possible to create pairs of barcodes such that the first barcode and the second barcode are related. This can be informative. Conveniently, the paired Using barcode sequences, sequencing data is generated from a library of template nucleic acids. It is possible to do this. For example, a first template nucleic acid containing a first barcode sequence, and 1 Identifying a second template nucleic acid, which includes a second barcode sequence that pairs with the first barcode sequence. This means that the first and second template nucleic acids represent sequences adjacent to each other in the sequence representation of the target nucleic acid. This means that, using this method, target nucleic acids can be targeted without requiring a reference genome. The array representation can be (newly) assembled in a de novo environment.

[0066] In one aspect, the present invention provides a shotgun sequence of a specific DNA fragment. This invention relates to a method and composition for producing rallies.

[0067] In one embodiment, a cloned bead index is used to immobilize an oligonucleotide. Cleotide sequences: Generated using random or specific primers and unique indices. The target nucleic acid is added to the clone indexed beads. In some embodiments, The target nucleic acid is DNA. In one embodiment, the target DNA is denatured. NA is a primer containing a unique index immobilized on a solid surface (e.g., beads). - Hybridizes to one primer, and then hybridizes to another primer having the same index. The primers on the beads amplify the DNA. One or more further amplification Width rounding may be performed. In one embodiment, amplification is performed on a 3' random n-mer array. This may be performed by whole-genome amplification using bead-immobilized primers having [specific properties]. In one embodiment, the random n-mer prevents primer-primer interactions during amplification. Therefore, pseudo-complementary bases (2-thiothymine, 2-aminodA, N4-ethylcytosine, etc.) Includes (Hoshika, S; Chen, F; Leal, NA; Benner, SA, A ngew.Chem.Int.Ed.49(32)5554-5557(2010)). Figure 23 shows an exemplary method for generating a shotgun sequence library of specific DNA fragments. This scheme involves a sequence library with a clone index and amplification. A library of things can be generated. In one embodiment, such a live Rally can be generated by dislocation. Use index information as a guide. This allows the sequence information of a clone-indexed library to be used to determine continuity. It can be performed. Figure 24 shows the sequencer with clone index. This shows an exemplary scheme for assembling sequence information from a rally.

[0068] The method of the above embodiment has several advantages. Intramolecular amplification on beads is performed on beads It is much faster than inter-amplification. Therefore, the product on the beads has the same index. This allows for the creation of a shotgun library of specific DNA fragments. Random primers amplify the template at random locations, so the same index A shotgun library containing an index can be generated from a specific molecule. The array information can be assembled using the attached array. The advantage is that the reaction can be multiplexed in a single reaction (one-pot reaction), and many individual reactions This eliminates the need to use gel. Many indexed clone beads are prepared. can be done, and thus many different fragments can be uniquely labeled, and parental alleles can be discriminated for the same genomic region. By using a number of indexes, the probability that the paternal DNA copy and the maternal DNA copy receive the same index for the same genomic region is low. The method utilizes the fact that the intra reaction is much faster than the inter reaction, and the beads essentially create a substantial partition in a large physical compartment. In some embodiments of all the above aspects of the present invention, the method may be used for cfDNA in a cell free DNA (cfDNA) assay. In some embodiments, cfDNA is obtained from plasma, placental fluid. In one embodiment, plasma can be obtained from undiluted whole blood using a membrane based sedimentation assisted plasma separator (Liu et al. Anal Chem. 2013 Nov

[0069] 5;85(21):10463-70). In one embodiment, the plasma collection zone in the plasma separator may contain a solid support containing a transposome. The solid support containing the transposome may capture cfDNA from the isolated plasma when the plasma is separated from whole blood, and can perform concentration of cfDNA and / or tagging of DNA. In some embodiments, the tagging may further introduce unique barcodes, thereby enabling subsequent demultiplexing to be performed after sequencing of the library pool. In some embodiments, cfDNA is obtained from plasma, placental fluid.

[0070] In one embodiment, plasma can be obtained from undiluted whole blood using a membrane based sedimentation assisted plasma separator (Liu et al. Anal Chem. 2013 Nov 5;85(21):10463-70). In one embodiment, the plasma collection zone in the plasma separator may contain a solid support containing a transposome. The solid support containing the transposome may capture cfDNA from the isolated plasma when the plasma is separated from whole blood, and can perform concentration of cfDNA and / or tagging of DNA. In some embodiments, the tagging may further introduce unique barcodes, thereby enabling subsequent demultiplexing to be performed after sequencing of the library pool. In one embodiment, the plasma collection zone in the plasma separator may contain a solid support containing a transposome. The solid support containing the transposome may capture cfDNA from the isolated plasma when the plasma is separated from whole blood, and can perform concentration of cfDNA and / or tagging of DNA. In some embodiments, the tagging may further introduce unique barcodes, thereby enabling subsequent demultiplexing to be performed after sequencing of the library pool. In one embodiment, the plasma collection zone in the plasma separator may contain a solid support containing a transposome. The solid support containing the transposome may capture cfDNA from the isolated plasma when the plasma is separated from whole blood, and can perform concentration of cfDNA and / or tagging of DNA. In some embodiments, the tagging may further introduce unique barcodes, thereby enabling subsequent demultiplexing to be performed after sequencing of the library pool. In some embodiments, the tagging may further introduce unique barcodes, thereby enabling subsequent demultiplexing to be performed after sequencing of the library pool.

[0071] In some embodiments, the sampling zone of the separator is a PCR master mix (plastic It may contain an imager, nucleotide, buffer, metal, and polymerase. In this embodiment, the master mix is ​​reconstituted when the plasma comes out of the separator. As such, it may be in a dry state. In some embodiments, the primer is a random It is a primer. In some embodiments, the primer targets a specific gene. It may also be a specific primer. As a result of PCR amplification of cfDNA, the isolated This will involve generating a library directly from plasma.

[0072] In some embodiments, the sampling zone of the separator is the RT-PCR master mix Includes (primers, nucleotides, buffers, metals), reverse transcriptase, and polymerase. It may be so. In some embodiments, the primer is a random primer or This is an oligo-dT primer. In some embodiments, the primer is a specific gene It may also be a specific primer for the offspring. The obtained cDNA is then used for sequencing. It can be used. Alternatively, cDNA can be used as a solid support for sequence library preparation. The material may also be treated with immobilized transpososomes.

[0073] In some embodiments, the plasma separator uses a barcode (1D or 2D barcode) It may include. In some embodiments, the separation device has a blood collection device. This is also acceptable. This will send the blood directly to the plasma separator and library preparation device. In some embodiments, the apparatus may have a downstream sequence analyzer. In some embodiments, the sequence analyzer is a single-use sequencer. In this state, the sequencer creates a sequence of samples before sequencing them together. Alternatively, the sequencer can deliver the sample to the sequencing region. It may also have a random access function.

[0074] In some embodiments, the plasma collection zone is configured to concentrate cell-free DNA. It may also contain a silica substrate.

[0075] Simultaneous detection of phasing and methylation 5-methylcytosine (5-Me-C) and 5-hydroxymethylcytosine (5-Hydroxymethylcytosine (5-Hydroxymethylcytosine) Xy-C) is also known as epigenetic (Epi) modification and is involved in cell metabolism, differentiation, and It plays an important role in cancer growth. Surprisingly and unexpectedly, the inventors of this invention have discovered Furthermore, detection of phasing and simultaneous methylation is possible using the method and composition of this application. We have discovered something. The present invention relates to CPT sequencing on beads (CPT-s Combining eq (indexed linked library) with DNA methylation detection This makes it possible to use individual libraries generated on beads as sulfurous acid. Treatment with a hydrogen salt converts unmethylated C to U, but methylated C is not converted to U. This makes it possible to detect 5-Me-C. Further detection using heterojunction SNPs Through phasing analysis, epi-methylation-phasing blocks were identified as multiple megabases. It can be established in the domain.

[0076] In some embodiments, the size of the DNA being analyzed is approximately 100 base pairs to several millimeters. It is possible up to G bases. In some embodiments, the size of the DNA to be analyzed is about 100 bases, 200 bases, 300 bases, 400 bases, 500 bases, 600 bases, 700 base s, 800 bases, 900 bases, 1000 bases, 1200 bases, 1300 bases, 1500 base s, 2000 bases, 3000 bases, 3500 bases, 4000 bases, 4500 bases, 500 0 bases, 5500 bases, 6000 bases, 6500 bases, 7000 bases, 7500 bases, 8 000 bases, 8500 bases, 9000 bases, 9500 bases, 10,000 bases, 10,5 00 bases, 11,000 bases, 11,500 bases, 12,000 bases, 12,500 bases , 13,000 bases, 14,000 bases, 14,500 bases, 15,000 bases, 15, 500 bases, 16,000 bases, 16,500 bases, 17,000 bases, 17,500 base s, 18,000 bases, 18,500 bases, 19,000 bases, 19,500 bases, 20 ,000 bases, 20,500 bases, 21,000 bases, 21,500 bases, 22,000 bases, 22,500 bases, 23,000 bases, 23,500 bases, 24,000 bases, 2 4,500 bases, 25,000 bases, 25,500 bases, 26,000 bases, 26,50 0 bases, 27,000 bases, 27,500 bases, 28,000 bases, 28,500 bases, 29,500 bases, 30,000 bases, 30,500 bases, 31,000 bases, 31,5 00 bases, 32,000 bases, 33,000 bases, 34,000 bases, 35,000 bases , 36,000 bases, 37,000 bases, 38,000 bases, 39,000 bases, 40, 000 bases, 42,000 bases, 45,000 bases, 50,000 bases, 55,000 base s, 60,000 bases, 65,000 bases, 70,000 bases, 75,000 bases, 80 ,000 bases, 85,000 bases, 90,000 bases, 95,000 bases, 100,00 0 bases, 110,000 bases, 120,000 bases, 130,000 bases, 140,000 0 bases, 150,000 bases, 160,000 bases, 170,000 bases, 180,000 0 bases, 200,000 bases, 225,000 bases, 250,000 bases, 300,00 0 bases, 350,000 bases, 400,000 bases, 450,000 bases, 500,00 0 bases, 550,000 bases, 600,000 bases, 650,000 bases, 700,00 0 bases, 750,000 bases, 800,000 bases, 850,000 bases, 900,00 0 bases, 1,000,000 bases, 1,250,000 bases, 1,500,000 bases, 2,000,000 bases, 2,500,000 bases, 3,000,000 bases, 4,00 0,000 bases, 5,000,000 bases, 6,000,000 bases, 7,000,00 0 bases, 8,000,000 bases, 9,000,000 bases, 10,000,000 bases , 15,000,000 bases, 20,000,000 bases, 30,000,000 bases, 40,000,000 bases, 50,000,000 bases, 75,000,000 bases, 1 It is 00,000,000 base pairs or more.

[0077] 5-hydroxy-C, DNA oxidation product, DNA alkylation product, histone terminus (h Other Epi modifications, such as istone-foot printing, are also disclosed in this application. Analysis can also be performed during phasing using methods and compositions.

[0078] In some embodiments, the DNA is first indexed on a solid support. It is converted into a combined library. Each individual index is much smaller than the original DNA. Because the individual libraries are smaller, the libraries are less likely to become fragmented. Even if small fractions of a library with DEX are lost, the fading information is still indexed. It remains maintained throughout the entire length of the cusp-attached DNA molecule. For example, a 100kb molecule. In this case, conventional bisulfite conversion (BSC) fragments the material by half, and continuity is lost. It is now limited to 50kb. In the method disclosed herein, a 100kb library is Initially indexed, even if the individual library fractions disappear, the continuity remains. It is approximately 100kb (a rare occurrence where the entire library disappears from one end of the DNA molecule). (Except for). Furthermore, the methods disclosed herein do not require conventional bisulfite conversion approaches. In contrast to what is required, the yield is higher because no further purification steps are needed. Therefore, it has further advantages. In the method disclosed herein, the beads are bisulfite It is simply washed after the conversion. Furthermore, the DNA remains bound to the solid phase. (Indexed library) Minimal loss and easy buffer replacement with minimal effort. It can be done.

[0079] Figure 43 shows an exemplary scheme for the simultaneous detection of phasing and methylation. Operation Flow This involves DNA tagging on beads and gap-filling ligation of 9-base pair repeat regions. Removal of Tn5 by SDS, and modification of bisulfite in individual libraries on beads. It consists of substitution. To ensure that adjacent complementary libraries do not re-anneal, The bisulfite conversion is carried out under modified conditions, thereby reducing the efficiency of the bisulfite conversion. BSC converts unmethylated C to U, while methylated C remains unconverted.

[0080] Figure 44 shows another exemplary scheme for the simultaneous detection of phasing and methylation. Transition After preparing the sequencing library, a single-strand template is prepared using a gap. The fractions of the packed ligated library are decomposed. The single-strand template already has a template, This can reduce library loss or improve bisulfite conversion efficiency. Because it is a main chain, it requires milder conditions for bisulfite conversion. One implementation Morphologically, 3' thioprotected transposons (Exo-resistant) and unprotected transposons The mixture is used on the same beads. An enzyme, for example, Exo I, is used to remove the non-thioprotective layer. Libraries can be broken down and converted into single-stranded libraries. Thioprotection Lansposon: By using a mixture in which unprotected transposons are 50:50, Convert 50% of the library into a single-stranded library (50% is one of the libraries Transposons are protected, except for one transposon (complementary chain) which is not. ), 25% do not convert (both transposons are thioprotected), and 25% both Convert and remove the entire library (both transposons are not protected).

[0081] In performing bisulfite conversion of DNA bound to a solid phase such as streptavidin magnetic beads One challenge is that the DNA-bound beads need to be heated at high temperatures for a long time using sodium bisulfite. The process involves damaging both the DNA and the beads. To aid in recovery, the carrier DNA (i.e., lambda DNA) was reacted before the bisulfite treatment. Add to the mixture. Even if carrier DNA is present, approximately 80% of the original DNA will be lost. This was predicted. As a result, CPTSeq continuity block is a result of the conventional CPTSeq protocol. It has fewer members than Ru.

[0082] Therefore, in this specification, we improve the DNA yield of the Epi-CPTSeq protocol. We propose several strategies for this. The first strategy is streptavidin By making the transposome complex more densely packed on the beads, the library insertion size can be reduced. It relies on creating a library. By reducing the library size, a smaller proportion The library elements are decomposed by bisulfite treatment.

[0083] A second strategy for improving DNA yield in the Epi-CPTSeq protocol is The recovery strategy involves using enzymes to restore damaged library elements. The purpose of this is to remove the 3' common sequence necessary for library amplification during bisulfite treatment, specifically the 3' portion. This involves re-adding the bead-bonded library elements that have been decomposed and lost. 3' Co After adding the sequence, these elements can be amplified and sequenced using PCR. Yes, it is possible. Figures 67 and 68 show an exemplary scheme of this strategy. Double-stranded CPT Seq library elements are denatured and converted to bisulfite (upper panel). Bisulfite During the conversion, one of the DNA strands is damaged (middle panel), and the PCR common sequence on the 3' end is lost. The template rescue strategy recovers the 3' common sequence (green) necessary for PCR amplification. (bottom row). In one example, the 3' phosphorylated atenuator oligo, i.e., seek In the presence of a sequence containing an encoding adapter followed by an oligo dT stretch, terminology Use naltransferase (Figure 68A). Simply put, TdT is 10-15 The stretch of dA is annealed to the oligo dT portion of the attenuator oligo to break down the It is added to the 3' end of the Braley element. This DNA hybrid formation results in Td The T reaction is terminated, and the 3' end of the damaged library element is joined by DNA polymerase. It provides a template for effective extension.

[0084] In another operational flow (Figure 68B), the TdT tailing reaction is performed on the single-stranded oligo dT portion and Partial double-stranded atenucle having a 5' phosphorylated double-stranded sequencing adapter portion This reaction is carried out in the presence of a ter-oligonucleotide. At the end of the TdT reaction, the last added dA and 5' phosphate The nick between the chemical atenuator oligo and the nucleotide is sealed with DNA ligase.

[0085] Both of the described operational flows rely on the controllable TdT tailing reaction developed in recent years. This is based on and described in U.S. Patent Application Publication No. 2015 / 0087027. Common sequencing adapters also allow for the recent introduction of ssDNA casting for MMLV RTs. Due to its type-switching activity, it can be attached to the 3' end of a damaged library element. It can be done. In other words, MMLV RT and template switch oligo (TS oligo: template The switch_oligo is added to the damaged DNA (Figure 68C). In this step, reverse transcriptase adds several trailing segments to the 3' end of a single-stranded DNA fragment. The additional nucleotides are added, and these bases are located at the 3' end of one of the TS oligos. It pairs with the Rigo(N) sequence. Then, due to the template switching activity of reverse transcriptase, the an The sequence of the modified common primers is added to the 3' end of the BSC damaged library element. This restores the ability to be amplified by PCR using common sequencing primers.

[0086] As part of the third strategy, Epicentre's EpiGenome kit Using the "post-bisulfite conversion" library construction method, during bisulfite conversion, the 3' end Library elements that have lost their common array at the edges can be recovered. (See Figure 69) Thus, this library rescue method randomizes a common sequence followed by a short stretch. It utilizes 3' phosphorylated oligonucleotides with a specific sequence. These short, random sequences are bisulfite. The salt-treated single-stranded DNA is hybridized, and then the common sequence is processed by DNA polymerase. It is duplicated in the corrupted library chain.

[0087] Figure 74 shows a fourth strategy for improving the bisulfite sequencing method on beads. —This indicates that the first common sequence, including the capture tag, is covalently bonded to the 5' end of the DNA. The common arrangement is unilateral transposition (as shown), adapter ligation, or U.S. Patent Application Publication Terminal transferase (T) described in Specification No. 2015 / 0087027 dT) Adapter ligation, including various other methods, can be used to bind to DNA. can.

[0088] Next, the DNA is denatured (for example, by incubation at high temperature) and then bound to a solid support. To do so, if biotin is used as a capture tag on CS1, for example, DNA is stored. It can be bonded using leptavidin magnetic beads (shown in the figure) on a solid support. Once combined, buffer swapping can be easily performed.

[0089] In the next step, we will perform bisulfite conversion of ssDNA. In single-stranded form, DN A should be readily available for bisulfite conversion, and Promega's Methy Using an improved version of the Edge BSC kit, conversion efficiency was observed up to 95% (Figure 75).

[0090] After bisulfite conversion, the second common sequence is attached to the 3' end of the ssDNA bound to the solid support. To covalently bond the oligo to ssDNA, several methods for covalently bonding the oligo to ssDNA are described above. We have been using the TdT attenuator / adapter ligation method, achieving >95% Ligation efficiency was achieved. As a result, the proposed methyl sequencing (Met The yield of the final library using the hylSeq operation flow is higher than that of existing methods. It should be.

[0091] In the final step, PCR is performed to amplify the library, and the library is then supported by solids. Remove from the host. PCR primers require additional common components such as sequencing adapters. The columns can be designed to be appended to the ends of the MethylSeq library. ru.

[0092] Preparation of libraries of various sizes in a single assay The accuracy of genome assembly depends on the use of techniques with various length scales. For example, Shotgun (hundreds of bp) - Mate pair (~3Kb) to Hi-C (Mb scale) All of these are methods for improving assembly and contig lengths over time. The challenge is to accomplish this. Therefore, multiple assays are required, making the multilayer approach difficult to handle and costly. The composition and method disclosed herein allows for multiple lengths in a single assay. It can handle scaling.

[0093] In some embodiments, library preparation is performed using solid supports of different sizes, for example. This can be achieved in a single assay using beads. The size of each bead is... The physical size of the library determines the library size, and a specific library size or size range This will create a barrier. All beads of various sizes will be transferred to the library. It has a clone index. Therefore, libraries of various sizes have their own unique index. Generated with different library scale lengths, each with a specific length scale. By preparing libraries simultaneously in the same physical partition, costs are reduced and the entire operation flow is streamlined. To improve. In some embodiments, each specific solid support size, for example, bead size The index receives a unique index. In some other embodiments, the same solid support It is also possible to prepare multiple different indices of the same bead size, for example, multiple indices of the same bead size. Therefore, multiple DNA molecules can be indexed and categorized according to their size range. (Figure) 45 is a single assay using various sizes of clone indexed beads. This shows an example scheme for generating a library of a certain size.

[0094] In some embodiments, the size of the generated library is approximately 50 bases, 75 base, 100 bases, 150 bases, 200 bases, 250 bases, 300 bases, 350 bases, 4 00 bases, 500 bases, 600 bases, 700 bases, 800 bases, 900 bases, 1000 salt Base, 1200 bases, 1300 bases, 1500 bases, 2000 bases, 3000 bases, 350 0 bases, 4000 bases, 4500 bases, 5000 bases, 5500 bases, 6000 bases, 6 500 bases, 7000 bases, 7,500 bases, 8000 bases, 8500 bases, 9000 salts Base, 9500 bases, 10,000 bases, 10,500 bases, 11,000 bases, 11,5 00 bases, 12,000 bases, 12,500 bases, 13,000 bases, 14,000 bases , 14,500 bases, 15,000 bases, 15,500 bases, 16,000 bases, 16, 500 bases, 17,000 bases, 17,500 bases, 18,000 bases, 18,500 salts Base, 19,000 bases, 19,500 bases, 20,000 bases, 20,500 bases, 21 ,000 bases, 21,500 bases, 22,000 bases, 22,500 bases, 23,000 bases, 23,500 bases, 24,000 bases, 24,500 bases, 25,000 bases, 2 5,500 bases, 26,000 bases, 26,500 bases, 27,000 bases, 27,50 0 bases, 28,000 bases, 28,500 bases, 29,500 bases, 30,000 bases, 30,500 bases, 31,000 bases, 31,500 bases, 32,000 bases, 33,0 00 bases, 34,000 bases, 35,000 bases, 36,000 bases, 37,000 bases , 38,000 bases, 39,000 bases, 40,000 bases, 42,000 bases, 45, 000 bases, 50,000 bases, 55,000 bases, 60,000 bases, 65,000 salts Base, 70,000 bases, 75,000 bases, 80,000 bases, 85,000 bases, 90 ,000 bases, 95,000 bases, 100,000 bases, 110,000 bases, 120, 000 bases, 130,000 bases, 140,000 bases, 150,000 bases, 160, 000 bases, 170,000 bases, 180,000 bases, 200,000 bases, 225, 000 bases, 250,000 bases, 300,000 bases, 350,000 bases, 400, 000 bases, 450,000 bases, 500,000 bases, 550,000 bases, 600, 000 bases, 650,000 bases, 700,000 bases, 750,000 bases, 800, 000 bases, 850,000 bases, 900,000 bases, 1,000,000 bases, 1, 250,000 bases, 1,500,000 bases, 2,000,000 bases, 2,500, 000 bases, 3,000,000 bases, 4,000,000 bases, 5,000,000 salts Base, 6,000,000 bases, 7,000,000 bases, 8,000,000 bases, 9, 000,000 bases, 10,000,000 bases, 15,000,000 bases, 20,0 00,000 bases, 30,000,000 bases, 40,000,000 bases, 50,00 0,000 bases, 75,000,000 bases, 100,000,000 bases, or more It is above.

[0095] In some embodiments, the libraries of multiple length scales described above are combined into one large Instead of having a large length scale, it is used for assembling pseudogenes, paralogs, etc. This is possible. In some embodiments, a library of multiple length scales can be used as a single library. Prepare simultaneously in a single process. The advantage is that at least one length scale is a pseudo-gene or gene. This involves coupling to a specific region that has only a gene and not both. Therefore, this length scale Mutations detected by this tool are those that uniquely determine whether the mutation is in a gene or a pseudo-gene. This is possible. The same applies to copy number variations, paralogs, etc. The advantage of assembly is This involves the use of different length scales. By using the method disclosed herein, different An indexed join library for different length scales, It can be generated in a single assay without the need for separate library preparations. Figure 4 Figure 6 shows an exemplary scheme for determining gene mutations in libraries of different length scales. vinegar.

[0096] Analysis of gene mutations The compositions and methods disclosed herein relate to the analysis of gene mutations. Exemplary gene mutations Other variations include, but are not limited to, deletions, interchromosomal translocations, duplications, paralogs, and interchromosomal defects. Genetic fusion is one example. In some embodiments, the compositions and methods disclosed herein The law concerns the determination of phasing information for gene mutations. The following table shows exemplary interchromosomal mutations. This shows genetic fusion.

[0097] [Table 1]

[0098] Table 2 shows exemplary deletions in chromosome 1.

[0099] [Table 2]

[0100] In some embodiments, the target nucleic acid is subjected to a process before exposing the target nucleic acid to transposomes. Fragmentation is possible. Examples of fragmentation methods include, but are not limited to, those described above. However, ultrasonic treatment, mechanical shearing, and restricted digestion are also mentioned. Fragmentation of target nucleic acids prior to tagging and tagging can lead to the creation of pseudogenes (e.g., CYP2D6 ) is advantageous for assembly / phasing. Long islands of indexed bonded leads (> 30kb) would extend to pseudo-genes A and A', as shown in Figure 64. High sequence phase Because they are of the same sex, determining which mutations belong to gene A and gene A' is a challenge. Shorter mutations will bind to one mutation of the pseudogene in a unique surrounding sequence. Such shorter islands fragment the target nucleic acid before tagging. It can be achieved more effectively.

[0101] Binding transposomes In some embodiments, the transposase is present in large quantities in the transposomal complex. It is a body that, for example, forms dimers, tetramers, etc. within the transposome complex. Surprisingly and unexpectedly, the researchers found that monomers in the multimeric transposome complex To bind the transposase, or the transposase in the multimeric transpososome complex There are several advantages to ligating the transposon ends of somatic monomers. I found that, firstly, the binding of transposases or transposons is more stable. This results in a complex, and a large fraction is in an activated state. Secondly, at lower concentrations The potential use of transposomes in fragmentation by rearrangement reactions. There is. Thirdly, through binding, the mosaic ends (MEs) of the transposomal complex This reduces the amount of exchange, which in turn reduces the mixing of barcodes or adapter molecules. Such ME terminal replacement occurs when the complex breaks apart and reorganizes, or when the transistor Sposomes are immobilized on a solid support by streptavidin / biotin, and streptavidin If the toavidin / biotin interaction is broken and reorganized, or if contamination occurs... It may occur if there is a possibility. The inventors of this application have found that under various reaction conditions, ME I noticed there was a significant swap or exchange at the end. In the embodiment, the exchange may reach up to 15%. The exchange is performed on a high salt concentration buff. This is pronounced with α and decreases with glutamate buffer. Figures 57 and 58 show ME exchange. Here are some possible mechanisms.

[0102] In some embodiments, the sub-units of transposases in the transpososome complex The knits can be joined together by shared and non-shared means. Several implementations In this state, the transposase monomer forms a transpososome complex (transposase monomer). It can be bonded (before the addition of the sposon). In some embodiments, The sposase monomer can be bound after transposome formation.

[0103] In some embodiments, a native amino acid residue is used at the polymer interface as cysteine ​​(Cys) The amino acids may be substituted to promote the formation of disulfide bonds. For example, Tn5 tran In sposase, Asp468, Tyr407, Asp461, Lys459, S By substituting er458, Gly462, Ala466, and Met470 with Cys, monomers are formed. The disulfide bonds between the b units may be promoted. This is shown in Figures 59 and 60. Mos -1 Regarding transposases, exemplary amino acids that can be substituted with cysteine ​​are However, this does not apply to Leu21, Leu32, Ala35, His20, P he17, Phe36, Ile16, Thr13, Arg12, Gln10, Glu9 Examples are shown in Figure 61. In some embodiments, cysteine-substituted ami Modified transposases having no acid residues react with maleimide or pyridyldithiol reactive groups. The chemical crosslinking agent used can be used to chemically crosslink the materials together. Example of chemical crosslinking The agent is Pierce Protein Biology / ThermoFisher S It is commercially available from Scientific Corporation (Grand Island, New York, USA).

[0104] In some embodiments, the transposomal multimer complex is covalently bonded to a solid support. They can be combined. Examples of solid supports include, but are not limited to, nanoparticles. Examples include beads, flow cell surfaces, and column matrices. In some embodiments... Furthermore, the solid surface may be coated with amine groups. (Cysteine-substituted amino acids) Modified transposases having residues of amine sulfide are used to target such amine groups. Lyl crosslinking agent (i.e., succinimidyl-4-(N-maleimidomethyl)cyclohexane- Chemical crosslinking can be achieved using 1-carboxylate (SMCC). (Example) A typical scheme is shown in Figure 62. In some embodiments, maleimide-PEG-bio A tin crosslinking agent may be used to bind dTnp to a streptavidin-coated solid surface.

[0105] In some embodiments, the transposase gene is expressed in large quantities as a single polypeptide. It can be modified to express somatic proteins, such as Tn5 or Mos-1. The gene is a single polypeptide that expresses two Tn5 or Mos-1 proteins. It can be modified in the same way. Similarly, the Mu transposase gene is a single polypeptide It can be modified to encode four mu transposase units.

[0106] In some embodiments, the transposon ends of transposomal monomer units are connected. They can be combined to form a conjugated transpososome multimer complex. By joining the ends, the primer site can be inserted, and sequencing A primer, amplified primer, or any role DNA can filter the target DNA. It can act on gDNA without lagging. Such functional insertion is important. Is it necessary to extract the information from intact molecules or is subsampling important? This is an advantage in pneumatic assays or binding tag assays. The transposon end of the Mu transposome is a "loop-shaped" Mu transposer It can be bound to the ze / transposon structure. Since Mu is a tetramer, it can be bound to it. Not limited to, but including the combination of R2UJ and / or R1UJ with R2J and / or R1J. This allows for various structures. In these structures, R2UJ and R1UJ are R 2J and R1J can be coupled or not coupled, respectively. Figure 63 shows the transformer. This shows a Mu transposome complex with the transponder terminus bound. In some embodiments, , the transposon end of Tn5 or the transposon end of Mos-1 transpososome They can be combined.

[0107] As used herein, the term "transposon" refers to a transposon in an in vitro rearrangement reaction. Nuclei required to form a complex with a transposase or integrase enzyme that functions This refers to double-stranded DNA that shows only the rheotide sequence ("transposon terminal sequence"). Lansposons are transposases or integrators that recognize and bind to transposons. Along with the enzyme, the "complex" or "synaptic complex" or "transposomal complex" This forms a "transposome composition". The complex undergoes an in vitro transposition reaction. It is possible to insert or transpose transposons into the target DNA that is incubated together. It is possible. A transposon is a "transposition transposon sequence," that is, a "transposition chain," and This shows two complementary sequences consisting of "non-transferable transposon sequences," or "non-transferable chains." Example For example, a hyperfunctional Tn5 transposase that is active in in vitro transposition reactions. (For example, EZ-Tn5(trademark) transposase, Madison, Wisconsin, USA) Forming a complex with EPICENTRE Biotechnologies The transposon exhibits the following "transposition transposon sequence": 5'AGATGTGTATAAGAGACAG3' and non-transfer chains exhibiting the following "non-transferable transposon sequences": 5'CTGTCT CTTATACACATCT3' Includes.

[0108] The 3' end of the transfer chain binds to or transfers to the target DNA in the in vitro transfer reaction. The non-transfer chain exhibits a transposon sequence complementary to the transposon terminal sequence. In in vitro transposition reactions, it does not bind to or transfer to the target DNA. Several implementations Morphologically, transposon sequences contain one or more of the following sequences: Good: barcode, adapter sequence, tag sequence, primer binding sequence, capture sequence, unique A unique molecular identifier (UMI) sequence.

[0109] As used herein, the term "adapter" refers to a barcode, primer-binding sequence, capture Capture sequence, sequence complementary to the capture sequence, unique molecular identification (UMI) sequence, affinity moiety, restriction site This refers to nucleic acid sequences that can contain this information.

[0110] As used herein, the term "continuity information" means two or more shared information. This refers to the spatial relationship between the DNA fragments shown above. The modes of information sharing include adjacency and compartmentalization. This can be related to metric spatial relationships. Information regarding these relationships is sequential. This allows for the hierarchical assembly or mapping of sequence reads derived from DNA fragments. To simplify. This continuity information improves the efficiency and accuracy of such assemblies or mappings. It is good. Because traditional shotgun sequencing is used in conjunction with traditional In assembly or mapping methods, each sequence read originates from two or more DNs. Regarding the spatial relationship of A fragments, the relative genomic origin or coordinate of individual sequence reads This is because it does not take that into consideration. Therefore, according to the embodiments described herein, the continuity Methods for capturing information include the short-distance continuity method, which determines adjacent spatial relationships, and the partition spatial relationships. This may be done using the medium-range continuity method or the long-range continuity method that determines the metric-spatial relationship. These methods improve the accuracy and quality of DNA sequence assembly or mapping. This method may be used in conjunction with any sequencing method, such as the sequencing method described above. stomach.

[0111] Continuity information is derived from two or more DNA fragments from which each sequence read originates. Regarding spatial relationships, this includes the relative genomic origin or coordinates of individual sequence reads. In this embodiment, the continuity information includes sequence information from non-overlapping sequence reads.

[0112] In some embodiments, the continuity information of the target nucleic acid sequence indicates haplotype information. In some embodiments, the continuity information of the target nucleic acid sequence indicates genomic variation.

[0113] As used herein, the term "maintenance of the continuity of the target nucleic acid" means the fragmentation of nucleic acids and In relation to this, the intention is to maintain the order of the nucleic acid sequences of fragments from the same target nucleic acid. It tastes good.

[0114] As used herein, the term “at least part” and / or its grammatical equivalent shall be “total amount” It can refer to any portion of the total amount. For example, "at least a portion" can refer to a small portion of the total amount. Approximately 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20% %, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70 This refers to %, 75%, 80%, 85%, 90%, 95%, 99%, 99.9%, or 100%. It is possible.

[0115] As used herein, the term "approximately" means ±10%.

[0116] As used herein, the term “sequencing read” and / or its grammatical equivalent This is a physical or chemical test performed to obtain a signal indicating the order of monomers in a polymer. It can refer to a repeating step process. The signal is single monomer resolution or lower The order of monomers can be shown with high resolution. In certain embodiments, the step is to show the nucleus This process is initiated against an acid target and performed to obtain a signal indicating the sequence of bases in a nucleic acid target. Yes, it is possible. The process can be carried out to typical completion, which is usually from the end of the process. The signal is such that it is not possible to distinguish the target base any further with a reasonable level of certainty. It is defined as "up to a certain point in time." If desired, for example, until the desired amount of sequence information is obtained, It can be completed more quickly. Sequencing reads are used for a single target nucleic acid molecule. or simultaneously target nucleic acid molecules having the same sequence, or target molecules having different sequences. This can be performed simultaneously on nucleic acid groups. In some embodiments, sequencing Greed is a signal that is triggered by one or more target nucleic acid molecules that initiate signal acquisition. The process ends when no further results can be obtained. For example, the sequencing lead is a solid-phase group. Starting with one or more target nucleic acid molecules present in the substance, one or more targets The process can be terminated once the target nucleic acid molecule has been removed from the substrate. Alternatively, sequencing can be used. The system stopped detecting the target nucleic acid that was present on the substrate when the sequencing run started. This can be terminated. An example of a sequencing method is that the whole The following is incorporated herein by reference: It is.

[0117] As used herein, the terms “sequencing indication” and / or their grammatical equivalents are defined as follows: This can refer to information indicating the order and type of monomer units in a polymer. For example, information This can indicate the order and type of nucleotides in a nucleic acid. The information can be used, for example, to describe, Any, including images, electronic media, a series of symbols, a series of numbers, a series of letters, a series of colors, etc. The information may be in various formats. The information may be at a single monomer resolution or lower. Exemplary polymers are nucleic acids such as DNA or RNA, which have nucleotide units. A series of "A", "T", "G", and "C" letters represent DNA at single nucleotide resolution. This is a well-known sequence representation for DNA that correlates with the actual sequence of the molecule. Other exemplary These polymers are proteins that have amino acid units, and polysaccharides that have sugar units.

[0118] solid support Throughout this specification, solid supports and solid surfaces are used interchangeably. In this embodiment, the solid support or its surface is the inner or outer surface of a tube or container, etc. It is planar. In some embodiments, the solid support is made up of microspheres or beads. Includes. In this specification, "microsphere," "beads," "particles," or their grammar. The equivalent refers to small dispersed particles. Suitable bead compositions are not limited to these. However, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic Polymers, paramagnetic materials, triasols, carbon graphite, Titanium dioxide, latex, or cross-linked dextran such as Sepharose, cellulose, na Examples include iron, crosslinked micelles, and teflon, and similarly, as solid supports, generally as specified herein. Any other materials can be used to explain. Bangs Laboratori es Inc. (Fischer's, Indiana) "Microsphere Detection Guide (Micros The "Phenomenon Detection Guide" is a helpful guide. Specific implementation In some cases, the microspheres are magnetic microspheres or beads. In the application form, the beads may be color-coded. For example, Luminex (textiles) MicroPlex® microspheres (Austin, Sussex) may also be used. .

[0119] The beads do not need to be spherical; irregular particles may be used. Alternatively, or in addition, The beads may be porous. The size of the beads is in nanometers, or about 10 nm, in diameter. From millimeters, or 1 mm, the beads range from approximately 0.2 microns to approximately 200 microns. Preferably, beads of about 0.5 to about 5 microns are particularly preferred, but in some embodiments... In some embodiments, smaller or larger beads may be used. The beads have diameters of approximately 0.1 μm, 0.2 μm, 0.3 μm, 0.4 μm, 0.5 μm, 0 .6μm, 0.7μm, 0.8μm, 0.9μm, 1μm, 1.5μm, 2μm, 2.5 μm, 2.8μm, 3μm, 3.5μm, 4μm, 4.5μm, 5μm, 5.5μm, 6 μm, 6.5μm, 7μm, 7.5μm, 8μm, 8.5μm, 9μm, 9.5μm, 1 0μm, 10.5μm, 15μm, 20μm, 25μm, 30μm, 35μm, 40μm , 45μm, 50μm, 55μm, 60μm, 65μm, 70μm, 75μm, 80μm Even if it is 85 μm, 90 μm, 95 μm, 100 μm, 150 μm, or 200 μm good.

[0120] transposome "Transpososomes" are formed by the integration of integrases or transposases (integrases) Nucleic acids including the recognition site of the enzyme (transposase) and the transposase recognition site Includes. In embodiments provided herein, the transposase catalyzes the rearrangement reaction. It can form a functional complex with a transposase recognition site that is capable of doing so. Transposases are involved in a process sometimes called "tagmentation," where they transpose the transposases. It may bind to the transposase recognition site and potentially insert the transposase recognition site into the target nucleic acid. In some such embedded events, one chain of the transposase recognition site, It may be transferred to the target nucleic acid. In one example, the transposome has two subunits. A dimeric transposase containing a transposon sequence, and two discontinuous transposon sequences. Another example In this context, the transposome contains a dimeric transposase comprising two subunits. It includes a transposase and a continuous transposon sequence.

[0121] Some embodiments describe a hyperfunctional Tn5 transposase and a Tn5-type transposase. -ase recognition site (Goryshin and Reznikoff, J.Biol.Che m., 273:7367 (1998), or MuA transposase, and R1 and Mu transposase recognition site containing the R2 terminal sequence (Mizuuchi, K., Cell ,35:785,1983;Savilahti,H,etal.,EMBOJ.,14 This may include the use of (4893,1995). (Enhanced Tn5 Transposer) Exemplary transposase recognition sites that form complexes with ze (e.g., EZ-Tn5(trademark) Transposase, Madison, Wisconsin, Epicentre Biotec Hnologies Inc. has identified the following 19-base transfer strand (sometimes "M" or "ME") and Non-transfer chains: 5′AGATGTGTATAAGAGACAG3′, 5′CTGT Includes CTCTTATACACATCT3′. The ME sequence is also optimized by those skilled in the art. And it may be used.

[0122] Transfers that can be used in conjunction with specific embodiments of the compositions and methods provided herein A further example of a positional system is Staphylococcus aureus. reus)Tn552(Colegio et al.,J.Bacteriol.,1 83:2384-8,2001;Kirby C et al.,Mol.Microb iol.,43:173-86,2002), Ty1(Devine & Boeke, Nu Cleic Acids Res., 22:3765-72, 1994, and International Publication No. 95 / 23875), Transposon Tn7 (Craig, NL, Science. 271:1512,1996;Craig,NL,Review in:Curr T op Microbiol Immunol.,204:27-48,1996), Tn / O and IS10 (Kleckner N, et al., Curr Top Micr obiol Immunol.,204:49-82,1996), Mariner Transport Zase (Lampe DJ, et al., EMBO J., 15:5470-9,1 996), Tc1(Plasterk RH,Curr.Topics Microb iol.Immunol.,204:125-43,1996), P element (Glo or,GB,Methods Mol.Biol.,260:97-114,2004 ), Tn3(Ichikawa & Ohtsubo,J Biol.Chem.265 :18829-32,1990), bacterial insertion sequence (Ohtsubo & Sekine, Cu rr.Top.Microbiol.Immunol.204:1-26,1996), Retrovirus (Brown, et al., Proc Natl Acad Sci) USA, 86:2525-9, 1989), and yeast retrotransposons (Boek e&Corces, Annu Rev Microbiol.43:403-34,19 89) is one example. Further examples include IS5, Tn10, Tn903, IS911, Oxalis (Sleeping Beauty), SPIN 、 hat, Piggyback (Pi ggyBac), Hermes, TcBuster, Ae Buster 1 (AeBuster1), Tol2, and modified transposase family Enzyme (Zhang et al., (2009) PLoS Genet. 5: e1000) 689.Epub 2009 Oct 16;Wilson C.et al(2007 (J. Microbiol. Methods 71:332-5) is one example.

[0123] Further integrases that can be used with the methods and compositions provided herein Examples include retroviral integrase and retroviral integrase Examples of integrase recognition sequences include those for HIV-1, HIV-2, and SIV. Examples include PFV-1 and RSV integrases.

[0124] barcode Generally, barcodes are used to identify one or more specific nucleic acids. It can contain one or more nucleotide sequences that can be artificial. It may be an array, or a naturally occurring array generated during transposition, for example, previously juxtaposed The same flanking genomic DNA sequence (g-code) at the end of the DNA fragment These may also be the case. In some embodiments, the barcode is not present in the target nucleic acid sequence. It is an artificial sequence that can be used to identify one or more target nucleic acid sequences.

[0125] The barcodes are at least approximately 1, 2, 3, 4, 5, 6, 7, 8, and 9. , 10 pieces, 11 pieces, 12 pieces, 13 pieces, 14 pieces, 15 pieces, 16 pieces, 17 pieces, 18 pieces, 19 pieces It may contain 20 or more consecutive nucleotides. Several implementations In this state, the barcodes are at least approximately 10, 20, 30, 40, 50, and 6 Contains 0, 70, 80, 90, 100, or more consecutive nucleotides. In some embodiments, at least one barcode in a group of nucleic acids including barcodes However, some parts are different. In some embodiments, at least about 10 of the barcodes %, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99 The percentages are different. In a further such embodiment, all the barcodes are different. In nucleic acid groups including barcodes, the diversity of different barcodes is random or It can be generated non-randomly.

[0126] In some embodiments, the transposon array comprises at least one barcode Includes several embodiments, such as transpososomes containing two discontinuous transposon sequences. In this, the first transposon sequence includes the first barcode, and the second transposon The sequence includes a second barcode. In some embodiments, the transposon sequence This includes a barcode containing a first barcode sequence and a second barcode sequence. In some implementation configurations, the first barcode sequence is paired with the second barcode sequence. They can be identified or designed to pair with each other. For example, they can be known to pair with each other. Using a reference table containing multiple first and second barcode sequences, a known first barcode sequence We can see that the column will be paired with a known second barcode sequence.

[0127] In another example, the first barcode sequence contains the same sequence as the second barcode sequence. It may be. In another example, the first barcode sequence is the inverse complementary of the second barcode sequence. It may include an array. In some embodiments, a first barcode array and a second The barcode sequences are different. The first and second barcode sequences are bicode (bi-c It may include ode.

[0128] In some embodiments of the compositions and methods described herein, the barcode is cast Used in the preparation of nucleic acid types. Naturally, with a vast number of available barcodes, each Template nucleic acid molecules may contain unique identifiers. Unique identification can be used for several purposes. For example, uniquely identified molecules can be used for various purposes. For example, haplotype sequencing, parental allele identification, metagenomic sequencing In sequencing and genome sample sequencing, samples with multiple chromosomes This technology can be applied to the identification of individual nucleic acid molecules in genomes, cells, cell types, cellular disease states, and species. It is possible.

[0129] Exemplary barcode sequences include, but are not limited to, TATAGCCT, ATA GAGGC, CCTATCCT, GGCTCTGA, AGGCGAAG, TAATCTT Examples include A, CAGGACGT, and GTACTGAC.

[0130] Primer site In some embodiments, the transposon sequence is a "sequencing adapter" "or "sequencing adapter site," in other words, the hybridized primer It may include a region containing one or more parts that can be moved. In the embodiment, the transposon sequence is useful for amplification and sequencing, etc. Both can include the first primer site. An example sequence of the sequence binding site is: This is not limited to, but includes AATGATAACGGCGACCACCGAGATCTACAC (P5 sequence) and CAAGCAGAAGACGGCATACGAGAT (P7 sequence) are mentioned. It can be done.

[0131] target nucleic acid Any target nucleic acid can be listed. For example, DN A, RNA, peptide nucleic acids, morpholino nucleic acids, locked nucleic acids, glycol nucleic acids, threo - Nucleic acids, mixed samples of nucleic acids, polyploid DNA (i.e., plant DNA), mixtures thereof, And hybrids thereof can be cited. In a preferred embodiment, genome D NA or an amplified copy thereof is used as the target nucleic acid. In another preferred embodiment, c DNA, mitochondrial DNA, or chloroplast DNA is used. In some embodiments, Therefore, the target nucleic acid is mRNA.

[0132] In some embodiments, the target nucleic acid is derived from a single cell or a fraction of a single cell. In some embodiments, the target nucleic acid is derived from a single organelle. Symbolic single organelles include, but are not limited to, a single nucleus, a single mitochondria. Examples include ribosomal Derived from formalin-fixed paraffin-embedded (FFPE) samples. In some embodiments... In this case, the target nucleic acid is a cross-linked nucleic acid. In some embodiments, the target nucleic acid is a cross-linked nucleic acid. It crosslinks with a protein. In some embodiments, the target nucleic acid is crosslinked DNA. In some embodiments, the target nucleic acid is histone-protected DNA. In some embodiments, histones are removed from the target nucleic acid. In some embodiments, The target nucleic acid is derived from a nucleosome. In some embodiments, the target nucleic acid is It is derived from nucleosomes from which the nuclear protein has been removed.

[0133] The target nucleic acid may include any nucleotide sequence. In some embodiments, The target nucleic acid contains homopolymer sequences. The target nucleic acid may also contain repeating sequences. Yes, it is possible. Repeating sequences can be, for example, 2 nucleotides, 5 nucleotides, 10 nucleotides, etc. , 20 nucleotides, 30 nucleotides, 40 nucleotides, 50 nucleotides, 100 Nucleotides, 250 nucleotides, 500 nucleotides, 1000 nucleotides, or It can be any length, including any length greater than that. The repeating sequence can be continuous or non-continuous. Next, for example, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 15 times, The repetition can be any number of times, including 20 or more.

[0134] Some embodiments described herein can use a single target nucleic acid. Other embodiments may use multiple target nucleic acids. In such embodiments, Multiple target nucleic acids include multiple identical target nucleic acids, and several target nucleic acids are the same. This can refer to multiple target nucleic acids, each with different target nucleic acids, or all of which are different target nucleic acids. Embodiments using multiple target nucleic acids include, for example, one or more chambers or arrays. The process can be carried out in a multiplexed manner so that the reagent is simultaneously delivered to the target nucleic acid on the surface. In that embodiment, the multiple target nucleic acids include the genomes of virtually all specific organisms. Examples include: Multiple target nucleic acids, for example, at least about 1% of the genome. 5%, 10%, 25%, 50%, 75%, 80%, 85%, 90%, 95%, or 99% This includes at least a portion of the genome of a specific organism. In this state, the portion in question accounts for approximately 1%, 5%, 10%, 25%, 50%, and 75% of the genome. It may have an upper limit of %, 80%, 85%, 90%, 95%, or 99%.

[0135] Target nucleic acids can be obtained from any source. For example, target nucleic acids can be obtained from a single organism. nucleic acid molecules obtained from the body, or nucleic acids obtained from natural sources containing one or more living organisms. They may also be prepared from a group of molecules. While not limited to this, the source of nucleic acid molecules can also be cell microsaturated. Examples include organs, cells, tissues, organs, or living organisms. Used as a source of target nucleic acid molecules. Cells that can be used are prokaryotic cells (bacterial cells, for example, Escherichia coli). a) Bacillus, Serratia, Salmonea almonella, staphylococcus, streptococcus reptococcus, Clostridium, Chlamydia (Chlamydia), Neisseria, Treponema Mycoplasma, Borrelia ), Legionella, Pseudomonas Mycobacterium, Helicobacter bacter), Erwinia, Agrobacterium (cterium), rhizobia (Rhizobia), and Streptomyces g enera; Crenarchaeota, Nanoarchaeological phylum (Oarchaeota), or other ancient organisms such as the phylum Euryarchaeotia. Mycete cells; or fungi (e.g., yeast), plants, protozoa and other parasites, and animals (insects) For example, fruit flies (Drosophila spp.), nematodes (for example, line Insects (Caenorhabditis elegans), and mammals (e.g., rats) The target may be eukaryotic cells such as those of mice, monkeys, non-human primates, and humans. Nucleic acids and template nucleic acids are prepared using various methods well known in the art to achieve the desired result. Specific sequences can be enriched. An example of such a method is explained in detail by reference. It is provided in International Publication No. 2012 / 108864, which is incorporated into the detailed document. In several embodiments, nucleic acids are further concentrated during the preparation of the template library. This may also be the case. For example, nucleic acids may be before transposome insertion, after transposome insertion, and The nucleic acid may be enriched for a specific sequence after amplification.

[0136] Furthermore, in some embodiments, the target nucleic acid and / or template nucleic acid are highly purified. For example, nucleic acids can be used in the methods provided herein with minimal impurities. Also approximately 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% It may be omitted. In some embodiments, known in the art It is beneficial to use methods that maintain the quality and size of the target nucleic acid, for example, DNA isolation and / or direct transposition may be performed using an agarose plug. Furthermore, it can be performed directly in cells using cell populations, lysates, and unpurified DNA.

[0137] In some embodiments, the target nucleic acid may be obtained from a biological sample or a patient sample. Good. As used herein, the terms “biological sample” or “patient sample” refer to tissue and This includes samples of bodily fluids, etc. "Bodily fluids" are not limited to these, but include blood, serum, plasma, etc. , saliva, cerebrospinal fluid, pleural fluid, tears, lactal duct fluid, li Examples include phlegm, sputum, urine, amniotic fluid, and semen. The sample is a body fluid that is "acellular". It may contain. "Cellless fluid" contains less than approximately 1% (mass / mass) of total cellular material. Plasma or serum is an example of cell-free body fluid. Samples are of natural or synthetic origin (i.e., cell-free). It may include specimens of cell samples prepared to achieve the following:

[0138] In some embodiments of the methods disclosed above, the target nucleic acid is trans Before exposure to posomes, flag (e.g., by sonication, restricted digestion, or other mechanical means) It can be made into a mental form.

[0139] As used herein, the term "plasma" refers to the cell-free fluid found in blood. This involves extracting all of the blood from the blood by methods known in the art (e.g., centrifugation and filtration). It can also be obtained from blood by removing cellular material.

[0140] Unless otherwise specified, the terms "a" or "an" throughout this specification mean "one or It means "more than that".

[0141] Terms such as "for example", "e.g.", and "etc." has) ``include ``include ``include `` or When these variations are used herein, these terms are not considered limiting terms. It does not mean "not limited to this" or "without limitation." This is how it is interpreted.

[0142] The following examples are for illustrative purposes only and are not intended to be used in any way provided herein. This does not limit the invention. [Examples]

[0143] Example 1: Yield of DNA clusters from the bead-based tagging process The yield of DNA clusters from the bead-based tagging process in Figure 3 was evaluated, and Figure 4 shows... The results are shown in the table. In this example, 50 ng, 250 ng, and 1000 ng of human NA128 were used. 78 DNA was tagged using tagging beads (2.8 μm beads) from the same batch. The second batch of 50 ng of NA12878 DNA was tagged. Tagmentation was performed using t-type beads (complete repeat: 2.8 μm beads). Bead bonding. Tagmented DNA samples were amplified by PCR and purified. Each purified sample was then divided into fixed volumes (5.4 μL). The prepared PCR product (unquantified) was diluted 270-fold to create a stock sample solution of approximately 50 pM. For each sample, a 50 pM stock solution was prepared at 15 pM, 19 pM, 21 pM, and 2 pM concentrations. The sample was diluted to 4 pM. The diluted sample was used for cluster generation and sequencing. Loaded into the flow cell. According to the data, starting with the same dilution (~50 pM), The number of stars is three different input levels using the same set of beads (i.e., 50n). For g, 250 ng, and 1000 ng, the values ​​are found to be between 100% and 114%. The cluster count with a 50ng complete repeat (using beads from different batches) was 81%. Yes, there were different dilutions (15 pM, 19 pM, 21 pM, and 24 pM) that were approximately 10% apart. It produces the same number of clusters within. According to the data, the beads greatly control the yield. This shows that the results are reproducible with different DNA inputs and different repeats.

[0144] Example 2: Reproducibility of the bead-based tagment process Figure 5 shows the reproducibility of the bead-based tagment process in Figure 3. In this embodiment, "the same Six different indexed beads (indexed beads) made with transposome density Prepare the solution (1-6 μm beads; 2.8 μm beads) with 50 ng and 500 ng of input NA1 2878 DNA was used to prepare tagged DNA. Tagged DNA was processed with P CR amplification and purification were performed. Twelve purified PCR products were poured into two HiSeq lanes. The mixture was divided into two pools (pool 1 and pool 2), each containing six different types of liquids. Each sample contains 3-50 ng and 3-500 ng. Data Table 500 shows each sample. This shows the median and mean insertion sizes for indexed samples.

[0145] Example 3: Insertion size of pool 1 and insertion size of pool 2 The insertion sizes for pool 1 and pool 2 of the indexed sample in Figure 5 are as follows: These are shown in Figure 6A (plot 600) and Figure 6B (plot 650), respectively. The data is... Furthermore, the insertion size is uniform across the six different indexed bead preparations. This demonstrates that bead-based tagging is a mechanism that controls insertion size and DNA yield. It brings about.

[0146] Example 4: Reproducibility of the total number of leads Regarding the experiment shown in Figure 5, the reproducibility of the total number of leads and the percentage of aligned leads. This is shown in Figure 7 (bar graph 700). With both inputs (50ng and 500ng), The total number of codes is the same for the same indexed bead preparation. 6 types of codes Of the four types of bead preparations with dex (index 1, 2, 3, and 6), the most They showed similar yields, and in indexed bead preparations 4 and 5, the index sequence was There was some variation, which may indicate that it is related to the actual data.

[0147] In one application, the bead-based tagging process includes a tagging step. Exome enrichment assays, for example, Illumina's Nextera® It may be used in rapid capture and enrichment protocols. Current exome enrichment assays (i.e., Illu mina's Nextera® rapid capture and concentration protocol is solution-based Tagming (Nextera) is used for fragmentation of genomic DNA. Subsequently, Use gene-specific primers to pull down the target specific gene fragment. After two enrichment cycles, the pulled-down fragments were subjected to PCR and sequencing. It is concentrated by synthesizing.

[0148] To evaluate the use of a bead-based tagging process in exome enrichment assays. Human NA12878 DNA, 25ng, 50ng, 100ng, 150ng, 20 Tagging was performed using 0 ng and 500 ng of input DNA. Ibrary (NA00536) according to the standard protocol, input DN 50ng Prepared from A. Each DNA input had a different index (unique identifier). To conform to the standard method, and to ensure that a sufficient number of fragments are available for pulldown. To ensure this, concentrated polymerase master mix (EPM: enhanced We performed 10 cycles of PCR using polymerase mastermix. The amplification protocol involves 3 minutes at 72°C, 30 seconds at 98°C, followed by 10 seconds at 98°C, repeated 10 times. The samples were then heated at 65°C for 30 seconds and at 72°C for 1 minute. After that, the samples were kept at 10°C. Next, the sample is processed through an exome enrichment pull-down process and sequencing. Ta.

[0149] Example 5: Control library and bead base in exome enrichment assay Insertion size of tagging library Figures 8A, 8B, and 8C show the control in the exome enrichment assay, respectively. Plotting insertion size 800 in the library, bead-based tagging library - Plot 820 of the insertion size and summary data table 840 are shown. Depending on the data In contrast, bead-based tagging libraries are more widely used than control libraries. It has a wide insertion size distribution, but the insertion size is polarized regardless of the DNA input of the sample. It's clear that they're close.

[0150] Example 6: Read Sequence Quality Figures 9A, 9B, and 9C show the exome enrichment process of Figures 8A, 8B, and 8C, respectively. In Sei, duplicates that have passed through a filter (dups PF: duplicates) Bar graph of the percentage of passing filters (900, PCT selected) 920 bar graphs of bases, and PCT usable bases on tar The bar graph for get shows 940. Referring to Figure 9A, the percentage of duplicates PF is: This is a measure that indicates how many reads are replicated in other parts of the flow cell. To ensure that all clusters provide useful data for the results, Ideally, it should be lower (as shown here).

[0151] Figure 9B shows PCT selected bases, which were concentrated during the concentration process. This is a measure of the proportion of leads that are aligned in or near the intended site where they should be. Ideally, This value will be close to 1, reflecting the success of the concentration process. Furthermore, this value is... This indicates that leads that should not be concentrated have not completed the process.

[0152] Figure 9C shows PCT usable bases on target, enriched and This is a measure of the proportion of reads that are actually sequenced on the target specific base within a given region. Ideally, all concentrated reads would be aligned on the target base within the concentrated read. However, due to the randomness of tagging and the varying insertion lengths, the target region is distributed Reads that have not been sorted may become concentrated.

[0153] Two techniques may be used to optimize the insertion size distribution. For example, SPRI You may use cleanup to remove fragments that are too small or too large. RI cleanup involves size and retention of desired precipitated or non-precipitated DNA (i.e., first stage). The filter precipitates only DNA larger than the desired size, separating the smaller, soluble fragments. Selective DNA precipitation based on retention allows for the removal of fragments larger or smaller than the desired size. This is the process of removing the lagging. After that, the smaller fragments are further precipitated, At this time, remove any unwanted, extremely small fragments (still in the solution) and precipitate the DNA. The DNA is retained, washed, and then resolubilized to obtain DNA within the desired size range. Another example is... Then, using the spacing of activated transposomes on the bead surface, the insertion size distribution is determined. It may be controlled. For example, the gap on the bead surface may be controlled by inactive transposomes (e.g., It may be filled with transpososomes (which have inert transposons).

[0154] The continuity of the bead-based tagging process was evaluated. Table 3 shows the shared index. This indicates the number of times 0, 1, 2, or 3 reads occurred within a 1000bp window. The beads were generated using nine different indexed transpososomes, and a small amount of human DN was added. Used for tagging A. Generated reads, aligned them, and shared the same index. The number of reads within a 1000bp or 10KB window was analyzed. Index Among the reads in the small window that is shared, some may be generated by chance. The prediction of how many times this is likely to occur is shown in the "Random" column of Tables 3 and 4. The number of "Beads" columns is 1000bp (Table 3) or 10Kp (Table 4) that share an index. This shows the actual number of windows. As shown in Tables 3 and 4, the same index has 1000 windows. The actual number of times found within a bp or 10Kp window is higher than the prediction in the random case. This is also remarkably common. The "0" frame is the one that a specific 1000bp window maps to. This shows all number of iterations that do not contain DEX reads. The numbers indicate that only a very small amount of human genome is being sought. Because they are oriented and most windows do not have leads to align with them, This is the largest. "1" means that only one read wins 1000bp (or 10Kp). This is the number of times to map to the dow; "2" means within a 1000bp (or 10Kp) window. This is the number of times two reads share an index, etc. This data is 1400 In the extreme case, the same piece of DNA (over 10Kp) undergoes approximately 15,000 tagging cycles. In the tagging event, the same bead is tagged at least two to five times. This suggests that because fragments share an index, they are even Therefore, it's unlikely that it exists there, and it must originate from the same bead.

[0155] [Table 3]

[0156] Table 4 shows the number of reads (up to 5) within a 10Kp window that shares an index. .

[0157] [Table 4]

[0158] Example 7: Isolation of free transpososomes from CPT-DNA After rearrangement, the reaction mixture containing CPT-DNA and free transpososomes, Sepha Cryl S-400 and Sephacryl S-200 size exclusion chromatography The results were obtained using column chromatography with - shown in Figure 22. CPT-DNA was NC It will be labeled as P DNA.

[0159] Example 8: Optimization of the capture probe density on beads The densities of capture probes A7 and B7 were optimized on 1 μm beads, and the results are shown in Figure 25. Lane 1 (A7) and Lane 3 (B7) have a high probe density, and Lane 2 (A 7) and lane 4 (B7) have an estimated 10,000 to 100,000 beads per 1um. It had probe density. The ligation product of the capture probe against the target molecule was agaro -Evaluated using a gel. The probe density per bead was approximately 10,000 to 100,000. It exhibited better ligation efficiency than higher probe densities.

[0160] Example 9 Indication of CPT-DNA on beads by intramolecular hybridization Feasibility study of preparing sequencing libraries with ks. Transposomes capture A7' and B7' elements that are complementary to the A7 and B7 capture sequences on the beads. By mixing a transposon with a grafting sequence with a hyperfunctional Tn5 transposase... The following was prepared: High molecular weight genomic DNA was mixed with transpososomes, and CPT-DNA was used. It generates. Separately, beads are immobilized oligonucleotides: P5-A7, P7-B 7. Prepare with P5-A7 + P7-B7. Here, P5 and P7 are primer conjugates. The sequence is such that A7 and B7 are complementary capture sequences to the A7' and B7' sequences, respectively. P5-A7 alone, P7-B7 alone, P5-A7 + P7-B7, or P5-A7 beads and Beads containing a mixture of P7-B7 beads are treated with CPT-DNA, and the reaction mixture is Gauze is added to improve the efficiency of hybridization of immobilized oligonucleotides to transposed DNA. The decision was made. The results are shown in Figure 26. The sequencing library was prepared on an agarose gel. As indicated by the high molecular weight band, P5-A7 and P7-B7 are both in a single bead. This is only produced when it is immobilized on top (lane 4). The result is a highly efficient intramolecular high It exhibits hybridization, and also C on the beads due to intramolecular hybridization. We demonstrated the feasibility of preparing an indexed sequencing library of PT-DNA. Ta.

[0161] Example 10 Feasibility Test of Clone Indexing Several transposome sets were prepared. In one set, enhanced transpososomes were found. n5 transposase was mixed with the transposon sequence Tnp1, which has 5' biotin. Prepare transposome 1. In another set, a unique index having 5' biotin Transosome 2 is prepared using Tnp2 having TX2. In another set, For the preparation of Lansposome 3, enhanced Tn5 transposase and 5' biotin were used. It mixes with the transposon sequence Tnp3. In another set, the unique index Transposomal 4 is prepared with Tnp4 containing 4 and 5' biotin. Transosomes 1 and 2, and transpososomes 3 and 4, were each separately subjected to streptavidin. Mix with the beads to produce Bead Set 1 and Bead Set 2. Next, the two sets of beads Mix the ingredients and incubate with genomic DNA and tagging buffer. This promotes the tagging of genomic DNA. Subsequently, PCR amplification of the tagged sequence is performed. The amplified DNA is sequenced, and the insertion of index sequences is analyzed. Tagmen When the tethering is limited to beads, the majority of fragments are Tnp1 / Tnp2 and Tn This will be coded using p3 / Tnp4 indices. Intramolecular hybridization If present, the fragments are Tnp1 / Tnp4, Tnp2 / Tnp3, Tnp1 It will be coded with / Tnp3 and Tnp2 / Tnp4 indices. 5 cycles The sequencing results after PCR for 10 cycles are shown in Figure 27. The result contains all four types of transposons, which are mixed and immobilized on beads. The majority of sequences have Tnp1 / Tnp2 or Tnp3 / Tnp4 indices. This demonstrates that clonal indexing is feasible. The control is: This indicates that the index is not distinguished.

[0162] Example 11 Indexed clonal bead rearrangement in a single reaction Prepare 96 types of indexed transposomal beads. The attached transposome has a Tn5 mosaic terminal sequence (ME) at its 5' end and an index A mixture of transposons containing oligonucleotides having a sequence was prepared. INDEX-equipped transpososomes are subjected to streptavidin-biotin interactions. They were immobilized on the beads. The transposomes on the beads were washed, and all 96 types of individual transposomes on the beads were removed. Each indexed transposome was pooled. The ME sequence was complementary and indexed. Oligonucleotides having a kus sequence are annealed to immobilized oligonucleotides, and unique We created transposons with the following indexes: 96 types of clone indexes. Mix the transposome bead set together and mix to create a high molecular weight (HMW: high mole) bead. (Cycular weight) Genomic DNA along with Nextera tagging buff The sample was incubated in a single tube in the presence of ferrite.

[0163] The beads are washed and treated with a 0.1% SDS reaction solution, which is then used to transport the beads. Remove the ze. Amplify the tagged DNA with indexed primers and PE Hi Using the Seq flow cell v2, sequencing is performed using the TrueSeq v3 cluster kit. We then perform sequencing and analyze the sequencing data.

[0164] Observe the clusters of reads, or islands. The nearest neighbor distance between reads for each sequence is The plot shows the main peaks, one from within the cluster (proximal) and the other from the class. This primarily shows the results from between the terminals (distal). Schematic diagrams of the method and results are shown in Figures 30 and 31. As shown, the island size ranges from approximately 3 to 10 kb. The percentage of covered bases is approximately 5. The percentage is approximately 10%. The insertion size of genomic DNA is about 200-300 base pairs.

[0165] Example 12 Library size for transposomes on beads First, the first oligonucleotide having the ME' sequence, ME-barcode-P5 / P Mix a second oligonucleotide having sequence 7 with Tn5 transposase. Transposomes were assembled in solution by [method]. In the first set, ME' [component] The first oligonucleotide, which has a row, is biotinylated at its 3' end. Then, an oligonucleotide having the ME-barcode-P5 / P7 sequence is subjected to a 5' end. Otinization is performed. Each of the obtained transtrans at various concentrations (10 nM, 50 nM, and 200 NM) Streptoavidin beads are added to the sporosome set, and the transpososomes The treptavidin beads are immobilized. The beads are washed, and the HMW genome DN is removed. Add A and perform tagging. In some cases, 0.1% of the tagged DNA is used. In other cases, the tagged DNA is not processed with SDS. A is amplified by PCR for 5-8 cycles and then sequenced. A schematic diagram is shown in Figure 32.

[0166] As shown in Figure 33, SDS processing improves amplification efficiency and sequencing quality. Oligonucleotides containing 3' biotin offer better performance against transpososomes. It has ibrary size.

[0167] Figure 34 shows the effect of transposome surface density on insertion size. 5' biotin Transpososomes with this feature have smaller libraries and more self-insertion. This shows the by-product.

[0168] Example 13: Titration of input DNA Various amounts of target HMW DNA are used in a 50 mM Tn5 transposon density In addition to the beads with the lawn index, leave them at 37°C for 15 or 60 minutes, or at room temperature. Incubated for 60 minutes. Transpososomes contain oligonucleotides with 3' biotin. It contained ocidal compounds. Tagming was performed, and the reaction mixture was treated with 0.1% SDS and then amplified by PCR. The width was increased. The amplified DNA was sequenced. Figure 35 shows the inverse of the size distribution. This shows the effect of the input DNA. The response with 10 pg of input DNA shows a minimum signal. The size distribution patterns were shown for 20pg, 40pg, and 200pg DNA inputs. The same was true in To.

[0169] Example 14: Island size and distribution using solution-based and bead-based methods The size and distribution of the islands were compared using solution-based and bead-based methods. In the approach, there are 96 types of transposons, each with its own unique index. Assemble the transpososomes in a 96-well plate. Add HMW genomic DNA. Next, the tagging reaction is performed. The reaction product is treated with 0.1% SDS and amplified by PCR. The amplified DNA was sequenced.

[0170] In the bead-based approach, each transposon has its own unique index. 96 types of transposomes were assembled into a 96-well plate. The solution contained 3'-terminal biotin. Streptoavidin beads were placed in each of the 96 wells of a plate. Add and incubate so that the transposomes are immobilized on streptavidin beads. The beads are washed and pooled, HMW genomic DNA is added, and tagging is performed. The t-conjugation reaction is carried out in a single reaction vessel (one-pot). The reaction product is treated with 0.1% SDS. The product was then amplified by PCR. The amplified product was sequenced.

[0171] In the negative control, we first have 96 types, each with its own unique index. Mix all transposon sequences. Oligonucleotides contain 3'-terminal biotin. Therefore, transpososomes are prepared from individual mixed-index transposons. Leptavidin beads are added to the mixture. HMW genomic DNA is added and tagged. The reaction is carried out. The reaction product is treated with 0.1% SDS and amplified by PCR. The amplified product is then sieved. Quensing occurred.

[0172] The number of reads within an island is plotted against the size of the island. The results shown in Figure 36 indicate that the island (proximal reads) (The method) is similar to solution-based methods, but can be observed with one-pot clone indexed beads. This indicates that indexed transposons can be used before transposome formation. When mixed, no islands (proximal reads) were observed. Transpososome formation occurred before transpososome formation. By mixing posons, each bead has different index / transposomes. This results in beads that are not clones.

[0173] Example 15: Analysis of structural mutations using CPT-seq Detection of 60kb heterojunction defects Extract the sequencing data as a fastq file and separate it (demultiplex). (S) Perform the process to generate individual fastq files for each barcode. The fastq files from sequencing are separated according to the index and duplicates are removed. Align to the removed reference genome. Index any read within the scan window. Scan chromosomes using a 5kb / 1kb window to record the number of chromosomes. Statistical Due to the heterozygous deletion region, it contains only half the amount of DNA compared to the adjacent region in the library. It cannot be used for generation. Therefore, the number of indices should be approximately half that of adjacent areas. NA12878 chr1 60kb heterojunction deletion with 9216 indexes By scanning the CPT sequencing data into 5kb windows, Figure 47 As shown in A and 47B.

[0174] Detection of gene fusions The fastq files from CPT sequencing are separated according to the index. Align to a reference genome with duplicates removed. Scan chromosomes in 2kb windows. Each 2kb window is 36864 vectors, and reads from the unique index are Each element records how many were found within a 2kb window. Across the genome, For every 2kb window pair (X, Y), a weighted jackard (weighted-Ja The ccard) index is calculated. This index is effectively the (X, This indicates the distance between Y's. These indices are displayed as a heatmap as shown in Figure 48. Each data point represents a pair of 2kb scan windows, and the square in the upper left corner is, Both are X and Y from region 1, the bottom right is X and Y from region 2, and the top right is region These are X and Y from the region spanning from region 1 to region 2. The gene fusion signal in this case is It is represented as a horizontal line in the center.

[0175] Detection of defects The fastq files from CPT sequencing are separated according to the index. Align to a reference genome with duplicates removed. Scan chromosomes in 1kb windows. Figure 49 shows the results of detecting gene deletions.

[0176] Example 16: Detection of phasing and methylation Optimization of bisulfite conversion efficiency For indexed coupled CPT-seq libraries on beads, ME (Moss The conversion was evaluated in the elemental region and the gDNA region. Promega's Meth The ylEdge bisulfite conversion system was optimized to improve efficiency.

[0177] [Table 5]

[0178] The efficiency of the bisulfite conversion process was determined by analyzing the ME sequence. This is shown in Figure 50. (Beads) Of the indexed bound libraries attached to the substrate, 95% were converted to bisulfite (BS). C) was done. Similar PCR yields were observed between bisulfite conditions. > Stricter bisulfite The library did not appear to decompose even with hydrogen salt treatment. This is shown in Figure 51. BSCs were observed in approximately 95% of indexed combined libraries. BSCs were improved (C The variables studied to determine >U were temperature and NaOH concentration (denaturation). Good results were obtained with 1M NaOH, or with 0.3M NaOH and ℃.

[0179] After sequencing of BSC-converted CPT-seq in a bead library, the expected sequence - Quensing read structure was observed. The measured percentage of bases (%) was plotted on the IVC plot. This is shown in Figure 52.

[0180] Figure 53 shows the indexed conjugated library after PCR following bisulfite conversion. The image shows the results of agarose gel electrophoresis. The expected size range of 200-500 bp was observed. A burr was observed. In the reaction without DNA, an indexed binding library was used. It does not produce.

[0181] Example 17 Target phasing The whole-genome indexed conjugated CPT-seq library was enriched. Figure 54 shows the results. A bound CPT-seq library with whole genome indexes before enrichment without isoselection. The bioanalyzer trace is shown. Figure 55 shows the agarose gel of the concentrated library. The analysis is presented below.

[0182] The enrichment statistics for the HLA region are shown below.

[0183] [Table 6]

[0184] Figure 56 shows the results of applying targeted haplotyping to the HLA region of a chromosome. A diagram illustrating the enrichment of a whole-genome indexed bound read library is shown on the left. - indicates an indexed short library. A raster is a region that is clone-indexed on a single bead with the same index. It is an "island," and therefore exhibits read proximal relationship ("island" property) on a genome scale. Library enrichment in a specific domain (International Publication No. 2012 / 108864 "Selection of Nucleic Acids") Selective enrichment of nucleic acid See (s) shown on the right. The read is enriched in the HLA region. Furthermore, the read is When reads are sorted by DEX and aligned to the genome, the reads are indexed binding leads. This again shows the "island" structure, which indicates that continuity information is maintained from the code.

[0185] Example 18: Index Replacement To evaluate the exchange of mosaic ends (MEs) of transpososome complexes, different indices were used. Beads containing ribs were prepared. After mixing, the library was sequenced, and each rib By reporting on the Braley index, we decided to replace the index. The percentage of "swapped" is calculated as (D4+D5+E3+E5+F4) / (Total 96) The calculation was performed using the total number of species. This is shown in Figure 65.

[0186] Example 19: Condensing transposome complexes with streptavidin beads. Reduced library insertion size Streptoavidin magnetic beads are used in TsTn5 transformers at 1x, 6x, and 12x concentrations. Loaded together with the posome complex. For each bead type, Epi-CPTSeq Rotol was performed. For analysis, the final PCR product was sent to Agilent BioAnaly. It was loaded onto zer. This is shown in the figure. The Epi-CPTSeq library fragment is: They were relatively small, and the more TsTn5 was loaded onto the beads, the more they were produced.

[0187] Example 20 Fragmentation of DNA library during bisulfite conversion After bisulfite conversion, the DNA is damaged, resulting in the loss of the common sequence necessary for PCR amplification. CS2) decreases. DNA fragment CPTSeq and Epi-CPTSeq (Me -CPTSeq) libraries were analyzed with BioAnalyzer. Bisulfite mutations Due to DNA damage during exchange, the Epi-CPTSeq library, as shown in Figure 70, Compared to CPTSeq libraries, it offers a 5x lower yield and a smaller library size. It has a distribution.

[0188] Example 21: TdT-mediated ssDNA ligation reaction DNA end repair by terminal transferase (TdT) mediated ligation Feasibility was tested. In short, a 5 pmol ssDNA template was used, TdT (10 / 50 U) , attenuator / adapter double strand (0 / 15 / 25 pmol), and DNA rigger Incubated with Ze (0 / 10U) at 37°C for 15 minutes. Extension / Ligation The DNA products were analyzed using TBE-Urea gel, and the results are shown in Figure 71. All reactions As a result of adding the component, almost complete ligation of the adapter molecule occurred (Lane 5). ~8).

[0189] DNA end repair by terminal transferase (TdT) mediated ligation Feasibility was tested using a sodium bisulfite conversion bead-bound library. Figure 72 As shown below, simply tag the DNA on the beads (first two lanes), and Prome Processed with ga's MethylEdge bisulfite conversion kit (lanes 3 and 4), D The NA rescue protocol was performed (lanes 5 and 6). The DNA library after the rescue reaction was collected. The quantity and size have clearly increased. Also, the abundance of inserted transposons (SIs) has increased. This is increasing, indicating efficient ligation of adapter molecules.

[0190] The results of the methyl-CPTSeq assay are shown in Figure 73.

Claims

1. (a) A step of bringing a target nucleic acid into contact with a plurality of transposomal complexes, Each transposome complex contains a transposon and a transposase. The transposon comprises a transfer chain and a non-transfer chain, and at least one of the transposons in the transposome complex comprises an adapter sequence capable of hybridizing to a complementary capture sequence, the steps include: (b) A step of fragmenting the target nucleic acid into a plurality of fragments and inserting a plurality of transfer chains using the transposome complex, wherein adjacent fragments are linked to each other via the transposome complex to produce a plurality of linked fragments, and the continuity of the fragments of the target nucleic acid is maintained by the transposase, (c) A step of contacting the plurality of linked fragments of the target nucleic acid with a plurality of solid supports, each of the plurality of solid supports comprising a plurality of immobilized oligonucleotides, each of the immobilized oligonucleotides comprising a complementary capture sequence and a first barcode sequence, wherein the first barcode sequence from each solid support in the plurality of solid supports is different from the first barcode sequence from the other solid supports in the plurality of solid supports. (d) A step of transferring the barcode sequence information to the linked plurality of fragments of the target nucleic acid, wherein at least two fragments of the same target nucleic acid receive the same barcode information. (e) The step of subjecting the target nucleic acid fragment containing the barcode to bisulfite treatment to generate a bisulfite-treated target nucleic acid fragment containing the barcode, (f) A step of determining the sequence information of the bisulfite-treated target nucleic acid fragment and the barcode sequence, (g) The step of determining the continuity information of the target nucleic acid by identifying the barcode sequence, The sequence information indicates the methylation state of the target nucleic acid, and the continuity information indicates the haplotype information. A method for simultaneously determining the phasing information and methylation state of a target nucleic acid.

2. The method according to claim 1, wherein a single barcode sequence is present in each of the plurality of immobilized oligonucleotides on the individual solid support.

3. The method according to claim 1 or 2, wherein different barcode sequences are present in the plurality of immobilized oligonucleotides on each of the individual solid supports.

4. The method according to any one of claims 1 to 3, wherein the transfer of the barcode sequence information to the target nucleic acid fragment is performed by ligation.

5. The method according to any one of claims 1 to 4, wherein the transfer of the barcode sequence information to the target nucleic acid fragment is performed by polymerase elongation.

6. The method according to any one of claims 1 to 4, wherein the transfer of the barcode sequence information to the target nucleic acid fragment is performed by both ligation and polymerase elongation.

7. The method according to claim 5 or 6, wherein the polymerase elongation is performed by using the ligated immobilized oligonucleotide as a template and elongating the 3' end of the unligated transposon chain with DNA polymerase.

8. The method according to any one of claims 1 to 7, wherein at least a portion of the adapter array further comprises a second barcode array.

9. The method according to any one of claims 1 to 8, wherein the transposome complex is a polymer, and the adapter sequence of the transposon of each monomer unit is different from that of other monomer units of the same transposome complex.

10. The method according to any one of claims 1 to 9, wherein the adapter sequence further comprises a first primer binding sequence.

11. The method according to any one of claims 1 to 10, wherein the immobilized oligonucleotide on each of the solid supports further comprises a second primer-binding sequence.

12. The method according to any one of claims 1 to 11, wherein the transposome complex is a polymer, and the monomeric units of the transposome are bound to each other within the same transposome complex.

13. The method according to claim 12, wherein the transposase of a transposomal monomer unit is bound to another transposase of another transposomal monomer unit of the same transposomal complex.

14. The method according to claim 13, wherein a transposon of a transposomal monomer unit is bound to a transposon of another transposomal monomer unit of the same transposomal complex.

15. The method according to any one of claims 1 to 14, wherein the continuity information of the target nucleic acid represents haplotype information.

16. The method according to any one of claims 1 to 14, wherein the continuity information of the target nucleic acid indicates a genomic mutation.

17. The method according to claim 16, wherein the genomic mutation is selected from the group consisting of deletions, transpositions, interchromosomal gene fusions, duplications, and paralogs.

18. The method according to any one of claims 1 to 17, wherein the oligonucleotide immobilized on the individual solid supports includes a partially double-stranded region and a partially single-stranded region.

19. The method according to claim 18, wherein the partially single-stranded region of the oligonucleotide includes the second barcode sequence and the second primer-binding sequence.

20. The method according to any one of claims 1 to 19, wherein the target nucleic acid fragment including the barcode is amplified before the sequence information of the target nucleic acid fragment is determined.

21. The method according to claim 20, wherein steps (a) to (d) and the subsequent amplification are carried out in a single reaction compartment before determining the sequence information of the target nucleic acid fragment.

22. The method according to claim 20, wherein a third barcode sequence is inserted into the target nucleic acid fragment during the amplification.

23. The step of combining the target nucleic acid fragments including the barcode from a plurality of first sets of reaction compartments into a pool of target nucleic acid fragments including the barcode, The steps include: redistributing the pool of target nucleic acid fragments including the barcode into a plurality of second sets of reaction compartments; The steps include introducing a third barcode into the target nucleic acid fragment by amplifying the target nucleic acid fragment in the reaction compartment of the second set before sequencing, and The method according to any one of claims 1 to 22, further comprising:

24. The method according to any one of claims 1 to 23, further comprising the step of pre-fragmenting the target nucleic acid before contacting the target nucleic acid with a transposome complex.

25. The method according to claim 24, wherein the step of pre-fragmenting the target nucleic acid is performed by a method selected from the group consisting of sonication and restriction digestion.