Sequence-specific targeted transposition, selection, and sorting of nucleic acids
Patent Information
- Application Number
- JP2023503428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-02
- Filing Date
- 2021-08-17
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2041-08-17
AI Technical Summary
【0557】 本方法の他の利点は、本明細書に記載される。
Smart Images

Figure 0007914079000003 
Figure 0007914079000004 
Figure 0007914079000005
Abstract
Description
[Technical Field]
[0001] (Cross-reference of related applications) This application claims priority to U.S. Provisional Patent Applications No. 63 / 066,905 and No. 63 / 066,906, filed on 18 August 2020, No. 63 / 162,775, filed on 18 March 2021, No. 63 / 163,381, filed on 19 March 2021, No. 63 / 168,753, filed on 31 March 2021, and No. 63 / 228,344, filed on 2 August 2021, each of which is incorporated herein by reference in whole for any purpose.
[0002] (Sequence Listing) This application is filed together with a sequence listing in electronic format. The sequence listing was created on July 28, 2021, and is provided as a file titled "2021-07-28_01243-0020-00PCT_Seq_List_ST25", with a size of 4,096 bytes. The information in the electronic format of the sequence listing is incorporated herein by reference in its entirety.
[0003] (Field of Invention) explanation This disclosure relates to sequence-specific targeted transposition of nucleic acids. Sequence-specific targeted transposition can be mediated using targeted transposomal complexes. This disclosure relates to a method comprising initial sequencing, selection, and resequencing for evaluating a desired sample. As described herein, initial sequencing can identify the desired sample in a pool of mixed samples, which can then be depleted of unwanted samples or enriched based on unique sample barcodes. The desired sample can then be resequenced. [Background technology]
[0004] Library construction of selected regions of target nucleic acids may be desired for many different applications. For example, the ability to construct libraries from selected regions of genomic DNA is desirable when platform output is limited (e.g., PacBio, ONT, or iSeq). Furthermore, libraries of selected regions of genomic DNA are advantageous when a very high range of applicability is required, such as for screening rare somatic mutations in liquid biopsy samples.
[0005] Current methods for obtaining libraries from selected regions of genomic DNA include oligonucleotide hybridization-based enrichment kits (e.g., TruSeq Exome, Nextera Flex for Enrichment). Furthermore, CRISPR-based systems for constructing such libraries have recently become available. In particular, CRISPR-based systems are used to extract tens to hundreds of kilobase regions suitable for long-read technologies such as PacBio and ONT.
[0006] This disclosure describes novel methods for preparing targeted libraries of desired regions of genomic DNA. These methods combine different targeting techniques with transpososomes in many unique ways. Furthermore, this disclosure describes means for preparing targeted libraries from cell-free DNA (cfDNA) that do not require histone removal before tagmentation.
[0007] This disclosure also describes single-cell analysis methods that may be used to elucidate intercellular differences that are difficult to determine when studying bulk populations of cells. Characterization of rare cells may be important for several applications, including oncology (liquid or tumor biopsy, minimal residual disease or early disease detection, tumor evolution, or tumor resistance), immunology (immune or T cell receptor repertoire), and metagenomics (non-culturable biological genome assembly). Figure 1 provides some representative examples of metagenomic and oncological samples that may be of interest, in which rare cells are of very high interest. Current methods in single-cell sequencing allow for parallel “omics” analysis of millions of single cells to study the genomic, transcriptome, or epigenomic characteristics of individual cells, for example.
[0008] However, characterizing rare cells in a population based on comprehensive sequencing is costly and difficult if the desired sample is not selected. Furthermore, enrichment methods based on cell sorting are limited by the availability of discriminable cell features. For example, FACS can enrich cells for specific cell size, morphology, and surface protein expression, but other features may not be discriminable by FACS. Enriching cells based on specific “omics” features (e.g., enrichment based on the presence of species, cell type, or mutants) would be extremely useful. These features can be known in advance (based on the latest technology) or newly (determined by initial sequence analysis). It is also very valuable to perform follow-up, comprehensive / orthogonal “omics” characterization by resequencing single-cell-derived samples identified as desired after initial sequencing.
[0009] A methodology for selecting, enriching, and characterizing individual cell DNA libraries from a “single-cell sequence library” or “sc library” consisting of multiple cellular DNA libraries, including libraries prepared from different single cells, is disclosed herein. Initial sequencing of the sc library (i.e., sequencing of all DNA libraries derived from individual cells) can be performed, and individual cells can be classified with respect to specific “omics” features of interest using bioinformatics analysis. Using this method, libraries prepared from different individual cells are identified by unique cellular DNA barcodes (UBCs). The “omics” features used for selection may define cell type (e.g., expression, epigenetic pattern, or immunorecombination), species type (e.g., using bacterial 16s, 18s, or ITS rRNA / rDNA sequences), or disease status / risk (e.g., germline or somatic mutations important in cancer) using a relatively small, targeted sequencing panel. In other words, the footprint of the initial sequencing can be small, and resequencing can be made more comprehensive and focused on the target cells. Thus, those skilled in the art can query millions or billions of cells for exemplary features using a single initial sequencing to classify samples into desired and unwanted samples, followed by targeted resequencing of the desired samples.
[0010] Alternatively, initial sequencing can be performed to identify new, exemplary "omics" cellular features for follow-up analysis. For example, initial sequencing can be performed to identify new cellular features that can then be used for classification.
[0011] The enrichment or depletion in this method can be carried out by known nucleic acid-targeted enrichment methods (e.g., hybrid capture, specific sample barcode-specific amplification, or CRISPR digestion). Individual cellular DNA from the target cells can then be resequenced and characterized in isolation from a complete sc library. Therefore, this method can enable more comprehensive and / or orthogonal resequencing and analysis after the initial sequencing performed for cell selection. [Overview of the Initiative] [Means for solving the problem]
[0012] This disclosure describes several different targeted transposome complexes, each comprising one or more elements that guide the transposome complex to bind to one or more target nucleic acid sequences in a target nucleic acid. Several methods of using these targeted transposome complexes are also described herein.
[0013] This description also describes a method for characterizing a desired sample in a mixed pool of samples containing both desired and unwanted samples.
[0014] Embodiment 1. A targeted transposome complex comprising a transposase, a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence, a recombinase-coated targeted oligonucleotide, the targeted oligonucleotide being capable of binding to one or more target nucleic acid sequences, and a second transposon comprising a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence.
[0015] Embodiment 2. The transposome complex according to Embodiment 1, wherein the sequence of the targeted oligonucleotide is fully or partially complementary to one or more target nucleic acid sequences.
[0016] Embodiment 3. The transposome complex according to any one of Embodiments 1 or 2, wherein one or more targeting oligonucleotides are linked to the 5' end of the adapter sequence.
[0017] Embodiment 4. The transposome complex according to any one of Embodiments 1 to 3, wherein one or more targeting oligonucleotides are directly linked to the 5' end of the adapter sequence.
[0018] Embodiment 5. The transposome complex according to any one of Embodiments 1 to 4, wherein one or more targeting oligonucleotides are linked to the 5' end of the adapter sequence via a linker.
[0019] Embodiment 6. The transposome complex according to Embodiments 1 to 5, wherein the linker is an oligonucleotide linker.
[0020] Embodiment 7. The transposome complex according to Embodiments 1 to 6, wherein the linker is a non-oligonucleotide linker.
[0021] Embodiment 8. The transposome complex according to Embodiments 1 to 7, wherein both the 5' end of the adapter sequence and the targeting oligonucleotide are biotinylated and linked via streptavidin.
[0022] Embodiment 9. The transposome complex according to any one of Embodiments 1 to 8, wherein the adapter sequence comprises a primer sequence, an index tag sequence, a capture sequence, a barcode sequence, a cleavage sequence, or a sequencing-related sequence, or a combination thereof.
[0023] Embodiment 10. The transposome complex according to Embodiments 1 to 9, wherein the adapter sequence comprises a P5 or P7 sequence.
[0024] Embodiment 11. The transposome complex according to any one of Embodiments 1 to 10, wherein the recombinase is UVSX, Rec233, or RecA.
[0025] Embodiment 12. The transposome complex according to any one of Embodiments 1 to 11, wherein the transposome complex is in a solution.
[0026] Embodiment 13. The method according to any one of Embodiments 1 to 12, wherein the transposome complex is immobilized on a solid support.
[0027] Embodiment 14. The transposome complex according to Embodiments 1 to 13, wherein the solid support is a bead.
[0028] Embodiment 15. A kit or composition comprising: a first transposome complex described in any one of Embodiments 1 to 14, which is a targeted transposome complex; a second transposome complex comprising a first transposon comprising a transposase, a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence; and a second transposon comprising a 5' transposon terminal sequence, wherein the 5' transposon terminal sequence is complementary to the 3' transposon terminal sequence.
[0029] Embodiment 16. A kit or composition comprising two transposome complexes according to any one of Embodiments 1 to 14, each of which is a targeted transposome complex, wherein the two targeted transposome complexes contain different targeted oligonucleotides.
[0030] Embodiment 17. A method for the targeted generation of 5' tagged fragments of a target nucleic acid, comprising: combining a sample containing a double-stranded nucleic acid with a transposome complex described in any one of Embodiments 1 to 14, which is a targeted transposome complex; initiating strand entry of the nucleic acid with a recombinase; and fragmenting the nucleic acid into multiple fragments with a transposase, wherein the fragmentation is performed by ligating the 3' end of a first transposon to the 5' end of the fragments to produce multiple 5' tagged fragments.
[0031] Embodiment 18. A method for generating a library of tagged nucleic acid fragments, comprising: a sample containing a double-stranded nucleic acid; a first transposome complex as described in any one of Embodiments 1 to 14, which is a targeted transposome complex; a second transposome complex comprising a transposase; a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence; and a second transposon comprising a 5' transposon terminal sequence wherein the 5' transposon terminal sequence is paired with the 3' transposon terminal sequence. A method comprising: combining a second transposomal complex containing a complementary second transposon; initiating strand entry of nucleic acid by a recombinase; and fragmenting the nucleic acid into multiple fragments by a transposase, wherein the 3' end of each first transposon is ligated to the 5' end of a target fragment to produce multiple first 5'-tagged target fragments generated from the first transposomal complex and multiple second 5'-tagged target fragments generated from the second transposomal complex.
[0032] Embodiment 19. A method for generating a library of tagged nucleic acid fragments, comprising: combining a sample containing a double-stranded nucleic acid with a first transposomal complex described in any one of Embodiments 1 to 14, which is a targeted transposomal complex, and a second transposomal complex described in any one of Embodiments 1 to 14, which is a targeted transposomal complex; initiating strand entry of the nucleic acid with a recombinase; and fragmenting the nucleic acid into a plurality of fragments with a transposase, wherein the 3' end of each first transposon is ligated to the 5' end of a target fragment to produce a plurality of first 5'-tagged target fragments generated from the first transposomal complex and a plurality of second 5'-tagged target fragments generated from the second transposomal complex.
[0033] Embodiment 20. The method according to any one of Embodiments 17 to 19, or the kit or composition according to Embodiment 15 or Embodiment 16, wherein the 5' adapter sequences contained in the first transposome complex and the second transposome complex are different.
[0034] Embodiment 21. The method according to Embodiment 19, wherein the targeted oligonucleotides contained in the first transposome complex, which is a targeted transposome complex, and the second transposome complex, which is a targeted transposome complex, are different.
[0035] Embodiment 22. The method according to Embodiment 21, wherein targeted oligonucleotides of a first transposome complex, which is a targeted transposome complex, and a second transposome complex, which is a targeted transposome complex, bind to different target sequences within a given target region in a target nucleic acid.
[0036] Embodiment 23. The method according to Embodiment 22, wherein the targeted oligonucleotides of a first transposome complex, which is a targeted transposome complex, and a second transposome complex, which is a targeted transposome complex, are bound to the reverse strand of a double-stranded nucleic acid.
[0037] Embodiment 24. The method according to any one of Embodiments 17 to 23, wherein the initiation of nucleic acid strand entry by recombinase occurs in the presence of a recombinase loading factor, and optionally, the recombinase loading factor is removed or inactivated before fragmentation.
[0038] Embodiment 25. The method according to any one of Embodiments 17 to 24, wherein the initiation of strand penetration occurs via substitution loop formation.
[0039] Embodiment 26. The method according to any one of Embodiments 17 to 25, wherein strand entry is initiated within 40, 30, 20, 15, 10, or 5 bases from the binding site of a targeted oligonucleotide to one or more sequences of interest.
[0040] Embodiment 27. The method according to any one of Embodiments 17 to 26, wherein the temperature used to initiate strand entry is different from the optimal temperature for fragmentation by the transposase.
[0041] Embodiment 28. The method according to Embodiment 27, wherein the temperature used to initiate strand entry is below the optimal temperature for fragmentation by the transposase.
[0042] Embodiment 29. The method according to Embodiment 28, wherein strand penetration is initiated at 27°C to 47°C.
[0043] Embodiment 30. The method according to Embodiment 29, wherein strand penetration is initiated at 32°C to 42°C.
[0044] Embodiment 31. The method according to Embodiment 30, wherein strand penetration is initiated at 37°C.
[0045] Embodiment 32. The method according to any one of Embodiments 28, wherein the fragmentation is performed at 45°C to 65°C.
[0046] Embodiment 33. The method according to any one of Embodiments 32, wherein the fragmentation is performed at 50°C to 60°C.
[0047] Embodiment 34. The method according to any one of Embodiments 33, wherein the fragmentation is performed at 55°C.
[0048] Embodiment 35. The method according to any one of Embodiments 17 to 34, wherein the transposase cofactor is added to the transpososome complex after the initiation of entry and before fragmentation.
[0049] Embodiment 36. Cofactor is Mg ++ The method according to embodiment 35.
[0050] Embodiment 37. Mg ++ The method according to Embodiment 36, wherein the concentration is 10 mM to 18 mM.
[0051] Embodiment 38. The method according to any one of Embodiments 17 to 37, wherein the fragmentation occurs within 40, 30, 20, 15, 10, or 5 bases from one or more target sequences in the nucleic acid sequence to which the targeted oligonucleotide is bound.
[0052] Embodiment 39. The method according to any one of Embodiments 17 to 38, further comprising treating a plurality of 5'-tagged target fragments with polymerase and ligase to extend and ligate the chains to produce a fully double-stranded tagged fragment.
[0053] The method according to any one of Embodiments 17 to 39, further comprising sequencing one or more of the 5'-tagged target fragments or fully double-stranded tagged fragments.
[0054] Embodiment 41. A method for preserving continuity information when sequencing a target nucleic acid, comprising: producing a tagged fragment of a target nucleic acid according to the method of any one of Embodiments 17 to 40; sequencing the 5' tagged fragment or a fully double-stranded tagged fragment to provide the fragment's sequence; grouping the fragment's sequences that contain the same targeting oligonucleotide sequence; and determining that, if they contain the same targeting oligonucleotide sequence, the group of sequences were in close proximity within the target nucleic acid.
[0055] Embodiment 42. A method for preserving continuity information when sequencing a target nucleic acid, comprising generating tagged fragments of a target nucleic acid according to the method of any one of Embodiments 17 to 40, wherein one or more adapter sequences include a unique molecular identifier (UMI) associated with a single targeted oligonucleotide sequence; sequencing 5' tagged fragments or fully double-stranded tagged fragments to provide a sequence of fragments; grouping fragment sequences containing the same UMI sequence; and determining that, if they contain the same UMI sequence, the group of sequences were in close proximity within the target nucleic acid.
[0056] Embodiment 43. A method for targeted generation of 5'-tagged fragments of nucleic acids, comprising: hybridizing one or more targeted oligonucleotides to a sample containing a single-stranded nucleic acid, wherein each of the one or more targeted oligonucleotides can bind to a target sequence in the nucleic acid; applying a transposome complex comprising a transposase; a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence; a second transposon comprising a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence; and fragmenting the nucleic acid into multiple fragments by the transposase, thereby ligating the 3' end of the first transposon to the 5' end of the fragments to produce multiple 5'-tagged fragments.
[0057] Embodiment 44. The method according to Embodiment 43, wherein double-stranded DNA is denatured to produce single-stranded DNA.
[0058] Embodiment 45. The method according to any one of Embodiments 43 to 44, wherein a region of a double-stranded nucleic acid that can be fragmented is generated by hybridizing a targeted oligonucleotide to a sample containing a single-stranded nucleic acid.
[0059] Embodiment 46. The method according to any one of Embodiments 43 to 45, wherein two or more targeted oligonucleotides having different sequences are hybridized.
[0060] Embodiment 47. The method according to any one of Embodiments 43 to 45, wherein multiple copies of a single targeted oligonucleotide are hybridized.
[0061] Embodiment 48. The method according to Embodiment 47, wherein a single targeted oligonucleotide is long enough to allow the binding of two transposomal complexes to a double-stranded nucleic acid produced by hybridizing the single targeted oligonucleotide to a sample containing a single-stranded nucleic acid.
[0062] Embodiment 49. The method according to Embodiment 47 or Embodiment 48, wherein the single targeted oligonucleotide comprises 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 base pairs.
[0063] Embodiment 50. The method according to any one of Embodiments 43 to 49, wherein the fragmentation occurs within one or more target sequences in a nucleic acid sequence that is bound by one or more targeted oligonucleotides.
[0064] Embodiment 51. The method according to any one of Embodiments 43 to 50, further comprising treating a plurality of 5'-tagged target fragments with polymerase and ligase to extend and ligate the chains to produce a fully double-stranded tagged fragment.
[0065] Embodiment 52.5': The method according to any one of Embodiments 43 to 51, further comprising sequencing one or more of a 5'-tagged target fragment or a fully double-stranded tagged fragment.
[0066] Embodiment 53. A targeted transposome complex comprising a first transposon comprising a transposase, a 3' transposon terminal sequence, and a 5' adapter sequence; a catalytically inactive endonuclease associated with a guide RNA, wherein the guide RNA can lead to the binding of the endonuclease to one or more target nucleic acid sequences; and a second transposon comprising a complement of the transposon terminal sequence.
[0067] Embodiment 54. The transposome complex according to Embodiment 53, wherein a catalytically inactive endonuclease binds to the nucleic acid but does not initiate cleavage.
[0068] Embodiment 55. The transposome complex according to Embodiment 53 or Embodiment 54, wherein the guide RNA is a single guide RNA.
[0069] Embodiment 56. A transposome complex according to any one of Embodiments 53 to 55, wherein a catalytically inactive endonuclease is associated with a transposase.
[0070] Embodiment 57. The transposome complex according to Embodiment 56, wherein a catalytically inactive endonuclease is linked to the transposase.
[0071] Embodiment 58. A transposome complex according to any one of Embodiments 53 to 57, wherein the transposase and a catalytically inactive endonuclease are included in the CRISPR-related transposase.
[0072] Embodiment 59. The CRISPR-related transposase is derived from the cyanobacterium Scytonema hofmanni (ShCAST), and is optionally selected. a. ShCAST can be coupled to a guide RNA, and at least one of the gRNA and transposase can be selectively biotinylated. The biotinylated gRNA and transposase can then be coupled to streptavidin-coated beads. b.ShCAST contains Cas12K, c. The transposase comprises a Tn5 or Tn7-like transposase, and optionally the first transposon comprises at least one of a P5 adapter and a P7 adapter, the transposome complex according to Embodiment 58.
[0073] Embodiment 60. The transposome complex according to Embodiment 57, wherein a catalytically inactive endonuclease is ligated to the 5' end of the transposase.
[0074] Embodiment 61. The transposome complex according to Embodiment 57, wherein a catalytically inactive endonuclease is ligated to the 3' end of the transposase.
[0075] Embodiment 62. The transposome complex according to Embodiment 57, wherein the transposase is ligated to the 5' end of a catalytically inactive endonuclease.
[0076] Embodiment 63. The transposome complex according to Embodiment 57, wherein the transposase is ligated to the 3' end of a catalytically inactive endonuclease.
[0077] Embodiment 64. A transposome complex according to any one of Embodiments 53 to 63, wherein a catalytically inactive endonuclease and a transposase are contained in the fusion protein.
[0078] Embodiment 65. The transposome complex according to Embodiment 64, wherein a catalytically inert and a transposase are linked via a linker.
[0079] Embodiment 66. A transposome complex according to any one of Embodiments 53 to 56, wherein a catalytically inactive endonuclease and a transposase are contained in separate proteins.
[0080] Embodiment 67. The transposome complex according to Embodiment 66, wherein separate catalytically inactive endonucleases and transposases can associate with each other via binding partner pairing, the first binding partner binding to the catalytically inactive endonuclease and the second binding partner binding to the transposase.
[0081] Embodiment 68. The transposome complex according to Embodiment 67, wherein the binding partners are biotin and streptavidin / avidin.
[0082] Embodiment 69. A transposome complex according to any one of Embodiments 55 to 68, wherein the single guide RNA is contained in an oligonucleotide comprising a first and / or second transposon.
[0083] Embodiment 70. The transposome complex according to Embodiment 69, wherein the oligonucleotide comprises a 5' single guide RNA and a 3' first and / or second transposon.
[0084] Embodiment 71. A transposome complex according to any one of Embodiments 53 to 70, wherein the single guide RNA contains fewer than 20 nucleotides.
[0085] Embodiment 72. The transposome complex according to Embodiment 71, wherein the single guide RNA sequence comprises 15, 16, 17, 18, or 19 nucleotides.
[0086] Embodiment 73. A transposome complex according to any one of Embodiments 53 to 72, wherein the single guide RNA includes a hairpin secondary structure.
[0087] Embodiment 74. The transposome complex according to any one of Embodiments 53 to 73, wherein the catalytically inactive endonuclease is a Cas9 protein.
[0088] Embodiment 75. The transposome complex according to Embodiment 74, wherein the Cas9 protein is Cas9 of Streptococcus canis.
[0089] Embodiment 76. A transposome complex according to any one of Embodiments 53 to 75, wherein Cas9 of Streptococcus canis has minimal sequence constraints.
[0090] Embodiment 77. A targeted transposome complex comprising a transposase, a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence, a zinc finger DNA-binding domain to which the zinc finger DNA-binding domain can bind to one or more target nucleic acid sequences, and a second transposon comprising a complement of the transposon terminal sequence.
[0091] Embodiment 78. The targeted transposome complex according to Embodiment 77, wherein the zinc finger DNA-binding domain is contained in a zinc finger nuclease.
[0092] Embodiment 79. The targeted transposome complex according to Embodiment 78, wherein the zinc finger nuclease is catalytically inactive.
[0093] Embodiment 80. A targeted transposome complex according to any one of Embodiments 77 to 79, wherein one or more target nucleic acid sequences are contained in histone-associated DNA.
[0094] Embodiment 81. The targeted transposome complex of Embodiment 80, wherein the DNA associated with histones is cell-free DNA.
[0095] Embodiment 82. A targeted transposome complex according to any one of Embodiments 77 to 81, wherein the first transposon comprises an affinity element.
[0096] Embodiment 83. The targeted transposome complex according to Embodiment 82, wherein the affinity element is bound to the 5' end of the first transposon.
[0097] Embodiment 84. A targeted transposome complex according to any one of Embodiments 82 to 83, wherein the first transposon comprises a linker.
[0098] Embodiment 85. The targeted transposome complex according to Embodiment 84, wherein the linker has a first end bound to the 5' end of the first transposon and a second end bound to an affinity element.
[0099] Embodiment 86. A targeted transposome complex according to any one of Embodiments 77 to 85, wherein the second transposon comprises an affinity element.
[0100] Embodiment 87. The targeted transposome complex according to Embodiment 86, wherein the affinity element is bound to the 3' end of the second transposon.
[0101] Embodiment 88. A targeted transposome complex according to any one of Embodiments 82 to 85, wherein the second transposon comprises a linker.
[0102] Embodiment 89. The targeted transposome complex according to Embodiment 88, wherein the linker has a first end bound to the 3' end of the second transposon and a second end bound to an affinity element.
[0103] Embodiment 90. A targeted transposome complex according to any one of Embodiments 82 to 89, wherein the affinity element is biotin.
[0104] Embodiment 91. A targeted transposome complex according to Embodiments 77-90, wherein the complex comprises a zinc finger DNA-binding domain array.
[0105] Embodiment 92. The transposome complex according to Embodiments 77-91, wherein the zinc finger DNA-binding domain is associated with the transposase.
[0106] Embodiment 93. The transposome complex according to Embodiment 92, wherein the zinc finger DNA-binding domain is linked to the transposase.
[0107] Embodiment 94. The transposome complex according to Embodiment 93, wherein the zinc finger DNA-binding domain is ligated to the 5' end of the transposase.
[0108] Embodiment 95. The transposome complex according to Embodiment 93, wherein the zinc finger DNA-binding domain is ligated to the 3' end of the transposase.
[0109] Embodiment 96. The transpososome complex of Embodiment 94 or 95, wherein the transposase is ligated to the 5' end of the zinc finger DNA-binding domain.
[0110] Embodiment 97. The transpososome complex of Embodiment 94 or 95, wherein the transposase is ligated to the 3' end of the zinc finger DNA-binding domain.
[0111] Embodiment 98. A transposome complex according to any one of Embodiments 77 to 97, wherein a zinc finger DNA-binding domain and a transposase are contained in the fusion protein.
[0112] Embodiment 99. A transposome complex according to any one of Embodiments 77 to 98, wherein a zinc finger DNA-binding domain and a transposase are linked via a linker.
[0113] Embodiment 100. A transposome complex according to any one of Embodiments 77 to 92, wherein the zinc finger DNA-binding domain and the transposase are contained in separate proteins.
[0114] Embodiment 101. A transposome complex according to Embodiment 100, wherein separate zinc finger DNA-binding domains and transposases can associate with each other via pairing of binding partners, the first binding partner binding to a catalytically inactive endonuclease and the second binding partner binding to the transposase.
[0115] Embodiment 102. The transposome complex according to Embodiment 101, wherein the binding partners are (i) biotin and (ii) streptavidin or avidin.
[0116] Embodiment 103. A transposome complex according to any one of Embodiments 53 to 102, wherein the adapter sequence includes a primer sequence, an index tag sequence, a capture sequence, a barcode sequence, a cleavage sequence, or a sequencing-related sequence, or a combination thereof.
[0117] Embodiment 104. The transposome complex according to Embodiments 53 to 103, wherein the adapter sequence includes a P5 or P7 sequence.
[0118] Embodiment 105. A transposome complex according to any one of Embodiments 53 to 104, wherein the transposome complex is in a solution.
[0119] Embodiment 106. The transposome complex according to any one of Embodiments 53 to 105, wherein the transposome complex is immobilized on a solid support.
[0120] Embodiment 107. The transposome complex according to Embodiment 106, wherein the solid support is a bead.
[0121] Embodiment 108. A kit or composition comprising: a first transposome complex as described in any one of Embodiments 53 to 107, which is a targeted transposome complex; a second transposome complex comprising: a first transposon comprising a transposase and a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence; and a second transposon comprising a 5' transposon terminal sequence, wherein the 5' transposon terminal sequence is complementary to the 3' transposon terminal sequence.
[0122] Embodiment 109. A kit or composition according to Embodiment 108, comprising two transposome complexes according to any one of Embodiments 53 to 107, each of which is a targeted transposome complex, wherein the two targeted transposome complexes contain different guide RNAs.
[0123] Embodiment 110. A kit or composition comprising two transposome complexes as described in either Embodiment 108 or 109, each of which is a targeted transposome complex, wherein the two targeted transposome complexes contain different zinc finger DNA-binding domains.
[0124] Embodiment 111. A method for the targeted generation of 5' tagged fragments of a target nucleic acid, comprising: combining a sample containing a double-stranded nucleic acid with a transposome complex described in any one of Embodiments 53 to 107, which is a targeted transposome complex; and fragmenting the nucleic acid into multiple fragments with a transposase, by ligating the 3' end of a first transposon to the 5' end of the fragments to produce multiple 5' tagged fragments.
[0125] Embodiment 112. A method for generating a library of tagged nucleic acid fragments, comprising: combining a sample containing a double-stranded nucleic acid with a first transpososome complex described in any one of Embodiments 53 to 107, which is a targeted transpososome complex, and a second transpososome complex comprising a transposase, a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence, and a second transposon comprising a 5' transposon terminal sequence, wherein the 5' transposon terminal sequence is complementary to the 3' transposon terminal sequence; and fragmenting the nucleic acid into a plurality of fragments by the transposase, ligating the 3' end of each first transposon to the 5' end of a target fragment to produce a plurality of first 5' tagged target fragments generated from the first transpososome complex and a plurality of second 5' tagged target fragments generated from the second transpososome complex.
[0126] Embodiment 113. A method for generating a library of tagged nucleic acid fragments, comprising: combining a sample containing a double-stranded nucleic acid with a first transposome complex described in any one of Embodiments 53 to 107, which is a targeted transposome complex, and a second transposome complex described in any one of Embodiments 53 to 107, which is a targeted transposome complex; and fragmenting the nucleic acid into a plurality of fragments using a transposase, wherein the 3' end of each first transposon is ligated to the 5' end of a target fragment to produce a plurality of first 5'-tagged target fragments generated from the first transposome complex and a plurality of second 5'-tagged target fragments generated from the second transposome complex.
[0127] Embodiment 114. The method according to any one of Embodiments 111 to 113, wherein the first and / or second targeted transposomal complex comprises a zinc finger DNA-binding domain.
[0128] Embodiment 115. The method according to Embodiment 114, wherein the zinc finger DNA-binding domain is contained in a zinc finger nuclease.
[0129] Embodiment 116. The method according to Embodiment 115, wherein the zinc finger nuclease is catalytically inactive.
[0130] Embodiment 117. The method according to any one of Embodiments 111 to 116, wherein the first transposon contained in the targeted transposomal complex includes an affinity element.
[0131] Embodiment 118. The method according to Embodiment 117, wherein the affinity element is bonded to the 5' end of the first transposon.
[0132] Embodiment 119. The method according to any one of Embodiments 118, wherein the first transposon contained in the targeted transposome complex includes a linker.
[0133] Embodiment 120. The method according to Embodiment 119, wherein the linker has a first end bonded to the 5' end of the first transposon and a second end bonded to an affinity element.
[0134] Embodiment 121. The method according to any one of Embodiments 111 to 120, wherein the second transposon includes an affinity element.
[0135] Embodiment 122. The method according to Embodiment 121, wherein the affinity element is bonded to the 3' end of the second transposon.
[0136] Embodiment 123. The method according to Embodiment 121, wherein the second transposon includes a linker.
[0137] Embodiment 124. The method according to Embodiment 123, wherein the linker has a first end bonded to the 3' end of the second transposon and a second end bonded to an affinity element.
[0138] Embodiment 125. The method according to any one of Embodiments 117 to 124, wherein the affinity element is biotin.
[0139] Embodiment 126. The method according to any one of Embodiments 111 to 125, wherein the double-stranded nucleic acid comprises DNA.
[0140] Embodiment 127. The method of Embodiment 126, wherein the DNA includes DNA associated with histones.
[0141] Embodiment 128. The method according to Embodiment 127, wherein the DNA associated with histones is cell-free DNA.
[0142] Embodiment 129. The method according to Embodiment 127 or Embodiment 128, wherein the cell-free DNA is not treated with a protease before being combined with the zinc finger DNA-binding domain.
[0143] Embodiment 130. The method according to any one of Embodiments 111 to 129, further comprising adding an affinity binding partner to a solid support after fragmentation, wherein the tagged target fragment is bound to the solid support.
[0144] Embodiment 131. The method according to Embodiment 130, wherein fragmentation is stopped before the addition of affinity elements onto a solid support.
[0145] Embodiment 132. The method according to Embodiment 131, wherein fragmentation is stopped by the addition of a solution containing proteinase K and / or SDS.
[0146] Embodiment 133. The method according to any one of Embodiments 111 to 132, comprising combining a sample containing double-stranded nucleic acid with one or more targeted transpososome complexes, wherein the sample is combined with a zinc finger DNA-binding domain or a catalytically inactive endonuclease, the zinc finger DNA-binding domain or the catalytically inactive endonuclease is bound to a first binding partner, and further comprising adding a transposase and first and second transposons, the transposase being bound to a second binding partner, the transposase being able to bind to the zinc finger DNA-binding domain or the catalytically inactive endonuclease by pairing of the first and second binding partners.
[0147] Embodiment 134. The method according to Embodiment 133, wherein the sample is combined with a zinc finger DNA-binding domain.
[0148] Embodiment 135. The method according to Embodiment 134, wherein the zinc finger DNA-binding domain is contained in a zinc finger nuclease.
[0149] Embodiment 136. The method according to Embodiment 135, wherein the zinc finger nuclease is catalytically inactive.
[0150] Embodiment 137. The method according to any one of Embodiments 133 to 136, wherein the double-stranded nucleic acid comprises DNA.
[0151] Embodiment 138. The method according to Embodiment 137, wherein the double-stranded nucleic acid comprises DNA associated with histones.
[0152] Embodiment 139. The method according to Embodiment 138, wherein the DNA associated with histones is cell-free DNA.
[0153] Embodiment 140. The method according to Embodiment 139, wherein the cell-free DNA has not been treated with a protease before being combined with the zinc finger DNA-binding domain.
[0154] Embodiment 141. The method according to any one of Embodiments 133 to 140, wherein the method includes washing after combination and before addition.
[0155] Embodiment 142. The method according to any one of Embodiments 133 to 141, wherein a targeted first transposome complex and a targeted second transposon complex bind to the reverse strand of a double-stranded nucleic acid, the first transposome complex binds to the first transposome complex binding site, and the second transposome complex binds to the second transposome complex binding site.
[0156] Embodiment 143. The method according to Embodiment 142, wherein the first 5'-tagged target fragment and the second 5'-tagged target fragment include nucleic acid sequences contained in a region of double-stranded nucleic acid between the first transposome complex binding site and the second transposome complex binding site.
[0157] Embodiment 144. The method according to Embodiment 143, wherein the first 5' tagged target fragment and the second 5' tagged fragment are at least partially complementary.
[0158] Embodiment 145. The method according to any one of Embodiments 133 to 144, wherein the transposome complex is stoichiometrically equivalent to the target DNA.
[0159] Embodiment 146. The method according to any one of Embodiments 133 to 145, wherein a divalent cation is not present in the combination.
[0160] Embodiment 147.Ca 2+ and / or Mn 2+ However, the method according to any one of embodiments 133 to 145 present in the combination.
[0161] Embodiment 148. The method according to any one of Embodiments 133 to 145, further comprising adding one or more divalent cations to the sample after combination and before fragmentation.
[0162] Embodiment 149. The divalent cation is Mg 2+ The method described in Embodiment 148.
[0163] Embodiment 150. The method according to any one of Embodiments 133 to 149, further comprising treating the sample with an exonuclease after combination and before fragmentation.
[0164] Embodiment 151. After treating the sample with exonuclease and before fragmentation, Mg 2+ The method according to Embodiment 150, comprising adding [a certain substance].
[0165] Embodiment 152. The method according to any one of Embodiments 133 to 151, further comprising releasing the tagged fragment using proteinase K and / or SDS.
[0166] Embodiment 153. The method according to any one of Embodiments 111 to 152, or the kit or composition according to Embodiments 108 to 110, wherein the 5' adapter sequences contained in the first transposome complex and the second transposome complex are different.
[0167] Embodiment 154. The method according to any one of Embodiments 111 to 153, wherein the catalytically inactive endonuclease or zinc finger DNA-binding domains contained in the first transposome complex, which is a targeted transposome complex, and the second transposome complex, which is a targeted transposome complex, are different.
[0168] Embodiment 155. The method according to Embodiments 111 to 154, wherein catalytically inactive endonuclease or zinc finger DNA-binding domains of a first transposome complex, which is a targeted transposome complex, and a second transposome complex, which is a targeted transposome complex, bind to different target sequences in a given target region in a target nucleic acid.
[0169] Embodiment 156. The method according to any one of Embodiments 111 to 155, wherein the fragmentation is performed at 45°C to 65°C.
[0170] Embodiment 157. The method according to Embodiment 156, wherein fragmentation is performed at 50°C to 60°C.
[0171] Embodiment 158. The method according to any one of Embodiments 157, wherein the fragmentation is performed at 55°C.
[0172] Embodiment 159. The method according to any one of Embodiments 111 to 158, further comprising treating a plurality of 5' tagged target fragments with polymerase and ligase to extend and ligate the chains to produce a fully double-stranded tagged fragment.
[0173] The method according to any one of Embodiments 111 to 159, further comprising sequencing one or more of the 5'-tagged target fragments or fully double-stranded tagged fragments.
[0174] Embodiment 161. A method for characterizing a desired sample in a mixed pool of samples containing both desired and unwanted samples, comprising: first sequencing a library containing multiple nucleic acid samples from the mixed pool to produce sequence data from double-stranded nucleic acids, wherein each nucleic acid library includes nucleic acids derived from a single sample and a unique sample barcode for distinguishing the nucleic acids derived from the single sample from nucleic acids derived from other samples in the library; performing a selection step on the library, which includes analyzing the sequence data to identify the unique sample barcode associated with the sequence data derived from the desired sample; enriching the nucleic acid samples derived from the desired sample and / or depleting the nucleic acid samples derived from unwanted samples; and resequencing the nucleic acid library.
[0175] Embodiment 162. The method according to Embodiment 161, wherein the sample mixing pool includes a cell mixing pool, a nuclear mixing pool, or a high molecular weight DNA mixing pool.
[0176] Embodiment 163. The method according to Embodiment 161 or Embodiment 162, wherein the sample is a cell, nucleus, or high molecular weight DNA.
[0177] Embodiment 164. The method according to any one of Embodiments 161 to 163, wherein the unique sample barcode is a unique cell barcode.
[0178] Embodiment 165. The method according to any one of Embodiments 161 to 164, wherein the concentration step includes hybrid capture, capture by a catalytically inactive endonuclease, or specific amplification of a sample barcode.
[0179] Embodiment 166. The method according to Embodiment 165, wherein the unique sample barcode-specific amplification is unique sample barcode-targeted PCR amplification.
[0180] Embodiment 167. The method according to any one of Embodiments 161 to 164, wherein the depletion step includes hybrid capture, capture by a catalytically inactive endonuclease, CRISPR digestion, or cleavage by a complex containing ShCAST (Scytonema hofmanni CRISPR-associated transposase) coupled to a guide RNA (gRNA).
[0181] Embodiment 168. The method according to Embodiment 167, wherein hybrid capture includes hybridizing the hybrid capture oligonucleotide to a unique sample barcode.
[0182] Embodiment 169. The method according to Embodiment 168, wherein the hybrid captured oligonucleotide is directly or indirectly bound to a solid support.
[0183] Embodiment 170. The method according to Embodiment 169, wherein the hybrid capture oligonucleotide is bound to a solid support via a biotin-streptavidin interaction.
[0184] Embodiment 171. The method according to Embodiment 167, wherein CRISPR digestion is cleavage by a catalytically active endonuclease.
[0185] Embodiment 172. The method according to Embodiment 171, wherein the endonuclease is Cas9.
[0186] Embodiment 173. The method according to Embodiment 172, wherein Cas9 is Cas9 of Streptococcus canis.
[0187] Embodiment 174. The method according to Embodiment 173, wherein Cas9 of Streptococcus canis has minimal sequence constraints.
[0188] Embodiment 175. The method according to any one of Embodiments 171 to 174, wherein the endonuclease is a higher fidelity variant.
[0189] Embodiment 176. The method according to Embodiment 171, comprising cleavage by a complex containing ShCAST coupled to gRNA.
[0190] Embodiment 177. A transposome complex according to any one of Embodiments 171 to 176, wherein the endonuclease is contained in the fusion protein together with the FokI nuclease.
[0191] Embodiment 178. The method according to any one of Embodiments 171 to 177, wherein an endonuclease is associated with a guide RNA that binds to one or more unique sample barcodes.
[0192] Embodiment 179. The method according to Embodiment 178, wherein the guide RNA is directed to a unique sample barcode associated with the nucleic acid of an unwanted sample.
[0193] Embodiment 180. The method according to Embodiment 178, wherein the guide RNA is directed to a unique sample barcode associated with the nucleic acid of the desired sample.
[0194] Embodiment 181. A transposome complex according to any one of Embodiments 178 to 180, wherein the guide RNA is a single guide RNA.
[0195] Embodiment 182. The transposome complex according to Embodiment 181, wherein the single guide RNA contains fewer than 20 nucleotides.
[0196] Embodiment 183. The transposome complex according to Embodiment 182, wherein the single guide RNA sequence comprises 15, 16, 17, 18, or 19 nucleotides.
[0197] Embodiment 184. A transposome complex according to any one of Embodiments 178-183, wherein a single guide RNA comprises a hairpin secondary structure.
[0198] Embodiment 185. The method according to any one of Embodiments 171 to 184, wherein the endonuclease is directly or indirectly bound to a solid support.
[0199] Embodiment 186. The method according to Embodiment 185, wherein the endonuclease is bound to a solid support via a biotin-streptavidin interaction.
[0200] Embodiment 187. The method according to any one of Embodiments 161 to 186, wherein the desired sample is a rare sample present in the sample mixture pool at a concentration of 1%, 0.1%, 0.01%, 0.001%, 0.0001%, 0.00001%, 0.000001%, 0.00000001%, 0.00000001%, or 0.000000001% or less.
[0201] Embodiment 188. The method according to Embodiments 161 to 186, wherein the desired sample is a desired cell present in a mixed pool of the sample at a concentration of 1%, 0.1%, 0.01%, 0.001%, 0.0001%, 0.00001%, 0.0000001%, 0.00000001%, or 0.000000001% or less.
[0202] Embodiment 189. The method according to any one of Embodiments 161 to 188, wherein the method includes an amplification step before resequencing.
[0203] Embodiment 190. The method according to Embodiment 189, wherein the amplification step uses a universal primer.
[0204] Embodiment 191. The method according to any one of Embodiments 161 to 190, wherein the nucleic acid library is prepared by tagmentation.
[0205] Embodiment 192. The method according to any one of Embodiments 161 to 191, wherein the method includes a step of spatially separating nucleic acid samples before incorporating a unique sample barcode.
[0206] Embodiment 193. The method according to any one of Embodiments 161 to 192, wherein the method includes tagging multiple nucleic acid samples derived from a mixed pool of samples before sequencing.
[0207] Embodiment 194. The method according to any one of Embodiments 161 to 193, wherein a unique sample barcode is incorporated into each nucleic acid sample.
[0208] Embodiment 195. The method according to any one of Embodiments 161 to 194, wherein the i5 and i7 sequences are incorporated into each nucleic acid sample.
[0209] Embodiment 196. The method according to any one of Embodiments 161 to 195, wherein a universal primer is incorporated into each nucleic acid sample.
[0210] Embodiment 197. The method according to any one of Embodiments 196, wherein the universal primer is a P5 and / or P7 primer.
[0211] Embodiment 198. The method according to any one of Embodiments 161 to 197, wherein the unique sample barcode is a single continuous barcode.
[0212] Embodiment 199. The method according to any one of Embodiment 198, wherein the unique sample barcode is a plurality of discontinuous barcodes.
[0213] Embodiment 200. The method according to Embodiment 199, wherein multiple discontinuous barcodes are separated by a fixed arrangement.
[0214] Embodiment 201. The method according to any one of Embodiments 161 to 200, wherein the amplification and resequencing steps are repeated once.
[0215] Embodiment 202. The method according to any one of Embodiments 161 to 200, wherein the amplification and resequencing steps are repeated two or more times.
[0216] Embodiment 203. The method according to any one of Embodiments 161 to 202, wherein the nucleic acid is DNA.
[0217] Embodiment 204. The method according to any one of Embodiments 161 to 202, wherein the nucleic acid is RNA.
[0218] Embodiment 205. The method according to Embodiment 204, wherein the nucleic acid is rRNA.
[0219] Embodiment 206. The method according to Embodiment 205, wherein the nucleic acid is 16s rRNA.
[0220] Embodiment 207. The method according to Embodiment 205, wherein the nucleic acid is 18s rRNA.
[0221] Embodiment 208. The method according to Embodiment 203, wherein the nucleic acid is rDNA.
[0222] Embodiment 209. The method according to any one of Embodiments 161 to 208, wherein the nucleic acid is an internal transcription spacer nucleic acid.
[0223] Embodiment 210. The method according to any one of Embodiments 161 to 209, wherein the initial sequencing step does not include whole-genome sequencing, and the resequencing step includes whole-genome sequencing.
[0224] Embodiment 211. The method according to any one of Embodiments 161 to 209, wherein the initial sequencing step comprises targeted sequencing and the resequencing step comprises whole-genome sequencing.
[0225] Embodiment 212. The method according to Embodiment 211, wherein the initial sequencing step includes targeted sequencing using one or more gene-specific primers.
[0226] Embodiment 213. The method according to Embodiment 212, wherein the gene-specific primer includes a universal primer tail.
[0227] Embodiment 214. The method according to any one of Embodiments 161 to 210, wherein the initial sequencing step comprises ribosome sequencing and the resequencing step comprises whole-genome sequencing.
[0228] Embodiment 215. The method according to Embodiment 214, wherein ribosome sequencing includes 16s, 18s, or internal transcription spacer sequencing.
[0229] Embodiment 216. The method according to any one of Embodiments 161 to 215, wherein the desired sample is a cell or a nucleus.
[0230] Embodiment 217. The method according to Embodiment 216, wherein the desired sample is cells.
[0231] Embodiment 218. The method according to any one of Embodiments 161 to 217, wherein the desired sample is a cell-derived nucleus.
[0232] Embodiment 219. The method according to any one of Embodiments 161 to 217, wherein the desired sample is a human cell or a nucleus derived from a human cell.
[0233] Embodiment 220. The method according to any one of Embodiments 161 to 217, wherein the desired sample is cancer cells or nuclei derived from cancer cells.
[0234] Embodiment 221. The method according to any one of Embodiments 161 to 220, wherein the desired cells or nuclei are of a specific desired cell type or are derived therefrom.
[0235] Embodiment 222. The method according to any one of Embodiments 161 to 221, wherein the desired sample has variations compared to other samples in the pool.
[0236] Embodiment 223. The method according to any one of Embodiments 161 to 222, wherein the desired sample is cancer cells or immune cells, or derived therefrom.
[0237] Embodiment 224. The method according to Embodiment 223, wherein the desired sample is a cancer stem cell or derived therefrom.
[0238] Embodiment 225. The method according to Embodiment 223, wherein the desired sample is cancer cells in a liquid or tumor biopsy sample, or derived therefrom.
[0239] Embodiment 226. The method according to Embodiment 220, wherein the desired sample is or is derived from drug-resistant cancer cells.
[0240] Embodiment 227. The method according to Embodiment 220, wherein the desired sample is a cancer cell having at least one mutation compared to other cancer cells in the cell pool, or is derived therefrom.
[0241] Embodiment 228. The method according to any one of Embodiments 161 to 227, wherein the method is used to track the evolution of cancer.
[0242] Embodiment 229. The method according to any one of Embodiments 161 to 228, wherein the desired sample is a cell having a somatic driver mutation or is derived therefrom.
[0243] Embodiment 230. The method according to any one of Embodiments 161 to 218, wherein the method is used for metagenomics.
[0244] Embodiment 231. The method according to Embodiment 230, wherein the method is used for sequencing microorganisms derived from environmental samples.
[0245] Embodiment 232. The method according to Embodiment 231, wherein the method does not involve culturing microorganisms derived from an environmental sample.
[0246] Embodiment 233. The method according to any one of Embodiments 230 to 232, wherein the microorganism includes bacteria, fungi, archaea, algae, protozoa, or viruses.
[0247] Embodiment 234. The method according to any one of Embodiments 161 to 233, wherein the desired sample has a single nucleotide variant (SNV).
[0248] Embodiment 235. The method according to any one of Embodiments 161 to 234, wherein the desired sample has copy number variations (CNV).
[0249] Embodiment 236. The method according to any one of Embodiments 161 to 235, wherein a desired sample has a desired methylation pattern.
[0250] Embodiment 237. The method according to any one of Embodiments 161 to 236, wherein a desired sample has a desired expression pattern.
[0251] Embodiment 238. The method according to any one of Embodiments 161 to 237, wherein a desired sample has a desired epigenetic pattern.
[0252] Embodiment 239. The method according to any one of Embodiments 161 to 229 or 234 to 238, wherein the desired sample has the desired immunorecombination.
[0253] Embodiment 240. The method according to any one of Embodiments 161-229 or 234-239, wherein the method includes characterization of the TCR repertoire.
[0254] Embodiment 241. The method according to any one of Embodiments 161 to 240, wherein the desired sample has a specific species type.
[0255] Embodiment 242. The method according to any one of Embodiments 230 to 238, wherein the desired sample is a pathogen.
[0256] Embodiment 243. The method according to Embodiment 242, wherein the desired sample is a bacterium, fungus, archaea, algae, protozoa, or virus, or derived therefrom.
[0257] Embodiment 244. The method according to any one of Embodiments 161 to 243, wherein the method does not use a concentration method based on cell sorting.
[0258] Embodiment 245. The method according to Embodiment 244, wherein the method does not use FACS.
[0259] Embodiment 246. The method according to Embodiment 245, wherein the method does not use FACS based on cell size, morphology, or surface protein expression.
[0260] Embodiment 247. The method according to any one of Embodiments 161 to 246, wherein the method does not use microfluidics.
[0261] Embodiment 248. The method according to any one of Embodiments 161 to 247, wherein the method does not use whole-genome amplification.
[0262] Embodiment 249. a. ShCAST includes Cas12K, b. The transposase contains a Tn5 or Tn7-like transposase and / or The method according to Embodiment 176, wherein at least one of c.gRNA and transposase is biotinylated, and at least one of the biotinylated gRNA and transposase can be coupled to streptavidin-coated beads.
[0263] Embodiment 250. The method according to Embodiment 176 or 249, wherein the depletion of nucleic acid samples derived from unwanted samples is carried out in a fluid having conditions for restricting the binding of transposases contained in the complex to double-stranded nucleic acids.
[0264] Embodiment 251. The method according to Embodiment 250, wherein the condition for restricting the binding of the transposase contained in the complex to the double-stranded nucleic acid is a magnesium concentration of 15 mM or less.
[0265] Embodiment 252. The method according to Embodiment 250 or 251, wherein the condition for restricting the binding of the transposase contained in the complex to the double-stranded nucleic acid is a transposase concentration of 50 nM or less.
[0266] Embodiment 253. Depleting nucleic acid samples derived from unnecessary samples, a. To bind the complex to a double-stranded nucleic acid under conditions that inhibit nucleic acid binding by the transposase contained in the complex, b. The method according to Embodiment 176 or 249, comprising promoting the cleavage of nucleic acids by the complex after binding.
[0267] Embodiment 254. The method according to Embodiment 253, wherein (1) the transposase is not present during binding, and (2) the addition of the transposase promotes cleavage.
[0268] Embodiment 255. The method according to Embodiment 253, wherein (1) the transposase is at a low concentration during binding, and (2) the transposase is added to promote cleavage.
[0269] Embodiment 256. The method according to any one of Embodiments 252 to 255, comprising (1) the transposase being reversibly inactivated during binding, and (2) activating the transposase by promoting cleavage.
[0270] Embodiment 257. The method according to Embodiment 256, wherein (1) a transposase is reversibly inactivated due to the absence of one or more transposons, and (2) activating the transposase provides one or more transposons.
[0271] Embodiment 258. A composition comprising (1) a target nucleic acid comprising one or more target nucleic acid sequences, and (2) a plurality of targeted transposome complexes according to Embodiment 59, each comprising a ShCAST coupled to a gRNA, wherein the ShCAST has an amplification adapter coupled thereto, and each of the targeted transposome complexes is hybridized to a target nucleic acid sequence.
[0272] Embodiment 259. The composition according to Embodiment 258, wherein ShCAST comprises Cas12K and further comprises a fluid having conditions that promote hybridization of Cas12K contained in the complex to one or more sequences of interest and inhibit the binding of transposases contained in the complex.
[0273] Embodiment 260. The composition according to Embodiment 259, further comprising fluid conditions such that there is no sufficient amount of magnesium ions for transposase activity, and optionally, the magnesium concentration is 15 mM or less.
[0274] Embodiment 261. The composition according to Embodiment 258, comprising a fluid having conditions that promote the activity of a transposase, wherein the transposase can add an amplification adapter at a position in a target nucleic acid.
[0275] Embodiment 262. The composition according to Embodiment 261, wherein the fluid conditions include the presence of a sufficient amount of magnesium ions for transposase activity, and optionally, the magnesium concentration is 15 mM or higher.
[0276] Embodiment 263. The composition according to any one of Embodiments 258 to 262, wherein ShCAST contains Cas12K.
[0277] Embodiment 264. The composition according to any one of Embodiments 258 to 263, wherein the transposase comprises a Tn5 or Tn7-like transposase.
[0278] Embodiment 265. The composition according to any one of Embodiments 258 to 264, wherein the adapter includes at least one of a P5 adapter and a P7 adapter.
[0279] Embodiment 266. The composition according to any one of Embodiments 258 to 265, wherein the target nucleic acid comprises double-stranded DNA.
[0280] Embodiment 267. The composition according to any one of Embodiments 258 to 266, wherein at least one of the gRNA and transposase is biotinylated, and the composition further comprises beads coated with streptavidin to which at least one of the biotinylated gRNA and transposase is coupled.
[0281] Embodiment 268. The method according to any one of Embodiments 111 to 113, wherein the first and / or second targeted transposome complex comprises the targeted transposome complex described in Embodiment 59.
[0282] Embodiment 269. The method according to Embodiment 268, wherein the method is carried out in a fluid having conditions for restricting the binding of transposases contained in the complex.
[0283] Embodiment 270. The method of Embodiment 269, wherein the condition for limiting binding of the transposase comprised in the complex is a magnesium concentration of 15 mM or less.
[0284] Embodiment 271. The method of Embodiment 269 or 270, wherein the condition for limiting binding of the transposase comprised in the complex is a transposase concentration of 50 nM or less.
[0285] Embodiment 272. The method comprises a. binding the complex to a double-stranded nucleic acid under a condition that inhibits binding of the double-stranded nucleic acid by the transposase comprised in the complex; and b. after the binding, promoting cleavage of the double-stranded nucleic acid by the complex; the method according to Embodiment 268, comprising the foregoing.
[0286] Embodiment 273. The method of Embodiment 272, wherein (1) no transposase is present during the binding, and (2) promoting cleavage comprises adding a transposase.
[0287] Embodiment 274. The method according to any one of Embodiments 271 to 273, wherein (1) the transposase is at a low concentration during binding, and (2) promoting cleavage comprises adding a transposase.
[0288] Embodiment 275. The method according to any one of Embodiments 271 to 274, wherein (1) the transposase is reversibly inactivated during binding, and (2) promoting cleavage comprises activating the transposase.
[0289] Embodiment 276. The method of Embodiment 275, wherein (1) the transposase is reversibly inactivated due to the lack of one or more transposons, and (2) activating the transposase comprises providing one or more transposons.
[0290] Embodiment 277. The method according to any one of Embodiments 268 to 276, wherein the transposase adds an amplification adapter to a position in the double-stranded nucleic acid.
[0291] Additional objectives and benefits are partially described below, some of which are evident from the description or can be learned through practice. These objectives and benefits will be realized and achieved by the elements and combinations specifically indicated in the attached claims.
[0292] Please understand that the general explanation above and the detailed explanation below are illustrative and explanatory, and do not limit the scope of the claims.
[0293] The accompanying drawings incorporated herein and constituting part of this specification illustrate one (or several) embodiments and serve to illustrate the principles described herein together with the specification. [Brief explanation of the drawing]
[0294] [Figure 1] This document provides an exemplary sample population that can be used with this method. In metagenomics samples, the target rare sample may be a bacterium expressing the presence of a specific plasmid (shaded inset) or a rare virus (black inset) within the sample. In oncology samples, the target rare sample may be a cell expressing a somatic driver mutation (inset). Generally, since the data from abundant samples is overwhelmingly dominant in sequencing results, it can be difficult to evaluate data from these rare samples. [Figure 2]A typical method for metagenomics use is described. A single-cell library (sc library) containing multiple libraries derived from a single cell is generated. Using the method of the present invention, fragments in each single-cell-derived library are uniquely tagged, for example, with a unique cell barcode (UBC). After initial sequencing to identify UBCs associated with the desired sample (e.g., from the target rare cell), selection and resequencing of the desired sample are performed. This method avoids the loss or overwhelmance of data from the target cell by the large amount of sequence data generated from a rich sample. Without the quality control method of the present invention, the target rare sample may be lost from bioinformatics analysis. [Figure 3] This document describes typical methods for sorting and selection of libraries derived from rare single cells based on sequencing. After the library is constructed, initial sequencing (e.g., 16s sequencing) can be performed to determine the desired samples. These desired samples may be libraries generated from rare cells within the entire single-cell population. Then, based on the UBC associated with the desired single-cell library fragments, selection of the desired samples is performed by either enrichment or depletion. Selection can be carried out by several different means, for example, using unique sample barcode-specific PCR, hybrid capture, or capture with catalytically inactive Cas9. After selecting the desired samples, comprehensive sequencing can be performed to better understand the characteristics of the rare cells of interest. [Figure 4] This document describes a selection method for use with libraries generated from mixed populations via the Sci-RNA3 method. A similar method may be used with libraries generated by other means. [Figure 5] This document demonstrates a method for generating a library and obtaining sequential barcodes using a modified SCI-seq method. [Figure 6] This document describes a method for generating a library using a synthetic ligated DNA library constructed with physically addressable barcodes. [Figure 7] This document describes a method for performing initial targeting sequencing. [Figure 8] This paper describes various methods for increasing the specificity of endonucleases (such as Cas9) that can be used for selection. [Figure 9] This paper provides an overview of recombinase-mediated targeted transposition. Targeted oligonucleotides coated with recombinase (Rec) can bind to targeted genomic DNA. The recombinase mediates strand entry to localize the transposome to the region of interest. Subsequent transposition allows for the insertion of a P5 / P7 sequence into the genomic DNA, subsequently generating a fragment of the region of interest. [Figure 10] This outlines targeted transposition based on targeted oligonucleotides. Single-stranded genomic target DNA can be denatured, and then the targeted oligonucleotide can hybridize (hyb) to one or more target nucleic acid sequences within the single-stranded DNA (ssDNA). Transposases and transposons can then be added. Since transposases bind to regions of double-stranded nucleic acids, the transposition is targeted to the region to which the targeted oligonucleotide is bound. In contrast, transposases do not bind to other regions of ssDNA. The transposition allows for the insertion of a P5 / P7 sequence into genomic DNA, and subsequently, fragments of the target region can be generated. [Figure 11]A method for generating a library is described using a targeted transpososome complex containing a fusion protein of a catalytically inactive endonuclease (in this embodiment, inactivated or dCas9) linked to a transposase (Tn5 in this embodiment). A single guide RNA (sgRNA) associated with dCas9 targets the fusion protein and binds to a specific nucleotide sequence within the target nucleic acid. This binding can be performed under conditions where dCas9 binding is active but the transposase is inactive (e.g., in the presence of Ca2+ and / or Mn2+). After binding of the fusion protein, transposase-mediated tagmentation can be activated with Mg2+ to generate tagged library fragments using a protocol similar to that for Nextera preparations. The resulting fragments can then be sequenced. [Figure 12A] Various means for generating a targeted transpososome complex containing catalytically inactive endonucleases and transposases are presented. The targeted transpososome complex may include a fusion protein in which the endonuclease and transposase are expressed as a single protein (A). This fusion protein may include a linker between the endonuclease and transposase. Alternatively, a binding pair (such as streptavidin and biotin) may be used to associate the transposase and endonuclease (B). In any embodiment described herein, the truncated guide RNA may be truncated, such as containing 17 nucleotides (e.g., containing fewer than 20 nucleotides), as this can increase the specificity for one or more desired sequences in the target nucleic acid. A single guide RNA (sgRNA) can associate with a transposon; for example, sgRNA associates with a transposon containing a transposon terminal sequence and Tn5 adapters, such as A14 and B15 (C). The association between sgRNA and transposon can be mediated by a complementary sequence region. Furthermore, a contiguous sgRNA transfer chain oligonucleotide (single oligonucleotide) may be used (D). [Figure 12B]Various means for generating a targeted transpososome complex containing catalytically inactive endonucleases and transposases are presented. The targeted transpososome complex may include a fusion protein in which the endonuclease and transposase are expressed as a single protein (A). This fusion protein may include a linker between the endonuclease and transposase. Alternatively, a binding pair (such as streptavidin and biotin) may be used to associate the transposase and endonuclease (B). In any embodiment described herein, the truncated guide RNA may be truncated, such as containing 17 nucleotides (e.g., containing fewer than 20 nucleotides), as this can increase the specificity for one or more desired sequences in the target nucleic acid. A single guide RNA (sgRNA) can associate with a transposon; for example, sgRNA associates with a transposon containing a transposon terminal sequence and Tn5 adapters, such as A14 and B15 (C). The association between sgRNA and transposon can be mediated by a complementary sequence region. Furthermore, a contiguous sgRNA transfer chain oligonucleotide (single oligonucleotide) may be used (D). [Figure 12C]Various means for generating a targeted transpososome complex containing catalytically inactive endonucleases and transposases are presented. The targeted transpososome complex may include a fusion protein in which the endonuclease and transposase are expressed as a single protein (A). This fusion protein may include a linker between the endonuclease and transposase. Alternatively, a binding pair (such as streptavidin and biotin) may be used to associate the transposase and endonuclease (B). In any embodiment described herein, the truncated guide RNA may be truncated, such as containing 17 nucleotides (e.g., containing fewer than 20 nucleotides), as this can increase the specificity for one or more desired sequences in the target nucleic acid. A single guide RNA (sgRNA) can associate with a transposon; for example, sgRNA associates with a transposon containing a transposon terminal sequence and Tn5 adapters, such as A14 and B15 (C). The association between sgRNA and transposon can be mediated by a complementary sequence region. Furthermore, a contiguous sgRNA transfer chain oligonucleotide (single oligonucleotide) may be used (D). [Figure 12D]Various means for generating a targeted transpososome complex containing catalytically inactive endonucleases and transposases are presented. The targeted transpososome complex may include a fusion protein in which the endonuclease and transposase are expressed as a single protein (A). This fusion protein may include a linker between the endonuclease and transposase. Alternatively, a binding pair (such as streptavidin and biotin) may be used to associate the transposase and endonuclease (B). In any embodiment described herein, the truncated guide RNA may be truncated, such as containing 17 nucleotides (e.g., containing fewer than 20 nucleotides), as this can increase the specificity for one or more desired sequences in the target nucleic acid. A single guide RNA (sgRNA) can associate with a transposon; for example, sgRNA associates with a transposon containing a transposon terminal sequence and Tn5 adapters, such as A14 and B15 (C). The association between sgRNA and transposon can be mediated by a complementary sequence region. Furthermore, a contiguous sgRNA transfer chain oligonucleotide (single oligonucleotide) may be used (D). [Figure 13] This paper describes various embodiments that can increase the specificity of targeted transposomal complexes containing catalytically inactive endonucleases. Truncate guide RNAs can increase specificity to a particular sequence of interest in the target nucleic acid, and endonucleases with minimal sequence constraints to a specific protospacer-adjacent motif (PAM) can enable a larger target design space. Hairpin secondary structures, such as toehold-blocked guide RNA, can also be used to increase specificity. [Figure 14A]This paper describes the use of a targeted transpososome complex containing a dCas9-transposase fusion protein to mediate the fragmentation of a concentrated target region. The fusion protein scans the target nucleic acid (such as DNA) to search for a target sequence that is in close proximity to the PAM and binds to the dCas9 guide RNA (A). Once the target sequence is found, dCas9 binding with high specificity can be achieved by tagmentation (e.g., contact without divalent ions, i.e., Ca2+ or Mn2+, to allow sgRNA-Cas9 binding and conformational change without enabling tagmentation by the transposase). After enabling dCas9 binding, tagmentation via a transposase (such as Tn5) is initiated by adding Mg2+. Further specificity may be possible by removing the non-Cas9 protected region of the target DNA by exonuclease treatment before adding Mg2+. After cleavage, the DNA fragments can be released by proteinase K and / or SDS. These methods can yield a high percentage of fragments in a library containing the enriched target region. After DNA release, extension and gap-filling ligation can be performed (C). [Figure 14B]This paper describes the use of a targeted transpososome complex containing a dCas9-transposase fusion protein to mediate the fragmentation of a concentrated target region. The fusion protein scans the target nucleic acid (such as DNA) to search for a target sequence that is in close proximity to the PAM and binds to the dCas9 guide RNA (A). Once the target sequence is found, dCas9 binding with high specificity can be achieved by tagmentation (e.g., contact without divalent ions, i.e., Ca2+ or Mn2+, to allow sgRNA-Cas9 binding and conformational change without enabling tagmentation by the transposase). After enabling dCas9 binding, tagmentation via a transposase (such as Tn5) is initiated by adding Mg2+. Further specificity may be possible by removing the non-Cas9 protected region of the target DNA by exonuclease treatment before adding Mg2+. After cleavage, the DNA fragments can be released by proteinase K and / or SDS. These methods can yield a high percentage of fragments in a library containing the enriched target region. After DNA release, extension and gap-filling ligation can be performed (C). [Figure 14C]Describes the use of a targeted transpososome complex comprising a fusion protein of dCas9 and transposase to mediate fragmentation of an enriched target region. The fusion protein scans a target nucleic acid (such as DNA) to find a target sequence that binds to the guide RNA of dCas9 in immediate proximity to a PAM (A). Once the target sequence is found, binding of dCas9 with high specificity can be achieved by tagmentation (for example, contacting is first performed in the absence of divalent ions, namely Ca²⁺ or Mn²⁺, to allow sgRNA-Cas9 binding and conformational change without enabling tagmentation by the transposase). After allowing dCas9 binding, tagmentation mediated by a transposase (such as Tn5) is initiated by adding Mg²⁺. Exonuclease treatment prior to Mg²⁺ addition can enable further specificity by removing non-Cas9 protected regions of the target DNA. After cleavage, DNA fragments can be released by proteinase K and / or SDS. These methods can result in a high proportion of fragments in a library comprising an enriched target region. After release of DNA, extension and gap-filling ligation can be performed (C). [Figure 15] Shows the use of zinc finger nuclease (ZNF)-associated transpososomes to generate targeted libraries from cell-free DNA (cfDNA) in plasma. A zinc finger DNA-binding domain or ZNF can target the transpososome complex to a site within cfDNA, even when the cfDNA is associated with histones. [Figure 16A] Schematically shows an exemplary composition (A) and operation in a process flow (B) for ShCAST (Scytonema hofmanni CRISPR-associated transposase) targeted library preparation and enrichment. [Figure 16B] Schematically shows an exemplary composition (A) and operation in a process flow (B) for ShCAST (Scytonema hofmanni CRISPR-associated transposase) targeted library preparation and enrichment. [Modes for carrying out the invention]
[0295] Table 2 below provides a description of the labeled components.
[0296] [Table 2]
[0297] Array description Table 1 provides a list of specific sequences referenced herein.
[0298] [Table 1] [Modes for carrying out the invention]
[0299] Various targeted transposome complexes are described herein. As used herein, “targeted transposome complex” refers to a transposome complex that is targeted to one or more desired nucleic acid sequences in a target nucleic acid.
[0300] I. Targeted transposome complexes This application describes several different targeted transposome complexes in which transpososomes are targeted to a desired nucleic acid sequence in a target nucleic acid. In some embodiments, the targeted transposome complex includes elements that can bind to one or more desired nucleic acid sequences in the target nucleic acid. Based on this binding, the targeted transposome complex can mediate transposition at a desired region in the target nucleic acid.
[0301] A targeted transposome complex may be any transposome complex having non-random binding to a target nucleic acid. Therefore, a targeted transposome complex may differ from an untargeted transposome complex that randomly binds to a sequence in the target nucleic acid. For example, a targeted transposome complex may contain elements that bind to one or more target nucleic acid sequences in the target nucleic acid. A targeted library can be generated using a method that utilizes these targeted transposome complexes, where the fragments contain the target region in the target nucleic acid.
[0302] Several different types of targeted transposomal complexes are described herein.
[0303] A. Transposome complex Generally, the transposon complex of the present invention comprises a transposase and first and second transposons, along with one or more elements that mediate targeting to one or more target nucleic acid sequences.
[0304] When used herein, a “transposomal complex” comprises at least one transposase (or other enzyme described herein) and a transposon recognition sequence. In some such systems, the transposase binds to the transposon recognition sequence to form a functional complex capable of catalyzing a transposition reaction. In some embodiments, the transposon recognition sequence is a double-stranded transposon terminal sequence. The transposase binds to a transposase recognition site in a target nucleic acid and inserts the transposon recognition sequence into the target nucleic acid. In some such insertion events, one strand of the transposon recognition sequence (or terminal sequence) is transcribed into the target nucleic acid, resulting in a cleavage event. Exemplary transposition procedures and systems can be readily adapted for use with transposases.
[0305] "Transposase" means an enzyme that forms a functional complex comprising a transposon end-containing composition (e.g., a transposon, a transposon end, or a transposon end composition) and can catalyze the insertion or rearrangement of the transposon end-containing composition into a double-stranded target nucleic acid. The transposases presented herein may also include integrases from retrotransposons and retroviruses.
[0306] Representative transposases that can be used in the specific embodiments provided herein include (or are encoded by) Tn5 transposase, Sleeping Beauty (SB) transposase, Vibrio harveyi, MuA transposase and Mu transposase recognition sites including R1 and R2 terminal sequences, Staphylococcus aureus Tn552, Ty1, Tn7 transposase, TN / O and IS10, Mariner transposase, Tc1, P Element, Tn3, bacterial insertion sequences, retroviruses, and yeast retrotransposons. More examples include IS5, Tn10, Tn903, IS911, and designed versions of transposase family enzymes. The methods described herein also include combinations of transposases, not just single transposases.
[0307] In some embodiments, the transposase is Tn5, Tn7, MuA, or Vibrioharvey transposase, or active variants thereof. In other embodiments, the transposase is Tn5 transposase or a variant thereof. In other embodiments, the transposase is Tn5 transposase or a variant thereof. In other embodiments, the transposase is Tn5 transposase or an active variant thereof. In some embodiments, the Tn5 transposase is hyperactive Tn5 transposase, or an active variant thereof. In some aspects, the Tn5 transposase is the Tn5 transposase described in International Publication No. 2015 / 160895, which is incorporated herein by reference. In some aspects, the Tn5 transposase is hyperactive Tn5 having mutations at positions 54, 56, 372, 212, 214, 251, and 338 relative to wild-type Tn5 transposase. In some embodiments, the Tn5 transposase is a highly active Tn5 having the following mutations from the wild-type Tn5 transposase: E54K, M56A, L372P, K212R, P214R, G251R, and A338V. In some embodiments, the Tn5 transposase is a fusion protein. In some embodiments, the Tn5 transposase fusion protein contains a fusion elongation factor Ts(Tsf) tag. In some embodiments, the Tn5 transposase is an overactive Tn5 transposase containing mutations at amino acids 54, 56, and 372 from the wild-type sequence. In some embodiments, the overactive Tn5 transposase is a fusion protein, and optionally, the fusion protein is the elongation factor Ts(Tsf). In some embodiments, the recognition site is a Tn5 transposase recognition site (Goryshin and Reznikoff, J. Biol. Chem., 273:7367, 1998). In one embodiment, a transposase recognition site that forms a complex with a highly active Tn5 transposase is used (e.g., EZ-Tn5™ transposase, Epicentre Biotechnologies, Madison, Wis.). In some embodiments, the Tn5 transposase is a wild-type Tn5 transposase.
[0308] When used throughout, the term “transposase” refers to an enzyme that forms a functional complex with a transposon-containing composition (e.g., transposon, transposon composition) and can catalyze the insertion or rearrangement of the transposon-containing composition into a double-stranded target nucleic acid incubated together in an in vitro rearrangement reaction. The transposases of the presented method may also include integrases derived from retrotransposons and retroviruses. Exemplary transposases that can be used in the provided method include wild-type or mutant Tn5 transposase and MuA transposase.
[0309] A transposition reaction is a reaction in which one or more transposons are inserted into a target nucleic acid at a random or nearly random site. Essential components of a transposition reaction are transposases and DNA oligonucleotides that represent the transposon's nucleotide sequence, including the transposition transposon sequence and its complement (i.e., the non-transposition transposon terminal sequence), as well as other components necessary for functional transposition or the formation of a transposon complex. The methods of this disclosure are exemplified by using transposition complexes formed by a highly active Tn5 transposase and a Tn5-type transposon terminus, or by a MuA or HYPERMu transposase and a Mu transposon terminus containing R1 and R2 terminal sequences (see, for example, Goryshin, I. and Reznikoff, WS, J. Biol. Chem., 273:7367, 1998; and Mizuuchi, Cell, 35:785, 1983; Savilahti, H, et al., EMBO J., 14:4893, 1995, which are incorporated herein by reference in their entirety). However, any transposition system capable of inserting transposon terminus in a random or near-random manner with sufficient efficiency to tag a target nucleic acid for the intended purpose may be used in the methods provided.Other examples of known transposition systems usable in the provided method include, but are not limited to, Staphylococcus aureus Tn552, Tyl, transposons Tn7, Tn / O and IS10, Mariner transposase, Tel, P Element, Tn3, bacterial insertion sequences, retroviruses, and yeast retrotransposons (e.g., Colegio OR et al, J. Bacteriol., 183:2384-8, 2001; Kirby C et al, Mol. Microbiol., 43:173-86, 2002; Devine SE, and Boeke J D., Nucleic Acids Res., 22:3765-72, 1994; International Publication No. 95 / 23875; Craig, NL, Science. 271:1512, 1996; Craig, NL, Review in: Curr Top Microbiol). Immunol.,204:27-48,1996;Kleckner N,et al.,Curr Top Microbiol Immunol.,204:49-82,1996;Lampe DJ,et al.,EMBO J.,15:5470-9,1996;Plasterk RH,Curr Top Microbiol Immunol,204:125-43,1996;Gloor,GB,Methods Mol.Biol,260:97-1 14,2004;Ichikawa H,and Ohtsubo E.,J Biol.Chem.265:18829-32,1990;Ohtsubo,F and Sekine,Y,Curr.Top.Microbiol.Immunol.204:1-26,1996;Brown PO,et al,Proc Natl Acad Sci USA, 86:2525-9, 1989; Boeke JD and Corces VG, Annu Rev Microbiol. 43:403-34, 1989; the entire article is incorporated herein by reference).
[0310] Methods for inserting transposons into target sequences can be carried out in vitro using any suitable transposon system that is available or can be developed based on the knowledge of the art. Generally, a suitable in vitro transposition system for use in the methods of the present disclosure requires at least a transposase enzyme of sufficient purity, sufficient concentration, and sufficient in vitro transposition activity, and a transposon that forms a functional complex with the respective transposase capable of catalyzing the transposition reaction. Suitable transposase transposon terminal sequences that can be used include, but are not limited to, wild-type, derivative, or mutant transposon terminal sequences that form a complex with a transposase selected from the wild-type, derivative, or mutant of the transposase.
[0311] In some embodiments, the transposase includes a Tn5 transposase. In some embodiments, the Tn5 transposase is a highly active Tn5 transposase.
[0312] In some embodiments, the transposome complex comprises a dimer of two molecules of transposase. In some embodiments, the transposome complex is homodimer, and the two molecules of transposase are each bound to a first and second transposon of the same type (for example, the sequences of the two transposons bound to each monomer are the same and form a “homodimer”). In some embodiments, the compositions and methods described herein employ two populations of transposome complexes. In some embodiments, the transposons in each population are the same. In some embodiments, the transposome complexes in each population are homodimer, and the first population has a first adapter sequence in each monomer, while the second population has different adapter sequences in each monomer.
[0313] The term “transposon terminus” refers to double-stranded nucleic acid DNA that exhibits only the nucleotide sequence (“transposon terminus sequence”) necessary to form a complex with a transposase or integrase enzyme that functions in an in vitro transposition reaction. In some embodiments, a transposon terminus can form a functional complex with a transposase in a transposition reaction. As a non-limiting example, as described in the disclosure of U.S. Patent Application Publication 2010 / 0120098, which is incorporated entirely herein by reference, a transposon terminus may include a 19-bp outer terminus (“OE”), an inner terminus (“IE”), or a “mosaic terminus” (“ME”) transposon terminus recognized by wild-type or mutant Tn5 transposase, or R1 and R2 transposon terminus. A transposon terminus may include any nucleic acid or nucleic acid analog suitable for forming a functional complex with a transposase or integrase enzyme in an in vitro transposition reaction. For example, transposon ends may contain DNA, RNA, modified bases, non-native bases, and modified backbones, and may contain nicks in one or both strands. The term “DNA” is used throughout this disclosure in relation to the composition of transposon ends, but it should be understood that any suitable nucleic acid or nucleic acid analog may be available at the transposon ends.
[0314] The term "transfer strand" refers to the transposed portion of both transposon ends. Similarly, the term "non-transfer strand" refers to the non-transfer portion of both transposon ends. The 3' end of the transfer strand is bound to or transferred to the target DNA in an in vitro transposition reaction. The non-transfer strand, which exhibits a transposon end sequence complementary to the transposed transposon end sequence, is not bound to or transferred to the target DNA in an in vitro transposition reaction.
[0315] In some embodiments, the transfer and non-transfer strands are covalently bonded. For example, in some embodiments, the transfer and non-transfer strand sequences are provided on a single oligonucleotide, for example, in a hairpin configuration. Thus, although the free end of the non-transfer strand is not directly bound to the target DNA by the transposition reaction, the non-transfer strand is indirectly attached to the DNA fragment because it is linked to the transfer strand by the loop of the hairpin structure. Further examples of transposome structures and methods for preparing and using transpososomes can be found in the disclosure of U.S. Patent Application Publication No. 2010 / 0120098, the contents of which are incorporated herein by reference in their entirety.
[0316] In some embodiments, the transposome complex includes a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence. In some embodiments, the transposome complex includes a second transposon comprising a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence.
[0317] Therefore, in some embodiments, the transposon composition includes a transposition chain having one or more other nucleotide sequences, such as adapter sequences, on the 5' side of the transposition transposon sequence. In some embodiments, the adapter sequence is a tag sequence. In addition to the transposition transposon sequence, the tag may have one or more other tag portions or tag domains.
[0318] As used herein, “tagmentation” refers to the use of transposases on fragments and tagged nucleic acids. Tagmentation involves the modification of DNA by a transposomal complex containing a transposase enzyme complexed with one or more tags (such as adapter sequences) containing transposon terminal sequences (referred to herein as transposons). Thus, tagmentation can simultaneously result in DNA fragmentation and ligation of adapters to the 5' ends of both strands of a double-stranded fragment.
[0319] Although several targeted transposome complexes are described in the present application, it is understood that some methods can use both targeted transposome complexes and non-targeted transposome complexes.
[0320] B. Immobilized Transposome Complexes In some embodiments, the transposome complex is immobilized on a solid support.
[0321] In some embodiments, the transposome complex is present at a density of at least 10 2 per 1 mm 3 , 10 4 , 10 5 , 10 6 complexes on the solid support.
[0322] In some embodiments, the length of double-stranded fragments in an immobilized library is adjusted by increasing or decreasing the density of transposome complexes on the solid support.
[0323] Many different types of immobilized transposomes can be used in these methods, as described in U.S. Patent No. 9,683,230, which is incorporated herein by reference in its entirety.
[0324] In the methods and compositions presented herein, the transposome complex is immobilized on a solid support. In some embodiments, the transposome complex and / or the capture oligonucleotide is immobilized on the support via one or more polynucleotides, such as a polynucleotide containing the transposon terminal sequence. In some embodiments, the transposome complex is immobilized via a linker molecule, which can couple a transposase enzyme to the solid support. In some embodiments, both the transposase enzyme and the polynucleotide are immobilized on the solid support. When referring to the immobilization of molecules (e.g., nucleic acids) on a solid support, the terms “immobilized” and “bound” are used interchangeably herein, and both terms are intended to encompass direct or indirect, covalent or non-covalent bonding unless otherwise indicated either explicitly or by context. In some embodiments, covalent bonding may be used, but generally, what is required is that the molecule (e.g., nucleic acid) remains immobilized or bound to the support under conditions in which the support is intended to be used, for example in applications requiring nucleic acid amplification and / or nucleic acid sequencing.
[0325] Certain embodiments may utilize a solid support comprising an inert substrate or matrix (e.g., a glass slide, polymer beads, etc.) that has been "functionalized" by applying a layer or coating of an intermediate material containing reactive groups that enable covalent bonding to biomolecules such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass, particularly the polyacrylamide hydrogels described in International Publication No. 2005 / 065814 and U.S. Patent Application Publication No. 2008 / 0280773, the contents of which are incorporated herein by reference in their entirety. In such embodiments, biomolecules (e.g., polynucleotides) may be directly covalently attached to the intermediate material (e.g., hydrogel), or the intermediate material may be noncovalently attached to the substrate or matrix (e.g., a glass substrate). The term “covalent bonding to a solid support” should be interpreted as appropriate to encompass this type of arrangement.
[0326] The terms “solid surface,” “solid support,” and other grammatical equivalents herein refer to any material that is suitable for, or can be modified to be suitable for, the binding of transposome complexes. As is understood in the art, the number of possible substrates is very large. Possible substrates include, but are not limited to, glass and modified or functionalized glass, plastics (acrylic, polystyrene, and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon®, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica, or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glass, plastics, optical fiber bundles, and various other polymers. In some embodiments, particularly useful solid supports and solid surfaces are placed within a flow cell apparatus. An exemplary flow cell is described below in further detail.
[0327] In some embodiments, the solid support includes a patterned surface suitable for immobilizing transposome complexes in a regular pattern. “Patterned surface” refers to an arrangement of different regions within or on the exposed layer of the solid support. For example, one or more regions may be features containing one or more transposome complexes. These features may be separated by interstitial regions where transposome complexes are absent. In some embodiments, the pattern may be an xy format of features in rows and columns. In some embodiments, the pattern may be a repeating arrangement of features and / or interstitial regions. In some embodiments, the pattern may be a random arrangement of features and / or interstitial regions. In some embodiments, the transposome complexes are randomly distributed on the solid support. In some embodiments, the transposome complexes are distributed on the patterned surface. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Patent Application No. 13 / 661524 or U.S. Patent Publication No. 2012 / 0316086(A1), each of which is incorporated herein by reference.
[0328] In some embodiments, the solid support comprises an array of wells or depressions on its surface. This can be processed using a variety of techniques, including but not limited to photolithography, stamping, molding, and microetching, as is generally known in the art. As is understood in the art, the technique used depends on the composition and shape of the array substrate.
[0329] The composition and geometric shape of the solid support may vary depending on its use. In some embodiments, the solid support is a planar structure such as a slide, chip, microchip, and / or array. Thus, the surface of the substrate may be in the form of a planar layer. In some embodiments, the solid support includes one or more surfaces of a flow cell. As used herein, the term “flow cell” refers to a chamber containing a solid surface through which one or more fluid reagents can flow. Examples of flow cells and associated fluid systems and detection platforms readily usable in the methods of this disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, U.S. Patent No. 7,057,026, International Publication No. 91 / 06678, International Publication No. 07 / 123744, U.S. Patent No. 7,329,492, International Publication No. 7,211,414, International Publication No. 7,315,019, International Publication No. 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.
[0330] In some embodiments, the solid support or its surface is non-planar, such as the inner or outer surface of a tube or container. In some embodiments, the solid support comprises microspheres or beads. "Microsphere," "beads," "particles," or grammatical equivalents as used herein mean small discrete particles. Suitable bead compositions include, but are not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, triasol, carbon graphite, titanium dioxide, latex, or cross-linked dextran such as Sepharose, cellulose, nylon, cross-linked micelles, and Teflon, as well as any other materials outlined herein for the solid support. Bangs Laboratories, Fishers, Ind.'s "Microsphere Selection Guide" is a useful guide. In certain embodiments, the microspheres are magnetic microspheres or beads.
[0331] The beads do not need to be spherical. Irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead size ranges from nanometers, i.e., 100 nm, to millimeters, i.e., 1 mm, and the beads are 0.2 to 200 micrometers, or 0.5 to 5 micrometers, although in some embodiments smaller or larger beads may be used.
[0332] The density of these surface-bound transpososomes can be adjusted by changing the density of the first polynucleotide or by the amount of transposase added to the solid support. For example, in some embodiments, the transposomal complexes are present on the solid support at a density of at least 10³, 10⁴, 10⁵, or 10⁶ complexes per mm².
[0333] The attachment of nucleic acids to a support, whether rigid or semi-rigid, can occur via covalent or non-covalent bonds. Exemplary bonding is described in U.S. Patents 6,737,236, 7,259,258, 7,375,234, and 7,427,678, and U.S. Patent Application Publication 2011 / 0059865(A1), each of which is incorporated herein by reference. In some embodiments, nucleic acids or other reactive elements may be bonded to a gel or other semi-solid support, which is then bonded or adhered to a solid-phase support. In such embodiments, the nucleic acids or other reactive elements are understood to be the solid phase.
[0334] In some embodiments, the solid support includes beads, a planar support, a patterned surface, or a well. In some embodiments, the planar support is the inner or outer surface of a tube.
[0335] In some embodiments, the solid support has a library of tagged DNA fragments that have been prepared and immobilized thereon.
[0336] In some embodiments, the solid support comprises a capture oligonucleotide and a first polynucleotide immobilized thereon, the first polynucleotide comprising a 3' portion containing a transposon terminal sequence and a first tag.
[0337] In some embodiments, the solid support further comprises a transposase conjugated to a first polynucleotide to form a transposomal complex.
[0338] In some embodiments, the solid support comprises a capture oligonucleotide and a second polynucleotide immobilized thereon, the second polynucleotide comprising a 3' portion containing a transposon terminal sequence and a second tag.
[0339] In some embodiments, the solid support further comprises a transposase conjugated to a second polynucleotide to form a transposomal complex.
[0340] In some embodiments, the kit includes a solid support as described herein. In some embodiments, the kit further includes a transposase. In some embodiments, the kit further includes a reverse transcription polymerase. In some embodiments, the kit further includes a second solid support for immobilizing DNA.
[0341] A wide variety of different means for immobilizing the transposome complex are described, for example, in International Publication No. 2018 / 156519, and are incorporated herein by reference. In some embodiments, the first transposon contained in the targeted transposome complex includes an affinity element. In some embodiments, the affinity element is bound to the 5' end of the first transposon. In some embodiments, the first transposon includes a linker. In some embodiments, the linker has a first end bound to the 5' end of the first transposon and a second end bound to the affinity element.
[0342] In some embodiments, the targeted transposome complex further comprises a second transposon complementary to at least a portion of the terminal sequence of the first transposon. In some embodiments, the second transposon comprises an affinity element. In some embodiments, the affinity element is bound to the 3' end of the second transposon. In some embodiments, the second transposon comprises a linker. In some embodiments, the linker has a first end bound to the 3' end of the second transposon and a second end bound to the affinity element.
[0343] In some embodiments, the affinity element is biotin.
[0344] C. Liquid-phase transposome complex The targeted transposome complex may be a liquid-phase transposome complex. These liquid-phase transposome complexes may be mobile and may not be immobilized on a solid support. In some embodiments, a liquid-phase targeted transposome complex is used to generate tagged fragments in solution.
[0345] Furthermore, the method may include a step involving a liquid-phase transposome complex. For example, the method presented herein may further include a step of providing a transposome complex in a solution and contacting the liquid-phase transposome complex with an immobilized fragment under conditions in which DNA is fragmented by the transposome complex solution, thereby obtaining an immobilized nucleic acid fragment having one end in solution. In some embodiments, the transposome complex in solution may include a second tag, and as a result, the method produces an immobilized nucleic acid fragment having a second tag, the second tag being in solution. The first and second tags may be different or the same.
[0346] In some embodiments, the method further includes the step of contacting a liquid-phase transposome complex with an immobilized DNA fragment under conditions in which the DNA fragment is further fragmented by the liquid-phase transposome complex, thereby obtaining an immobilized nucleic acid fragment having one end in solution.
[0347] In some embodiments, the liquid-phase transposome complex includes a second tag, thereby generating an immobilized nucleic acid fragment having the second tag in solution. In some embodiments, the first and second tags are different. In some embodiments, at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% of the liquid-phase transposome complex includes the second tag.
[0348] In some embodiments, one form of surface-bound transposomes is primarily present on a solid support. For example, in some embodiments, at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% of the tags present on such a solid support contain the same tag domain. In such embodiments, after the initial tagmentation reaction with surface-bound transposomes, at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% of the bridge structure contain the same tag domain at each end of the bridge. A second tagmentation reaction can be carried out by adding transposomes from a solution that further fragment the bridge. In some embodiments, most or all of the liquid-phase transposomes contain tag domains different from those present on the bridge structure generated in the first tagmentation reaction. For example, in some embodiments, at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% of the tags present in the liquid-phase transposome include tag domains different from the tag domains present on the bridge structure generated in the first tagmentation reaction.
[0349] In some embodiments, the template length is longer than what can be adequately amplified using standard cluster chemistry. For example, in some embodiments, the template length is at least 100bp, 200bp, 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp, 1000bp, 1100bp, 1200bp, 1300bp, 1400bp, 1500bp, 1600bp, 1700bp, 1800bp, 1900bp, 2000bp, 2100bp, 2200bp, 2300bp, 2400bp, 2500bp, 2600bp. These are 2700bp, 2800bp, 2900bp, 3000bp, 3100bp, 3200bp, 3300bp, 3400bp, 3500bp, 3600bp, 3700bp, 3800bp, 3900bp, 4000bp, 4100bp, 4200bp, 4300bp, 4400bp, 4500bp, 4600bp, 4700bp, 4800bp, 4900bp, 5000bp, 10000bp, 30000bp, or 100,000bp. In such embodiments, the second tagmentation reaction can be carried out by adding transposomes from a solution that further fragments the bridge, as described in U.S. Patent No. 9683230, which is incorporated entirely herein. Therefore, the second tagmentation reaction can remove the internal span of the bridge, leaving short, surface-fixed fragments that can be converted into clusters ready for further sequencing steps. In certain embodiments, the mold length may be within a range defined by upper and lower limits selected from those illustrated above.
[0350] D. Adapter and Tag In some embodiments, the first transposon includes a 3' transposon terminal sequence and a 5' adapter sequence. In some embodiments, the 5' adapter sequence is a tag sequence. Fragmentation mediated by a transposome complex including the first transposon containing the 3' transposon terminal sequence and 5' tag can be used in a method for generating a library of tagged fragments.
[0351] In some embodiments, the adapter sequence includes a primer sequence, an index tag sequence, a capture sequence, a barcode sequence, a cleavage sequence, or a sequencing-related sequence, or a combination thereof. When used herein, the sequencing-related sequence may be any sequence related to a subsequent sequencing step. The sequencing-related sequence may act to simplify downstream sequencing steps. For example, the sequencing-related sequence may be a sequence that is otherwise incorporated via a step of ligating the adapter to a nucleic acid fragment. In some embodiments, the adapter sequence includes a P5 or P7 sequence (or its complement) to facilitate binding to a flow cell in a particular sequencing method.
[0352] As used herein, the term “tag” refers to a portion or domain of a polynucleotide that indicates a sequence for a desired intended purpose or use. A tag domain may contain any sequence provided for any desired purpose. For example, in some embodiments, a tag domain contains one or more restriction endonuclease recognition sites. In some embodiments, a tag domain contains one or more regions suitable for hybridization with primers for cluster amplification reactions. In some embodiments, a tag domain contains one or more regions suitable for hybridization with primers for sequencing reactions. It will be understood that any other suitable features may be incorporated into the tag domain. In some embodiments, a tag domain contains a sequence having a length of 5 bp to 200 bp. In some embodiments, a tag domain contains a sequence having a length of 10 bp to 100 bp. In some embodiments, a tag domain contains a sequence having a length of 20 bp to 50 bp. In some embodiments, the tag domain includes an array having a length of 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 bp.
[0353] The tag may include one or more functional sequences or components (e.g., primer sequences, anchor sequences, universal sequences, spacer regions, or index tag sequences) as needed or desired.
[0354] In some embodiments, the tag includes a region for cluster amplification. In some embodiments, the tag includes a region for priming the sequencing reaction.
[0355] In some embodiments, the method further includes amplifying a fragment on a solid support by reacting a polymerase with an amplification primer corresponding to a portion of a first transposon. In some embodiments, the portion of the first transposon includes the amplification primer. In some embodiments, the tag of the first transposon includes the amplification primer.
[0356] In some embodiments, the tag includes an A14 primer sequence. In some embodiments, the tag includes a B15 primer sequence.
[0357] In some embodiments, the transposomes on individual beads have a unique index, and when a large number of such indexed beads are used, a phased transfer is produced.
[0358] Targeted transpososome complex containing targeted oligonucleotides coated with E. recombinase In some embodiments, the targeted transposome complex includes a targeted oligonucleotide. As used herein, “targeted oligonucleotide” is an oligonucleotide capable of binding to one or more target nucleic acid sequences. In some embodiments, the targeted oligonucleotide is coated with a recombinase. The targeted oligonucleotide can be used to induce the binding of the transposome complex to one or more target nucleic acid sequences within the target nucleic acid.
[0359] In some embodiments, the targeted transposome complex comprises a transposase and a targeted oligonucleotide coated with a recombinase, the targeted oligonucleotide comprising a first transposon capable of binding to one or more target nucleic acid sequences, and a second transposon comprising a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence.
[0360] 1. Targeted oligonucleotides A targeted oligonucleotide can be any type of nucleic acid having affinity for one or more target nucleic acid sequences in the target nucleic acid. In some embodiments, a targeted oligonucleotide can hybridize to the target nucleic acid based on a complementary sequence to the sequences contained in the target nucleic acid.
[0361] In some embodiments, the targeted oligonucleotide comprises nucleic acid sequences that are fully or partially complementary to one or more sequences contained in the target nucleic acid. In some embodiments, the sequence of the targeted oligonucleotide is fully or partially complementary to one or more target nucleic acid sequences.
[0362] In some embodiments, the targeted oligonucleotide is 80%, 85%, 90%, 95%, 97%, 99%, or 100% complementary to the sequence contained in the target nucleic acid.
[0363] Those skilled in the art can use a database of any number of sequences to develop targeted oligonucleotides that bind to a target nucleic acid sequence in a target nucleic acid. For example, those skilled in the art can select a target nucleic acid sequence in a given gene and develop a targeted oligonucleotide complementary to the target sequence. In this way, the transposome complex is targeted to a given gene.
[0364] In some embodiments, one or more targeted oligonucleotides are ligated to the 5' end of the adapter sequence. In some embodiments, one or more targeted oligonucleotides are directly ligated to the 5' end of the adapter sequence. In some embodiments, one or more targeted oligonucleotides are ligated to the 5' end of the adapter sequence via a linker. In some embodiments, the linker is an oligonucleotide linker. In some embodiments, the linker is a non-oligonucleotide linker. In some embodiments, both the 5' ends of the adapter sequence and the targeted oligonucleotides are biotinylated and ligated via streptavidin.
[0365] 2. Recombinase Recombinases can mediate strand entry into nucleic acids. This strand entry can be the entry of a recombinase into a double-stranded nucleic acid, such as double-stranded target DNA.
[0366] By coating targeted oligonucleotides with a recombinase, these coated oligonucleotides can mediate strand entry into double-stranded nucleic acids and subsequent binding of the targeted oligonucleotide to one or more target nucleic acid sequences. Recombinase-mediated insertion of oligonucleotides into double-stranded target nucleic acids is described in Strand Invasion Based Amplification (SIBA, see, e.g., Hoser et al. PLoS ONE 9(11):e112656). The recombinase can dissociate the double-stranded region of the double-stranded nucleic acid, enabling the binding of the targeted oligonucleotide to the single-stranded region of the target nucleic acid. As shown in Figure 9, the binding of the recombinase-coated targeted oligonucleotide can localize the transposome to the desired region in the target nucleic acid.
[0367] In some embodiments, the recombinase is UVSX, Rec233, or RecA.
[0368] F. Targeted transposomal complex containing catalytically inactive endonuclease Targeted transposomal complexes comprising a catalytically inactive endonuclease are described herein. In some embodiments, the catalytically inactive endonuclease plays a role in targeting the transposomal complex.
[0369] In some embodiments, the targeted transposomal complex includes a catalytically inactive endonuclease. As used herein, “catalystically inactive endonuclease” is an endonuclease that can bind to nucleic acids but does not mediate cleavage (this may mean that the endonuclease has no cleavage activity, or that the endonuclease has only minimal cleavage activity so that the amount of nucleic acid lost by cleavage does not substantially interfere with tagmentation). Catalytically inactive endonucleases are sometimes called inactivated endonucleases (e.g., “dCas” proteins). An exemplary catalytically inactive endonuclease is dCas9, as shown in Figure 11. Typically, endonucleases can bind to nucleic acids and mediate cleavage. Therefore, a catalytically inactive endonuclease is an endonuclease that retains nucleic acid binding function without having cleavage activity. A catalytically inactive endonuclease can be used to target a transposomal complex to one or more target nucleic acid sequences in a target nucleic acid. A typical catalytically inactive Cas9 protein is disclosed in U.S. Patent No. 1,0457,969, which is incorporated herein by reference in its entirety.
[0370] In some embodiments, the targeted transpososome complex comprises a first transposon comprising a transposase, a 3' transposon terminal sequence, and a 5' adapter sequence; a second transposon comprising a catalytically inactive endonuclease associated with a guide RNA, wherein the guide RNA can lead to the binding of the endonuclease to one or more target nucleic acid sequences; and a complement of the transposon terminal sequence.
[0371] As used herein, “guide RNA” is an RNA sequence that confers specificity to an endonuclease that binds to a target nucleic acid. A catalytically inactive endonuclease can be targeted to one or more desired nucleic acid sequences by the guide RNA.
[0372] A range of guide RNAs can be used with catalytically inactive endonucleases. In some embodiments, the guide RNA includes transactivated CRISPR RNA (tracrRNA) and CRISPR RNA (crRNA). In some embodiments, the guide RNA includes tracrRNA only. In some embodiments, the guide RNA is a single guide RNA (or sgRNA) containing both tracrRNA and crRNA.
[0373] Those skilled in the art can use one of the many available design tools (e.g., those available from Synthego or Benchling) to develop guide RNAs with specificity to bind to one or more target sequences. Guide RNA selection is also based on the presence of protospacer-adjacent motifs (PAMs) within the target nucleic acid, although endonucleases with minimal PAM specificity have been described to allow greater flexibility in the designed guide RNA (as shown in Figure 13).
[0374] As described herein, single guide RNA sequences may be included in oligonucleotides, including transposons. Such oligonucleotides can be developed using standard molecular biological techniques.
[0375] In some embodiments, the catalytically inactive endonucleases associate with the transposase. In some embodiments, the catalytically inactive endonucleases are ligated to the transposase. In some embodiments, the catalytically inactive endonucleases are ligated directly or indirectly to the transposase.
[0376] In some embodiments, the transposase and catalytically inactive endonuclease are included in the CRISPR-related transposase. As used herein, "CRISPR-related transposase" refers to a multiprotein complex comprising the endonuclease and the transposase.
[0377] Other systems in which Tn7-like transposons utilize a nuclease-deficient CRISPR-Cas system to generate CRISPR-related transposases have also been described (see Klompe et al., Nature 571:219-225 (2019)). The targeted transpososomes described herein may include any type of CRISPR-Cas system.
[0378] Catalytically inactive endonucleases can also be ligated to transposases in many different ways. In some embodiments, the catalytically inactive endonuclease is ligated to the 5' end of the transposase. In some embodiments, the catalytically inactive endonuclease is ligated to the 3' end of the transposase. In some embodiments, the transposase is ligated to the 5' end of the catalytically inactive endonuclease. In some embodiments, the transposase is ligated to the 3' end of the catalytically inactive endonuclease.
[0379] In some embodiments, catalytically inactive endonucleases and transposases are contained within a fusion protein, as shown in Figure 12A. A fusion protein means that catalytically inactive endonucleases and transposases are contained within a single protein. In some embodiments, the fusion protein containing catalytically inactive endonucleases and transposases is expressed as a single protein using a nucleic acid construct expressed by a host cell.
[0380] In some embodiments, the catalytically inert and the transposase are directly linked. In some embodiments, the catalytically inert and the transposase are linked via a linker.
[0381] In some embodiments, catalytically inactive endonucleases and transposases are contained within separate proteins. In some embodiments, catalytically inactive endonucleases and transposases are expressed as separate proteins in the host cell.
[0382] In some embodiments, separate catalytically inactive endonucleases and transposases can associate with each other through binding partner pairing, where the first binding partner binds to the catalytically inactive endonuclease and the second binding partner binds to the transposase. In some embodiments, the binding partners are biotin and streptavidin / avidin, as shown in Figure 12B.
[0383] In some embodiments, the sgRNA is contained within an oligonucleotide containing a first and / or second transposon. In some embodiments, the oligonucleotide contains a 5' single guide RNA and a 3' first and / or second transposon. In some embodiments, the sgRNA and the first and / or second transposon associate with each other via complementary sequence pairing (Figure 12C). In some embodiments, the sgRNA and the first and / or second transposon are contained within separate oligonucleotides. In some embodiments, the sgRNA is contained within a successive sgRNA transfer chain oligonucleotide (Figure 12D).
[0384] Several different methods for increasing the specificity of catalytically inactive endonucleases are shown in Figures 12A-12D and 13. Any method for increasing the specificity of catalytically inactive endonucleases can also be used to increase the specificity of catalytically active endonucleases.
[0385] In some embodiments, the single guide RNA contains fewer than 20 nucleotides (e.g., the embodiment with 17 nucleotides in Figure 12B or the embodiment with 18 nucleotides in Figure 13). Such single guide RNAs containing fewer than 20 nucleotides may be referred to as cleavage-type guide RNAs. In some embodiments, the single guide RNA sequence contains 15, 16, 17, 18, or 19 nucleotides. Shorter single guide RNAs reduce the likelihood of the single guide RNA binding to sequences in the target nucleic acid that are not completely or highly complementary to the sgRNA sequence.
[0386] In some embodiments, the single guide RNA includes a hairpin secondary structure (Kocak et al., Nat Biotechnol. 37(6):657-666 (2019)). In some embodiments, the hairpin secondary structure is used to block binding to the target nucleic acid in the absence of a trigger strand, such as a guide RNA blocked by a toehold (Siu et al. Nat Chem Biol 15(3):217-220 (2019)).
[0387] In some embodiments, the catalytically inactive endonuclease is a Cas9 protein (which may be referred to as inactivated Cas9 or dCas9). A wide variety of different Cas9 proteins may be included in the targeted transpososome complex described herein. Furthermore, those skilled in the art will recognize the catalytic domain of an endonuclease and may be able to design mutations from a wild-type endonuclease to produce a catalytically inactive endonuclease (see Maeder et al., Nat Methods 10(10):977-979 (2013)). Such designed catalytically inactive endonucleases may be tested to confirm their lack of cleavage activity.
[0388] In some embodiments, the Cas9 protein is Streptococcus canis Cas9, as shown in Figure 13. In some embodiments, Streptococcus canis Cas9 has minimal sequence constraints (see Chatterjee et al., Sci.Adv.4:eaau0766 (2018)). In some embodiments, Streptococcus canis Cas9 has a reduced need for specific protospacer facing motifs (PAMs) that are close to the sequence in the target nucleic acid that can bind to the guide RNA. For example, Streptococcus canis Cas9 may require an NNG PAM sequence instead of an NRG PAM sequence (as shown in Figure 13), which reduces the need for specific PAMs and increases the ability to select the desired sequence for binding to the guide RNA. Lower sequence constraints in endonucleases with minimal sequence constraints can allow for an improved target design space because they reduce the requirement for specific PAM sequences that are close to the desired sequence in the target nucleic acid.
[0389] In some embodiments, the CRISPR-associated transposase is derived from the cyanobacterium Scytonema hofmanni (ShCAST). ShCAST is a four-protein system for RNA-directed (sgRNA) DNA transposition mediated by a Tn7-like transposase subunit and a VK-type CRISPR effector (Cas12k) (see Strecker et al., Science. 365(6448):48-53 (2019), including embodiments shown in Strecker's Figure 5, all of which are incorporated by reference for the teaching of ShCAST). These systems, including the Tn7-like transposon, and the CRISPR-Cas system have been suggested to hijack the CRISPR effector, generate an R-loop at the target site, and facilitate the diffusion of the transposon via plasmids and phages. ShCAST can result in insertion into a specific site in a target nucleotide via an RNA-guided Tn7-like transposon. Therefore, in some embodiments, the targeted transpososome complex includes a catalytically inactive endonuclease and a transposase within ShCAST to enable targeted transposition.
[0390] 1. Targeted transposomal complex containing Cas endonuclease In some embodiments, the targeted transposomal complex contains a Cas endonuclease.
[0391] As used herein, terms such as “CRISPR-Cas system,” “Cas-gRNA ribonucleoprotein,” and “Cas-gRNA RNP” refer to an enzyme system comprising a guide RNA (gRNA) sequence containing an oligonucleotide sequence complementary or substantially complementary to a sequence in a target nucleic acid, and a Cas protein. CRISPR-Cas systems can generally be classified into three major types, which are further subdivided into 10 subtypes based on core element content and sequence; see, for example, Makarova et al., “Evolution and classification of the CRISPR-Cas systems,” Nat Rev Microbiol. 9(6):467-477 (2011). Cas proteins can have various activities, such as nuclease activity. Thus, CRISPR-Cas systems provide a mechanism for targeting a specific sequence (e.g., via gRNA), as well as a specific enzymatic activity against the sequence (e.g., via the Cas protein).
[0392] A type I CRISPR-Cas system may include Cas3 proteins having distinct helicase and DNase activities. For example, in the 1-E system, crRNAs are incorporated into a multi-subunit effector complex called the cascade (CRISPR-associated complex for antiviral defense), which binds to target DNA and induces degradation by the Cas3 protein; e.g., Brouns et al., "Small CRISPR RNAs guide antiviral defense in prokaryotes", Science 321(5891):960-964(2008); Sinkunas et al., "Cas3 is a single-stranded DNA nuclease and ATP-dependent helicase in the CRISPR-Cas immune system", EMBO J 30:1335-1342(2011); and Beloglazova et al., "Structure and activity of the Cas3 HD nuclease MJ0384, an effector enzyme of the CRISPR interference", EMBO J See 30:4616-4627 (2011). The Type II CRISPR-Cas system includes the signature Cas9 protein, a single protein (approximately 160 kDa) capable of generating crRNA to cleave target DNA. The Cas9 protein typically contains two nuclease domains: a RuvC-like nuclease domain near the amino terminus and an HNH (or McrA-like) nuclease domain near the center of the protein. Each nuclease domain of the Cas9 protein is specialized to cleave one strand of the double helix; see, for example, Jinek et al., "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Science 337(6096):816-821 (2012). The Type III CRISPR-Cas system includes the polymerase and RAMP modules.Type III systems can be further divided into subtypes III-A and III-B. Type III-A CRISPR-Cas systems have been shown to target plasmids, and the polymerase-like protein of the Type III-A system is involved in the cleavage of target DNA; see, for example, Marraffini et al., "CRISPR interference limits horizontal gene transfer in Staphylococci by targeting DNA," Science 322(5909):1843-1845 (2008). Type III-B CRISPR-Cas systems have also been shown to target RNA; see, for example, Hale et al., "RNA-guided RNA cleavage by a CRISPR-RNA-Cas protein complex," Cell 139(5):945-956 (2009). CRISPR-Cas systems include engineered and / or programmed nuclease systems derived from naturally occurring CRISPR-Cas systems. The CRISPR-Cas system may include engineered and / or mutated Cas proteins. The CRISPR-Cas system may include engineered and / or programmed guide RNAs.
[0393] In some embodiments, the Cas protein in one of the Cas-gRNA RNPs of the present invention may include Cas9, or another suitable Cas capable of cleaving the target nucleic acid at a sequence complementary to the gRNA, in the manner described in the following references, the entire contents of each of these references incorporated herein by reference: Nachmanson et al., "Targeted genome fragmentation with CRISPR / Cas9 enables fast and efficient enrichment of small genomic regions and ultra-accurate sequencing with low DNA input (CRISPR-DS)", Genome Res. 28(10):1589-1599 (2018); Vakuraskas et al., "A high-fidelity Cas9 mutant delivered as a ribonucleoprotein complex enables efficient gene editing in human hematopoietic stem and progenitor cells", Nature Medicine 24:1216-1224 (2018); Chatterjee et al., "Minimal PAM specificity of a highly similar SpCas9 "ortholog", Science Advances 4(10):eaau0766,1-10(2018); Lee et al., "CRISPR-Cap: multiplexed double-stranded DNA enrichment based on the CRISPR system", Nucleic Acids Research 47(1):1-13(2019). The Cas9-crRNA complex isolated from the S. thermophilus CRISPR-Cas system, as well as the complex assembled in vitro from separate elements, is demonstrated to bind to both synthetic oligodeoxynucleotides and plasmid DNA having nucleotide sequences complementary to crRNA.Cas9 has two nuclease domains, the RuvC and the HNH active site / nuclease domain, and these two nuclease domains have been shown to be involved in the cleavage of opposite DNA strands. In some cases, the Cas9 protein is derived from the Cas9 protein of the S. thermophilus CRISPR-Cas system. In some cases, the Cas9 protein is a multi-domain protein with approximately 1,409 amino acid residues.
[0394] In other embodiments, Cas may be manipulated so that the gRNA does not cleave the target nucleic acid at a complementary sequence, thereby preparing inactivated Cas (dCas) in the manner described in the following references, the entire contents of which are incorporated herein by reference: Guilinger et al., "Fusion of catalytically inactive Cas9 to Fokl nuclease improves the specificity of genome modification", Nature Biotechnology 32:577-582 (2014); Bhatt et al., "Targeted DNA transposition using a dCas9-transposase fusion protein", https: / / doi.org / 10.1101 / 571653, pages 1-89 (2019); Xu et al., "CRISPR-assisted targeted enrichment-sequencing (CATE-seq)", available at URL www.biorxiv.org / content / 10.1101 / 672816v1,1-30(2019); and Tijan et al., "dCas9-targeted locus-specific protein isolation method identifies histone gene regulators", PNAS 115(12):E2734-E2741(2018). Cas lacking nuclease activity may be called inactivated Cas (dCas). In some embodiments, dCas may include nuclease-deficient mutants of the Cas9 protein in which both the RuvC and HNH active site / nuclease domain are mutated. Nuclease-deficient mutants of the Cas9 protein (dCas9) bind to double-stranded DNA but do not cleave it. Another variant of the Cas9 protein has two inactivating nuclease domains, one with a first mutation in the domain that cleaves a strand complementary to crRNA, and the other with a second mutation in the domain that cleaves a strand non-complementary to crRNA.In some embodiments, the Cas9 protein has a first mutation D10A and a second mutation H840A.
[0395] In some embodiments, the Cas protein includes a cascade protein. The E. coli cascade complex recognizes double-stranded DNA (dsDNA) targets in a sequence-specific manner. The E. coli cascade complex is a 405 kDa complex containing five functionally essential CRISPR-associated (Cas) proteins (CasA1B2C6D1E1, also called cascade proteins) and a 61-nucleotide crRNA. The crRNA guides the cascade complex to the dsDNA target sequence by forming an R-loop by base-pairing with the complementary DNA strand while substituting the non-complementary strand. The cascade recognizes the target DNA without consuming ATP, suggesting that continuous invader DNA surveillance can occur without energy investment; see, for example, Matthijs et al., "Structural basis for CRISPR RNA-guided DNA recognition by Cascade," Nature Structural & Molecular Biology 18(5):529-536 (2011). In some embodiments, the Cas protein includes the Cas3 protein. Exemplarily, E. coli Cas3 can catalyze ATP-independent annealing of RNA and DNA to form R-loops and RNA base pair hybrids on double-stranded DNA. The Cas3 protein may use longer gRNAs than Cas9; see, for example, Howard et al., "Helicase disassociation and annealing of RNA-DNA hybrids by Escherichia coli Cas3 protein," Biochem J. 439(1):85-95 (2011). Such longer gRNAs may allow easier access of other elements to the target DNA, such as access of primers extended by polymerase. Another feature offered by the Cas3 protein is that it does not require a PAM sequence like Cas9, and therefore offers greater flexibility in targeting desired sequences.R-loop formation by Cas3 can utilize magnesium as a cofactor; see, for example, Howard et al., "Helicase disassociation and annealing of RNA-DNA hybrids by Escherichia coli Cas3 protein," Biochem J.439(1):85-95(2011). It will be understood that any suitable cofactor, such as a cation, may be used in conjunction with the Cas protein used in the compositions and methods of the present invention.
[0396] It should also be understood that any CRISPR-Cas system capable of disrupting double-stranded polynucleotides and creating loop structures could be used. For example, Cas proteins include, but are not limited to, those listed in the following references, and the entire contents of each of these references are incorporated herein by reference: Haft et al., "A guild of 45 CRISPR-associated (Cas) protein families and multiple CRISPR / Cas subtypes exist in prokaryotic genomes", PLoS Comput Biol. 1(6):e60, 1-10 (2005); Zhang et al., "Expanding the catalog of cas genes with metagenomes", Nucl. Acids Res. 42(4):2448-2459 (2013); and Strecker et al., "RNA-guided DNA insertion with CRISPR-associated transposases", Science 365(6448):48-53 (2019) (Cas proteins may include Cas12k). Some of these CRISPR-Cas systems can utilize specific sequences to recognize and bind to target sequences. For example, Cas9 can utilize the presence of the 5'-NGG protospacer adjacent motif (PAM).
[0397] The CRISPR-Cas system may also include an engineered and / or programmed guide RNA (gRNA). As used herein, the terms “guide RNA” and “gRNA” (which may also be referred to in the art as single guide RNA or sgRNA) are intended to mean RNA containing a sequence that is complementary or substantially complementary to a region of the target DNA sequence and that guides the Cas protein to that region. In addition to a nucleotide sequence that is complementary or substantially complementary to a region of the target DNA sequence, the guide RNA may also include a nucleotide sequence.Methods for designing gRNAs are well known in the art, and non-limiting examples are provided in the following references, each of which is incorporated herein by reference: Stevens et al., "A novel CRISPR / Cas9 associated technology for sequence-specific nucleic acid enrichment", PLoS ONE 14(4):e0215441, pages 1-7 (2019); Fu et al., "Improving CRISPR-Cas nuclease specificity using truncated guide RNAs", Nature Biotechnology 32(3):279-284 (2014); Kocak et al., "Increasing the specificity of CRISPR systems with engineered RNA secondary structures", Nature Biotechnology 37:657-666 (2019); Lee et al., "CRISPR-Cap: multiplexed double-stranded DNA enrichment based on the CRISPR system", Nucleic Acids Research 47(1):e1,1-13(2019); Quan et al., “FLASH: a next-generation CRISPR diagnostic for multiplexed detection of antimicrobial resistance sequences”, Nucleic Acids Research 47(14):e83,1-9(2019); and Xu et al., “CRISPR-assisted targeted "enrichment-sequencing(CATE-seq)", https: / / doi.org / 10.1101 / 672816,1-30(2019).
[0398] In some embodiments, the gRNA includes a chimeric CRISPR RNA (crRNA) fused to a trans-activated CRISPR RNA (tracrRNA). Such a chimeric single-guide RNA (sgRNA) is described in Jinek et al., "A programmable dual-RNA-guided endonuclease in adaptive bacterial immunity," Science 337(6096):816-821 (2012). The Cas protein can be directed by the chimeric sgRNA to any genomic locus followed by a 5'-NGG protospacer fringe motif (PAM). In one non-limiting example, the crRNA and tracrRNA can be synthesized by in vitro transcription using a synthetic double-stranded DNA template containing a T7 promoter. The tracrRNA may have a fixed sequence, but the target sequence may determine part of the crRNA sequence. Equimolar concentrations of crRNA and tracrRNA may be mixed and heated at 55°C for 30 seconds. Cas9 can be added at the same molar concentration at 37°C and incubated with the RNA mix for 10 minutes. The resulting Cas9-gRNA RNP can then be added to the target DNA in a 10-20 molar excess. Binding can occur within 15 minutes. Other suitable reaction conditions can be readily used.
[0399] 2. Targeted transposomal complex containing ShCAST In some embodiments, the targeted transposomal complex is contained within ShCAST.
[0400] Several examples herein provide compositions comprising a target nucleic acid (e.g., a double-stranded nucleic acid) containing one or more sequences of interest. The compositions may comprise a plurality of complexes, each comprising a ShCAST (Scytonema hofmanni CRISPR-associated transposase) coupled to a guide RNA (gRNA). The ShCAST may have an amplification adapter coupled thereto. Each complex may hybridize to a corresponding subsequence in the target nucleic acid (e.g., one or more sequences of interest). Such complexes are disclosed in U.S. Provisional Patent Applications 63 / 162,775 and 63 / 163,381, each of which is incorporated herein by reference in whole.
[0401] In some embodiments, the composition comprises (1) a target nucleic acid comprising one or more target nucleic acid sequences, and (2) a plurality of targeted transposome complexes described herein, each comprising a ShCAST coupled to a gRNA, wherein the ShCAST has an amplification adapter coupled thereto, and each of the targeted transposome complexes is hybridized to a target DNA sequence.
[0402] In some embodiments, ShCAST comprises a catalytically inactive endonuclease (such as Cas12K) and a transposase (such as Tn5). In some embodiments, nucleic acid cleavage by ShCAST can be considered a two-step process involving 1) binding to the nucleic acid based on the association of the catalytically inactive endonuclease with one or more gRNAs bound to a sequence of interest, and 2) cleavage by the transposase. In some embodiments, the frequency of preparation of target fragments (i.e., fragments generated from cleavage after association of the catalytically inactive endonuclease with the gRNA) is increased by limiting the nonspecific binding of the transposase to the nucleic acid.
[0403] In some embodiments, the composition further comprises a fluid having conditions that promote hybridization of the complex to the subsequence and inhibit the binding of the transposase. In some examples, the fluid conditions include the absence of a sufficient amount of magnesium ions for transposase activity.
[0404] By inhibiting transposase binding, cleavage by ShCAST is limited to the site where Cas12K in ShCAST is associated with the gRNA bound to the target sequence in the nucleic acid. In this way, nonspecific cleavage (caused by nonspecific binding of transposase to nucleic acids) is restricted, and most of the nucleic acid cleavage occurs within or near the target sequence.
[0405] In some embodiments, the conditions for limiting the binding of transposases contained in the complex are a magnesium concentration of 15 mM or less and / or a transposase concentration of 50 nM or less. Such compositions that inhibit transposase binding can work to inhibit nonspecific cleavage by transposases contained in ShCAST, most of which occur based on the binding of CasK12 to gRNA bound to the target sequence in the nucleic acid.
[0406] In some examples, the composition further comprises a fluid having conditions that promote the activity of a transposase, which attaches an amplification adapter to a position in the target nucleic acid. In some examples, the fluid conditions include the presence of a sufficient amount of magnesium ions for the activity of the transposase. Such embodiments that promote the activity of the transposase may be for preparing a fragment in or near the target sequence bound by gRNA, for example, by tagmentation. Such conditions may involve a magnesium concentration of 15 mM or higher.
[0407] In some embodiments, ShCAST comprises Cas12K. In some examples, the transposase comprises Tn5 or Tn7 transposase. In some embodiments, the adapter comprises at least one of P5 adapters and P7 adapters. In some embodiments, the target nucleic acid comprises double-stranded DNA.
[0408] In some examples, at least one of the gRNA and transposase is biotinylated. The composition may further include streptavidin-coated beads to which at least one of the biotinylated gRNA and transposase is coupled.
[0409] For example, Figures 16A and 16B schematically illustrate exemplary compositions and processes for ShCAST (Scytonema hofmanni CRISPR-associated transposase) targeted library preparation and enrichment. ShCAST6000 comprises Cas12k6001 and Tn7-like transposase 6002, which can insert DNA6003 into a specific site in the E. coli genome using RNA guide 6004. Several examples provided herein utilize ShCAST or a modified version of ShCAST incorporating Tn5 transposase (ShCAST-Tn5) for targeted amplification of specific genes. Thus, the library preparation and enrichment steps are combined, thus simplifying, improving, and facilitating automation of the targeted library sequencing workflow.
[0410] Exemplary, gRNA 6004 may be designed to target a specific gene (sequence), and the spacing of the gRNA may control the insertion size. In some examples, gRNA 6004 and / or ShCAST / ShCAST-Tn5 6002 may be coupled to tag 6005, and may be biotinylated, for example. Complex 6000 can be obtained by incorporating a transposable element having gRNA 6004 and adapter 6003 (e.g., Illumina adapter) onto the ShCAST transposase 6002 in the manner illustrated in Figure 16A. The resulting ShCAST / ShCAST-Tn5 complex 6000 may be mixed with genomic DNA (target nucleic acid) 6011 under fluid conditions that inhibit tagmentation (e.g., low magnesium or magnesium-free), while allowing the complex to bind to each sequence in the target DNA. The complex may then be isolated using a substrate coupled to a tag partner, such as streptavidin beads 6012, to which the tagged (e.g., biotinylated) gRNA and / or ShCAST / ShCAST-Tn5 will form a coupling. For example, any unbound DNA may be washed away to reduce or minimize off-target tagmentation. The fluid conditions can then be modified (e.g., by significantly increasing the magnesium) to promote tagmentation. A gap-filling ligation step followed by thermal dissociation can be used to release the library from the beads in preparation for sequencing.
[0411] In compositions and operations as illustrated in Figures 16A and 16B, it should be noted that the transposase portion 6002 of complex 6000 may be capable of random insertion into DNA. Such insertions can be inhibited or minimized by mixing the ShCAST / ShCAST-Tn5 complex with genomic DNA under fluid conditions (e.g., low-magnesium or magnesium-free) that allow the target to bind by inhibiting tagmentation.
[0412] In some embodiments, the method is designed to limit off-target tagging. In some embodiments, a low concentration of Tn5 in the targeted transposition method using ShCAST limits off-target tagging. In some embodiments, a low concentration of Tn5 limits the number of times ShCAST binds nonspecifically to nucleic acids.
[0413] In some embodiments, the gRNA targets the binding of ShCAST (and, i.e., the transposase) to one or more target loci within the target sequence, thereby allowing the user to generate a PCR product that can be amplified using both forward and reverse primers. In some embodiments, different gRNAs bind to different sequences within the target locus, i.e., different gRNAs bind to two or more target sequences within the target locus. Such target loci may be, for example, a sequence within or adjacent to the target gene.
[0414] The fragments generated using the method of the present invention require tagmentation by two transpososome complexes for the preparation of fragments that all have appropriate adapters at both ends. If the fragment is generated using one targeted transpososome complex that targets the locus of interest (by gRNA) and the other transpososome complex binds randomly, the fragment may be too large to be properly amplified using the present method. In some embodiments, if the transposase concentration is very low, it is unlikely that random binding will occur across the genome next to another Tn5 that is close enough to generate an amplified / sequencingable fragment. Alternatively, binding and cleavage by ShCAST may be performed at low temperatures (e.g., below 37°C). Therefore, fragments generated via off-target binding and tagmentation by ShCAST are unlikely to be amplified PCR products. Only when the transposases are clustered relatively close together (similar to ShCAST complexes targeted using gRNA designed to target the locus of interest) will fragments be generated that can undergo PCR enrichment.
[0415] For further details regarding ShCAST, including Cas12k and Tn7, see Strecker et al., Science. 365(6448):48-53(2019), which is incorporated herein by reference in its entirety.
[0416] Targeted transpososomes containing a G. zinc finger DNA-binding domain In some embodiments, the targeted transposome complex includes a zinc finger DNA-binding domain. This zinc finger DNA-binding domain may play a role in targeting the transposome complex to a desired sequence in the target nucleic acid.
[0417] In some embodiments, the zinc finger DNA-binding domain is designed to bind to one or more target sequences in a target nucleic acid. Methods for designing zinc finger DNA-binding domains to bind to specific sequences are well known in the art (see Wei et al., BMC Biotechnology 8:28 (2008)).
[0418] In some embodiments, the targeted transposome complex comprises a transposase, a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence, a zinc finger DNA-binding domain on which the zinc finger DNA-binding domain can bind to one or more target nucleic acid sequences, and a second transposon comprising a complement to the transposon terminal sequence.
[0419] In some embodiments, the complex includes a zinc finger DNA-binding domain array. As used herein, “zinc finger DNA-binding array” is a domain comprising two or more zinc finger DNA-binding domains.
[0420] In some embodiments, the zinc finger DNA-binding domain is associated with the transposase. In some embodiments, the zinc finger DNA-binding domain is ligated to the transposase.
[0421] In some embodiments, the zinc finger DNA-binding domain is ligated to the 5' end of the transposase. In some embodiments, the zinc finger DNA-binding domain is ligated to the 3' end of the transposase. In some embodiments, the transposase is ligated to the 5' end of the zinc finger DNA-binding domain. In some embodiments, the transposase is ligated to the 3' end of the zinc finger DNA-binding domain. In some embodiments, the zinc finger DNA-binding domain and the transposase are contained within a fusion protein.
[0422] In some embodiments, the zinc finger DNA-binding domain and the transposase are linked via a linker.
[0423] In some embodiments, the zinc finger DNA-binding domain and the transposase are contained within separate proteins. In some embodiments, the separate zinc finger DNA-binding domain and the transposase can associate with each other through pairing of binding partners, where the first binding partner binds to a catalytically inactive endonuclease and the second binding partner binds to the transposase.
[0424] II. Kits or compositions containing targeted transposomes Various kits or compositions may contain targeted transposome complexes.
[0425] In some embodiments, the kit or composition comprises a first transposome complex which is a targeted transposome complex, a transposase, a first transposon which comprises a 3' transposon terminal sequence and a 5' adapter sequence, and a second transposon which comprises a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence, and a second transposome complex which comprises a first transposon which comprises a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence.
[0426] In some embodiments, the first transposome complex, which is a targeted transposome complex, comprises a targeted oligonucleotide coated with a recombinase. In some embodiments, the kit or composition comprises two transposome complexes, each of which is a targeted transposome complex, and the two targeted transposome complexes comprises different targeted oligonucleotides.
[0427] In some embodiments, the kit or composition comprises two transposome complexes, each being a targeted transposome complex, and the two targeted transposome complexes contain different guide RNAs.
[0428] In some embodiments, the kit or composition comprises two transposome complexes, each being a targeted transposome complex, the two targeted transposome complexes comprising different zinc finger DNA-binding domains.
[0429] III. Methods using targeted transpososome complexes for targeted transposition Methods using targeted transposome complexes can mediate transpositions within a region of the target nucleic acid adjacent to the site where the targeted transposome complex binds to the target nucleic acid. In other words, targeted transposome complexes can mediate sequence-specific targeted transpositions of nucleic acids. Sequence-specific transpositions can be used to fragment the target nucleic acid and generate tagged fragments containing specific portions of the target nucleic acid. Representative methods using targeted transposome complexes are shown in Figures 14A–14C, where the targeted transposome complex includes a non-cleaving endonuclease variant such as dCas9.
[0430] Generally, transposomal complexes mediate transpositions by randomly bound double-stranded nucleic acids. However, for some applications, those skilled in the art may prefer to prepare a library containing fragments that include a desired portion of the target nucleic acid. This desired portion can be called the enriched target region, as shown in Figure 14A.
[0431] A library generated through a method that increases the probability of a library containing a fragment with a specific portion of a target nucleic acid may be called a “targeted library.” Targeted libraries can be generated using this method, which uses targeted transposome complexes. As used herein, “untargeted library” refers to a library containing random fragments of a target nucleic acid (e.g., a library generated using random fragments by a standard tagging method, etc.).
[0432] In some embodiments, when targeted transposomes are used, the frequency of transposition around the desired site in the target nucleic acid is higher. In some embodiments, the targeted library produced by this method may also contain fragments containing other parts of the target nucleic acid. In other words, the targeted library may also contain fragments containing other parts of the target nucleic acid.
[0433] In some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% of the tagged fragments included in the library of fragments generated by the method include fragments of a desired portion of the target nucleic acid.
[0434] In some embodiments, the library of fragments produced by the method using a targeted transposome complex contains 2, 5, 10, 20, 50, 100, or 1000 times more tagged fragments containing a desired portion of the target nucleic acid compared to a library not produced by a targeted transposome complex or by other enrichment methods (i.e., an untargeted or unenriched library). In some embodiments, the untargeted or unenriched library may be produced by a method using a transposome complex that randomly binds to and fragments the target nucleic acid.
[0435] In some embodiments, the library of fragments generated by this method is enriched 2-fold, 5-fold, 10-fold, 20-fold, 50-fold, 100-fold, or 1000-fold in tagged fragments containing a desired portion of the target nucleic acid. In other words, the library of fragments generated by this method using a targeted transposome complex may have a higher frequency of tagged fragments containing a desired portion of the target nucleic acid compared to the frequency of these fragments in an untargeted or unenriched library.
[0436] Targeted libraries offer several significant advantages. They focus on specific regions within target nucleic acids, generating smaller, more manageable datasets for downstream applications such as sequencing. Using targeted libraries also reduces sequencing costs and data analysis burdens, as well as shortens processing time, compared to using untargeted libraries.
[0437] Libraries containing selected regions of target nucleic acids ("targeted libraries") can be important in certain applications. Generally, methods for targeted analysis of specific genes of interest (i.e., custom content) target within genes, or mitochondrial DNA, can also be suitable for generating targeted libraries. Targeted libraries may be desired when platform output is limited or when very high coverage is required. For example, targeted libraries can enable deep sequencing at high coverage levels for the identification of rare variants.
[0438] In some embodiments, the method using targeted transposome complexes allows for the use of lower concentrations of transposome complexes with respect to the amount of target nucleic acid compared to untargeted transposome complexes. In some embodiments, the targeted transposome complexes are used at a stoichiometric ratio approximately equal to that of the target DNA.
[0439] In other words, a molar excess of targeted transposome complexes may not be necessary to generate a library containing sufficient fragments of the desired region derived from the target nucleic acid. In comparison, obtaining sufficient fragments in an untargeted library (i.e., a library generation method in which transposome complexes are not targeted to one or more target nucleic acid sequences) may require more transposome complexes because the fragments generated using the untargeted library are randomly generated. Therefore, using targeted transpososomes allows more fragments in the library to contain the desired sequence, thereby enabling the use of smaller amounts of targeted transposome complexes and smaller amounts of target nucleic acid.
[0440] The targeted transposome complexes described herein may be used in conjunction with untargeted transposome complexes. In some embodiments, a method for generating a library of tagged nucleic acid fragments includes combining a sample containing double-stranded nucleic acid with a first transpososome complex which is a targeted transpososome complex, and a second transpososome complex which includes a transposase and a second transpososome complex which includes a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence, and a second transposon comprising a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence, and fragmenting the nucleic acid into a plurality of fragments by the transposase, ligating the 3' end of each first transposon to the 5' end of a target fragment to produce a plurality of first 5' tagged target fragments generated from the first transpososome complex and a plurality of second 5' tagged target fragments generated from the second transpososome complex.
[0441] The method may also use two targeted transposome complexes.
[0442] In some embodiments, a method for generating a library of tagged nucleic acid fragments includes combining a sample containing double-stranded nucleic acid with a first transposomal complex which is a targeted transposomal complex and a second transposomal complex which is a targeted transposomal complex, and fragmenting the nucleic acid into a plurality of fragments by transposase, wherein the 3' end of each first transposon is ligated to the 5' end of a target fragment to produce a plurality of first 5'-tagged target fragments generated from the first transposomal complex and a plurality of second 5'-tagged target fragments generated from the second transposomal complex.
[0443] The targeted transposomes used in the method may be any of those described herein, such as those containing a catalytically inactive endonuclease or those containing a zinc finger DNA-binding domain.
[0444] The methods described herein can be designed to facilitate the combination of a targeted transposomal complex with a target nucleic acid prior to fragmentation. In some embodiments, agents that promote the fragmentation activity of the transposase are absent or present at low levels during the combination step. In some embodiments, divalent cations are absent during combination. In some embodiments, Ca 2+ and / or Mn 2+ It is present in the combination. In some embodiments, Ca 2+ and / or Mn 2+ It is present in the combination, but Mg 2+ It does not exist.
[0445] In some embodiments, the method further includes adding one or more divalent cations to the sample after combination and before fragmentation. In some embodiments, the divalent cation is Mg 2+ That is the case.
[0446] In some embodiments, the method further includes treating the sample with an exonuclease after combination and before fragmentation. The exonuclease can promote the degradation of single-stranded DNA. In some embodiments, the method includes treating the sample with an exonuclease and before fragmentation, Mg 2+ This further includes adding [something].
[0447] In some embodiments, the method includes releasing the tagged fragment using proteinase K and / or SDS.
[0448] The method of the present invention can be used to tag both ends of the generated fragment with an adapter. This can be achieved by using a method that utilizes a first transposome complex and a second transposome complex. In some embodiments, the method incorporates different tags into each end of the fragment generated by fragmentation. In some embodiments, the 5' adapter sequences contained in the first transposome complex and the second transposome complex are different.
[0449] A. A method using a targeted transpososome complex containing a recombinase-coated targeted oligonucleotide. In some embodiments, the method uses a targeted transposome complex containing a targeted oligonucleotide coated with a recombinase. An exemplary embodiment is shown in Figure 9.
[0450] In some embodiments, the method for the targeted generation of 5'-tagged fragments of a target nucleic acid involves combining a sample containing a double-stranded nucleic acid with a transposomal complex, which is a targeted transposomal complex. In some embodiments, the targeted transposomal complex contains a targeted oligonucleotide coated with a recombinase. In some embodiments, strand entry of the nucleic acid is initiated by the recombinase. In some embodiments, after strand entry, the nucleic acid is fragmented into multiple fragments by a transposase by ligating the 3' end of a first transposon to the 5' end of a fragment to generate multiple 5'-tagged fragments.
[0451] In some embodiments, a method for generating a library of tagged nucleic acid fragments comprises a sample containing double-stranded nucleic acid, a first transposome complex which is a targeted transposome complex containing a recombinase-coated targeted oligonucleotide, a second transposome complex which comprises a transposase, a first transposon which comprises a 3' transposon terminal sequence and a 5' adapter sequence, and a second transposon which comprises a 5' transposon terminal sequence and the 5' transposon terminal sequence is a 3' transposon A method comprising: combining a second transposon complementary to the n-terminal sequence with a second transposomal complex, initiating strand entry of nucleic acid by a recombinase, and fragmenting the nucleic acid into multiple fragments by a transposase, wherein the 3' end of each first transposon is ligated to the 5' end of the target fragment to produce multiple first 5'-tagged target fragments generated from the first transposomal complex and multiple second 5'-tagged target fragments generated from the second transposomal complex.
[0452] In some embodiments, a method for generating a library of tagged nucleic acid fragments includes combining a sample containing a double-stranded nucleic acid with a first transpososome complex which is a targeted transpososome complex containing a recombinase-coated targeted oligonucleotide and a second transpososome complex which is a targeted transpososome complex containing a recombinase-coated targeted oligonucleotide; initiating strand entry of the nucleic acid with the recombinase; and fragmenting the nucleic acid into multiple fragments with the transposase, wherein the 3' end of each first transposon is ligated to the 5' end of a target fragment to produce a plurality of first 5'-tagged target fragments generated from the first transpososome complex and a plurality of second 5'-tagged target fragments generated from the second transpososome complex.
[0453] In some embodiments, the 5' adapter sequences contained in the first transposome complex and the second transposome complex are different.
[0454] In some embodiments, the targeted oligonucleotides contained in the first and second transposome complexes are different. In some embodiments, the targeted oligonucleotides of the first and second transposome complexes bind to different target sequences within a given target region in the target nucleic acid. In this way, the first and second transposome complexes can generate fragments containing the desired target sequence. Those skilled in the art can design targeted oligonucleotides that bind at, near, or beyond the terminal of the target sequence to generate these target sequence-containing fragments. In this way, a targeted library with an increased frequency of fragments containing the target sequence can be generated.
[0455] In some embodiments, the second transposome complex binds to the reverse strand of the double-stranded nucleic acid compared to the first transposome complex.
[0456] In some embodiments, the initiation of nucleic acid strand entry by recombinase occurs in the presence of a recombinase loading factor. In some embodiments, the recombinase loading factor is removed or inactivated before fragmentation.
[0457] In some embodiments, the initiation of strand penetration occurs via substitution loop formation.
[0458] In some embodiments, strand entry begins within 40, 30, 20, 15, 10, or 5 bases of the binding site of the targeted oligonucleotide to one or more target sequences. In other words, strand entry can occur very close to the binding site of the targeted oligonucleotide.
[0459] In some embodiments, the method proceeds through different steps based on temperature changes during the method. In some embodiments, the temperature used to initiate strand entry is different from the optimal temperature for fragmentation by transposase. In some embodiments, the temperature used to initiate strand entry is below the optimal temperature for fragmentation by transposase. In some embodiments, initiating strand entry at a lower temperature facilitates proper targeting of the transposomal complex based on the recombinase-coated targeted oligonucleotide before fragmentation is initiated by rising temperature. These temperature changes may help facilitate the binding of the targeted transposomal complex to the desired sequence in the target nucleic acid before fragmentation.
[0460] In some embodiments, strand penetration begins at 27°C to 47°C. In some embodiments, strand penetration begins at 32°C to 42°C. In some embodiments, strand penetration begins at 37°C.
[0461] In some embodiments, fragmentation is carried out at 45°C to 65°C. In some embodiments, fragmentation is carried out at 50°C to 60°C. In some embodiments, fragmentation is carried out at 55°C.
[0462] In some embodiments, the initiation of strand entry occurs while the reaction solution lacks components for transposase activity. For example, in some embodiments, the transposase cofactor is added to the transpososome complex after the initiation of entry and before fragmentation. In some embodiments, the cofactor is Mg ++ In some embodiments, Mg ++ The concentration is between 10 mM and 18 mM.
[0463] A method using a targeted transpososome complex containing a targeted oligonucleotide coated in recombinase can increase the probability of fragmentation occurring near the site where the targeted oligonucleotide is bound to the target nucleic acid. In some embodiments, fragmentation occurs within 40, 30, 20, 15, 10, or 5 bases from one or more target sequences in the nucleic acid sequence bound by the targeted oligonucleotide.
[0464] B. Method using hybridization of targeted oligonucleotides to single-stranded nucleic acids Transposases can mediate the rearrangement and fragmentation of double-stranded nucleic acids. Therefore, selective generation of double-stranded nucleic acid regions via the binding of targeted oligonucleotides to single-stranded nucleic acids (e.g., single-stranded DNA) can be used in methods for generating tagged fragments. An exemplary method using targeted oligonucleotides is shown in Figure 10.
[0465] A method for the targeted generation of 5' tagged fragments of nucleic acids may involve hybridizing one or more targeted oligonucleotides to a sample containing single-stranded nucleic acids. In some embodiments, double-stranded target nucleic acids may be denatured to generate single-stranded nucleic acids. In some embodiments, double-stranded DNA is reorganized to generate single-stranded DNA. In some embodiments, denaturation is carried out by increasing the temperature. In some embodiments, the temperature of the double-stranded nucleic acid is increased to the melting temperature (T) of the nucleic acid. m Denaturation occurs by raising the temperature above 70°C. In some embodiments, a sample containing double-stranded DNA is heated to a temperature above 70°C to promote the denaturation of double-stranded DNA to single-stranded DNA. In some embodiments, double-stranded nucleic acids are treated with urea and / or pH changes to produce single-stranded DNA.
[0466] In some embodiments, hybridization of one or more targeted oligonucleotides to a sample containing single-stranded nucleic acids is performed by lowering the temperature of the sample containing single-stranded nucleic acids to enable the binding of one or more targeted oligonucleotides to the single-stranded nucleic acids.
[0467] In some embodiments, one or more targeted oligonucleotides can each bind to a target sequence in the nucleic acid. In some embodiments, the targeted oligonucleotides are fully or partially complementary to the target sequence in the nucleic acid.
[0468] In some embodiments, hybridization of one or more targeted oligonucleotides to single-stranded nucleic acids generates a double-stranded nucleic acid region. Transposases do not bind to the single-stranded nucleic acid region, but transposases can bind to the double-stranded region resulting from the hybridization of the targeted oligonucleotide to the single-stranded nucleic acid. In some embodiments, hybridization of the targeted oligonucleotide to a sample containing single-stranded nucleic acid generates a double-stranded nucleic acid region that can be fragmented.
[0469] In some embodiments, the method includes hybridizing one or more targeted oligonucleotides to a sample, followed by applying a transposome complex. In some embodiments, the transposome complex includes a transposase and a second transposome complex comprising a first transposon comprising a 3' transposon terminal sequence and a 5' adapter sequence, and a second transposon comprising a 5' transposon terminal sequence, the 5' transposon terminal sequence being complementary to the 3' transposon terminal sequence. In some embodiments, the method then includes fragmenting the nucleic acid into multiple fragments by the transposase by ligating the 3' end of the first transposon to the 5' end of the fragment to generate multiple 5' tagged fragments.
[0470] In some embodiments, two or more targeted oligonucleotides having different sequences are hybridized. In some embodiments, the method using two or more targeted oligonucleotides can mediate fragmentation at two or more sites in the target nucleic acid. For example, two or more targeted oligonucleotides may be bound at the ends of a target region in the target nucleic acid so that fragmentation generates a fragment containing the region of interest. In other words, the method using two or more targeted oligonucleotides can generate a targeted library.
[0471] In some embodiments, multiple copies of a single targeted oligonucleotide are hybridized.
[0472] In some embodiments, only one type of targeted oligonucleotide is hybridized. In this way, the target nucleic acid is fragmented within a specific region. In some embodiments, the single targeted oligonucleotide is long enough to allow the binding of two transposomal complexes to a double-stranded nucleic acid produced by hybridizing the single targeted oligonucleotide to a sample containing a single-stranded nucleic acid. In some embodiments, the single targeted oligonucleotide contains 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 base pairs.
[0473] In some embodiments, fragmentation occurs within one or more target sequences in a nucleic acid sequence that are bound by one or more targeted oligonucleotides.
[0474] How to use C.ShCAST In some embodiments, ShCAST (Scytonema hofmanni CRISPR-related transposase) targeted library preparation and enrichment may be used, as summarized in Figures 16A and 16B.
[0475] Targeted gene sequencing using a separate enrichment step after library preparation can be time-consuming. For example, such a separate enrichment step may involve hybridizing oligonucleotide probes to library DNA and isolating the hybridized DNA on streptavidin-coated beads. Although efficiency and required time have been significantly improved, such separate enrichment protocols can still require approximately two hours and numerous reagents, making automation of such protocols difficult.
[0476] In comparison, the ShCAST-based method described herein allows for the preparation and enrichment of libraries for targeted sequencing of specific genes using a single step for both preparation and enrichment.
[0477] In some embodiments, the first and / or second targeted transposome complex comprises a targeted transposome complex containing ShCAST.
[0478] In some embodiments, at least one of the gRNA and transposase is biotinylated, and the composition further comprises streptavidin-coated beads coupled with at least one of the biotinylated gRNA and transposase. In this way, tagged fragments generated using a targeted transpososome complex containing ShCAST can be immobilized on streptavidin-coated beads.
[0479] In some embodiments, some or all steps of the method are carried out in a reaction fluid that restricts or inhibits nonspecific binding of nucleic acids by the transposase contained in ShCAST. In some embodiments, restricting or inhibiting nonspecific binding of the transposase contained in ShCAST reduces off-target rearrangement reactions mediated by the transposase contained in ShCAST. Such off-target rearrangements can occur when the transposase contained in ShCAST randomly binds to the nucleic acid itself; instead, ShCAST is targeted to the sequence of interest by gRNA bound to the sequence of interest. When off-target cleavage is reduced, most fragments are generated from cleavage mediated by targeted transposomal complexes. In this way, most tagged fragments are prepared from one or more loci of interest (containing one or more sequences of interest that can bind to one or more gRNAs). Furthermore, if tagged fragments are prepared from two targeted transposomal complexes, they are likely to be of a size that can be sequenced and / or amplified. In contrast, if one or both of the transposomal complexes used to prepare the fragments are not properly targeted (for example, if the transposase in ShCAST binds directly to nucleic acids without targeting by gRNA), the fragments may be too large to amplify and / or sequence.
[0480] In some embodiments, the method is carried out in a fluid having conditions to restrict direct binding of the complex by the transposase. In some embodiments, the conditions to restrict direct binding of the complex by the transposase are a magnesium concentration of 15 mM or less, and / or a Cas12K and / or transposase concentration of 50 nM or less.
[0481] In some embodiments, different steps of the method are carried out under different conditions. In some embodiments, complex binding is carried out under conditions that inhibit the transposase's binding to double-stranded nucleic acids. In this way, direct, untargeted binding of ShCAST to nucleic acids by the transposase is limited, and most ShCASTs bind to nucleic acids based on the association of Cas12K with gRNAs that target one or more target sequences in the nucleic acid.
[0482] In some embodiments, the conditions may be modified to promote cleavage by the transposase contained in ShCAST after binding. In some embodiments, the method includes binding the complex to a double-stranded nucleic acid under conditions that inhibit the binding of the double-stranded nucleic acid by the transposase contained in the complex, and then promoting cleavage of the double-stranded nucleic acid by the complex after binding.
[0483] In some embodiments, the transposase is either absent or present at low concentrations during binding, and promoting cleavage involves adding the transposase.
[0484] In some embodiments, an activatable transposase is included in ShCAST. As used herein, “activatable transposase” is one that can be reversibly inactivated and subsequently activated. For example, a reversibly inactivated transposase may lack an element for proper cleavage of nucleic acids, which may be added during a later step in the method.
[0485] In some embodiments, the transposase is reversibly inactivated during binding, and cleavage is facilitated by activation of the transposase.
[0486] In some embodiments, a transposase is reversibly inactivated by the absence of one or more transposons, and activating a transposase involves providing one or more transposons.
[0487] In some embodiments, the transposase adds amplification adapters to positions within the double-stranded nucleic acid. As used herein, “amplification adapter” is any sequence useful for amplification (e.g., a binding site for an amplification primer). Thus, the resulting tagged fragment can be amplified without the need to incorporate additional amplification adapters. In some embodiments, the amplification adapters may be added to the fragment after the tagged fragment has been prepared (e.g., by ligation of the amplification adapters).
[0488] D. Methods including pairing of binding partners A high-resolution sequencing library can be generated when a first pair of binding partners binds to a catalytically inactive endonuclease or zinc finger DNA-binding domain, and a second binding partner binds to a transposase.
[0489] Methods involving pairing of binding partners may be similar to the cut-and-tag method (see Kaya-Okur et al., Nature Communications 10:1930 (2019)). In such methods, a catalytically inactive endonuclease or zinc finger DNA-binding domain containing a first binding partner is bound to the target nucleic acid. In some embodiments, the reaction product is washed after this binding. A transposase containing a second binding partner is then added. The transposase localizes to the catalytically inactive endonuclease or zinc finger DNA-binding domain based on the affinity of the second binding partner to the first binding partner. These methods allow the transposase to bind to a site already bound by the catalytically inactive endonuclease or zinc finger DNA-binding domain.
[0490] In some embodiments, the method is carried out under conditions that restrict the binding of catalytically inactive endonuclease or zinc finger DNA binding domains. These conditions can restrict off-target transposase binding. In some embodiments, low concentrations of magnesium or low concentrations of catalytically inactive endonuclease or zinc finger DNA binding are used to reduce off-target transposase binding. In some embodiments, the likelihood of generating an amplified PCR product from off-target binding is reduced. In some embodiments, limited off-target transposase binding means that random (i.e., untargeted) transposase binding occurs infrequently, generally resulting in fragments that are too large for amplification and / or sequencing. In contrast, the use of targeted transpososome complexes can be designed to prepare fragments of appropriate size for amplification and / or sequencing.
[0491] As used herein, the first binding partner and the second binding partner may be referred to as “tags.” In some embodiments, the first tag is coupled to a first Cas-gRNA ribonucleoprotein (including RNP, Cas and its gRNA), and the second tag is coupled to a second Cas-gRNA RNP. In some examples, the method includes coupling the first tag to a first tag partner coupled to a substrate, and coupling the second tag to a second tag partner coupled to a substrate. In some examples, the coupling is performed after the first and second Cas-gRNA RNPs have been hybridized to the first and second subsequences, respectively. In some examples, the first and amplification adapters are attached after the first and second tags have been attached to the first and second tag partners, respectively.
[0492] In some examples, the first and second tags contain biotin. In some examples, the first and second tag partners contain streptavidin. In some examples, the substrate contains beads. In some examples, the Cas-gRNA RNP contains Cas12k. In some examples, the transposase contains a Tn5 or Tn7-like transposase.
[0493] In some embodiments, combining a sample containing double-stranded nucleic acid with one or more targeted transpososome complexes includes combining the sample with a zinc finger DNA-binding domain or a catalytically inactive endonuclease, wherein the zinc finger DNA-binding domain or the catalytically inactive endonuclease is bound to a first binding partner, and adding a transposase and first and second transposons, wherein the transposase is bound to a second binding partner, and the transposase can bind to the zinc finger DNA-binding domain or the catalytically inactive endonuclease by pairing of the first and second binding partners.
[0494] In some embodiments, the method includes washing after combination and before addition. In some embodiments, cell-free DNA is not treated with protease before combination with the zinc finger DNA-binding domain.
[0495] E. Method for generating targeted fragments with two targeted transposome complexes In some embodiments, a polynucleotide (e.g., a target nucleic acid) can be cleaved at any suitable pair of positions to form a fragment. After forming the fragment using the method disclosed herein, any suitable amplification primer can be coupled to the ends of the resulting fragment. The fragment can then be amplified and sequenced.
[0496] In a method using first and second transposome complexes that are both targeted, the complexes may be designed to produce a specific desired fragment. In some embodiments, a method using first and second transposome complexes that are both targeted can generate targeted or enriched libraries. These targeted or enriched libraries may contain a higher proportion of library fragments containing enriched target regions. These enriched target regions may be, for example, the gene of interest for sequencing.
[0497] In some embodiments, the targeted first transposome complex and the targeted second transposon complex bind to the reverse strand of a double-stranded nucleic acid, with the first transposome complex binding to the first transposome complex binding site and the second transposome complex binding to the second transposome complex binding site. In some embodiments, the first 5'-tagged target fragment and the second 5'-tagged target fragment include nucleic acid sequences contained in the region of the double-stranded nucleic acid between the first transposome complex binding site and the second transposome complex binding site. In some embodiments, the first 5'-tagged target fragment and the second 5'-tagged fragment are at least partially complementary.
[0498] In some embodiments, the catalytically inactive endonuclease or zinc finger DNA-binding domains contained in the first and second targeted transposome complexes are different. A typical method using two targeted transposome complexes containing catalytically inactive endonucleases is shown in Figure 11.
[0499] In some embodiments, catalytically inactive endonuclease or zinc finger DNA-binding domains of a first transposome complex, which is a targeted transposome complex, and a second transposome complex, which is a targeted transposome complex, bind to different target sequences within a given target region in a target nucleic acid.
[0500] F. Sample and target nucleic acid In some embodiments, the sample comprises a target nucleic acid. In some embodiments, the sample comprises DNA. In some embodiments, the DNA is genomic DNA. In some embodiments, the target nucleic acid is double-stranded DNA.
[0501] In some embodiments, the target nucleic acid is single-stranded DNA. Single-stranded DNA cannot be fragmented by transposases, but the methods described herein describe means for generating a region of double-stranded DNA, such as by hybridizing a targeted oligonucleotide to single-stranded DNA.
[0502] A biological sample can be of any type, including nucleic acids. For example, a sample may include nucleic acids in various purified states, including purified nucleic acids. However, the sample does not need to be completely purified and may include, for example, nucleic acids mixed with proteins, other nucleic acid species, other cellular components, and / or any other impurities. In some embodiments, the biological sample includes a mixture of nucleic acids, proteins, other nucleic acid species, other cellular components, and / or any other impurities present in proportions similar to those found in vivo. For example, in some embodiments, the components are found in proportions similar to those found in intact cells. In some embodiments, the biological sample has a 260 / 280 absorbance ratio of 2.0, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, 0.7, or 0.60 or less. In some embodiments, the biological sample has a 260 / 280 absorbance ratio of at least 2.0, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, 0.7, or 0.60. The method provided herein allows nucleic acids to be bound to a solid support so that other contaminants can be removed simply by washing the solid support after surface binding tagmentation has occurred. The biological sample may include, for example, crude cell lysates or whole cells. For example, crude cell lysates applied to a solid support in the method herein do not require one or more of the conventional separation steps used to isolate nucleic acids from other cellular components. Exemplary separation steps are described in Maniatis et al., Molecular Cloning: A Laboratory Manual, 2nd Edition, 1989, and Short Protocols in Molecular Biology, ed. Ausubel, et al., and are incorporated herein by reference.
[0503] In some embodiments, the sample applied to the solid support has an absorbance ratio of 1.7 or less (260 / 280).
[0504] Therefore, in some embodiments, the biological sample may include, for example, blood, plasma, serum, lymph, mucus, sputum, urine, semen, cerebrospinal fluid, bronchial aspirate, feces, and their macerated tissue or lysates, or any other biological specimen containing nucleic acids.
[0505] In some embodiments, the sample is blood. In some embodiments, the sample is cell lysate. In some embodiments, the cell lysate is crude cell lysate. In some embodiments, the method further includes the step of applying the sample to a solid support and then lysing the cells in the sample to produce cell lysate.
[0506] In some embodiments, the sample is a biopsy sample. In some embodiments, the biopsy sample is a liquid or solid sample. In some embodiments, a biopsy sample from a cancer patient is used to evaluate the target sequence to determine whether the subject has a specific mutation or variant in the predicted gene.
[0507] One advantage of the methods and compositions presented herein is that biological samples can be added to a flow cell, and the subsequent dissolution and purification steps can all occur within the flow cell without any further transfer or handling steps, simply by flowing the necessary reagents through the flow cell.
[0508] In some embodiments, protective elements may be incorporated into polynucleotides (such as target nucleic acids or double-stranded fragments produced by tagmentation). For example, protective elements may be added to target nucleic acids before tagmentation or to double-stranded nucleic acid fragments after tagmentation in any of the methods described herein. As used herein, the term “protective element” is intended to mean an element that inhibits modification of the polynucleotide terminal when used in relation to the 5' or 3' terminal of a polynucleotide. Exemplarily, protective elements may inhibit the action of one or more enzymes at the terminal of a polynucleotide, such as the action of 5' or 3' exonucleases. Non-limiting examples of protective elements include hairpin sequences ligated to the 5' and 3' chains at the terminals of a double-stranded polynucleotide, modified bases (e.g., including phosphorothioate bonds or 3' phosphates), or dephosphorylated bases.
[0509] G. Gap filling ligation In some embodiments, gaps in the DNA sequence remaining after a transposition event may also be filled using a strand displacement extension reaction, such as one containing Bst DNA polymerase and a dNTP mixture. In some embodiments, gap-filling ligation is performed using an extension-ligation mix buffer.
[0510] In some embodiments, the method includes treating a plurality of 5'-tagged target fragments with polymerase and ligase to extend and ligate the chains to produce fully double-stranded tagged fragments.
[0511] Next, a library of double-stranded DNA fragments can be selectively amplified (for example, using cluster amplification) and then sequenced using sequencing primers.
[0512] H. Amplification This disclosure further relates to the amplification of tagged fragments produced according to the methods provided herein. In some embodiments, the immobilized tagged fragments are amplified on a solid support. In some embodiments, the solid support is the same solid support on which tagmentation occurs, bonded to the surface. In such embodiments, the methods and compositions provided herein allow the sample preparation to proceed on the same solid support from the initial sample introduction step to amplification and optionally to a sequencing step.
[0513] For example, in some embodiments, immobilized tagged fragments are amplified using cluster amplification methodologies, as illustrated by the disclosures of U.S. Patents 7,985,565 and 7,115,400, the contents of each of these are incorporated herein by reference in their entirety. The incorporated materials of U.S. Patents 7,985,565 and 7,115,400 describe a solid-phase nucleic acid amplification method that enables the immobilization of amplification products onto a solid support to form an array consisting of clusters or “colonies” of immobilized nucleic acid molecules. Each cluster or colony on such an array is formed from multiple identical immobilized polynucleotide chains and multiple identical immobilized complementary polynucleotide chains. The array thus formed is generally referred herein to as a “clustered array.” Products of solid-phase amplification reactions, such as those described in U.S. Patents 7,985,565 and 7,115,400, are so-called "bridged" structures formed by annealing a pair of immobilized polynucleotide chains and an immobilized complementary chain, both chains immobilized on a solid support at their 5' ends via covalent bonds in some embodiments. Cluster amplification is an example of a method for producing immobilized amplicons using an immobilized nucleic acid template. Immobilized amplicons can also be produced from immobilized DNA fragments produced according to the methods provided herein using other suitable methods. For example, one or more clusters or colonies can be formed by solid-phase PCR, regardless of whether one or both primers of each pair of amplification primers are immobilized.
[0514] In other embodiments, the tagged fragment is amplified in solution. For example, in some embodiments, the tagged fragment is cleaved or otherwise released from the solid support, and then the amplification primer is hybridized to the free molecule in solution. In other embodiments, the amplification primer is hybridized to the tagged fragment for one or more initial amplification steps, and then subsequent amplification steps are performed in solution. In some embodiments, an immobilized nucleic acid template can be used to generate a liquid-phase amplicon.
[0515] It will be understood that tagged fragments can be amplified using any amplification method described herein or generally known in the art, together with universal or target-specific primers. Suitable amplification methods include, but are not limited to, polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription amplification (TMA), and nucleic acid sequence-based amplification (NASBA), as described in U.S. Patent No. 8,003,354, which is incorporated herein by reference in its entirety. One or more nucleic acids of interest can be amplified using the above amplification methods. For example, immobilized DNA fragments may be amplified using PCR, including multiplex PCR, SDA, TMA, and NASBA. In some embodiments, primers specifically directed to the nucleic acid of interest are included in the amplification reaction.
[0516] Other suitable methods for amplifying polynucleotides include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998), incorporated herein by reference), and oligonucleotide ligation assay (OLA) techniques (generally U.S. Patents Nos. 7,582,420, 5,185,243, 5,679,524, and 5,573,907; European Patent No. 0320308(B1), 0336731(B1), 0439182(B1); International Publication No. 90 / 01069, 89 / 12696, and 89 / 09835, all incorporated herein by reference). It should be understood that these amplification methods may be designed to amplify immobilized DNA fragments. For example, in some embodiments, the amplification method may include a ligation probe amplification or oligonucleotide ligation assay (OLA) reaction containing a primer specifically directed to the nucleic acid of interest. In some embodiments, the amplification method may include a primer extension ligation reaction containing a primer specifically directed to the nucleic acid of interest. As non-limiting examples of primer extension and ligation primers that can be specifically designed to amplify the nucleic acid of interest, the amplification may include primers used in GoldenGate assays (Illumina, Inc., San Diego, CA), as exemplified in U.S. Patents 7,582,420 and 7,611,869, each of which is incorporated herein by reference in whole.
[0517] Examples of isothermal amplification methods that may be used in the methods of this disclosure include, but are not limited to, multi-substitution amplification (MDA) as exemplified by Dean et al., Proc. Natl. Acad. Sci. USA 99:5261-66 (2002), or isothermal chain substitution nucleic acid amplification as exemplified by, for example, U.S. Patent No. 6,214,587 (each of which is incorporated herein by reference in whole). Other non-PCR methods that may be used in this disclosure include, for example, strand substitution amplification (SDA) described in Walker et al., Molecular Methods for Virus Detection, Academic Press, Inc., 1995, U.S. Patents 5,455,166 and 5,130,238, and Walker et al., Nucl. Acids Res. 20:1691-96 (1992), or superbranched strand substitution amplification described in, for example, Lage et al., Genome Research 13:294-307 (2003), each of which is incorporated herein by reference in its entirety. Isothermal amplification can be used with large fragments of strand substitution Phi 29 polymerase or Bst DNA polymerase, 5'→3' exo-, for random primer amplification of genomic DNA.
[0518] The use of these polymerases takes advantage of their high processability and chain displacement activity. Due to their high processability, the polymerases are capable of producing fragments 10–20 kb in length. As described above, smaller fragments can be produced under isothermal conditions using polymerases with lower processability and chain displacement activity, such as Klenow polymerase. Further descriptions of the amplification reactions, conditions, and components are detailed in the disclosure of U.S. Patent No. 7,670,810, which is incorporated herein by reference in its entirety.
[0519] Another nucleic acid amplification method useful in this disclosure is tagged PCR using a population of two-domain primers, each consisting of a constant 5' region followed by a random 3' region, as described, for example, in Grothues, et al. Nucleic Acids Res. 21(5):1321-2 (1993), which is incorporated herein by reference in its entirety. The first round of amplification is performed to enable numerous initiations on thermally denatured DNA based on individual hybridization from randomly synthesized 3' regions. Due to the nature of the 3' region, the start sites are considered to be random throughout the genome. Subsequently, unbound primers can be removed, and further replication can be performed using primers complementary to the constant 5' region.
[0520] I. Sequencing and Resequencing Initial sequencing (and possible resequencing) can be performed using a variety of different methods.
[0521] This disclosure further relates to sequencing of tagged fragments generated according to the methods provided herein. In some embodiments, the methods include sequencing one or more of 5' tagged fragments or fully double-stranded tagged fragments.
[0522] Tagged fragments produced by transposome-mediated tagmentation can be sequenced according to any suitable sequencing method, including direct sequencing such as synthetic sequencing, ligation sequencing, hybridization sequencing, and nanopore sequencing. In some embodiments, the tagged fragments are sequenced on a solid support. In some embodiments, the solid support for sequencing is the same solid support on which surface-bound tagmentation occurs. In some embodiments, the solid support for sequencing is the same solid support on which amplification occurs.
[0523] One representative sequencing method is synthetic sequencing (SBS). In SBS, the sequence of nucleotides in a nucleic acid template (e.g., a target nucleic acid or its amplicon) is determined by monitoring the extension of nucleic acid primers along the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme). In certain polymer-based embodiments of SBS, fluorescently labeled nucleotides are attached to the primers in a template-dependent manner (thus extending the primers) so that the sequence of the template can be determined by detecting the order and type of nucleotides attached to the primers.
[0524] The flow cell provides a convenient solid support for containing amplified DNA fragments produced by the methods of the present disclosure. One or more amplified DNA fragments in such format can be subjected to SBS or other detection techniques, which involve repeated delivery of reagents during a cycle. For example, to initiate a first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc., can be flowed into / passed through a flow cell containing one or more amplified nucleic acid molecules. The site where the labeled nucleotides are incorporated by primer extension can be detected. Optionally, the nucleotides may further include reversible termination properties, where the nucleotide terminates further primer extension once it is attached to the primer. For example, a nucleotide analog with a reversible terminator moiety can be attached to the primer so that further extension does not occur until a deblocking agent is delivered to remove that moiety. Thus, in embodiments using reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washing can be performed between various delivery steps. Next, the cycle is repeated n times to extend the primer with n nucleotides, thereby enabling the detection of a sequence of length n. Exemplary SBS procedures, fluid systems, and detection platforms that can be readily adapted for use with the amplicons produced by the method of this disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, U.S. Patent No. 7,057,026, International Publication No. 91 / 06678, International Publication No. 07 / 123744, U.S. Patent No. 7,329,492, International Publication No. 7,211,414, International Publication No. 7,315,019, International Publication No. 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.
[0525] Other sequencing procedures using cyclic reactions, such as pyrosequencing, can be used. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into nascent nucleic acid chains (Ronaghi et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 11(1), 3-11 (2001); Ronaghi et al. Science 281(5375), 363 (1998); U.S. Patents 6,210,891, 6,258,568, and 6,274,320, each incorporated herein by reference). In pyrosequencing, the released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of the generated ATP can be detected via luciferase-generating photons. Therefore, the sequencing reaction can be monitored via a luminescence detection system. Excitation radiation sources used in fluorescence-based detection systems are not required for the pyrosequencing procedure. Useful fluid systems, detectors, and procedures that can be adapted for the application of pyrosequencing to amplicons produced according to this disclosure are described, for example, in International Publication 2012058096, U.S. Patent Application Publication 2005 / 0191698(A1), U.S. Patent No. 7,595,883, and U.S. Patent No. 7,244,559, each of which is incorporated herein by reference.
[0526] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. For example, nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interactions between fluorophore-supported polymerase and γ-phosphate-labeled nucleotides, or using zero-mode waveguides (ZMWs). Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference.
[0527] Some SBS embodiments include the detection of protons released during the incorporation of nucleotides into the extension product. For example, sequencing based on the detection of released protons can be performed using electrodetectors and related technologies commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies), or sequencing methods and systems described in U.S. Patent Publications 2009 / 0026082(A1), 2009 / 0127589(A1), 2010 / 0137143(A1), or 2010 / 0282617(A1), the entirety of which are incorporated herein by reference. The methods herein for amplifying target nucleic acids using binding equilibrium exclusion can be readily applied to substrates used for proton detection. More specifically, the methods herein can be used to generate a clonal population of amplicons used for proton detection.
[0528] Another useful sequencing technique is nanopore sequencing (see, for example, Deamer et al. Trends Biotechnol. 18, 147-151 (2000); Deamer et al. Acc. Chem. Res. 35: 817-825 (2002), Li et al. Nat. Mater. 2: 611-615 (2003), the disclosures of which are incorporated herein by reference). In some embodiments of nanopores, the target nucleic acid or individual nucleotides removed from the target nucleic acid passes through the nanopore. As the nucleic acid or nucleotides pass through the nanopore, each nucleotide species can be identified by measuring the variation in the electrical conductance of the pore. (U.S. Patent No. 7,001,792, Soni et al. Clin. Chem. 53, 1996-2001 (2007), Healy, Nanomed. 2, 459-481 (2007), Cockroft et al. J. Am. Chem. Soc. 130, 818-820 (2008), these disclosures are incorporated herein by reference).
[0529] Exemplary methods for array-based expression and genotyping analysis applicable to the detections described herein are described in U.S. Patent Nos. 7,582,420, 6,890,741, 6,913,884, or 6,355,431, or U.S. Patent Publication Nos. 2005 / 0053980(A1), 2009 / 0186349(A1), or 2005 / 0181440(A1), each of which is incorporated herein by reference.
[0530] An advantage of the methods described herein is that they provide the rapid and efficient parallel detection of multiple target nucleic acids. Therefore, this disclosure provides an integrated system that allows for the preparation and detection of nucleic acids using techniques known in the art, such as those exemplified above. Accordingly, the integrated system of this disclosure may include a fluid component capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, and the system may include components such as pumps, valves, reservoirs, and fluid lines. A flow cell may constitute and / or be used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent Application Publications 2010 / 0111768(A1) and 2012 / 0270305(A1), each of which is incorporated herein by reference. As exemplified with respect to flow cells, one or more fluid components of the integrated system may be used in amplification and detection methods. Taking an embodiment of nucleic acid sequencing as an example, one or more fluid components of the integrated system may be used for the delivery of sequencing reagents in the amplification method described herein and the sequencing method as exemplified above. Alternatively, the integrated system may include separate fluid systems for carrying out amplification and detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and determining nucleic acid sequences include, but are not limited to, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and the apparatus described in U.S. Patent Application Publication 2012 / 0270305, incorporated herein by reference.
[0531] J. Preservation of continuity information when sequencing target nucleic acids In some embodiments, continuity information is stored based on targeted oligonucleotides.
[0532] In some embodiments, a method for preserving continuity information when sequencing a target nucleic acid includes producing tagged fragments of the target nucleic acid using a method comprising a targeted transposome complex containing a recombinase-coated targeted oligonucleotide; sequencing the 5' tagged fragments or fully double-stranded tagged fragments to provide the fragment sequences; grouping the fragment sequences containing the same targeted oligonucleotide sequence; and determining that, if they contain the same targeted oligonucleotide sequence, the group of sequences were in close proximity within the target nucleic acid.
[0533] Continuity information may also be stored based on adapter sequences containing unique molecular identifier (UMI) sequences. In some embodiments, a method for storing continuity information when sequencing a target nucleic acid is to produce tagged fragments of the target nucleic acid using a targeted transposome complex containing a recombinase-coated targeted oligonucleotide, wherein one or more adapter sequences contain a unique molecular identifier (UMI) associated with a single targeted oligonucleotide sequence; to sequence the 5' tagged fragments or fully double-stranded tagged fragments to provide the fragment sequences; to group the fragment sequences containing the same UMI sequence; and to determine that, if they contain the same UMI sequence, the group of sequences were in close proximity within the target nucleic acid.
[0534] Targeted transposomes may also be used in a method for generating a physical map of immobilized polynucleotides. This method can advantageously utilize clusters that are likely to contain linked sequences (i.e., first and second portions derived from the same target polynucleotide molecule) for identification. Thus, the relative proximity of any two clusters obtained from the immobilized polynucleotide provides useful information for aligning the sequence information obtained from the two clusters. Specifically, the distance between any two given clusters on a solid surface is positively correlated with the probability that the two clusters originate from the same target polynucleotide molecule, as described in detail in International Publication No. 2012 / 025250, which is incorporated herein by reference in its entirety.
[0535] For example, in some embodiments, long DNA molecules extending across the surface of a flow cell are tagged in situ, resulting in lines of connected DNA bridges across the surface of the flow cell. Furthermore, a physical map of the immobilized DNA is created. Thus, the physical map correlates the physical relationships of clusters after the immobilized DNA has been amplified. Specifically, the physical map calculates the probability that sequence data obtained from any two clusters are linked, as described in the incorporated material in International Publication No. 2012 / 025250.
[0536] In some embodiments, a physical map is generated by imaging DNA to establish the position of immobilized DNA molecules across a solid surface. In some embodiments, the immobilized DNA is imaged by adding an imaging agent to a solid support and detecting a signal from the imaging agent. In some embodiments, the imaging agent may be a detectable label. Suitable detectable labels include, but are not limited to, protons, haptens, radionuclides, enzymes, fluorescent labels, chemiluminescent labels, and / or chromogens. For example, in some embodiments, the imaging agent is an insertion dye or a non-insertion DNA binder. Any suitable insertion dye or non-insertion DNA binder known in the art may be used, including, but are not limited to, those described in U.S. Patent Application Publication 2012 / 0282617, which is incorporated entirely herein by reference.
[0537] In some embodiments, the immobilized DNA double strands are further fragmented to release the free ends before strand exchange and cluster formation. The cleavage of the bridge structure can be carried out using any suitable methodology known in the art, as exemplified by the incorporated material in International Publication No. 2012 / 025250. For example, cleavage may occur by incorporating a modified nucleotide such as uracil, as described in International Publication No. 2012 / 025250, by incorporating a restriction endonuclease site, or by applying a liquid-phase transposome complex to the bridge DNA structure, as described elsewhere herein.
[0538] In certain embodiments, multiple nucleic acids are flowed onto a flow cell comprising multiple nanochannels, each nanochannel having multiple transposome complexes immobilized thereon. As used herein, the term “nanochannel” refers to a narrow channel through which long linear nucleic acid molecules flow. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 900 or fewer, or up to 1000, individual long-chain target DNA molecules of the target nucleic acid flow into each nanochannel. In some embodiments, individual nanochannels are separated by a physical barrier that prevents individual long chains of target DNA from interacting with multiple nanochannels. In some embodiments, the solid support comprises at least 10, 50, 100, 200, 500, 1000, 3000, 5000, 10000, 30000, 50000, 80000, or 100000 nanochannels. In some embodiments, transposomes bound to the surface of the nanochannels tag the DNA. Sequential mapping can then be performed, for example, by tracking clusters along the length of one of these channels. In some embodiments, the long chain of the target DNA is at least 0.1kb, 1kb, 2kb, 3kb, 4kb, 5kb, 6kb, 7kb, 8kb, 9kb, 10kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 55kb, 60kb, 65kb, 70kb, 75kb, 80kb, 85kb, 90kb, 95kb, 10 The length may be 0kb, 150kb, 200kb, 250kb, 300kb, 350kb, 400kb, 450kb, 500kb, 550kb, 600kb, 650kb, 700kb, 750kb, 800kb, 850kb, 900kb, 950kb, 1000kb, 5000kb, 10000kb, 20000kb, 30000kb, or 50000kb.In some embodiments, the long chain of the target DNA is 0.1kb, 1kb, 2kb, 3kb, 4kb, 5kb, 6kb, 7kb, 8kb, 9kb, 10kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 55kb, 60kb, 65kb, 70kb, 75kb, 80kb, 8 The lengths are 5kb, 90kb, 95kb, 100kb, 150kb, 200kb, 250kb, 300kb, 350kb, 400kb, 450kb, 500kb, 550kb, 600kb, 650kb, 700kb, 750kb, 800kb, 850kb, 900kb, 950kb or less, or 1000kb or less. For example, a flow cell having 1000 or more nanochannels containing immobilized tagmentation products mapped within the nanochannels can be used to sequence the genome of an organism with short, "positioned" reads. In some embodiments, the immobilized tagmentation products mapped within the nanochannels can be used to analyze haplotypes. In some embodiments, the immobilized tagmentation products mapped within the nanochannels can be used to analyze phasing challenges.
[0539] IV. Method using a sample containing cell-free DNA and a targeted transposomal complex The targeted transposomes described herein can be used for targeted transposition within simplified library preparation and enrichment protocols. In some embodiments, the simplified protocols require less time or user steps compared to existing protocols. In some embodiments, one or more target nucleic acid sequences are contained in histone-associated DNA. In some embodiments, the histone-associated DNA is cell-free DNA.
[0540] In some embodiments, simplified library preparation and enrichment protocols are intended for use with cell-free DNA (cfDNA), such as the exemplary method shown in Figure 15. Current library preparation for cfDNA generally involves several steps, including cfDNA extraction from plasma (30 min), end repair (30 min), A-tailing (30 min), ligation of non-random unique molecular identifiers (UMIs) (30 min), ligation of adapters (30 min), and SPRI cleanup and subsequent PCR amplification (~30 min). cfDNA extraction from plasma in standard methods may include a protease step (e.g., proteinase K, as described in Illumina Document #1000000001856 v06 (April 2020) providing a protocol for VeriSeq NIPT). Based on these steps, cfDNA library preparation is a time-consuming and inefficient process that is difficult to automate.
[0541] Cell-free DNA (cfDNA) in plasma is known to exist associated with histones (see Marshman et al., Cell Death and Disease (2016) 7, e2518, and Rumore and Steinman J. Clin Inv. 86: 69-74 (1990)). A key challenge when performing direct tagmentation in plasma samples is removing histones from cfDNA. Methods for histone removal may include a protease step, which can also degrade proteins involved in tagmentation. For example, cfDNA extraction from plasma using the VeriSeq Non-Invasive Prenatal Testing (NIPT) method (Illumina) includes a protease step (proteinase K, as described in the VeriSeq NIPT solution insert, Illumina Document #1000000001856 v06 (April 2020)) followed by several washing steps before library preparation. Targeting specific sequences of interest (such as genes within the genome) with transpososomes significantly simplified workflows using cfDNA-containing samples without requiring histone removal.
[0542] Zinc finger DNA-binding domains allow zinc finger nucleases to target specific regions of the genome being edited (see Costa et al., Genome Editing Using Engineered Nucleases and Their Use in Genomic Screening, PMID:29165977, in Assay Guidance Manual (Markossian et al., editors) (2017)). In particular, ZFNs have the ability to efficiently cleave histone-bound DNA, while Cas9 nucleases are strongly inhibited when DNA is bound to histones (see Yarringon et al., PNAS 115(38):9351-9358 (2018)).
[0543] In some embodiments, the histone-bound DNA may be contained within a nucleosome. As used herein, “nucleosome” refers to a structure consisting of a segment of DNA wrapped around eight histone proteins. In some embodiments, the histone-bound DNA is cell-free DNA. Exemplary cell-free DNA may be cfDNA found in a blood sample from a pregnant woman (cfDNA may be of fetal origin) or a patient with cancer or suspected cancer (cfDNA may be of tumor cell origin).
[0544] In some embodiments, targeted transpososomes are targeted to one or more regions in cfDNA by a zinc finger DNA-binding domain. In some embodiments, histone-bound DNA (such as cfDNA) is tagged using targeted transpososomes containing a zinc finger DNA-binding domain.
[0545] In some embodiments, the method further includes the step of adding an affinity-binding partner to a solid support after fragmentation, so that the tagged target fragment is bound to the solid support. In some embodiments, fragmentation is stopped before adding the affinity element to the solid support. In some embodiments, fragmentation is stopped by adding a solution containing proteinase K and / or SDS.
[0546] For example, a transposomal complex containing a zinc finger DNA-binding domain can target a specific sequence of interest within cfDNA, as shown in Figure 15. In some embodiments, the zinc finger DNA-binding domain contained in the targeted transposon can bind to a sequence located within or near an oncogene to generate a targeted library from cfDNA in a sample derived from a cancer patient, allowing for the assessment of whether gain-of-function mutations are present in the cfDNA. Alternatively, the zinc finger DNA-binding domain contained in the targeted transposon can bind to a sequence located within or near a tumor suppressor gene to generate a specific library from cfDNA, allowing for the assessment of whether loss-of-function mutations (i.e., activating mutations) are present in the cfDNA. In this way, such targeted transposons can be used to generate targeted libraries for assessing changes in cancer cells associated with more aggressive tumors or a worse prognosis.
[0547] Similarly, cfDNA-derived targeted libraries can be used to evaluate specific gene sequences associated with genetic disorders. These genetic disorders may be known genetic diseases caused by known alterations in gene sequences, such as Tay-Sachs disease, cystic fibrosis, and many others well known to those skilled in the art. In some embodiments, a zinc finger DNA-binding domain contained in a targeted transposome may bind to a sequence contained within or near a gene associated with a genetic disorder to generate a targeted library. In some embodiments, the targeted library may be used for sequencing a region of a target gene for SNPs or other mutations in prenatal testing using maternal plasma containing fetal cfDNA.
[0548] V. Methods for sorting and selecting single-cell nucleic acids This specification describes a method for enabling cell sorting based on "omics" characteristics by utilizing sc-NGS (single-cell next-generation sequencing) in combination with nucleic acid selection techniques. The method may include enriching or depleting sc library members by targeting unique cell barcodes. This workflow, which includes a two-step sequencing process, provides an easy-to-manage methodology in which the initial sequencing creates a cell database used to determine which cells to obtain additional "omics" data from in a second, more comprehensive sequencing run, after the selection of desired cells. Figure 3 provides an overview of such a sorting and selection method, in which the desired cell barcode ID is determined using the first 16s sequencing, followed by enrichment of desired samples or depletion of unwanted samples. After enrichment / depletion, the desired samples can undergo comprehensive sequencing.
[0549] In some embodiments, cell selection is achieved by depleting unwanted samples, such as abundant cells of little interest, from the sc library based on the UBCs assigned to them. This depletion-based secondary sequencing can characterize the desired sample, i.e., a DNA library generated from target cells that may be scarce in the library. In some embodiments, cell selection is achieved by enriching the desired sample using the UBCs assigned from the sc library. These desired samples may be scarce or low in abundance in the sample.
[0550] VI. Method for characterizing a desired sample in a mixed pool of samples A method for characterizing a desired sample in a mixed pool of samples containing both desired and unwanted samples is described herein. In some embodiments, the method first involves sequencing a library containing multiple nucleic acid samples derived from the mixed pool of samples to generate sequence data from double-stranded nucleic acids. In some embodiments, each nucleic acid library includes a single-sample nucleic acid and a unique sample barcode to distinguish the nucleic acid derived from a single sample from nucleic acids derived from other samples in the library.
[0551] This method may be a cost-effective way to characterize a single cell within a given population based on a barcode associated with a cell possessing a desired genomic feature (where the desired genomic feature may be the presence of a specific gene mutation, the methylation status of a given gene, etc.). This desired genomic feature can be determined from initial sequencing, followed by a selection step, and then resequencing to provide further information about the single cell of interest. Typical methods for incorporating barcodes are shown in Figures 5 and 6.
[0552] In some embodiments, the method also includes performing a selection step on the library, which includes analyzing sequence data to identify unique sample barcodes associated with sequence data from a desired sample, concentrating nucleic acid samples from the desired sample and / or depleting nucleic acid samples from unwanted samples, and resequencing the nucleic acid library.
[0553] In some embodiments, resequencing is orthogonal resequencing. As used herein, “orthogonal resequencing” refers to resequencing that analyzes different physiological features compared to initial sequencing. For example, initial sequencing may assess methylation status, while resequencing may be a comprehensive whole-genome sequencing of cells having a desired methylation pattern. In other words, initial sequencing and resequencing may assess the same features of a mixed pool of samples, but initial sequencing and resequencing may also assess different features of a desired sample.
[0554] The advantage of this method is that it avoids certain steps that are typically used to generate sequence data for a desired sample. In other words, the method of the present invention may be faster or easier than other methods, or it may avoid steps that may bias the results. In some embodiments, the method does not use enrichment methods based on cell sorting. In some embodiments, the method does not use FACS. In some embodiments, the method does not use FACS based on cell size, morphology, or surface protein expression. In some embodiments, the method does not use microfluidics. In some embodiments, the method does not use whole-genome amplification. Avoiding these steps in this method can reduce the time and cost required to generate comprehensive sequence data for a desired sample. Furthermore, avoiding these steps can avoid biases that may arise from specific methods (e.g., relying on surface protein expression to sort cells using FACS).
[0555] Furthermore, this sequencing and analysis method can be performed using a sequencing system without the need for FACS equipment or other such devices.
[0556] In some embodiments, the initial sequencing results can be used to guide the selection process without the initial sequencing being biased by a prior sorting step. Using this method, those skilled in the art can select multiple single-cell libraries by initial sequencing for a desired characteristic, use these initial sequencing results to determine which cells are the desired cells, and then select and resequence the desired cells.
[0557] Other advantages of this method are described herein.
[0558] A. Library preparation The initial sequencing step of these methods may be any means of generating a library containing multiple nucleic acid samples derived from a mixed pool of samples. In some embodiments, the library is a single-cell library (sc library). As used herein, “single-cell library” or “sc library” means a library generated from a single cell within a mixed population of cells. However, the library may also be a library derived from a single nucleus, virus, or high molecular weight (HMW) DNA within a mixed population. Thus, the methods of the present invention may be used with various mixed populations, and any method described for use with an sc library may be used for other types of libraries.
[0559] In some embodiments, this method is performed after library indexing but before comprehensive library sequencing.
[0560] In some embodiments, the nucleic acid library includes nucleic acids derived from a single sample, each containing a unique sample barcode for distinguishing the nucleic acid derived from that single sample from nucleic acids derived from other samples in the library. A wide variety of means for generating such libraries are well known in the art. The advantage of the present method is that it can be used in conjunction with libraries generated by many different methods. Thus, those skilled in the art can, based on their preferences, select a particular method for generating a library containing multiple nucleic acid samples derived from a mixed pool of samples and perform initial sequencing. The disclosed method can then be used for selection based on unique sample barcodes and subsequent resequencing.
[0561] Typical methods of SC sequencing include the method described in International Publication No. 2016 / 130704, which is incorporated herein by reference. In some embodiments, the method includes a step of spatially separating nucleic acid samples before incorporating unique sample barcodes.
[0562] These methods are applicable to any sc library generation and sequencing method using unique cell barcodes (UBCs) or unique sample barcodes. Exemplary sc library generation / sequencing methods include Biorad ddSEQ (e.g., using the Illumina Bio-Rad SureCell WTA 3' Library Prep Kit), various 10X Genomics systems (e.g., Chromium Single Cell Expression), Drop-Seq (see Macosko et al., Cell 161(5):1202-1214 (2015)), InDrop® (1CellBio), Tapestri® Platform (MissionBio), Split-Seq (see Rosenburg et al., Science 360(6385):176-182 (2018)), or Illumina's Single Cell Combinatorial Indexing Sequencing (SCI-seq, Cao et al., Science) Examples include 357(6352):661-667(2017), all of which are incorporated by reference for the disclosure of library generation and sequencing methods.
[0563] In some embodiments, the method includes tagging multiple nucleic acid samples derived from a mixed pool of samples before sequencing. In some embodiments, the library is generated using tagging. In some embodiments, tagging incorporates a unique sample barcode into each nucleic acid sample.
[0564] In some embodiments, the universal primer is incorporated into each nucleic acid sample in the nucleic acid library. In some embodiments, the universal primer is incorporated into each nucleic acid sample during library preparation. In some embodiments, the universal primer is P5 and P7 primer. In some embodiments, the P5 and P7 sequences are incorporated into each nucleic acid sample in the nucleic acid library.
[0565] In some embodiments, the i5 and i7 sequences are incorporated into each nucleic acid sample in the nucleic acid library. In some embodiments, the i5 and i7 sequences are incorporated into each nucleic acid sample during library preparation.
[0566] B. Initial sequencing In some embodiments, untargeted initial sequencing may be useful for characterizing multiple single cells, after which selection and resequencing can be performed to further analyze the target single cell using the population. In some embodiments, initial sequencing identifies unique sample barcodes associated with unwanted samples. In some embodiments, initial sequencing identifies unique sample barcodes associated with desired samples.
[0567] In some embodiments, targeted initial sequencing can determine the target cells within a population of single cells (i.e., determine the desired sample), and then a library generated from these target cells can be selected and resequenced to provide further information.
[0568] In some embodiments, the initial sequencing step includes targeted sequencing, and the resequencing step includes whole-genome sequencing. In some embodiments, the initial sequencing may be gene-specific sequencing. In some embodiments, the initial sequencing may be 16s sequencing.
[0569] In some embodiments, the initial sequencing step includes targeted sequencing using one or more gene-specific primers (as illustrated in Figure 7). In some embodiments, the gene-specific primers include a universal primer tail.
[0570] In some embodiments, the initial sequencing step does not include whole-genome sequencing, while the resequencing step does. In other words, the initial sequencing may be less comprehensive, and the resequencing may be more comprehensive. Such an approach can dramatically reduce the time and cost required to generate comprehensive data on desired samples by avoiding the resequencing of unnecessary samples.
[0571] In some embodiments, the initial sequencing step includes ribosome sequencing, and the resequencing step includes whole-genome sequencing. In some embodiments, the ribosome sequencing includes 16s, 18s, or internal transcription spacer sequencing. In some embodiments, the internal transcription spacer region is located between the 16s rRNA gene and the 23s rRNA gene. In some embodiments, ribosome sequencing is used to determine the species in a sample, including a mixed pool of samples containing samples from different species. For example, ribosome sequencing can be used to determine bacterial species in a metagenomics sample. In some embodiments, resequencing includes whole-genome sequencing of the target species after enriching these desired samples from the target species or after depleting unwanted samples from non-target species.
[0572] In some embodiments, initial sequencing is performed to characterize a cell population, followed by resequencing. For example, initial sequencing may identify cells of a desired cell type within a blood sample, and resequencing may then focus specifically on these cells.
[0573] 1. Initial targeting sequencing In some embodiments, initial sequencing is targeted sequencing. As used herein, targeted sequencing refers to sequencing of a region of a target nucleic acid. For example, targeted sequencing may be sequencing of a specific gene within a target genome.
[0574] Figure 7 shows an example of a method for performing targeted initial sequencing. An sc library containing multiple cell nucleic acid libraries may be prepared, each library being tagged with one or more UBCs. Fragments in each cell nucleic acid library contain a P5 sequence at one end and a P7 sequence at the other end. Gene-specific primers with a P7 tail can be used in conjunction with P5 primers to generate a target gene specific to the amplified product from the sc library. In this way, the fragment containing the gene of interest is specifically amplified and can then be used for initial sequencing based on the primer sequences of read 1 and read 2 contained in the amplified fragment. Analysis of the initial sequencing results may identify UBCs associated with cell nucleic acid libraries derived from cells expressing the target sequence for the target gene. Selection can then be performed to proceed with subsequent sequencing of the desired sample.
[0575] In some embodiments, targeted initial sequencing identifies 16s rRNA sequences associated with a target bacterial taxa or species. In some embodiments, targeted initial sequencing identifies cells in a cancer biopsy that express a mutation in the KRAS G12 gene. After initial sequencing and identification of the desired sample, the desired sample can be enriched, or unwanted samples can be depleted. The selected cellular nucleic acid library can be used for deeper sequencing or whole-genome analysis to better understand the sequence of the target single cell.
[0576] A similar approach can be used for any gene of interest. Furthermore, initial sequencing can analyze mRNA expression levels or methylation status in different regions of the target nucleic acid to classify cell types corresponding to different barcodes. If epigenetic factors are evaluated during initial sequencing, resequencing may then provide comprehensive whole-genome sequencing of cells exhibiting the desired phenotype.
[0577] 2. Representative sequencing information obtained from initial sequencing In these methods, initial sequencing may provide sequence information for classification based on "omics" features. In some embodiments, initial sequencing provides information about genomic features such as the sequences or variants of one or more genes. In some embodiments, DNA derived from a sample is sequenced to generate genomic data. In some embodiments, initial sequencing provides information about transcriptome features such as the expression of different genes. In some embodiments, RNA derived from a sample is sequenced to generate transcriptome data. In some embodiments, initial sequencing provides data about methylation marks or patterns. In some embodiments, DNA derived from a sample is used for methylation analysis. In some embodiments, methylation analysis is bisulfite sequencing. In some embodiments, single cells can be selected, and then a single-cell-derived sample can be used for bisulfite sequencing and methylation analysis. For any of these initial sequencing methods, sequencing may be whole-genome sequencing or targeted sequencing.
[0578] In some embodiments, initial sequencing is used to generate metagenomics data. In some embodiments, initial sequencing is used to identify species in a mixed pool of samples containing samples from several species. In some embodiments, initial sequencing is used to identify abundant species in a mixed pool of samples containing samples from several species. Re-sequencing may then generate further sequence data for the desired species. In some embodiments, the species are bacterial species. In some embodiments, the mixed pool of samples contains a mixed pool of bacteria isolated from a patient.
[0579] Initial sequence data can be analyzed using any bioinformatics approach. The analysis of initial sequencing results depends on how the user wants to use the method. In other words, the user can choose the most appropriate method for analyzing initial sequencing results based on how they want to characterize the sample into desired and unwanted samples. For example, if the user wants methylation analysis to be a selection criterion, they will use methylation analysis.
[0580] Furthermore, one clear advantage of this method is that the initial sequencing can be an unbiased analysis of a mixed population, followed by resequencing of the desired sample determined via the initial sequencing. For example, a user may have a metagenomics sample from an infected patient, but the user has no information about the bacterial species contained in that sample. Using this method, the initial 16s sequencing can identify the bacterial species in the sample, allowing the user to identify the sample from a known pathogenic bacterial species. In this case, the desired sample would be one of these potentially pathogenic bacterial species, while the unwanted sample might be an abundant species in the sample known to be non-pathogenic. Resequencing can then be performed to provide more information about the desired sample, such as whether the potentially pathogenic bacteria express genes associated with antibiotic resistance. These results can then be used to determine the best antimicrobial therapy for the subject. This method is particularly powerful because it eliminates the need for the user to make predictions about the presumed pathogenic species, which can bias the results if the infection is caused by rare bacteria. Such methods may also be particularly useful for evaluating samples in which pathogenic bacteria are not adequately cultured. In such cases, the methods of the present invention may enable the identification of potentially pathogenic bacteria and clinically relevant evaluation, whereas culture-based methods for evaluating the same patient sample may miss the presence of these unculturable pathogenic bacteria.
[0581] 3. Amplification and resequencing In some embodiments, the method includes one or more amplification steps after initial sequencing. In some embodiments, the method includes an amplification step before resequencing.
[0582] In some embodiments, amplification is used for selection. In some embodiments, the desired sample is enriched via PCR amplification of the desired sample using a unique sample barcode, as discussed below.
[0583] In some embodiments, amplification is performed after selection. In some embodiments, the desired sample is concentrated or unwanted samples are depleted before the amplification step. In such cases, the amplification may be unbiased, and all remaining samples in the library after selection are amplified. In some embodiments, the amplification step uses universal primers.
[0584] In some embodiments, the amplification and resequencing process is repeated once. In some embodiments, the amplification and resequencing process is repeated two or more times. In some embodiments, the amplification and resequencing process is repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55 times or more, or at any interval made up of the enumerated integers.
[0585] In some embodiments, the sample is amplified on a solid support.
[0586] C. Sample In some embodiments, the method includes the step of first sequencing a library containing a number of individual nucleic acid libraries generated from a mixed pool of nucleic acid samples.
[0587] 1. Mixing pool of samples The sample mixture pool can be any heterogeneous group of samples. For example, the sample mixture pool may be a blood sample containing different individual cells, a tissue sample containing different individual cells (i.e., a tumor sample), or an environmental sample containing different bacterial species.
[0588] In some embodiments, the sample mixture pool includes a mixture pool of cells, a mixture pool of nuclei, or a mixture pool of high molecular weight DNA (HMW DNA). In some embodiments, the sample is cells, nuclei, or HMW DNA. In some embodiments, the HMW DNA is viral DNA. The high molecular weight DNA includes an average fragment length of 20 kb or more. In some embodiments, the DNA includes an average fragment length of 25, 30, 35, 40, 45, 50 kb or more.
[0589] In some embodiments, a single sample is a single cell. In some embodiments, a multiple nucleic acid sample derived from a mixed pool is multiple nucleic acids derived from a mixed pool of cells.
[0590] In some embodiments, the sample mixture pool is collected from a patient. In some embodiments, the mixture pool is derived from blood or other tissue samples, or from biopsy samples taken from a tumor.
[0591] In some embodiments, the sample mixing pool is an environmental sample. In some embodiments, the mixing pool is derived from a mixing pool of different species of bacteria or other microorganisms.
[0592] In some embodiments, the sample mixing pool contains both the desired sample and the unwanted sample.
[0593] 2. Desired sample As used herein, “desired sample” refers to a sample that a person skilled in the art would like to evaluate. This definition does not mean that the desired sample itself is desirable, as a user may want to study malignant cells or other substances that are harmful to the subject being evaluated.
[0594] For example, a person skilled in the art may be interested in only specific individual cell libraries within multiple single-cell libraries. A user may want to study cells with a particular "omics" profile, such as cells expressing gene mutations that confer resistance to cancer drug therapy. Using this method, a person skilled in the art could monitor patients for the potential development of resistance to specific drug therapies.
[0595] In many cases, the desired sample is contained within a pool of samples that also contains other unwanted (i.e., undesirable) samples. The desired sample may be a sample with a specific profile, and it may be contained within a pool of samples that also contains unwanted samples. For example, the desired sample may express a specific gene mutation that is not expressed by unwanted samples from the mixed pool of samples. Alternatively, the desired sample may be a pathogenic bacterium contained within a sample that also contains abundant non-pathogenic bacteria.
[0596] In the methods described herein, any features that can be analyzed by sequencing can be used to characterize the desired sample. Therefore, an advantage of the methods of the present invention is that they can be used with a wide range of different samples.
[0597] In some embodiments, the desired sample is a cell or a nucleus. In some embodiments, the desired sample is a cell. In some embodiments, the desired sample is a nucleus derived from a cell.
[0598] In some embodiments, the desired sample is a human cell or a nucleus derived from a human cell. In some embodiments, the desired sample is a cancer cell or a nucleus derived from a cancer cell. In some embodiments, the desired cell or nucleus is of a specific desired cell type or derived therefrom. In some embodiments, the desired sample has mutations compared to other samples in the pool. In some embodiments, the desired sample is a cancer cell or an immune cell or derived therefrom.
[0599] In some embodiments, the desired sample is or is derived from cancer cells. In some embodiments, the desired sample is or is derived from cancer stem cells. In some embodiments, the desired sample is or is derived from cancer cells in a liquid or tumor biopsy sample. In some embodiments, the desired sample is or is derived from drug-resistant cancer cells.
[0600] In some embodiments, the desired sample is a cancer cell having at least one mutation compared to other cancer cells in the cell pool, or is derived therefrom. In some embodiments, the method is used to track cancer evolution. In some embodiments, cancer evolution may be the emergence of resistance to a given chemotherapy. In some embodiments, the desired sample is a cell having a somatic driver mutation, or is derived therefrom.
[0601] In some embodiments, the desired sample is a metagenomics sample. In some embodiments, the desired sample is a microorganism derived from an environmental sample. In some embodiments, the desired sample is a microorganism not cultured from an environmental sample. In some embodiments, the microorganism includes bacteria, fungi, archaea, algae, protozoa, or viruses. In some embodiments, the desired sample is a pathogen.
[0602] In some embodiments, the desired sample has mutations in its nucleic acid compared to other samples. In some embodiments, the desired sample has single nucleotide variants (SNVs). In some embodiments, the desired sample has copy number variations (CNVs).
[0603] In some embodiments, the desired sample has a desired methylation pattern. In some embodiments, the desired sample has a desired expression pattern. In some embodiments, the desired sample has a desired epigenetic pattern. In some embodiments, the desired sample has a desired immunorecombination.
[0604] In some embodiments, the sample has a specific species type. In some embodiments, the specific species type is a human species. In some embodiments, the specific species type is a specific species of bacteria.
[0605] Some representative uses of this method with different types of samples are described below.
[0606] a) Rare specimens In some embodiments, the desired sample is rare within the starting population. For example, the desired sample may be a single-cell derived sample that was rare in the population of cells used to generate the sc library. Therefore, if sequence data from the entire pool of individual cell-derived libraries in a mixed pool of cells is evaluated, the desired sequence data from the rare cell may be overwhelmed by the abundant sequence data from unwanted cells.
[0607] As used herein, the desired sample is a “rare sample” present in a mixed pool of samples at a concentration of 1%, 0.1%, 0.01%, 0.001%, 0.0001%, 0.00001%, 0.000001%, 0.0000001%, 0.00000001%, or 0.000000001% or less. In some embodiments, the desired sample is the desired cells. In some embodiments, the desired cells are present in a mixed pool of cells at a concentration of 1%, 0.1%, 0.01%, 0.001%, 0.0001%, 0.00001%, 0.000001%, 0.0000001%, 0.00000001%, or 0.000000001% or less. Rare cells can be characterized by any features that can be evaluated by initial sequencing, such as features based on the cell's genomic or epigenetic composition. For example, a rare cell may have DNA that contains mutations compared to the DNA of other cells in the sample. In some embodiments, a rare cell may have a different DNA methylation pattern compared to other cells in the sample. Any features that can be analyzed using sequence data in the methods described herein may be used to characterize a rare sample.
[0608] In some embodiments, the initial sequencing in the method of the present invention can be used to identify libraries produced from rare cells. The selection step can be performed to enrich the desired sample (i.e., the library derived from the target rare cells) or to deplete the unwanted sample (i.e., the library derived from abundant unwanted cells). After selection, the resulting library can be re-sequenced by deeper sequencing to evaluate the characteristics of the desired rare cells.
[0609] 3. Unnecessary samples As used herein, “unwanted sample” refers to a sample that a person skilled in the art would not wish to sequence. Unwanted samples may contain beneficial cells but are of no interest to the user. For example, a user may want to evaluate liver cancer cells derived from a biopsy but not cells containing normal, non-cancerous liver tissue. A person skilled in the art may also only want to sequence a sample derived from cells expressing a specific gene mutation and not from samples derived from other cells in the sample. Sequencing unwanted samples can waste time, resources, and sequencing capacity if there is no option to concentrate the desired sample or deplete the unwanted sample.
[0610] D. Nucleic acid These methods can be used to evaluate nucleic acids. In some embodiments, these nucleic acids are derived from a single cell. In some embodiments, the nucleic acid is DNA. In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid is ribosomal RNA (rRNA). In some embodiments, the nucleic acid is 16s rRNA. In some embodiments, the nucleic acid is 18s rRNA.
[0611] In some embodiments, the nucleic acid is ribosomal DNA (rDNA).
[0612] In some embodiments, the nucleic acid is an internal transcription spacer nucleic acid.
[0613] E. Unique sample barcode and unique cell barcode As used herein, “unique sample barcode” refers to a barcode unique to an individual sample within a pool of samples. In some embodiments, initial sequencing of the library includes sequencing a library containing multiple nucleic acid samples derived from a mixed pool of samples. This mixed pool of samples may be any heterogeneous group of samples, such as blood samples containing different individual cells. In some embodiments, the unique sample barcode can distinguish nucleic acids derived from a desired single sample from nucleic acids derived from other samples in the library.
[0614] A unique sample barcode may consist of a single barcode sequence, or it may consist of multiple barcode sequences. As used herein, “barcode sequence” refers to a sequence that can be used to distinguish samples. For example, a unique sample barcode may be unique to a given desired sample in a mixed pool of samples, based on the multiple barcodes contained in the unique sample barcode, even if a given barcode sequence can be associated with multiple samples. In such a case, a particular combination of barcode sequences within a unique sample barcode may be unique, but one or more barcode sequences within a unique sample barcode may be shared with other samples.
[0615] In some embodiments, the unique sample barcode is a unique cell barcode. As used herein, “unique cell barcode” or “UBC” refers to a barcode unique to a single cell in a mixed pool of cells. When analyzing sequence data, the UBC can be used to identify sequences that were originally present in the same single cell in the starting mixed pool of cells.
[0616] In some embodiments, the unique sample barcode is specific to the nuclear type, HMW DNA, etc., and the present invention is not limited to single-cell use.
[0617] To enable robust enrichment methods, a specific, unique sample barcode design may be desirable. For example, when using a hybrid capture approach, enrichment specificity depends on the ability to design probes that specifically hybridize to the desired unique sample barcode. Similar considerations apply to unique sample barcode-targeted PCR amplification. For this reason, it may be desirable to have a unique sample barcode that exists as a continuous nucleic acid sequence attached to a cellular DNA library. Alternatively, it may be desirable to have fixed sequences between barcode sequences within the unique sample barcode, so that the user knows which primers bind to which combinations of barcode sequences within the unique sample barcode.
[0618] A unique sample barcode can be used in combination with other known barcodes or adapter sequences. For example, a library fragment may include a unique sample barcode and one or more commercially available adapters. In some embodiments, the i5 and / or i7 adapter sequences (Illumina) are included in the library fragment.
[0619] 1. Barcode type In some embodiments, the barcode is a physically addressable barcode. “Physically addressable” means that the barcode contains one or more nucleic acid sequences that can be bound to another drug. In some embodiments, the physically addressable barcode can be bound to complementary nucleic acid sequences. In some embodiments, the physically addressable barcode can be bound by a primer or a capture oligonucleotide. For example, the physically addressable barcode may be bound to a sequencing primer to enable sequencing of library fragments. In another example, the physically addressable barcode may be bound to a capture oligonucleotide to enable immobilization of library fragments onto a flow cell.
[0620] In some embodiments, the barcode is a unique sample barcode.
[0621] In some embodiments, the unique sample barcode is a single continuous barcode. In some embodiments, the unique sample barcode includes two or more barcode sequences without nucleic acid sequences between different barcode sequences. For example, multiple barcode sequences (BC1~BC X ) can be added in different processes, in which case the nucleic acid sequence is not incorporated between the barcode sequences. As shown in the exemplary method in Figure 5, BC1 can be incorporated during tagmentation, and BC2~BC X It can be incorporated via ligation. As shown in the exemplary method in Figure 6, BC1 can be incorporated during tagmentation, followed by one or more ligations of well-specific BCs, and then pooled. The preparation of a single continuous barcode can facilitate the design of primers that can bind to unique sample barcodes.
[0622] In some embodiments, a unique sample barcode is a series of discontinuous barcodes. In some embodiments, the series of discontinuous barcodes are separated by nucleic acid sequences. In some embodiments, the series of discontinuous barcodes are separated by fixed sequences. For example, a series of barcode sequences (BC1~BC X ) can be added in different processes, in which case nucleic acid sequences are incorporated between barcode sequences. Since the barcode and fixed sequences of such multiple discontinuous barcodes are known, it is possible to design primers that can bind to unique sample barcodes.
[0623] F. endonuclease Different endonucleases may be used in the methods of the present invention. As used herein, the term “endonuclease” is used to refer to an enzyme capable of cleaving nucleic acids. An endonuclease may refer to either a catalytically active endonuclease or a catalytically inactive endonuclease. Some characteristics of endonucleases, such as the ability to target specific target sequences based on a guide RNA associated with the endonuclease, are common to both catalytically active and catalytically inactive endonucleases. In some embodiments, the endonuclease associates with a guide RNA that binds to one or more unique sample barcodes. Various different endonucleases that may be used to improve specificity (i.e., to improve targeting and reduce off-target activity) are shown in Figure 8.
[0624] In some embodiments, the endonuclease is a catalytically inactive endonuclease. As used herein, a “catalystically inactive endonuclease” is an endonuclease that can bind to nucleic acids but does not mediate nucleic acid cleavage. Catalytically inactive endonucleases are sometimes also called inactivated endonucleases (e.g., “dCas” proteins). An exemplary catalytically inactive endonuclease is dCas9, as shown in Figure 3 (dCas9 bound to biotin) and Figure 8 (dCas9 contained in a fusion protein with FokI). Typically, endonucleases can bind to nucleic acids and then mediate cleavage. Therefore, a catalytically inactive endonuclease is an endonuclease that retains nucleic acid binding function without possessing cleavage activity. Catalytically inactive endonucleases may be used for the selection step of the present method. In some embodiments, catalytically inactive endonucleases are used to deplete unwanted samples. In some embodiments, catalytically inactive endonucleases are used to concentrate a desired sample. In some embodiments, catalytically inactive endonucleases are directly or indirectly bound to a solid support. In some embodiments, catalytically active endonucleases are bound to a solid support via biotin-streptavidin interactions.
[0625] Furthermore, those skilled in the art will recognize the catalytic domain of endonucleases and may be able to design mutations from wild-type endonucleases to produce catalytically inactive endonucleases (see Maeder et al., Nat Methods 10(10):977-979 (2013)). Such designed catalytically inactive endonucleases can be tested to confirm their lack of cleavage activity. A representative catalytically inactive Cas9 protein is disclosed in U.S. Patent No. 1,0457,969, which is incorporated herein by reference in its entirety.
[0626] In some embodiments, the endonucleases are catalytically active endonucleases, meaning they can cleave nucleic acids. In some embodiments, the catalytically active endonucleases are used to deplete unwanted samples.
[0627] In some embodiments, the endonuclease associates with a guide RNA. The endonuclease can be targeted by the guide RNA to one or more target nucleic acid sequences. In some embodiments, the target nucleic acid sequences are one or more unique sample barcodes.
[0628] In some embodiments, the endonuclease has minimal PAM specificity (shown in Figure 8), which allows for greater flexibility in designing the guide RNA.
[0629] In some embodiments, the endonuclease associates with a guide RNA that binds to one or more unique sample barcodes. In some embodiments, the guide RNA is directed towards a unique sample barcode associated with the nucleic acid of an unwanted sample. In some embodiments, the guide RNA is directed towards a unique sample barcode associated with the nucleic acid of a desired sample.
[0630] In some embodiments, the endonuclease is derived from the cyanobacterium Scytonema hofmanni (ShCAST). ShCAST is a four-protein system for RNA-directed (sgRNA) DNA transposition mediated by a Tn7-like transposase subunit and a VK-type CRISPR effector (Cas12k) (see Strecker et al., Science. 365(6448):48-53 (2019), including the embodiment shown in Figure 5 of Strecker). Other systems in which a Tn7-like transposon utilizes a nuclease-deficient CRISPR-Cas system to generate CRISPR-associated transposases have also been described (see Klompe et al., Nature 571:219-225 (2019)).
[0631] Figure 8 shows several different methods for increasing the specificity of endonucleases. The methods described herein may use any type of endonuclease and / or guide RNA that can improve specificity. In some embodiments, the improved specificity of the endonuclease is due to improved binding of the endonuclease to one or more unique sample barcodes. Such improved binding may be a higher percentage of binding to one or more unique sample barcodes of interest (i.e., specific binding) compared to binding to other sequences (i.e., nonspecific binding).
[0632] In some embodiments, catalytically active endonucleases are those that have higher specificity for cleaving nucleic acids. In some embodiments, this higher specificity is not solely due to higher specificity in binding to target sequences in nucleic acids. In some embodiments, these catalytically active endonucleases with higher specificity can cleave unwanted samples and deplete them from the sample.
[0633] In some embodiments, the catalytically active endonucleases are higher-fidelity mutants. A "high-fidelity" endonuclease refers to one that exhibits reduced off-target activity compared to the wild-type endonuclease.
[0634] In some embodiments, a catalytically active endonuclease is contained within the fusion protein together with a FokI nuclease. In some embodiments, the fusion protein contains Cas9 and a FokI nuclease (see Guilinger et al., Nat Biotechnol. 32(6):577-582 (2014)). Such a fusion protein may act to require the binding of two separate fusion proteins containing a catalytically inactive Cas9 fused in close proximity to the FokI nuclease (as shown in Figure 8), after which the dimerized FokI nuclease can cleave the target nucleic acid. In some embodiments, the two fusion proteins bind to different target sequences. In some embodiments, the two fusion proteins bind to two different unique sample barcodes.
[0635] G. Concentration By using a variety of different concentration methods, it is possible to select the desired sample while excluding unwanted samples. In this way, only the desired sample is resequenced without resequencing the unwanted sample.
[0636] In some embodiments, depletion refers to the physical separation of unwanted samples from the desired sample. In some embodiments, depletion includes capturing the desired sample on a solid support and discarding any uncaptured sequences. Such a capture step can avoid capturing unwanted samples, which are then discarded. After such a concentration step, only the desired sample remains in the library.
[0637] In some embodiments, the enrichment step includes hybrid capture, unique sample barcode-specific amplification, or capture by a catalytically inactive endonuclease. In some embodiments, a unique sample barcode is used to direct the enrichment of a desired sample. In some embodiments, a unique sample barcode is used to direct the enrichment of a desired sample from one or more single cells derived from a mixed pool of cells.
[0638] In some embodiments, multiple concentration steps are performed. In some embodiments, multiple steps include the same type of concentration. For example, two or more hybrid capture steps are performed, in which different hybrid capture oligonucleotides may be used in different steps.
[0639] In some embodiments, the enrichment process may involve multiple steps of different types. For example, enrichment may be performed by hybrid capture followed by PCR amplification.
[0640] In some embodiments, sequencing may be performed between multiple concentration steps. Such sequencing results may indicate a desired sample to be further concentrated.
[0641] In some embodiments, selection is carried out by combining a concentration step and a depletion step. In other words, any combination of selection steps described herein can be combined by the user.
[0642] 1. Hybrid Capture In some embodiments, the enrichment step includes hybrid capture. In some embodiments, the hybrid capture step includes hybridizing hybrid capture oligonucleotides to unique sample barcodes. This step can be carried out using several hybrid capture oligonucleotides bound to a set of unique sample barcodes, where the unique sample barcodes represent the unique sample barcodes of several desired samples. For example, initial sequence data may indicate that a set of single cells in a mixed pool of cells express a given gene mutation, and the unique sample barcodes associated with these single cells can be used for hybrid capture to enrich a nucleic acid library from these particular single cells. After enrichment, resequencing can be performed to generate additional sequence data for the single cells of interest. This method avoids the generation of additional sequence data for unwanted cells because samples from unwanted cells are not enriched during the hybrid capture step.
[0643] In some embodiments, a unique sample barcode is selected to hybridize with a known panel of hybrid capture oligonucleotides. Alternatively, a custom panel of hybrid capture oligonucleotides may be generated based on a unique sample barcode used when preparing the nucleic acid library.
[0644] In some embodiments, hybrid capture oligonucleotides are bound to affinity elements. In some embodiments, affinity elements are used to enable the capture of oligonucleotides bound to specific unique sample barcodes, thereby enabling the enrichment of libraries containing these unique sample barcodes. In some embodiments, the affinity element is biotin. Various affinity elements are known to those skilled in the art, such as magnetic microparticles that can be bound by certain types of capture beads.
[0645] In some embodiments, the hybrid capture oligonucleotide is bound directly or indirectly to a solid support. In some embodiments, the hybrid capture oligonucleotide is bound to the solid support via biotin-streptavidin interactions. In some embodiments, the solid support is a bead.
[0646] 2. Capture by catalytically inactive endonucleases Similar to hybrid capture, catalytically inactive endonucleases associated with specific guide RNAs can be used for enrichment. These catalytically inactive endonucleases can be targeted to specific, unique sample barcodes using the guide RNA. In some embodiments, capture by catalytically inactive endonucleases involves binding the catalytically inactive endonuclease to the unique sample barcode via the guide RNA.
[0647] In some embodiments, catalytically inactive endonucleases are bound to affinity elements. In some embodiments, affinity elements are used to enable the capture of catalytically inactive endonucleases bound to specific unique sample barcodes, thereby enabling the enrichment of libraries containing these unique sample barcodes. In some embodiments, the affinity element is biotin. Various affinity elements are known to those skilled in the art, such as magnetic nanoparticles that can be bound by certain capture beads.
[0648] In some embodiments, a catalytically inactive endonuclease is bound directly or indirectly to a solid support. In some embodiments, the catalytically inactive endonuclease is bound to the solid support via biotin-streptavidin interactions. In some embodiments, the solid support is a bead.
[0649] 3. PCR amplification In some embodiments, enrichment is performed by PCR amplification. In some embodiments, enrichment is performed by PCR amplification targeting specific sample barcodes. In some embodiments, primers that bind to specific specific sample barcodes enable amplification of the desired sample based on specific sample barcodes known to be associated with the desired sample from initial sequencing. In contrast, primers that bind to other specific sample barcodes associated with unwanted samples are not included in the amplification reaction. In this way, the desired sample can be selected.
[0650] H. Depletion Using a variety of depletion methods, unwanted samples can be removed without removing the desired samples. In this way, only the desired samples are resequenced, without resequencing the unwanted samples.
[0651] In some embodiments, the depletion step includes hybrid capture, capture by a catalytically inactive endonuclease, or CRISPR digestion.
[0652] In some embodiments, a unique sample barcode is used to direct the depletion of unwanted samples. In some embodiments, a unique sample barcode is used to direct the depletion of unwanted samples from one or more single cells derived from a mixed pool of cells.
[0653] In some embodiments, multiple depletion steps are performed. In some embodiments, the multiple steps include the same type of depletion. In some embodiments, the multiple steps of concentration include different types of depletion. For example, CRISPR digestion may be performed after depletion by hybrid capture. In some embodiments, sequencing may be performed between depletion steps. For example, the method may include initial targeted sequencing, depletion of unwanted samples, another targeted sequencing, further depletion of unwanted samples, and comprehensive resequencing.
[0654] 1. Depletion by physically separating unwanted samples from the desired sample. In some embodiments, depletion refers to the physical separation of unwanted samples from the desired samples. In some embodiments, depletion includes capturing unwanted samples on a solid support and removing them. After such a depletion step, only the desired samples remain in the library.
[0655] In some embodiments, hybrid capture may be carried out as described for the concentration of a desired sample, except that unwanted samples isolated by hybrid capture are removed from further resequencing (instead of being retained for resequencing, as in the case of the desired sample in the concentration embodiment).
[0656] In some embodiments, capture by catalytically inactive endonuclease capture may be carried out as described for the concentration of a desired sample, except that unwanted samples isolated by capture by catalytically inactive endonuclease capture are removed from further resequencing (instead of being retained for resequencing, as in the case of the desired sample in the concentration embodiment).
[0657] 2. Depletion due to cutting of unnecessary samples In some embodiments, depletion involves cleavage that makes it impossible to properly sequence unwanted samples. In other words, depletion may refer to a situation where, based on sample cleavage, unwanted samples have little to no ability to be properly sequenced. In some embodiments, nucleic acids derived from unwanted samples are present in the library and selection, but depletion refers to a reduction in the ability of these unwanted samples to be sequenced.
[0658] For example, cutting sequences within or near one or more unique sample barcodes associated with unwanted samples can separate the nucleic acid sequences necessary for sequencing from the remaining unwanted samples. In this way, the unwanted samples can no longer generate sequencing results in resequencing after depletion. In some embodiments, such cutting separates nucleic acid sequences from the remaining unwanted samples. In some embodiments, the separated nucleic acid sequences are adapter sequences. In some embodiments, such adapter sequences may be primer sequences or sequences for immobilizing nucleic acids on flow cells used for sequencing. For example, separating sequencing primer binding sites from the remaining unwanted samples can prevent the unwanted samples from being sequenced via a selected sequencing method. Those skilled in the art can identify such sequences that can be separated to mediate depletion based on the platform used for sequencing and the composition of the initially generated library.
[0659] In some embodiments, the depletion step includes CRISPR digestion. As used herein, CRISPR (clustered regularly interspaced short palindromic repeats) refers to a family of DNA sequences found in the genomes of prokaryotes such as bacteria and archaea. As used herein, CRISPR digestion refers to any digestion of one or more nucleic acids based on CRISPR sequences. Endonucleases, such as Cas9, can utilize CRISPR sequences to cleave nucleic acids at a defined sequence. In some embodiments, the endonuclease is a catalytically active endonuclease.
[0660] In some embodiments, CRISPR digestion is directed towards a unique sample barcode associated with the nucleic acid of the unwanted sample. In some embodiments, CRISPR digestion includes cutting the unwanted sample. In some embodiments, CRISPR digestion depletes the unwanted sample by separating the nucleic acid sequence required for sequencing from the remaining unwanted sample.
[0661] a) Method for cutting unwanted samples using ShCAST In some embodiments, the depletion method is carried out using ShCAST cutting. In some embodiments, cutting makes it impossible to amplify and / or sequence unwanted samples.
[0662] In some embodiments, ShCAST comprises Cas12K, the transposase comprises a Tn5 or Tn7-like transposase, and / or at least one of the gRNA and transposase is biotinylated, and at least one of the biotinylated gRNA and transposase can be coupled to streptavidin-coated beads. In some embodiments, the biotinylated gRNA and / or transposase allows for the capture of unwanted samples onto the streptavidin beads. In this way, unwanted samples can be removed from the reaction mixture while retaining the desired samples.
[0663] In some embodiments, a fluid (also known as a reaction fluid) is used to restrict the binding of the transposase contained in ShCAST. In some embodiments, restricting or inhibiting the binding of the transposase reduces off-target rearrangement reactions mediated by the transposase contained in ShCAST. When off-target cleavage is reduced, the depletion step may be more selective in depleting only the unwanted sample without affecting the desired sample.
[0664] In some embodiments, depletion of nucleic acid samples derived from unwanted samples is carried out in a fluid having conditions for limiting cleavage by the complex. Those skilled in the art are aware of several means for limiting cleavage by transposase-mediated transposition reactions, and any means known in the art can be used. For example, transposase activity is dose-dependent (i.e., the lower the concentration of transposase, the more limited the number of transposition reactions). In addition, transposase is magnesium-dependent. In some embodiments, the conditions for limiting cleavage by the complex are a magnesium concentration of 15 mM or less, and / or a Cas12K and / or transposase concentration of 50 nM or less.
[0665] In some embodiments, the cleavage of nucleic acids by ShCAST takes into account the timing of the process. For example, the user may want to limit the binding and / or cleavage of nucleic acids by ShCAST in the initial reaction steps to allow for higher selectivity (e.g., cleaving unwanted samples and not cleaving desired samples). In later reaction steps, the user may want to promote the cleavage of nucleic acids by transposases contained in the complex for efficient cleavage of unwanted samples. In other words, the user may want the cleavage of nucleic acids by transposases to occur with relatively high efficiency, while the binding of transposases is relatively selective. Therefore, initial conditions during the hybridization of the complex to nucleic acids may inhibit the binding of transposases contained in the complex to the nucleic acids and / or inhibit cleavage by transposases contained in the complex. Conditions in subsequent steps may promote the cleavage of nucleic acids by transposases.
[0666] In some embodiments, depleting nucleic acid samples derived from unwanted samples includes (1) binding the complex to double-stranded nucleic acids under conditions that inhibit nucleic acid cleavage by the complex, and (2) promoting nucleic acid cleavage by the complex after binding.
[0667] In some embodiments, binding occurs under conditions that (1) inhibit the binding of the complex to the target nucleic acid and (2) inhibit the cleavage of the target nucleic acid by the complex. In other words, the initial conditions can inhibit both the binding of the complex and the cleavage by the complex.
[0668] In some embodiments, different means of selective activation of the transposase may be used. In some embodiments, during binding, the transposase contained in ShCAST is inactive or has low activity depending on the reaction conditions used. In some embodiments, the reaction conditions are changed after ShCAST has bound to the nucleic acid, allowing for more efficient cleavage by the transposase after more selective binding of ShCAST. In such embodiments, a reversibly inactivated transposase may be used, and the user can control the time the transposase is active by using a selective activation step. Although means of selective activation of such transposases are described for ShCAST, these methods can be used in conjunction with other methods of incorporating transposases.
[0669] In some embodiments, the transposase is reversibly inactivated during binding, and cleavage is facilitated by activation of the transposase.
[0670] In some embodiments, the magnesium concentration during bonding is low (e.g., less than 15 mM), and promoting cleavage involves increasing the magnesium concentration.
[0671] In some embodiments, the transposase is not present during binding, and promoting cleavage involves adding the transposase.
[0672] In some embodiments, a transposase is reversibly inactivated by the absence of one or more transposons, and activating a transposase involves providing one or more transposons.
[0673] VII. Typical Uses of the Method The method of the present invention can be used in a variety of sequencing applications. The specific uses described herein are not intended to limit the invention, and those skilled in the art can imagine a wide range of ways in which the method can be used to improve the results of various sequencing applications.
[0674] A. Correction of library quality control In some embodiments, the method may be used for quality control (QC) of libraries containing multiple nucleic acid samples derived from a mixed pool of samples. In some embodiments, a concentration or depletion step is used for quality control. In some embodiments, the quality control step is corrective in that it reduces signals from unwanted samples. Figure 2 provides an overview of how current single-cell methods may lose rare cell-derived information from metagenomics samples without the quality control steps described herein.
[0675] As used herein, “quality control” or “QC” refers to a selection process based on the properties of the library obtained from various individuals within the library, and not on factors related to the original mixed population of samples. In other words, the QC method does not necessarily identify desired or unwanted samples in a single-cell library based on biological differences between samples in the mixed pool of original samples used to prepare the library; rather, it identifies desired or unwanted samples based on factors related to the prepared library.
[0676] For example, a given library prepared from a single cell may be of lower quality due to random variations in the library generation process, rather than due to biological differences between this cell and other cells in the original mixed pool of cells. Unwanted samples may include single-cell libraries with an insufficient number of fragments, or those with fragments of undesirable size. Any factor that can degrade the quality of sequencing results may result in certain nucleic acid libraries being classified as unwanted samples. In other words, those skilled in the art can use the method of the present invention to correct low-quality library preparation (where some samples are noise and scattered, related to unique sample barcodes), remove unwanted samples from the library, and then perform resequencing. This resequencing can then focus on libraries that are potentially capable of generating sequence data of sufficient quality.
[0677] In some embodiments, initial sequencing identifies desired and undesirable libraries based on the quality of the sequencing results.
[0678] In some embodiments, the initial sequencing reaction identifies unique sample barcodes associated with single-cell libraries that are unwanted samples, because these libraries are of low quality. In some embodiments, libraries of unwanted samples are identified by initial sequencing, and these libraries are depleted from the sc library before resequencing. In some embodiments, libraries of desired samples are identified by initial sequencing to identify higher-quality libraries, and these libraries are enriched from the sc library before resequencing.
[0679] In some embodiments, the quality control process improves the quality of the library used for resequencing. In this way, resequencing can focus on deeper sequencing of the higher-quality library. In some embodiments, the QC process can avoid wasting time and reagents by avoiding deeper sequencing of low-quality libraries (i.e., unwanted samples).
[0680] B. Oncological use In some embodiments, this method is used to evaluate or monitor a disease. In some embodiments, the disease is cancer.
[0681] In some embodiments, cancer is a blood or solid tumor. In some embodiments, cancer can be assessed based on a biopsy from a solid tumor or blood sample. In some embodiments, the method is used to assess heterogeneous tumors or to assess circulating cancer cells (CTCs). CTCs are a predictive marker of tumor prognosis and can be useful in assessing a subject's response to a given treatment (such as chemotherapy or immunotherapy).
[0682] In some embodiments, this method is used to evaluate cells in the tumor microenvironment, which may or may not be cancer cells. These non-cancer cells may be stromal cells, vascular cells, or any other type of cell that is not cancerous itself but may be in close proximity to cancer cells. Cells in the tumor microenvironment are known to influence tumor growth and metastasis.
[0683] In some embodiments, initial sequencing evaluates libraries within the sc library via targeted sequencing against mutant cells. These mutant cells may have single nucleotide polymorphisms, insertions, deletions, and / or copy number variations in their nucleic acids. These mutant cells may also have differences in other factors, such as methylation changes. In some embodiments, these variants are CTCs. Based on the initial sequencing, a selection step can be performed to enrich or deplete the mutant cells to obtain sc libraries containing the desired cellular nucleic acid libraries. These libraries can then be used in a resequencing step for deeper genomic characterization of the mutant cells.
[0684] In some embodiments, initial sequencing is targeted sequencing of somatic driver mutation regions. Somatic driver mutations are mutations that give a proliferative advantage to cells expressing them, and these cells can be positively selected during cancer evolution. In some embodiments, initial sequencing assigns cancer / molecular types to individual cell nucleic acid libraries tagged with a given unique sample barcode within multiple cell nucleic acid libraries. In some embodiments, deeper resequencing is performed after selection of libraries tagged with unique sample barcodes associated with driver mutations.
[0685] In some embodiments, the somatic driver mutation is a mutation in KRAS G12. In some embodiments, initial sequencing is targeted sequencing of KRAS G12. In some embodiments, analysis is performed to determine the UBC barcode of individual cell nucleic acid libraries containing the KRAS G12 mutation (shown in Figure 7). In some embodiments, after selecting these libraries of interest, resequencing is performed to better understand the profiles of cells containing KRAS G12, which may be deeper sequencing or whole-genome sequencing. Similar protocols can be used to select and evaluate sequence data from cells expressing any other mutations of interest.
[0686] In some embodiments, the method is used to track tumor evolution. As used herein, “tumor evolution” refers to changes in the characteristics of cancer cells over time, and tracking tumor evolution may include characterizing cellular evolutionary patterns. For example, tumors are heterogeneous, and over time, as certain traits are selected, changes in tumor characteristics are considered due to this heterogeneity within the tumor. Changes in tumor characteristics may allow the tumor to evolve to have faster growth or metastasis, or to have resistance to a given treatment.
[0687] For example, if a target tumor acquires resistance to a given chemotherapy, treatment with this drug may no longer act to slow or halt tumor growth. Methods described herein can be selected to deeply sequence cells of interest in order to assess the presence or acquisition of resistance to a given treatment. In this way, the treatment plan for the target can be optimized to focus on treatments that are likely to be effective for the target and to avoid treatments that are less likely to be effective.
[0688] C. Metagenomic use This method may be used for metagenomics. As used herein, “metagenomics” refers to the study of genetic material recovered directly from environmental samples. In some embodiments, these environmental samples include two or more types of microorganisms. As used herein, microorganisms may include bacteria, viruses, fungi, or other small organisms. For example, a metagenomic sample may include a group of microorganisms (such as various bacteria).
[0689] In some embodiments, metagenomics analysis avoids culturing organisms. In other words, metagenomic samples can be evaluated without first culturing and artificially growing them. By avoiding culturing, selective pressure on organisms that do not grow sufficiently in culture can be avoided. Furthermore, avoiding culturing can be particularly important when there is little knowledge about the target microorganism, such as appropriate culture conditions. Otherwise, the target microorganism may be selected by the culture conditions and lost from the mixed population before sequencing if other microbial cultures grow better.
[0690] Conventional methods make de novo assembly and species identification of rare and unculturable microorganisms virtually impossible (see Malmstrom and Eloe-Fadrosh mSystems 4:e00118-19 (2019)). Conventional approaches have involved isolating a single amplified genome (SAG) by cell distribution (i.e., FACS, microfluidics), followed by cell lysis and whole-genome analysis (Approach 1). Another approach was metagenomic assembled genome (MAG) analysis, short / long-read shotgun sequencing using differential binning by coverage, and tetranucleotide frequency analysis (Approach 2). Yet another approach is the “mini-metagenomic” hybrid approach (Quake lab, MetaSort) (Approach 3).
[0691] However, these approaches in the field are best suited for the assembly and identification of abundant species in low-diversity samples. Diversity can refer to the number of different species in a sample. In other words, conventional metagenomics methods have been limited in their use for the assembly and identification of rare species that are not commonly found in highly diverse samples.
[0692] For example, Approach 1 can only be used with empirical knowledge of classifiable phenotypes to deplete abundant species and enrich rare species. Furthermore, cell distribution in Approach 1 cannot be performed if there are no enrichable or divisible features. In addition, all prior art methods may involve exorbitant sequencing costs to fully characterize microbiome samples.
[0693] In contrast, this method can be used to select desired samples based on initial sequencing. These desired samples may be cellular nucleic acid libraries derived from target microorganisms within a metagenomics sample. After selection, resequencing can be performed by enrichment or depletion to provide deeper sequence data for these target microorganisms.
[0694] In some embodiments, the method uniquely barcodes the DNA (RNA) of each organism in a microbiome sample, and as a result, after initial sequencing and analysis, it is physically possible to enrich the desired cellular nucleic acid library or deplete the unwanted cellular nucleic acid library.
[0695] In some embodiments, initial sequencing focuses on targeted sequencing. In some embodiments, initial sequencing is ribosomal RNA or DNA (rRNA or rDNA) sequencing. In some embodiments, initial sequencing is 16S, 18S, or internal transcription spacer sequencing. In some embodiments, initial sequencing assigns taxonomic identification to cellular RNA / DNA tagged with a given barcode within multiple cellular nucleic acid libraries. In some embodiments, this targeted sequencing is prokaryotic 16S rDNA or rRNA sequencing. Sequencing of the variable region of 16S rRNA is frequently used for phylogenetic classification, such as genus or species, in diverse microbial populations.
[0696] In some embodiments, an initial sequencing reaction is performed, followed by analyses such as the determination of rich species / taxa from 16s rDNA analysis (see Figure 7 for an example of such targeted sequencing). For example, the initial sequencing may be 16s rRNA sequencing of the entire cellular nucleic acid library, followed by whole-genome sequencing of the desired cellular nucleic acid library after a selection step. Such a method can save time and cost by concentrating deep sequencing on libraries derived from the target microorganism.
[0697] In some embodiments, initial sequencing is performed using continuity-preserved transposition sequencing. In some embodiments, continuity-preserved transposition sequencing is used when the sample contains a significant amount of intact single chromosomes or high molecular weight genomes after extraction.
[0698] In some embodiments, metagenomics can be used to evaluate samples taken from patients. In some embodiments, the sample may be taken from a patient exhibiting symptoms of an unknown infection. In some embodiments, the sample may be a microbiome sample (such as a fecal sample for evaluating the microbiome of interest). As used herein, a microbiome sample refers to an aggregate of microorganisms present on or in human tissue or biological fluids.
[0699] D. Immunological use In some embodiments, the method is used for immunological analysis. In some embodiments, the method is used to evaluate T cell clonal types. The composition of T cell clonal types in a given individual may be referred to as the T cell repertoire. In some embodiments, initial sequencing characterizes the TCR repertoire. In some embodiments, a selection step depletes abundant T cell clonal types. In some embodiments, resequencing is used for deeper sequencing of less common T cell clonal types. [Examples]
[0700] Example 1. Enrichment from a Sci-RNA3 library or other Sci-RNA libraries A wide range of different methods for generating single-cell libraries (SC libraries) are known in the art. The method of the present invention can be used in conjunction with any of these different methods for generating SC libraries based on specific indices contained in library fragments.
[0701] For example, a single-cell sequencing library can be generated using sci-RNA-seq3, as shown in Figure 4 (see Cao et al., Nature 566(7745):496-502(2019)). This method utilizes the RT index (BCRT) and ligation adapter index (BCLIG) along with the i5 and i7 indices. The i5 and i7 indices are a commercially available set of 96 unique adapters (Illumina).
[0702] The RT index can be combined with the hairpin adapter index (oligoTp). Multiple indices enable read demultiplexing, such as removing duplicates based on reads having the same UMI, RT index, ligation adapter index, and tagmentation site. Figure 4 shows different indices (i.e., barcodes) used as black ovals: BCRT (10 nucleotides), BCLIG (10 nucleotides), i5 (8 nucleotides), and i7.
[0703] Various different methods can be used for enrichment along with the sc-RNA-seq3 (Sci-RNA3) library generated by the sc-RNA-seq3 method.
[0704] Firstly, a probe capture approach that avoids i7 selection can be used. Based on the nucleotides contained in the i5, BCLIG, and BCRT indices, a total of 28 bases represent specific hybridization bases for developing a capture probe, and a total of 67 nucleotides are available for hybridization (including 33 nucleotides for the R1 primer and 6 nucleotides for the fixed region). In this calculation, the capture probe includes a universal sequence for binding to the UMI sequence.
[0705] Secondly, a nested PCR approach may be used. In this approach, PCR for enriching the desired sample is performed using i7 primers along with primers that bind to selected i5, BCLIG, and BCRT indices. In this approach, the library can be designed by swapping the positions of BCRT and UMI in the library fragments so that the nested PCR approach using BCRT retains the UMI sequence in the resulting PCR product.
[0706] Thirdly, a combination approach can be used. In the combination approach, the probe capture and concentration step is followed by the i7-specific PCR concentration step.
[0707] These specific approaches utilize sci-RNA-seq3 library designs, but barcodes / indices used in other types of sc libraries can also be leveraged for the enrichment process. These sc libraries include BioRad-ddSEQ, 10X Genomics, InDrop, Drop-Seq, and Split-Seq. As shown in Figure 4, enrichment protocols can be designed using the specific barcode structure of the library (including the number of nucleotides in different barcode regions). Those skilled in the art can use information on various methods to design the most appropriate approach for enrichment based on the specific sc library used for initial sequencing.
[0708] Example 2. Modified SCI-seq approach for generating library fragments containing sequential barcodes Using a modified SCI-seq approach, single-cell RNA / DNA NGS libraries containing sequential barcodes can be generated, as shown in Figure 5.
[0709] In the first step, tagmentation is performed using a transposomal complex containing a Tn5 transposase with a transposon containing the BC1 sequence to incorporate the BC1 barcode. Cells or nuclei are distributed into the reaction wells. If the starting target nucleic acid is RNA, cDNA synthesis is performed to generate the first and second strands. Tagmentation is performed using well-specific barcodes (BC1 barcodes). DNA is pooled from the entire well. Gap repair is performed (3' fill-in), followed by 5' phosphorylation and generation of the 3'A tail end.
[0710] In the second step, T / A ligation is performed using one or more barcodes (BC2, ..., BCx). These barcodes do not have to be random. In this step, nuclei or cells are redistributed into reaction wells, followed by ligation of the T-tail adapter using a well-specific barcode (BC2 barcode). DNA is pooled from the entire well, followed by 5' phosphorylation and 3'A tail generation. Alternatively, the library fragment may have C / G overhangs for subsequent C / G ligation (used in all other barcoding steps). These steps are repeated in multiple barcoding steps as needed.
[0711] In the third step, T / A ligation is performed to generate the desired fragment with a BCn barcode. In this step, nuclei or cells are redistributed into reaction wells, and T-tailed Y-type adapters are ligated with well-specific barcodes. Then, DNA is pooled from the entire well and PCR is performed using a sample index.
[0712] During sc library generation, the library does not need to be fully constructed. Short, asymmetric ends can improve the specificity of hybridization and / or PCR results.
[0713] The resulting library can then be used for initial sequencing and subsequently enriched or depleted based on the serial barcodes present in the library fragments. The presence of serial barcodes can improve subsequent enrichment by PCR, as primers can be designed across the complete serial barcode.
[0714] Example 3. Method for use with microbial cells distributed in a metagenomics sample. This method can be used for metagenomics, such as genome assembly of organisms that are not cultured. These organisms may be microbial cells, such as those found in samples taken from patients.
[0715] For this method, cells are distributed into wells, and BC1 (only) is inserted by tagmentation. After pooling the DNA, it is extended and blunt-ended to produce A-tails. The sample is then distributed into appropriate DNA dilutions.
[0716] Next, T / A ligation is performed using a T-tail adapter containing BC2. DNA is pooled, extended, and blunt-ended to generate an A-tail. These steps are repeated to incorporate the desired number of barcodes (BCn).
[0717] For final ligation, a fork-type adapter is added, followed by PCR to add the i5 / i7 and P5 / P7 sequences. The P5 and P7 sequences are useful for sequencing methods using the Illumina platform, but other sequences may be added if sequencing is performed on other platforms.
[0718] An initi...
Claims
1. A targeted transposomal complex, a. Transposase and, b. The first transposon, i. 3' transposon terminal sequence and ii. A first transposon including a 5' adapter array, c. A catalytically inactive endonuclease associated with a guide RNA, wherein the catalytically inactive endonuclease is derived from the cyanobacterium Scytonema hofmanni (ShCAST), ShCAST contains Cas12K, the guide RNA is capable of leading the binding of the endonuclease to one or more target nucleic acid sequences, the guide RNA is biotinylated, and the biotinylated guide RNA is coupled to streptavidin-coated beads, d. A second transposon comprising a complement to the transposon terminal sequence, A targeted transposomal complex.
2. a. The transposase comprises a Tn5 or Tn7-like transposase and / or b. The targeted transposome complex according to claim 1, wherein the first transposon comprises at least one of a P5 adapter and a P7 adapter.
3. The targeted transposome complex according to claim 1, wherein the catalytically inactive endonuclease and transposase are contained in the fusion protein.
4. The targeted transposome complex according to claim 1, wherein the catalytically inactive endonuclease and transposase are linked via a linker.
5. The targeted transposome complex according to claim 1, wherein the sequence in the guide RNA that binds to the target nucleic acid sequence contains fewer than 20 nucleotides.
6. The targeted transposome complex according to claim 5, wherein the sequence in the guide RNA that binds to the target nucleic acid sequence comprises 15, 16, 17, 18, or 19 nucleotides.
7. The targeted transposome complex according to claim 1, wherein the guide RNA includes a hairpin secondary structure.
8. The targeted transposome complex according to claim 1, wherein the 5' adapter sequence includes a primer sequence, an index tag sequence, a capture sequence, a barcode sequence, a cleavage sequence, or a sequencing-related sequence, or a combination thereof.
9. The targeted transposome complex according to claim 1, wherein the targeted transposome complex is in a solution.
10. The targeted transposome complex according to claim 9, wherein the solution contains a magnesium concentration of 15 mM or more.
11. The targeted transposome complex according to claim 9, wherein the solution contains a magnesium concentration of 15 mM or less.
12. A targeted transposomal complex, a. Transposase and, b. The first transposon, i. 3' transposon terminal sequence and ii. A first transposon including a 5' adapter array, c. A catalytically inactive endonuclease associated with a guide RNA, wherein the catalytically inactive endonuclease is derived from the cyanobacterium Scytonema hofmanni (ShCAST), and ShCAST contains Cas12K, and the guide RNA is capable of leading the binding of the endonuclease to one or more target nucleic acid sequences. d. A second transposon comprising a complement to the transposon terminal sequence, A targeted transposome complex immobilized on a solid support.
13. The targeted transposome complex according to claim 12, wherein the solid support is a bead.
14. The targeted transposome complex according to claim 12, wherein the catalytically inactive endonuclease and transposase are contained in the fusion protein.
15. The targeted transposome complex according to claim 12, wherein the catalytically inactive endonuclease and transposase are linked via a linker.
Citation Information
Patent Citations
Methods and compositions for analyzing cellular constituents
JP2018509140A
Polynucleotide enrichment using crispr-CAS systems
WO2016014409A1
Use of transposase and y adapters to fragment and tag DNA
WO2017171985A1
Tagmentation using immobilized transposomes with linkers
WO2018156519A1
RNA-guided DNA integration using tn7-like transposons
WO2020181264A1