Modified transposase for improved insertion sequence bias and increased DNA input tolerance
By mutation of Tn5 transposase in specific amino acid sequences, its insertion sequence bias and DNA input tolerance are improved, and the shortcomings of transposases in the prior art are solved in these aspects, achieving more efficient nucleic acid sample fragmentation and labeling effects.
Patent Information
- Application Number
- CN202510139371.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2014-11-17
- Filing Date
- 2015-04-15
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, transposases have shortcomings in insertion sequence bias and DNA input tolerance, resulting in inefficiency in fragmentation and tagging of nucleic acid samples.
The insertion sequence bias and DNA input tolerance of the transposase are improved by mutations of specific amino acid sequences, such as replacement or insertion mutations at Asp248, Asp119, Trp125, Lys120, Lys212, Pro214, Gly251, Ala338, Glu146 and Glu190 positions.
The improved transposase can fragment DNA more evenly and maintain high fragment size consistency at different DNA inputs, improving the labeling efficiency of nucleic acid samples and the quality of sequencing results.
Smart Images

Figure BDA0005264299730000191 
Figure BDA0005264299730000201 
Figure BDA0005264299730000211
Abstract
Description
[0001] This application is a divisional application of an application with an application date of April 15, 2015, application number 202110295229.2, and invention name “Modified transposase for improving insertion sequence bias and increasing DNA input tolerance”.
[0002] The application with the application date of April 15, 2015, application number 202110295229.2, and invention name “Modified transposase for improving insertion sequence bias and increasing DNA input tolerance” is a divisional application of the application with the application date of April 15, 2015, application number 201580027103.X, and invention name “Modified transposase for improving insertion sequence bias and increasing DNA input tolerance”.
[0003] Related Applications
[0004] This application claims priority to U.S. Provisional Application No. 61 / 979,871, filed April 15, 2014; U.S. Provisional Application No. 62 / 062,006, filed October 9, 2014; and U.S. Provisional Application No. 62 / 080,882, filed November 17, 2014, which are hereby incorporated by reference in their entireties. Technical Field
[0005] The present application relates to, but is not limited to, modified transposases for improved insertion sequence bias and increased DNA import tolerance. background
[0007] Transposases are useful in in vitro transposition systems. They allow large-scale fragmentation and tagging of genomic DNA and are used to prepare libraries of tagged DNA fragments from target DNA for use in nucleic acid analysis methods such as next generation sequencing and amplification methods. There remains a need for modified transposases that have improved properties and produce tagged DNA fragments that qualitatively and quantitatively represent the target nucleic acid in the sample from which the target nucleic acid was produced.
[0008] Sequence Listing
[0009] This application is submitted with a sequence listing in electronic format. The sequence listing is provided as a file named IP-1198-TW_Sequence_Listing.TXT, created on April 13, 2015, and is 106Kb in size. The information in the electronic sequence listing is incorporated herein by reference in its entirety.
[0010] Overview
[0011] Provided herein are transposases for improved fragmentation and tagging of nucleic acid samples.The inventors have unexpectedly identified certain altered transposases that exhibit improved insertion sequence bias and have numerous other associated advantages.
[0012] Provided herein are mutant Tn5 transposases modified relative to wild-type Tn5 transposase. In some embodiments, the mutant transposase may include a mutation at position Asp248. In certain aspects, the mutation at position Asp248 is a substitution mutation. In certain aspects, the substitution mutation at position Asp248 may include a mutation to a residue selected from the group consisting of Tyr, Thr, Lys, Ser, Leu, Ala, Trp, Pro, Gln, Arg, Phe, and His.
[0013] In some aspects, the mutation at position Asp248 is an insertion mutation after position Asp248. In some aspects, the insertion mutation may include inserting a hydrophobic residue after position Asp248. In some aspects, the insertion mutation may include inserting a valine residue after position Asp248.
[0014] Also provided herein is a mutant Tn5 transposase modified relative to a wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Asp119. In certain aspects, the mutation at position Asp119 is a substitution mutation. In certain aspects, the substitution mutation at position Asp119 may include a mutation to a hydrophobic residue. In certain aspects, the substitution mutation at position Asp119 may include a mutation to a hydrophilic residue. In certain aspects, the substitution mutation at position Asp119 may include a mutation to a residue selected from the group consisting of Leu, Met, Ser, Ala, and Val.
[0015] Also provided herein is a mutant Tn5 transposase modified relative to a wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Trp125. In certain aspects, the mutation at position Trp125 is a substitution mutation. In certain aspects, the substitution mutation at position Trp125 may include a mutation to a methionine residue.
[0016] Also provided herein is a mutant Tn5 transposase modified relative to a wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Lys120. In certain aspects, the mutation at position Lys120 is a substitution mutation. In certain aspects, the substitution mutation at position Lys120 may include a mutation to a bulky aromatic residue. In certain aspects, the substitution mutation at position Lys120 may include a mutation to a residue selected from the group consisting of: Tyr, Phe, Trp, and Glu.
[0017] Also provided herein is a mutant Tn5 transposase modified relative to a wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Lys212 and / or Pro214 and / or Ala338. In some aspects, one or more mutations at position Lys212 and / or Pro214 and / or Ala338 are substitution mutations. In some aspects, the substitution mutation at position Lys212 comprises a mutation to arginine. In some aspects, the substitution mutation at position Pro214 comprises a mutation to arginine. In some aspects, the substitution mutation at position Ala338 comprises a mutation to valine. In some embodiments, the transposase may also comprise a substitution mutation at Gly251. In some aspects, the substitution mutation at position Gly251 comprises a mutation to arginine.
[0018] Also provided herein is a mutant Tn5 transposase modified relative to a wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Glu146 and / or Glu190 and / or Gly251. In some aspects, one or more mutations at position Glu146 and / or Glu190 and / or Gly251 are substitution mutations. In some aspects, the substitution mutation at position Glu146 may include a mutation to glutamine. In some aspects, the substitution mutation at position Glu190 may include a mutation to glycine. In some aspects, the substitution mutation at position Gly251 may include a mutation to arginine.
[0019] Also provided is an altered transposase comprising a substitution mutation to a semi-conserved domain comprising an amino acid sequence of SEQ ID NO: 21, wherein the substitution mutation comprises a mutation to any residue other than Trp, Asn, Val or Lys at position 2. In certain embodiments, the mutation comprises a substitution to Met at position 2.
[0020] In any of the above described embodiments, the mutant Tn5 transposase may further comprise a substitution mutation at a position functionally equivalent to Glu54 and / or Met56 and / or Leu372 in the Tn5 transposase amino acid sequence. In certain embodiments, the transposase comprises a substitution mutation homologous to Glu54Lys and / or Met56Ala and / or Leu372Pro in the Tn5 transposase amino acid sequence.
[0021] Also provided herein is a mutant Tn5 transposase comprising the amino acid sequence of any one of SEQ ID NOs: 2-10, 12-20, and 25-26.
[0022] Also provided herein is a fusion protein comprising a mutant Tn5 transposase as defined in any of the above embodiments fused to an additional polypeptide. In some embodiments, the polypeptide domain fused to the transposase may comprise a purification tag, an expression tag, a solubility tag, or a combination thereof. In some embodiments, the polypeptide domain fused to the transposase may comprise, for example, maltose binding protein (MBP). In some embodiments, the polypeptide domain fused to the transposase may comprise, for example, elongation factor Ts (Tsf).
[0023] Also provided herein are nucleic acid molecules encoding mutant Tn5 transposases as defined in any of the above embodiments. Also provided herein are expression vectors comprising the nucleic acid molecules described above. Also provided herein are host cells comprising the vectors described above.
[0024] Also provided herein is a method for in vitro transposition, the method comprising: allowing the following components to interact: (i) a transposome complex comprising a mutant Tn5 transposase according to any of the embodiments described above, and (ii) a target DNA.
[0025] Also provided herein is a method for sequencing a target DNA using the Tn5 transposase described above. In some embodiments, the method may include (a) incubating the target DNA with a transposome complex, the transposome complex comprising (1) a mutant Tn5 transposase according to any of the embodiments described above and (2) a first polynucleotide, the first polynucleotide comprising (i) a 3' portion, the 3' portion comprising a transposon end sequence and (ii) a first tag, the first tag comprising a first sequencing tag domain, the incubation being performed under conditions where the target DNA is fragmented and the 3' transposon end sequence of the first polynucleotide is transferred to the 5' end of the fragment, thereby generating double-stranded fragments, the 5' end of the double-stranded fragments being tagged via the first tag, and the 5'-end of the fragments being tagged via the first tag. (a) providing a single-stranded nick at the 3' end of the 5'-tagged strand; (b) incubating the fragments with an enzyme that modifies nucleic acids under conditions whereby a second tag is attached to the 3' end of the 5'-tagged strand; (c) optionally, amplifying the fragments by providing a polymerase and an amplification primer corresponding to a portion of the first polynucleotide, thereby generating a representative library of di-tagged fragments having a first tag at the 5' end and a second tag at the 3' end; (d) providing a first sequencing primer comprising a portion corresponding to a first sequencing tag domain; and (e) extending the first sequencing primer and concurrently detecting the identity of nucleotides adjacent to the first sequencing tag domain of the representative library of di-tagged fragments.
[0026] Also provided herein are kits for performing in vitro transposition reactions. In some embodiments, the kit may include a transposome complex comprising (1) a mutant Tn5 transposase according to any of the embodiments described above; and (2) a polynucleotide comprising a 3' portion comprising a transposon end sequence.
[0027] This article also provides the following:
[0028] Item 1. A modified transposase, comprising a wild-type or mutant transposase fused to a polypeptide fusion domain.
[0029] Item 2. The modified transposase as described in Item 1, wherein the polypeptide fusion domain is fused to the N-terminus of the transposase.
[0030] Item 3. The modified transposase as described in Item 1, wherein the polypeptide fusion domain is fused to the C-terminus of the transposase.
[0031] Item 4. The modified transposase as described in Item 1, wherein the polypeptide fusion domain comprises a purification tag.
[0032] Item 5. The modified transposase as described in Item 1, wherein the polypeptide fusion domain comprises a tag that increases solubility.
[0033] Item 6. The modified transposase as described in Item 1, wherein the polypeptide fusion domain comprises a domain selected from the following: maltose binding protein (MBP), elongation factor Ts (Tsf), 5-methylcytosine binding domain and protein A.
[0034] Item 7. A labeling reaction buffer, wherein the labeling reaction buffer is 2+ Compared to the same enzyme in a buffer containing the transposase and a concentration of Co sufficient to reduce GC bias, 2+ .
[0035] Item 8. The labeling reaction buffer as described in Item 7, wherein the Co 2+ At a concentration in the range of about 5 mM to about 40 mM.
[0036] Item 9. The labeling reaction buffer as described in Item 7, wherein the Co 2+ At a concentration in the range of about 5 mM to about 20 mM.
[0037] Item 10. The labeling reaction buffer as described in Item 7, wherein the Co 2+ At a concentration in the range of about 8 mM to about 12 mM.
[0038] Item 11. The labeling reaction buffer as described in Item 7, wherein the Co 2+ At a concentration of about 10 mM.
[0039] Item 12. A method for in vitro transposition, the method comprising: allowing the following components to interact: (i) a transposome complex comprising a transposase, (ii) a target DNA; and (iii) a reaction buffer, the reaction buffer being reacted with a precipitate containing Mg. 2+ Compared with the same enzyme in a buffer containing Co at a concentration effective to reduce GC bias, 2 + .
[0040] Item 13. A method as described in Item 12, wherein the Co 2+ At a concentration in the range of about 5 mM to about 40 mM.
[0041] Item 14. A method as described in Item 12, wherein the Co 2+ At a concentration in the range of about 5 mM to about 20 mM.
[0042] Item 15. A method as described in Item 12, wherein the Co 2+ At a concentration in the range of about 8 mM to about 12 mM.
[0043] Item 16. A method as described in Item 12, wherein the transposase comprises a mutant Tn5 transposase described herein.
[0044] Item 17. A method as described in Item 12, wherein the transposase includes Mos-1 transposase.
[0045] Item 18. A method as described in Item 12, wherein the transposase comprises a highly active Tn5 transposase.
[0046] Item 19. A mutant Tn5 transposase, wherein the mutant Tn5 transposase exhibits increased tolerance to DNA input compared to wild-type Tn5 transposase.
[0047] Item 20. The mutant Tn5 transposase as described in Item 19, wherein the mutant transposase comprises the amino acid sequence of SEQ IDNO:19.
[0048] Item 21. The mutant Tn5 transposase of Item 20, wherein the mutant Tn5 transposase exhibits increased DNA input tolerance at a ratio of nM mutant Tn5 transposase: ng input DNA of ≥2.4.
[0049] Item 22. A method for producing uniform fragment sizes across a range of target DNA input amounts, the method comprising:
[0050] (a) providing a target DNA; wherein the amount of the target DNA is selected from a range of target DNA amounts, and
[0051] (b) incubating the target DNA with a transposase, wherein the amino acid sequence of the transposase comprises at least 400 amino acids from the C-terminus of SEQ ID NO: 19, and wherein the transposase fragments the target DNA to produce uniform fragment sizes.
[0052] Item 23. A method as described in Item 22, wherein the target DNA is genomic DNA.
[0053] Item 24. A method as described in Item 23, wherein the genomic DNA is prokaryotic DNA.
[0054] Item 25. A method as described in Item 24, wherein the prokaryotic DNA is bacterial DNA.
[0055] Item 26. A method as described in Item 23, wherein the genomic DNA is eukaryotic genomic DNA.
[0056] Item 27. A method as described in Item 26, wherein the eukaryotic genomic DNA is human genomic DNA.
[0057] Item 28. A method as described in any one of Items 22-27, wherein the amount of the target DNA is in the range of about 1 ng to 200 ng.
[0058] Item 29. A method for sequencing a target DNA across a range of target DNA input amounts, the method comprising:
[0059] (a) incubating the target DNA with a transposome complex, the transposome complex comprising
[0060] (1) a mutant Tn5 transposase comprising the amino acid sequence of SEQ ID NO: 19; and
[0061] (2) a first polynucleotide, wherein the first polynucleotide comprises:
[0062] (i) a 3' portion, said 3' portion comprising a transposon end sequence, and
[0063] (ii) a first tag, wherein the first tag comprises a first sequencing tag domain,
[0064] wherein the amount of the target DNA is selected from a range of target DNA amounts, and the incubation is performed under conditions whereby the target DNA is fragmented and the 3' transposon end sequence of the first polynucleotide is transferred to the 5' end of the fragment,
[0065] thereby generating a double-stranded fragment in which the 5' end is tagged via the first tag and a single-stranded nick is present at the 3' end of the 5'-tagged strand;
[0066] (b) incubating the fragments with a nucleic acid modifying enzyme under conditions whereby a second tag is attached to the 3' end of the 5'-tagged strand,
[0067] (c) optionally, amplifying the fragments by providing a polymerase and an amplification primer corresponding to a portion of the first polynucleotide, thereby generating a representative library of di-tagged fragments having the first tag at the 5' end and the second tag at the 3' end;
[0068] (d) providing a first sequencing primer, wherein the first sequencing primer comprises a portion corresponding to the first sequencing tag domain; and
[0069] (e) extending the first sequencing primer and detecting in parallel the identity of nucleotides adjacent to the first sequencing tag domain of the representative library of di-tagged fragments.
[0070] Item 30. A method as described in Item 29, wherein the amount of the target DNA is in the range of about 1 ng to 200 ng.
[0071] Item 31. A method as described in Item 30, wherein the ratio of nM of mutant Tn5 transposase: ng of input DNA is ≥2.4.
[0072] Item 32. A method as described in Item 22, wherein the method is used for exome enrichment.
[0073] Item 33. The mutant Tn5 transposase as described in Item 19, wherein the mutant transposase has the amino acid sequence of SEQ ID NO:26.
[0074] Item 34. A mutant Tn5 transposase, comprising an amino acid sequence of at least 400 amino acids of any one of SEQ ID NOs: 2-10, 12-20 and 25-26.
[0075] The details of one or more embodiments are described in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1A Schematic diagram showing the structural alignment of the catalytic core domains of Tn5 transposase (1MUH), Hermes transposase (2BW3), HIV integrase (1ITG), Mu transposase (1BCM) and Mos1 transposase (3HOS). The numbers shown represent the numbers of the amino acid residues in Tn5 transposase.
[0078] Figure 1B Schematic diagram showing the structural alignment of the catalytic core domains of Tn5 transposase (1MUH, pink), Hermes transposase (2BW3, black), HIV integrase (1ITG, brown), Mu transposase (1BCM), and Mos1 transposase (3HOS, yellow). The position of W125 in Tn5 transposase is shown as a stick representation.
[0079] Figure 2 IVC graph showing the altered sequence insertion bias of the D248Y mutant Tn5 transposase compared to the Tn5 control.
[0080] Figure 3 IVC graph showing the altered sequence insertion bias of the D119L mutant Tn5 transposase compared to the Tn5 control.
[0081] Figure 4 IVC graph showing the altered sequence insertion bias of the W125M mutant Tn5 transposase compared to the Tn5 control.
[0082] Figure 5 IVC graph showing the altered sequence insertion bias of the ia248V insertion mutant Tn5 transposase compared to the Tn5 control.
[0083] Figure 6 IVC graph showing the altered sequence insertion bias of the K120Y, K120F and K120W Tn5 transposases compared to the Tn5 control.
[0084] Figure 7 IVC graph showing the altered sequence insertion bias of three mutant Tn5 transposases compared to a Tn5 control.
[0085] Fig. 8A Graph showing AT / GC dropout in a B. cereus library created by three mutant Tn5 transposases compared to a Tn5 control. Figure 8B Graph showing estimated library size of Bacillus cereus libraries created by three mutant Tn5 transposases compared to a Tn5 control.
[0086] Fig.9AFigure 2 is a graph showing uniformity of coverage in rapid capture enrichment experiments in libraries created from two mutant Tn5 transposases compared to a Tn5 control. Fig. 9B Figure 2 is a graph showing 10X and 20X target coverage and average target coverage in rapid capture enrichment experiments in libraries created from two mutant Tn5 transposases compared to a Tn5 control.
[0087] Fig. 10A Figure 2 is a graph showing the percentage of unique reads and hybrid selection library size passing filters in rapid capture enrichment experiments in libraries created from two mutant Tn5 transposases compared to a Tn5 control. Fig. 10B Figure 2 is a graph showing the penalties for achieving 10X, 20X, and 30X coverage in a rapid capture enrichment experiment in libraries created from two mutant Tn5 transposases compared to a Tn5 control.
[0088] Fig.11 Shown is a bar graph of the number of unique molecules in TS-Tn5059 and TS-Tn5 tagmented DNA libraries prepared using different tagmentation buffers.
[0089] Fig.12 Shown is a bar graph of the GC dropout percentage in TS-Tn5059 and TS-Tn5 tagged DNA libraries prepared using different tagmentation buffers.
[0090] Fig.13 Shown is a bar graph of the percentage of AT dropout in TS-Tn5059 and TS-Tn5 tagged DNA libraries prepared using different tagmentation buffers.
[0091] Fig.14 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5059 library prepared using standard buffer (TD) and cobalt buffer (Co) formulations.
[0092] Fig.15 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5059 library prepared using cobalt-DMSO (Co-DMSO), NF2, and HMW buffer formulations.
[0093] Fig.16 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5 library prepared using standard buffer formulation (TD) and cobalt buffer (Co).
[0094] Fig.17Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5 libraries prepared using cobalt-DMSO (Co-DMSO), NF2, and HMW buffer formulations.
[0095] Fig.18A , 18B , 18C and 18D show the bias map of sequence content in TS-Tn5 library, the bias map of sequence content in TS-TN5-Co library, the bias map of sequence content in TS-Tn5-Co-DMSO library and the bias map of sequence content in TS-Tn5-NF2 library, respectively.
[0096] Fig.19A , 19B , 19C and 19D show the bias map of the sequence content in the TS-Tn5059 library, the bias map of the sequence content in the TS-TN5059-Co library, the bias map of the sequence content in the TS-Tn5059-Co-DMSO library, and the bias map of the sequence content in the TS-Tn5059-NF2 library, respectively.
[0097] Fig. 20 Shown are bar graphs of the average total number of reads and average diversity in MBP-Mos1 tagged libraries prepared using different tagmentation buffers.
[0098] Fig.21 Bar graphs showing GC and AT shedding in the MBP-Mos1-tagged library.
[0099] Fig.22A , 22B , 22C and 22D show the bias map of sequence content in the Mos1-HEPES library, the bias map of sequence content in the Mos1-HEPES-DMSO library, the bias map of sequence content in the Mos1-HEPES-DMSO-Co library, and the bias map of sequence content in the Mos1-HEPES-DMSO-Mn library, respectively.
[0100] Fig.23 A flow chart showing an example of a method for preparing and enriching a genomic DNA library for exome sequencing.
[0101] Fig.24A A graph showing coverage in a tagged Bacillus cereus genomic DNA library prepared using the TS-Tn5059 transposome.
[0102] Fig. 24B A graph showing coverage in a tagged Bacillus cereus genomic DNA library prepared using the NexteraV2 transposome.
[0103] Fig.25A Diagram showing the gap positions and gap lengths in a tagged Bacillus cereus genomic DNA library prepared using the TS-Tn5059 transposome.
[0104] Fig.25B Diagram showing the gap positions and gap lengths in a tagged Bacillus cereus genome tagged DNA library prepared using NexteraV2 transposomes.
[0105] Fig.26 Shown is a graph of the bioanalyzer traces of fragment size distribution in a tagged genomic DNA library prepared using TDE1 (Tn5 version-1) and TS-Tn5 of TS-Tn5059 normalized to 40 nM (1X normalized concentration) against 25 ng human gDNA.
[0106] Fig. 27 Shown is the analysis of size distribution in tagged genomic DNA libraries prepared with 25 ng human gDNA using TDE1 (Tn5 version-1) and TS-Tn5 of TS-Tn5059 normalized to 40 nM (1X normalized concentration).
[0107] Fig.28 Graph showing Bioanalyzer traces of fragment size distribution in a tagged genomic DNA library prepared using a range of DNA inputs.
[0108] Fig.29A A graph showing a Bioanalyzer trace of fragment size distribution in a TS-Tn5059-tagged library prepared by a first user and using Coriel human DNA.
[0109] Fig.29B A graph showing a Bioanalyzer trace of fragment size distribution in a TS-Tn5059-tagged library prepared by a second user and using Coriel human DNA.
[0110] Fig.30 Graph showing bioanalyzer traces of fragment size distribution in a library tagged with Tn5 version 1 (TDE1) prepared using 25 ng - 100 ng gDNA at 6x "normalized" concentration of TDE1.
[0111] Fig.31 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5-tagged libraries prepared using 25 ng-100 ng gDNA at 6x "normalized" concentrations of TS-Tn5.
[0112] Fig.32Graph showing bioanalyzer traces of fragment size distribution in a TS-Tn5059-tagged library prepared using 10 ng-100 ng gDNA at 6x "normalization" concentration of TS-Tn5059.
[0113] Fig.33 Graph showing bioanalyzer traces of fragment size distribution in a TS-Tn5059-tagged library prepared using a wide range of gDNA (5 ng - 500 ng) at 6x "normalized" concentrations of TS-Tn5059. Detailed Description
[0115] In some sample preparation methods for DNA sequencing, each template contains adapters at either end of the insert and many steps are typically required to both modify the DNA or RNA and purify the desired modification reaction products. These steps are typically performed in solution before adding the appropriate fragments to a flow cell where they are coupled to a surface by a primer extension reaction that copies the hybridized fragments to the ends of primers covalently attached to the surface. These 'seeding' templates are then subjected to several cycles of amplification to produce monoclonal clusters of copied templates. However, as disclosed in US2010 / 0120098, the contents of which are incorporated herein in their entirety, the number of steps required to convert DNA in solution into adapter-modified templates in preparation for cluster formation and sequencing can be minimized by using transposase-mediated fragmentation and tagging (referred to herein as tagmentation). For example, tagmentation can be used to fragment DNA, for example, as used in Nextera TM As exemplified in the workflow of a DNA sample preparation kit (Illumina, Inc.), genomic DNA can be fragmented by an engineered transposome that simultaneously fragments and tags input DNA, thereby creating a population of fragmented nucleic acid molecules containing unique adapter sequences at the ends of the fragments. However, there is a need for transposases that exhibit improved insertion bias.
[0116] Thus, provided herein are transposases for improving the fragmentation and tagging of nucleic acid samples. The inventors have unexpectedly identified certain altered transposases that exhibit improved insertion sequence bias and have many other related advantages. One embodiment of the altered transposases provided herein is a transposase that exhibits improved insertion bias.
[0117] As used herein, the term "normalized transposome activity" refers to the minimum concentration of transposomes that produces the following bioanalyzer fragment size distribution in a 50 μl reaction: total area under the curve: 100-300 bp = 20%-30%; 301-600 bp = 30%-40%; 601-7,000 bp = 30-40%; 100-7,000 bp ≥ 90%.
[0118] This minimum concentration is referred to as 1X.
[0119] As used throughout the application, the concentration of transposomes and normalized activity are used interchangeably. In addition, as used throughout the application, the concentration of transposomes and the concentration of transposase are used interchangeably.
[0120] As used herein, the term "insertion bias" refers to the sequence preference of a transposase for an insertion site. For example, if the background frequency of A / T / C / G in a polynucleotide sample is evenly distributed (25% A, 25% T, 25% C, 25% G), any over-representation of a nucleotide relative to the other three nucleotides at a transposase binding site or cleavage site reflects the insertion bias at that site. Insertion bias can be measured using any of a variety of methods known in the art. For example, the insertion site can be sequenced, and the relative abundance of any particular nucleotide at each position in the insertion site can be compared, as generally stated in Example 1 below.
[0121] "Improvement of insertion bias" refers to that the frequency of a specific base at one or more positions of the binding site of the altered transposase is reduced or increased to be closer to the background frequency of the base in the polynucleotide sample. The improvement can be an increase in the frequency at the position relative to the frequency in the position of the unchanged transposase. Alternatively, the improvement can be a reduction in the frequency at the position relative to the frequency in the position of the unchanged transposase. Thus, for example, if the background frequency of T nucleotides in the polynucleotide sample is 0.25, and the altered transposase reduces the frequency of T nucleotides at the specific position in the transposase binding site from a frequency higher than 0.25 to a frequency closer to 0.25, the altered transposase has an improvement in insertion bias. Similarly, for example, if the background frequency of T nucleotides in the polynucleotide sample is 0.25, and the altered transposase increases the frequency of T nucleotides at the specific position in the transposase binding site from a frequency lower than 0.25 to a frequency closer to 0.25, the altered transposase has an improvement in insertion bias.
[0122] One way to measure insertion bias is by large-scale sequencing of the insertion site and measuring the frequency of bases at each position in the binding site relative to the insertion site, as described, for example, in Green et al. Mobile DNA (2012) 3:3, which is incorporated herein by reference in its entirety. A typical tool for displaying the abundance at each position is an intensity versus circulation distribution plot, for example, as in Figure 2 As described in Example 1 below, the fragment ends generated by transposon-mediated tagging and fragmentation can be sequenced at scale, and the base distribution frequency at each position of the insertion site can be measured to detect bias at one or more positions of the insertion site. Thus, for example, Figure 2 As shown in , the base distribution at position (1) with a frequency of 0.55 for the 'G' nucleotide and 0.16 for the 'A' nucleotide reflects a strong preference for G and a bias away from A at this position. As another example, and in contrast, Figure 3 As shown in , the base distribution of each of the four bases at position (20) is essentially 0.25 reflecting little or no sequence bias at this position.
[0123] In some embodiments provided herein, the altered transposase provides a reduction in insertion bias at one or more sites that are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more than 20 bases upstream or downstream of the insertion site. In some embodiments, the altered transposase provides a reduction in insertion bias at one or more sites that are 1 to 15 bases downstream of the insertion site. In some embodiments, the altered transposase provides a reduction in insertion bias at one or more sites that are 1 to 15 bases upstream of the insertion site.
[0124] As described in more detail below, the inventors have unexpectedly discovered that one or more mutations to residues at certain positions in the amino acid sequence of the transposase result in improved sequence insertion bias during the transposition event. These altered transposases provide improved performance in the tagmentation of high-diversity and low-diversity nucleic acid samples, resulting in greater uniformity of coverage and less dropout of multiple regions being sequenced.
[0125] As used herein, the term "DNA input tolerance" refers to the ability of a transposase to generate uniform DNA fragment sizes across a range of input DNA amounts.
[0126] As used herein, the symbols for elongation factor: TS and Tsf are used interchangeably.
[0127] In some embodiments, the input DNA is genomic DNA. In some embodiments, the range of input DNA can be 0.001 μg to 1 mg, 1 ng to 1 mg, 1 ng to 900 ng, 1 ng to 500 ng, 1 ng to 300 ng, 1 ng to 250 ng, 1 ng to 100 ng, 5 ng to 250 ng, or 5 ng to 100 ng, and the concentration of the transposase is between 5 nM and 500 nM. In some embodiments, the concentration of the transposase for the above-mentioned input DNA ranges is about 25 nM, 30 nM, 35 nM, 40 nM, 50 nM, 60 nM, 65 nM, 70 nM, 75 nM, 80 nM, 90 nM, 95 nM, 100 nM, 125 nM, 130 nM, 140 nM, 150 nM, 175 nM, 180 nM, 190 nM, In some embodiments, the normalized concentration of the transposase for the above-mentioned input DNA ranges or the concentration of the normalized transposase is selected from the range of about 0.1X to 10X, 1X to 10X, 3X to 8X, 4X to 7X. In some embodiments, the normalized concentration of transposase or the normalized transposase for the above-mentioned input DNA range is about 0.1X, 0.2X, 0.3X, 0.4X, 0.5X, 0.6X, 0.7X, 0.8X, 0.9X, IX, 1.1X, 1.2X, 1.3X, 1.4X, 1.5X, 1.6X, 1.7X, 1.8X, 1.9X, 2X, 2.1X, 2.2X, 2.3X, 2.4X, 2.5X, 2.6X, 2.7X, 2.8X, 2.9X, 3X, 3.1X, 3.2X, 3.3X, 3.4X, 3.6X, 3.7X, 3.8X, 3.9X, 4.1X, 4.2X, 4.3X, 4.4X, 4.5X, 4.6X, 4.7X, 4.8X, 4.9X, 5.1X, 5.2X, 5.3X, 5.4X, 5.6X, 5.7X, 5.8X, 5.9X, 3.4X, 3.5X, 3.6X, 3.7X, 3.8X, 3.9X, 4X, 4.1X, 4.2X, 4.3X, 4.4X, 4.5X, 4.6X, 4.7X, 4.8X, 4.9X, 5X, 5.1X, 5.2X, 5.3X, 5.4X, 5.5X, 5.6X, 5.7X, 5.8X, 5.9X, 6X, 6.1X, 6.2X, 6.3X, 6.4X, 6.5X, 6.6X, 6.7X, 6.8X, 6.9X, 7X or 7.5X, 8X, 8.5X, 9X, 9.5X, 10X.
[0128] In some embodiments, the amount of input DNA is 1 ng, 2 ng, 3 ng, 4 ng, 5 ng, 6 ng, 7 ng, 8 ng, 9 ng, 10 ng, 11 ng, 12 ng, 13 ng, 15 ng, 20 ng, 25 ng, 30 ng, 35 ng, 40 ng, 45 ng, 50 ng, 55 ng, 60 ng, 65 ng, 70 ng, 75 ng, 80 ng, 85 ng, 90 ng, 95 ng, 100 ng, 110 ng, 115 ng, 120 ng, 125 ng, 130 ng, 135 ng, 140 ng, 150 ng, 155 ng, 160 ng , 165ng, 170ng, 180ng, 185ng, 190ng, 195ng, 200ng, 210ng, 220ng, 225ng, 230ng, 235ng, 240ng, 245ng, 250ng, 260ng, 270ng, 280ng, 29 0ng, 300ng, 325ng, 350ng, 375ng, 400ng, 425ng, 450ng, 475ng, 500ng, 525ng, 550ng, 600ng, 650ng, 700ng, 750ng, 800ng, 850ng or 900ng. In some embodiments, the concentration of the transposase for the above-mentioned amount of input DNA is about 25nM, 30nM, 35nM, 40nM, 50nM, 60nM, 65nM, 70nM, 75nM, 80nM, 90nM, 95nM, 100nM, 125nM, 130nM, 140nM, 150nM, 175nM, 180nM, 19 ...95nM, 1 In some embodiments, the normalized concentration of the transposase for the above-mentioned amount of input DNA or the concentration of the normalized transposase is selected from the following ranges: about 0.1X to 10X, 1X to 10X, 3X to 8X, 4X to 7X.In some embodiments, the normalized concentration of transposase for the above-mentioned amount of input DNA or the normalized transposase is about 0.1X, 0.2X, 0.3X, 0.4X, 0.5X, 0.6X, 0.7X, 0.8X, 0.9X, 1X, 1.1X, 1.2X, 1.3X, 1.4X, 1.5X, 1.6X, 1.7X, 1.8X, 1.9X, 2X, 2.1X, 2.2X, 2.3X, 2.4X, 2.5X, 2.6X, 2.7X, 2.8X, 2.9X, 3X, 3.1X, 3.2X, 3.3X, 3.4X, 3.6X, 3.7X, 3.8X, 3.9X, 4.1X, 4.2X, 4.3X, 4.4X, 4.5X, 4.6X, 4.7X, 4.8X, 4.9X, 5.1X, 5.2X, 5.3X, 5.4X, 5.6X, 5.7X, 5.8X, 5.9X, .4X, 3.5X, 3.6X, 3.7X, 3.8X, 3.9X, 4X, 4.1X, 4.2X, 4.3X, 4.4X, 4.5X, 4.6X, 4.7X, 4.8X, 4.9X, 5X, 5.1X, 5.2X, 5.3X, 5.4X, 5.5X, 5.6X, 5.7X, 5.8X, 5.9X, 6X, 6.1X, 6.2X, 6.3X, 6.4X, 6.5X, 6.6X, 6.7X, 6.8X, 6.9X, 7X or 7.5X, 8X, 8.5X, 9X, 9.5X, 10X.
[0129] In some embodiments, the ratio of nM concentration of the transposase to the ng amount of input DNA is about 0.5 to 5, 1 to 5, 2 to 5, 2.1 to 3, or 2.1 to 2.5.
[0130] As used herein, the term "genomic DNA" refers to nucleic acid present in a cell that contains one or more genes encoding a variety of cellular proteins. In some embodiments, the genomic DNA is from a prokaryotic organism, e.g., bacteria and archaea. In some embodiments, the genomic DNA is from a eukaryotic organism, e.g., humans, plants, fungi, amoeba.
[0131] As used herein, the term "mutant" or "modified" refers to a gene or gene product that exhibits a change in sequence and / or functional properties (i.e., a characteristic of an alteration) when compared to a wild-type gene or gene product. "Mutant" or "modified" also refers to a sequence at one or more specific nucleotide positions or a sequence at one or more specific codon positions or a sequence at one or more specific amino acid positions that exhibits a change in sequence and / or functional properties (i.e., a characteristic of an alteration) when compared to a wild-type gene or gene product.
[0132] "Including" as used herein has the same meaning as the term comprising.
[0133] As used herein, "about" means plus or minus 10% of a quantitative aspect.
[0134] As described in more detail below, the inventors have surprisingly found that one or more mutations to residues at certain positions in the amino acid sequence of the transposase result in increased DNA input tolerance, such that the mutant transposase produces uniform DNA fragment sizes across a range of input DNA amounts compared to the wild-type transposase. In one embodiment, the TS-Tn5059 transposase exhibits increased DNA input tolerance compared to other transposases, e.g., TS-Tn5 and Tn5 version 1 (TDE1).
[0135] In some embodiments, TS-Tn5059 exhibits increased DNA input tolerance compared to other transposases, wherein the range of input DNA is between 1 ng and 200 ng genomic DNA, and the concentration of TS-Tn5059 is between 100-300 nM. In some embodiments in which TS-Tn5059 exhibits increased DNA input tolerance compared to other transposases, the range of input DNA is between 5 ng and 200 ng genomic DNA, and the concentration of TS-Tn5059 is between 100-250 nM. In some embodiments in which TS-Tn5059 exhibits increased DNA input tolerance compared to other transposases, the range of input DNA is between 5 ng and 100 ng genomic DNA, and the concentration of TS-Tn5059 is between 240 nM and 250 nM.
[0136] Improved insertion bias together with increased DNA input tolerance provides better performance than current Rapid Capture Protocol (Illumina, Inc.) A faster and more flexible sample preparation and exome enrichment protocol.
[0137] As used herein, the term "tagmentation" refers to the modification of DNA by a transposome complex comprising a transposase complexed with an adaptor comprising a transposon end sequence. Tagging results in simultaneous fragmentation of DNA and ligation of adaptors to the 5' ends of both strands of the duplex fragments.
[0138] As used herein, a "transpososome complex" or "transposome" comprises at least one transposase and a transposase recognition site. In some such systems, the transposase can form a functional complex with a transposon recognition site that can catalyze a transposition reaction. The transposase can bind to the transposase recognition site and insert the transposase recognition site into the target nucleic acid in a process referred to herein as tagmentation. In some such insertion events, one strand of the transposase recognition site can be transferred into the target nucleic acid.
[0139] The altered transposase provided herein can form part of a transpososome complex. Exemplary transposition complexes include, but are not limited to, highly active Tn5 transposases and Tn5-type transposase recognition sites. The highly active Tn5 transposase may be included in U.S. Patent No. 5,925,545, U.S. Patent No. 5,965,443, U.S. Patent No. 7,083,980 and U.S. Patent No. 7,608,434 and Goryshin and Reznikoff, J. Biol Chem., 273: Those described in 7367 (1998), each of which is incorporated herein by reference in its entirety. However, it should be understood that the altered transposase provided herein can be used in any transposition system that can be used in the methods provided herein that can insert the transposon end in a random or nearly random manner with an efficiency sufficient to label the target nucleic acid for its intended purpose.
[0140] For example, the altered transposase provided may include at least one amino acid substitution mutation at one or more positions that are functionally equivalent to a site in the Tn5 amino acid sequence. The region homologous to Tn5 is shown herein (as illustrated in FIG. 1 ) and allows identification of functionally equivalent sites in other transposases, such as Hermes transposase, HIV integrase, Mu transposase, and Mos1 transposase. Similarly, functionally equivalent sites in other transposases or integrases will be apparent to one of ordinary skill in the art by, for example, performing a sequence alignment of the Tn5 amino acid sequence and identifying conservative or semi-conservative residues or domains. Therefore, it should be understood that the transposition system that can be used for certain embodiments provided herein includes any known transposase having sites that are functionally equivalent to those of Tn5. For example, such a system can include a MuA transposase and a Mu transposase recognition site comprising R1 and R2 terminal sequences (Mizuuchi, K., Cell, 35:785, 1983; Savilahti, H, et al., EMBO J., 14:4893, 1995).
[0141] Further examples of transposition systems included in certain embodiments provided herein include Staphylococcus aureus Tn552 (Colegio et al., J. Bacteriol., 183:2384-8, 2001; Kirby C et al., Mol. Microbiol., 43:173-86, 2002), Ty1 (Devine & Boeke, Nucleic Acids Res., 22:3765-72, 1994 and International Publication WO 95 / 23875), transposon Tn7 (Craig, NL, Science. 271:1512, 1996; Craig, NL, in: Review in Curr Top Microbiol Immunol., 204:27-48, 1996), Tn / O and IS10 (Kleckner N, et al., Curr Top Microbiol Immunol., 204:27-48, 1996), Immunol., 204: 49-82, 1996), Mariner transposase (Lampe DJ, et al., EMBO J., 15: 5470-9, 1996), Tc1 (Plasterk RH, Curr. Topics Microbiol. Immunol., 204: 125-43, 1996), P element (Gloor, GB, Methods Mol. Biol., 260: 97-114, 2004), Tn3 (Ichikawa & Ohtsubo, J Biol. Chem. 265: 18829-32, 1990), bacterial insertion sequence (Ohtsubo & Sekine, Curr. Top. Microbiol. Immunol. 204: 1-26, 1996), retrovirus (Brown, et al., Proc Natl Acad Sci USA, 86:2525-9, 1989), and yeast retrotransposons (Boeke & Corces, Annu Rev Microbiol. 43:403-34, 1989). More examples include IS5, Tn10, Tn903, IS911, and engineered versions of transposase family enzymes (Zhang et al., (2009) PLoS Genet. 5:e1000689. Epub 2009 Oct. 16; Wilson C. et al. (2007) J. Microbiol. Methods 71:332-5). The above cited references are incorporated herein by reference in their entirety.
[0142] In short, a "transposition reaction" is a reaction in which one or more transposons are inserted into a target nucleic acid at a random site or a nearly random site. The essential components in a transposition reaction are a transposase and a DNA oligonucleotide presenting the nucleotide sequence of the transposon, including the transferred transposon sequence and its complement (i.e., the non-transferred transposon end sequence), and other components required to form a functional transposition complex or transposome complex. The DNA oligonucleotide may also contain additional sequences (e.g., adapter or primer sequences) as needed or desired.
[0143] The connector added to the 5' and / or 3' end of the nucleic acid may include a universal sequence. A universal sequence is a region of nucleotide sequence that is common to two or more nucleic acid molecules, i.e., shared. Optionally, two or more nucleic acid molecules also have a region of sequence differences. Thus, for example, a 5' connector may include the same or universal nucleotide sequence, and a 3' connector may include the same or universal sequence. The universal sequence that may be present in different members of a plurality of nucleic acid molecules may allow the use of a single universal primer that is complementary to the universal sequence to replicate or amplify a plurality of different sequences.
[0144] Transposase mutant
[0145] Therefore, mutant transposases modified relative to wild-type transposases are provided herein. The altered transposases may include at least one amino acid substitution mutation at one or more positions functionally equivalent to those residues listed in Table 1 below. Table 1 lists substitution mutations at transposase residues that have been shown to result in improved insertion bias. As listed in Table 1, substitution mutations provided herein may be in any functional transposase backbone, such as the wild-type Tn5 transposase exemplified herein as SEQ ID NO: 1, or transposases with additional mutations to other sites, the additional mutations included in those mutations found in the transposase sequence known as the hyperactive Tn5 transposase, such as, for example, one or more mutations shown in the incorporated materials of U.S. Patent No. 5,925,545, U.S. Patent No. 5,965,443, U.S. Patent No. 7,083,980, and U.S. Patent No. 7,608,434, and as exemplified herein as SEQ ID NO: 11.
[0146] Table 1 - Examples of mutations leading to improved insertion bias
[0147]
[0148]
[0149]
[0150] As understood in the art, the reference numbers listed in the above table refer to the amino acid positions of the wild-type Tn5 sequence (SEQ ID NO: 1). One of ordinary skill in the art will appreciate that numbering may change due to N-terminal truncations, insertions or fusions. Although the position numbers may have changed, the functional positions of the amino acids listed above will remain the same. For example, the first 285 amino acid residues of the sequence listed in SEQ ID NO: 25 comprise an N-terminal fusion of E. coli TS, followed by amino acid residues 2-476 of SEQ ID NO: 11. Thus, for example, Pro 656 of SEQ ID NO: 25 corresponds functionally to Pro 372 of SEQ ID NO: 11.
[0151] Thus, in certain embodiments, relative to the wild-type transposase, the altered transposase provided herein comprises at least one amino acid substitution mutation at one or more positions that are functionally equivalent to, for example, Asp248, Asp119, Trp125, Lys120, Lys212, Pro214, Gly251, Ala338, Glu146 and / or Glu190 in the amino acid sequence of Tn5 transposase.
[0152] In some embodiments, the mutant transposase may comprise a mutation at position Asp248. The mutation at position Asp248 may be, for example, a substitution mutation or an insertion mutation. In certain embodiments, the mutation is a substitution mutation to any residue other than Asp. In certain embodiments, the substitution mutation at position Asp248 comprises a mutation to a residue selected from the group consisting of Tyr, Thr, Lys, Ser, Leu, Ala, Trp, Pro, Gln, Arg, Phe and His.
[0153] In certain embodiments, the mutation at position Asp248 is an insertion mutation after position Asp248. In certain aspects, the insertion mutation may include inserting any residue after Asp248. In certain aspects, the insertion mutation may include inserting a hydrophobic residue after position Asp248. Hydrophobic residues are known to those skilled in the art and include, for example, Val, Leu, Ile, Phe, Trp, Met, Ala, Tyr and Cys. In certain aspects, the insertion mutation may include inserting a valine residue after position Asp248.
[0154] Some embodiments provided herein include mutant Tn5 transposases modified relative to wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Asp119. In some aspects, the mutation at position Asp119 is a substitution mutation. In some aspects, the substitution mutation at position Asp119 may include a mutation to a hydrophobic residue. Hydrophobic residues are known to those skilled in the art and include, for example, Val, Leu, Ile, Phe, Trp, Met, Ala, Tyr and Cys. In some aspects, the substitution mutation at position Asp119 may include a mutation to a hydrophilic residue. Hydrophilic residues are known to those skilled in the art and include, for example, Arg, Lys, Asn, His, Pro, Asp and Glu. In some aspects, the substitution mutation at position Asp119 may include a mutation to a residue selected from the group consisting of: Leu, Met, Ser, Ala and Val.
[0155] Some embodiments provided herein include a mutant Tn5 transposase modified relative to a wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Trp125. In certain aspects, the mutation at position Trp125 is a substitution mutation. In certain aspects, the substitution mutation at position Trp125 may include a mutation to a methionine residue.
[0156] Some embodiments provided herein include mutant Tn5 transposases modified relative to wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Lys120. In some aspects, the mutation at position Lys120 is a substitution mutation. In some aspects, the substitution mutation at position Lys120 may include a mutation to a bulky aromatic residue. Residues characterized as bulky aromatic residues are known to those skilled in the art and include, for example, Phe, Tyr, and Trp. In some aspects, the substitution mutation at position Lys120 may include a mutation to a residue selected from the group consisting of: Tyr, Phe, Trp, and Glu.
[0157] Some embodiments provided herein include mutant Tn5 transposases modified relative to wild-type Tn5 transposase, the mutant transposase comprising mutations at positions Lys212 and / or Pro214 and / or Ala338. In certain aspects, one or more mutations at positions Lys212 and / or Pro214 and / or Ala338 are substitution mutations. In certain aspects, the substitution mutation at position Lys212 comprises a mutation to arginine. In certain aspects, the substitution mutation at position Pro214 comprises a mutation to arginine. In certain aspects, the substitution mutation at position Ala338 comprises a mutation to valine. In some embodiments, the transposase may also comprise a substitution mutation at Gly251. In certain aspects, the substitution mutation at position Gly251 comprises a mutation to arginine.
[0158] Some embodiments provided herein include mutant Tn5 transposases modified relative to wild-type Tn5 transposase, the mutant transposase comprising a mutation at position Glu146 and / or Glu190 and / or Gly251. In some aspects, one or more mutations at position Glu146 and / or Glu190 and / or Gly251 are substitution mutations. In some aspects, the substitution mutation at position Glu146 may include a mutation to glutamine. In some aspects, the substitution mutation at position Glu190 may include a mutation to glycine. In some aspects, the substitution mutation at position Gly251 may include a mutation to arginine.
[0159] In any of the above described embodiments, the mutant Tn5 transposase may further comprise a substitution mutation at a position functionally equivalent to Glu54 and / or Met56 and / or Leu372 in the Tn5 transposase amino acid sequence. In certain embodiments, the transposase comprises a substitution mutation homologous to Glu54Lys and / or Met56Ala and / or Leu372Pro in the Tn5 transposase amino acid sequence.
[0160] Some embodiments provided herein include a mutant Tn5 transposase comprising the amino acid sequence of any one of SEQ ID NOs: 2-10 and 12-20.
[0161] Also provided herein are altered transposases comprising substitution mutations to a semiconservative domain. As used herein, the term "semiconservative domain" refers to a portion of a transposase that is fully conserved or at least partially conserved between multiple transposases and / or between multiple species. A semiconservative domain comprises amino acid residues located in the catalytic core domain of a transposase. It has been unexpectedly found that mutations in one or more residues in a semiconservative domain affect transposase activity, resulting in improvements in insertion bias.
[0162] In some embodiments, the semiconservative domain comprises amino acids having a sequence listed in SEQ ID NO: 21. SEQ ID NO: 21 corresponds to residues 124-133 of the Tn5 transposase amino acid sequence listed herein as SEQ ID NO: 1. A structural alignment showing conservation among various transposases in a semiconservative domain is shown in Figure 1. The transposase sequences shown in Figure 1 include the catalytic core domains of Tn5 transposase (1MUH), Hermes transposase (2BW3), HIV integrase (1ITG), Mu transposase (1BCM), and Mos1 transposase (3HOS).
[0163] It has been unexpectedly found that mutations to one or more residues in the semi-conserved domains result in improved insertion bias. For example, in some embodiments of the altered transposases provided herein, the substitution mutation comprises a mutation to any residue other than Trp at position 2 of SEQ ID NO: 21. In certain embodiments, the altered transposase comprises a mutation to Met at position 2 of SEQ ID NO: 21.
[0164] "Functionally equivalent" means that in the case of studies using completely different transposases, the control transposase will contain amino acid substitutions at amino acid positions that are believed to have the same functional role enzymatically in the other transposases. As an example, a mutation from lysine to methionine at position 288 in Mu transposase (K288M) will be functionally equivalent to a substitution from tryptophan to methionine at position 125 in Tn5 transposase (W125M).
[0165] Typically, functionally equivalent substitution mutations in two or more different transposases occur at homologous amino acid positions in the amino acid sequence of the transposase. Therefore, the term "functionally equivalent" as used herein also includes mutations that are "positionally equivalent" or "homologous" to a given mutation, regardless of whether the specific function of the mutated amino acid is known. It is possible to identify positionally equivalent or homologous amino acid residues in the amino acid sequences of two or more different transposases based on sequence alignment and / or molecular modeling. The example of sequence alignment and molecular modeling for identifying positionally equivalent and / or functionally equivalent residues is shown in Figure 1. Therefore, for example, as shown in Figure 1, the residues in the semi-conservative domain are identified as positions 124-133 of the Tn5 transposase amino acid sequence. The corresponding residues in the Hermes transposase, HIV integrase, Mu transposase and Mos1 transposase transposases are marked by vertical alignment in the figure, and are considered to be positionally equivalent to and functionally equivalent to the corresponding residues in the Tn5 transposase amino acid sequence.
[0166] The altered transposase described above may comprise additional substitution mutations known to enhance one or more aspects of transposase activity. For example, in some embodiments, in addition to any of the above mutations, the altered Tn5 transposase may further comprise substitution mutations at positions functionally equivalent to Glu54 and / or Met56 and / or Leu372 in the Tn5 transposase amino acid sequence. Any of a variety of substitution mutations leading to improved activity at one or more positions functionally equivalent to the positions of Glu54 and / or Met56 and / or Leu372 in the Tn5 transposase amino acid sequence may be made, as known in the art and exemplified by the disclosures of U.S. Pat. Nos. 5,925,545, 5,965,443, 7,083,980, and 7,608,434, and in Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998), the contents of each of which are incorporated by reference in their entirety. In some embodiments, the transposase comprises substitution mutations homologous to Glu54Lys and / or Met56Ala and / or Leu372Pro in the Tn5 transposase amino acid sequence. For example, the substitution mutation may include a substitution mutation homologous to Glu54Lys and / or Met56Ala and / or Leu372Pro in the amino acid sequence of Tn5 transposase.
[0167] Mutant transposase
[0168] Various types of mutagenesis are optionally used in the present disclosure to, for example, modify transposases to produce variants, according to, for example, transposase patterns and pattern predictions as discussed above, or using random or semi-random mutation methods. Generally, any available mutagenesis program can be used to prepare transposon mutants. Such mutagenesis programs optionally include selection of mutant nucleic acids and polypeptides for one or more activities of interest (e.g., improved insertion bias). Available programs include, but are not limited to: site-directed mutagenesis, random point mutagenesis, in vitro or in vivo homologous recombination (DNA reorganization and combined overlapping PCR), mutagenesis using uracil templates, oligonucleotide-directed mutagenesis, thiophosphate-modified DNA mutagenesis, mutagenesis using gapped duplex DNA, point mismatch repair, mutagenesis using repair-deficient host strains, restricted selection and restricted purification, deletion mutagenesis, mutagenesis by full gene synthesis, degenerate PCR, double-strand break repair, and many other programs known to those skilled in the art. The starting transposase for mutagenesis can be any transposase noted herein, including useful transposase mutants such as, for example, those identified in U.S. Pat. No. 5,925,545, U.S. Pat. No. 5,965,443, U.S. Pat. No. 7,083,980, and U.S. Pat. No. 7,608,434, and the disclosure of Goryshin and Reznikoff, J. Biol Chem., 273:7367 (1998), the contents of each of which are incorporated by reference in their entirety.
[0169] Optionally, mutagenesis can be guided by known information from naturally occurring transposase molecules or known altered or mutant transposases (e.g., using existing mutant transposases as mentioned in the aforementioned references), e.g., sequence, sequence comparisons, physical properties, crystal structures, and / or the like. However, in another class of embodiments, the modifications can be essentially random (e.g., as in classical or "family" DNA shuffling, see, e.g., Crameri et al. (1998) "DNA shuffling of a family of genes from diverse species accelerates directed evolution" Nature 391: 288-291).
[0170] Additional information on mutant forms can be found in: Sambrook et al., Molecular Cloning--A Laboratory Manual (3rd Edition), Volumes 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY, 2000 ("Sambrook"); Current Protocols in Molecular Biology, FM Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (supplemented through 2011) ("Ausubel")) and PCR Protocols A Guide to Methods and Applications (Innis et al., eds.) Academic Press Inc. San Diego, Calif. (1990) ("Innis"). Additional details on mutant forms are provided in the following publications and references cited: Arnold, Protein engineering for unusual environments, Current Opinion in Biotechnology 4:450-455 (1993); Bass et al., Mutant Trp repressors with new DNA-binding specificities, Science 242:240-245 (1988); Bordo and Argos (1991) Suggestions for "Safe" Residue Substitutions in Site-directed Mutagenesis 217:721-729; Botstein & Shortle, Strategies and applications of in vitro mutagenesis, Science 229:1193-1201 (1985); Carter et al., Improved oligonucleotide site-directed mutagenesis using M13 vectors, Nucl. Acids Res.13:4431 - 4443(1985); Carter, Site - directed mutagenesis, Biochem. J. 237:1 - 7(1986); Carter, Improved oligonucleotide - directed mutagenesis using M13 vectors, Methods in Enzymol. 154:382 - 403(1987); Dale et al., Oligonucleotide - directed random mutagenesis using the phosphorothioate method, Methods Mol. Biol. 57:369 - 374(1996); Eghtedarzadeh & Henikoff, Use of oligonucleotides to generate large deletions, Nucl. Acids Res. 14:5115(1986); Fritz et al.; Oligonucleotide - directed construction of mutations: a gapped duplex DNA procedure without enzymatic reactions in vitro, Nucl. Acids Res. 16:6987 - 6999(1988); Grundstrom et al., Oligonucleotide - directed mutagenesis by microscale `shot - gun` gene synthesis, Nucl. Acids Res. 13:3305 - 3316(1985); Hayes(2002)Combining Computational and Experimental Screening for rapid Optimization of Protein Properties PNAS 99(25)15926 - 15931; Kunkel, The efficiency of oligonucleotide directed mutagenesis, in Nucleic Acids & Molecular Biology(Eckstein, F. and Lilley, D.M.J.Editor, Springer Verlag, Berlin)) (1987); Kunkel, Rapid and efficient site-specific mutagenesis without phenotypic selection, Proc. Natl. Acad. Sci. USA 82:488-492 (1985); Kunkel et al., Rapid and efficient site-specific mutagenesis without phenotypic selection, Methods in Enzymol. 154, 367-382 (1987); Kramer et al., The gapped duplex DNA approach to oligonucleotide-directed mutation construction, Nucl. Acids Res. 12:9441-9456 (1984); Kramer & Fritz Oligonucleotide-directed construction of mutations via gapped duplex DNA, Methods in Enzymol. 154:350-367 (1987); Kramer et al., Point Mismatch Repair, Cell 38:879-887 (1984); Kramer et al., Improved enzymatic in vitro reactions in the gapped duplex DNA approach to oligonucleotide-directed construction of mutations, Nucl. Acids Res. 16:7207 (1988); Ling et al., Approaches to DNA mutagenesis: an overview, Anal Biochem. 254(2):157-178 (1997); Lorimer and Pastan Nucleic Acids Res. 23, 3067-8 (1995); Mandecki, Oligonucleotide-directed double-strand break repair in plasmids of Escherichia coli: a method for site-specific mutagenesis, Proc. Natl.Acad. Sci. USA, 83: 7177 - 7181 (1986); Nakamaye & Eckstein, Inhibition of restriction endonuclease Nci I cleavage by phosphorothioate groups and its application to oligonucleotide-directed mutagenesis, Nucl. Acids Res. 14: 9679 - 9698 (1986); Nambiar et al., Total synthesis and cloning of a gene coding for the ribonuclease S protein, Science 223: 1299 - 1301 (1984); Sakamar and Khorana, Total synthesis and expression of a gene for the a-subunit of bovine rod outer segment guanine nucleotide-binding protein (transducin), Nucl. Acids Res. 14: 6361 - 6372 (1988); Sayers et al., Y-T Exonucleases in phosphorothioate-based oligonucleotide-directed mutagenesis, Nucl. Acids Res. 16: 791 - 802 (1988); Sayers et al., Strand specific cleavage of phosphorothioate-containing DNA by reaction with restriction endonucleases in the presence of ethidium bromide, (1988) Nucl. Acids Res. 16: 803 - 814; Sieber, et al., Nature Biotechnology, 19: 456 - 460 (2001); Smith, In vitro mutagenesis, Ann. Rev. Genet. 19: 423 - 462 (1985); Methods in Enzymol. 100: 468 - 500 (1983); Methods in Enzymol.154:329-350(1987); Stemmer, Nature 370, 389-91(1994); Taylor et al., The use of phosphorothioate-modified DNA in restriction enzyme reactions to prepare nicked DNA, Nucl. Acids Res. 13:8749-8764(1985); Taylor et al., The rapid generation of oligonucleotide-directed mutations at high frequency using phosphorothioate-modified DNA, Nucl. Acids Res. 13:8765-8787(1985); Wells et al., Importance of hydrogen-bond formation in stabilizing the transition state of subtilisin, Phil. Trans. R. Soc. Lond. A 317:415-423(1986); Wells et al., Cassette mutagenesis: an efficient method for generation of multiple mutations at defined sites, Gene 34:315-323(1985); Zoller & Smith, Oligonucleotide-directed mutagenesis using M13-derived vectors: an efficient and general procedure for the production of point mutations in any DNA fragment, Nucleic Acids Res. 10:6487-6500(1982); Zoller & Smith, Oligonucleotide-directed mutagenesis of DNA fragments cloned into M13 vectors, Methods in Enzymol.100:468-500 (1983); Zoller & Smith, Oligonucleotide-directed mutagenesis: a simple method using two oligonucleotide primers and a single-stranded DNA template, Methods in Enzymol. 154:329-350 (1987); Clackson et al. (1991) "Making antibody fragments using phage display libraries" Nature 352:624-628; Gibbs et al. (2001) "Degenerate oligonucleotide gene shuffling (DOGS): a method for enhancing the frequency of recombination with family shuffling" Gene 271:13-20; and Hiraga and Arnold (2003) "General method for sequence-independent site-directed chimeragenesis: J. Mol. Biol. 330:287-296. Additional details on many of the above methods can be found in Methods in Found in Enzymology Volume 154, which also describes useful controls for solving problems with various mutagenesis methods.
[0171] Preparation and isolation of recombinant transposase
[0172] Typically, nucleic acids encoding transposases as provided herein can be prepared by cloning, recombination, in vitro synthesis, in vitro amplification and / or other available methods. A variety of recombinant methods can be used to express expression vectors encoding transposases as provided herein. Methods for preparing recombinant nucleic acids, expressing and separating expression products are well known and described in the art. Many exemplary mutations and combinations of mutations are described herein, as well as strategies for designing desired mutations. Methods for preparing and selecting mutations in the catalytic domain of transposases are found herein and are exemplified in U.S. Patent No. 5,925,545, U.S. Patent No. 5,965,443, U.S. Patent No. 7,083,980 and U.S. Patent No. 7,608,434, which are incorporated by reference in their entirety.
[0173] Additional useful references for mutagenesis, recombination, and in vitro nucleic acid manipulation methods (including cloning, expression, PCR, etc.) include Berger and Kimmel, Guide to Molecular Cloning Techniques, Methods in Enzymology Vol. 152 Academic Press, Inc., San Diego, Calif. (Berger); Kaufman et al. (2003) Handbook of Molecular and Cellular Methods in Biology and Medicine 2nd Edition Ceske (ed.) CRC Press (Kaufman); and The Nucleic Acid Protocols Handbook Ralph Rapley (ed.) (2000) Cold Spring Harbor, Humana Press Inc (Rapley); Chen et al. (eds.) PCR Cloning Protocols, 2nd Edition (Methods in Molecular Biology, Vol. 192) Humana Press; and in Viljoen et al. (2005) Molecular Diagnostic PCR Handbook Springer, ISBN 6002668. 1402034032 in.
[0174] In addition, a large number of kits are commercially available for purifying plasmids or other relevant nucleic acids from cells (see, e.g., EasyPrep and EasyPrep, both from Pharmacia Biotech). TM 、FlexiPrep TM ; StrataClean from Stratagene TM ; and QIAprep from Qiagen TM ). Any isolated and / or purified nucleic acid can be further manipulated to produce other nucleic acids, used to transfect cells, incorporated into related vectors to infect organisms for expression, etc. Typical cloning vectors contain transcription and translation terminators, transcription and translation initiation sequences, and promoters useful for regulating the expression of specific target nucleic acids. The vector optionally includes a universal expression cassette comprising at least one independent terminator sequence, a sequence (e.g., a shuttle vector) that allows the cassette to replicate in eukaryotes or prokaryotes or both, and a selection marker for both prokaryotic and eukaryotic systems. The vector is suitable for replication and integration in prokaryotes, eukaryotes, or both.
[0175] In some embodiments, the transposase provided herein is expressed as a fusion protein. Fusion proteins can enhance the following features: for example, the solubility, expression and / or purification of the transposase. As used herein, the term "fusion protein" refers to a single polypeptide chain having at least two polypeptide domains that are not usually present in a single natural polypeptide. Therefore, naturally occurring proteins and point mutants thereof are not "fusion proteins" as used herein. Preferably, the polypeptide of interest is fused to at least one polypeptide domain via a peptide bond, and the fusion protein may also include a connection region of amino acids between the amino acid portions derived from the separate proteins. The polypeptide domain fused to the polypeptide of interest can enhance the solubility and / or expression of the polypeptide of interest, and purification tags may also be provided to allow purification of the recombinant fusion protein from the host cell or culture supernatant or both. Polypeptide domains that increase solubility during expression, purification and / or storage are well known in the art and include, for example, maltose binding protein (MBP) and elongation factor Ts (Tsf), as exemplified by Fox, JD and Waugh DS, E. coli Gene Expression Protocols Methods in Molecular Biology, (2003) 205:99-117 and Han et al. FEMS Microbiol. Lett. (2007) 274:132-138, each of which is incorporated herein by reference in its entirety. The polypeptide domain fused to the polypeptide of interest can be fused at the N-terminus or C-terminus of the polypeptide of interest. The term "recombinant" refers to the artificial combination of two otherwise separate sequence region segments by, for example, chemical synthesis or by manipulating the separated segments of amino acids or nucleic acids by genetic engineering techniques.
[0176] In one embodiment, the present invention provides a transposase fusion protein comprising a modified Tn5 transposase and an elongation factor Ts (Tsf). The Tsf-Tn5 fusion protein can be assembled into a functional dimer transposome complex comprising a fused transposase and free transposon ends. Compared to unfused Tn5 proteins, the Tsf-Tn5 fusion protein has increased solubility and thermal stability.
[0177] In another embodiment, the invention provides a transposase fusion protein comprising a modified Tn5 transposase and a protein domain that recognizes 5-methylcytosine. The 5-methylcytosine-Tn5 fusion protein can be assembled into a functional dimer transposome complex comprising a fused transposase and free transposon ends. The 5-methylcytosine binding protein domain can, for example, be used to target the Tn5 transposome complex to methylated regions of the genome.
[0178] In yet another embodiment, the invention provides a transposase fusion protein comprising a modified Tn5 transposase and a protein A antibody binding domain. The protein A-Tn5 fusion protein can be assembled into a functional dimeric transposome complex comprising a fused transposase and free transposon ends. The antibody binding domain of protein A can, for example, be used to target the Tn5 transposome complex to a region of the genome that binds the antibody.
[0179] The invention provides a transposase fusion protein comprising all or part of a modified Tn5 transposase and an elongation factor Ts (Tsf). Tsf is a protein tag that can be used to enhance the solubility of a heterologous protein expressed in a bacterial expression system, for example, an E. coli expression system. The ability of Tsf to increase the solubility of a heterologous protein may be due to the inherent high folding efficiency of the Tsf protein. In a protein purification experiment (data not shown), Tsf was purified as a complex with protein Tu. In order for Tu to combine in the complex, Tsf needs to be correctly folded. The purification of the Tsf-Tu complex shows that Tsf is correctly folded. The exemplary nucleic acid and corresponding amino acid sequence of the E. coli elongation factor TS are shown as SEQ ID NOs: 22 and 23, respectively.
[0180] In one example, a TS-Tn5 fusion protein is constructed by fusing TS to the N-terminus of a highly active Tn5 transposase. Exemplary amino acid sequences of TS-fusions with mutant Tn5 transposase proteins are shown as SEQ ID NOs: 25 and 26, respectively. SEQ ID NO: 24 corresponds to a nucleic acid sequence encoding the TS-mutant protein fusion of SEQ ID NO: 25.
[0181] Although all or part of the TS can be fused at the N-terminus or C-terminus, it will be understood by those skilled in the art that a linker sequence can be inserted between the TS sequence and the N-terminus or C-terminus of the transposase. In some embodiments, the linker sequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids in length. In some embodiments, one or more amino acids of the transposase portion of the fusion protein can be deleted or replaced with a linker sequence. In some embodiments, the first methionine of the transposase portion of the fusion protein can be replaced with two amino acids, for example, Gly-Thr as shown in SEQ ID NO: 25 and 26.
[0182] The TS-Tn5 fusion construct was expressed in E. coli and evaluated for expression, solubility and thermostability. The fusion of TS to the N-terminus of Tn5 increased the solubility of Tn5. The increase in solubility can be associated with the increased robustness of the transposome complex and the reduction of protein aggregation. Compared with the thermostability of the unfused Tn5 transposome, the thermostability of the TS-Tn5 transposome is substantially improved. In an example, compared with the unfused Tn5 control, the heat-induced aggregation of Tn5 is substantially reduced in the Tsf-Tn5 fusion construct.
[0183] In one application, the TS-Tn5 fusion protein is used to construct a directional RNA-seq library (e.g., TotalScript TM RNA-Seq Kit, Illumina) for sequencing on a next-generation sequencing platform (e.g., Illumina GA or HiSeq platform).
[0184] In another application, TS-Tn5 fusion proteins can be used in normalization methods. In another application, the TS solubility tag can be used to express and purify other modified (mutant) Tn5 transposases.
[0185] Antibodies specific for the TS fusion tag can be used in a pull-down process to capture transposome-tagged sequences. In one example, the TS fusion tag antibody is rabbit polyclonal.
[0186] In another application, the TS-Tn5 transposome and anti-TS antibodies can be used in a mixed transposome method. For example, a transposome reaction is performed using a Tsf-Tn5 transposome and a Tn5 transposome (i.e., not tagged with TS). The anti-TS antibody is used to specifically pull down the sequence tagged with the Tsf-Tn5 transposome.
[0187] The present invention provides a transposase fusion protein comprising a modified Tn5 transposase and a protein domain that recognizes 5-methylcytosine. The 5-methylcytosine binding protein domain can, for example, be used to target the Tn5 transpososome complex to methylated regions of the genome. In one application, the 5-methylcytosine binding domain-Tn5 transpososome complex can be used to generate methyl-rich fragmented and tagged (tagged) libraries for methylation analysis.
[0188] In some embodiments, the polypeptide domain fused to the transposase comprises an antibody binding domain of protein A. Protein A is a relatively small, compact molecule with robust folding characteristics. The antibody binding domain of protein A can, for example, be used to target the Tn5 transposome complex to regions of the genome that bind antibodies. For example, antibodies specific for 5-methylcytosine can be used to bind to and identify methylated regions of the genome. Subsequently, regions of the genome that bind antibodies can be targeted using the protein A-Tn5 fusion transposome complex for fragmentation and tagging (i.e., tagging).
[0189] In some embodiments, the polypeptide domain fused to the transposase comprises a purification tag. As used herein, the term "purification tag" refers to any peptide sequence suitable for purifying or identifying a polypeptide. The purification tag specifically binds to another part with affinity for the purification tag. Such parts that specifically bind to the purification tag are usually attached to a matrix or resin, such as agarose beads. Parts that specifically bind to the purification tag include antibodies, other proteins (e.g., protein A or streptavidin), nickel ions or cobalt ions or resins, biotin, amylose, maltose, and cyclodextrin. Exemplary purification tags include histidine (HIS) tags (such as hexa-histidine peptides) that will bind metal ions such as nickel ions or cobalt ions. Other exemplary purification tags are myc tags (EQKLISEEDL), Strep tags (WSHPQFEK), Flag tags (DYKDDDDK), and V5 tags (GKPIPNPLLGLDST). The term "purification tag" also includes "epitope tags", i.e., peptide sequences specifically recognized by antibodies. Exemplary epitope tags include FLAG tags specifically recognized by monoclonal anti-FLAG antibodies. The peptide sequence recognized by the anti-FLAG antibody consists of the sequence DYKDDDDK or a substantially identical variant thereof. In some embodiments, the polypeptide domain fused to the transposase comprises two or more tags, such as a SUMO tag and a STREP tag as exemplified in Example 1 below. The term "purification tag" also includes substantially identical variants of purification tags. As used herein, "substantially identical variants" refer to derivatives or fragments of purification tags that are modified (e.g., via amino acid substitutions, deletions, or insertions) compared to the original purification tag but retain the properties of the purification tag specifically binding to the part that specifically recognizes the purification tag.
[0190] In some embodiments, the polypeptide domain fused to the transposase comprises an expression tag. As used herein, the term "expression tag" refers to any peptide or polypeptide that can be attached to a second polypeptide and is believed to support the solubility, stability and / or expression of a recombinant polypeptide of interest. Exemplary expression tags include Fc-tags and SUMO-tags. In principle, any peptide, polypeptide or protein can be used as an expression tag.
[0191] Other useful references, e.g., for cell isolation and culture (e.g., for subsequent nucleic acid isolation), include, Freshney (1994) Culture of Animal Cells, a Manual of Basic Technique, 3rd Edition, Wiley-Liss, New York, and references cited therein; Payne et al. (1992) Plant Cell and Tissue Culture in Liquid Systems John Wiley & Sons, Inc. New York, NY; Gamborg and Phillips (eds.) (1995) Plant Cell, Tissue and Organ Culture; Fundamental Methods Springer Lab Manual, Springer-Verlag (Berlin Heidelberg New York) and Atlas and Parks (eds.) The Handbook of Microbiological Media (1993) CRC Press, Boca Raton, Fla.
[0192] Nucleic acids encoding recombinant transposases disclosed herein are also features of the embodiments presented herein. Specific amino acids can be encoded by multiple codons, and certain translation systems (e.g., prokaryotes or eukaryotic cells) typically exhibit codon biases, e.g., different organisms typically preferentially encode one of several synonymous codons for the same amino acid. Therefore, nucleic acids provided herein are optionally "codon optimized," meaning that synthetic nucleic acids include codons preferred by a specific translation system used to express a transposase. For example, when it is desired to express a transposase in a bacterial cell (or even a specific bacterial strain), nucleic acids can be synthesized to include the codons most frequently found in the genome of the bacterial cell for efficient expression of the transposase. When it is desired to express a transposase in a eukaryotic cell, a similar strategy can be adopted, e.g., nucleic acids can include codons preferred by the eukaryotic cell.
[0193] A variety of protein isolation and detection methods are known and can be used, for example, to isolate a transposase from a recombinant culture of cells expressing a recombinant transposase provided herein. A variety of protein separation and detection methods are well known in the art, including, for example, those set forth in: R. Scopes, Protein Purification, Springer-Verlag, NY (1982); Deutscher, Methods in Enzymology Vol. 182: Guide to Protein Purification, Academic Press, Inc. NY (1990); Sandana (1997) Bioseparation of Proteins, Academic Press, Inc.; Bollag et al. (1996) Protein Methods, 2. Sup. nd Edition Wiley-Liss, NY; Walker (1996) The Protein Protocols Handbook Humana Press, NJ, Harris and Angal (1990) Protein Purification Applications: A Practical Approach IRL Press at Oxford, Oxford, England; Harris and Angal Protein Purification Methods: A Practical Approach IRL Press at Oxford, Oxford, England; Scopes (1993) Protein Purification: Principles and Practice 3. Sup.rd Edition Springer Verlag, NY; Janson and Ryden (1998) Protein Purification: Principles, High Resolution Methods and Applications, 2nd Edition Wiley-VCH, NY; and Walker (1998) Protein Protocols on CD-ROM Humana Press, NJ; and references cited therein. Additional details on protein purification and detection methods can be found in Satinder Ahuja, ed., Handbook of Bioseparations, Academic Press (2000).
[0194] How to use
[0195] The altered transposases provided herein can be used in sequencing procedures, such as in vitro transposition techniques. Briefly, in vitro transposition can be initiated by contacting the transpososome complex with the target DNA. Exemplary transposition procedures and systems that can be readily adapted for use with the transposases of the present disclosure are described, for example, in WO 10 / 048605; US 2012 / 0301925; US 2013 / 0143774, each of which is incorporated herein by reference in its entirety.
[0196] For example, in some embodiments, the transposases provided herein can be used in a method of generating a library of tagged DNA fragments (e.g., for use as templates for next-generation sequencing or amplification) from a target DNA comprising any dsDNA of interest, the method comprising: incubating the target DNA in an in vitro transposition reaction with at least one transposase and a transposon end composition that forms a transposition complex with the transposase, under conditions and for a sufficient time for multiple insertions into the target DNA to occur, the transposon end composition comprising (i) a transferred strand that exhibits a transferred transposon end sequence and, optionally, an additional sequence 5'- to the transferred transposon end sequence, and (ii) a non-transferred strand that exhibits a sequence 5'- to the transferred transposon end sequence. The invention relates to a method for annealing 5'-tagged DNA fragments to a sequence complementary to a sequence at the end of the transposon, wherein each of the plurality of insertions results in the first tag comprising or consisting of the transferred strand being attached to the 5'-end of a nucleotide in the target DNA, thereby fragmenting the target DNA and generating a population of annealed 5'-tagged DNA fragments, each of the annealed 5'-tagged DNA fragments having the first tag at the 5'-end; and then, ligating the 3'-ends of the 5'-tagged DNA fragments to the first tag or to the second tag, thereby generating a library of tagged DNA fragments (e.g., comprising tagged circular ssDNA fragments or 5'- and 3'-tagged DNA fragments (or "di-tagged DNA fragments")).
[0197] In some embodiments, the amount of transposase and transposon end composition or transposome composition used in the in vitro transposition reaction is between about 1 picomoles and about 25 picomoles per 50 nanograms of target DNA per 50 microliters of reaction. In some preferred embodiments of any of the methods of the invention, the amount of transposase and transposon end composition or transposome composition used in the in vitro transposition reaction is between about 5 picomoles and about 50 picomoles per 50 nanograms of target DNA per 50 microliters of reaction. In some preferred embodiments of any of the methods of the invention, wherein the transposase is a hyperactive Tn5 transposase and the transposon end composition comprises a MEDS transposon end composition, or wherein the transposome composition comprises the hyperactive Tn5 transposase and a transposon end composition comprising a MEDS transposon end, the amount of the transposase and transposon end composition or the transposome composition used in the in vitro transposition reaction is between about 5 picomoles and about 25 picomoles per 50 nanograms of target DNA per 50 microliters of reaction. In some preferred embodiments of any of the methods of the invention, wherein the transposase is a hyperactive Tn5 transposase or a MuA transposase, the final concentration of the transposase and the transposon end composition or the transposome composition used in the in vitro transposition reaction is at least 250 nM; in some other embodiments, the final concentration of the hyperactive Tn5 transposase or the MuA transposase and their respective transposon end composition or the transposome composition is at least 500 nM.
[0198] In some embodiments, the invention provides a method for preparing and enriching a genomic DNA library for exome sequencing. In various embodiments, the method of the present invention uses the Tn5 transposase of change, for example, TS-Tn5059, for preparing a genomic library. In one embodiment, the method of the present invention provides the preparation of a genomic library with a bias of reduction, the bias of this reduction being driven by the insertion sequence bias of the reduction of the transposase (for example, TS-TN5059) of the change. Genomic DNA uses the labeling of TS-Tn5059 to provide a more complete genome coverage of a genome spanning a wide GC / AT range.
[0199] In another embodiment, the invention provides a method for preparing a genomic library using an altered transposase with increased DNA input tolerance. Tagging of genomic DNA using TS-Tn5059 provides uniform insert size across a range of DNA input amounts. In some embodiments, the invention provides a method for exome sequencing.
[0200] Labeling reaction conditions
[0201] Reaction conditions and buffers for tagmentation reactions are provided herein. In some embodiments, a divalent cation is included in the tagmentation reaction buffer. In certain embodiments, the divalent cation can be, for example, Co 2+ , Mn 2 + Mg 2+ 、Cd 2+ or Ca 2+ The divalent cation may be included in the form of any suitable salt, such as a chloride salt, e.g., CoCl 2 、MnCl 2 MgCl 2 , magnesium acetate 、 CdCl 2 or CaCl 2 In certain embodiments, the tagmentation buffer comprises CoCl 2 , as exemplified in the Examples below. As demonstrated by experimental evidence in Example 5, the addition of CoCl to the tagging buffer formulation 2 Unexpected improvement in sequence bias during tagmentation.
[0202] In certain embodiments, the tagmentation buffer can have a concentration of divalent cations of, i.e., about or greater than 0.01 mM, 0.02 mM, 0.05 mM, 0.1 mM, 0.2 mM, 0.5 mM, 1 mM, 2 mM, 5 mM, 8 mM, 10 mM, 12 mM, 15 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60, mM, 70, mM, 80, mM, 90 mM, 100 mM, or a range between any of these values, e.g., 0.01 mM to 0.05 mM, 0.02 mM to 0.5 mM, 8 mM to 12 mM, and the like. In some embodiments, the tagmentation buffer can have, i.e., about, or greater than 0.01 mM, 0.02 mM, 0.05 mM, 0.1 mM, 0.2 mM, 0.5 mM, 1 mM, 2 mM, 5 mM, 8 mM, 10 mM, 12 mM, 15 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM CoCl. 2The tagging buffer may have a concentration of, i.e., about, or greater than 0.01 mM, 0.02 mM, 0.05 mM, 0.1 mM, 0.2 mM, 0.5 mM, 1 mM, 2 mM, 5 mM, 8 mM, 10 mM, 12 mM, 15 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM MnCl. 2 The tagmentation buffer may have a concentration of, i.e., about, or greater than 0.01 mM, 0.02 mM, 0.05 mM, 0.1 mM, 0.2 mM, 0.5 mM, 1 mM, 2 mM, 5 mM, 8 mM, 10 mM, 12 mM, 15 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM MgCl. 2 The tagmentation buffer may have a concentration of, i.e., about, or greater than 0.01 mM, 0.02 mM, 0.05 mM, 0.1 mM, 0.2 mM, 0.5 mM, 1 mM, 2 mM, 5 mM, 8 mM, 10 mM, 12 mM, 15 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM CdCl 2 The tagging buffer may have a concentration of, i.e., about, or greater than 0.01 mM, 0.02 mM, 0.05 mM, 0.1 mM, 0.2 mM, 0.5 mM, 1 mM, 2 mM, 5 mM, 8 mM, 10 mM, 12 mM, 15 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM CaCl. 2 The concentration of divalent cations may be from 0.01 mM to 0.05 mM, from 0.02 mM to 0.5 mM, from 8 mM to 12 mM, or a range between any of these values, e.g., from 0.01 mM to 0.05 mM, from 0.02 mM to 0.5 mM, from 8 mM to 12 mM, and the like.
[0203] In some embodiments, the fragmentation of genomic DNA by transposase or tagmentation reaction can be performed at a temperature ranging from 25° C. to 70° C., 37° C. to 65° C., 50° C. to 65° C., or 50° C. to 60° C. In some embodiments, the fragmentation of genomic DNA by transposase or tagmentation reaction can be performed at 37° C., 40° C., 45° C., 50° C., 51° C., 52° C., 53° C., 54° C., 55° C., 56° C., 57° C., 58° C., 59° C., 60° C., 61° C., 62° C., 63° C., 64° C., or 65° C.
[0204] Nucleic acid encoding an altered transposase
[0205] Also provided herein are nucleic acid molecules encoding the altered transposases provided herein. For any given altered transposase that is a mutant version of a transposase, the amino acid sequence of the altered transposase and preferably also the wild-type nucleotide sequence encoding the transposase are known, and it is possible to obtain the nucleotide sequence encoding the mutant according to the basic principles of molecular biology. For example, in view of the fact that the wild-type nucleotide sequence encoding the Tn5 transposase is known, it is possible to derive the nucleotide sequence of any given mutant version of the Tn5 encoding one or more amino acid substitutions using the standard genetic code. Similarly, for other transposases of mutant versions, nucleotide sequences can be easily derived. Then, nucleic acid molecules with the desired nucleotide sequence can be constructed using standard molecular biology techniques known in the art.
[0206] According to the embodiments provided herein, the defined nucleic acid includes not only the same nucleic acid, but also any small base variations, including, in particular, substitutions in the case of synonymous codons (different codons specifying the same amino acid residue) resulting in degenerate codons attributed to conservative amino acid substitutions. The term "nucleic acid sequence" also includes the complementary sequence to any single-stranded sequence given for base variations.
[0207] The nucleic acid molecules described herein may also be advantageously included in suitable expression vectors to express the transposase protein encoded therein in a suitable host. The incorporation of cloned DNA into suitable expression vectors for subsequent transformation of the cells and subsequent selection of the transformed cells is well known to those skilled in the art, as provided in Sambrook et al. (1989), Molecular cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, which is incorporated by reference in its entirety.
[0208] Such expression vectors include vectors having nucleic acids according to the embodiments provided herein that are operably connected to regulatory sequences that can affect the expression of the DNA fragment, such as promoter regions. The term "operably connected" refers to the juxtaposition of components described therein that allow them to function in a desired manner. Such vectors can be transformed into suitable host cells to provide expression of the protein according to the embodiments provided herein.
[0209] Nucleic acid molecules can encode mature proteins or nucleic acid molecules of proteins with precursor sequences (prosequences), including nucleic acid molecules encoding leader sequences on preproteins (preproteins) that are then cut by host cells to form mature proteins. The vector can be, for example, a plasmid, virus or phage vector equipped with a replication origin and optionally a promoter for expressing the nucleotides and optionally a regulator of the promoter. The vector can contain one or more selectable markers, such as, for example, antibiotic resistance genes.
[0210] The regulatory elements required for expression include promoter sequences that bind RNA polymerase and instruct the initiation of transcription at an appropriate level and translation initiation sequences for ribosome binding. For example, bacterial expression vectors may include promoters such as lac promoters and Shine-Dalgarno sequences and start codon AUG for translation initiation. Similarly, eukaryotic expression vectors may include heterologous or homologous promoters of RNA polymerase II, downstream polyadenylation signals, start codon AUG and termination codons for detaching from ribosomes. Such vectors may be commercially available or assembled by the sequence described by methods well known in the art.
[0211] The transcription of the DNA encoding the transposase in higher eukaryotes can be optimized by including an enhancer sequence in the vector. An enhancer is a DNA cis-acting element that acts on a promoter to increase the level of transcription. The vector will generally also contain an origin of replication in addition to a selectable marker.
[0212] Example 1
[0213] General assay methods and conditions
[0214] The following paragraphs describe the general assay conditions used in the examples presented below.
[0215] Tagmentation of human gDNA for WGS using TN5
[0216] This section describes the tagmentation assay used in the examples below to monitor the insertion bias of transposases. Briefly, 50 ng of human genomic DNA was incubated with 5 μL of TDE1 in 10 mM Tris-acetate, pH 7.6, 25 mM magnesium acetate at 55 °C for 5 min. Then, 1 / 5 of the reaction volume of 125 mM HEPES, pH 7.5, 1 M NaCl, 50 mM MgCl 2 , followed by addition of Tn5 transposase to 100 nM and incubation for 60 min at 30° C. The reactions were then cleaned-up and amplified as described in Illumina's Nextera protocol and submitted for sequencing using a HiSeq 2000 instrument.
[0217] Tn5 transposome assembly
[0218] Tn5 was incubated with 20 μM annealed transposon in 25 mM HEPES, pH 7.6, 125 mM KCl, 18.75 mM NaCl, 0.375 mM EDTA, 31.75% glycerol for 30 min at room temperature.
[0219] Transposon assembly
[0220] Transposons were annealed individually to 40 μM in 10 mM TrisHCl, pH 7.5, 50 mM NaCl, 1 mM EDTA by heating the reaction to 94°C and slowly cooling it to room temperature. For Tn5-ME-A, Tn5 Mosaic End Sequence A14 (Tn5MEA) was annealed to the Tn5 non-transferred sequence (NTS), and for Tn5-ME-B, Tn5 Mosaic End Sequence B15 (Tn5MEB) was annealed to the Tn5 non-transferred sequence (NTS). These sequences are shown below:
[0221] Tn5MEA:5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3'
[0222] Tn5MEB:5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3'
[0223] Tn5 NTS:5'-CTGTCTCTTATACACATCT-3'
[0224] 2. Cloning and Expression of Transposase
[0225] This section describes the methods used to clone and express the various transposase mutants used in the Examples below.
[0226] The gene encoding the backbone gene sequence of the transposase was mutagenized using standard site-directed mutagenesis methods. For each mutation made, the proper sequence of the mutant gene was confirmed by sequencing the cloned gene sequence.
[0227] The Tn5 transposase gene was cloned into a modified pET11a plasmid. The modified plasmid contained the purification tag Strep Tag-II from pASK5plus and SUMO from pET-SUMO. BL21 (DE3) pLysY competent cells (New England Biolabs) were used for expression. The cells were grown at 25 °C to an OD of 600nm 0.5, and then 100 μM IPTG was used for induction. Expression was performed at 18 ° C for 19 h. Microfluidizer was used to lyse the cell pellet in 100mM TrisHCl, pH 8.0, 1MNaCl, 1mMEDTA in the presence of protease inhibitors. After lysis, the cell lysate was incubated with deoxycholate at a final concentration of 0.1% for 30 minutes. Polyethyleneimine was added to 0.5%, and then centrifuged at 30,000xg for 20 min. The supernatant was collected and mixed with an equal volume of saturated ammonium sulfate solution, and then stirred on ice for at least 1 h. The solution was then centrifuged at 30,000xg for 20 min, and the precipitate was resuspended in 100mM Tris, pH8.0, 1M NaCl, 1mM EDTA and 1mM DTT. The resuspended and filtered solution was then applied to a Streptactin column using an AKTA purification system. Use 100mM Tris, pH 8.0, 1mM EDTA, 1mM DTT, 4M NaCl, followed by 100mM Tris, pH 7.5, 1mM EDTA, 1mM DTT, 100mM NaCl washing column. Use 100mM Tris, pH 7.5, 1mM EDTA, 1mM DTT, 100mM NaCl and 5mM desthiobiotin to be eluted. Use 100mM Tris pH 7.5, 100mM NaCl, 0.2mM EDTA and 2mM DTT to load the eluate onto the heparin retention column. After washing with the same buffer, use 100mM Tris pH 7.5, 1M NaCl, 0.2mM EDTA and 2mM DTT to elute the Tn5 variant with a salt gradient. Fractions were collected, pooled, concentrated, and glycerol was added to give a final concentration of 50% before long-term storage at -20°C.
[0228] 3. IVC Analysis of Insertion Bias in E. coli Genomic DNA
[0229] Analysis of insertion sequence bias using IVC-graph data _ (intensity to cycle) is carried out. After sequencing the DNA library created using the respective DNA transposase, available data is generated. In brief, mutagenesis and expression are carried out as described above. Transposase is incubated with transposon A and B described above, and incubated with Escherichia coli genomic DNA to produce a DNA library. According to the manufacturer's instructions, each library is sequenced at least 35 cycles on the Illumina genome analyzer system (Illumina, Inc., San Diego, CA) running MiSeq fast chemistry. Illumina RTA software is used to produce base calls (base call) and intensity values in each cycle. In order to produce IVC figures, sequencing reads are compared with the Escherichia coli reference genome, and the intensity (occurrence) of each of the four bases in each cycle is used as the score mapping of the total intensity values of all comparison sequencing reads.
[0230] Example 2
[0231] Identification of insertion bias in Tn5 transposase mutants
[0232] This example describes the use of IVC-plot data (intensity versus cycle) to analyze insertion sequence bias. This data is available after sequencing a DNA library created using the corresponding DNA transposase. The analysis requires only a few sequencing reads (20k-40k) to produce stable results and can be performed in E. coli cell lysates used to express the corresponding Tn5 transposase variants and is suitable for HTS (high throughput screening) purposes.
[0233] Representative results for various single amino acid substitution Tn5 variants are shown in Figure 2-7 The variant shown indicates, for example, a loss of symmetry by substitution at position 248 ( Figure 2 ), the flattened IVC graph by substitution at position 119 ( Figure 3 ), reduced IVC in the lower half of the IVC diagram by substitution at position 125 or insertion after position 248 ( Figure 4-5 ), and increasing the repeat from 9 bp to 10 bp by using a different aromatic amino acid at position K120 ( Figure 6 , indicating a change in symmetry from 1 and 9 bp to 1 and 10 bp). These results indicate that the indicated mutations provide improved insertion bias compared to the wt control.
[0234] Example 3
[0235] Whole genome sequencing of bacterial gDNA
[0236] These experiments were performed to compare a) estimated library size / diversity and b) AT / GC-dropout of various transposase mutants. These experiments were performed with purified and activity-normalized Tn5 transposase variants. These experiments required 500k-1M sequencing reads / experiment.
[0237] Results were obtained by tagmentation of Bacillus cereus gDNA using the indicated purified Tn5 transposase variants. Enzymes were normalized to activity and set to the same as Nextera TM The activity of the commercial TDE1 enzyme sold together with the kit matches.
[0238] The mutant indicated as Tn5001 has the same amino acid sequence as SEQ ID NO:11 and is used as a "wt" control. As shown in the above table, the mutants indicated as Tn5058, Tn5059 and Tn5061 have the same amino acid sequence as SEQ ID NO:18, 19 and 20, respectively. The experiment was performed in triplicate, and the data show the mean and standard deviation of the collected data. "Estimated library size" was calculated without using optical replication, providing repeatable results.
[0239] like Fig. 8A As shown in , Tn5058 and Tn5059 have significantly reduced AT-shedding compared to Tn5001, while maintaining low GC-shedding. Figure 8B As shown in , the library size increased significantly by 1.7x (Tn5058) and 2.2x (Tn5059). These results indicate that mutants Tn5058 and Tn5059 greatly improved sequence insertion bias compared to the wild-type transposase, leading to further experiments described in Example 4 below.
[0240] Example 4
[0241] Nextera Rapid Capture Enrichment Assay for Human gDNA
[0242] The following experiments were performed with the same purified and activity-normalized Tn5 transposase variants described above in Example 3. These experiments typically required 40M-100M sequencing reads per experiment, and the sequencing data was analyzed to compare a) diversity, b) enrichment, c) coverage, d) coverage uniformity, e) penalties.
[0243] Capture was performed using Nextera Rapid Exome Capture (Illumina) according to the manufacturer's instructions. TM ) CEX pool capture probe was performed in triplicate. Fig.9AAs shown in , the indicated mutants produced significant improvements in coverage uniformity compared to the wt control Tn5001. Fig. 9B As indicated in , despite the lower average target coverage, the mutants tested produced statistically significant improvements in target coverage at 10x and 20x.
[0244] like Fig. 10A As shown in , the mutants shown produced an increase in the number of unique reads and the size of the pooled selection library. Fig. 10B As shown in , the mutants produce lower penalties compared to Tn5001, which is a multiple greater sequencing required to reach 10x, 20x or 30x coverage depth. These results indicate that the tested mutants provide reduced insertion bias and more uniform coverage compared to the control.
[0245] Example 5
[0246] Effect of labeling buffer composition on Tn5 activity
[0247] The following experiments were performed to characterize the effects of Tn5 tagmentation buffer composition and reaction conditions on library output and sequencing metrics.
[0248] In order to evaluate the influence of labeling buffer composition and reaction conditions on library output and sequencing metrics, a DNA library of Tn5 labeling was constructed using Bacillus cereus genomic DNA. Two different Tn5 transposases were used to construct a labeling library, i.e., mutant Tn5 ("TS-Tn5059") and control high-activity Tn5 ("TS-Tn5"). TS is a fusion tag for purifying Tn5 and Tn5059 proteins. Relative to the high-activity Tn5 amino acid sequence (SEQ ID NO:11), Tn5059 has 4 additional mutations K212R, P214R, G251R and A338V. TS-Tn5059 comprises a TS tag at the N-terminal of Tn5059. In some embodiments, the C-terminal of the TS-tag can be fused to the N-terminal of Tn5059 by replacing the joint of the first methionine residue. In some embodiments, the joint is Gly-Thr.
[0249] TS-Tn5059 is used at a final concentration of 10nM, 40nM and 80nM. TS-Tn5 is used at a final concentration of 4nM, 15nM and 30nM. The enzyme concentrations of TS-Tn5059 and TS-Tn5 are normalized (using standard buffer preparation) to provide approximately the same level of labeling activity, i.e., 10nM, 40nM and 80nM TS-Tn5059 has approximately the same activity as the Tn5 at 4nM, 15nM and 30nM, respectively. Each labeling library is prepared using the bacillus cereus genomic DNA of 25ng input. The genomic content of bacillus cereus is approximately 40% GC and approximately 60% AT.
[0250] The labeling buffer was prepared as a 2x formulation. The 2x formulation was as follows: standard buffer (TD; 20 mM Tris acetate, pH 7.6, 10 mM magnesium acetate, and 20% dimethylformamide (DMF); cobalt buffer (Co; 20 mM acetate, pH 7.6, and 20 mM CoCl 2 ); Cobalt + DMSO buffer (Co-DMSO; 20 mM Tris acetate, pH 7.6, 20 mM CoCl 2 and 20% dimethyl sulfoxide (DMSO)); high molecular weight buffer (HMW; 20 mM Tris acetate, pH 7.6, and 10 mM magnesium acetate); NF2 buffer (NF2; 20 mM Tris acetate, pH 7.6, 20 mM CoCl 2 and 20% DMF). Fresh preparation of CoCl 2 For each library, the tagmentation reaction was performed by mixing 20 μL of Bacillus cereus genomic DNA (25 ng), 25 μL of 2x tagmentation buffer, and 5 μL of enzyme (10x Ts-Tn5059 or 10x Ts-Tn5) in a total reaction volume of 50 μL. The reaction was incubated at 55°C for 5 minutes. After the tagmentation reaction, the DNA was purified according to the standard Nextera TM Samples were processed using the sample preparation protocol. Libraries were sequenced using Illumina SBS (sequencing-by-synthesis) chemistry on a MiSeq instrument. Sequencing runs were 2x71 cycles using the V2 MiSeq kit. Fragment size distribution in each library was evaluated on a bioanalyzer.
[0251] Fig.11Shown is a bar graph of the number of unique molecules in TS-Tn5059 and TS-Tn5 labeled DNA libraries prepared using different labeling buffers. The number of unique molecules in the library is an indication of library diversity (complexity). Each bar on the figure represents a labeled library. The experiment was repeated three times (n=3). The control library (i.e., the library prepared using standard labeling buffer) is named "enzyme-enzyme concentration-DNA input". For example, the first bar in the bar graph is labeled "TS-Tn5059-10nM-25ng" and indicates the control library prepared using TS-Tn5059 and 25ng of input DNA at a final concentration of 10nM in the standard buffer preparation. The library prepared using the modified labeling buffer preparation is named "enzyme-enzyme concentration-buffer additive-DNA input". For example, the fourth bar in the bar graph is labeled "TS-Tn5059-10nM-Co-25ng" and indicates the use of a final concentration of 10nM in the standard buffer preparation and 25ng of input DNA. 2 The library was prepared with TS-Tn5059 at a final concentration of 10 nM in the modified tagmentation buffer. The data show that the library was 2 Compared to libraries prepared in buffer containing 10 mM CoCl 2 TS-Tn5059 and TS-Tn5 tagged libraries prepared with different tagging buffers (i.e., Co, Co-DMSO, and NF2 buffer) had higher average diversity.
[0252] Fig.12 Bar graphs showing the percentage of GC dropout in TS-Tn5059 and TS-Tn5 tagged DNA libraries prepared using different tagmentation buffers. Fig.11 GC dropout can be defined as the percentage of GC-rich regions in the genome that are dropped (missing) from the tagged library. The data show that for the control TS-Tn5059 and TS-Tn5 libraries prepared using standard tagmentation buffer, the percentage of GC dropout is relatively low. The data also show that compared to the library prepared without the addition of CoCl 2 Compared to libraries prepared in buffer containing 10 mM CoCl 2The TS-Tn5059 and TS-Tn5 labeled libraries prepared with the labeling buffer (i.e., Co, Co-DMSO and NF2 buffer) have a higher GC dropout percentage (i.e., up to about 6%). The increase of GC dropout in the library prepared using the buffer comprising Co is improved by the increase in the concentration of TS-Tn5059 and TS-Tn5. For example, compared with the TS-Tn5059-10nm-25ng control library, the GC dropout percentage in the TS-Tn5059-10nm-Co-25ng library is relatively high. When the concentration of TS-Tn5059 is increased to 40nM (i.e., TS-Tn5059-40nm-Co-25ng) and 80nM (i.e., TS-Tn5059-80nm-Co-25ng), the percentage of GC dropout is reduced.
[0253] Fig.13 Bar graphs showing the percentage of AT shedding in TS-Tn5059 and TS-Tn5 tagged DNA libraries prepared using different tagmentation buffers. Fig.11 AT dropout can be defined as the percentage of AT-rich regions in the genome that are dropped (missed) from the tagged library. The data show that for the control TS-Tn5059 and TS-Tn5 libraries prepared using standard tagmentation buffer, a certain amount (i.e., about 1% to about 3% and about 7% to about 3%, respectively) of AT dropout was observed. The data also show that compared with the control TS-Tn5059 and TS-Tn5 libraries prepared using standard tagmentation buffer, the AT dropout was not significantly increased compared with the control TS-Tn5059 and TS-Tn5 libraries prepared using standard tagmentation buffer. 2 Compared to libraries prepared in buffer (i.e., standard buffer or HMW), libraries prepared in buffer containing 10 mM CoCl 2 TS-Tn5059-tagged libraries prepared with different tagmentation buffers (i.e., Co, Co-DMSO, and NF2 buffers) had lower AT dropout percentages. 2 Compared to libraries prepared in buffer containing 10 mM CoCl 2 TS-Tn5 libraries prepared with different tagmentation buffers (i.e., Co, Co-DMSO, and NF2 buffer) had lower AT shedding percentages.
[0254] Reference now Fig.12 and 13 , add CoCl to the labeling buffer 2(10 nM) (i.e., Co, Co-DMSO, and NF2 buffer) can "flip" the percentage of GC and AT shedding in the tagged library. For example, the percentage of GC shedding in the TS-Tn5059-10 nm-Co-25 ng library is relatively high ( Fig.12 ); while the percentage of AT shedding in the TS-Tn5059-10nm-Co-25ng library was relatively low (or absent) compared to the TS-Tn5059-10nm-25ng control library ( Fig.13 ).
[0255] Fig.14 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5059 library prepared using standard buffer (TD) and cobalt buffer (Co) formulations. Fig.14 Shown are curve 410 for the fragment size distribution in the Ts-Tn5059-10nM-Co-25ng library, curve 415 for the fragment size distribution in the Ts-Tn5059-40nM-Co-25ng library, curve 420 for the fragment size distribution in the Ts-Tn5059-80nM-Co-25ng library, curve 425 for the fragment size distribution in the Ts-Tn5059-10nM-TD-25ng library, curve 430 for the fragment size distribution in the Ts-Tn5059-10nM-TD-25ng library, and curve 435 for the fragment size distribution in the Ts-Tn5059-80nM-TD-25ng library. Fig.14 Also shown is a curve 440 which is a standard ladder of DNA fragment sizes in base pairs (bp). The fragment sizes in the ladder (from left to right) are shown in Table 2.
[0256]
[0257] Data show that the concentration of TS-Tn5059 used in the tagging reaction is increased from 10nM to 40nM and 80nM to shift the fragment size distribution to a smaller fragment size. In the library prepared using the standard buffer (TD) preparation, the displacement of the fragment size distribution is more obvious. For example, the fragment size distribution in the library prepared using the cobalt buffer (Co) preparation is about 3,000bp in the library (curve 410) prepared using 10nM TS-Tn5059, and about 1,000 to about 2,000bp in the library (curves 415 and 420, respectively) prepared using 40nM and 80nM TS-Tn5059. For the library (curve 435) prepared using the standard buffer (TD) preparation and 80nM TS-Tn5059, the fragment size distribution is about 200bp to about 1,000bp.
[0258] Fig.15 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5059 library prepared using cobalt-DMSO (Co-DMSO), NF2, and HMW buffer formulations. Fig.15 The following are shown: Curve 510 is a curve of the fragment size distribution in the Ts-Tn5059-10nM-Co-DMSO-25ng library; Curve 515 is a curve of the fragment size distribution in the Ts-Tn5059-10nM-NF2-25ng library; Curve 520 is a curve of the fragment size distribution in the Ts-Tn5059-10nM-HMW-25ng library; Curve 525 is a curve of the fragment size distribution in the Ts-Tn5059-40nM-Co-DMSO-25ng library; Curve 526 is a curve of the fragment size distribution in the Ts-Tn5059-40nM-NF2-25 Curve 530 is a curve of the fragment size distribution in the Ts-Tn5059-40nM-HMW-25ng library, curve 535 is a curve of the fragment size distribution in the Ts-Tn5059-80nM-Co-DMSO-25ng library, curve 540 is a curve of the fragment size distribution in the Ts-Tn5059-80nM-NF2-25ng library, and curve 550 is a curve of the fragment size distribution in the Ts-Tn5059-80nM-HMW-25ng library. Fig.15 It also shows Fig.14 A curve 440 is shown which is a standard ladder of DNA fragment sizes in base pairs (bp).
[0259] The data show that, in general, increasing the concentration of TS-Tn5059 used in the tagmentation reaction from 10 nM to 40 nM and 80 nM shifted the fragment size distribution to smaller fragment sizes. Compared to libraries prepared using Co-DMSO (e.g., curves 510 and 525), when using a library that does not contain CoCl 2 The change in fragment size distribution is more pronounced in libraries prepared with HMW buffer (eg, curves 520 and 535).
[0260] Fig.16 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5 library prepared using standard buffer formulation (TD) and cobalt buffer (Co). Fig.16 Shown are curve 610 for the fragment size distribution in the Ts-Tn5-4nM-Co-25ng library, curve 615 for the fragment size distribution in the Ts-Tn5-15nM-Co-25ng library, curve 620 for the fragment size distribution in the Ts-Tn5-30nM-Co-25ng library, curve 625 for the fragment size distribution in the Ts-Tn5-4nM-TD-25ng library, curve 630 for the fragment size distribution in the Ts-Tn5-15nM-TD-25ng library, and curve 635 for the fragment size distribution in the Ts-Tn5-30nM-TD-25ng library. Fig.16 It also shows Fig.14 A curve 440 is shown which is a standard ladder of DNA fragment sizes in base pairs (bp).
[0261] The data show that increasing the concentration of TS-Tn5 used in the tagmentation reaction from 4 nM to 15 nM and 30 nM shifts the fragment size distribution to smaller fragment sizes. The shift in fragment size distribution is more pronounced in libraries prepared using the standard buffer (TD) formulation. This observation is consistent with Fig.14 The fragment size distribution in the TS-Tn5059 library was similar.
[0262] Fig.17 Graph showing bioanalyzer traces of fragment size distribution in TS-Tn5 libraries prepared using cobalt-DMSO (Co-DMSO), NF2, and HMW buffer formulations. Fig.17The following are shown: Curve 710 is a curve of the fragment size distribution in the Ts-Tn5-4nM-Co-DMSO-25ng library; Curve 715 is a curve of the fragment size distribution in the Ts-Tn5-4nM-NF2-25ng library; Curve 720 is a curve of the fragment size distribution in the Ts-Tn5-4nM-HMW-25ng library; Curve 725 is a curve of the fragment size distribution in the Ts-Tn5-15nM-Co-DMSO-25ng library; Curve 726 is a curve of the fragment size distribution in the Ts-Tn5-15nM-NF2-25 Curve 730 is a curve of the fragment size distribution in the Ts-Tn5-15nM-HMW-25ng library, curve 735 is a curve of the fragment size distribution in the Ts-Tn5-15nM-HMW-25ng library, curve 740 is a curve of the fragment size distribution in the Ts-Tn5-30nM-Co-DMSO-25ng library, curve 745 is a curve of the fragment size distribution in the Ts-Tn5-30nM-NF2-25ng library, and curve 750 is a curve of the fragment size distribution in the Ts-Tn5-30nM-HMW-25ng library. Fig.17 It also shows Fig.14 A curve 440 is shown which is a standard ladder of DNA fragment sizes in base pairs (bp).
[0263] The data show that increasing the concentration of TS-Tn5 used in the tagmentation reaction from 4 nM to 15 nM and 30 nM shifts the fragment size distribution to smaller fragment sizes. The shift in fragment size distribution is more pronounced in libraries prepared using the standard buffer (TD) formulation. This observation is consistent with Fig.15 The fragment size distribution in the TS-Tn5059 library was similar.
[0264] Usually, now refer to Figures 14 to 17 In the presence of 10 nM CoCl 2 The fragment sizes of TS-Tn5059 and TS-Tn5 libraries prepared with tagmentation buffers containing CoCl (e.g., Co, Co-DMSO, and NF2 buffer) were larger than those prepared with buffers without CoCl. 2 Fragment sizes in TS-Tn5059 and TS-Tn5 libraries prepared with different tagmentation buffers (i.e., TD and HMW buffers).
[0265] Fig.18A , 18B, 18C and 18D show the bias plots of sequence content in the TS-Tn5 library, the bias plots of sequence content in the TS-TN5-Co library, the bias plots of sequence content in the TS-Tn5-Co-DMSO library and the bias plots of sequence content in the TS-Tn5-NF2 library, respectively. The bias plot (or intensity versus cycle number (IVC) plot) plots the ratio of observed bases (A, C, G or T) as a function of the SBS cycle number and shows the preferred sequence context that Tn5 has during tagmentation.
[0266] Fig.18A , 18B , 18C, and 18D each show a curve 810 for A content versus cycle number, a curve 815 for C content versus cycle number, a curve 820 for G content versus cycle number, and a curve 825 for T content versus cycle number. Fig.18A In the TS-TN5 library, the curve 820 representing the base G shows that about 38% of the bases observed in the first cycle are G; the curve 825 representing the base T shows that about 15% of the bases observed in the first cycle are T, and so on.
[0267] refer to Fig.18A , the data show that Tn5 sequence bias is observed in about the first 15 SBS cycles in the TS-Tn5 library prepared using a standard tagging buffer formulation. After about 15 cycles, the sequence bias gradually decreases, and the A, T, C and G contents reflect the desired genome composition. For Bacillus cereus, the genome is about 40% GC and about 60% AT, which is represented in the bias graph from about the 16th or 17th cycle to the 35th cycle, where curve 810 (i.e., A) and curve 825 (i.e., T) converge at about 30% (A+T~60%); and curve 815 (i.e., C) and curve 820 (i.e., G) converge at about 20% (C+G~40%).
[0268] refer to Fig.18B , 18C The data also showed that Tn5 sequence bias was observed within approximately the first 15 SBS cycles in the TS-Tn5-Co, TS-Tn5-Co-DMSO, and Ts-Tn5-NF2 libraries, which were constructed using a CoCl-containing 2 Again, after about 15 cycles, the sequence bias gradually decreased, and as in reference Fig.18AAs described, the A, T, C, and G contents reflect the desired genome composition. However, in the TS-Tn5-Co, TS-Tn5-Co-DMSO, and Ts-Tn5-NF2 libraries, curve 810 (i.e., A) and curve 825 (i.e., T) begin to shift toward the desired genome composition at about the 10th to 15th cycles; and curve 815 (i.e., C) and curve 820 (i.e., G) begin to shift toward the desired genome composition at about the 10th to 15th cycles. In addition, when compared with Fig.18A The bias between cycles 2-8 was reduced when compared to the control. The data show that the addition of CoCl 2 Improved Tn5 sequence bias during tagmentation.
[0269] Fig.19A , 19B , 19C and 19D show the bias map of the sequence content in the TS-Tn5059 library, the bias map of the sequence content in the TS-TN5059-Co library, the bias map of the sequence content in the TS-Tn5059-Co-DMSO library, and the bias map of the sequence content in the TS-Tn5059-NF2 library, respectively. Fig.19A , 19B The bias graphs in 19C and 19D each show curve 910 which is a curve of A content with respect to cycle number, curve 915 which is a curve of C content with respect to cycle number, curve 920 which is a curve of G content with respect to cycle number, and curve 925 which is a curve of T content with respect to cycle number.
[0270] refer to Fig.19A , the data show that in the TS-Tn5059 tagged library, Tn5059 sequence bias was observed within about the first 15 SBS cycles. After about 15 cycles, the sequence bias decreased, and as in reference Fig.18A As described, the A, T, C, and G content reflects the expected genome composition. Fig.18A Compared to the Tn5 sequence bias shown in Figure 1, mutant Tn5059 shows a reduced sequence bias. In the TS-Tn5059 library, curve 910 (i.e., A) and curve 925 (i.e., T) begin to shift toward the desired genome composition at about the 10th cycle to the 15th cycle; and curve 915 (i.e., C) and curve 920 (i.e., G) begin to shift toward the desired genome composition at about the 10th cycle to the 15th cycle.
[0271] refer to Fig.19B , 19CThe data also showed that Tn5059 sequence bias was observed within approximately the first 15 SBS cycles in the TS-Tn5059-Co, TS-Tn5059-Co-DMSO, and Ts-Tn5059-NF2 libraries, which were constructed using a CoCl-containing 2 Again, after about 15 cycles, the sequence bias gradually decreased, and as in reference Fig.18A Described, A, T, C and G content reflects the desired genome composition. However, in the TS-Tn5059-Co library, curve 910 (i.e., A) and curve 925 (i.e., T) start to shift towards the desired genome composition at about the 5th cycle; and curve 915 (i.e., C) and curve 920 (i.e., G) start to shift towards the desired genome composition at about the 5th cycle. In the TS-Tn5059-Co-DMSO and Ts-Tn5059-NF2 libraries, curve 910 (i.e., A) and curve 925 (i.e., T) start to shift towards the desired genome composition before about the 5th cycle; and curve 915 (i.e., C) and curve 920 (i.e., G) start to shift towards the desired genome composition before about the 5th cycle.
[0272] Example 6
[0273] Effect of labeling buffer composition on Mos1 activity
[0274] The following experiments were performed to characterize the effects of Tn5 tagmentation buffer composition and reaction conditions on library output and sequencing metrics.
[0275] Mos1 tagged DNA library was constructed using Bacillus cereus genomic DNA. The Mos1 transposase used to construct the tagged library was MBP-Mos1 fusion protein. Maltose binding protein (MBP) was a fusion tag used to purify Mos1 protein. MBP-Mos1 was used at a final concentration of 100 μM. Each tagged library was prepared using 50 ng of input Bacillus cereus genomic DNA.
[0276] The labeling buffer was prepared as a 2x formulation. The 2x formulations were as follows: standard buffer (TD; 20 mM Tris acetate, pH 7.6, 10 mM magnesium acetate, and 20% dimethylformamide (DMF); TD+NaCl (TD-NaCl; 20 mM Tris acetate, pH 7.6, 10 mM magnesium acetate, 20% DMF, and 200 mM NaCl); high molecular weight buffer (HMW; 20 mM Tris acetate, pH 7.6 and 10 mM magnesium acetate); HEPES (50 mM HEPES, pH 7.6, 10 mM magnesium acetate, 20% DMF); HEPES-DMSO (50 mM HEPES pH 7.6, 10 mM magnesium acetate, and 20% DMSO); HEPES-DMSO-Co (50 mM HEPES, pH 7.6, 20% DMSO and 20 mM CoCl 2 ) and HEPES-DMSO-Mn (50 mM HEPES, pH 7.6, 20% DMSO and 20 mM manganese (Mn)). CoCl 2 of labeling buffer.
[0277] For each library, the tagmentation reaction was performed by mixing 20 μL of Bacillus cereus genomic DNA (50 ng), 25 μL of 2x tagmentation buffer, and 5 μL of enzyme (10xMBP-Mos1) in a total reaction volume of 50 μL. The reaction was incubated at 30°C for 60 minutes. After the tagmentation reaction, the DNA was purified according to the standard Nextera TM Sample preparation protocol samples were processed. Libraries were sequenced using Illumina SBS (sequencing-by-synthesis) chemistry on a MiSeq instrument. Sequencing runs were 2x71 cycles.
[0278] Fig. 20Bar graphs showing the average total number of reads and average diversity in MBP-Mos1-tagged libraries prepared using different tagging buffers are shown. The total number of reads is the total number of reads from the flow cell. Diversity is the number of unique molecules in the library and is used as an indication of library complexity. Each pair of bars on the graph represents a tagged library. The experiment was repeated three times (n=3). The first two bars, EZTn5-std-Bacillus cereus and NexteraV2-30C, are comparative libraries prepared using Tn5 and standard buffer preparations at 55°C and 30°C, respectively. Libraries prepared using MBP-Mos1 for tagging reactions are named "enzyme-enzyme concentration-buffer". For example, the third pair of bars is labeled "MBPMos1-100μM-TD" and indicates a library prepared using MBP-Mos1 at a final concentration of 100μM in a standard tagging buffer (TD). Data display. Effect of different buffers on the diversity of libraries prepared by Mos1 tagmentation at relatively the same number or sequencing reads. In particular, HEPES-DMSO-Mn buffer helps to increase the diversity of the library.
[0279] Fig.21 Bar graphs showing GC and AT dropout in MBP-Mos1 tagged libraries. GC and AT dropout can be defined as the percentage of GC-rich regions and the percentage of AT-rich regions in the genome that are dropped (deleted) from the tagged library, respectively. Fig. 20 The named libraries described in . The data show that libraries prepared using EZTn5 and NexteraV2 (i.e., Tn5 transposase) have essentially no GC fall-off, but about 7% and about 5% of AT-rich regions, respectively, fall off from the tagged libraries. Libraries prepared using MBP-Mos1 and standard tagging buffer (MBPMos1-100μM-TD) have essentially no AT fall-off, but about 2% or less of GC-rich regions fall off from the tagged libraries. The GC fall-off percentage in the MBP-Mos1-tagged library is affected by the composition of the tagging buffer. The GC fall-off percentage increased in the MBP-Mos1-tagged library prepared using HMW, HEPES, HEPES-DMSO, HEPES-DMSO-Co, and HEPES-DMSO-Mn buffers.
[0280] Fig.22A , 22B , 22C and 22D The bias plots of sequence content in the Mos1-HEPES library, the bias plots of sequence content in the Mos1-HEPES-DMSO library, the bias plots of sequence content in the Mos1-HEPES-DMSO-Co library, and the bias plots of sequence content in the Mos1-HEPES-DMSO-Mn library are shown respectively. Fig.22A , 22BThe bias graphs in 22C and 22D each show curve 1210 which is a curve of A content with respect to cycle number, curve 1215 which is a curve of C content with respect to cycle number, curve 1220 which is a curve of G content with respect to cycle number, and curve 1225 which is a curve of T content with respect to cycle number.
[0281] refer to Fig.22A and 22B , the data show that in the Mos1-HEPES and Mos1-HEPES-DMSO labeled libraries, Mos1 sequence bias was observed within the first few SBS cycles. In the first SBS cycle, the detection of T was about 100% throughout the flow cell. In the second SBS cycle, the detection of A was about 100% throughout the flow cell. After about 4 cycles, the sequence bias decreased, and the A, T, C, and G contents reflected the expected genomic composition.
[0282] refer to Fig. 22C and 22D , data also show that in the library of Mos1-HEPES-DMSO-Co and Mos1-HEPES-DMSO-Mn labeling, Mos1 sequence bias is observed in the first few SBS cycles. The library of Mos1-HEPES-DMSO-Co and Mos1-HEPES-DMSO-Mn labeling is a library prepared by using labeling buffers that replace magnesium (Mg) with cobalt (Co) or manganese (Mn), respectively. Again, after about 4 cycles, sequence bias decreases, and A, T, C and G content reflect the desired genome composition. However, in Mos1-HEPES-DMSO-Co and Mos1-HEPES-DMSO-Mn libraries, curves 1210 (i.e., A) and curves 1225 (i.e., T) in the 1st cycle and the 2nd cycle, the displacement towards the desired genome composition is observed. In the Mos1-HEPES-DMSO-Mn library, the displacement towards the desired genome composition is more obvious.
[0283] Example 7
[0284] TS-Tn5059 library preparation and exome enrichment protocol
[0285] In one embodiment, the methods of the invention provide an improved workflow for preparing and enriching Tn5 transposome-based exome libraries.
[0286] Fig.23 A flow chart of an example of a method 1260 for preparing and enriching a genomic DNA library for exome sequencing is shown. Method 1260 uses the TS-Tn5059 transposon and the current Modifications of certain process steps of the rapid capture protocol to provide improved library yields across a range of DNA input amounts and sequencing metrics. For example, method 1260 uses a "double-sided" solid phase reversible immobilization (SPRI) protocol (Agencourt AMPure XP beads; Beckman Coulter, Inc.) to purify tagged DNA and provides a first DNA fragment size selection step and a second DNA fragment size selection step prior to PCR amplification. In another example, a pre-concentration process is used to concentrate the fragmented DNA library prior to exome enrichment. Method 1260 includes, but is not limited to, the following steps.
[0287] In step 1270, the genomic DNA is tagged (tagged and fragmented) by the transposome pieces. The transposome simultaneously fragments the genomic DNA and adds adapter sequences to the ends, allowing for subsequent amplification by PCR. In one example, the transposome is TS-Tn5059. Upon completion of the tagging reaction, a tagging stop buffer is added to the reaction. The tagging stop buffer may be modified to ensure that the TS-Tn5059 transposome complex from the tagged DNA is fully denatured (e.g., the concentration of SDS in the stop buffer is increased from 0.1% to 1.0% SDS combined with high temperature heating).
[0288] At step 1275, a first cleanup is performed to purify the tagged DNA from the transposomes and provide a first DNA fragment size selection step. The DNA fragment size can be selected by changing the volume to volume ratio of SPRI beads to DNA (e.g., 1X SPRI = 1:1 volume of SPRI:DNA). For example, in the first size selection, the volume ratio of SPRI beads to DNA is selected to bind DNA fragments larger than a certain size (i.e., larger DNA fragments are removed from the sample), while DNA fragments smaller than a certain size are retained in the supernatant. The supernatant with size-selected DNA fragments therein is transferred to a clean reaction vessel for subsequent processing. SPRI beads with large DNA fragments thereon may be discarded. In one embodiment, the concentration of SPRI beads may vary from 0.8X to 1.5X. In one embodiment, the concentration of SPRI beads is 0.8X.
[0289] At step 1280, a second cleanup is performed to further select for DNA fragments within a certain size range. For example, the volume ratio of SPRI beads to DNA is selected to bind DNA fragments larger than a certain size (i.e., DNA fragments within a desired size range bind to the SPRI beads). Smaller DNA fragments are retained in the supernatant and discarded. The bound DNA fragments are then eluted from the SPRI beads for subsequent processing.
[0290] In optional step 1285, the DNA fragment size distribution is determined. The DNA fragment size distribution is determined, for example, using a bioanalyzer.
[0291] In step 1290, the purified tagged DNA is amplified via a limited cycle PCR procedure. The PCR step also adds index 1 (i7) and index 2 (i5) and sequencing, as well as common adapters (P5 and P7) required for subsequent cluster generation and sequencing. Since the desired DNA fragment size range is selected using the double-sided SPRI method (i.e., steps 1275 and 1280), only tagged DNA fragments within the desired size range can be used for PCR amplification. Therefore, library yield is significantly increased, and subsequent sequencing metrics (e.g., read enrichment percentage) are improved.
[0292] At step 1295, the amplified tagged DNA library is purified using a bead-based purification method.
[0293] In optional step 1300, the size distribution of DNA fragments after PCR is determined. The size distribution of DNA fragments is determined, for example, using a bioanalyzer.
[0294] In step 1310, the tagged DNA library is pre-concentrated, followed by subsequent hybridization, for exome enrichment. For example, the tagged DNA library is pre-concentrated from about 50 μl to about 10 μl. Since the tagged DNA library is pre-concentrated, the hybridization kinetics are faster and the hybridization time is reduced.
[0295] At step 1320, a first hybridization for exome enrichment is performed. For example, the DNA library is mixed with a biotinylated capture probe targeting a region of interest. The DNA library is denatured at about 95°C for about 10 minutes and hybridized with the probe at about 58°C for about 30 minutes, with a total reaction time of about 40 minutes.
[0296] In step 1325, streptavidin beads are used to capture biotinylated probes that hybridize to the targeted region of interest. Non-specifically bound DNA is removed from the beads using two heated washes. The enriched library is then eluted from the beads and prepared for a second round of hybridization.
[0297] In step 1330, a second hybridization for exome enrichment is performed using the same probes and blocks as the first hybridization. For example, the eluted DNA library from step 155 is denatured at about 95°C for about 10 minutes and hybridized at about 58°C for about 30 minutes, with a total reaction time of about 40 minutes. The second hybridization is used to ensure high specificity of the capture region.
[0298] In step 1335, streptavidin beads are used to capture biotinylated probes that hybridize to the targeted region of interest. Non-specifically bound DNA is removed from the beads using two heated washes. The exome-enriched library is then eluted from the beads and amplified by 10 cycles of PCR in preparation for sequencing.
[0299] At step 1340, the exome-enriched capture sample (ie, the exome-enriched DNA library) is purified using a bead-based purification scheme.
[0300] At step 1345, the exome-enriched DNA library is PCR amplified for sequencing.
[0301] The amplified enriched DNA library is optionally purified using a bead-based purification method at step 1350. For example, a 1X SPRI bead protocol is used to remove unwanted products (e.g., excess primers) that may interfere with subsequent cluster amplification and sequencing.
[0302] Method 100 provides library preparation and exome enrichment in about 11 hours. If optional steps 1285 and 1300 are omitted, method 1260 provides library preparation and exome enrichment in about 9 hours.
[0303] Example 8
[0304] TS-Tn5059 insertion bias
[0305] Transposases may have a certain insertion site (DNA sequence) bias in the tagging reaction. DNA sequence bias may cause certain regions of the genome (e.g., GC-rich or AT-rich) to fall off from the tagged library. For example, the Tn5 transposase has a certain bias for GC-rich regions of the genome; therefore, AT regions of the genome may fall off in the Tn5-tagged library. In order to provide more complete genome coverage, minimal sequence bias is desired.
[0306] To evaluate the impact of the TS-Tn5059 transposome on library output and sequencing metrics, standard Nextera TM The DNA library preparation kit was used to prepare TS-Tn5059-tagged DNA libraries for whole genome sequencing and Bacillus cereus genomic DNA sequencing. TS-Tn5059 was used at a final concentration of 40 nM. The reference control library was prepared using standard reaction conditions of 25 nM Nextera V2 transposomes. The library was evaluated by sequencing by synthesis (SBS).
[0307] Fig.24AA graph showing coverage in a tagged Bacillus cereus genomic DNA library prepared using the TS-Tn5059 transposome. As the GC content increases, the TS-Tn5059 transposome becomes tolerant to increasing levels of bias. Fig. 24B A graph showing coverage in a tagged Bacillus cereus genomic DNA library prepared using the NexteraV2 transposome. Fig. 24B It is shown that as the GC content increases, the Nextera V2 coverage of GC-rich regions becomes skewed, with an increasing bias. The data show that the tagged DNA library prepared using TS-Tn5059 has improved and more uniform coverage across a wide GC / AT range with lower insertion bias compared to the tagged library prepared using NexteraV2.
[0308] Fig.25A Diagram showing the gap positions and gap lengths in a tagged Bacillus cereus genomic DNA library prepared using the TS-Tn5059 transposome. Fig.25B A graph showing the locations of gaps and lengths of gaps in a tagged Bacillus cereus genome tagged DNA library prepared using the NexteraV2 transposome. The number of gaps in the TS-Tn5059 tagged library was 27. The number of gaps in the NexteraV2 tagged library was 208. The data show that the tagged DNA library prepared using the TS-Tn5059 transposome has more uniform coverage with fewer gaps than the tagged library prepared using the NexteraV2 transposome.
[0309] Example 9
[0310] TS-Tn5059 DNA input tolerance
[0311] Preparation of tagged DNA libraries uses an enzymatic DNA fragmentation step (e.g., transposome-mediated tagging) and can therefore be more sensitive to DNA input than, for example, mechanical fragmentation methods. In one example, the present The rapid capture enrichment protocol has been optimized for 50 ng of total genomic DNA input. Higher amounts of genomic DNA input can result in incomplete tagmentation and larger insert sizes, which can affect subsequent enrichment performance. Lower amounts of genomic DNA input or low-quality genomic DNA in the tagmentation reaction can produce smaller than desired insert sizes. Smaller inserts can be lost during subsequent cleanup steps and result in lower library diversity.
[0312] To evaluate the effect of different DNA input amounts on fragment (insert) size distribution, TS-Tn5059-tagged DNA libraries were prepared at various enzyme concentrations using different amounts of input genomic DNA, and the fragment sizes were compared with those obtained with other transposases, whose activities were normalized to the activity of 40 nM TS-Tn5059 and 25 ng genomic DNA input.
[0313] The size distribution of fragments generated by 40 nM TS-Tn5059, normalized TDE1 (Tn5 version-1), and normalized TS-Tn5 using 25 ng human genomic DNA was similar, as shown in Figure 2. Fig.26 and 27 As shown in .
[0314] However, TS-Tn5059 showed increased tolerance to DNA input at higher enzyme concentrations and over a wide range of input DNA amounts. Fig.28 Figure 2 shows a graph of the bioanalyzer traces of fragment size distribution in tagged genomic DNA libraries prepared using a range of DNA inputs. A 240 nM TS-Tn5059 tagged library was prepared by tagmentation, 1.8X SPRI cleanup, followed by a bioanalyzer trace. A reference control library was prepared using the current The tagged libraries were prepared using the TS-Tn5059 Rapid Capture Kit ("Nextera") and the Agilent QXT Kit ("Agilent QXT"). The tagged libraries were prepared using 25ng, 50ng, 75ng and 100ng of Bacillus cereus genomic DNA. The data show that the tagged DNA libraries prepared using the TS-Tn5059 transposomes have a more consistent fragment size distribution across the 25ng to 100ng DNA input range compared to the libraries prepared using the Nextera or Agilent QXT transposomes. When the amount of DNA input was increased from 25ng to 100ng, the yield of the tagged DNA in the TS-Tn5059 tagged libraries increased, while the fragment size distribution remained substantially the same. In contrast, at 75ng and 100ng DNA input, the Nextera and Agilent QXT fragment libraries showed a substantial shift in the DNA fragment size distribution to larger fragment sizes.
[0315] Fig.29A Graph showing the bioanalyzer traces of fragment size distribution in the TS-Tn5059-tagged library prepared by the first user using different human Coriel DNA inputs ranging from 5 ng to 100 ng. Fig.29BA graph showing a bioanalyzer trace of fragment size distribution in a TS-Tn5059 tagged library prepared by a second user. The tagged library was prepared using 5 ng, 10 ng, 25 ng, 50 ng, 75 ng and 100 ng of Bacillus cereus genomic DNA. Fig.29A and Fig.29B Both show line 1710 of the fragment size distribution in the tagged library prepared using 5 ng DNA input, line 1715 of the fragment size distribution in the tagged library prepared using 10 ng DNA input, line 1720 of the fragment size distribution in the tagged library prepared using 25 ng DNA input, line 1725 of the fragment size distribution in the tagged library prepared using 50 ng DNA input, line 1730 of the fragment size distribution in the tagged library prepared using 75 ng DNA input, and line 1735 of the fragment size distribution in the tagged library prepared using 100 ng DNA input. The data show that the fragment size distribution in the TS-Tn5059 tagged DNA library is consistent with the DNA input range of 5 ng to 100 ng. The consistency of the fragment size distribution is observed for different users.
[0316] In another example, Table 3 shows the median library insert size in TS-Tn5059-tagged DNA libraries spanning a DNA input range from 5 ng to 200 ng.
[0317]
[0318]
[0319] In yet another example, Table 4 shows the library insert size and exome enrichment sequencing metrics for TS-Tn5059 tagged DNA libraries prepared using 25 ng, 50 ng, 75 ng, and 100 ng DNA input. The data show that the read enrichment percentage (%) is about 80%. Using the current The read enrichment percentage of the tagged libraries prepared by the rapid capture enrichment protocol was about 60% (data not shown). The data also showed consistent insert sizes across the DNA input range from 15 ng to 100 ng.
[0320]
[0321] In another example, Table 5 shows the pre-enrichment library yields across a DNA input range from 25 ng to 100 ng in TS-Tn5059-tagged libraries.
[0322]
[0323] In another example, Table 6 shows the exon group enrichment sequencing metric of the DNA library of TS-Tn5059 labeling. Starting with 50ng of input DNA, 500ng, 625ng and 750ng of library DNA input were used to prepare the library for exon group enrichment. Data show that the exon group enrichment metric is consistent across a certain range of pre-enrichment library input amounts (i.e., 500ng to 750ng).
[0324]
[0325]
[0326] In yet another example, Table 7 shows the use of the current Rapid capture enrichment hybridization protocol ("NRC") and Fig.23 Exome enrichment sequencing metrics for the tagged DNA library prepared by enrichment steps 1310 to 1350 of method 1260. The data show that compared with the current Compared with libraries prepared by hybridization in the rapid capture enrichment protocol ("NRC"), Fig.23 Method 1260 improves and / or maintains exome enrichment metrics in a TS-Tn5059-tagged library prepared.
[0327]
[0328] In a separate experiment, TS-Tn5059 showed increased tolerance to DNA input at higher concentrations (normalized to 6X concentration) compared to Tn5 version-1 and TS-Tn5 transposase normalized to the same concentration. The results are shown in Figure 30-33 At 6X "normalized" concentration, Tn5 version 1 ( Fig.30 ) and TS-Tn5( Fig.31 ) Both showed fragment size distribution shifts with gDNA inputs varying between 25-100 ng. In contrast, TS-Tn5059 at 6X normalized concentration did not show significant size shifts with DNA inputs between 10-100 ng ( Fig.32 When the DNA input was increased to 200-500 ng, the fragment size distribution began to shift ( Fig.33 The results of the increased DNA input tolerance of TS-Tn5059 are summarized in Table 8 below.
[0329] Table 8: Ratios of TS-Tn5059 (nM): gDNA (ng) in the final 50 uL reaction
[0330]
[0331] Thus, for TS-Tn5059, at ratios ≥ 2.4 (nM TS-Tn5059: ng input DNA), there was no size shift, indicating increased DNA input tolerance.
[0332] Throughout this application, various publications, patents and / or patent applications are mentioned. The disclosures of these publications are hereby incorporated into this application by reference in their entirety.
[0333] Herein the term comprising is intended to be open ended and include not only the stated elements but also any additional elements.
[0334] A number of embodiments have been described. However, it will be appreciated that various changes may be made. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A fusion protein comprising a mutant Tn5 transposase having transposase activity and a polypeptide fusion domain, wherein the fusion protein comprises SEQ ID NO:
25. 2 . A transposome complex, comprising the fusion protein of claim 1 and a polynucleotide comprising a transposon end.
3. The transposome complex according to claim 2, wherein the polynucleotide further comprises a tag.
4. The transposome complex of claim 3, wherein the tag comprises a sequencing tag domain. 5 . A kit for performing an in vitro transposition reaction, comprising the transposome complex according to any one of claims 2 to 4 .
6. A method for preparing the fusion protein according to claim 1, wherein the method include: A host cell transformed with the nucleic acid of SEQ ID NO: 24 is cultured, wherein the host cell expresses the nucleic acid.
7. A method for in vitro transposition, the method comprising contacting the transposome complex of any one of claims 2 to 4 with a target DNA.
8. The method of claim 7, wherein the polynucleotide comprising the transposon end further comprises a first tag; wherein the polynucleotide comprises two strands; wherein (i) one of the strands is a transferred strand comprising the first tag at the 5' end of the transposon end, and (ii) the other strand is a non-transferred strand complementary to the transposon end of the transferred strand.
9. The method of claim 8, wherein the contacting comprises fragmenting the target DNA and ligating the transferred strands to the 5′ ends of the fragments of the target DNA resulting from the fragmentation, thereby generating a population of 5′-tagged DNA fragments comprising the first tag at the 5′ ends; and wherein the method further comprises ligating a second tag to the 3′ ends of the 5′-tagged DNA fragments to generate a library of di-tagged DNA fragments.
10. The method of claim 7, wherein the polynucleotide comprising the transposon end further comprises a first tag, the first tag comprising a sequencing tag domain, wherein the first tag is at the 5' end of the transposon end; wherein the contacting occurs under conditions where the target DNA is fragmented and the transposon end of the polynucleotide is transferred to the 5' end of the DNA fragment, thereby generating double-stranded DNA fragments, wherein the 5' end of the DNA fragment is tagged with the first tag and the 3' end of the DNA fragment has a single-stranded nick.
Citation Information
Patent Citations
Modified transposases for improved insertion sequence bias and increased DNA input tolerance
CN113005108B
Transposon end compositions and methods for modifying nucleic acids
US20100120098A1
Methods and compositions for DNA fragmentation and tagging by transposases
US20120301925A1
Methods and compositions for generating polynucleic acid fragments
US20130143774A1
System for in vitro transposition
US5925545A