Enzymatic DNA synthesis
The use of AbiK and Abi-P2 enzymes in a reversible terminator system allows for the efficient enzymatic synthesis of long DNA strands, addressing the limitations of existing methods and enabling applications like RNA-based vaccines and de novo DNA synthesis.
Patent Information
- Application Number
- JP2025539651
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-04
- Filing Date
- 2024-01-04
- Publication Date
- 2026-01-16
AI Technical Summary
Current DNA synthesis methods struggle to produce strands longer than 200 nucleotides, which is insufficient for synthesizing average protein-coding genes and eukaryotic genomes, relying on fragment stitching techniques that are inefficient and limited by polymerase precision and inability to incorporate modified nucleotides.
Employing AbiK and Abi-P2 enzymes to synthesize single-stranded DNA segments through a reversible terminator system, allowing for the enzymatic production of DNA strands up to 10,000 nucleotides long, using reverse transcriptases that can incorporate a wide range of nucleotide modifications.
Enables the efficient and cost-effective synthesis of long DNA segments with reduced chemical waste, overcoming the limitations of chemical synthesis and polymerase precision, and facilitating applications such as RNA-based vaccines and de novo DNA synthesis.
Smart Images

Figure 2026501690000001_ABST
Abstract
Description
[Background technology]
[0001] Despite its maturity, synthesizing DNA strands longer than 200 nucleotides is extremely difficult, with most DNA synthesis companies offering lengths up to 120 nucleotides. By comparison, the average protein-coding gene is around 2,000–3,000 nucleotides long, and the average eukaryotic genome is several billion nucleotides in size. Therefore, all major gene synthesis companies today rely on various “synthesize and stitch” techniques, in which overlapping fragments of 40–60 residues are synthesized by PCR and stitched together (see Young, L. et al. (2004) Nucleic Acid Res. 32, e59). [Prior art documents] [Non-patent literature]
[0002] [Non-Patent Document 1] Young, L. et al. (2004) Nucleic Acid Res.32,e59 Summary of the Invention
[0003] Provided herein are useful compositions and methods for producing synthetic de novo DNA molecules. One embodiment provides a method for synthesizing a single-stranded DNA segment, the method comprising the steps of: a) contacting a nucleotide with the enzyme AbiK and / or Abi-P2, wherein the enzyme and the nucleotide form a complex; b) removing unbound nucleotides from a); c) constructing a single-stranded DNA segment from the 3'-end of the nucleotide bound to the enzyme / nucleotide of b) by contacting the enzyme / nucleotide of b) with another nucleotide; d) removing unbound nucleotides from c); and e) repeating the contacting and removing free nucleotides until a desired sequence of the single-stranded DNA segment of a desired length is formed. In one embodiment, the enzyme is AbiK. In another embodiment, the enzyme AbiK has an amino acid sequence provided by at least one of SEQ ID NO:2, SEQ ID NO:12, SEQ ID NOs:22-29, or a sequence 70% identical to, or having a template modeling (TM) value of at least about 0.5 compared to, SEQ ID NO:2, SEQ ID NO:12, SEQ ID NO:22-39. In one embodiment, the enzyme is Abi-P2. In one embodiment, the enzyme Abi-P2 has an amino acid sequence provided by SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO:11, or a sequence 70% identical to, or having a template modeling (TM) value of at least about 0.5 compared to, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO:11.
[0004] In one embodiment, the nucleotides are linked to a reversible strand nucleotide terminator. In one embodiment, the reversible strand nucleotide terminator comprises at least one 3'-O-blocked reversible terminator and / or a 3'-O-unblocked reversible terminator. In one embodiment, the reversible strand nucleotide terminator is removed after removing unbound nucleotides and before contacting with additional nucleotides to construct the single-stranded DNA segment.
[0005] In one embodiment, the nucleotides are independently selected from A, T, G, C, or nucleotide analogs. In one embodiment, the nucleotide analogs comprise one or more non-natural nucleotides / nucleotide analogs, including 5-bromouracil (5BU), fluorescent analogs (2-aminopurine (2-AP)), 3-MI, 6-MI, 6-MAP, pyrrolo-dC, furan-modified bases, d5SICS, dNaM, 2-amino-8-(2-thienyl)purine (s), pyridin-2-one (y), 7-(2-thienyl)imidazo[4,5-b]pyridine (Ds), pyrrole-2-carbaldehyde (Pa), 4-[3-(6-aminohexanamido)-1-propynyl]-2-nitropyrrole (Px); xanthine, 5-(2,4 diaminopyrimidine), or combinations thereof.
[0006] In one embodiment, the single-stranded DNA segment is from 1 to about 10,000 nucleotides in length. In another embodiment, the single-stranded DNA segment comprises natural or unnatural nucleotides. In one embodiment, the protein or nucleotide is attached to a solid substrate, which in one embodiment is paper, ceramic, gold, glass, metal, plastic, polystyrene, a protein affinity tag resin, a protein purification column, silicon, or a combination thereof.
[0007] One embodiment provides a DNA synthesis device, the device comprising a reaction chamber; and an enzyme / nucleotide complex described herein (at least partially disposed within the reaction chamber). In one embodiment, the reaction chamber is a well, channel, cartridge, or hole. In another embodiment, the device is a flow cytometry device, a microarray, a thermal cycler, a droplet sorter, or a 96-well plate. In one embodiment, the device is an automated device. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 provides a schematic diagram of de novo sequence extension of DNA. [Figure 2] Figure 2 provides the sequence assignment of proteins AbiK and AbiP2 using MUSCLE. Exactly matching amino acids are indicated with a "*" and similar amino acids are annotated with a ":". [Figure 3] Figures 3A–3C provide structural overlays calculated using TM-align (4) for (A) AbiK (red) and Abi-P2 (blue), (B) AbiK (red) and TdT (blue), and (C) AbiK (red) and MuLV (blue). [Figure 4] 4A-4C show that proteins incorporate two different reversible terminator groups. [Figure 5] FIG. 5 provides a schematic illustration of how reversible terminators can be used to block and unblock reactions. DETAILED DESCRIPTION OF THE INVENTION
[0009] The methods provided herein are useful for producing synthetic de novo DNA molecules, such as those required for RNA-based vaccines, de novo synthesis of long DNA segments, or synthesis of DNA with one or more unnatural nucleotides.
[0010] definition The following definitions are included to provide a clear and consistent understanding of the specification and claims. As used herein, the described terms have the following meanings. All other terms and phrases used in this specification have their ordinary meanings as would be understood by one of ordinary skill in the art. Such ordinary meanings may be obtained by reference to a specialized dictionary, such as Hawley's Condensed Chemical Dictionary, 14th Edition, by R.J. Lewis, John Wiley & Sons, New York, NY, 2001.
[0011] References within this specification to "one embodiment," "an embodiment," and the like indicate that the described embodiment may include a particular aspect, feature, structure, portion, or characteristic, but not all embodiments necessarily include that aspect, feature, structure, portion, or characteristic. Moreover, such phrases may, but need not, refer to similar embodiments referenced elsewhere in this specification. Furthermore, when a particular aspect, feature, structure, portion, or characteristic is described in connection with one embodiment, it is within the knowledge of one of ordinary skill in the art that the aspect, feature, structure, portion, or characteristic also affects or relates to other embodiments, whether or not explicitly described.
[0012] The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a compound" includes a plurality of such compounds, and a compound X includes a plurality of compounds X. It should be further understood that the claims may be drafted to exclude optional elements. Accordingly, such statements are intended to serve as an antecedent basis for using exclusive terminology, such as "solely," "only," etc., in connection with the description of any element and / or claim element described herein or for using a "negative" limitation.
[0013] The term "and / or" means any one of multiple items, any combination of multiple items, or all of the multiple items with which the term is associated. The phrase "one or more" is readily understood by those of ordinary skill in the art, especially when read in the context of its use. For example, one or more substituents on a phenyl ring refers to 1-5, e.g., 1-4 if the phenyl ring is disubstituted.
[0014] As used herein, "or" should be understood to have a meaning similar to "and / or," as defined above. For example, when separating a list of multiple items, "and / or" or "or" should be interpreted to be inclusive, e.g., including not only the inclusion of at least one, but also two or more items and optionally additional unlisted items. Unless the term clearly indicates to the contrary, terms such as "only one of" or "exactly one of," or, when used in the claims, "consisting of," refer to the inclusion of multiple or exactly one of the listed elements. Generally, as used herein, when preceded by exclusive terms such as "either," "one of," "only one of," or "exactly one of," the term "or" should be interpreted only as indicating exclusive alternatives (i.e., "one or the other, but not both").
[0015] As used herein, the terms "including," "includes," "having," "has," "with," or variations thereof, are intended to be inclusive in the same manner as the term "comprising."
[0016] The term "about" may refer to a variation of ±5%, ±10%, ±20%, or ±25% of the specified value. For example, in some embodiments, "about 50%" may have a variation of 45% to 55%. For integer ranges, the term "about" may include one or two integers greater than and / or less than the integers recited at each end of the range. Unless otherwise indicated herein, the term "about" is intended to include values (e.g., weight percent) that are near the recited range and are equivalent with respect to the function of an individual component, composition, or embodiment. The term "about" may modify the endpoints of the recited range, as discussed at the top of this paragraph.
[0017] As would be understood by one of ordinary skill in the art, all numbers, including those expressing quantities of ingredients, properties such as molecular weight, reaction conditions, and the like, are approximations and are understood in all instances as optionally modified by the term "about." These numerical values may vary depending upon the desired properties to be obtained by the skilled artisan utilizing the teachings of the descriptions herein. It is also understood that such numerical values inherently contain variability necessarily resulting from the standard deviation found in their respective testing measurements.
[0018] As will be understood by those skilled in the art, for any and all purposes, particularly for providing a written description, all ranges described herein encompass any and all possible subranges and combinations thereof, as well as the individual numerical values (e.g., integers) comprising the range. A stated range (e.g., weight percent or carbon group) includes each specific numerical value, integer, decimal, or unit within that range. Any recited range can be readily recognized as fully descriptive and allowing for similar ranges that are resolved to at least 1 / 2, 1 / 3, 1 / 4, or 1 / 10 equivalents. As a non-limiting example, each range discussed herein can be readily resolved to a lower third, middle third, upper third, etc. Additionally, as will be understood by those skilled in the art, all terms such as "up to," "at least," "greater than," "less than," "more than," "or more than," and the like are inclusive of the recited numbers, and such terms refer to ranges that can then be resolved into subranges as discussed above. Similarly, all ratios described herein also include all subratios within the broader ratio. Thus, specific values recited for radicals, substituents, and ranges are for illustrative purposes only and do not exclude other stated values for radicals and substituents or other values within stated ranges.
[0019] Those skilled in the art will also readily recognize that when members are grouped together in a familiar manner, such as a Markush group, this encompasses not only the entire group listed as a whole, but also each member of the group individually and all possible subgroups of that main group.
[0020] Additionally, for all purposes, the invention encompasses not only the main group, but also the main group when one or more of the group members are absent. Thus, the invention contemplates the explicit exclusion of any one or more of the described group members. Thus, a disclaimer may be applied to any of the disclosed categories or embodiments, thereby excluding any one or more of the described elements, species, or embodiments from such category or embodiment (e.g., for use in an express negative limitation).
[0021] The term "contacting" refers to the act of touching, contacting, or bringing into close proximity or proximity, including causing a biological response, chemical response, or physical change, e.g., at the cellular or molecular level (e.g., in solution, in a reaction mixture, in vitro).
[0022] The use of the word "detect" and its grammatical variations is intended to indicate the measurement of a chemical species without quantification, while the words "determine" or "measure," along with their grammatical variations, are intended to indicate the measurement of a chemical species with quantification. The terms "detect" and "identify" are used interchangeably herein.
[0023] As used herein, an "essentially pure" preparation of a particular DNA or protein is one that is at least about 90%, at least about 95%, e.g., at least about 99% by weight of the DNA or protein in the preparation. Alternatively, purity is defined as the amount of correct DNA sequence, e.g., at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% correct DNA sequence.
[0024] A "fragment" or "segment" is a portion of a larger DNA sequence comprising at least two nucleotides. The terms "fragment" and "segment" are used interchangeably herein.
[0025] As used herein, a "functional" biological molecule is a biological molecule in a form that exhibits a property that characterizes it. For example, a functional enzyme is an enzyme that exhibits a characteristic enzymatic activity that characterizes that enzyme.
[0026] Methods for conventional molecular biology techniques are described herein. Such techniques are generally known in the art and are described in detail in methodological treatises such as, for example, "Molecular Cloning: A Laboratory Manual," 2nd ed., vol. 1-3, ed. Sambrook et al., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989; and "Current Protocols in Molecular Biology," ed. Ausubel et al., Greene Publishing and Wiley-Interscience, New York, 1992 (with periodic updates). Methods for chemical synthesis of nucleic acids are described, for example, in Beaucage and Carruthers, Tetra. Letts. 22:1859-1862, 1981, and Matteucci et al., J. Am. Chem. Soc. 103:3185, 1981.
[0027] Methods of enzymatic DNA synthesis Currently, DNA is synthesized using chemical synthesis. This process produces short (less than 60 bases) molecules without a significant increase in cost or time. Enzymatic DNA synthesis using AbiK has been shown to very rapidly produce molecules several hundred bases long. This results in a system with much less chemical waste and longer DNA molecules at significantly reduced cost.
[0028] Provided herein are compositions and methods for enzymatic de novo DNA synthesis, including Abi family proteins for use in enzymatic DNA synthesis, as illustrated in Figure 1. AbiK, Abi-P2, a protein having 70% or more sequence identity to AbiK or Abi-P2, a protein having a TM value of greater than about 0.5 relative to the AbiK or Abi-P2 protein, or a combination thereof, are used in the methods to synthesize de novo DNA sequences (Figure 1).
[0029] In step 1), the addition of a single nucleotide onto a DNA strand is achieved by flowing the nucleotide through a flow cell. Addition of a single nucleotide is achieved by using a nucleotide containing a reversible chain terminator (Figure 1 and shown below). In step 2), continuation of DNA synthesis is controlled by a specific wavelength or appropriate chemical conditions that remove the chain terminator, thereby resulting in an activated nucleotide. After the chain terminator is removed, addition of the next nucleotide is achieved by repeating steps 1 and 2.
[0030] In a flow cell setup, proteins can be immobilized on a substrate surface, or DNA can be immobilized on a substrate surface, as shown in Figure 1. The flow cell efficiently moves fluid in and out of chambers containing proteins and nucleic acids.
[0031] To release the DNA molecule from the enzyme / DNA complex, a restriction enzyme site can be synthesized at the end of the DNA strand and cleaved at the appropriate time, for example, by flowing the enzyme through a flow cell. For example, BsaI cleaves outside its recognition site, thereby ensuring that no unwanted DNA sequences remain in the product, thereby releasing the DNA from its protein substrate. Alternatively, the protein can be degraded by injecting proteinase K, thereby also releasing the DNA. Alternatively, the synthesized DNA can be used to double-strand, amplify, and elute using PCR techniques.
[0032] The method can produce nucleic acid segments of 1 to about 10,000 nucleotides in length, including segments of about 2 to about 5,000 nucleotides in length, about 2 to about 2,000 nucleotides in length, and about 2 to about 1,000 nucleotides in length, and can be used to produce nucleic acid segments of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 1 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 6 00, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500 or 10,000 nucleotides in length.
[0033] Reversible transcriptase / Abi polymerase Companies such as DNA Script, Molecular Assemblies, Nucera, Kern Systems, and Camena have been built around enzymatic DNA synthesis. However, all such companies use polymerase-type terminal transferases or DNA assembly techniques (which involve using pre-synthesized fragments as starting blocks) for synthesis. Efforts to use these enzymes have been hampered primarily by the inability of polymerases to tolerate modified nucleotides. Provided herein is the first approach to use reverse transcriptases, rather than polymerases, for DNA synthesis. Polymerases are specialized for replicating DNA with great precision, which explains their inability to utilize modified nucleotides. Reverse transcriptases, on the other hand, are much less discriminatory and can incorporate a wide range of nucleotide modifications. This makes reverse transcriptases an ideal candidate for use in enzymatic DNA synthesis. Although we refer to this protein as a reverse transcriptase herein, it is not truly a reverse transcriptase because it does not need to replicate an RNA template. This protein is also not a polymerase because it does not replicate DNA. De novo synthesized DNA (random or otherwise) has no biologically distinct family. The proteins may also be named de novo DNA enzymes or de novo polymerases.
[0034] Bacteria possess several antiphage defense strategies that act at various stages of phage infection (1). Some mechanisms prevent phage entry by blocking phage absorption to the cell surface or by inhibiting viral DNA injection into the cell. Other mechanisms target and degrade the genomic DNA of invading phages (e.g., restriction-modification and CRISPER-Cas systems (2, 3)). Furthermore, bacterial cells can respond to infection using virulence-antivirulence or defective infection systems that trigger dormancy, temporary growth arrest, or cell death (4, 5).
[0035] Abortive infection (Abi) is a process of programmed cell death that prevents virion release and subsequent spread to other bacterial cells in the population. Most Abi systems identified to date have been characterized in Escherichia coli and Lactococcus lactis (1). The protein effectors involved in these systems are usually plasmid-encoded. They have a range of activities and are thought to initiate cell death by various mechanisms. Some of these are three systems that involve the activity of proteins related to reversible transcriptases (RTs): AbiA, AbiK, and Abi-P2 (8). Methods for determining Abi (e.g., AbiK) activity have been described (see, e.g., Figel et al. Nucleic Acids Res. 2022 Sep 23;50(17):10026-10040. doi:10.1093 / nar / gkac772 and Wang et al. Nucleic Acids Res. 2011 Sep 1;39(17):7620-9. doi:10.1093 / nar / gkr397. Epub 2011 Jun 15).
[0036] AbiK The AbiK system of L. lactis is encoded by a single, constitutively transcribed gene located on the native pSRQ800 plasmid (6), the sequences of which are known to those skilled in the art and are provided herein below. This gene is also commercially available (e.g., L1-AbiK; a synthetic gene can be purchased from BioBasic to produce a recombinant protein that can be cloned into an expression vector and purified for use in the methods provided herein; ncbi.nlm.nih.gov / nuccore / U35629.2?from=3297&to=5096&strand=2).
[0037] SEQ ID NO: 1
[0038] [ka]
[0039] SEQ ID NO: 2
[0040] [ka]
[0041] ORF7(AbiK-like) [Staphylococcus aureus] (24636606) SEQ ID NO: 12
[0042] [ka]
[0043] SEQ ID NO: 13 (https: / / www.ncbi.nlm.nih.gov / nuccore / AB057421.1?report=fasta)
[0044] [ka]
[0045] Further AbiK-like proteins: >WP_253019558.1 MULTISPECIES:RNA-directed DNA polymerase (Lactococcus lactis) SEQ ID NO: 22
[0046] [ka]
[0047] >WP_256969077.1RNA-directed DNA polymerase (Enterococcus faecalis) SEQ ID NO: 23
[0048] [ka]
[0049] >WP_227259543.1 RNA-directed DNA polymerase (Vagococcus fluvialis) SEQ ID NO: 24
[0050] [ka]
[0051] >WP_081041261.1RNA-directed DNA polymerase (Lactococcus lactis) SEQ ID NO: 25
[0052] [ka]
[0053] >WP_259683536.1RNA-directed DNA polymerase (Lactococcus cremoris) SEQ ID NO: 26
[0054] [ka]
[0055] >WP_270321368.1 RNA-directed DNA polymerase (Lactococcus petaurid) SEQ ID NO: 27
[0056] [ka]
[0057] >MDY5176804.1RNA-directed DNA polymerase (Lactococcus lactis) SEQ ID NO: 28
[0058] [ka]
[0059] >WP_165719325.1 RNA-directed DNA polymerase (Lactococcus petaurid) SEQ ID NO: 29
[0060] [ka]
[0061] >WP_081199340.1RNA-directed DNA polymerase (Lactococcus lactis) SEQ ID NO: 30
[0062] [ka]
[0063] >WP_259751088.1RNA-directed DNA polymerase (Lactococcus cremoris) SEQ ID NO: 31
[0064] [ka]
[0065] >WP_270695689.1RNA-directed DNA polymerase (Lactococcus cremoris) SEQ ID NO: 32
[0066] [ka]
[0067] >WP_311792723.1 RNA-directed DNA polymerase (Lactococcus petaurid) SEQ ID NO: 33
[0068] [ka]
[0069] >WP_270502917.1 RNA-directed DNA polymerase, partial sequence (Lactococcus petaurid) SEQ ID NO: 34
[0070] [ka]
[0071] >WP_270228270.1RNA-directed DNA polymerase, partial sequence (Lactococcus garvieae) SEQ ID NO:5
[0072] [ka]
[0073] >WP_300697012.1 RNA-directed DNA polymerase (Uncultured Clostridium sp.) SEQ ID NO: 36
[0074] [ka]
[0075] >WP_285118722.1 RNA-directed DNA polymerase, partial sequence (Lactococcus petaurid) SEQ ID NO: 37
[0076] [ka]
[0077] >OHY29199.1 Hypothetical protein BI362_11970 (Streptococcus parauberis) SEQ ID NO: 38
[0078] [ka]
[0079] >WP_205288036.1RNA-directed DNA polymerase (Lactococcus cremoris) SEQ ID NO: 39
[0080] [ka]
[0081] AbiP-2 Abi-P2 proteins are encoded by P2-like prophage genes, such as enterobacterial phages P2-EC30 (orf570) and P2-EC58 (orf544). (9) These genes and the proteins they encode are known to those of skill in the art and are commercially available (e.g., a synthetic gene encoding Abi-P2 (residues 1-541) of E. coli prophage EC30 can be purchased from BioBasic to produce recombinant protein that can be cloned into an expression vector and purified for use in the methods provided herein).
[0082] SEQ ID NO: 3
[0083] [ka]
[0084] SEQ ID NO:4
[0085] [ka]
[0086] Abi-P2 Synechococcus sp(427996238) SEQ ID NO:5
[0087] [ka]
[0088] SEQ ID NO:6
[0089] [ka]
[0090] Abi-P2 Oenococcus oeni(118587264) SEQ ID NO:7
[0091] [ka]
[0092] SEQ ID NO:8
[0093] [ka]
[0094] Hypothetical reverse transcriptase [Enterobacteriaceae phage P2-EC58] (84619238) SEQ ID NO:9
[0095] [ka]
[0096] SEQ ID NO: 10
[0097] [ka]
[0098] Hypothetical protein LMHG_01920 [Listeria monocytogenes FSL N1-017] (133728112) SEQ ID NO: 11
[0099] [ka]
[0100] In one embodiment, the Abi nucleic acid sequence can optionally encode at least one tag, or the Abi protein sequence can optionally include at least one tag, to aid in purification, immobilization, and / or solubility. Exemplary tags include His (HHHHHH; SEQ ID NO:21), Myc (EQKLISEEDL; SEQ ID NO:14), HA (e.g., YPYDVPDYA; SEQ ID NO:15), GST (MSPILGYWKIKGLVQPTRLLLEYLEEKYEEHLYERDEGDKWRNKKFELGLEFPNLPYYIDGDVKLTQSMAIIRYIADKHNMLGGCPKERAEISMLEGAVLDIRYGVSRIAYSKDFETLKVDFLSKLP EMLKMFEDRLCHKTYLNGDHVTHPDFMLYDALDVVLYMDPMCLDAFPKLVCFKKRIEAIPQIDKYLKSSKYIAWPLQGWQATFGGGDHPPK; SEQ ID NO: 16), FLAG (e.g., DYKDDDD; SEQ ID NO: 17), CPB (KRRWKKNFIAVSAANRFKKISSSGAL; SEQ ID NO: 18), S-tag (KETAAAKFERQHMDS; SEQ ID NO: 19), or V5 (GKPIPNP LLGLDST; SEQ ID NO: 20), MBP (MKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNLGIEGR;SEQ ID NO: 40), or Tsf(MAEITASLVKELRERTGAGMMDCKKALTEANGDIELAIENMRKSGAIKAAKKAGNVAADGVIKTKIDGNYGIILEVNCQTDFVAKDAGFQAFADKVLDAAVAGKITDVEVLKAQFEEERVALVAKIGENINIRRVAALEGDVLGSYQHGARIGVLVAAKGADEELVKHIAMHVAASKPEFIKPEDVSAEVVEKEYQVQLDIAMQSGKPKEIAEKMVEGRMKKFTGEVSLTGQPFVMEPSKTVGQLLKEHNAEVTGFIRFEVGEGIEKVETDFAAEVAAMSKQS; SEQ ID NO: 41). A tag can be fused to the N-terminus or C-terminus of the protein, or can be contained internally.
[0101] In one embodiment, the Abi nucleic acid and / or protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 1-13 and 22-39.
[0102] As used herein, "homologous" refers to the similarity of subunit sequences between two polymeric molecules, e.g., between two nucleic acid molecules (e.g., two DNA molecules or two RNA molecules), or between two polypeptide molecules. If both subunit positions of two molecules are occupied by the same monomeric subunit, e.g., if each position of two DNA molecules is occupied by adenine, the two DNA molecules are homologous at that position. Homology between two sequences is a direct function of the number of matching or homologous positions; for example, if half of the positions in two compound sequences (e.g., five positions out of a polymer length of 10 subunits) are homologous, the two sequences are 50% homologous; if 90% of the positions (e.g., 9 out of 10) are matching or homologous, the two sequences share 90% homology. For example, 3'ATTGCC5' and 3'TATGGC5' share 50% homology.
[0103] As used herein, "homology" is used synonymously with "identity." The determination of percent identity between two nucleotide sequences or two amino acid sequences can be accomplished using a mathematical algorithm. For example, a mathematical algorithm useful for comparing two sequences is the algorithm of Karlin and Altschul (1990, Proc. Natl. Acad. Sci. USA 87:2264-2268), as modified by Karlin and Altschul (1993, Proc. Natl. Acad. Sci. USA 90:5873-5877). This algorithm is incorporated into the NBLAST and XBLAST programs of Altschul et al. (1990, J. Mol. Biol. 215:403-410), which can be accessed from the National Center for Biotechnology Information (NCBI) World Wide Web site via the universal resource locator (URL), for example, using the BLAST tool on the NCBI website. BLAST nucleotide searches may be performed with the NBLAST program (designated "blastn" on the NCBI website) using the following parameters to obtain nucleotide sequences homologous to the nucleic acids described herein: gap penalty=5; gap extension penalty=2; mismatch penalty=3; match reward=1; expectation value 10.0; and word size=11. BLAST protein searches may be performed with the XBLAST program (designated "blastn" on the NCBI website) or the NBCI "blastp" program using the following parameters to obtain amino acid sequences homologous to the protein molecules described herein: expectation value 10.0, BLOSUM62 scoring matrix. For comparison, Gapped BLAST may be utilized as described in Altschul et al. (1997, Nucleic Acids Res. 25:3389-3402) to obtain gaps in alignments.Alternatively, PSI-Blast or PHI-Blast may be used to perform an iterative search, thereby detecting distance relationships (Id.) between multiple molecules and relationships between multiple molecules that share common patterns. When utilizing BLAST, Gapped BLAST, PSI-Blast, and PHI-Blast programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) may be used.
[0104] The percent identity between two sequences can be determined using techniques similar to those described above, with or without allowing gaps. When calculating percent identity, exact matches are specifically counted.
[0105] In another embodiment, the Abi protein has a TM (template modeling) value of at least about 0.5 compared to any one of SEQ ID NOs: 2, 4, 5, 7, 9, 11, or 12.
[0106] Proteins fold into complex three-dimensional structures that determine their function. The amino acid sequence homology between AbiK and AbiP2 is only 32% as calculated by BLAST (Figure 2) (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi?PROGRAM=blastp&PAGE_TYPE=BlastSearch&BLAST_SPEC=blast2seq&LINK_LOC=blasttab&LAST_PAGE=blastp&BLAST_INIT=blast2seq). Proteins are typically defined as "similar" if they share a percent similarity / identity of 70% or more. The overall view of this sequence alignment is shown below in Figure 2.
[0107] However, protein sequences, despite their diversity, can still fold into very similar structures and perform similar functions (Pearson. Curr Protec Bioinformatics. 2013. "An Introduction to Sequence Similarity ("Homology") Searching"). A different measure can also be used to assess protein structural similarity, while also accounting for differences in protein length: Tm (template modeling) (Zhand and Skolnik. Nucleic Acids Res. 2005;33(7):2302-2309 ("TM-Align: An Algorithm for Aligning Protein Structures Based on TM-score")). TM-Align is described by Zhand and Skolnik, and the TM-Align program is available at http: / / bioinformatics.buffalo.edu / TM-align and https: / / zhanggroup.org / TM-align.
[0108] Tm values range between 0 and 1. A Tm of 1 indicates a perfect match, and a Tm of 0.2 is defined as a comparison of any two randomly selected, unrelated structures. Tm values above approximately 0.5 are likely to be statistically relevant. Tm is commonly used in structural biology to determine relationships between multiple proteins based on their structure, and has recently been cited in AlphaFold, an international protein prediction competition (Critical Assessment of Techniques for Protein Structure Prediction - CASP) (Jumper et al. Nature. 596, 583-589 (2021) "Highly accurate protein structure prediction with AlphaFold") and (https: / / predictioncenter.org / casp14 / results.cgi?view=tb-sel)).
[0109] The amino acid sequence of AbiK shares only 32% sequence similarity with that of AbiP2 (Figure 2). The calculated Tm value for the AbiK (PDB: 7R07) structure compared with that of AbiP2 (PDB: 7R08) was 0.73 (Figure 3) (https: / / zhanggroup.org / TM-align / tmp / 148448.html). This result contrasts with the evaluation of the Tm value of AbiK compared with terminal deoxynucleotidyl transferase (TdT) (PDB: 1K3J) or murine leukemia virus (MLV). TdT, another protein used in the enzymatic synthesis of DNA, has a calculated Tm of 0.34, while MLV is another reversible transcriptase (PDB: 4MH8) with a Tm of 0.35.
[0110] Because it is clear that AbiK and AbiP2 share similar biological roles in phage defense and have a common mechanism for template-free DNA synthesis, their relationship can be expressed not only by sequence homology but also by template modeling values. Evaluation of structural scores provides a determination of protein family relationships. Therefore, Tm values of 0.5 or greater (including about 0.5, about 0.55, about 0.6, about 0.65, about 0.7, about 0.75, about 0.8, about 0.85, about 0.9, about 0.95, and about 1) can be used to determine structural similarity / identity between two proteins.
[0111] Nucleotides / Nucleic Acids As used herein, the term "nucleic acid" encompasses RNA as well as single-stranded DNA, double-stranded DNA, and cDNA. Additionally, the terms "nucleic acid," "DNA," "RNA," and similar terms also include nucleic acid analogs, i.e., analogs having other than a phosphodiester backbone. For example, peptide nucleic acids (PNAs), morpholino nucleic acids, and locked nucleic acids (LNAs), as well as glycol nucleic acids (GNAs), threose nucleic acids (TNAs), and hexitol nucleic acids (HNAs). Each of these is distinguished from natural DNA or RNA by modifications to the backbone of the molecule.
[0112] By "nucleic acid" is intended any nucleic acid, whether composed of deoxyribonucleosides or ribonucleosides, whether composed of phosphodiester or modified linkages (e.g., phosphotriester, phosphoramidate, siloxane, carbonate, carboxymethyl ester, acetamidate, carbamate, thioether, bridged phosphoramidate, bridged methylene phosphonate, bridged phosphoramidate, bridged phosphoramidate, bridged methylene phosphonate, phosphorothioate, methylphosphonate, phosphorodithioate, bridged phosphorothioate, or sulfone linkages and combinations of such linkages).
[0113] The term nucleic acid also specifically includes nucleic acids composed of bases other than the five biological bases (adenine, guanine, thymine, cytosine, and uracil), including, but not limited to, unnatural nucleotides / nucleotide analogs (e.g., 5-bromouracil (5BU), fluorescent analogs (2-aminopurine (2-AP)), 3-MI, 6-MI, 6-MAP, pyrrolo-dC, furan-modified bases, d5SICS, dNaM, 2-amino-8-(2-thienyl)purine (s), pyridin-2-one (y), 7-(2-thienyl)imidazo[4,5-b]pyridine (Ds), pyrrole-2-carbaldehyde (Pa), 4-[3-(6-aminohexanamido)-1-propynyl]-2-nitropyrrole (Px); xanthine, 5-(2,4 diaminopyrimidine), and many others).
[0114] Conventional notation is used herein to describe polynucleotide sequences: the left-hand end of a single-stranded polynucleotide sequence is designated the 5'-end; the left-hand direction of a double-stranded polynucleotide sequence is designated as the 5'-direction. The direction of 5' to 3' addition of nucleotides to a nascent RNA transcript is designated as the transcription direction. The DNA strand having the same sequence as the mRNA is designated as the "coding strand," the sequence on the DNA strand that is 5' to a reference point on the DNA is designated as the "upstream sequence," and the sequence on the DNA strand that is 3' to a reference point on the DNA is designated as the "downstream sequence."
[0115] Reversible chain-terminating nucleotide Reversible terminators are modified nucleotide analogs that can reversibly terminate extension and are widely used in various sequencing technologies. The term "reversible" refers to the terminator, which is derived from a blocking group in the molecule that can be removed chemically or photochemically, thereby further extending the DNA molecule. Such reversible terminators can include 3'-O-blocked reversible terminators and / or 3'-O-unblocked reversible terminators. They also include any reversible terminators available to those skilled in the art, such as those described in Chen et al., Genomics Proteomics Bioinformatics 11 (2013) 34-40 (incorporated herein by reference).
[0116] In some embodiments, the reversible chain terminator nucleotide is at least one of the following:
[0117] [ka]
[0118] [ka]
[0119] [ka]
[0120] Or,
[0121] [ka]
[0122] (Disclosed in Ju et al. 2006. Applied Biological Sciences. 103(52):19635-19640) Reversible chain terminator nucleotides can be removed by specific wavelengths or appropriate chemical conditions.
[0123] Figure 4 shows that AbiK can prevent DNA synthesis using two reversible terminators: Group A: 3'-O-azidomethyl-dCTP (Jena Biosciences) and Group B: 3''-aminoxytriphosphate (Firebird Biomolecular Sciences), demonstrating the incorporation of reversible terminators. For the data generated in Figure 4, the conditions for incorporation of reversible terminators included 500 nM AbiK, 0.1 mM dNTP mix, 5 μM fluorescein-dUTP, 2 mM MgCl, 50 mM Tris pH 8.0, 100 mM NaCl, 5 mM DTT, and 5 mM blocker. Reactions were incubated at 37°C for 60 minutes, followed by an additional 15-minute incubation with 2 μL of proteinase K. 8 μL of the reaction mixture was then added to 8 μL of 2x urea loading buffer and loaded onto a 10% urea loading gel. The gel was run for 45 minutes at 55°C in TBE loading buffer. The gel was exposed for 2 minutes and then fluorescently imaged using a gel imager.
[0124] Figure 5 shows that this reaction can be blocked and unblocked using reversible terminators. With the experimental setup described in Figure 4 (minus all added reaction components), the reaction tube was left at room temperature while the reaction proceeded. The reaction was performed at room temperature to ensure that DNA molecules were synthesized that were short enough to show a clear size difference on the gel. After 2 minutes, two-thirds of the reaction volume was aliquoted into a separate tube, and then reversible terminator group A was added to a final concentration of 1 mM. The reaction was then incubated at 37°C for 15 minutes. During this time, blocked reactions are expected to remain unextended, while unblocked reactions are expected to extend in length. After 15 minutes, TCEP was added to half of the blocked reactions to a concentration of 50 mM. TCEP should unblock the reversible terminators, and if unblocking is successful, the synthesized DNA will continue to grow longer. All reactions were allowed to proceed for an additional 15 minutes at 37°C. The samples were then treated with proteinase K to degrade the proteins and run on a gel as illustrated in Figure 4 .
[0125] solid support Abi enzymes or nucleic acids can be attached to a solid support, for example, via a linker or tag that attaches the protein to a surface. The solid support can be any substrate to which DNA or protein can be attached, including, but not limited to, paper, ceramic, gold, glass, metal, plastic, polystyrene, protein affinity tag resins, protein purification columns, and / or silicon. One of skill in the art can attach the substrate using methods available to one of skill in the art.
[0126] References 1.Rostol, JT and Marraffini, L. (2019)(Ph)ighting phages:how bacteria resist their parasites.Cell Host Microbe,25,184-194. 2.Tock,M.R. and Dryden,D.T.(2005)The biology of restriction and anti-restriction.Curr.Opin.Microbiol.,8,466-472. 3.Nussenzweig,P.M. and Marraffini,L.A.(2020)Molecular mechanisms of CRISPR-Cas immunity in bacteria.Annu.Rev.Genet.,54,93-120. 4.Page,R. and Peti,W.(2016)Toxin-antitoxin systems in bacterial growth arrest and persistence.Nat.Chem.Biol.,12,208-214. 5.Lopatina,A.,Tal,N. and Sorek,R.(2020)Abortive infection:bacterial suicide as an antiviral immune strategy.Annu.Rev.Virol.,7,371-384. 6.Emond,E.,Holler,B.J.,Boucher,I.,Vandenbergh,P.A.,Vedamuthu,E.R.,Kondo,J.K. and Moineau,S.(1997)Phenotypic and genetic characterization of the bacteriophage abortive infection mechanism AbiK from Lactococcus lactis.Appl.Environ.Microbiol.,63,1274-1283. 7. Odegrip, R., Nilsson, AS and Hagg ard-Ljungquist, E. (2006) Identification of a gene encoding a functional reverse transcriptase within a highly variable locus in the P2-Like coliphages. J. Bacteriol., 188, 1643-1647. 8.Ma▲l▼gorzata Figiel, Marta Gapinska, Mariusz Czarnocki-Cieciura, Weronika Zajko, Ma▲l▼gorzata Sroka, Krzysztof Skowronek and Marcin Nowotny.(2022)Mechanism of protein-primed template-independent DNA synthesis by Abi polymerases.Nucleic Acids Research.,50(17),10026-10040. The embodiments have been described in sufficient detail to enable one skilled in the art to practice the invention. Other embodiments may be utilized, and changes to the formulation and methods of use may be made without departing from the scope of the invention. The detailed description is not to be taken in a limiting sense, and the scope of the present invention is defined solely by the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0127] It will be appreciated by those skilled in the art that modifications could be made to the above-described embodiments without departing from the broad inventive concepts thereof. It will be understood, therefore, that the invention is not limited to the particular embodiments disclosed, but that it is intended to cover modifications within the spirit and scope of the invention as defined by this specification.
[0128] All publications, patents and patent applications, Genbank sequences, websites, and other publications referenced throughout this disclosure are incorporated herein by reference to the same extent as if each individual publication, patent and patent application, Genbank sequence, website, and other publication was specifically and individually incorporated by reference. In the event that the definition of a term incorporated by reference conflicts with a term defined herein, this specification controls.
Claims
1. 1. A method for synthesizing a single-stranded DNA segment, comprising: a) contacting an AbiK and / or Abi-P2 enzyme with a nucleotide, wherein the enzyme and the nucleotide form a complex; b) removing unbound nucleotides from a); c) constructing a single-stranded DNA segment from the 3' end of the nucleotide bound to the enzyme / nucleotide of b) by contacting the enzyme / nucleotide of b) with another nucleotide; d) removing unbound nucleotides from c); e) repeating the contacting step and the removing step of free nucleotides until the desired sequence of the single-stranded DNA segment of the desired length is formed; A method comprising:
2. The method of claim 1, wherein the enzyme is AbiK.
3. 3. The method of claim 1 or 2, wherein the AbiK enzyme has an amino acid sequence provided in one of SEQ ID NO:2, SEQ ID NO:12, SEQ ID NOs:22-39, or a sequence that is 70% identical to SEQ ID NO:2, SEQ ID NO:12, SEQ ID NOs:22-39, or a sequence that has a template modeling (TM) value of at least about 0.5 compared to SEQ ID NO:2, SEQ ID NO:12, SEQ ID NOs:22-39.
4. The method of claim 1, wherein the enzyme is Abi-P2.
5. 3. The method of claim 1 or 2, wherein the Abi-P2 enzyme has an amino acid sequence as provided in one of SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO:11, or a sequence that is 70% identical to SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO:11, or a sequence that has a template modeling (TM) value of at least about 0.5 compared to SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO:
11.
6. The method of any one of claims 1 to 5, wherein the nucleotides are attached to reversible chain nucleotide terminators.
7. 7. The method of claim 6, wherein the reversible chain nucleotide terminators comprise at least one 3'-O-blocked reversible terminator and / or a 3'-O-unblocked reversible terminator.
8. the reversible chain nucleotide terminator is 【Chemistry 32】 , 【Transformation 33】 , 【Transformation 34】 Or, 【Chemistry 35】 The method of claim 6, comprising at least one of:
9. 9. The method of claims 6-8, wherein the reversible chain nucleotide terminator is removed after removing unbound nucleotides and before contacting with another nucleotide to construct the single-stranded DNA segment.
10. The method of any one of claims 1 to 9, wherein the nucleotides are independently selected from A, T, G, C, or nucleotide analogs.
11. 11. The method of claim 10, wherein the nucleotide analog comprises one or more unnatural nucleotides / nucleotide analogs including 5-bromouracil (5BU), fluorescent analogs (2-aminopurine (2-AP), 3-MI, 6-MI, 6-MAP, pyrrolo-dC, furan-modified bases, d5SICS, dNaM, 2-amino-8-(2-thienyl)purine (s), pyridin-2-one (y), 7-(2-thienyl)imidazo[4,5-b]pyridine (Ds), pyrrole-2-carbaldehyde (Pa), 4-[3-(6-aminohexanamido)-1-propynyl]-2-nitropyrrole (Px); xanthine, 5-(2,4 diaminopyrimidine), or combinations thereof.
12. 12. The method of any one of claims 1 to 11, wherein the single-stranded DNA segment is from 1 to about 10,000 nucleotides in length.
13. The method of any one of claims 1 to 12, wherein the single-stranded DNA segment comprises natural and unnatural nucleotides.
14. The method of any one of claims 1 to 13, wherein the enzyme or nucleotide is attached to a solid substrate.
15. 15. The method of claim 14, wherein the solid substrate is paper, ceramic, gold, glass, metal, plastic, polystyrene, protein affinity tag resin, protein purification column, silicon, or a combination thereof.
16. a reaction chamber; and the enzyme / nucleotide complex of any one of claims 1 to 15, at least partially disposed within the reaction chamber. DNA synthesizer.
17. 17. The method of claim 16, wherein the reaction chamber is a well, a channel, a cartridge, or a hole.
18. 18. The method of claim 16 or 17, wherein the device is a flow cytometry device, a microarray, or a 96-well plate.
19. The method according to any one of claims 16 to 18, wherein the device is an automated device.