Nucleic Acid Ligation Method

The biocatalytic ligation method using a fusion polypeptide with PPK and ligase domains addresses the high ATP requirement in nucleic acid ligation, achieving efficient and cost-effective ligation with ATP regeneration.

JP2025542208APending Publication Date: 2025-12-25NOVARTIS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025535921
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-20
Filing Date
2023-12-19
Publication Date
2025-12-25

Smart Images

  • Figure 2025542208000120
    Figure 2025542208000120
  • Figure 2025542208000121
    Figure 2025542208000121
  • Figure 2025542208000122
    Figure 2025542208000122
Patent Text Reader

Abstract

The present disclosure relates to biocatalytic ligation methods for generating oligonucleotides and fusion polypeptides for use in the methods. In particular, the disclosure relates to biocatalytic ligation methods that incorporate ATP regeneration and fusion polypeptides that include a polyphosphate kinase domain and an ATP-dependent nucleic acid ligase domain.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of European Patent Application No. 22215207.6, filed December 20, 2022, the contents of which are incorporated herein by reference in their entirety.

[0002] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format and is incorporated herein by reference in its entirety. The XML copy, created November 29, 2023, is titled PAT059356-WO-PCT SL.xml and is 186KB in size.

[0003] The present disclosure relates to the field of biotechnology, and in particular to a biocatalytic ligation method for generating oligonucleotides and a fusion polypeptide for use in said method. In particular, the present disclosure relates to a biocatalytic ligation method incorporating ATP regeneration and a fusion polypeptide comprising a polyphosphate kinase domain and an ATP-dependent nucleic acid ligase domain. [Background technology]

[0004] Therapeutic oligonucleotides, including small interfering RNA (siRNA) and inhibitory antisense oligonucleotides (ASO), have the potential to treat a variety of life-threatening diseases. In recent years, the number of approved oligonucleotide-based drugs has increased significantly, and the number of therapeutic oligonucleotides in clinical investigation has also increased significantly (Roberts, T.C., Langer, R., & Wood, M.J. Nature Reviews Drug Discovery 2020 19:10 19, 673-694 (2020)).

[0005] In support of green synthesis initiatives across the pharmaceutical industry, there is a great need for next-generation oligonucleotide synthesis methods that are sustainable and economical at the scale required to reach broader patient populations (Mishra, M. et al. Current Research in Green and Sustainable Chemistry 4, (2021)).

[0006] To this end, biocatalysis is being applied more frequently in the production of active pharmaceutical ingredients (APIs) because enzymes are capable of highly selective transformations under mild reaction conditions and in aqueous media (Mann, G, & Stanger, F, V, Chimia (Aarau) 74, 407-417 (2020)). Biocatalysis of short oligonucleotide fragments offers a sustainable and economical alternative to the currently used solid-phase chemical synthesis of full-length therapeutic oligonucleotides.

[0007] Shorter oligonucleotides can be synthesized more easily and with higher purity than longer oligonucleotides, simplifying downstream processing and reducing solvent waste. These short oligonucleotide fragments can then be combined using nucleic acid ligase to generate the oligonucleotide product. Nucleic acid ligases have shown remarkable tolerance to non-natural DNA / RNA containing pharmaceutically relevant chemical modifications (Kestemont, D., Herdewijn, P., & Renders, M., Curr Protoc Chem Biol 11, e62 (2019); Kestemont, D., et al., Chemical Communications 54, 6408-6411 (2018); and Nandakumar, J. & Shuman, S., Molecular Cell 16, 211-221 (2004)). The use of dsRNA ligases to synthesize siRNA products starting from short fragments (≦9 nts) containing a wide range of chemical modifications, including 2′-OMe, 2′-F modified nucleotides, phosphorothioate backbone modified nucleotides, and terminal fragments functionalized with bulky N-acetylgalactosamine (GalNAc) moieties, has been previously described (Mann, G. et al., Tetrahedron Letters 93,153696(2022)).

[0008] A major drawback of existing nucleic acid ligation reactions is their dependence on the expensive cofactor ATP. Because one molecule of ATP is converted to AMP per ligation reaction, increasing the substrate (oligonucleotide fragment) concentration necessitates increasing the ATP concentration to achieve complete ligation. In practice, excess (i.e., greater than the stoichiometric amount) of ATP is typically required to achieve complete ligation. This requirement for high ATP concentrations presents many limitations in terms of sustainability, process costs, difficulties with downstream processing, and potential cofactor by-product inhibition (Mordhorst, S. & Andexer, J.N. Natural Product Reports 37, 1316-1333 (2020)). Summary of the Invention

[0009] There is an urgent unmet need for efficient and cost-effective biocatalytic methods for producing oligonucleotides; and enzymes for use in said methods. [Brief explanation of the drawings]

[0010] [Figure 1] Chemical mechanism of ligation by ATP-dependent ligase. [Figure 2] Dose response of polyphosphate kinase 12 (PPK12) showing ATP synthesis activity starting from AMP via an ADP intermediate. Approximately 65% ​​of ATP was achieved as a maximum conversion, which is consistent with previous reports. [Figure 3] dsRNA ligase catalyzed the ligation of short chemically modified oligonucleotides to generate siRNA products. ATP was converted to AMP during the ligation reaction. Polyphosphate kinase converted the generated AMP back to ATP using polyphosphate as the phosphate donor for the reaction. [Figure 4] Enzyme dose response showing ligation activity with and without ATP regeneration as indicated in the legend. [Figure 5-1] Figure 5: Example chromatograms of ligation reaction analysis highlighting how identified substrate, intermediate, and product peaks correspond to calculated pseudo% conversion, called arbitrary units (AU). A) Product standard, B) 1.25 g / L Ligase + 10 mM ATP, C) 1.25 g / L Cligase 11 + 2.5 mM AMP + 20 mM polyphosphate, D) 1.25 g / L Ligase + 1.25 g / L Kinase + 2.5 mM AMP + 20 mM polyphosphate, E) 1.25 g / L Cligase 9 + 2.5 mM AMP + 20 mM polyphosphate, F) 1.25 g / L Cligase 2 + 2.5 mM AMP + 20 mM polyphosphate, G) 1.25 g / L Ligase + 2.5 mM AMP + 20 mM polyphosphate, H) Substrate only (no enzyme control). [Figure 5-2] (As mentioned above.) [Figure 5-3] (As mentioned above.) [Figure 5-4] (As mentioned above.) [Figure 5-5] (As mentioned above.) [Figure 5-6] (As mentioned above.) [Figure 5-7] (As mentioned above.) [Figure 5-8] (As mentioned above.) [Figure 6] SDS-PAGE of enzyme stock diluted to 0.131 g / L (left panel) and 0.261 g / L (right panel). The band observed at approximately 37 kDa corresponds to the ligase / kinase enzyme; the band at approximately 76 kDa corresponds to various Cligase constructs; and the band at approximately 25 kDa corresponds to chloramphenicol acetyltransferase (antibiotic resistance). Lanes contain the following proteins: R) protein standard; 1)–11) Cligases 1–11, respectively; 12) dsRNA ligase; 13) a mixture of dsRNA ligase and PPK; and 14) PPK12. [Figure 7] Enzyme dose response comparing the ligation and ATP synthesis activities of Clignase and the unfused enzyme. [Figure 8] Enzyme dose response comparing the ligation and ATP synthesis activities of Clignases 4 and 11 and the unfused enzyme. [Figure 9] Enzyme dose response comparing the ligation activity of the optimized bacteriophage RB69 RNA ligase 2 polypeptide of SEQ ID NO: 2 on 1 mM substrate in the presence of 10 mM ATP and with increasing concentrations of polyphosphate as indicated by the legend. [Figure 10] Comparison of the ligation activity of optimized bacteriophage RB69 RNA ligase 2 polypeptides of SEQ ID NO:2 and SEQ ID NO:88 with 1 and 3 mM substrate. [Figure 11-1]Figure 11: A. Enzyme dose response of Clignase 4.2 of SEQ ID NO: 90 comparing ligation and ATP synthesis activity in the presence of different polyphosphates with either 5 mM or 65 mM MgCl2 as indicated by the legend. B. Enzyme dose response of Clignase 4.2 of SEQ ID NO: 90 in the presence of 80 mM polyphosphate (Madrell's salt) and increasing concentrations of MgCl2 as indicated by the legend. [Figure 11-2] (As mentioned above.) [Figure 12] Enzyme dose response comparing the ligation and ATP synthesis activities of Clignases SEQ ID NO:90, SEQ ID NO:92, SEQ ID NO:94, SEQ ID NO:96, and SEQ ID NO:98 for 2.5 mM of each substrate oligonucleotide (Substrates 1-6, Table 2) in the presence of 1 mM AMP, 80 mM polyphosphate, 40 mM MgCl2, 5 mM DTT, and 100 mM MOPS buffer (pH 7.2). [Figure 13] Enzyme dose response comparing the ligation and ATP synthesis activities of Clignases SEQ ID NO: 11 and SEQ ID NO: 90 for 6 mM of each substrate oligonucleotide (Substrates 1-6, Table 2) in the presence of 0.25 mM AMP, 80 mM polyphosphate, 40 mM MgCl2, 5 mM DTT, and 100 mM MOPS buffer (pH 7.2). DETAILED DESCRIPTION OF THE INVENTION

[0011] The present disclosure provides a ligation reaction method that includes an ATP regeneration system. The ATP regeneration system overcomes the need for high concentrations of ATP and allows ligation reactions to be performed in the presence of AMP, a cheaper alternative. Advantageously, the methods described herein achieve complete ligation of oligonucleotide fragments in the presence of substoichiometric amounts of ATP or AMP. The methods described herein benefit from significantly lower costs and improved sustainability compared to methods performed in the absence of ATP regeneration, which require significantly higher ATP concentrations to achieve complete ligation.

[0012] The present disclosure also provides bifunctional fusion polypeptides comprising a PPK domain and a ligase domain. These fusion polypeptides are particularly well suited for industrial biocatalytic ligation methods because they can be produced more quickly, efficiently, and at lower cost than producing separate PPK and ligase enzymes. As demonstrated herein, linking a ligase and a PPK enzyme unexpectedly resulted in a functional fusion polypeptide that retained ligase and PPK activity. Furthermore, as demonstrated herein, linking a ligase and a PPK enzyme unexpectedly improved ligase activity compared to ligase activity in a reaction mixture containing the unlinked enzymes.

[0013] The present disclosure provides a method for generating an oligonucleotide from two or more oligonucleotide fragments, the method comprising contacting: i. two or more oligonucleotide fragments; ii. an ATP-dependent nucleic acid ligase; iii. a polyphosphate kinase (PPK); iv. adenosine triphosphate (ATP) and / or adenosine monophosphate (AMP); v. a polyphosphate; and vi. a divalent cation, thereby providing an oligonucleotide.

[0014] The present disclosure also provides the use of an ATP-dependent nucleic acid ligase and a PPK in generating an oligonucleotide from two or more oligonucleotide fragments.

[0015] In some embodiments, the two or more oligonucleotide fragments comprise two or more RNA oligonucleotide fragments. In some embodiments, the ATP-dependent nucleic acid ligase is an RNA ligase. In some embodiments, the RNA ligase is a double-stranded RNA ligase. In some embodiments, the RNA ligase is a member of the RNA ligase 2 family. In some embodiments, the RNA ligase is bacteriophage RB69 RNA ligase 2.

[0016] In some embodiments, the two or more oligonucleotide fragments comprise two or more DNA oligonucleotide fragments. In some embodiments, the ATP-dependent nucleic acid ligase is a DNA ligase. In some embodiments, the DNA ligase is a T4 DNA ligase.

[0017] In some embodiments, the PPK is PPK12 or ajPAP.

[0018] In some embodiments, the ATP-dependent nucleic acid ligase and the PPK are linked.

[0019] In some embodiments, the ATP-dependent nucleic acid ligase and the PPK are linked via a polypeptide linker, ie, the PPK is located at the N-terminus of the linker and the ATP-dependent nucleic acid ligase is located at the C-terminus of the linker.

[0020] In some embodiments, the ATP-dependent nucleic acid ligase comprises a purification tag. In some embodiments, the PPK comprises a purification tag. In some embodiments, the linker comprises a purification tag.

[0021] In some embodiments, the linker is a polypeptide linker comprising at least 3 amino acids, and optionally at least 6 amino acids.

[0022] In some embodiments, the linker comprises an amino acid sequence selected from: a) HHHHHH (SEQ ID NO: 19), optionally HHHHHHHHHHH (SEQ ID NO: 20); b) ENLYFQS (SEQ ID NO: 21); c) ENLYFQG (SEQ ID NO: 22); d) SSGSSG (SEQ ID NO: 23); e) GSAGSAAGSGEF (SEQ ID NO: 24); and / or f) GSSGSGSSSGGSSSSGSS (SEQ ID NO: 25).

[0023] In some embodiments, the polyphosphate is a polyphosphate salt. In some embodiments, the polyphosphate salt is sodium polyphosphate (Madrell's salt) or sodium hexametaphosphate (Graham's salt).

[0024] In some embodiments, the divalent cation cofactor is Mg 2+ or Mn 2+ In an embodiment, the method is carried out at a divalent cation concentration of 5 to 100 mM, optionally 30 to 50 mM.

[0025] In some embodiments, the method is carried out using sub-stoichiometric concentrations of ATP and / or AMP.

[0026] In some embodiments, the method further comprises purifying the oligonucleotide.

[0027] In some embodiments, the oligonucleotide is up to 60 nucleotides in length.

[0028] In some embodiments, each of the oligonucleotide fragments is 4 to 16 nucleotides in length, optionally 6 to 9 nucleotides in length.

[0029] In some embodiments, the oligonucleotide fragment is single-stranded.

[0030] In some embodiments, the oligonucleotide fragments are double-stranded, and optionally one or more of the double-stranded oligonucleotide fragments comprises one or two single-stranded overhangs.

[0031] In some embodiments, one or more of the oligonucleotide fragments comprises a chemical modification. In some embodiments, the chemical modification is selected from: (a) a modified backbone optionally selected from phosphorothioate (e.g., chiral phosphorothioate) or methylphosphonate internucleotide linkages; (b) optionally 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), 2'-O-methyl (2'-O-Me), 2'-O-methyl (2'-O ... modified nucleotides selected from 2'-O-methylacetamide (2'-O-NMA), locked nucleic acid (LNA), glycol nucleic acid (GNA), phosphoramidates (e.g., mesyl phosphoramidate), 2',3'-seconucleotide mimics, 2'-F-arabinonucleotides, abasic nucleotides, 2'-amino modified nucleotides, 2'-alkyl modified nucleotides, morpholino nucleotides, vinyl phosphonates (e.g., 5'-vinyl phosphonate), and cyclopropyl phosphonate deoxyribonucleotides; and / or (c) conjugation to a ligand (optionally wherein the ligand comprises one or more N-acetylgalactosamine (GalNAc) derivatives).

[0032] In some embodiments, the ATP-dependent nucleic acid ligase and / or PPK is immobilized. In some embodiments, the ATP-dependent nucleic acid ligase and / or PPK is immobilized on a solid material by chemical bonding or physical adsorption.

[0033] The present disclosure also provides compositions comprising: i. an ATP-dependent nucleic acid ligase; ii. a PPK; iii. ATP and / or AMP; iv. a divalent cation; and v. a polyphosphate. In some embodiments, the composition further comprises two or more oligonucleotide fragments.

[0034] The present disclosure also provides kits comprising: i. an ATP-dependent nucleic acid ligase; ii. a PPK; iii. ATP and / or AMP; iv. a polyphosphate; v. a divalent cation; and vi. instructions for use in a method of generating an oligonucleotide from two or more oligonucleotide fragments.

[0035] In some embodiments, the polyphosphate is a polyphosphate salt, hi some embodiments, the polyphosphate salt is Graham's salt or Maddrell's salt.

[0036] In some embodiments, the divalent cation is Mg 2+ or Mn 2+ In some embodiments, the concentration of the divalent cation is 5 to 100 mM, optionally 30 to 50 mM.

[0037] The present disclosure also provides a fusion polypeptide comprising: a) a PPK domain; and b) an ATP-dependent nucleic acid ligase domain.

[0038] In some embodiments, the fusion polypeptide comprises a linker.

[0039] In some embodiments, the PPK is PPK12 or ajPAP.

[0040] In some embodiments, the PPK domain comprises an amino acid sequence having at least 85% identity to the amino acid sequence of any one of SEQ ID NOs:5-7.

[0041] In some embodiments, the ATP-dependent nucleic acid ligase domain is an RNA ligase domain.

[0042] In some embodiments, the RNA ligase domain is a double-stranded RNA (dsRNA) ligase domain.

[0043] In some embodiments, the dsRNA ligase is a member of the RNA ligase 2 family.

[0044] In some embodiments, the dsRNA ligase is bacteriophage RB69 RNA ligase 2.

[0045] In some embodiments, the ATP-dependent nucleic acid ligase domain is a DNA ligase domain.

[0046] In some embodiments, the DNA ligase domain is a T4 DNA ligase domain.

[0047] In some embodiments, the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 1-4. In some embodiments, the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of SEQ ID NO: 88. In some embodiments, the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 1-4 or 88.

[0048] In some embodiments, the linker is positioned between the PPK domain and the ATP-dependent nucleic acid ligase domain.

[0049] In some embodiments, the PPK domain is located at the N-terminus of the linker and the ATP-dependent nucleic acid ligase domain is located at the C-terminus of the linker.

[0050] In some embodiments, the fusion polypeptide comprises a purification tag.

[0051] In some embodiments, the linker comprises a purification tag, hi some embodiments, the purification tag is located at the N-terminus and / or C-terminus of the fusion polypeptide.

[0052] In some embodiments, the linker is a polypeptide linker comprising at least 3 amino acids, optionally at least 6 amino acids. In some embodiments, the linker comprises an amino acid sequence selected from: a) HHHHHH (SEQ ID NO: 19), optionally HHHHHHHHHHH (SEQ ID NO: 20); b) ENLYFQS (SEQ ID NO: 21); c) ENLYFQG (SEQ ID NO: 22); d) SSGSSG (SEQ ID NO: 23); e) GSAGSAAGSGEF (SEQ ID NO: 24); and / or f) GSSGSGSSSGGSSSSGSS (SEQ ID NO: 25).

[0053] In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 8-18. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 90, 92, 94, 96, and 98. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 8-18, 90, 92, 94, 96, and 98.

[0054] In some embodiments, the ATP-dependent nucleic acid ligase and the PPK are provided as a fusion polypeptide as described herein.

[0055] The present disclosure also provides nucleic acid molecules encoding the fusion polypeptides described herein.

[0056] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 85% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 34 to 36. SEQ ID NOs: 34 to 36 are nucleic acid sequences that encode the amino acid sequences of SEQ ID NOs: 5 to 7, respectively.

[0057] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 85% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 30-33. In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 85% sequence identity to the nucleic acid sequence of SEQ ID NO: 87. In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 85% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 30-33 or 87. SEQ ID NOs: 30-33 and 87 are nucleic acid sequences that encode the amino acid sequences of SEQ ID NOs: 1-4 and 88, respectively.

[0058] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 85% sequence identity to the nucleic acid of any one of SEQ ID NOs: 37-47. In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 85% sequence identity to the nucleic acid of any one of SEQ ID NOs: 89, 91, 93, 95, or 97. In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 85% sequence identity to the nucleic acid of any one of SEQ ID NOs: 37-47, 89, 91, 93, 95, or 97. SEQ ID NOs: 37-47, 89, 91, 93, 95, and 97 are nucleic acid sequences that encode the amino acid sequences of SEQ ID NOs: 8-18, 90, 92, 94, 96, and 98, respectively.

[0059] The present disclosure also provides vectors comprising the nucleic acids described herein.

[0060] In some embodiments, the vector is selected from a plasmid, cosmid, bacteriophage, or viral vector.

[0061] The present disclosure also provides host cells comprising the nucleic acid molecules described herein or the vectors described herein. In some embodiments, the host cell is E. coli.

[0062] The present disclosure also provides the use of the fusion polypeptide described herein in ATP-dependent nucleic acid ligation reaction. In some embodiments, the ATP-dependent nucleic acid ligation reaction is an ATP-dependent RNA ligation reaction (for example, where the ATP-dependent nucleic acid ligase domain is a dsRNA ligase domain).

[0063] In some embodiments, the rate of nucleic acid ligation exceeds the rate of nucleic acid ligation of a control; the control comprises: (a) a first protein comprising a PPK domain of a fusion polypeptide; and (b) a second protein comprising an ATP-dependent nucleic acid ligase domain of a fusion polypeptide, wherein the first and second proteins are not linked.

[0064] In some embodiments, the rate of RNA ligation exceeds the rate of RNA ligation of a control; the control comprises: (a) a first protein comprising a PPK domain of a fusion polypeptide; and (b) a second protein comprising an ATP-dependent nucleic acid ligase domain of a fusion polypeptide, wherein the ATP-dependent nucleic acid ligase domain is a dsRNA ligase domain, and wherein the first and second proteins are not linked.

[0065] The present disclosure also provides the use of a fusion polypeptide described herein in a method for generating an oligonucleotide from two or more oligonucleotide fragments.

[0066] In some embodiments, the oligonucleotide is a therapeutic oligonucleotide.

[0067] In some embodiments, the oligonucleotide product is at least 80% pure, optionally the oligonucleotide product is at least 85% pure, at least 90% pure, at least 95% pure, and optionally the oligonucleotide product is at least 98% pure.

[0068] definition Unless otherwise clearly defined, the technical and scientific terms used in this disclosure have the meanings commonly understood by those skilled in the art to which this invention belongs. The following references provide those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The HarperCollins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless otherwise specified.

[0069] As used throughout this disclosure, the articles "a" and "an" refer to one or to more than one (e.g., to at least one) of the grammatical object of the article.

[0070] The term "and / or" means "and" or "or" unless otherwise indicated.

[0071] As used herein, the term "about" typically refers to the value immediately following the term "about." For example, "about 15 or more nucleotides" typically refers to 15 or more nucleotides. In some aspects and embodiments, the term "about" encompasses values ​​that are ±1, 2, or 3 of the stated value. For example, "about 15 or more nucleotides" can refer to 15 ±3 nucleotides, e.g., 12, 13, 14, 15, 16, 17, or 18 nucleotides.

[0072] As used herein, "ligation" refers to the enzymatic joining of two adjacent nucleotides, for example, via a phosphodiester bond. ATP-dependent nucleic acid ligases are enzymes that use ATP to catalyze the formation of a covalent bond between two adjacent nucleotides.

[0073] As used herein, the term "oligonucleotide" refers to a nucleic acid that typically contains up to 100 nucleotides. Unless otherwise clearly defined, the term "oligonucleotide" encompasses both single-stranded and double-stranded oligonucleotides. Oligonucleotides can contain DNA and / or RNA. For example, one portion of an oligonucleotide can be double-stranded DNA, while another portion can be double-stranded RNA that forms a DNA-RNA chimera.

[0074] The term "therapeutic oligonucleotide" refers to an oligonucleotide that can provide a therapeutic effect, for example, by interacting with a biomolecule and / or regulating gene expression. Therapeutic oligonucleotides include, but are not limited to, RNA interference (RNAi) agents and antisense oligonucleotides (ASOs). RNAi is a post-transcriptional targeted gene silencing technique that uses an RNAi agent to degrade messenger RNA (mRNA) containing the same sequence as the RNAi agent. ASOs are single-stranded nucleic acids that can be used to target mRNA derived from a gene of interest. ASOs can alter gene expression through many mechanisms, including direct steric blocking of mRNA and ribonuclease H (RNase H)-mediated degradation of mRNA.

[0075] Non-limiting examples of RNAi agents include siRNA (small interfering RNA), dsRNA (double-stranded RNA), shRNA (short hairpin RNA), and miRNA (microRNA). RNAi agents also include, by way of further non-limiting example, locked nucleic acids (LNA), morpholinos, UNAs, threose nucleic acids (TNAs), glycol nucleic acids (GNAs), peptide nucleic acids (PNAs), and fluoro-arabino nucleic acids (FANAs). RNAi agents also include molecules in which one or more strands are a mixture of RNA, DNA, LNAs, morpholinos, UNAs (unlocked nucleic acids), TNAs, GNAs, and / or FANAs. As a non-limiting example, one or both strands of an RNAi agent can be RNA, except that, for example, one or more RNA nucleotides are replaced by DNA, LNAs, morpholinos, UNAs, TNAs, GNAs, and / or FANAs. In some embodiments, one or both strands of the RNAi agent can be nicked, and both strands can be the same length, or one strand can be shorter than the other. The oligonucleotides of the invention can be any of the RNAi agents described herein.

[0076] As used herein, the term "oligonucleotide fragment" refers to a nucleic acid that can be ligated to one or more additional oligonucleotide fragments to provide an oligonucleotide product. Unless expressly defined otherwise, the term "oligonucleotide fragment" encompasses both single-stranded and double-stranded oligonucleotide fragments. Each oligonucleotide fragment corresponds to a portion of an oligonucleotide product. As used herein, a "terminal oligonucleotide fragment" refers to a nucleic acid that corresponds to a terminal (e.g., 5' or 3') portion of an oligonucleotide product.

[0077] The term "overhang" or "nucleotide overhang" as used herein refers to at least one unpaired nucleotide that protrudes from at least one end of the two strands of a double-stranded nucleotide. In some embodiments, when the 3' end of one strand extends beyond the 5' end of the other strand, or vice versa, this forms a nucleotide overhang, for example, an unpaired nucleotide forms an overhang. The overhang that is complementary to the overhang of a second oligonucleotide fragment can be referred to as a "sticky end". The oligonucleotide fragments described herein can have one or two sticky ends.

[0078] "Blunt" or "blunt-ended" refers to the absence of unpaired nucleotides at the ends of a double-stranded nucleic acid, i.e., the absence of nucleotide overhangs. A "blunt-ended" oligonucleotide or oligonucleotide fragment is an oligonucleotide that is double-stranded throughout its entire length, i.e., an oligonucleotide that has no nucleotide overhangs at either end of the molecule.

[0079] A double-stranded nucleic acid comprises two antiparallel and substantially complementary nucleic acid strands, called "sense" and "antisense" strands. In the context of a double-stranded RNAi agent, "antisense strand" refers to the strand of RNAi that comprises a region that is substantially complementary to a target sequence, such as an mRNA sequence. "Sense strand" refers to the strand of RNAi that comprises a region that is substantially complementary to a region of the antisense strand. The sense strand and antisense strand of an RNAi agent can be called passenger strand and guide strand, respectively.

[0080] "Substantially complementary" sequences may be perfectly complementary or may contain one or more mismatches upon hybridization, while retaining the ability to hybridize under conditions most relevant to their end use.

[0081] The stoichiometric concentration of cofactor is the theoretical concentration required to achieve complete ligation in a given ligation reaction.Those skilled in the art can easily derive the stoichiometric concentration of ATP required to achieve complete ligation based on the concentration of oligonucleotide fragments and the number of ligation reactions required to produce oligonucleotide products.For example, a ligation reaction that uses 1 mM substrate and requires 4 ligation reactions has a stoichiometric ATP concentration of 4 mM.

[0082] "Conversion" refers to the enzymatic conversion of a substrate to the corresponding product. "Percentage of conversion" or "conversion" refers to the proportion of oligonucleotide fragments that are converted to oligonucleotides under specific conditions within a given time period. Thus, the "enzymatic activity" or "activity" of a ligase can be expressed as the "percent conversion" of oligonucleotide fragments to oligonucleotide products.

[0083] Ideally, to compare activity between ligation reactions and account for natural variations in peak intensity between injections, the percent conversion to product would be calculated for each sample analyzed using the following formula:

number

number

[0084] "Improved enzyme properties" and the like refer to enzyme properties that are better or more desirable for a particular purpose compared to a reference, such as a ligase that is not linked to a PPK. As used herein, "improved enzyme properties" and the like typically refer to properties of the ligase domain of the fusion polypeptides described herein. Expected improved enzyme properties include, but are not limited to, enzyme activity (which may be expressed as a percentage of substrate conversion or any unit described herein), thermostability, solvent stability, pH activity profile, cofactor requirements, and tolerance to inhibitors (e.g., reaction component, substrate, or product inhibition). The fusion polypeptides described herein also exhibit improved ligase activity when immobilized compared to immobilized ligase polypeptides that are not linked to a PPK domain. The fusion polypeptides described herein may also exhibit improved soluble yields from host cells, resulting in increased enzyme activity when crude extracts (e.g., cell-free lyophilized extracts or cell lysates) are used.

[0085] An "isolated polypeptide" (e.g., an "isolated fusion polypeptide") refers to a polypeptide that has been substantially separated from other materials with which it is naturally associated, such as proteins, lipids, and polynucleotides. The term includes polypeptides that have been removed or purified from their naturally occurring environment or expression system (e.g., a host cell or in vitro synthesis). The polypeptide (e.g., a fusion polypeptide) may be present intracellularly, in cell culture medium, or prepared in various forms, such as a lysate or isolated preparation. Thus, in some embodiments, the fusion polypeptide is an isolated fusion polypeptide.

[0086] As used herein, a "crude extract" is a solution produced by lysing cells expressing a polypeptide of interest and removing cellular debris, for example, by centrifugation. The crude extract described herein can be a cell-free lyophilized extract or a cell lysate.

[0087] "Naturally-occurring" or "wild-type" refers to a form found in nature. For example, a naturally-occurring or wild-type polypeptide or polynucleotide sequence is a sequence present in an organism that can be isolated from a natural source and has not been intentionally modified by manual procedures.

[0088] As used herein, "polyphosphate kinase" or "PPK" refers to a wild-type or modified enzyme having polyphosphate kinase activity, i.e., an enzyme that catalyzes the formation of ATP from AMP and polyphosphate. PPKs are also referred to herein as phosphotransferases or polyphosphate nucleotide phosphotransferases.

[0089] The terms "polynucleotide," "nucleic acid molecule," and "nucleic acid" are used interchangeably herein.

[0090] The terms "protein," "polypeptide," and "peptide" are used interchangeably herein to refer to a polymer of at least two amino acids covalently linked by an amide bond, regardless of length or post-translational modification (e.g., glycosylation, phosphorylation, lipidation, myristoylation, ubiquitination, etc.). This definition includes D- and L-amino acids, as well as mixtures of D- and L-amino acids. Preferably, the amino acids have the L-configuration.

[0091] "Recombinant," "modified," or "non-naturally occurring," when used with reference to, for example, a cell, nucleic acid, or polypeptide, refers to material that does not otherwise occur in nature, or material that corresponds to a naturally occurring form of the material that is identical to but has been modified in a way that is produced or derived from synthetic material and / or by manipulation using recombinant techniques.

[0092] As used herein, oligonucleotides or oligonucleotide fragments containing chemical modifications refer to oligonucleotides and oligonucleotide fragments having modified nucleotides, oligonucleotides and oligonucleotide fragments having modified backbones, and / or oligonucleotides and oligonucleotide fragments conjugated to a ligand.

[0093] As used herein, an "unmodified nucleotide" is a nucleotide having a deoxyribose or ribose sugar and a nucleobase selected from adenine, cytosine, guanine, thymine, and uracil. As used herein, a "modified nucleotide" refers to a nucleotide containing a modified sugar and / or modified base. A modified sugar can be a modified deoxyribose sugar or a modified ribose sugar substituted at one or more positions with a non-hydrogen substituent. A modified base refers to any base other than adenine, cytosine, guanine, thymine, and uracil. Exemplary modified sugars and modified bases are described herein.

[0094] As used herein, an "unmodified backbone" consists of 3' to 5' phosphodiester linkages. As used herein, a "modified backbone" can include any non-natural internucleoside linkages, such as phosphorothioate linkages (e.g., chiral phosphorothioate linkages) and phosphorodithioate linkages. Exemplary backbone modifications are described herein.

[0095] The abbreviations used for the genetically encoded amino acids are conventional and are as follows:

[0096] [Table 1]

[0097] "-" is used for amino acid deletions, and " is used for stop codons. * When a three-letter abbreviation is used, the amino acid is referred to as the α-carbon (C α) can be in either the L- or D-configuration. For example, "Ala" refers to alanine without specifying the configuration at the α-carbon, while "D-Ala" and "L-Ala" refer to D-alanine and L-alanine, respectively.

[0098] When single-letter abbreviations are used, an uppercase letter indicates an amino acid in the L-configuration about the α-carbon, and a lowercase letter indicates an amino acid in the D-configuration about the α-carbon. For example, "A" indicates L-alanine and "a" indicates D-alanine. When polypeptide sequences are presented as a string of one-letter or three-letter abbreviations (or mixtures thereof), the sequences are presented in amino (N) to carboxy (C) orientation, according to common convention.

[0099] The abbreviations used for genetically encoded nucleotides are conventional and are as follows: adenosine (A); guanosine (G); cytidine (C); thymidine (T); and uridine (U). Unless specifically indicated, abbreviated nucleotides may be ribonucleotides or 2'-deoxyribonucleotides. Nucleotides may be identified individually or collectively as ribonucleotides or 2'-deoxyribonucleotides. When nucleic acid sequences are presented as strings of single-letter abbreviations, the sequences are presented in the 5' to 3' direction, according to common convention, and phosphodiester bonds are not shown.

[0100] Those skilled in the art are well aware that guanine, cytosine, adenine and uracil can be substituted by other moieties without substantially changing the base pairing properties of the oligonucleotide that contains the nucleotide with such a substituted moiety.For example, but not limited to, the nucleotide that contains inosine as its base can base pair with the nucleotide that contains adenine, cytosine or uracil.Therefore, the nucleotide that contains uracil, guanine or adenine can be substituted by the nucleotide that contains inosine, for example, in the nucleotide sequence of the oligonucleotide that is characterized in the present disclosure.In another example, the adenine and cytosine anywhere in oligonucleotide can be substituted by guanine and uracil, respectively, to form a wobble base pair with target mRNA.

[0101] Methods for determining the percentage of sequence identity are known in the art. For example, when assessing sequence identity, a sequence having a specified number of consecutive nucleotides or amino acids can be aligned with a nucleic acid or peptide sequence (having the same number of consecutive nucleotides or amino acids) from a corresponding portion of the nucleic acid or peptide sequence disclosed herein. The percentage of sequence identity can be calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue exists or where the nucleic acid base or amino acid residue is aligned with a gap to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Those skilled in the art will understand that there are many established algorithms available for aligning two sequences. Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math. 2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443, by the similarity search method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computer implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Package), or by visual inspection (see generally, Current Protocols in Molecular Biology, F.M. Ausubel et al. eds., Current Protocols, a Joint Venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)).Examples of algorithms suitable for determining percent sequence identity and percent sequence similarity are the BLAST and BLAST 2.0 algorithms described in Altschul et al., 1990, J. Mol. Biol. 215:403-410 and Altschul et al., 1977, Nucleic Acids Res. 3389-3402, respectively. Software for performing BLAST analyses is publicly available from the website of the National Center for Biotechnology Information. This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that, when aligned with words of the same length in a database sequence, match or satisfy a certain positive threshold score T. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. Word hits are extended in either direction along each sequence for as far as the cumulative alignment score can be increased. For nucleotide sequences, cumulative scores are calculated using the parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatching residues; always < 0). For amino acid sequences, a scoring matrix is ​​used to calculate the cumulative score. Extension of the word hits in each direction is halted when the cumulative alignment score falls by an amount X from its achieved maximum value; when the cumulative score falls to 0 or below due to the accumulation of alignment of one or more negative-scoring residues; or when either end of the sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses defaults of a word length (W) of 11, an expectation (E) of 10, M=5, and N=-4, and performs a comparison of both strands.The BLASTP program for amino acid sequences uses as defaults a word length of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc Natl Acad Sci USA 89:10915). Exemplary sequence alignments and determination of percent sequence identity can use the BESTFIT or GAP programs using the default parameters provided in the GCG Wisconsin software package (Accelrys, Madison WI).

[0102] It will be understood that an ATP-dependent nucleic acid ligase or ATP-dependent nucleic acid ligase domain (e.g., a dsRNA ligase domain) has ATP-dependent nucleic acid ligase activity regardless of the percent sequence identity to a reference sequence. Similarly, it will be understood that a PPK or PPK domain has PPK activity regardless of the percent sequence identity to a reference sequence.

[0103] "Suitable reaction conditions" refer to conditions in a reaction system (e.g., enzyme load, substrate load, temperature, pH, etc.) under which a substrate is converted into a desired product. Suitable reaction conditions can be easily identified by one skilled in the art. Examples of "suitable reaction conditions" are provided in the present disclosure and illustrated by the examples.

[0104] Ligation Reaction Enzymatic ligation of short oligonucleotide fragments provides a sustainable and economical alternative to solid-phase chemical synthesis of full-length (e.g., therapeutic) oligonucleotides. Herein, "nucleic acid ligase" refers to an ATP-dependent nucleic acid ligase (rather than, e.g., an NAD-dependent nucleic acid ligase). The main drawback of ligation reactions catalyzed by ATP-dependent nucleic acid ligases is the requirement for a stoichiometric amount of the expensive cofactor ATP. ATP-dependent nucleic acid ligases are ATP-dependent enzymes, and their catalytic mechanisms have been well characterized. First, the active site lysine attacks the α-phosphate of ATP, forming a lysine-AMP intermediate and releasing pyrophosphate. Next, AMP is transferred from the active site lysine to the 5' phosphate of the 3'-ligated fragment, forming an adenylated oligonucleotide intermediate. Finally, the 3'-OH of the 5'-ligated fragment attacks the 5'-phosphate of the adenylated intermediate, releasing AMP (Figure 1) (Nandakumar, J., Shuman, S. & Lima, CDCell 127, 71-84 (2006)). In this way, one molecule of ATP is converted to AMP per ligation reaction. In practice, excess ATP is required to achieve complete ligation.

[0105] As demonstrated herein, the requirement for high concentrations of ATP in ligation reactions can be overcome by incorporating an ATP regeneration system into the reaction. The ATP regeneration system described herein includes a PPK and polyphosphate. The PPK generates ATP from AMP using polyphosphate as a phosphate donor. ATP converted to AMP during the ligation reaction can be regenerated by the PPK to ATP, which can be used as a cofactor in subsequent ligation reactions. This cycle of ATP eliminates the need for high ATP concentrations in the starting reaction. Instead, the reaction can be performed using substoichiometric concentrations of ATP and / or using the cheaper alternative, AMP.

[0106] In some embodiments, oligonucleotide products are obtained with a percent conversion of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%. Conversion may be expressed in arbitrary units (AU), calculated as described herein.

[0107] Oligonucleotides and Oligonucleotide Fragments The method of the present invention produces an oligonucleotide by ligating two or more oligonucleotide fragments. The produced oligonucleotide (also referred to herein as "oligonucleotide product") is typically a nucleic acid containing up to 100 nucleotides. The oligonucleotide fragment may be referred to herein as the "substrate" of the ligation reaction.

[0108] In some embodiments, the oligonucleotide is a therapeutic oligonucleotide. In some embodiments, the therapeutic oligonucleotide is a small interfering RNA (siRNA) or an antisense oligonucleotide (ASO). In some embodiments, the oligonucleotide is an aptamer.

[0109] In some embodiments, the oligonucleotide is a double-stranded oligonucleotide. In some embodiments, the oligonucleotide comprises an overhang. In some embodiments, the oligonucleotide comprises a 3' overhang. In some embodiments, the oligonucleotide comprises a 5' overhang. In some embodiments, the overhang comprises 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides. In some embodiments, the oligonucleotide comprises a blunt end. In some embodiments, the oligonucleotide comprises two blunt ends.

[0110] In some embodiments, the oligonucleotide is a single-stranded oligonucleotide.

[0111] In some embodiments, the oligonucleotide is up to 20 nucleotides in length. In some embodiments, the oligonucleotide is up to 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides in length. In some embodiments, the oligonucleotide is up to 60 nucleotides in length.

[0112] In some embodiments, the oligonucleotide is at least 20 nucleotides in length, hi some embodiments, the oligonucleotide is at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or 100 nucleotides in length.

[0113] In some embodiments, the oligonucleotide is 10 to 100 nucleotides in length. In some embodiments, the oligonucleotide is 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 25, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, or 15 to 25 nucleotides in length. In some embodiments, the oligonucleotide is 15 to 25 nucleotides in length.

[0114] In some embodiments, the two or more oligonucleotide fragments comprise single-stranded oligonucleotide fragments.

[0115] In some embodiments, the two or more oligonucleotide fragments comprise double-stranded oligonucleotide fragments. In some embodiments, one or more of the oligonucleotide fragments comprise one or more mismatches. In some embodiments, one or more of the oligonucleotide fragments comprise an overhang. In some embodiments, one or more of the oligonucleotide fragments comprise a 3' overhang. In some embodiments, one or more of the oligonucleotide fragments comprise a 5' overhang. In some embodiments, one or more of the oligonucleotide fragments comprise a 3' overhang and a 5' overhang. In some embodiments, the overhang comprises 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides.

[0116] In some embodiments, the two or more oligonucleotide fragments comprise a first oligonucleotide fragment having an overhang complementary to an overhang of a second oligonucleotide fragment, in some embodiments, the two or more oligonucleotide fragments comprise a first oligonucleotide fragment having a 3' overhang and a 5' overhang, wherein the 3' overhang is complementary to the 5' overhang of the second oligonucleotide fragment and the 5' overhang is complementary to the 3' overhang of a third oligonucleotide.

[0117] In some embodiments, one or more of the oligonucleotide fragments comprise a blunt end. In some embodiments, one or more of the oligonucleotide fragments comprise a 3' overhang and a 5' blunt end. In some embodiments, one or more of the oligonucleotide fragments comprise a 5' overhang and a 3' blunt end.

[0118] In some embodiments, the two or more oligonucleotide fragments comprise two or more RNA oligonucleotide fragments. In some embodiments, the two or more RNA oligonucleotide fragments comprise double-stranded RNA (dsRNA) oligonucleotide fragments. In some embodiments, the two or more RNA oligonucleotide fragments comprise single-stranded RNA (ssRNA) oligonucleotide fragments.

[0119] In some embodiments, the two or more oligonucleotide fragments comprise two or more DNA oligonucleotide fragments. In some embodiments, the two or more DNA oligonucleotide fragments comprise double-stranded DNA (dsDNA) oligonucleotide fragments. In some embodiments, the two or more DNA oligonucleotide fragments comprise single-stranded DNA (ssDNA) oligonucleotide fragments.

[0120] In some embodiments, the two or more oligonucleotide fragments comprise DNA oligonucleotide fragments and RNA oligonucleotide fragments, hi some embodiments, the two or more oligonucleotide fragments comprise dsDNA oligonucleotide fragments and dsRNA oligonucleotide fragments.

[0121] In some embodiments, one or more of the oligonucleotide fragments comprise one or two strands that are RNA or a mixture of RNA, DNA, LNA, morpholino, UNA (unlocked nucleic acid), TNA (threose nucleic acid), GNA (glycol nucleic acid), and / or FANA (fluoro-arabino nucleic acid), modified RNA, etc. As a non-limiting example, one or both strands can be RNA, except that, for example, one or more nucleotides have been replaced with DNA, LNA, morpholino, UNA, TNA, GNA, and / or FANA, and / or modified RNA (e.g., any modified RNA disclosed herein or known in the art, e.g., 2'-modified RNA, including but not limited to, 2'-F, 2'-OMe, 2'-O-MOE RNA, etc.).

[0122] In some embodiments, two or more oligonucleotide fragments are the same length. In some embodiments, two or more oligonucleotide fragments are different lengths. In some embodiments, each of the two or more oligonucleotide fragments is 3 to 20 nucleotides in length. In some embodiments, each of the two or more oligonucleotide fragments is 4 to 16 nucleotides in length. In some embodiments, each of the two or more oligonucleotide fragments is 4 to 16, 4 to 15, 5 to 15, 6 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 5 to 9, or 6 to 9 nucleotides in length.

[0123] In some embodiments, each of the two or more oligonucleotide fragments is at least 3 nucleotides in length, hi some embodiments, each of the two or more oligonucleotide fragments is at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 nucleotides in length.

[0124] In some embodiments, the two or more oligonucleotide fragments comprise 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more oligonucleotide fragments.

[0125] In some embodiments, one or more ligation reactions are required to generate the oligonucleotide product, hi some embodiments, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more ligation reactions are required to generate the oligonucleotide product.

[0126] In some embodiments, one or more of the oligonucleotide fragments comprises a chemical modification. In some embodiments, one or more of the oligonucleotide fragments comprises at least one backbone modification. In some embodiments, one or more of the oligonucleotide fragments comprises at least one nucleotide modification. In some embodiments, one or more of the oligonucleotide fragments comprises at least one sugar modification (e.g., at the 2' or 4' position). In some embodiments, one or more of the oligonucleotide fragments comprises: (i) at least one backbone modification; (ii) at least one nucleotide modification; and / or (iii) at least one sugar modification. In some embodiments, the oligonucleotide comprises a chemical modification. In some embodiments, the oligonucleotide comprises at least one backbone modification. In some embodiments, the oligonucleotide comprises at least one nucleotide modification. In some embodiments, the oligonucleotide comprises at least one sugar modification (e.g., at the 2' or 4' position). In some embodiments, the oligonucleotide comprises: (i) at least one backbone modification; (ii) at least one nucleotide modification; and / or (iii) at least one sugar modification.

[0127] In some embodiments, one or more of the oligonucleotide fragments comprises a modification selected from the group consisting of 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'- 2'-O-DMAEOE), 2'-ON-methylacetamide (2'-O-NMA), locked nucleic acid (LNA), glycol nucleic acid (GNA), phosphoramidate (e.g., mesyl phosphoramidate), 2',3'-seconucleotide mimic, 2'-F-arabinonucleotide, abasic nucleotide, 2'-amino modified nucleotide, 2'-alkyl modified nucleotide, morpholino nucleotide, vinyl phosphonate (e.g., 5'-vinyl phosphonate), and cyclopropyl phosphonate deoxyribonucleotide. In some embodiments, one or more of the oligonucleotide fragments comprises a 2'-modification selected from the group consisting of 2'-OMe, 2'-F, and 2'-deoxy.

[0128] In some embodiments, the oligonucleotide comprises a modification selected from the group consisting of 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-D 2'-O-methylacetamide (2'-O-NMA), locked nucleic acids (LNA), glycol nucleic acids (GNAs), phosphoramidates (e.g., mesyl phosphoramidate), 2',3'-seconucleotide mimics, 2'-F-arabinonucleotides, abasic nucleotides, 2'-amino-modified nucleotides, 2'-alkyl-modified nucleotides, morpholino nucleotides, vinyl phosphonates (e.g., 5'-vinyl phosphonate), and cyclopropyl phosphonate deoxyribonucleotides. In some embodiments, the oligonucleotide comprises a 2'-modification selected from the group consisting of: 2'-OMe, 2'-F, and 2'-deoxy.

[0129] In some embodiments, one or more of the oligonucleotide fragments comprises at least one phosphorothioate or methylphosphonate internucleotide linkage. In some embodiments, the oligonucleotide comprises at least one phosphorothioate or methylphosphonate internucleotide linkage. In some embodiments, the oligonucleotide comprises at least one chiral phosphorothioate linkage.

[0130] In some embodiments, one or more of the oligonucleotide fragments are conjugated to at least one ligand, which can be conjugated in any configuration to the sense strand, the antisense strand, or both strands, e.g., at the 3' end, the 5' end, non-terminal, or in combination.

[0131] In some embodiments, the oligonucleotide is conjugated to at least one ligand, which can be conjugated in any configuration to the sense strand, the antisense strand, or both strands, e.g., at the 3' end, the 5' end, non-terminal, or in combination.

[0132] In some embodiments, the ligand comprises one or more N-acetylgalactosamine (GalNAc) derivatives. GalNAc is an amino sugar derivative of galactose and can be used as a targeting ligand in oligonucleotides intended to target the liver, binding to the asialoglycoprotein receptor on hepatocytes. In some embodiments, the ligand comprises one or more GalNAc derivatives conjugated via a bivalent or trivalent branched carrier. In some embodiments, the ligand is a peptide or peptidomimetic.

[0133] In some embodiments, the ligand is conjugated to the sense strand. In some embodiments, the ligand is conjugated to the 3' end of the sense strand. In some embodiments, the ligand is conjugated to the 5' end of the sense strand. In some embodiments, the ligand is conjugated to a non-terminal end of the sense strand.

[0134] In some embodiments, the ligand is conjugated to the antisense strand. In some embodiments, the ligand is conjugated to the 3' end of the antisense strand. In some embodiments, the ligand is conjugated to a non-terminal end of the antisense strand.

[0135] In some embodiments, one or more of the oligonucleotide fragments comprises at least one 2'-modified nucleotide selected from the group consisting of 2'-OMe, 2'-F, 2'-deoxy, 2'-deoxy-2'-fluoro, and 2'-O-MOE. In some embodiments, one or more of the oligonucleotide fragments is a dsRNA in which the sense strand is conjugated to one or more GalNAc ligands.

[0136] In some embodiments, the oligonucleotide comprises at least one 2'-modified nucleotide selected from the group consisting of 2'-OMe, 2'-F, 2'-deoxy, 2'-deoxy-2'-fluoro, and 2'-O-MOE. In some embodiments, the oligonucleotide is a dsRNA, wherein the sense strand is conjugated to one or more GalNAc ligands.

[0137] In some embodiments, the oligonucleotide is an RNAi agent comprising at least one 2'-modified nucleotide selected from the group consisting of 2'-OMe, 2'-F, 2'-deoxy, 2'-deoxy-2'-fluoro, and 2'-O-MOE. In some embodiments, the oligonucleotide is an RNAi agent in which the sense strand is conjugated to one or more GalNAc ligands.

[0138] In some embodiments, the method is performed using an oligonucleotide fragment concentration of at least 1 mM, at least 2 mM, at least 3 mM, at least 4 mM, at least 5 mM, at least 6 mM, at least 7 mM, at least 8 mM, at least 9 mM, or at least 10 mM. In some embodiments, the method is performed using at least 1 mM, at least 2 mM, at least 3 mM, at least 4 mM, at least 5 mM, at least 6 mM, at least 7 mM, at least 8 mM, at least 9 mM, or at least 10 mM of each oligonucleotide fragment. In some embodiments, the method is performed using equimolar amounts of each of two or more oligonucleotide fragments.

[0139] In some embodiments, the method produces at least 15 g of oligonucleotide product per liter of reaction mixture, hi some embodiments, the method produces at least 16 g, at least 17 g, at least 18 g, at least 19 g, at least 20 g, at least 30 g, at least 40 g, at least 50 g, at least 60 g, at least 70 g, at least 80 g, at least 90, or at least 100 g of oligonucleotide product per liter of reaction mixture.

[0140] In some embodiments, the method further comprises purifying the oligonucleotide product from the reaction mixture. In some embodiments, the oligonucleotide product is at least 80% pure, optionally at least 85% pure, at least 90% pure, at least 95% pure, optionally at least 98% pure, optionally at least 99% pure, optionally at least 99.5% pure, or optionally at least 99.9% pure. Pure oligonucleotide products typically do not contain oligonucleotide fragments, intermediate ligation products, or by-products resulting from non-specific ligation. Oligonucleotide products can be purified or isolated using any method known in the art, for example, by separating the oligonucleotides based on their size using, for example, gel extraction, ultrafiltration, and / or chromatography using a cellulose-based matrix.

[0141] The present disclosure also provides oligonucleotides produced by the methods described herein.

[0142] ATP-dependent nucleic acid ligase ATP-dependent nucleic acid ligases are a family of enzymes that catalyze the ligation of oligonucleotide fragments using ATP as a cofactor. The catalytic mechanism of ATP-dependent nucleic acid ligases has been well characterized as described above and outlined in Figure 1.

[0143] In some embodiments, the ATP-dependent nucleic acid ligase is an RNA ligase. In some embodiments, the RNA ligase is a dsRNA ligase. In some embodiments, the RNA ligase is a member of the RNA ligase 2 family. In some embodiments, the RNA ligase is bacteriophage RB69 RNA ligase 2 (UniProt ID: Q7Y4V8). In some embodiments, the RNA ligase is bacteriophage T4 RNA ligase 2 (UniProt ID: P32277).

[0144] In some embodiments, the RNA ligase comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1, the amino acid sequence of bacteriophage RB69 RNA ligase 2 (UniProt ID: Q7Y4V8). In some embodiments, the RNA ligase comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 1.

[0145] In some embodiments, the RNA ligase comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO:2, the amino acid sequence of optimized bacteriophage RB69 RNA ligase 2. In some embodiments, the RNA ligase comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:2.

[0146] In some embodiments, the RNA ligase comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 88, the amino acid sequence of optimized bacteriophage RB69 RNA ligase 2. In some embodiments, the RNA ligase comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 88.

[0147] In some embodiments, the RNA ligase comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 3, the amino acid sequence of bacteriophage T4 RNA ligase 2 (UniProt ID: P32277). In some embodiments, the RNA ligase comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 3.

[0148] In some embodiments, the ATP-dependent nucleic acid ligase is a DNA ligase, hi some embodiments, the DNA ligase is T4 DNA ligase or a variant thereof.

[0149] In some embodiments, the DNA ligase comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 4, the amino acid sequence of bacteriophage T4 DNA ligase (UniProt ID: P00970). In some embodiments, the RNA ligase comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 4.

[0150] In some embodiments, the ATP-dependent nucleic acid ligase is used in the form of a whole cell, a crude extract, an isolated polypeptide, or a purified polypeptide. In some embodiments, the ATP-dependent nucleic acid ligase polypeptide is used in an immobilized form, such as immobilized on a solid support material, as described herein.

[0151] Polyphosphate kinase "Polyphosphate kinases" or "PPKs" are a family of enzymes that catalyze the formation of ATP from AMP and polyphosphate.

[0152] In some embodiments, the PPK is PPK12. In some embodiments, the PPK comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 5, the amino acid sequence of PPK12. In some embodiments, the PPK comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 5.

[0153] In some embodiments, the PPK comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 6, the amino acid sequence of optimized PPK12. In some embodiments, the PPK comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 6.

[0154] In some embodiments, the PPK is Acinetobacter johnsonii polyphosphate:AMP phosphotransferase (AjPAP) (UniProt ID: Q83XD3). In some embodiments, the PPK comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO:7, the amino acid sequence of AjPAP. In some embodiments, the PPK comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:7.

[0155] In some embodiments, the PPK is used in the form of a whole cell, a crude extract, an isolated polypeptide, or a purified polypeptide. In some embodiments, the PPK polypeptide is used in an immobilized form, such as immobilized on a solid support material, as described herein.

[0156] Fusion Polypeptides The methods described herein may be carried out using an ATP-dependent nucleic acid ligase and a PPK, where the ATP-dependent nucleic acid ligase and the PPK are provided as separate polypeptides.

[0157] The methods described herein may also be performed using a fusion polypeptide comprising an ATP-dependent nucleic acid ligase linked to a PPK. Herein, the ATP-dependent nucleic acid ligase portion of the fusion polypeptide is referred to as the ATP-dependent nucleic acid ligase "domain," and the PPK portion of the fusion polypeptide is referred to as the PPK "domain."

[0158] An important consideration for the economic viability and sustainability of biocatalytic processes is the production costs associated with the enzymes used therein. To overcome the potential increase in ligation reaction costs associated with the inclusion of an ATP regeneration system, we developed a bifunctional fusion polypeptide containing a kinase domain and a ligase domain. The fusion polypeptide exhibits both kinase and ligase activity and can be produced via a single reaction (e.g., by recombinant expression), saving time, effort, and expense compared to producing separate kinase and ligase enzymes.

[0159] Advantageously, it was unexpectedly found that the fusion polypeptide exhibited higher ligase activity compared to the unligated enzyme. Without wishing to be bound by theory, the inventors believe that this improvement in ligase activity may be due to the generation of a local supply of ATP by the kinase and / or improved stability of the enzyme compared to the unligated enzyme.

[0160] The present disclosure also provides a fusion polypeptide comprising: (a) a PPK domain; and (b) an ATP-dependent nucleic acid ligase domain. The fusion polypeptides described herein can be used in the ATP-dependent RNA ligation reactions described herein.

[0161] In some embodiments, the ATP-dependent nucleic acid ligase domain is an RNA ligase domain. In some embodiments, the RNA ligase domain is a dsRNA ligase domain. In some embodiments, the dsRNA ligase domain is a member of the RNA ligase 2 family. In some embodiments, the dsRNA ligase domain comprises bacteriophage RB69 RNA ligase 2 (UniProt ID: Q7Y4V8). In some embodiments, the dsRNA ligase domain comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1, the amino acid sequence of bacteriophage RB69 RNA ligase 2 (UniProt ID: Q7Y4V8). In some embodiments, the dsRNA ligase domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:1.

[0162] In some embodiments, the dsRNA ligase domain comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO:2, the amino acid sequence of optimized bacteriophage RB69 RNA ligase 2. In some embodiments, the dsRNA ligase domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:2.

[0163] In some embodiments, the dsRNA ligase domain comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 88, the amino acid sequence of optimized bacteriophage RB69 RNA ligase 2. In some embodiments, the dsRNA ligase domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 88.

[0164] In some embodiments, the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 3, the amino acid sequence of bacteriophage T4 RNA ligase 2 (UniProt ID: P32277). In some embodiments, the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 3.

[0165] In some embodiments, the ATP-dependent nucleic acid ligase domain is a DNA ligase domain. In some embodiments, the DNA ligase domain is a T4 DNA ligase domain. In some embodiments, the T4 DNA ligase domain comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO:4, the amino acid sequence of bacteriophage T4 RNA ligase 2 (UniProt ID: P32277). In some embodiments, the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:4.

[0166] In some embodiments, the PPK domain is PPK12. In some embodiments, the PPK domain comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 5, the amino acid sequence of PPK12. In some embodiments, the PPK domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 5.

[0167] In some embodiments, the PPK domain comprises an amino acid sequence having at least 70% sequence identity to the amino acid sequence of optimized PPK12, SEQ ID NO: 6. In some embodiments, the PPK domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 6.

[0168] In some embodiments, the PPK domain is AjPAP (UniProt ID: Q83XD3). In some embodiments, the PPK domain comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 7, the amino acid sequence of AjPAP. In some embodiments, the PPK domain comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 7.

[0169] In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 70% sequence identity to a sequence selected from SEQ ID NOs: 8-18. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a sequence selected from SEQ ID NOs: 8-18.

[0170] In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 70% sequence identity to a sequence selected from SEQ ID NO: 90, 92, 94, 96, or 98. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a sequence selected from SEQ ID NO: 90, 92, 94, 96, or 98.

[0171] In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 70% sequence identity to a sequence selected from SEQ ID NOs: 8-18, 90, 92, 94, 96, or 98. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a sequence selected from SEQ ID NOs: 8-18, 90, 92, 94, 96, or 98.

[0172] In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 17. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 17.

[0173] In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 18. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 18.

[0174] In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 98. In some embodiments, the fusion polypeptide comprises an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 98.

[0175] In some embodiments, the methods described herein are carried out using a fusion polypeptide described herein.

[0176] In some embodiments, the compositions described herein comprise a fusion polypeptide described herein. In some embodiments, the kits described herein comprise a fusion polypeptide described herein.

[0177] In some embodiments, the fusion polypeptide is used in the form of a whole cell, a crude extract, an isolated fusion polypeptide, or a purified fusion polypeptide. In some embodiments, the fusion polypeptide is used in an immobilized form, such as immobilized on a solid support, as described herein.

[0178] In some embodiments, the rate of nucleic acid ligation exceeds the rate of nucleic acid ligation of a control, wherein the control comprises: (a) a first protein comprising a PPK domain of a fusion polypeptide; and (b) a second protein comprising an ATP-dependent nucleic acid ligase domain of a fusion polypeptide, wherein the first and second proteins are unlinked. In some embodiments, the rate of nucleic acid ligation exceeds the rate of nucleic acid ligation of the control by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50%.

[0179] In some embodiments, the rate of RNA ligation exceeds the rate of RNA ligation of a control, wherein the control comprises: (a) a first protein comprising a PPK domain of a fusion polypeptide; and (b) a second protein comprising a dsRNA ligase domain of a fusion polypeptide, wherein the first and second proteins are unlinked. In some embodiments, the rate of RNA ligation exceeds the rate of RNA ligation of the control by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50%.

[0180] In some embodiments, the rate of nucleic acid ligation is calculated as a percent conversion over a defined incubation time (e.g., 12, 18, or 24 hours) or using arbitrary units (AU) (calculated as described herein). Comparisons between RNA ligation rates are typically made between ligation reactions performed using similar or identical molar amounts of ATP-dependent nucleic acid ligase or ATP-dependent nucleic acid ligase domain.

[0181] Linkers and Purification Tags Purification tags are typically added to polypeptides to enable purification from crude biological sources using affinity techniques. Purification tags include, but are not limited to, polyhistidine, chitin-binding protein (CBP), maltose-binding protein (MBP), Strep-tag, and glutathione-S-transferase (GST). Polyhistidine tags bind to matrices with immobilized metal ions, thereby allowing the immobilization of polypeptides bound to them via metal affinity.

[0182] Linkers are typically short peptide sequences present between protein domains. Linkers can be configured to allow adjacent protein domains to move relative to one another, or they can be rigid to prevent undesired interactions between protein domains.

[0183] In some embodiments, the ATP-dependent nucleic acid ligase and the PPK are linked via a linker, e.g., a peptide linker. In some embodiments, the PPK is at the N-terminus of the linker and the ATP-dependent nucleic acid ligase is at the C-terminus of the linker. In some embodiments, the ATP-dependent nucleic acid ligase is at the N-terminus of the linker and the PPK is at the C-terminus of the linker.

[0184] In some embodiments, the linker is a polypeptide linker. In some embodiments, the linker comprises at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 6 amino acids, at least 7 amino acids, at least 8 amino acids, at least 9 amino acids, or at least 10 amino acids. In some embodiments, the linker comprises at least 6 amino acids.

[0185] In some embodiments, the linker is linear. In some embodiments, the linker comprises at least three amino acids, each amino acid selected from glycine, serine, or alanine. In some embodiments, the linker comprises at least one amino acid selected from glutamic acid, aspartic acid, lysine, or arginine.

[0186] In some embodiments, the linker comprises a poly-histidine tag. In some embodiments, the linker comprises an amino acid recognition sequence for tobacco etch virus (TEV) protease.

[0187] In some embodiments, the linker comprises an amino acid sequence selected from: (a) HHHHHH (SEQ ID NO: 19), optionally HHHHHHHHHHH (SEQ ID NO: 20); (b) ENLYFQS (SEQ ID NO: 21); (c) ENLYFQG (SEQ ID NO: 22); (d) SSGSSG (SEQ ID NO: 23); (e) GSAGSAAGSGEF (SEQ ID NO: 24); and / or (f) GSSGSGSSSGGSSSSGSS (SEQ ID NO: 25).

[0188] In some embodiments, the fusion polypeptide comprises a purification tag. In some embodiments, the linker comprises a purification tag. In some embodiments, the purification tag is located at the N-terminus or C-terminus of the fusion polypeptide. In some embodiments, the purification tag comprises a poly-histidine tag. In some embodiments, the purification tag comprises MHHHHHHENLYFQS (SEQ ID NO: 26). In some embodiments, the purification tag comprises GQTGHHHHHH (SEQ ID NO: 27). In some embodiments, the purification tag comprises a Myc-tag (EQKLISEEDL (SEQ ID NO: 28)). In some embodiments, the purification tag comprises a FLAG-tag (DYKDDDDK (SEQ ID NO: 29)).

[0189] immobilization In some embodiments, the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide is immobilized. In some embodiments, the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide is immobilized using affinity immobilization. In some embodiments, the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide is immobilized using metal affinity immobilization, for example, by contacting a His-tagged PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide with an immobilized metal such as nickel, zinc, cobalt, or copper.

[0190] In some embodiments, the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide is immobilized on a solid material by chemical bonding or physical adsorption. The terms "solid support," "solid material," and "solid support material" are used interchangeably herein. Immobilization of a polypeptide by physical absorption typically involves physically adsorbing or binding the polypeptide to a solid support material. Adsorption can occur through weak nonspecific forces such as van der Waals, hydrophobic interactions, and hydrogen bonds. Physical adsorption can be achieved by immersing the solid support material in a solution of the polypeptide and incubating it to allow time for physical adsorption to occur. Immobilization of a polypeptide by chemical bonding typically involves binding the polypeptide to the solid support material by a covalent bond.

[0191] When the ATP-dependent nucleic acid ligase is immobilized, the inventors believe that (due to the size of the oligonucleotide fragment substrate) the catalytic activity of the ATP-dependent ligase is optimal when a spacer is present between the ATP-dependent ligase and the immobilization moiety. When the fusion polypeptide is immobilized, the inventors believe that the catalytic activity of the ATP-dependent ligase domain is optimal when the PPK domain is linked to the immobilization moiety (optionally via a spacer). In some embodiments, the PPK is immobilized (optionally via a spacer) and the ATP-dependent nucleic acid ligase is not immobilized. In some embodiments, the spacer is a polypeptide (e.g., a polypeptide comprising 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 75 or more, or 100 or more amino acids).

[0192] In some embodiments, the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide is immobilized on a solid support such as a membrane, resin, solid carrier, or other solid phase material. The solid support can be composed of organic polymers such as polystyrene, polyethylene, polypropylene, polyfluoroethylene, polyethyleneoxy, polymethacrylate, and polyacrylamide, as well as copolymers and grafts thereof. The solid support can also be inorganic, such as glass, silica, controlled pore glass (CPG), reverse-phase silica, or metal, such as gold or platinum. The solid support can be in the form of beads, spheres, particles, granules, gels, membranes, or surfaces. Surfaces can be planar, substantially planar, or non-planar. The solid support can be porous or non-porous and can have swelling or non-swelling properties. The solid support can be configured in the form of a well, depression, or other container, vessel, feature, or location. Useful solid supports for immobilizing PPK, ATP-dependent nucleic acid ligase and / or fusion polypeptides for carrying out reactions include, but are not limited to, beads or resins such as polymethacrylate, e.g., epoxy-functionalized polymethacrylate, aminoepoxy-functionalized polymethacrylate, polymethacrylate, styrene / DVB copolymer, or octadecyl-functionalized polymethacrylate.

[0193] Exemplary solid supports include chitosan beads, Eupergit C, IB-150, IB-350, IB-C435, IB-A369, IB-A161, IB-A171, IBS500, IB-S861, SEPABEADS (Mitsubishi), such as Sepabeads EC-EP, Sepabeads EC-HFA, Sepabeads EC-HG, Sepabeads EC-BU, Sepabeads EC-OD, Sepabeads EC-CM, Sepabeads EC-IDA, Sepabeads EC-EA, Sepabeads EC-HA, Sepabeads EC-QA, Sepabeads EXE, Sepabeads EXA, Dilbeads-TA, Amberzyme Oxirane, Amberlite XAD-7HP, Amberlite FPA98Cl, Amberlite IRA958Cl, Amberlite IRA67, Amberlite FPA90Cl, Amberlite FPA40Cl, Amberlite XAD18, Accurel EP100, ECR8206F / 5730, ECR8206 / 5803, ECR8206M / 5749, ReliZyme EP403, ReliZyme EP113, Lewatit VP OC 1600, Diaion WA20, Diaion WA21J, Diaion WA30, Dowex 66, Diaion HPA-25L, Lewatit VP OC 1064 MD PH, Lewatit VP OC 1163, Lifetech ECR8304F, Lifetech ECR8309F, Lifetech ECR8315F, Lifetech ECR8204F, Lifetech Chromalite MIDA / M, Chromalite MIDA / M / Fe, Chromalite MIDA / M / Co, Chromalite MIDA / M / Ni, Chromalite MIDA / M / Cu and Chromalite MIDA / M / Zn.

[0194] cofactor The enzymatic activity of ATP-dependent nucleic acid ligase requires ATP as a cofactor. One molecule of ATP is converted to AMP per ligation reaction. The catalytic mechanism of ATP-dependent nucleic acid ligase and the role of ATP in nucleic acid ligation reactions are described above.

[0195] In some embodiments, the method is performed using the cofactor ATP. In some embodiments, the method is performed using a sub-stoichiometric concentration of the cofactor ATP.

[0196] In some embodiments, the method is performed using the cofactor AMP. In some embodiments, the method is performed using a sub-stoichiometric concentration of the cofactor AMP.

[0197] In some embodiments, the method is carried out using a mixture of the cofactors ATP and AMP, hi some embodiments, the method is carried out using sub-stoichiometric concentrations of the cofactors ATP and AMP.

[0198] Those skilled in the art can easily determine the stoichiometric concentration of cofactor required for a given ligation reaction based on the concentration of the oligonucleotide fragment and the number of ligation reactions required to produce the oligonucleotide product. For example, for a ligation reaction using 1 mM of oligonucleotide fragment that requires four ligation reactions to produce the oligonucleotide product, the stoichiometric concentration of ATP is 4 mM.

[0199] In some embodiments, the method is carried out using an ATP and / or AMP concentration of about 0.5 mM, about 1 mM, about 2 mM, about 3 mM, about 4 mM, about 5 mM, about 6 mM, about 7 mM, about 8 mM, about 9 mM, about 10 mM, about 12 mM, about 14 mM, about 16 mM, about 18 mM, or about 20 mM.

[0200] Polyphosphate Polyphosphate is used as a phosphate donor by PPK in the catalytic formation of ATP from AMP.

[0201] In some embodiments, the polyphosphate is a polyphosphate salt. In some embodiments, the polyphosphate salt is sodium polyphosphate (Madrell's salt) or sodium hexametaphosphate (Graham's salt). In some embodiments, the method is carried out using a stoichiometric excess of polyphosphate. In some embodiments, the method is carried out using a polyphosphate concentration of at least 5 mM, at least 10 mM, at least 15 mM, at least 20 mM, at least 25 mM, at least 30 mM, at least 35 mM, at least 40 mM, at least 45 mM, at least 50 mM, 55 mM, at least 60 mM, at least 65 mM, at least 70 mM, at least 75 mM, at least 80 mM, at least 85 mM, at least 90 mM, at least 95 mM, or at least 100 mM.

[0202] divalent cations The enzymatic activity of ATP-dependent nucleic acid ligase and PPK requires the presence of divalent cations.

[0203] In some embodiments, the divalent cation is Mg 2+ and / or Mn 2+In some embodiments, the method is carried out using a divalent cation concentration of 5-100 mM, 10-100 mM, 15-100 mM, 20-100 mM, 30-100 mM, 5-90 mM, 5-80 mM, 5-70 mM, 5-60 mM, 5-50 mM, or 30-50 mM. In some embodiments, the method is carried out using a divalent cation concentration of at least 5 mM, at least 10 mM, at least 15 mM, at least 20 mM, at least 25 mM, at least 30 mM, at least 35 mM, at least 40 mM, at least 45 mM, at least 50 mM, 55 mM, at least 60 mM, at least 65 mM, at least 70 mM, at least 75 mM, at least 80 mM, at least 85 mM, at least 90 mM, at least 95 mM, or at least 100 mM.

[0204] The divalent cation concentration will typically depend on the amount of ATP required to achieve complete ligation (which in turn depends on the starting concentration of ATP / AMP and the concentration and number of different oligonucleotide fragments).

[0205] nucleic acid The present disclosure further provides nucleic acid molecules encoding the fusion polypeptides described herein. The nucleic acid molecules encoding the fusion polypeptides described herein can be linked to one or more heterologous regulatory sequences that control gene expression to generate recombinant polynucleotides capable of expressing the fusion polypeptides. An expression construct containing a heterologous polynucleotide encoding a fusion polypeptide can be introduced into a suitable host cell to express the corresponding fusion polypeptide.

[0206] As will be apparent to those skilled in the art, the availability of protein sequences and knowledge of the codons corresponding to various amino acids provides an example of all possible polynucleotides encoding a protein sequence of interest. The degeneracy of the genetic code, in which the same amino acid is coded for by alternative or synonymous codons, allows for the generation of an extremely large number of nucleic acid molecules, all of which encode the fusion polypeptides disclosed herein. Thus, once a particular amino acid sequence is determined, one skilled in the art can generate any number of different nucleic acids by modifying one or more codons in a manner that does not alter the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates all possible variations of nucleic acid molecules that can be made by selecting combinations based on possible codon choices for any of the polypeptides disclosed herein, including the exemplary fusion polypeptides listed in Example 1, as well as the amino acid sequences of any of the fusion polypeptides of SEQ ID NOS: 8-18, 90, 92, 94, 96, and 98 in the Sequence Listing, which are incorporated by reference.

[0207] In various embodiments, codons are preferably selected to be compatible with the host cell in which the recombinant protein will be produced, for example, bacterially preferred codons are used to express genes in bacteria, yeast preferred codons are used to express genes in yeast, and mammalian preferred codons are used to express genes in mammalian cells.

[0208] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 1, the amino acid sequence of bacteriophage RB69 RNA ligase 2 (UniProt ID: Q7Y4V8).

[0209] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:30, which is a nucleic acid sequence that encodes the amino acid sequence of SEQ ID NO:1.

[0210] In some embodiments, the nucleic acid molecule encodes a fusion peptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:2, the amino acid sequence of optimized bacteriophage RB69 RNA ligase 2.

[0211] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:31, which is a nucleic acid sequence encoding the amino acid sequence of SEQ ID NO:2.

[0212] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:88, the amino acid sequence of optimized bacteriophage RB69 RNA ligase 2.

[0213] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:87, which is a nucleic acid sequence that encodes the amino acid sequence of SEQ ID NO:88.

[0214] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 3, the amino acid sequence of bacteriophage T4 RNA ligase 2 (UniProt ID: P32277).

[0215] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 32, which is a nucleic acid sequence encoding the amino acid sequence of SEQ ID NO: 3.

[0216] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 4, the amino acid sequence of bacteriophage T4 DNA ligase (UniProt ID: P00970).

[0217] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 33, which is a nucleic acid sequence encoding the amino acid sequence of SEQ ID NO: 4.

[0218] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 5, the amino acid sequence of PPK12.

[0219] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 34, which is a nucleic acid sequence that encodes the amino acid sequence of SEQ ID NO: 5.

[0220] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of optimized PPK12, SEQ ID NO: 6.

[0221] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:35, which is a nucleic acid sequence encoding the amino acid sequence of SEQ ID NO:6.

[0222] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 7, the amino acid sequence of AjPAP (UniProt ID: Q83XD3).

[0223] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:36, which is a nucleic acid sequence encoding the amino acid sequence of SEQ ID NO:7.

[0224] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs:8-18.

[0225] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 37-47.

[0226] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 90, 92, 94, 96, or 98.

[0227] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 89, 91, 93, 95, or 97.

[0228] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs:8-18, 90, 92, 94, 96, or 98.

[0229] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 37-47, 89, 91, 93, 95, or 97.

[0230] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NO:17.

[0231] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NO:46.

[0232] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NO:18.

[0233] In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NO:47.

[0234] In some embodiments, the nucleic acid molecule encodes a fusion polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NO:98.

[0235] In some embodiments, the nucleic acid molecule comprises a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NO:97.

[0236] Nucleic acid molecules encoding fusion polypeptides can be engineered in a variety of ways to allow expression of the fusion polypeptide, for example, by codon optimization to improve expression in a host cell, insertion into suitable expression elements with or without additional regulatory sequences, and transformation into a host cell suitable for fusion polypeptide expression and production.

[0237] Depending on the expression vector, manipulation of the nucleic acid molecule prior to insertion into the vector may be desirable or necessary. Techniques for modifying nucleic acid sequences using recombinant DNA methods are well known in the art. Guidance is provided, for example, in Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press; and Current Protocols in Molecular Biology, Ausubel, F. Eds., Greene Pub. Associates, 1998, updated 2010.

[0238] If the sequence of a polypeptide is known, encoding polynucleotides can be prepared by standard solid-phase methods according to known synthesis methods. In some embodiments, fragments of up to about 100 bases can be synthesized separately and then ligated (e.g., by enzymatic or chemical ligation or polymerase-mediated methods) to form any desired contiguous sequence. For example, nucleic acid molecules of the present disclosure can be prepared by chemical synthesis using the classical phosphoramidite method, as described, for example, in Beaucage et al., 1981, Tet Lett 22:1859-69, or Matthes et al., People, 1984, EMBO J. 3:801-05, which is commonly performed in automated synthesis methods. Nucleic acid molecules are synthesized according to the phosphoramidite method, purified, annealed, ligated, and cloned into a suitable vector, for example, in an automated DNA synthesizer. Furthermore, essentially any nucleic acid is available from any of a variety of commercial sources.

[0239] vector The present disclosure provides recombinant expression vectors comprising nucleic acid molecules encoding the PPKs, ATP-dependent nucleic acid ligases, and / or fusion polypeptides described herein. In some embodiments, the vectors are selected from plasmids, cosmids, bacteriophages, or viral vectors. Recombinant expression vectors typically include one or more expression control regions, such as promoters and terminators, origins of replication, etc.

[0240] Nucleic acid molecules encoding the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide described herein can be expressed by inserting the nucleic acid sequence or a nucleic acid construct containing this sequence into an appropriate expression vector. In creating an expression vector, a coding sequence is placed in a vector so that it is linked to appropriate control sequences for expression. Recombinant expression vectors can be any vector (e.g., a plasmid or virus) that can be conveniently used in recombinant DNA procedures and can result in the expression of a polynucleotide sequence. The vector is generally selected based on its compatibility with the host cell into which it will be introduced. The vector can be a linear or circular plasmid. Expression vectors can be autonomously replicating vectors, i.e., vectors that exist as extrachromosomal entities whose replication is independent of chromosomal replication, such as plasmids, extrachromosomal elements, minichromosomes, or artificial chromosomes. The vector can contain any tool that ensures self-replication. Alternatively, the vector can be a vector that, upon introduction into a host cell, integrates into the genome and replicates along with the chromosome into which it is integrated. Furthermore, a single vector or plasmid, or two or more vectors or plasmids, can be used that together contain the total DNA to be introduced into the genome of the host cell.

[0241] Many expression vectors useful for embodiments of the present disclosure are commercially available. Exemplary expression vectors can be prepared by inserting a polynucleotide encoding a fusion polypeptide into the plasmid pACYC-Duet-1 (Novagen), pBR322 vector (New England Biolabs), pUC19 vector (New England Biolabs), or pET T7 expression vector (Novagen).

[0242] host cell The present disclosure also provides host cells capable of expressing the fusion polypeptides described herein. In some embodiments, the host cells comprise a nucleic acid molecule described herein or a vector described herein. In some embodiments, the host cell is Escherichia coli.

[0243] In some embodiments, the nucleic acid molecule encoding the polypeptide is linked to one or more control sequences for expression of the polypeptide in a host cell. Host cells for expressing the polypeptide encoded by the expression vectors of the present disclosure are well known in the art and include, but are not limited to, bacterial cells such as E. coli, Streptomyces, and Salmonella typhimurium; fungal cells (e.g., Saccharomyces cerevisiae or Pichia pastoris); insect cells such as Drosophila S2 and Spodoptera Sf9; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. An exemplary host cell is E. coli BL21(DE3). The host cells may be wild-type or modified. Appropriate culture media and growth conditions for the above host cells are well known in the art.

[0244] Nucleic acid molecules or vectors used to express polypeptides can be introduced into cells by a variety of methods known in the art. Techniques include, among others, electroporation, bioparticle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion. Various methods for introducing polynucleotides into cells are known to those skilled in the art.

[0245] Host cells can be used to express and isolate the polypeptides described herein.

[0246] In some embodiments, the present disclosure also provides a process for producing a fusion polypeptide described herein, the process comprising culturing a host cell capable of expressing a polynucleotide encoding the fusion polypeptide under culture conditions suitable for expression of the fusion polypeptide. In some embodiments, the process for preparing a fusion polypeptide further comprises isolating the fusion polypeptide. The fusion polypeptide is expressed in a suitable cell and can be isolated (or recovered) from the host cell and / or culture medium using any one or more well-known techniques for purifying proteins, including, among others, lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography.

[0247] Reaction conditions As disclosed herein and illustrated in the Examples, the present disclosure contemplates a range of suitable reaction conditions that can be used in the methods herein, including, but not limited to, pH, temperature, buffer, substrate load, enzyme load, cofactor load, pressure, and reaction time. Additional suitable reaction conditions for the ligation reactions described herein can be readily optimized by routine experimentation, e.g., by performing the methods described herein under experimental reaction conditions of varying reagent concentrations, pH, and temperature to determine the rate of oligonucleotide product formation.

[0248] In any of the process embodiments disclosed herein, the reaction conditions can include a suitable pH. As noted above, the desired pH or desired pH range can be maintained using an acid or base, a suitable buffer, or a combination of a buffer and added acid or base. The pH of the reaction mixture can be controlled before and / or during the reaction. In some embodiments, suitable reaction conditions include a solution pH of about 4 to about 8, a pH of about 5 to about 7, a pH of about 6 to about 8, or a pH of about 7 to about 8. In some embodiments, the reaction conditions include a solution pH of about 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, or 8.

[0249] In any of the process embodiments disclosed herein, a temperature appropriate for the reaction conditions can be used, taking into account, for example, increased reaction rate at higher temperatures and enzyme activity over a sufficient reaction time. Thus, in some embodiments, suitable reaction conditions include temperatures of about 10°C to about 60°C, about 10°C to about 50°C, about 25°C to about 50°C, about 25°C to about 40°C, about 25°C to about 30°C, or about 10°C to about 30°C. In some embodiments, suitable reaction temperatures include temperatures of about 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, or 60°C. In some embodiments, the temperature during the enzymatic reaction can be maintained at a specific temperature throughout the reaction. In some embodiments, the temperature during the enzymatic reaction can be adjusted along a temperature profile over the course of the reaction.

[0250] The reaction can be carried out in any suitable buffer solution. In some embodiments, the buffer solution is selected from Tris buffer (e.g., Tris-HCl), phosphate buffer, HEPES, MOPS (3(N-morpholino)propanesulfonic acid), and triethanolamine (TEOA) buffer. In some embodiments, the buffer solution comprises acetate, citrate, prolamine, carbonate, or phosphate, or any combination thereof. In some embodiments, the buffer solution is phosphate buffered saline (PBS), e.g., PBS with a NaCl concentration of <100 mM.

[0251] In some embodiments, the reaction mixture further comprises a reducing agent, optionally DTT (dithiothreitol).

[0252] Suitable solvents include water, aqueous buffers, organic solvents, and / or co-solvent systems (generally comprising an aqueous solvent and an organic solvent). The aqueous solutions (water or aqueous co-solvent systems) can be pH buffered or unbuffered.

[0253] Suitable reaction conditions can include a combination of reaction parameters that provide for the biocatalytic conversion of oligonucleotide fragments to corresponding oligonucleotide products.

[0254] In carrying out the ligation reactions described herein, the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide can be added to the reaction mixture in different formulations: as frozen or lyophilized whole cells (FWC or LWC) transformed with a gene encoding the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide, and / or as a cell lysate or lyophilized cell lysate of such cells, so-called shake flask powder (SFP), from which cellular debris is removed and / or further purified as a fermentation powder (FP). Whole cells transformed with a gene encoding the PPK, ATP-dependent nucleic acid ligase, and / or fusion polypeptide, or their cell extracts, lysates, and isolated enzymes, can be used in a variety of different forms, including solids (e.g., lyophilized, spray-dried, etc.) or semi-solids (e.g., crude pastes). Cell extracts or lysates can be partially purified by precipitation (e.g., ammonium sulfate, polyethyleneimine, heat treatment, etc.) followed by desalting procedures (e.g., ultrafiltration, dialysis, etc.) prior to lyophilization. Any of the enzyme preparations can be immobilized on a solid-phase material (e.g., a resin).

[0255] In any of the embodiments of the processes disclosed herein in which the modified polypeptide is expressed in the form of a secreted polypeptide, the culture medium containing the secreted polypeptide can be used in the processes herein.

[0256] In any of the process embodiments disclosed herein, solid reactants (e.g., enzymes, salts, etc.) can be provided to the reactants in a variety of forms, including powders (e.g., lyophilized, spray-dried, etc.), solutions, emulsions, suspensions, etc. The reactants can be readily lyophilized or spray-dried using methods and equipment known to those of skill in the art. For example, a protein solution can be frozen in aliquots at −80° C. and then placed in a pre-cooled lyophilization chamber, after which a vacuum can be applied.

[0257] In any of the process embodiments disclosed herein, the order of addition of the reactants is not important: the reactants may be added together to the solvent at the same time, or some of the reactants may be added separately, or some may be added together at different times.

[0258] The method for carrying out a ligation reaction may further comprise isolating the oligonucleotide product of the enzymatic reaction. In particular, this step is typically carried out after the enzymatic reaction is completed. The product is typically separated from one or more components of the reaction mixture, particularly from substantially all other components. For example, the product is typically separated from remaining substrates, by-products, and / or enzymes. Isolation of the product can be achieved by means and techniques known in the art, for example, by separating oligonucleotides based on their size, such as by gel electrophoresis, gel extraction, the use of a cellulose matrix, ultrafiltration, and / or chromatography.

[0259] stoichiometry The ligation method can be carried out under any suitable reaction conditions, which can be readily identified by one of skill in the art.

[0260] Those skilled in the art will readily understand that the concentrations of the reagents will vary depending on the concentration of the oligonucleotide fragments used. For example, as the concentration of the oligonucleotide fragments increases, more ATP molecules are required to achieve complete ligation, and therefore higher concentrations of polyphosphate and divalent cations are required to enable ATP regeneration from AMP.

[0261] The methods described herein are typically performed using a concentration of AMP, ATP, or a combined concentration of AMP and ATP that is less than the stoichiometric concentration. In some embodiments, the methods are performed using an ATP / AMP concentration that is less than the concentration of the oligonucleotide fragments.

[0262] In some embodiments where one ligation reaction is required to produce the oligonucleotide product, the method is carried out using about 1 mM of oligonucleotide fragment and less than about 1 mM ATP and / or AMP, optionally ≦0.5 mM ATP and / or AMP.

[0263] In some embodiments where one ligation reaction is required to produce the oligonucleotide product, the method is carried out using about 2 mM of oligonucleotide fragments and less than about 2 mM ATP and / or AMP, optionally ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0264] In some embodiments where one ligation reaction is required to produce the oligonucleotide product, the method is carried out using about 3 mM of oligonucleotide fragments and less than about 3 mM ATP and / or AMP, optionally ≦2.5 mM, ≦2 mM, ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0265] In some embodiments where one ligation reaction is required to produce the oligonucleotide product, the method is carried out using about 4 mM of oligonucleotide fragments and less than about 4 mM ATP and / or AMP, optionally ≦3.5 mM, ≦3 mM, ≦2.5 mM, ≦2 mM, ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0266] In some embodiments where one ligation reaction is required to produce the oligonucleotide product, the method is carried out using about 5 mM of oligonucleotide fragments and less than about 5 mM ATP and / or AMP, optionally ≦4.5 mM, ≦4 mM, ≦3.5 mM, ≦3 mM, ≦2.5 mM, ≦2 mM, ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0267] In some embodiments where one ligation reaction is required to produce the oligonucleotide product, the method is carried out using about 6 mM oligonucleotide fragments and less than about 6 mM ATP and / or AMP, optionally <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0268] In some embodiments where two ligation reactions are required to generate an oligonucleotide product, the method is carried out using about 1 mM of oligonucleotide fragment and less than about 2 mM ATP and / or AMP, optionally ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0269] In some embodiments where two ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 2 mM of oligonucleotide fragments and less than about 4 mM ATP and / or AMP, optionally ≦3.5 mM, ≦3 mM, ≦2.5 mM, ≦2 mM, ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0270] In some embodiments where two ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 3 mM of oligonucleotide fragment and less than about 6 mM ATP and / or AMP, optionally ≦5.5 mM, ≦5 mM, ≦4.5 mM, ≦4 mM, ≦3.5 mM, ≦3 mM, ≦2.5 mM, ≦2 mM, ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0271] In some embodiments where two ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 4 mM of oligonucleotide fragments and less than about 8 mM ATP and / or AMP, optionally <7.5 mM, <7 mM, <6.5 mM, <6 mM, <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0272] In some embodiments where two ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 5 mM of oligonucleotide fragments and less than about 10 mM ATP and / or AMP, optionally <9.5 mM, <9 mM, <8.5 mM, <8 mM, <7.5 mM, <7 mM, <6.5 mM, <6 mM, <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0273] In some embodiments where two ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 6 mM of oligonucleotide fragments and less than about 12 mM ATP and / or AMP, optionally <11.5 mM, <11 mM, <10.5 mM, <10 mM, <9.5 mM, <9 mM, <8.5 mM, <8 mM, <7.5 mM, <7 mM, <6.5 mM, <6 mM, <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0274] In some embodiments where three ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 1 mM of oligonucleotide fragment and less than about 3 mM ATP and / or AMP, optionally ≦2.5 mM, ≦2 mM, ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0275] In some embodiments where three ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 2 mM of oligonucleotide fragments and less than about 6 mM ATP and / or AMP, optionally <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0276] In some embodiments where three ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 3 mM of oligonucleotide fragments and less than about 9 mM ATP and / or AMP, optionally <8.5 mM, <8 mM, <7.5 mM, <7 mM, <6.5 mM, <6 mM, <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0277] In some embodiments where three ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 4 mM of oligonucleotide fragments and less than about 12 mM ATP and / or AMP, optionally <11.5 mM, <11 mM, <10.5 mM, <10 mM, <9.5 mM, <9 mM, <8.5 mM, <8 mM, <7.5 mM, <7 mM, <6.5 mM, <6 mM, <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0278] In some embodiments where three ligation reactions are required to generate the oligonucleotide product, the method is carried out using about 5 mM of oligonucleotide fragments and less than about 15 mM ATP and / or AMP, optionally <14.5 mM, <14 mM, <13.5 mM, <13 mM, <12.5 mM, <12 mM, <11.5 mM, <11 mM, <10.5 mM, <10 mM, <9.5 mM, <9 mM, <8.5 mM, <8 mM, <7.5 mM, <7 mM, <6.5 mM, <6 mM, <5.5 mM, <5 mM, <4.5 mM, <4 mM, <3.5 mM, <3 mM, <2.5 mM, <2 mM, <1.5 mM, <1 mM, or <0.5 mM ATP and / or AMP.

[0279] In some embodiments where three ligation reactions are required to generate the oligonucleotide product, the method comprises ligating oligonucleotide fragments at about 6 mM and less than about 18 mM ATP and / or AMP, optionally <17.5 mM, <17 mM, <16.5 mM, <16 mM, <15.5 mM, <15 mM, <14.5 mM, <14 mM, <13.5 mM, <13 mM, <12.5 mM, <14 mM, <15 mM, <16 mM, <17 mM, <17 mM, <16.5 mM, <16 mM, <15.5 mM, <15 mM, <14.5 mM, <14 mM, <13.5 mM, <13 mM, <12.5 mM, <14 mM, <15 mM, <16 mM, <16 mM, <17 mM, <17 mM, <16 mM, <16 mM, <16 mM, <16 mM, <16 mM, <17.5 mM, <17 mM, <16.5 mM, <16 mM, <16 mM, <16 mM, <16 mM, <16 mM, <17.5 ... 5.5 mM, ≦5 mM, ≦4.5 mM, ≦4 mM, ≦3.5 mM, ≦3 mM, ≦2.5 mM, ≦2 mM, ≦1.5 mM, ≦1 mM, or ≦0.5 mM ATP and / or AMP.

[0280] In some embodiments, the method provides a PPK concentration of about 1 g / L, optionally 1.1 g / L, 1.15 g / L, 1.2 g / L, 1.25 g / L, 1.3 g / L, 1.35 g / L, 1.4 g / L, 1.45 g / L, 1.5 g / L, 1.55 g / L, 1.6 g / L, 1.65 g / L, 1.7 g / L, 1.75 g / L, 1.8 g / L, 1.8 This is done using PPKs of 5g / L, 1.9g / L, 1.95g / L, 2g / L, 2.1g / L, 2.2g / L, 2.3g / L, 2.4g / L, 2.5g / L, 2.6g / L, 2.7g / L, 2.8g / L, 2.9g / L, 3g / L, 3.25g / L, 3.5g / L, 3.75g / L, 4g / L, 4.5g / L or 5g / L.

[0281] In some embodiments, the method comprises using about 1 g / L of ATP-dependent nucleic acid ligase, optionally 1.1 g / L, 1.15 g / L, 1.2 g / L, 1.25 g / L, 1.3 g / L, 1.35 g / L, 1.4 g / L, 1.45 g / L, 1.5 g / L, 1.55 g / L, 1.6 g / L, 1.65 g / L, 1.7 g / L, 1.75 g / L, 1.8 g / L, 1.8 The reaction is carried out using 5g / L, 1.9g / L, 1.95g / L, 2g / L, 2.1g / L, 2.2g / L, 2.3g / L, 2.4g / L, 2.5g / L, 2.6g / L, 2.7g / L, 2.8g / L, 2.9g / L, 3g / L, 3.25g / L, 3.5g / L, 3.75g / L, 4g / L, 4.5g / L or 5g / L of ATP-dependent nucleic acid ligase.

[0282] In some embodiments, the method comprises administering a concentration of about 1 g / L of fusion polypeptide, optionally 1.1 g / L, 1.15 g / L, 1.2 g / L, 1.25 g / L, 1.3 g / L, 1.35 g / L, 1.4 g / L, 1.45 g / L, 1.5 g / L, 1.55 g / L, 1.6 g / L, 1.65 g / L, 1.7 g / L, 1.75 g / L, 1.8 g / L, 1.8 This is performed using 5g / L, 1.9g / L, 1.95g / L, 2g / L, 2.1g / L, 2.2g / L, 2.3g / L, 2.4g / L, 2.5g / L, 2.6g / L, 2.7g / L, 2.8g / L, 2.9g / L, 3g / L, 3.25g / L, 3.5g / L, 3.75g / L, 4g / L, 4.5g / L or 5g / L of the fusion polypeptide.

[0283] In some embodiments where four ligation reactions are required to produce an oligonucleotide product, the method is carried out using about 1.25 g / L of ligase, about 1.25 g / L of kinase, about 1 mM of oligonucleotide fragment, about 2.5 mM of ATP and / or AMP, and about 20 mM of polyphosphate. In some embodiments where four ligation reactions are required to produce an oligonucleotide product, the method is carried out using about 1.25 g / L of fusion polypeptide, about 1 mM of oligonucleotide fragment, about 2.5 mM of ATP and / or AMP, and about 20 mM of polyphosphate.

[0284] In some embodiments where four ligation reactions are required to produce an oligonucleotide product, the method is carried out using about 1.25 g / L of ligase, about 1.25 g / L of kinase, about 3 mM of oligonucleotide fragments, about 1 mM of ATP and / or AMP, and about 20 mM of polyphosphate. In some embodiments where four ligation reactions are required to produce an oligonucleotide product, the method is carried out using about 1.25 g / L of fusion polypeptide, about 3 mM of oligonucleotide fragments, about 1 mM of ATP and / or AMP, and about 20 mM of polyphosphate.

[0285] In some embodiments where four ligation reactions are required to generate oligonucleotide products, the method is carried out using about 1.25 g / L of ligase, about 1.25 g / L of kinase, at least 3 mM (e.g., at least 3.1 mM, at least 3.2 mM, at least 3.3 mM, at least 3.4 mM, at least 3.5 mM, at least 3.6 mM, at least 3.7 mM, at least 3.8 mM, at least 3.9 mM, at least 4 mM, at least 4.5 mM, or at least 5 mM) of oligonucleotide fragments, up to 1 mM (e.g., <0.9 mM, <0.8 mM, <0.7 mM, <0.6 mM, or <0.5 mM) of ATP and / or AMP, and about 20 mM polyphosphate. In some embodiments where four ligation reactions are required to produce the oligonucleotide product, the method is carried out using about 1.25 g / L of fusion polypeptide, at least 3 mM (e.g., at least 3.1 mM, at least 3.2 mM, at least 3.3 mM, at least 3.4 mM, at least 3.5 mM, at least 3.6 mM, at least 3.7 mM, at least 3.8 mM, at least 3.9 mM, at least 4 mM, at least 4.5 mM, or at least 5 mM) of oligonucleotide fragment, up to 1 mM (e.g., <0.9 mM, <0.8 mM, <0.7 mM, <0.6 mM, or <0.5 mM) of ATP and / or AMP, and about 20 mM polyphosphate.

[0286] References to concentrations of ATP and / or AMP will be understood to include the concentration of ATP alone, the concentration of AMP alone, or the combined concentration of ATP and AMP, e.g., "2.5 mM ATP and / or AMP" encompasses: (i) an ATP concentration of 2.5 mM; (ii) an AMP concentration of 2.5 mM; and (iii) a combined ATP and AMP concentration of 2.5 mM.

[0287] qualification In some embodiments, the oligonucleotide fragments and / or oligonucleotides contain modifications, such as chemical modifications. As used herein, the term "oligonucleotide fragment" refers to one or more oligonucleotide fragments. It will be understood that modifications present in an oligonucleotide fragment are typically present in an oligonucleotide produced from said oligonucleotide fragment. In some embodiments, modifications are introduced into and / or removed from the oligonucleotide product.

[0288] In some embodiments, the oligonucleotide fragments and / or oligonucleotides comprise chemical modifications. In some embodiments, the oligonucleotide fragments and / or oligonucleotides comprise at least one backbone modification. In some embodiments, the oligonucleotide fragments and / or oligonucleotides comprise at least one nucleotide modification. In some embodiments, the oligonucleotide fragments and / or oligonucleotides comprise at least one sugar modification (e.g., at the 2' or 4' position). In some embodiments, the oligonucleotide fragments and / or oligonucleotides comprise: (i) at least one backbone modification; (ii) at least one nucleotide modification; and / or (iii) at least one sugar modification.

[0289] Modifications include, but are not limited to, terminal modifications of terminal oligonucleotide fragments, such as 5'-end modifications (phosphorylation, conjugation, inverted linkage) or 3'-end modifications (conjugation, inverted linkage, etc.); base modifications, such as substitution of stabilizing bases, destabilizing bases, or bases that base-pair with an expanded repertoire of partners, removal of bases (abasic nucleotides), or conjugated bases; sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions; or backbone modifications, including modification or substitution of phosphodiester bonds.

[0290] In some embodiments, the terminal oligonucleotide fragment and / or oligonucleotide comprises a cap. Terms such as "cap" include chemical moieties attached to the terminus of a double-stranded nucleotide duplex, but are used herein to exclude chemical moieties that are nucleotides or nucleosides. A "3' cap" is attached to the 3' end of a nucleotide or oligonucleotide to protect the molecule from degradation, for example, from nucleases such as those found in serum or intestinal fluids. A non-nucleotide 3' cap can replace a TT or UU dinucleotide at the end of a blunt-ended oligonucleotide, rather than a nucleotide. In some embodiments, non-nucleotide 3' terminal caps are as disclosed, for example, in WO 2005 / 021749 and WO 2007 / 128477; and U.S. Pat. Nos. 8,097,716, 8,084,600, and 8,344,128. A "5' cap" is attached to the 5' end of a nucleotide or oligonucleotide. The cap should not interfere (or unduly interfere) with oligonucleotide activity.

[0291] In some embodiments, the oligonucleotide fragment and / or oligonucleotide contains one or more mismatches. A mismatch is defined herein as a difference in base sequence or length when two sequences are maximally aligned and compared. In the context of a double-stranded oligonucleotide (two sequences aligned antiparallel to each other), a mismatch is defined as a position where a base in one sequence is not complementary to a base in the other sequence. Thus, for example, when a first and a second sequence are aligned antiparallel to each other, a mismatch is counted if a position in the first sequence has a particular base (e.g., A) and the corresponding position in the second sequence has a base (e.g., G) that is not complementary to the base in the first sequence. However, it should be noted that on a given RNA strand, U can be substituted by T (as RNA or, preferably, DNA, e.g., 2'-deoxy-thymidine); a substitution of U by T is not a mismatch as used herein, since either U or T can pair with A on the opposite strand. Thus, an RNA oligonucleotide can contain one or more DNA bases, such as T. Mismatches between the DNA portion of an RNAi agent and the corresponding target mRNA are not counted when base pairing occurs (e.g., between A, G, C, or T of the DNA portion and the corresponding U, C, G, or A, respectively, in the mRNA).

[0292] A mismatch is also counted, for example, when a position in one sequence has a base (e.g., A) and the corresponding position on the other sequence does not have a base (e.g., the position is an abasic nucleotide that contains a phosphate-sugar backbone but no base). A single-stranded nick in either sequence (or in the sense or antisense strand) is not counted as a mismatch. Thus, as a non-limiting example, if one sequence (in the 5'→3' orientation) contains the sequence AG, but the complementary sequence (in the 3'→5' orientation) contains the sequence TC with a single-stranded nick between the T and C, no mismatch is counted. Nucleotide modifications in the sugar or phosphate are also not considered mismatches. Thus, if one sequence contains G and the complementary sequence contains a modified C (e.g., a 2'-modification) at the same position, no mismatch would be counted.

[0293] Therefore, if the sugar, phosphate, or backbone of an oligonucleotide is modified without modifying the base, the mismatch is not counted. Thus, in the context of double-stranded RNAi, a strand having a given sequence as RNA will have zero mismatches with its complementary sequence, such as PNA; or morpholino; or LNA; or TNA; or GNA; or FANA; or a mixture or chimera of RNA and DNA, TNA, GNA, FANA, morpholino, UNA, LNA, and / or PNA. There will be no mismatch between a nucleotide that is T and a nucleotide that is A with a 5'-modification and / or 2'-modification. The important feature of a mismatch (base substitution) is that it cannot base-pair with the corresponding base on the opposite strand. In addition, when counting the number of mismatches, terminal overhangs such as "UU" or "dTdT" are not counted. In such cases, a mismatch is defined as a position where a base in one sequence does not match a base in the other sequence.

[0294] Note that dTdT (2'-deoxy-thymidine-5'-phosphate and 2'-deoxy-thymidine-5'-phosphate), or in some cases TT or UU, can be added to one or both 3' ends of an oligonucleotide as a terminal dinucleotide cap or extension, but this cap or extension is not included in the calculation of the total number of mismatches and is not considered part of the target sequence. This is because the terminal dinucleotide protects the end from nuclease degradation but does not contribute to target specificity (Elbashir et al. 2001 Nature 411:494-498; Elbashir et al. 2001 EMBO J. 20:6877-6888; and Kraynack et al. 2006 RNA 12:163-176).

[0295] Several examples exist in the art describing sugar, base, phosphate, and backbone modifications that can be introduced into nucleic acid molecules to significantly enhance nuclease stability and efficacy. For example, oligonucleotides have been modified to enhance stability and / or enhance biological activity by modification with nuclease-resistant groups, such as 2'-amino, 2'-C-allyl, 2'-fluoro, 2'-O-methyl, 2'-O-allyl, 2'-H, and nucleotide base modifications. Sugar modifications of nucleic acid molecules have been extensively described in the art.

[0296] Additional modifications and conjugation of oligonucleotides have been described. Soutschek et al. 2004 Nature 432:173-178 showed that cholesterol was conjugated to the 3' end of the sense strand of siRNA molecules via a pyrrolidine linker, thereby producing a covalent and irreversible conjugate. Chemical modifications of oligonucleotides (including conjugation with other molecules) can also be made to improve in vivo pharmacokinetic retention time and efficiency.

[0297] In some embodiments, the oligonucleotide fragments and / or oligonucleotides comprise modified bases. The present disclosure encompasses oligonucleotides and oligonucleotide fragments in which a single nucleotide at a given position is replaced with a modified version of the same nucleotide.Thus, a nucleotide (A, G, C, or U) can be 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2,2-dimethylamino ... Adenine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid(v), wybutoxosin, pseudouracil, queosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil uracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 2,6-diaminopurine, 5-hydroxymethylcytosine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiothymine, 5-propynyl (-C=C-CH3) uracil and cytosine, and Other alkynyl derivatives of pyrimidine bases may be substituted with modified bases selected from 6-azo uracil, cytosine and thymine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methyladenine, 2-F-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine.

[0298] Additional modified variants include the addition of any other moiety (e.g., a radiolabel or other tag or conjugate) to the oligonucleotide or oligonucleotide fragment, provided that the base sequence is identical and the addition of the other moiety results in a "modified variant" (no mismatch).

[0299] In addition to these modifications and patterns (e.g., formats) of modifications, other modifications or sets of modifications of the provided sequences can be generated using common knowledge of nucleic acid modifications. These various embodiments and embodiments of the oligonucleotides of the present disclosure can be used for RNA interference.

[0300] In some embodiments, the oligonucleotides and / or oligonucleotide fragments comprise modifications that cause the oligonucleotides to have increased stability in a biological sample or environment (e.g., cytoplasm, interstitial fluid, serum, lung or intestinal lavage fluid).

[0301] In some embodiments, the oligonucleotide and / or oligonucleotide fragment contains a modification that facilitates cleavage by the RNA-induced silencing complex (i.e., a "RISC cleavage site"). The RISC cleavage site is the site on the target where cleavage occurs. In some embodiments, the antisense strand contains a RISC cleavage site. For RNAi agents having a duplex region 17-23 nucleotides in length, the cleavage site on the antisense strand is typically approximately 10, 11, and 12 nucleotides from the 5' end. As used herein, the term "cleavage region" refers to the region located immediately adjacent to the cleavage site. In some embodiments, the cleavage region includes three bases on either end of the cleavage region and immediately adjacent to the cleavage region. In some embodiments, the cleavage region includes two bases on either end of the cleavage region and immediately adjacent to the cleavage region. In some embodiments, the cleavage site specifically occurs at the site bounded by nucleotides 10 and 11 of the antisense strand, and the cleavage region includes nucleotides 11, 12, and 13 of the antisense strand.

[0302] In some embodiments, the oligonucleotide fragments and / or oligonucleotides comprise a modified backbone. As used herein, an unmodified backbone consists of a 3' to 5' phosphodiester linkage. A modified backbone may contain a non-natural internucleoside linkage. Oligonucleotide fragments and / or oligonucleotides with modified backbones include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.

[0303] Oligonucleotide fragments and / or oligonucleotides containing modified backbones include, but are not limited to, those that do not have a phosphorus atom in the backbone.Modified backbones include, but are not limited to, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates (e.g., 3'-alkylene phosphonates and chiral phosphonates), phosphinates, phosphoramidates (e.g., mesyl phosphoramidate, 3'-amino phosphoramidate, and aminoalkyl phosphoramidates), thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, and boranophosphates with normal 3'-5' linkages, 2'-5'-linked analogs thereof, and those with reversed polarity where adjacent pairs of nucleoside units are linked 3'-5' to 5'-3' or 2'-5' to 5'-2'.

[0304] Oligonucleotide fragments and / or oligonucleotides containing modified backbones that do not contain phosphorus atoms therein can have backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages. These include those with morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and others with mixed N, O, S, and CH moieties.

[0305] In some embodiments, the oligonucleotide and / or oligonucleotide fragment comprises at least one phosphonate linkage, wherein the phosphonate is a modified phosphonate selected from the group consisting of: phosphorothioate (which may be the Rp or Sp isomer): [ka] ;phosphorodithioate: [ka] Methylphosphonates: [ka] Methoxypropylphosphonate: [ka] 5'-(E)-vinylphosphonate: [ka] 5'-Methylphosphonate: [ka] (S)-5'-C-methyl bearing phosphonate: [ka] 5'-phosphorothioate; [ka] and peptide nucleic acids: [ka]

[0306] In some embodiments, the oligonucleotide and / or oligonucleotide fragment comprises at least one 5'-uridine-adenine-3'-(5'-ua-3'-) dinucleotide (wherein uridine is a 2'-modified nucleotide); at least one 5'-uridine-guanine-3'-(5'-ug-3'-) dinucleotide (wherein 5'-uridine is a 2'-modified nucleotide); at least one 5'-cytidine-adenine-3'-(5'-ca-3'-) dinucleotide (wherein 5'-cytidine is a 2'-modified nucleotide); or at least one 5'-uridine-uridine-3'-(5'-uu-3'-) dinucleotide (wherein 5'-uridine is a 2'-modified nucleotide). These dinucleotide motifs are particularly susceptible to serum nuclease degradation (e.g., RNase A). Chemical modification at the 2' position of the first pyrimidine nucleotide in the motif prevents or slows such cleavage. This modified recipe is also known by the term "endolite."

[0307] In some embodiments, the oligonucleotide and / or oligonucleotide fragment comprises a modified nucleobase, wherein the modified nucleobase is difluorotolyl, nitroindolyl, nitropyrrolyl, or nitroimidazolyl. In certain embodiments, the modified nucleobase is difluorotolyl. In some embodiments, the oligonucleotide and / or oligonucleotide fragment is double-stranded, only one of the two strands comprises a modified nucleobase. In some embodiments, the oligonucleotide and / or oligonucleotide fragment is double-stranded, both strands comprise a modified nucleobase.

[0308] In some embodiments, the oligonucleotide fragment and / or oligonucleotide comprises a modified sugar. Sugar modifications typically involve chemical modifications of the sugar moiety of RNA or DNA. Sugar modifications include, but are not limited to, one of the following at the 2' position: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl (where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C6). 10 Alkyl or C2-C 10 (It can be alkenyl or alkynyl). Exemplary modifications include O[(CH) n O] m CH3, O(CH2). n OCH3, O(CH2) n NH2, -O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON[(CH2) n CH3)]2, where n and m are from 1 to about 10. The oligonucleotide fragments used in the methods described herein can include one of the following at the 2' position: C1 to C 10lower alkyl, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH, OCN, Cl, Br, CN, CF, OCF, SOCH, SOCH, ONO, NO, N, NH, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving group, reporter group, intercalating agent, group for improving the pharmacokinetic properties of a therapeutic RNA, or group for improving the kinetic properties of a therapeutic RNA. In some embodiments, modifications include 2'-methoxyethoxy (also known as 2'-O-(2-methoxyethyl) or 2'-O-MOE), 2'-dimethylaminooxyethoxy (also known in the art as 2'-O-dimethylaminoethoxyethyl or 2'-DMAEOE). Further exemplary modifications include: 5'-Me-2'-F nucleotides, 5'-Me-2'-OMe nucleotides, 5'-Me-2'-deoxynucleotides, 2'-alkoxyalkyl; and 2'-NMA (N-methylacetamide).

[0309] Other modifications include 2'-methoxy (2'-OCH), 2'-aminopropoxy (2'-OCHCHCHNH), and 2'-fluoro (2'-F). Similar modifications can also be made at other positions on the RNA, particularly the 3' position of the sugar on the 3' terminal nucleotide or in 2'-5' linked dsRNA, and the 5' position of 5' terminal nucleotide.

[0310] In some embodiments, the oligonucleotide fragment and / or oligonucleotide comprises at least one 2'-modified nucleotide. In some embodiments, the 2'-modification is selected from the group consisting of 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE). ), 2'-ON-methylacetamide (2'-O-NMA), locked nucleic acid (LNA), glycol nucleic acid (GNA), phosphoramidates (e.g., mesyl phosphoramidate), '2',3'-seconucleotide mimics, 2'-F-arabinonucleotides, abasic nucleotides, 2'-amino-modified nucleotides, 2'-alkyl-modified nucleotides, morpholino nucleotides, vinyl phosphonates (e.g., 5'-vinyl phosphonate), and cyclopropyl phosphonate deoxyribonucleotides. In some embodiments, one or more of the oligonucleotide fragments comprises a 2'-modification selected from the group consisting of 2'-OMe, 2'-F, and 2'-deoxy. In some embodiments, the oligonucleotide and / or oligonucleotide fragment comprises one or more 3'-O-methyl nucleotides.

[0311] In some embodiments, the oligonucleotides and / or oligonucleotide fragments comprise a 2'-modification selected from the group consisting of 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxye. 2'-O-methyl (2'-O-DMAEOE), 2'-ON-methylacetamide (2'-O-NMA), locked nucleic acid (LNA), glycol nucleic acid (GNA), phosphoramidate (e.g., mesyl phosphoramidate), 2',3'-seconucleotide mimics, 2'-F-arabinonucleotides, abasic nucleotides, 2'-amino-modified nucleotides, 2'-alkyl-modified nucleotides, morpholino nucleotides, vinyl phosphonates (e.g., 5'-vinyl phosphonate), deoxyribonucleotides, and cyclopropyl phosphonates. In some embodiments, the oligonucleotide and / or oligonucleotide fragment comprises one or more 3'-O-methyl nucleotides.

[0312] In some embodiments, the oligonucleotides and / or oligonucleotide fragments comprise a bridged nucleic acid. In some embodiments, the bridged nucleic acid is a locked nucleic acid. In some embodiments, the bridged nucleic acid is a constrained ethyl-bridged nucleic acid, as described below. [ka]

[0313] In some embodiments, all pyrimidines (uridine and cytidine) are 2'-O-methyl modified nucleosides.

[0314] In some embodiments, the sense and / or antisense strands are conjugated to one or more diagnostic compounds, reporter groups, crosslinkers, moieties that confer nuclease resistance, modified or unmodified nucleobases, lipophilic molecules, cholesterol, lipids, lectins, steroids, uvaol, hesigenin, diosgenin, terpenes, triterpenes, sarsasapogenin, friedelin, epifriedelanol-derivatized lithocholic acid, vitamins, carbohydrates, dextran, pullulan, chitin, chitosan, synthetic carbohydrates, oligolactate 15-mers, natural polymers, low or medium molecular weight polymers, inulin, cyclodextrin, hyaluronic acid, proteins, protein-binding agents, integrin targeting molecules, polycations, peptides, polyamines, peptidomimetics, and / or transferrin.

[0315] In some embodiments, the antisense strand comprises at least one 2'-OMe modified nucleotide. In some embodiments, the antisense strand comprises at least one 2'-F modified nucleotide. In some embodiments, the antisense strand comprises at least one 2'-deoxy modified nucleotide. In some embodiments, the antisense strand comprises at least one 2'-OMe modified nucleotide, at least one 2'-F modified nucleotide, or at least one 2'-deoxy modified nucleotide, or any combination thereof. In some embodiments, the antisense strand comprises alternating 2'-OMe and 2'-F modified nucleotides. In some embodiments, the antisense strand comprises at least one 5'-vinyl phosphonate. In some embodiments, the antisense strand comprises at least one chiral phosphorothioate linkage. In some embodiments, the antisense strand comprises at least one GNA. In some embodiments, the sense strand comprises at least one 2'-OMe modified nucleotide. In some embodiments, the sense strand comprises at least one 2'-F modified nucleotide. In some embodiments, the sense strand comprises at least one 2'-deoxy modified nucleotide. In some embodiments, the sense strand comprises at least one 2'-OMe modified nucleotide, at least one 2'-F modified nucleotide, or at least one 2'-deoxy modified nucleotide, or any combination thereof. In some embodiments, the sense strand comprises alternating 2'-OMe and 2'-F modified nucleotides. In some embodiments, the antisense strand and the sense strand each comprise at least one 2'-OMe modified nucleotide. In some embodiments, the antisense strand and the sense strand each comprise at least one 2'-F modified nucleotide. In some embodiments, the antisense strand and the sense strand each comprise alternating 2'-OMe and 2'-F modified nucleotides. In some embodiments, the sense strand comprises at least one 5'-vinyl phosphonate. In some embodiments, the sense strand comprises at least one chiral phosphorothioate linkage. In some embodiments, the sense strand comprises at least one GNA.

[0316] In some embodiments, the sense strand comprises alternating 2'-OMe and 2'-F modified nucleotides along the entire length of the sense strand, hi some embodiments, the sense strand comprises alternating 2'-OMe and 2'-F modified nucleotides along the entire length of the sense strand, for example, across at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides of the sense strand.

[0317] In some embodiments, the antisense strand comprises alternating 2'-OMe and 2'-F modified nucleotides along the entire length of the antisense strand, hi some embodiments, the antisense strand comprises alternating 2'-OMe and 2'-F modified nucleotides along a portion of the length of the antisense strand, for example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides of the antisense strand.

[0318] In some embodiments, the sense and antisense strands each comprise alternating 2'-OMe and 2'-F modified nucleotides along the entire length of the sense and antisense strands, hi some embodiments, the sense and antisense strands comprise alternating 2'-OMe and 2'-F modified nucleotides along a portion of the length of the sense and antisense strands, for example, along at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides of the sense and antisense strands.

[0319] Ligand In some embodiments, one or more of the oligonucleotide fragments are conjugated to at least one ligand. In some embodiments, the oligonucleotide product is conjugated to at least one ligand. The ligand can be conjugated in any configuration to the sense strand, the antisense strand, or both strands, for example, at the 3' end, the 5' end, non-terminal, or a combination.

[0320] In some embodiments, the ligand comprises one or more N-acetylgalactosamine (GalNAc) derivatives, hi some embodiments, the ligand comprises one or more GalNAc derivatives conjugated via a bivalent or trivalent branched carrier.

[0321] In some embodiments, the ligand is [ka] is.

[0322] In some embodiments, the ligand is [ka] is.

[0323] In some embodiments, the ligand is [ka] is.

[0324] In some embodiments, the ligand is [ka] is.

[0325] In some embodiments, the ligand is [ka] is.

[0326] In some embodiments, the ligand is [ka] is.

[0327] In some embodiments, a ligand alters the distribution, targeting, or lifetime of a molecule into which it is incorporated. In some embodiments, a ligand provides enhanced affinity for a selected target, e.g., a molecule, a cell or cell type, a compartment, a receptor, e.g., a cell or organ compartment, a tissue, an organ, or a region of the body, compared to a species in which such ligand is not present. Ligands that provide enhanced affinity for a selected target are also referred to as targeting ligands.

[0328] Some ligands may have endosome-destabilizing properties. The endosome-destabilizing ligand promotes endosomal lysis and / or transport of the oligonucleotide or a composition comprising the oligonucleotide from the endosome into the cytoplasm of the cell. The endosome-destabilizing ligand may be a polyanionic peptide or peptidomimetic that exhibits pH-dependent membrane activity and fusogenicity. In some embodiments, the endosome-destabilizing ligand assumes its active conformation at endosomal pH. An "active" conformation is one in which the endosome-destabilizing ligand promotes endosomal lysis and / or transport of the oligonucleotide or a composition comprising the oligonucleotide from the endosome into the cytoplasm of the cell. Exemplary endosome-destabilizing ligands include GALA peptides (Subbarao et al., Biochemistry, 1987, 26:2964-2972), EALA peptides (Vogel et al., J. Am. Chem. Soc., 1996, 118:1581-1586), and their derivatives (Turk et al., Biochem. Biophys. Acta, 2002, 1559:56-68). Endosome-destabilizing components may contain chemical groups (e.g., amino acids) that undergo changes in charge or protonation in response to changes in pH. Endosome-destabilizing components may be linear or branched.

[0329] Ligands can improve the transport, hybridization and specificity properties, as well as the nuclease resistance of the resulting natural or modified oligonucleotides.

[0330] Ligands may generally include therapeutic modifiers, e.g., to enhance uptake; diagnostic compounds or reporter groups, e.g., to monitor distribution; cross-linking agents; and nuclease-resistance-conferring moieties. Common examples include lipids, steroids, vitamins, sugars, proteins, peptides, polyamines, and peptidomimetics.

[0331] Ligands can include naturally occurring substances such as proteins (e.g., human serum albumin (HSA), low-density lipoprotein (LDL), high-density lipoprotein (HDL), or globulins); carbohydrates (e.g., dextran, pullulan, chitin, chitosan, inulin, cyclodextrin, or hyaluronic acid); or lipids. Ligands can also be recombinant or synthetic molecules, such as synthetic polymers, e.g., synthetic polyamino acids, oligonucleotides (e.g., aptamers). Examples of polyamino acids include polylysine (PLL), poly-L-aspartic acid, poly-L-glutamic acid, styrene-maleic anhydride copolymer, poly(L-lactide-co-glycolide) copolymer, divinyl ether-maleic anhydride copolymer, N-(2-hydroxypropyl)methacrylamide copolymer (HMPA), polyethylene glycol (PEG), polyvinyl alcohol (PVA), polyurethane, poly(2-ethylacrylic acid), N-isopropylacrylamide polymer, or polyphosphazine. Examples of polyamines include polyethyleneimine, polylysine (PLL), spermine, spermidine, polyamines, pseudopeptide-polyamines, pseudopeptide polyamines, dendrimer polyamines, arginine, amidine, protamine, cationic lipids, cationic porphyrins, quaternary salts of polyamines, or alpha-helical peptides.

[0332] The ligand can also include a targeting group, such as a cell or tissue targeting agent, such as a lectin, glycoprotein, lipid, or protein, such as an antibody, which binds to a specific cell type. The targeting group can be thyrotropin, melanotropin, lectin, glycoprotein, surfactant protein A, mucin carbohydrate, polyvalent lactose, polyvalent galactose, N-acetyl-galactosamine, N-acetyl-glucosamine polyvalent mannose, polyvalent fucose, glycosylated polyamino acid, polyvalent galactose, transferrin, bisphosphonate, polyglutamate, polyaspartate, lipid, cholesterol, steroid, bile acid, folate, vitamin B12, biotin, RGD peptide, RGD peptidomimetic, or aptamer.

[0333] Other examples of ligands include dyes, intercalators (e.g., acridine), crosslinkers (e.g., psoralens, mitomycin C), porphyrins (TPPC4, texaphyrin, sapphyrin), polycyclic aromatic hydrocarbons (e.g., phenazine, dihydrophenazine), artificial endonucleases or chelators (e.g., EDTA), lipophilic molecules, such as cholesterol, cholic acid, adamantaneacetic acid, 1-pyrenebutyric acid, dihydrotestosterone, 1,3-bis-O(hexadecyl)glycerol, geranyloxyhexyl group, hexadecylglycerol, borneol, menthol, 1,3-propanediol, heptadecyl group, palmitic acid, myristic acid, O3-(oleoyl)lithocholic acid, O3- (oleoyl)cholenic acid, dimethoxytrityl, or phenoxazine) and peptide conjugates (e.g., antennapedia peptide, Tat peptide), alkylating agents, phosphate, amino, mercapto, PEG (e.g., PEG-40K), MPEG, [MPEG]2, polyamino, alkyl, substituted alkyl, radiolabeled markers, enzymes, haptens (e.g., biotin), transport / absorption enhancers (e.g., aspirin, vitamin E, folic acid), synthetic ribonucleases (e.g., imidazole, bis-imidazole, histamine, imidazole clusters, acridine-imidazole conjugates, tetraazamacrocyclic EU3+ complexes), dinitrophenyl, HRP, or AP.

[0334] Ligands can be proteins, e.g., glycoproteins, or peptides, e.g., molecules with specific affinity for a co-ligand, or antibodies, e.g., antibodies that bind to specific cell types, such as cancer cells, endothelial cells, or bone cells. Ligands can also include hormones and hormone receptors. Ligands can also include non-peptide species, such as lipids, lectins, carbohydrates, vitamins, cofactors, multivalent lactose, multivalent galactose, N-acetyl-galactosamine, N-acetyl-glucosamine, multivalent mannose, multivalent fucose, or aptamers. Ligands can be, for example, lipopolysaccharides, activators of p38 MAP kinase, or activators of NF-κB.

[0335] In some embodiments, the ligand is a lipid or lipid-based molecule. Such lipid or lipid-based molecule preferably binds to a serum protein, such as human serum albumin (HSA). The HSA-binding ligand allows the conjugate to be distributed to a target tissue. The lipid or lipid-based ligand can (a) increase the resistance of the conjugate to degradation, (b) increase targeting or transport to a target cell or cell membrane, and / or (c) be used to adjust binding to a serum protein, such as HSA. The lipid-based ligand can be used to regulate, for example, control, the binding of the conjugate to a target tissue.

[0336] In some embodiments, the ligand is a peptide or peptidomimetic. A peptidomimetic is a molecule that can fold into a defined three-dimensional structure similar to a natural peptide. The peptide or peptidomimetic moiety can be about 5 to 50 amino acids in length, e.g., about 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 amino acids in length. The peptide or peptidomimetic can be, for example, a cell-penetrating peptide, a cationic peptide, an amphipathic peptide, or a hydrophobic peptide (e.g., composed primarily of Tyr, Trp, or Phe). The peptide moiety can be a dendrimeric peptide, a constrained peptide, or a cross-linked peptide. In another alternative, the peptide moiety can include a hydrophobic membrane translocation sequence (MTS). The peptide moiety can be a "delivery" peptide, which can transport large polar molecules, including peptides, oligonucleotides, and proteins, across cell membranes. Peptides or peptidomimetics can be encoded by random sequences of DNA, such as peptides identified from phage-display libraries or one-bead-one-compound (OBOC) combinatorial libraries (Lam et al., Nature, 354:82-84, 1991).

[0337] As used herein, a "peptide moiety" can range in length from about 5 amino acids to about 50 amino acids. The peptide moiety can have structural modifications, such as to enhance stability or direct conformational properties. Any of the structural modifications described below can be utilized. Arginine-glycine-aspartic acid (RGD)-peptide moieties can be used to target tumor cells, such as endothelial tumor cells or breast cancer tumor cells (Zitzmann et al., Cancer Res., 62:5139-43, 2002). RGD peptides can facilitate targeting of oligonucleotides to tumors in a variety of other tissues, including the lung, kidney, spleen, or liver (Aoki et al., Cancer Gene Therapy 8:783-787, 2001). RGD peptides can be linear or cyclic and can be modified, e.g., glycosylated or methylated, to facilitate targeting to specific tissues. Peptides that target markers enriched in proliferating cells can be used. For example, RGD-containing peptides and peptidomimetics can target cancer cells, particularly those that express integrins. Thus, the ligand may comprise an RGD peptide, a cyclic peptide containing RGD, an RGD peptide containing D-amino acids, or a synthetic RGD mimic.

[0338] Peptide and peptidomimetic ligands include naturally occurring or modified peptides, such as D or L peptides; α, β, or γ peptides; N-methyl peptides; azapeptides; peptides with one or more amide bonds, i.e., peptides in which a peptide bond is replaced by one or more urea, thiourea, carbamate, or sulfonylurea bonds; or cyclic peptides.

[0339] Ligands can be attached to oligonucleotide fragments and / or oligonucleotides at various locations, e.g., the 3'-terminus, 5'-terminus, and / or internal ("non-terminal") positions. In some embodiments, the ligand is attached via an intervening tether, e.g., a carrier described herein. The ligand or tethered ligand can be present on the monomer when the monomer is incorporated into the oligonucleotide fragment and / or oligonucleotide. In some embodiments, the ligand can be incorporated via attachment to a "precursor" monomer after the "precursor" monomer has been incorporated into the oligonucleotide fragment and / or oligonucleotide. For example, a monomer having, e.g., TAP-(CH2)nNH2, an amino-terminal tether (i.e., no associated ligand) can be incorporated into a growing oligonucleotide fragment. In a subsequent operation, i.e., after incorporation of the precursor monomer into the oligonucleotide fragment, a ligand having an electrophilic group, e.g., a pentafluorophenyl ester or aldehyde group, can then be attached to the precursor monomer by coupling the electrophilic group of the ligand with the terminal nucleophilic group of the precursor monomer's tether.

[0340] In another example, a monomer bearing a chemical group suitable for participating in a click chemistry reaction, such as an azide or alkyne terminal tether / linker, may be incorporated. In a subsequent operation, i.e., after the precursor monomer is incorporated into an oligonucleotide fragment and / or oligonucleotide, a complementary chemical group, such as an alkyne or azide, can be attached to the precursor monomer by coupling the alkyne and azide together.

[0341] In some embodiments, the ligand is conjugated to the nucleobase, sugar moiety, or internucleoside linkage of the oligonucleotide fragment and / or oligonucleotide. Conjugation to a purine nucleobase or a derivative thereof can occur at any position, including endocyclic and exocyclic atoms. In some embodiments, the 2-, 6-, 7-, or 8-position of a purine nucleobase is bound to a conjugate moiety. Conjugation to a pyrimidine nucleobase or a derivative thereof can also occur at any position. In some embodiments, the 2-, 5-, and 6-positions of a pyrimidine nucleobase can be substituted with a conjugate moiety. Conjugation to the sugar moiety of a nucleoside can occur at any carbon atom. Exemplary carbon atoms of the sugar moiety that can be bound to a conjugate moiety include the 2', 3', and 5' carbon atoms. In abasic residues, for example, the 1' position can also be bound to a conjugate moiety. The internucleoside linkage can also bear a conjugate moiety. In the case of phosphorus-containing linkages (e.g., phosphodiesters, phosphorothioates (e.g., chiral phosphorothioates), phosphorodithioates, phosphoramidates, etc.), the conjugate moiety can be attached directly to the phosphorus atom or to an O, N, or S atom attached to the phosphorus atom. In the case of amine- or amide-containing internucleoside linkages (e.g., PNA), the conjugate moiety can be attached to the nitrogen atom or adjacent carbon atom of the amine or amide.

[0342] In some embodiments, the ligand is conjugated to the sense strand. In some embodiments, the ligand is conjugated to the 3' end of the sense strand. In some embodiments, the ligand is conjugated to the 5' end of the sense strand. In some embodiments, the ligand is conjugated to a non-terminal end of the sense strand.

[0343] In some embodiments, the ligand is conjugated to the antisense strand. In some embodiments, the ligand is conjugated to the 3' end of the antisense strand. In some embodiments, the ligand is conjugated to a non-terminal end of the antisense strand.

[0344] The ligand can be attached via the carrier. The carrier comprises (i) at least one "backbone attachment point," preferably two "backbone attachment points," and (ii) at least one "tethering attachment point." As used herein, "backbone attachment point" refers to a functional group of a nucleic acid, such as a hydroxyl group, or generally to an available and suitable bond for incorporation of the carrier into the backbone, such as a phosphate or modified phosphate, e.g., a sulfur-containing backbone. In some embodiments, a "tethering attachment point" (TAP) refers to a ring atom, e.g., a carbon atom or heteroatom (different from the atom providing the backbone attachment point), of the cyclic carrier to which the selected moiety is attached. The moiety can be, for example, a carbohydrate, such as a monosaccharide, disaccharide, trisaccharide, tetrasaccharide, oligosaccharide, or polysaccharide. In some cases, the selected moiety is linked to the cyclic carrier by an intervening tether. Thus, the cyclic carrier often contains a functional group, e.g., an amino group, or generally provides a suitable bond for incorporation or tethering of another chemical entity, e.g., a ligand, into the ring.

[0345] When the oligonucleotide fragment is a dsRNA, the sense and / or antisense strand may be conjugated to the ligand via a carrier, which may be a cyclic or acyclic group; preferably, the cyclic group is selected from pyrrolidinyl, pyrazolinyl, pyrazolidinyl, imidazolinyl, imidazolidinyl, piperidinyl, piperazinyl, [1,3]dioxolane, oxazolidinyl, isoxazolyl, morpholinyl, thiazolidinyl, isothiazolidinyl, quinoxalinyl, pyridazinonyl, tetrahydrofuryl, and decalin; preferably, the acyclic group is selected from a serinol skeleton or a diethanolamine skeleton.

[0346] In some embodiments, one or more oligonucleotide fragments contain the sequence "TT," "dTdT," "dTsdT," or "UU" as a single-stranded overhang at the 3' end, also referred to herein as a terminal dinucleotide or 3'-terminal dinucleotide. dT is 2'-deoxy-thymidine-5'-phosphate, and sdT is 2'-deoxythymidine-5'-phosphorothioate. The terminal dinucleotide "UU" is UU or 2'-OMe-U 2'-OMe-U, and the terminal TT and terminal UU may be in an inverted / reverse orientation. Terminal dinucleotides (e.g., UU) are modified variants of dithymidine dinucleotides, which are generally placed as overhangs to protect the ends of siRNA from nucleases (see, for example, Elbashir et al. 2001 Nature 411:494-498; Elbashir et al. 2001 EMBO J. 20:6877-6888; and Kraynack et al. 2006 RNA 12:163-176). These references reveal that terminal dinucleotides enhance nuclease resistance but do not contribute to target recognition.

[0347] In some embodiments, one or both terminal oligonucleotide fragments contain a 3' terminal cap instead of or in addition to a terminal dinucleotide to stabilize the termini from nuclease degradation, provided that the 3' terminal cap is capable of both stabilizing the oligonucleotide (e.g., against nucleases) and not unduly interfering with its desired activity.

[0348] Oligonucleotides produced by the methods described herein are also provided.

[0349] The oligonucleotides can be present in any suitable buffer solution. In some embodiments, the buffer solution is selected from Tris buffer (e.g., Tris-HCl), phosphate buffer, HEPES, MOPS (3(N-morpholino)propanesulfonic acid), and triethanolamine (TEOA) buffer. In some embodiments, the buffer solution comprises acetate, citrate, prolamine, carbonate, or phosphate, or any combination thereof. In some embodiments, the buffer solution is phosphate-buffered saline (PBS), e.g., PBS having a NaCl concentration of <100 mM. In some embodiments, the buffer solution further comprises an agent for adjusting the osmolality of the solution so that the osmolality is maintained at a desired value, e.g., the physiological value of human plasma. Solutes that can be added to the buffer solution to control osmolality include, but are not limited to, proteins, peptides, amino acids, non-metabolized polymers, vitamins, ions, sugars, metabolites, organic acids, lipids, or salts. In some embodiments, the agent for adjusting the osmolality of the solution is a salt. In some embodiments, the agent for adjusting the osmolality of the solution is sodium chloride or potassium chloride.

[0350] Additional Embodiments Embodiment 1. A method of generating an oligonucleotide from two or more oligonucleotide fragments, comprising: i. two or more oligonucleotide fragments; ii. ATP-dependent nucleic acid ligase; iii. Polyphosphate kinase (PPK); iv. adenosine triphosphate (ATP) and / or adenosine monophosphate (AMP); v. Polyphosphates; and vi. divalent cations; contacting the thereby providing an oligonucleotide. Embodiment 2. Use of an ATP-dependent nucleic acid ligase and a PPK in generating an oligonucleotide from two or more oligonucleotide fragments. Embodiment 3. The method of embodiment 1 or the use of embodiment 2, wherein the two or more oligonucleotide fragments comprise two or more RNA oligonucleotide fragments. Embodiment 4 The method or use of embodiment 3, wherein the ATP-dependent nucleic acid ligase is an RNA ligase. Embodiment 5 The method of embodiment 4, wherein the RNA ligase is a double-stranded RNA ligase. Embodiment 6 The method of embodiment 4 or embodiment 5, wherein the RNA ligase is a member of the RNA ligase 2 family. Embodiment 7. The method or use of embodiment 6, wherein the RNA ligase is bacteriophage RB69 RNA ligase 2. Embodiment 8. The method of embodiment 1 or the use of embodiment 2, wherein the two or more oligonucleotide fragments comprise two or more DNA oligonucleotide fragments. Embodiment 9. The method or use according to embodiment 8, wherein the ATP-dependent nucleic acid ligase is a DNA ligase. Embodiment 10. The method or use according to embodiment 9, wherein the DNA ligase is T4 DNA ligase. Embodiment 11. The method according to any one of embodiments 1 or 3 to 10 or the use according to any one of embodiments 2 to 10, wherein the PPK is PPK12 or ajPAP. Embodiment 12. The method according to any one of embodiments 1 or 3 to 11 or the use according to any one of embodiments 2 to 11, wherein the ATP-dependent nucleic acid ligase and the PPK are ligated. Embodiment 13 The method or use according to embodiment 12, wherein the ATP-dependent nucleic acid ligase and the PPK are linked via a polypeptide linker. Embodiment 14. The method or use according to embodiment 13, wherein the PPK is located at the N-terminus of the linker and the ATP-dependent nucleic acid ligase is located at the C-terminus of the linker. Embodiment 15. a. the ATP-dependent nucleic acid ligase comprises a purification tag; b. the PPK comprises a purification tag; and / or c. the linker comprises a purification tag; The method according to any one of embodiments 1 or 3 to 14 or the use according to any one of embodiments 2 to 14. Embodiment 16. The method or use according to any one of embodiments 13 to 15, wherein the linker is a polypeptide linker comprising at least 3 amino acids, optionally at least 6 amino acids. Embodiment 17. The method or use of embodiment 16, wherein the linker comprises an amino acid sequence selected from the following: a) HHHHHH (SEQ ID NO: 19), optionally HHHHHHHHHHH (SEQ ID NO: 20); b) ENLYFQS (SEQ ID NO: 21); c) ENLYFQG (SEQ ID NO: 22); d) SSGSSG (SEQ ID NO: 23); e) GSAGSAAGSGEF (SEQ ID NO: 24); and / or f) GSSGSGSSSGGSSSSGSS (SEQ ID NO: 25). Embodiment 18. The method of any one of embodiments 1 or 3-17, wherein the polyphosphate is a polyphosphate salt. Embodiment 19. The method of embodiment 18, wherein the polyphosphate salt is sodium polyphosphate (Madrell's salt) or sodium hexametaphosphate (Graham's salt). Embodiment 20. The divalent cation cofactor is Mg 2+ or Mn 2+ 20. The method of any one of embodiments 1 or 3 to 19, wherein Embodiment 21. The method of any one of embodiments 1 or 3 to 20, performed at a divalent cation concentration of 5 to 100 mM, optionally 30 to 50 mM. Embodiment 22. The method of any one of embodiments 1 or 3 to 21, wherein the method is performed with substoichiometric concentrations of ATP and / or AMP. Embodiment 23. The method of any one of embodiments 1 or 3 to 22, further comprising purifying the oligonucleotide. Embodiment 24. The method according to any one of embodiments 1 or 3 to 23 or the use according to any one of embodiments 2 to 17, wherein the oligonucleotide is up to 60 nucleotides in length. Embodiment 25. The method according to any one of embodiments 1 or 3 to 24 or the use according to any one of embodiments 2 to 17 or 24, wherein each of the oligonucleotide fragments is 4 to 16 nucleotides in length, optionally 6 to 9 nucleotides in length. Embodiment 26. The method according to any one of embodiments 1 or 3 to 25, or the use according to any one of embodiments 2 to 17, 24 or 25, wherein the oligonucleotide fragment is single-stranded. Embodiment 27. The method of any one of embodiments 1 or 3 to 25, or the use of any one of embodiments 2 to 17, 24 or 25, wherein the oligonucleotide fragments are double-stranded, and optionally one or more of the double-stranded oligonucleotide fragments comprises one or two single-stranded overhangs. Embodiment 28. The method according to any one of embodiments 1 or 3 to 27 or the use according to any one of embodiments 2 to 17 or 24 to 27, wherein one or more of the oligonucleotide fragments comprises a chemical modification. Embodiment 29. The method or use of embodiment 28, wherein the chemical modification is selected from the following: (a) optionally a modified backbone selected from phosphorothioate (e.g., chiral phosphorothioate) or methylphosphonate internucleotide linkages; (b) optionally 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), 2'-O-methylacetamide (2'-O- modified nucleotides selected from 5'-amino- and 5'-amino-modified nucleotides, 2'-alkyl-modified nucleotides, morpholino nucleotides, vinyl phosphonates (e.g., 5'-vinyl phosphonates), and cyclopropyl phosphonate deoxyribonucleotides; and / or (c) conjugation to a ligand, optionally comprising one or more N-acetylgalactosamine (GalNAc) derivatives. Embodiment 30. The method according to any one of embodiments 1 or 3 to 29 or the use according to any one of embodiments 2 to 17 or 24 to 29, wherein the ATP-dependent nucleic acid ligase and / or PPK is immobilized. Embodiment 31. The method or use according to embodiment 30, wherein the ATP-dependent nucleic acid ligase and / or PPK is immobilized on the solid material by chemical bonding or physical adsorption. Embodiment 32. i. ATP-dependent nucleic acid ligase; ii.PPK; iii. ATP and / or AMP; iv. divalent cations; and v. Polyphosphate A composition comprising: Embodiment 33. The composition of embodiment 32, further comprising two or more oligonucleotide fragments. Embodiment 34. i. ATP-dependent nucleic acid ligase; ii.PPK; iii. ATP and / or AMP; iv. Polyphosphates; v. Divalent cations; and vi. Instructions for use in a method for generating an oligonucleotide from two or more oligonucleotide fragments. Kit including: Embodiment 35. The composition of embodiment 32 or embodiment 33 or the kit of embodiment 34, wherein the polyphosphate is a polyphosphate salt. Embodiment 36. The composition or kit of embodiment 35, wherein the polyphosphate salt is selected from Graham's salt and Maddrell's salt. Embodiment 37. The divalent cation is Mg 2+ or Mn 2+ 37. The composition of any one of embodiments 32, 33, 35 or 36, or the kit of any one of embodiments 34 to 36, wherein Embodiment 38. The composition of any one of embodiments 32, 33, or 35 to 37, or the kit of any one of embodiments 34 to 37, wherein the concentration of the divalent cation is 5 to 100 mM, optionally 30 to 50 mM. Embodiment 39. a) the PPK domain; and b) ATP-dependent nucleic acid ligase domain A fusion polypeptide comprising: Embodiment 40. The fusion polypeptide of embodiment 39, comprising a linker. Embodiment 41. A fusion polypeptide according to embodiment 39 or embodiment 40, wherein the PPK is PPK12 or ajPAP. Embodiment 42. A fusion polypeptide according to any one of embodiments 39 to 41, wherein the PPK domain comprises an amino acid sequence having at least 85% identity to the amino acid sequence of any one of SEQ ID NOs: 5 to 7. Embodiment 43. A fusion polypeptide according to any one of embodiments 39 to 42, wherein the ATP-dependent nucleic acid ligase domain is an RNA ligase domain. Embodiment 44. The fusion polypeptide of embodiment 43, wherein the RNA ligase domain is a double-stranded RNA (dsRNA) ligase domain. Embodiment 45. A fusion polypeptide according to embodiment 43 or embodiment 44, wherein the dsRNA ligase is a member of the RNA ligase 2 family. Embodiment 46. The fusion polypeptide of embodiment 45, wherein the dsRNA ligase is bacteriophage RB69 RNA ligase 2. Embodiment 47. A fusion polypeptide according to any one of embodiments 39 to 42, wherein the ATP-dependent nucleic acid ligase domain is a DNA ligase domain. Embodiment 48. The fusion polypeptide of embodiment 47, wherein the DNA ligase domain is a T4 DNA ligase domain. Embodiment 49. A fusion polypeptide described in any one of embodiments 39 to 48, wherein the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 85% sequence identity with the amino acid sequence of any one of SEQ ID NOs: 1 to 4. Embodiment 50. A fusion polypeptide described in any one of embodiments 39 to 48, wherein the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 1 to 4 or 88. Embodiment 51. A fusion polypeptide according to any one of embodiments 40 to 50, wherein the linker is located between the PPK domain and the ATP-dependent nucleic acid ligase domain. Embodiment 52. A fusion polypeptide according to any one of embodiments 40 to 51, wherein the PPK domain is located at the N-terminus of the linker and the ATP-dependent nucleic acid ligase domain is located at the C-terminus of the linker. Embodiment 53. A fusion polypeptide according to any one of embodiments 39 to 52, comprising a purification tag. Embodiment 54. A fusion polypeptide according to any one of embodiments 40 to 53, wherein the linker comprises a purification tag. Embodiment 55. A fusion polypeptide according to embodiment 53 or embodiment 54, wherein the purification tag is located at the N-terminus and / or C-terminus of the fusion polypeptide. Embodiment 56. A fusion polypeptide according to any one of embodiments 40 to 55, wherein the linker is a polypeptide linker comprising at least 3 amino acids, optionally at least 6 amino acids. Embodiment 57. The fusion polypeptide of embodiment 56, wherein the linker comprises an amino acid sequence selected from the following: a) HHHHHH (SEQ ID NO: 19), optionally HHHHHHHHHHH (SEQ ID NO: 20); b) ENLYFQS (SEQ ID NO: 21); c) ENLYFQG (SEQ ID NO: 22); d) SSGSSG (SEQ ID NO: 23); e) GSAGSAAGSGEF (SEQ ID NO: 24); and / or f) GSSGSGSSSGGSSSSGSS (SEQ ID NO: 25). Embodiment 58. A fusion polypeptide according to any one of embodiments 39 to 57, comprising an amino acid sequence having at least 85% sequence identity with the amino acid sequence of any one of SEQ ID NOs: 8 to 18. Embodiment 59. A fusion polypeptide described in any one of embodiments 39 to 57, comprising an amino acid sequence having at least 85% sequence identity with the amino acid sequence of any one of SEQ ID NOs: 8 to 18, 90, 92, 94, 96, or 98. Embodiment 60. The method of any one of embodiments 1 or 3 to 31 or the use of any one of embodiments 2 to 17 or 24 to 31, wherein the ATP-dependent nucleic acid ligase and the PPK are provided as a fusion polypeptide as defined in any one of embodiments 39 to 59. Embodiment 61. A nucleic acid encoding a fusion polypeptide according to any one of embodiments 39 to 59. Embodiment 62. The nucleic acid molecule of embodiment 61, comprising a nucleic acid sequence having at least 85% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 34 to 36. Embodiment 63. A nucleic acid molecule according to embodiment 61 or embodiment 62, comprising a nucleic acid sequence having at least 85% sequence identity with any one of the nucleic acid sequences of SEQ ID NOs: 30 to 33. Embodiment 64. A nucleic acid molecule according to embodiment 61 or embodiment 62, comprising a nucleic acid sequence having at least 85% sequence identity with the nucleic acid sequence of any one of SEQ ID NOs: 30 to 33 or 87. Embodiment 65. A vector comprising the nucleic acid of any one of embodiments 61 to 64. Embodiment 66. The vector of embodiment 65, selected from a plasmid, cosmid, bacteriophage or viral vector. Embodiment 67. A host cell comprising a nucleic acid molecule according to any one of embodiments 61 to 64 or a vector according to embodiment 65 or embodiment 66. Embodiment 68. The host cell of embodiment 67, which is E. coli. Embodiment 69. Use of a fusion polypeptide according to any one of embodiments 39 to 59 in an ATP-dependent nucleic acid ligation reaction. Embodiment 70. The rate of nucleic acid ligation exceeds the rate of nucleic acid ligation of a control; the control is: (a) a first protein consisting of the PPK domain of embodiment 39; and (b) a second protein consisting of the ATP-dependent nucleic acid ligase domain of embodiment 39; Including, 70. The use of embodiment 69, wherein the first and second proteins are not linked. Embodiment 71. Use of a fusion polypeptide according to any one of embodiments 39 to 59 in a method for generating an oligonucleotide from two or more oligonucleotide fragments. Embodiment 72. The method according to any one of embodiments 1, 3 to 31 or 60, or the use according to any one of embodiments 2 to 17, 24 to 31 or 69 to 71, wherein the oligonucleotide is a therapeutic oligonucleotide. Embodiment 73. The method of any one of embodiments 1, 3 to 31, 60 or 72, or the use of any one of embodiments 2 to 17, 24 to 31 or 69 to 72, wherein the oligonucleotide product is at least 80% pure, optionally the oligonucleotide product is at least 85% pure, at least 90% pure, at least 95% pure, optionally the oligonucleotide product is at least 98% pure.

[0351] Various features and embodiments of the present disclosure are illustrated in the following representative examples, which are intended to be illustrative and not limiting. [Example]

[0352] The following examples (including experiments and results achieved) are provided for illustrative purposes only and are not to be construed as limiting the invention.

[0353] In the examples below, the following abbreviations apply: ppm (parts per million); M (mole); mM (millimole), uM and μM (micromolar); nM (nanomole); mol (mole); gm and g (gram); mg (milligram); ug and μg (microgram); L and I (liter); ml and mL (milliliter); cm (centimeter); mm (millimeter); um and μm (micrometer); sec. (second); min(s) (minute); h(s) and hr(s) (hour); U (unit); MW (molecular weight); rpm (revolutions per minute); psi and PSI (pounds per square inch); °C (degrees Celsius); RT and rt (room temperature); OD 600 (optical density at 600 nm), CAM and cam (chloramphenicol); DMSO (dimethyl sulfoxide); FP (fermentation powder); FWC (frozen whole cells), LWC (dried frozen whole cells), PMBS (polymyxin B sulfate); IPTG (isopropyl β-D-1-thiogalactopyranoside); AMP (adenosine monophosphate); ADP (adenosine diphosphate), ATP (adenosine triphosphate), PolyP (polyphosphate); LB (lysogeny broth); TB (terrific broth; 12 g / L bacto-tryptone, 24 g / L yeast extract, 4 mL / L glycerol, 65 mM potassium phosphate, pH 7.0, 1 mM MgSO4); HEPES (HEPES zwitterionic buffer; 4-(2-hydroxyethyl)-piperazineethanesulfonic acid); SFP (shake flask powder); CDS (coding sequence); DNA (deoxyribonucleic acid); RNA (ribonucleic acid); Escherichia coli (E. coli) W3110 (a commonly used laboratory E. coli (E. coli) strain, available from the Coli Genetic Stock Center [CGSC], New Haven, CT); HTP (high throughput); HPLC (high pressure liquid chromatography); GC (gas chromatography), MS (mass spectrometry), RF (rapid injection), FIOP (fold improvement over positive control); Microfluidics (Microfluidics, Corp., Westwood, MA); Sigma-Aldrich (Sigma-Aldrich, St.Louis, MO; Difco (Difco Laboratories, BD Diagnostic Systems, Detroit, MI); Agilent (Agilent Technologies, Inc., Santa Clara, CA); Corning (Corning, Inc., Palo Alto, CA); Dow Corning (Dow Corning, Corp., Midland, MI); and Gene Oracle (Gene Oracle, Inc., Mountain View, CA). .

[0354] Example 1 Results and Discussion Polyphosphate kinase (PPK12) has previously been shown to convert AMP to ATP (via an ADP intermediate) using polyphosphate as the phosphate donor (Tavanti, M., Hosford, J., Lloyd, R.C. & Brown, M.J.B. Green Chemistry 23, 828-837 (2021)). To verify the activity of PPK12, wild-type PPK12 was produced in Escherichia coli (E. coli) and applied to the reaction as a lyophilized cell-free extract, referred to herein as shake-flask powder (SFP). Starting with AMP and 100 equivalents of polyphosphate, ATP production was monitored at different doses of PPK12; the reaction reached equilibrium with a maximum conversion of 66% to ATP (Figure 2), consistent with previous reports.

[0355] To improve the catalytic activity of PPK12, a saturation mutagenesis library was designed, constructed, and screened. The best-performing enzyme identified in the screen (SEQ ID NO: 6) was approximately two-fold more active than the wild-type enzyme and was used in all subsequent experiments.

[0356] To develop an efficient biocatalytic process, it is important to maximize the space-time yield of the ligation reaction. To this end, ligation reactions were performed at a high substrate concentration of at least 1 mM. As shown in Figure 3, to generate a full-length oligonucleotide (in this example, siRNA) product from an oligonucleotide fragment, the ligase must catalyze four ligation reactions. Therefore, at least 4 mM ATP is required to process 1 mM of substrate. In practice, excess ATP is required to achieve complete ligation of the substrate. Using a dsRNA ligase (SEQ ID NO: 2) evolved to function at high substrate concentrations, complete conversion starting from 1 mM substrate oligonucleotide was obtained in the presence of 10 mM ATP (Figures 2 and 5). As expected, the addition of only 2.5 mM ATP resulted in only partial product formation (Figure 4). In contrast, no product was observed when AMP was supplied instead of ATP (Figure 4).

[0357] As demonstrated herein, ligation activity in the presence of AMP can be rescued by the addition of PPK12 and polyphosphate (Figures 4 and 5). In the presence of substoichiometric concentrations of AMP (2.5 mM), the oligonucleotide fragments were completely converted to oligonucleotide products, confirming that PPK12 converts AMP to ATP multiple times during the ligation reaction. These data indicate that PPK12 and polyphosphate function as an effective ATP regeneration system during the ligation reaction. Control experiments without polyphosphate, kinase, or ligase demonstrate that each of these components must be present to achieve ligation (Figure 4).

[0358] The use of substoichiometric amounts of AMP positively impacts the cost and sustainability of the ligation reaction. To offset the potential cost and resource demands associated with using two separate enzymes (i.e., ligase and kinase), a gene fusion of the kinase and ligase genes was generated and used to express a single fusion polypeptide containing both the kinase and ligase domains. A linker was included between the two enzymes to help each enzyme maintain activity without relying on the other. Producing a bifunctional biocatalyst from a single fermentation rather than two fermentations advantageously saves the time, effort, and expense associated with enzyme production.

[0359] To investigate whether modified PPKs could be linked to modified dsRNA ligases, 11 different fusion constructs were designed. These constructs, referred to as "Kligases," are provided in Table 1.

[0360] All Clignases were expressed as soluble proteins as determined by SDS-PAGE analysis (Figure 6). To evaluate the activity of each Clignase relative to its respective unligated enzyme, oligonucleotide product formation was monitored at various enzyme concentrations under the previously described ATP regeneration conditions. Oligonucleotide products were identified for all Clignases, demonstrating that both ATP synthesis and ligation activity were maintained (Figure 7, Figure 5). Relative activity varied among fusion constructs (Figure 7, Figure 5), with constructs containing kinases at the N-terminus (Clignases 1, 3, 4, 5, 10, and 11) typically exhibiting higher activity than unligated ligases. The most active Clignases (Clignases 10 and 11) exhibited approximately 1.5- to 2-fold higher ligase activity than unligated enzymes. Exemplary Clignases 4 and 11 exhibit higher ligase activity compared to unligated ligases at equimolar amounts (Figure 8). While not wishing to be bound by theory, it is possible that ligated enzymes are more stable. Alternatively, the close proximity of the kinase to the ligase may provide a local supply of ATP to the ligase, thereby accelerating the reaction.

[0361] [Table 2]

[0362] conclusion Enzymatic ligation of short oligonucleotide fragments offers a sustainable and economical alternative to solid-phase chemical synthesis of full-length (e.g., therapeutic) oligonucleotides. However, one drawback of ligase-catalyzed reactions is the requirement for stoichiometric amounts of the expensive cofactor, ATP. As demonstrated herein, the addition of a second enzyme, polyphosphate kinase (e.g., PPK12 or a modified variant of PPK12) and polyphosphate to the reaction facilitates cofactor regeneration, enabling complete ligation at high substrate concentrations using a ligase (e.g., dsRNA ligase or modified dsRNA ligase) in the presence of substoichiometric amounts of AMP, which is significantly cheaper than ATP. Kinases and ligases (e.g., PPK and dsRNA ligase) can be expressed together as a single polypeptide, further simplifying and reducing the cost of the biocatalytic process. Advantageously, linking the kinase and ligase can improve the activity of the ligase compared to the activity of the unlinked ligase.

[0363] method Oligonucleotide synthesis Substrate fragments and reference oligonucleotides were synthesized by a commercial partner. Oligonucleotide substrates 1-6 and products AS (antisense strand) and SS (sense strand) are provided in Table 2.

[0364] [Table 3]

[0365] DNA sequence encoding the enzyme >PPK12 with a C-terminal tag (GQTGHHHHHH; SEQ ID NO: 27) [ka] >Optimized PPK with C-terminal tag (GQTGHHHHHH; SEQ ID NO: 27) [ka] >dsRNA ligase with an N-terminal tag (MHHHHHHENLYFQS; SEQ ID NO: 26) [ka] >Optimized dsRNA ligase with an N-terminal tag (MHHHHHHENLYFQS; SEQ ID NO: 26) [ka]

[0366] Enzyme expression shake flask powder production DNA encoding the engineered enzyme was cloned into the pCK110900 vector and transformed into W3110 E. coli electrocompetent cells. Single bacterial colonies were picked and incubated with 1% glucose and 30 μg mL -1 Grown overnight in 25 mL of LB medium (Teknova) supplemented with chloramphenicol at 37°C, 200 rpm, and humidity of 85°C. After overnight growth, 5 mL of culture was used to elucidate 30 μg.mL -1 Inoculate a shake flask containing 250 mL of TB medium (Teknova) supplemented with chloramphenicol and measure the optical density (OD) of the culture. 600 The cells were grown at 30°C, 200 rpm, and 85% humidity until the RI (ratio of RI to RI) reached 0.7. Enzyme expression was induced by adding IPTG to a final concentration of 1 mM, and the cultures were further grown at 30°C, 200 rpm, and 85% humidity for 16 hours. Cells were harvested by centrifugation at 4000g for 5 minutes at 4°C, and the medium was discarded. The cell pellet was stored at -80°C until ready for use.

[0367] The cell pellet was thawed, resuspended in 30 mL of 50 mM Tris-HCl (pH 7.5) supplemented with 1 mM DTT, and lysed by passage through a microfluidizer (Microfluidics, LM20). The lysate was clarified by centrifugation at 40,000 g for 30 minutes at 4°C, frozen at -80°C, and then lyophilized. The resulting shake flask powder (SFP) was stored at -20°C until use.

[0368] ATP synthesis reaction The ATP synthesis activity of wild-type PPK12 was examined at various enzyme concentrations. A 5 g / L stock solution of wild-type PPK12 was prepared by dissolving 50 mg of SFP in 10 mL of 50 mM Tris-HCl (pH 8.0). Seven samples, 2-fold serial dilutions, were prepared by diluting the stock solution accordingly, starting with 0.5 g / L PPK12. A blank control containing 0 g / L PPK12 was included as the eighth sample.

[0369] ATP synthesis reactions were set up using 0.3 mM AMP, 30 mM polyphosphate (Madrell's salt), 5 mM MgCl, 5 mM DTT, and a 20% (v / v) dilution series of PPK12 (starting at 0.1 g / L PPK12 final reaction concentration) in 50 mM Tris-HCl (pH 8.0). Reactions were set up in a BioRad hardshell PCR plate and incubated at 30°C for 30 minutes in a thermocycler. After 30 minutes, the reaction was quenched by diluting the reaction mixture 2.5-fold with 10 mM EDTA (pH 7.0). Samples were analyzed by HPLC as described below.

[0370] HPLC analysis of the ATP synthesis reaction The ATP synthesis reaction was analyzed by HPLC using an Agilent Infinity Lab Poroshell 120 HILIC-Z column (50 x 2.1 mm, 1.7 μm) on a Thermo Scientific UHPLC Vanquish Horizon. Mobile phase A (MPA) consisted of 20 mM ammonium acetate (pH 9.0) (Honeywell HPLC grade) in Milli-Q grade water. Mobile phase B (MPB) consisted of acetonitrile (Superior HPLC grade) / 150 mM ammonium acetate (pH 10.5). Elution began with a linear gradient from 95% MPB to 72% MPB in 0.8 min, followed by a second linear gradient to 48% MPB in 0.1 min. 48% MPB was maintained for 0.3 min, then changed to 95% MPB in 0.1 min, which was maintained for 0.45 min to re-equilibrate the column for the next sample, resulting in a total run time of 1.75 min. The flow rate is set at 1.0 mL / min and the column temperature is set at 25° C. Detection is monitored by UV at λ=260 nm.

[0371] Ligation reaction and ATP regeneration Ligation activity and ATP regeneration activity were examined at different concentrations. A 50 g / L stock solution of dsRNA ligase, PPK, and Cligase was prepared by dissolving 15 mg of SFP in 300 μL of 50 mM Tris-HCl (pH 7.0). For ligation reactions using only dsRNA ligase or kinase, the stock solution was diluted to 25 g / L with 50 mM Tris-HCl (pH 7.0). For ligation reactions requiring both ligase and kinase, the two stock solutions were combined 1:1 to obtain a 25 g / L stock solution of each enzyme. For ligation reactions using Cligase, the stock solution was not diluted, given that Cligase has approximately twice the molecular weight of the ligase / kinase. Therefore, seven samples, 2-fold serial dilutions, were prepared by diluting the stock solutions, starting with 5 g / L ligase and / or kinase and 10 g / L Cligase. A blank control containing 0 g / L enzyme was included as the eighth sample.

[0372] Ligation reactions were set up as follows: oligonucleotide fragments 1–6 (Table 2) were adjusted to a final concentration of 1 mM in 50 mM Tris-HCl (pH 7.0), 5 mM DTT, and a 20% (v / v) appropriate enzyme dilution series. Reaction components, including ATP / AMP, MgCl2, and polyphosphate (Madrell's salt), were added during incubation according to Table 3. Otherwise, the remaining reaction volume consisted of 50 mM Tris-HCl (pH 7.0). Reactions were prepared in BioRad PCR plates and incubated at 30°C for 24 hours, followed by immediate heat shock at 95°C for 20 minutes to inactivate the enzyme. The inactivated reactions were then diluted 400-fold with 10 mM EDTA (pH 7.0) and analyzed by HPLC as described below.

[0373] [Table 4]

[0374] HPLC analysis of ligation reactions Ligation reactions were analyzed using an IP-RP-HPLC (ion-pairing reversed-phase high-performance liquid chromatography) analytical method focusing on product formation using a Waters® Acquity UHPLC® BEH C18 column (50 x 2.1 mm, 1.7 μm) on a Thermo Scientific UHPLC Vanquish Horizon. Mobile phase A (MPA) consisted of 200 mM HFIP (1,1,1,3,3,3-hexafluoro-2-propanol, Sigma-Aldrich, purity >99%) and 10 mM TEA (triethylamine; Sigma-Aldrich BioUltra purity) in Milli-Q grade water. Mobile phase B (MPB) consisted of methanol (Supelco, HPLC grade). Elution began with 16% MPB for 0.5 min, followed by a linear gradient from 16% MPB to 23% MPB in 3.0 min, followed by a second linear gradient from 23% MPB to 90% MPB in 0.1 min. 90% MPB was maintained for 0.20 min, then changed to 16% MPB at 0.20 min, and maintained at 16% MPB for 0.70 min, resulting in a runtime of 4.7 min. The flow rate was 0.5 mL / min. -1 The column temperature is set to 75° C. Detection is monitored by UV at λ=260 nm.

[0375] Caution when using arbitrary units (AU) to compare ligation activity Ideally, to compare activity between samples and account for natural variations in peak intensity between injections, we calculate the % conversion to product for each sample analyzed using the following formula:

number

number

[0376] [Table 5]

[0377] Example 2 Here, directed evolution of a PPK derived from the bacterium Erysipelotrichaceae is described, resulting in a series of engineered polypeptides with improved activity.

[0378] In this example, % conversion corresponds to the calculated conversion rate for each single sample, expressed as a percentage of ATP relative to the sum of the substrate (AMP), intermediate product (ADP), and product (ATP). The amounts of AMP, ADP, and ATP in each single sample are quantified by HPLC.

[0379] Preparation of isolated enzymes Polynucleotides encoding polypeptides having phosphotransferase activity were cloned into the pCK110900 vector system (see, e.g., U.S. Patent Application Publication No. 2006 / 0195947A1, incorporated herein by reference in its entirety), and then expressed in E. coli W3110fhuA under the control of the lac promoter. The expression vector also contained a P15a origin of replication and a chloramphenicol (CAM) resistance gene.

[0380] E. coli W3110fhuA cells were transformed with the pCK110900 plasmid, which contains a gene encoding a phosphotransferase. Transformed cells were plated onto lysogeny broth (LB) agar plates containing 1% glucose and 30 μg / mL CAM and grown overnight at 37°C. A single colony was then inoculated into 25 mL of LB supplemented with 30 μg / mL CAM and 1% glucose in a 250 mL baffled shake flask. The culture was grown overnight (16-20 h) and at an optical density (OD ) of 100 kJ / mL in a 37°C incubator with shaking at 250 rpm. 600 )>3.8) was grown. 5 mL of the overnight culture was inoculated into a 1 L shake flask containing 250 mL of Terrific Broth (TB) containing 30 μg / mL of CAM. The 250 mL culture was incubated at 30°C, 250 rpm, and OD 600The culture was incubated for 3–3.5 hours until the pH reached 0.6–0.8. Expression of the oxynitrilase gene was induced by adding isopropyl-β-D-thiogalactoside (IPTG) to a final concentration of 1 mM, and growth was continued for an additional 18–20 hours. Cells were harvested by transferring the culture to a centrifuge bottle, which was then centrifuged at 7,000 rpm for 5 minutes at 4°C. The supernatant was discarded, and the remaining cell pellet was lysed. For lysis, the cell pellet was resuspended in 30 mL of 50 mM Tris buffer, pH 7.5, and lysed using an LM20 MICROFLUIDIZER® Processor System (Microfluidics). Cell debris was removed by centrifugation at 14,000 rpm for 30 minutes at 4°C. The phosphotransferase enzyme was then isolated from the clarified lysate using standard techniques known in the art, including immobilized metal affinity chromatography.

[0381] Example 3 Identification of phosphotransferase activity for ATP generation To identify a single enzyme with phosphotransferase activity for ATP production, we first screened a set of phosphotransferases described in the literature. Screening of isolated phosphotransferases was performed in 200 μL reaction volumes in 1.5 mL Eppendorf tubes, each containing 100 mM Tris buffer (pH 7.5), 100 mM MgCl2, 100 mM polyphosphate (Madrell's salt), and 2 mM AMP (1), with 50% (v / v) of the isolated phosphotransferase enzyme. Reactions were incubated at 30°C and 900 rpm for 1 hour and analyzed using standard techniques known in the art, including HPLC. The phosphotransferase of SEQ ID NO:52 (PPK12) showed the highest activity for ATP formation. The activity of SEQ ID NO:52 for ATP production was subsequently confirmed using multiple enzyme preparations, including the isolated enzyme (Example 2), clarified lysate (Example 5), and shake-flask powder (SFP; Example 1).

[0382] Example 4 Preparation of cell pellets for HTP screening Single colonies were picked in a 96-well format and grown in 190 μL of LB medium containing 1% glucose and 30 μg / mL CAM at 30°C, 200 rpm, and 85% humidity. After overnight growth, 20 μL of the grown culture was transferred to a deep-well plate containing 380 μL of TB medium with 30 μg / mL CAM. The culture was grown at 30°C, 250 rpm, and 85% humidity for approximately 2.5 hours. The OD of the culture was 600 When the RI reached 0.4–0.8, expression of the ligase gene was induced by adding IPTG to a final concentration of 1 mM. After induction, growth continued for 18–20 h at 30°C, 250 rpm, and 85% humidity. Cells were harvested by centrifugation at 4,000 rpm and 4°C for 10 min, and the supernatant was then discarded. The cell pellet was stored at -80°C until ready for use.

[0383] Example 5 Lysis and clarification lysate preparation Prior to performing the assay, the cell pellets were thawed and resuspended in 300 μL of lysis buffer (1 g / L lysozyme, 0.5 g / L PMBS, and 0.1 μL / mL or 0.2 U / mL commercial DNAse (New England BioLabs, M0303L) in 50 mM Tris buffer, pH 7.5). The plates were shaken and agitated at medium speed on a microtiter plate shaker for 2.5 hours at room temperature. The plates were then centrifuged at 4,000 rpm for 10 minutes at 4°C, and the clarified supernatant was used in the HTP assay reaction for activity determination, as described in the Examples below.

[0384] Example 6 Analytical methods for activity and selectivity evaluation The improved activity of the engineered phosphotransferase was analyzed by high-pressure liquid chromatography (HPLC) using the method described in Table 5. An HPLC method with UV detection was developed to analyze the formation of the product ATP and separate the substrate AMP (1) and the intermediate ADP. The analytical method aimed for the shortest run time possible to allow good separation of the three compounds. The percentage of conversion was calculated based on the peak area of ​​each compound according to the following formula:

number

[0385] [Table 6]

[0386] The methods provided herein are useful for analyzing variants generated using the present invention, however, it is not intended that the present invention be limited to the methods described herein, as there are other suitable methods known in the art that are applicable to analyzing variants provided herein and / or generated using the methods provided herein.

[0387] Example 7 Round 1 evolution and screening of modified polypeptides derived from SEQ ID NO:52 for improved ATP product generation A modified polynucleotide (SEQ ID NO:48) encoding a polypeptide having oxynitrilase activity of SEQ ID NO:52 was used to generate the modified polypeptides in Table 6. These polypeptides exhibited improved phosphotransferase activity under desired conditions, e.g., improved formation of ATP compared to the starting polypeptide. Modified polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the "backbone" amino acid sequence of SEQ ID NO:52 as described below, along with the analytical methods described in Table 5.

[0388] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 48. Libraries of modified polypeptides were generated using a variety of well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using the HTP assay and analytical methods described below, which measure the ability of the polypeptides to generate ATP.

[0389] Enzyme assays were performed in a 96-well PCR plate with a total reaction volume of 80 μL per well. Reactions contained 0.0025% (v / v) undiluted phosphotransferase lysate prepared as described in Example 4, 0.3 mM AMP, 30 mM polyphosphate (Madrell's salt), 50 mM Tris buffer, pH 8.0, 5 mM MgCl, and 5 mM DTT. Reaction plates were heat-sealed and incubated in a thermocycler at 30° C. for 30 minutes.

[0390] After incubation, the reaction was quenched by adding 120 μL of 10 mM EDTA solution to each well of the plate. The plate was then centrifuged at 4,000 rpm for 5 minutes. A 150 μL aliquot of the supernatant was then removed from each well and added to a shallow-well 96-well plate. The samples were analyzed by HPLC to determine the activity of the enzyme variants using the analytical method described in Table 5. Selected phosphotransferase (PPK12) variants that exhibit improved ATP formation compared to SEQ ID NO: 52 are shown in Table 6. The level of increased activity was determined as the average of two replicates.

[0391] [Table 7]

[0392] Example 8 Identification of improved dsRNA ligation activity under ATP-regenerating conditions with increased polyphosphate concentrations ATP is consumed during the ligation reaction. In the presence of PPK, polyphosphate is consumed to generate ATP from AMP. Therefore, a high concentration of polyphosphate is required to promote ligation reactions at higher substrate concentrations under ATP-regenerating conditions. To process 6 mM of substrate oligonucleotide (Substrates 1–6, Table 2), either a stoichiometric amount of ATP (24 mM) or substoichiometric amounts of ATP and stoichiometric amounts of polyphosphate (48 mM) are required. However, increasing the polyphosphate concentration adversely affects the ligation reaction (Figure 9).

[0393] Therefore, to identify an enzyme with improved dsRNA ligation activity under ATP-regenerating conditions that can tolerate higher polyphosphate concentrations compared to the ligase of SEQ ID NO:2, we screened a collection of ligases in the presence of AMP, polyphosphate, and the optimized PPK of SEQ ID NO:62.

[0394] Enzyme assays were performed in a 96-well PCR plate with a total reaction volume of 100 μL per well. Reactions contained 10% (v / v) undiluted ligase lysate prepared as described in Example 5, 1 mM (each) substrate oligonucleotide (Substrates 1–6, Table 2), 50 mM Tris buffer, pH 7.0, 1 mM AMP, 5 mM MgCl2, 5 mM DTT, 30 mM polyphosphate (Madrell's salt), and 1 g / L PPK SFP of SEQ ID NO: 62 (prepared as described in Example 1). Reaction plates were heat-sealed and incubated at 30°C in a thermocycler for 24 hours.

[0395] After incubation, the plate was subjected to a heat inactivation step (95°C, 20 minutes) to quench the reaction and precipitate the proteinaceous contents of the added lysate. The plate was then centrifuged at 4,000 rpm for 5 minutes. Subsequently, a 50 μL aliquot of the supernatant was removed from each well and added to a deep-well 96-well plate containing 950 μL of 10 mM EDTA solution (pH 7.0). The samples were further diluted by transferring 50 μL of the diluted sample to a deep-well 96-well plate containing 950 μL of 10 mM EDTA solution (pH 7.0). The samples were analyzed by HPLC to determine the activity of the ligase variants using the analytical method described in Example 1.

[0396] The ligase of SEQ ID NO: 88 was identified as having higher ligation activity under the desired reaction conditions compared to the ligase of SEQ ID NO: 2. The higher activity of the ligase of SEQ ID NO: 88 was confirmed using SFP (Example 1) with both 1 mM and 3 mM substrate oligonucleotide (Figure 10).

[0397] Example 9 Optimization of ATP regeneration conditions Polynucleotide sequence SEQ ID NO:87, encoding the best-performing polypeptide of SEQ ID NO:88 with dsRNA ligation activity in the presence of high polyphosphate concentrations, was cloned in place of the ligase polynucleotide of SEQ ID NO:31 at the C-terminus of Cligase 4 polynucleotide construct SEQ ID NO:40 to generate novel polynucleotide SEQ ID NO:89, encoding a novel fusion polypeptide designated Cligase 4.2 (SEQ ID NO:90). To assess both the dsRNA ligation and phosphotransferase activities of Cligase 4.2, ligation reactions were set up under various ATP regeneration conditions.

[0398] Enzyme assays were performed in a 96-well PCR plate with a total reaction volume of 100 μL per well. Reactions contained 1 mM of each substrate oligonucleotide (Substrates 1–6, Table 2), 1 mM AMP, 5 mM DTT, either 50 mM Tris buffer, pH 7.0, 100 mM Tris buffer, pH 7.0, or 100 mM MOPS buffer, 5 mM, 20 mM, 40 mM, 60 mM, 80 mM, or 100 mM MgCl2, and either 0% (v / v) or 10% (v / v) MgCl2. The reaction plates contained either DMSO, 30 mM polyphosphate (Graham's salt) or either 30 mM, 50 mM, or 80 mM polyphosphate (Madrell's salt), and either 0 g / L, 0.625 g / L, 1.25 g / L, 2.5 g / L, 5 g / L, 10 g / L, 20 g / L, or 40 g / L of Clignase 4.2 (SEQ ID NO: 90) SFP (prepared as described in Example 1). The reaction plates were heat-sealed and incubated at 30°C in a thermocycler for 24 hours.

[0399] After incubation, the plate was subjected to a heat inactivation step (95°C, 20 minutes) to quench the reaction and precipitate the proteinaceous contents of the added SFP. The plate was then centrifuged at 4,000 rpm for 5 minutes. Subsequently, a 50 μL aliquot of the supernatant was removed from each well and added to a deep-well 96-well plate containing 950 μL of 10 mM EDTA solution (pH 7.0). The samples were further diluted by transferring 50 μL of the diluted sample to a deep-well 96-well plate containing 950 μL of 10 mM EDTA solution (pH 7.0). The samples were analyzed by HPLC to determine the activity of the enzyme variants using the analytical method described in Example 1.

[0400] Cligase 4.2 was shown to accept both Maddrell's and Graham's salts of polyphosphate (Fig. 11a). Increasing the MgCl concentration from 5 mM to ≥ 20 mM significantly improved dsRNA ligation activity (Fig. 11b).

[0401] Example 10 Round 1 evolution and screening of modified polypeptides derived from SEQ ID NO: 90 for improved siRNA production under ATP regenerating conditions A modified polynucleotide of SEQ ID NO: 89 encoding a polypeptide having dsRNA ligation activity and phosphotransferase activity of SEQ ID NO: 90 was used to generate a modified polypeptide that exhibits improved dsRNA ligase activity under ATP regenerating conditions.

[0402] Modified polypeptides having amino acid sequences of even-numbered sequence identifiers were generated from the "backbone" amino acid sequence of SEQ ID NO:90, as described below, along with the analytical methods described in Example 1. Directed evolution began with the polynucleotide set forth in SEQ ID NO:89. Libraries of modified polypeptides were generated using a variety of well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using the HTP assay and analytical methods described below, which measure the ability of the polypeptides to generate siRNA products.

[0403] Enzyme assays were performed in 96-well PCR plates with a total reaction volume of 100 μL per well. Reactions contained either 10% (v / v) or 20% (v / v) undiluted ligase lysate prepared as described in Example 5, 2.5 mM (each) of substrate oligonucleotide (Substrates 1–6, Table 2), 50 mM Tris buffer, pH 7.0, 1 mM AMP, 40 mM MgCl2, 5 mM DTT, and 80 mM polyphosphate (Madrell's salt). Reaction plates were heat-sealed and incubated at 30°C in a thermocycler for 24 hours.

[0404] After incubation, the plate was subjected to a heat inactivation step (95°C, 20 minutes) to quench the reaction and precipitate the proteinaceous contents of the added lysate. The plate was then centrifuged at 4,000 rpm for 5 minutes. Subsequently, a 50 μL aliquot of the supernatant was removed from each well and added to a deep-well 96-well plate containing 950 μL of 10 mM EDTA solution (pH 7.0). The samples were further diluted by transferring 50 μL of the diluted sample to a deep-well 96-well plate containing 950 μL of 10 mM EDTA solution (pH 7.0). The samples were then analyzed by HPLC to determine the activity of the Klignase variants using the analytical method described in Example 1.

[0405] Lysates containing modified polypeptides having SEQ ID NOs: 92, 94, 96, and 98 were identified as having increased dsRNA ligation activity under ATP-regenerating conditions compared to the ligase of SEQ ID NO: 90. The higher activity of the clignases of SEQ ID NOs: 92, 94, 96, and 98 was confirmed using SFPs prepared as described in Example 1 (FIG. 12).

[0406] Example 11 Comparison of catalytic activity of modified polypeptides SEQ ID NO: 98 and SEQ ID NO: 11 To further validate the improved dsRNA ligase activity of the novel Cligase of SEQ ID NO:98 compared to one of the best-performing initial Cligase constructs, Cligase 4 of SEQ ID NO:11, two polypeptides were produced as SFPs as described in Example 1 and evaluated for their ability to convert 6 mM of each substrate oligonucleotide (Substrates 1-6, Table 2) in the presence of 0.25 mM AMP, 40 mM MgCl, 80 mM polyphosphate (Madrell's salt), 5 mM DTT, and 100 mM MOPS buffer (pH 7.2). Enzyme assays were performed in 96-well PCR plates with a total reaction volume of 100 μL per well and contained either 0 g / L, 0.313 g / L, 0.625 g / L, 1.25 g / L, 2.5 g / L, 5 g / L, 10 g / L, or 20 g / L of Cligase. The reaction plate was heat sealed and incubated in a thermocycler at 30°C for 24 hours.

[0407] After incubation, the plates were subjected to a heat inactivation step (95°C, 20 minutes) to quench the reaction and precipitate the proteinaceous contents of the added SFPs. The plates were then centrifuged at 4,000 rpm for 5 minutes. A 50 μL aliquot of the supernatant was removed from each well of each plate and subsequently diluted in 950 μL of 10 mM EDTA solution. The samples were further diluted by transferring 20 μL of the diluted sample to a deep-well 96-well plate containing 180 μL of 10 mM EDTA solution (pH 7.0). The samples were further diluted by transferring 18 μL of the diluted sample to a deep-well 96-well plate containing 198 μL of 10 mM EDTA solution (pH 7.0). The samples were analyzed via HPLC using the analytical method described in Example 1.

[0408] The comparative data in Figure 13 shows that the clignase of SEQ ID NO:98 exhibits improved ligation activity under ATP regenerating conditions compared to the clignase of SEQ ID NO:11.

[0409] array >Amino acid sequence of bacteriophage RB69 RNA ligase 2 (UniProt ID: Q7Y4V8) [ka] > Optimized bacteriophage RB69 RNA ligase 2 amino acid sequence [ka] > Amino acid sequence of bacteriophage T4 RNA ligase 2 (Uniprot ID: P32277) [ka] >Amino acid sequence of bacteriophage T4 DNA ligase (UniProt ID: P00970) [ka] >PPK12 amino acid sequence [ka] > Optimized PPK12 amino acid sequence [ka] >Amino acid sequence of AjPAP (UniProt ID: Q83XD3) [ka] >Clignase 1: Amino acid sequence of kinase-His6-TEV-ligase [ka] >Clignase 2: Amino acid sequence of kinase-His10-ligase [ka] >Clignase 3: Amino acid sequence of His6-TEV-kinase-SSGSSG-ligase [ka] >Cligase 4: Amino acid sequence of His6-TEV-kinase-GSAGSAAGSGEF-ligase [ka] >Cligase 5: Amino acid sequence of His6-TEV-kinase-GSSGSGSSSGGSSSSGSS-ligase [ka] >Cligase 6: His6-TEV-ligase-GSAGSAAGSGEF-kinase amino acid sequence [ka] >Clignase 7: His6-TEV-ligase-GSSGSGSSSGGSSSSGSS-kinase amino acid sequence [ka] >Cligase 8: Amino acid sequence of ligase-GSAGSAAGSGEF-kinase-His6 [ka] >Cligase 9: Amino acid sequence of ligase-GSSGSGSSSGGSSSSGSS-kinase-His6 [ka] >Cligase 10: Amino acid sequence of kinase-His6-SSGSSG-His6-TEV-ligase [ka] >Cligase 11: Amino acid sequence of kinase-SSGSSG-His6-SSGSSG-TEV-ligase [ka] >Linker amino acid sequence HHHHHH (SEQ ID NO: 19) >Linker amino acid sequence HHHHHHHHHH (SEQ ID NO: 20) >Linker amino acid sequence ENLYFQS (SEQ ID NO: 21) >Linker amino acid sequence ENLYFQG (SEQ ID NO: 22) >Linker amino acid sequence SSGSSG (SEQ ID NO: 23) >Linker amino acid sequence GSAGSAAGSGEF (SEQ ID NO: 24) >Linker amino acid sequence GSSGSGSSSGGSSSSGSS (SEQ ID NO: 25) >Linker amino acid sequence MHHHHHHENLYFQS (SEQ ID NO: 26) >Linker amino acid sequence GQTGHHHHHH (SEQ ID NO: 27) >Linker amino acid sequence EQKLISEEDL (SEQ ID NO: 28) >Linker amino acid sequence DYKDDDDK (SEQ ID NO: 29) >A nucleic acid sequence encoding SEQ ID NO: 1 [ka] >Nucleic acid sequence encoding SEQ ID NO: 2 [ka] >A nucleic acid sequence encoding SEQ ID NO:3 [ka] >A nucleic acid sequence encoding SEQ ID NO: 4 [ka] >A nucleic acid sequence encoding SEQ ID NO: 5 [ka] >A nucleic acid sequence encoding SEQ ID NO:6 [ka] >A nucleic acid sequence encoding SEQ ID NO: 7 [ka] >Nucleic acid sequence encoding SEQ ID NO: 8 [ka] >A nucleic acid sequence encoding SEQ ID NO: 9 [ka] >A nucleic acid sequence encoding SEQ ID NO: 10 [ka] >Nucleic acid sequence encoding SEQ ID NO: 11 [ka] >Nucleic acid sequence encoding SEQ ID NO: 12 [ka] >Nucleic acid sequence encoding SEQ ID NO: 13 [ka] >Nucleic acid sequence encoding SEQ ID NO: 14 [ka] >A nucleic acid sequence encoding SEQ ID NO: 15 [ka] >Nucleic acid sequence encoding SEQ ID NO: 16 [ka] >A nucleic acid sequence encoding SEQ ID NO: 17 [ka] >Nucleic acid sequence encoding SEQ ID NO: 18 [ka] >Nucleic acid sequence encoding SEQ ID NO: 52 [ka] > Nucleic acid sequence encoding optimized PPK12 with a C-terminal tag (GQTGHHHHHH; SEQ ID NO: 27) [ka] >A nucleic acid encoding a dsRNA ligase having an N-terminal tag (MHHHHHHENLYFQS; SEQ ID NO: 26) [ka] > Nucleic acid sequence encoding optimized dsRNA ligase with an N-terminal tag (MHHHHHHENLYFQS; SEQ ID NO: 26) [ka] Polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 54 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 56 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 58 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 60 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 62 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 64 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 66 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 68 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >A nucleic acid sequence encoding SEQ ID NO: 70 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 72 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 74 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 76 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 78 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >A nucleic acid sequence encoding SEQ ID NO: 80 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 82 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 84 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 86 [ka] > Engineered variants of polyphosphate-nucleotide phosphotransferase from Erysipelotrichaceae bacteria [ka] >Nucleic acid sequence encoding SEQ ID NO: 88 [ka] > Optimized bacteriophage RB69 RNA ligase 2 amino acid sequence [ka] >Nucleic acid sequence encoding SEQ ID NO: 90 [ka] > Amino acid sequence of Clignase 4.2 [ka] >Nucleic acid sequence encoding SEQ ID NO: 92 [ka] > Amino acid sequence of Clignase 4.2.1 [ka] >Nucleic acid sequence encoding SEQ ID NO: 94 [ka] > Amino acid sequence of Clignase 4.2.2 [ka] >Nucleic acid sequence encoding SEQ ID NO: 96 [ka] > Amino acid sequence of Clignase 4.2.3 [ka] >Nucleic acid sequence encoding SEQ ID NO: 98 [ka] > Amino acid sequence of Clignase 4.2.4 [ka]

Claims

1. 1. A method for generating an oligonucleotide from two or more oligonucleotide fragments, comprising: i. two or more oligonucleotide fragments; ii. ATP-dependent nucleic acid ligase; iii. Polyphosphate kinase (PPK); iv. adenosine triphosphate (ATP) and / or adenosine monophosphate (AMP); v. polyphosphate; and vi. a divalent cation; contacting the thereby providing an oligonucleotide.

2. Use of an ATP-dependent nucleic acid ligase and a PPK in generating oligonucleotides from two or more oligonucleotide fragments.

3. (a) the two or more oligonucleotide fragments comprise two or more RNA oligonucleotide fragments; optionally the ATP-dependent nucleic acid ligase is an RNA ligase; and optionally: (i) the RNA ligase is a double-stranded RNA ligase; and / or (ii) the RNA ligase is a member of the RNA ligase 2 family, and optionally the RNA ligase is bacteriophage RB69 RNA ligase 2; (b) the two or more oligonucleotide fragments comprise two or more DNA oligonucleotide fragments; optionally the ATP-dependent nucleic acid ligase is a DNA ligase; optionally the DNA ligase is T4 DNA ligase; and / or (c) the PPK is PPK12 or ajPAP,

4. (A) the ATP-dependent nucleic acid ligase and the PPK are linked, optionally the ATP-dependent nucleic acid ligase and the PPK are linked via a polypeptide linker; and optionally: (i) the PPK is located at the N-terminus of the linker and the ATP-dependent nucleic acid ligase is located at the C-terminus of the linker; and / or (ii) the linker is a polypeptide linker comprising at least three amino acids, optionally at least six amino acids; optionally the linker is: a) HHHHHH (SEQ ID NO: 19), optionally HHHHHHHHHHH (SEQ ID NO: 20); b) ENLYFQS (SEQ ID NO: 21); c) ENLYFQG (SEQ ID NO: 22); d) SSGSSG (SEQ ID NO: 23); e) GSAGSAAGSGEF (SEQ ID NO: 24); and / or f) GSSGSGSSGGSSSSGSS (SEQ ID NO: 25); comprising an amino acid sequence selected from and / or (B) (i) the ATP-dependent nucleic acid ligase comprises a purification tag; (ii) the PPK comprises a purification tag; and / or (iii) the linker comprises a purification tag; 4. The method according to claim 1 or 3 or the use according to claim 2 or 3.

5. (a) the polyphosphate is a polyphosphate, optionally the polyphosphate is sodium polyphosphate (Madrell's salt) or sodium hexametaphosphate (Graham's salt); and / or (b) the divalent cation cofactor is Mg 2+ or Mn 2+ and / or (c) the method is carried out at a divalent cation concentration of 5 to 100 mM, optionally 30 to 50 mM; and / or (d) the method is carried out using substoichiometric concentrations of ATP and / or AMP; and / or (e) the method further comprises purifying the oligonucleotide.

5. The method of any one of claims 1, 3 or 4.

6. (A) said oligonucleotides are up to 60 nucleotides in length; optionally each of said oligonucleotide fragments is 4 to 16 nucleotides in length, optionally 6 to 9 nucleotides in length; and / or (B) the oligonucleotide fragment is: (a) single stranded; or (b) are double-stranded, and optionally one or more of the double-stranded oligonucleotide fragments contain one or two single-stranded overhangs; and / or (C) one or more of said oligonucleotide fragments comprises a chemical modification; optionally said chemical modification is: (a) optionally a modified backbone selected from phosphorothioate (e.g., chiral phosphorothioate) or methylphosphonate internucleotide linkages; (b) optionally 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), 2'-O-N-methylacetamide (2'-O- modified nucleotides selected from: NMA), locked nucleic acids (LNA), glycol nucleic acids (GNA), phosphoramidates (e.g., mesyl phosphoramidates), 2',3'-seconucleotide mimics, 2'-F-arabinonucleotides, abasic nucleotides, 2'-amino modified nucleotides, 2'-alkyl modified nucleotides, morpholino nucleotides, vinyl phosphonates (e.g., 5' vinyl phosphonate), and cyclopropyl phosphonate deoxyribonucleotides; and / or (c) conjugation to a ligand, optionally wherein the ligand comprises one or more N-acetylgalactosamine (GalNAc) derivatives. and / or (D) the ATP-dependent nucleic acid ligase and / or the PPK is immobilized; optionally, the ATP-dependent nucleic acid ligase and / or the PPK is immobilized on a solid material by chemical bonding or physical adsorption. A method according to any one of claims 1 or 3 to 5 or a use according to any one of claims 2 to 4.

7. i. ATP-dependent nucleic acid ligase; ii. PPK; iii. ATP and / or AMP; iv. a divalent cation; and v. polyphosphate; and optionally further comprising two or more oligonucleotide fragments.

8. i. ATP-dependent nucleic acid ligase; ii. PPK; iii. ATP and / or AMP; iv. polyphosphate; v. divalent cations; and vi. Instructions for use in a method for generating an oligonucleotide from two or more oligonucleotide fragments Kit including:

9. (a) the polyphosphate is a polyphosphate; optionally the polyphosphate is selected from Graham's salt and Maddrell's salt; and / or (b) the divalent cation is Mg 2+ or Mn 2+ and / or (c) the concentration of the divalent cation is 5 to 100 mM, optionally 30 to 50 mM; The composition of claim 7 or the kit of claim 8.

10. a) a PPK domain; and b) ATP-dependent nucleic acid ligase domain A fusion polypeptide comprising:

11. (A) the PPK is PPK12 or ajPAP; and / or (B) the PPK domain comprises an amino acid sequence having at least 85% identity to the amino acid sequence of any one of SEQ ID NOs: 5-7; and / or (C) the ATP-dependent nucleic acid ligase domain comprises: (i) an RNA ligase domain; optionally the RNA ligase domain is a double-stranded RNA (dsRNA) ligase domain; and / or the dsRNA ligase is a member of the RNA ligase 2 family, optionally the dsRNA ligase is bacteriophage RB69 RNA ligase 2; or (ii) a DNA ligase domain; optionally, the DNA ligase domain is a T4 DNA ligase domain; and / or (D) the ATP-dependent nucleic acid ligase domain comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 1-4 or 88; and / or (E) the fusion polypeptide comprises a purification tag, optionally located at the N-terminus and / or C-terminus of the fusion polypeptide; and / or (F) the fusion polypeptide comprises an amino acid sequence having at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 8-18, 90, 92, 94, 96, or 98; and / or (G) the fusion polypeptide comprises a linker; optionally: (i) the linker is located between the PPK domain and the ATP-dependent nucleic acid ligase domain; and / or (ii) the PPK domain is located at the N-terminus of the linker and the ATP-dependent nucleic acid ligase domain is located at the C-terminus of the linker; and / or (iii) the linker comprises a purification tag; optionally, the purification tag is located at the N-terminus and / or C-terminus of the fusion polypeptide; and / or (iv) the linker is a polypeptide linker comprising at least 3 amino acids, optionally at least 6 amino acids, and optionally the linker is: a) HHHHHH (SEQ ID NO: 19), optionally HHHHHHHHHHH (SEQ ID NO: 20); b) ENLYFQS (SEQ ID NO: 21); c) ENLYFQG (SEQ ID NO: 22); d) SSGSSG (SEQ ID NO: 23); e) GSAGSAAGSGEF (SEQ ID NO: 24); and / or f) GSSGSGSSGGSSSSGSS (SEQ ID NO: 25) The fusion polypeptide of claim 10, comprising an amino acid sequence selected from:

12. 12. The method of any one of claims 1 or 3 to 6 or the use of any one of claims 2 to 4 or 6, wherein the ATP-dependent nucleic acid ligase and the PPK are provided as a fusion polypeptide as defined in any one of claims 10 or 11.

13. 12. A nucleic acid molecule encoding a fusion polypeptide according to claim 10 or claim 11, optionally comprising: (a) any one of SEQ ID NOs: 34-36; and / or (b) any one of SEQ ID NOs: 30 to 33 or 87 A nucleic acid molecule comprising a nucleic acid sequence having at least 85% sequence identity with the nucleic acid sequence of

14. 14. A vector comprising the nucleic acid of claim 13, optionally wherein the vector is selected from a plasmid, cosmid, bacteriophage or viral vector.

15. 15. A host cell comprising the nucleic acid molecule of claim 13 or the vector of claim 14, optionally wherein the host cell is E. coli.

16. (a) an ATP-dependent nucleic acid ligation reaction; optionally, the rate of nucleic acid ligation exceeds the rate of nucleic acid ligation of a control; said control being: (i) a first protein consisting of the PPK domain of claim 10; and (ii) a second protein consisting of the ATP-dependent nucleic acid ligase domain of claim 10; Including, the first and second proteins are unlinked; or (b) A method for generating an oligonucleotide from two or more oligonucleotide fragments 12. Use of a fusion polypeptide according to claim 10 or claim 11 in

17. (a) the oligonucleotide is a therapeutic oligonucleotide; and / or (b) the oligonucleotide product is at least 80% pure, optionally the oligonucleotide product is at least 85% pure, at least 90% pure, at least 95% pure, optionally the oligonucleotide product is at least 98% pure; A method according to any one of claims 1, 3 to 6 or 12 or a use according to any one of claims 2 to 4, 6, 12 or 16.