Novel method for producing oligonucleotides

By providing a method of complementary template oligonucleotide and oligonucleotide collection, combined with enzymatic or chemical ligation, the problems of high purification costs and equipment scale limitations in oligonucleotide synthesis are solved, and a high-efficiency and large-scale preparation of chemically modified oligonucleotides, especially therapeutic oligonucleotides, are achieved.

CN115161366BActive Publication Date: 2025-08-05GLAXOSMITHKLINE INTPROP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210531352.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-07-11
Filing Date
2017-07-07
Publication Date
2025-08-05
Estimated Expiration
2037-07-07

AI Technical Summary

Technical Problem

The existing oligonucleotide synthesis methods have problems such as high chromatography purification cost, low yield, limited equipment scale, and unsuitable DNA polymerase for the synthesis of therapeutic oligonucleotides, which are especially difficult to achieve during large-scale preparation.

Method used

Using a method, it includes providing a collection of template oligonucleotides and oligonucleotides complementary to the product sequence, denaturing the template and impurity oligonucleotide strands by changing the conditions after contact under annealing conditions and separating impurities, and finally isolating the product, using enzymatic or chemical ligation to form the oligonucleotide product.

Benefits of technology

Efficient purification and large-scale preparation of chemically modified oligonucleotides, especially therapeutic oligonucleotides, have improved yields and overcome equipment size limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115161366B_ABST
    Figure CN115161366B_ABST
Patent Text Reader

Abstract

The present invention relates to a novel method for producing oligonucleotides, which is suitable for producing chemically modified oligonucleotides, such as those used in therapy. Disclosed herein is a novel method for producing oligonucleotides, which is suitable for producing chemically modified oligonucleotides, such as those used in therapy, comprising: a) providing a template oligonucleotide (I) complementary to the sequence of the product, the template having properties allowing it to be separated from the product; b) providing a collection of oligonucleotides (II); c) contacting (I) and (II) under conditions that allow annealing; d) changing the conditions to separate any impurities, including denaturing the annealed template and impurity oligonucleotide chains and separating the impurities; and e) changing the conditions to separate the product, including denaturing the annealed template and product oligonucleotide chains and separating the product.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the PCT application with an international application date of July 7, 2017, international application number PCT / EP2017 / 067049, application number 201780043407.4 entering the Chinese national phase, and invention name “Novel method for producing oligonucleotides”. Technical Field

[0002] The present invention relates to novel methods for producing oligonucleotides, which are suitable for producing chemically modified oligonucleotides, such as those used in therapy. Background of the Invention

[0003] The chemosynthesis of oligonucleotide and oligonucleotide via phosphoramidite chemical modification are fully determined, and have become the candidate method for synthesizing these determined sequence biopolymers for decades.Described synthetic method is often run as solid phase synthesis, wherein various Nucleotide are added successively, and the interpolation of every kind of Nucleotide needs the circulation of several chemical steps to add and deprotect the oligonucleotide (" oligos ") in growing in the preparation of subsequent steps.When the continuous interpolation of Nucleotide ends, described oligos is discharged from solid support, further deprotection occurs, and then rough oligonucleotide is further purified by column chromatography.

[0004] Although this method can be considered routine and can be automated, it has several disadvantages, particularly when the goal is to produce oligonucleotides on a large scale, such as is required for oligonucleotide therapeutics. These disadvantages include, but are not limited to:

[0005] 1) Practical limitations inherent in the application of chromatography make it unsuitable for purification of large quantities of oligonucleotides. Large-scale application of chromatography is expensive and difficult to achieve due to limitations on column size and performance.

[0006] 2) The number of errors accumulates with the length of the synthesized oligonucleotide. Therefore, the linear, continuous nature of the current method leads to a geometric decrease in yield. For example, if the yield per cycle of nucleotide addition is 99%, then the yield of a 20-mer would be 83%.

[0007] 3) Scale limitations of oligonucleotide synthesizers and downstream purification and separation equipment: Currently, the maximum amount of product that can be produced in a single batch is on the order of 10 kg.

[0008] Therefore, there is a need to reduce (or ideally eliminate) column chromatography and to perform syntheses in a non-fully continuous manner in order to increase yields.

[0009] DNA polymerase is often used to synthesize oligonucleotides used in molecular biology and similar applications. However, due to the relatively short oligonucleotide length and the need to distinguish between nucleotides with different deoxyribose or ribose modifications, DNA polymerase is not suitable for synthesizing therapeutic oligonucleotides. For example, therapeutic oligonucleotides are often within the range of 20-25 nucleotides. DNA polymerase requires at least 7 or 8 nucleotides and optimally 18-22 nucleotides as primers in each direction, so if the primer is similar in size to the desired product, little is gained in trying to synthesize therapeutic oligos. In addition, DNA polymerase requires all nucleotides to be present in the reactant, and it relies on Watson-Crick base pairing to compare new nucleotides. Thus, it is not possible to distinguish between any ordering of deoxyribose or ribose modifications, such as those required by gapmers, and the result will be a mixture of deoxyribose or ribose modifications at a given position. SUMMARY OF THE INVENTION

[0010] The present invention provides a method for producing a single-stranded oligonucleotide product having at least one modified nucleotide residue, the method comprising:

[0011] a) providing a template oligonucleotide (I) complementary to the sequence of said product, said template having properties allowing its separation from said product;

[0012] b) providing a collection of oligonucleotides (II);

[0013] c) contacting (I) and (II) under conditions allowing annealing;

[0014] d) changing the conditions to separate any impurities, including denaturing the annealed template and impurity oligonucleotide chains and separating the impurities; and

[0015] e) changing the conditions to separate the products, including denaturing the annealed template and product oligonucleotide chains and separating the products.

[0016] Such methods can be used to separate single-stranded oligonucleotide products from impurities, for example as a purification method.

[0017] The present invention further provides a method for producing a single-stranded oligonucleotide product having at least one modified nucleotide residue, the method comprising:

[0018] a) providing a template oligonucleotide (I) complementary to the sequence of said product, said template having properties allowing its separation from said product;

[0019] b) providing a collection of oligonucleotides (II) comprising oligonucleotides that are segments of the product sequence, wherein at least one segment comprises at least one modified nucleotide residue;

[0020] c) contacting (I) and (II) under conditions allowing annealing;

[0021] d) ligating the segment oligonucleotides to form the product;

[0022] e) changing the conditions to separate any impurities, including denaturing the annealed template and impurity oligonucleotide chains and separating the impurities; and

[0023] f) changing the conditions to separate the products, including denaturing the annealed template and product oligonucleotide chains and separating the products.

[0024] Such methods can be used to generate a single-stranded oligonucleotide product and separate it from impurities, for example, as a preparation and purification method.

[0025] The invention also includes modified oligonucleotides prepared by such methods and ligases for use in such methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a schematic diagram of the method of the present invention, including steps of changing the conditions to remove impurities, such as a washing step.

[0027] Figure 2 is a schematic diagram of the method of the present invention, including the steps of joining / ligating segment oligonucleotides to form a product and varying the conditions to remove impurities.

[0028] Figure 3 Schematic diagram of various template configurations.

[0029] Figure 4 is a basic schematic diagram of the method of the present invention implemented in a flow system.

[0030] Figure 5 Detailed schematic diagram of the method of the present invention implemented in a flow system: a) connected chemical section, b) connected purification section, and c) alternating connected chemical sections, and d) alternating purification sections. NB Sections a) and b) (alternatively c) and d)) can be performed in a single step, e.g., the collection vessel in a) = the output from the connection step in b).

[0031] Figure 6 are examples of impurities that may be generated in the process of the present invention.

[0032] Figure 7 The chromatogram shows the results of a ligation reaction using commercially available NEB T4 ligase (SEQ ID NO: 3) and unmodified DNA (Example 1).

[0033] Figure 8 The chromatogram shows the results of a ligation reaction using PERLOZA-conjugated T4 ligase (SEQ ID NO: 4) and unmodified DNA (Example 1).

[0034] Figure 9 The chromatograms show the results of ligation reactions using PERLOZA-conjugated T4 ligase (expressed according to Example 2) and 2'-OMe-modified oligonucleotide fragments.

[0035] Figure 10 HPLC traces from Example 4: Upper trace (a) - no ligase control. Lower trace (b) - clone A4. Product and template coelute in this HPLC method. Two intermediate ligation fragments (segments) can be seen in the clone A4 trace at 10.3 and 11.2 minutes.

[0036] Figure 11 is a schematic diagram of the "three-template hub'" used in Example 13, which comprises a support material called "hub'" and three template sequences.

[0037] Figure 12 is a schematic diagram of the semi-continuous connection rig used in Example 13.

[0038] Figure 13 is a schematic diagram of the blind-end filtration equipment used in Example 14.

[0039] Figure 14 is a schematic diagram of the lateral flow filtration equipment used in Example 14.

[0040] Figure 15 The chromatograms show the results from a blind-end filtration experiment using a 10 kDa MWCO NADIR membrane at 60°C in Example 14. Panel a) is the chromatogram of the retentate solution after 2 diafiltration volumes, which remains in the filtration unit and contains primarily the three-template hub; and chromatogram b) belongs to the permeate after 2 diafiltration volumes, i.e., the product-rich solution.

[0041] Figure 16 The chromatograms show the results from a lateral flow filtration experiment using a 5 kDa MWCO Snyder membrane at 50°C and 3.0 bar pressure in Example 14. Panel a) is a chromatogram of the retentate solution after 20 diafiltration volumes, which primarily contains the tri-template hub and product; and panel b) is a chromatogram of the permeate after 20 diafiltration volumes, which primarily contains the segmented oligonucleotides.

[0042] Figure 17The chromatogram shows the results from a lateral flow filtration experiment in Example 14 using a 5 kDa MWCOSnyder membrane at 80°C and 3.1 bar pressure. Figure 17 Shown are a) the chromatogram of the retentate solution after 20 diafiltration volumes, which contains only the tri-template hub; and b) the chromatogram of the permeate solution after 2 diafiltration volumes, which contains only the product. Detailed Description of the Invention

[0043] definition

[0044] As used herein, the term "oligonucleotide" or "oligos" for short refers to a polymer of nucleotide residues, which are deoxyribonucleotides (wherein the oligonucleotide obtained is DNA), ribonucleotides (wherein the oligonucleotide obtained is RNA), or mixtures thereof. An oligonucleotide can be composed entirely of nucleotide residues found in nature, or can contain at least one modified nucleotide or at least one connection between nucleotides. An oligonucleotide can be single-stranded or double-stranded. The oligonucleotides of the present invention can be conjugated to another molecule, such as N-acetylgalactosamine (GalNAc) or its polymers (GalNAc clusters).

[0045] As used herein, the term "therapeutic oligonucleotide" refers to an oligonucleotide with therapeutic use. Such oligonucleotides typically contain one or more modified nucleotide residues or are connected. Therapeutic oligonucleotides work via one of several different mechanisms, including but not limited to antisense, splicing-exchange or exon skipping, immunostimulation, and RNA interference (RNAi), such as via microRNA (miRNA) and small interfering RNA (siRNA). A therapeutic oligonucleotide can be an aptamer. Therapeutic oligonucleotides often, but not always, have a definite sequence.

[0046] As used herein, the term "template" refers to an oligonucleotide whose sequence is 100% complementary to the sequence of a target (or product) oligonucleotide.

[0047] Unless otherwise indicated, the term "complementary" as used herein means 100% complementary.

[0048] As used herein, the term "product" refers to a desired oligonucleotide having a specific sequence, also referred to herein as a "target oligonucleotide."

[0049] As used herein, term " set (pool) " represents one group of oligonucleotide, and it may be different in sequence, may be shorter than or longer than target sequence, and may not have the sequence identical with target sequence.The set of oligonucleotide can be the product of oligonucleotide synthesis (such as solid phase chemical synthesis via phosphoramidite chemistry), and described oligonucleotide synthesis is used with or without purification.The set of oligonucleotide can be made up of the section of target sequence.Each section itself can exist as the set of this section, and can be the product of oligonucleotide synthesis (such as solid phase chemical synthesis via phosphoramidite chemistry).

[0050] As used herein, the term "annealing" refers to the hybridization of complementary oligonucleotides in a sequence-specific manner. "Conditions that allow annealing" will depend on the T values of the complementary oligonucleotides that are hybridized. m , and will be readily apparent to those skilled in the art. For example, the annealing temperature may be lower than the T of the hybridized oligonucleotide. m Alternatively, the annealing temperature can be close to the T of the hybridized oligonucleotide. m , such as ±1, 2 or 3°C. Generally, the annealing temperature is no higher than the T of the hybridized oligonucleotide. m Above 10° C. The specific conditions for annealing are as described in the Examples section.

[0051] As used herein, the term "denaturation" relevant to double-stranded oligonucleotides is used to refer to that the complementary chain no longer anneals. Denaturation occurs as a result of changing conditions, and is sometimes referred to as separating oligonucleotide chains in this article. Such chain separation can be completed, for example, by increasing the temperature, changing the pH, or changing the salt concentration of the buffer solution. Denaturing double-stranded oligonucleotides ("duplexes") can produce single-stranded oligonucleotides, which will be products or impurities "released" from the template.

[0052] Term used herein " impurity " refers to the oligonucleotide that does not have the product sequence of expectation.These oligonucleotides can comprise the oligonucleotide that is shorter than product (for example short 1,2,3,4,5 or more nucleotide residues) or longer than product (for example long 1,2,3,4,5 or more nucleotide residues).In the case that production method is included in the step of forming connection between section, impurity comprises the oligonucleotide that remains when one or more connections fail to form.Impurity also comprises the oligonucleotide that has wherein mixed wrong nucleotide, thereby produces mispairing when contrasting with template.Impurity can have one or more above-mentioned features. Figure 6 Some possible impurities are shown.

[0053] As used herein, the term "segment" is a smaller portion of a longer oligonucleotide, particularly a smaller portion of a product or target oligonucleotide. For a given product, when all of its segments anneal to its template and ligate together, the product is formed.

[0054] As used herein, the term "enzymatic ligation" refers to enzymatically forming a connection between two adjacent nucleotides. The connection can be a naturally occurring phosphodiester bond (PO) or a modified connection including, but not limited to, phosphorothioate (PS) or phosphoramidate (PA).

[0055] The term "ligase" as used herein refers to an enzyme that catalyzes the connection of two oligonucleotide molecules, i.e., covalently linked, such as by forming a phosphodiester bond between the 3' end of an oligonucleotide (or segment) and the 5' end of the same or another oligonucleotide (or segment). These enzymes are often referred to as DNA ligases or RNA ligases and utilize cofactors: ATP (eukaryotic, viral and archaebacterial (archael) DNA ligases) or NAD (prokaryotic DNA ligases). Although they are present in all organisms, DNA ligases show a wide diversity (Nucleic Acids Research, 2000, Vol. 28, No. 21, 4051-4058) of amino acid sequence, molecular size and performance. They are typically members of the enzyme class EC 6.5 defined by the International Union of Biochemistry and Molecular Biology, i.e., ligases for forming phosphoester bonds. Included within the scope of the present invention are ligases capable of ligating an unmodified oligonucleotide to another unmodified oligonucleotide, ligases capable of ligating an unmodified oligonucleotide to a modified oligonucleotide ( Right now Ligases that ligate a modified 5' oligonucleotide to an unmodified 3' oligonucleotide, and ligate an unmodified 5' oligonucleotide to a modified 3' oligonucleotide), and ligases that are capable of ligating a modified oligonucleotide to another modified oligonucleotide.

[0056] As used herein, a "thermostable ligase" is a ligase that is active at elevated temperatures, ie, above human body temperature, ie, above 37° C. A thermostable ligase may be active at, for example, 40° C. to 65° C., or 40° C. to 90° C., etc.

[0057] As used herein, the term "modified nucleotide residue" or "modified oligonucleotide" refers to a nucleotide residue or oligonucleotide that contains at least one aspect of its chemical properties that differs from a naturally occurring nucleotide residue or oligonucleotide. Such modifications may occur at any portion of the nucleotide residue, such as the sugar, base, or phosphate. Examples of modifications of nucleotides are disclosed below.

[0058] As used herein, the term "modified ligase" refers to a ligase that differs from a naturally occurring "wild-type" ligase by one or more amino acid residues. Such ligases do not occur in nature. Such ligases are useful in the novel methods of the present invention. Examples of modified ligases are disclosed below. The terms "modified ligase" and "mutant ligase" are used interchangeably.

[0059] As used herein, the term "gapmer" refers to an oligonucleotide having an internal "gap segment" flanked by two external "wing segments," wherein the gap segment consists of a plurality of nucleotides that support RNase H cleavage, and each wing segment consists of one or more nucleotides that are chemically different from the nucleotides within the gap segment.

[0060] As used herein, the term "support material" refers to a high molecular weight compound or material that increases the molecular weight of the template, thereby allowing it to be retained when separating impurities and product from a reaction mixture.

[0061] " identity percentage " between the query nucleic acid sequence used herein and the target nucleic acid sequence is, after execution by pairwise BLASTN comparison, when target nucleic acid sequence and query nucleic acid sequence have 100% query coverage, " identity " value expressed as percentage by BLASTN algorithm calculation.Use the default setting of the BLASTN algorithm available on the website of National Center for Biotechnology Institute (National Centerfor Biotechnology Institute), the filter of low complexity region is closed, perform such pairwise BLASTN comparison between query nucleic acid sequence and target nucleic acid sequence.Importantly, query nucleic acid sequence can be described by the nucleotide sequence identified in one or more claims herein.Query sequence can have 100% identity with target sequence, or it can comprise the nucleotide change of up to a certain integer number relative to target sequence, make % identity less than 100%.For example, query sequence has at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity with target sequence.

[0062] Summary of the invention.

[0063] In one aspect of the present invention, there is provided a method for producing a single-stranded oligonucleotide product having at least one modified nucleotide residue, the method comprising:

[0064] a) providing a template oligonucleotide (I) complementary to the sequence of said product, said template having properties allowing its separation from said product;

[0065] b) providing a collection of oligonucleotides (II);

[0066] c) contacting (I) and (II) under conditions allowing annealing;

[0067] d) changing the conditions to separate any impurities, including denaturing the annealed template and impurity oligonucleotide chains and separating the impurities; and

[0068] e) changing the conditions to separate the products, including denaturing the annealed template and product oligonucleotide chains and separating the products.

[0069] Such methods can be used to purify products from impurities (e.g., oligonucleotide collections produced by chemical synthesis via phosphoramidite chemistry, such as solid phase chemical synthesis via phosphoramidite chemistry).

[0070] In another embodiment of the present invention, a method is provided, comprising:

[0071] a) providing a template oligonucleotide (I) complementary to the sequence of said product, said template having properties allowing its separation from said product;

[0072] b) providing a collection of oligonucleotides (II) comprising short oligonucleotides that are segments of said target sequence, wherein at least one segment comprises at least one modified nucleotide residue;

[0073] c) contacting (I) and (II) under conditions allowing annealing;

[0074] d) ligating the segment oligonucleotides;

[0075] e) changing the conditions to separate any impurities, including denaturing the annealed template and impurity oligonucleotide chains and separating the impurities; and

[0076] f) changing the conditions to separate the products, including denaturing the annealed template and product oligonucleotide chains and separating the products.

[0077] In one embodiment of the invention, there are substantially no nucleotides in the reaction vessel. In one embodiment of the invention, there are no nucleotides in the reaction vessel. Specifically, the reaction vessel does not comprise a collection of nucleotides, i.e., the reactants contain substantially no, preferably no, nucleotides.

[0078] Another embodiment of the present invention provides the method as previously described, wherein said denaturation results from an increase in temperature.

[0079] One embodiment of the present invention provides the method disclosed above, wherein the segment oligonucleotides are linked by enzymatic ligation. In another embodiment, the enzymatic ligation is performed by a ligase.

[0080] Another embodiment of the present invention provides a method disclosed above, wherein the segment oligonucleotides are connected by chemical ligation. In another embodiment, the chemical ligation is a click chemistry reaction. In one embodiment of the present invention, the chemical ligation of the segment oligonucleotides occurs in a templated reaction that produces a phosphoramidate bond as described in detail by Kalinowski et al. in ChemBioChem 2016, 17, 1150–1155.

[0081] Another embodiment of the present invention provides the method disclosed above, wherein the segment oligonucleotide is 3-15 nucleotides long. In another embodiment of the present invention, the segment is 5-10 nucleotides long. In another embodiment of the present invention, the segment is 5-8 nucleotides long. In another embodiment of the present invention, the segment is 5, 6, 7 or 8 nucleotides long. In a specific embodiment, there are three segment oligonucleotides: a 5' segment of 7 nucleotides long, a central segment of 6 nucleotides long, and a 3' segment of 7 nucleotides long, which when linked together form an oligonucleotide of 20 nucleotides long ("20-mer"). In a specific embodiment, there are three segment oligonucleotides: a 5' segment of 6 nucleotides long, a central segment of 8 nucleotides long, and a 3' segment of 6 nucleotides long, which when linked together form an oligonucleotide of 20 nucleotides long ("20-mer"). In a specific embodiment, there are three segment oligonucleotides: a 5' segment of 5 nucleotides in length, a central segment of 10 nucleotides in length, and a 3' segment of 5 nucleotides in length, which when linked together form a 20 nucleotide long oligonucleotide ("20-mer"). In a specific embodiment, there are four segment oligonucleotides: a 5' segment of 5 nucleotides in length, a 5'-central segment of 5 nucleotides in length, a central-3' segment of 5 nucleotides in length, and a 3' segment of 5 nucleotides in length, which when linked together form a 20 nucleotide long oligonucleotide ("20-mer").

[0082] One embodiment of the present invention provides the method described above, wherein the product is 10-200 nucleotides long. In another embodiment of the present invention, the product is 15-30 nucleotides long. In one embodiment of the present invention, the product is 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides long. In one embodiment of the present invention, the product is 20 nucleotides long, a "20-mer". In one embodiment of the present invention, the product is 21 nucleotides long, a "21-mer". In one embodiment of the present invention, the product is 22 nucleotides long, a "22-mer". In one embodiment of the present invention, the product is 23 nucleotides long, a "23-mer". In one embodiment of the present invention, the product is 24 nucleotides long, a "24-mer". In one embodiment of the present invention, the product is 25 nucleotides long, a "25-mer". In one embodiment of the present invention, the product is 26 nucleotides long, a "26-mer". In one embodiment of the invention, the product is 27 nucleotides long, a "27-mer". In one embodiment of the invention, the product is 28 nucleotides long, a "28-mer". In one embodiment of the invention, the product is 29 nucleotides long, a "29-mer". In one embodiment of the invention, the product is 30 nucleotides long, a "30-mer".

[0083] In one embodiment of the invention, the method is a method for producing a therapeutic oligonucleotide. In one embodiment of the invention, the method is a method for producing a single-stranded therapeutic oligonucleotide. In one embodiment of the invention, the method is a method for producing a double-stranded therapeutic oligonucleotide.

[0084] Another embodiment of the present invention provides the method disclosed above, wherein the property that allows the template to be separated from the product is that the template is attached to a support material. In another embodiment of the present invention, the support material is a soluble support material. In another embodiment of the present invention, the soluble support material is selected from the group consisting of polyethylene glycol, a soluble organic polymer, DNA, protein, dendrimers, polysaccharides, oligosaccharides and carbohydrates. In one embodiment of the present invention, the support material is polyethylene glycol (PEG). In another embodiment of the present invention, the support material is an insoluble support material. In another embodiment of the present invention, the support material is a solid support material. In another embodiment, the solid support material is selected from the group consisting of glass beads, polymer beads, fiber supports, membranes, streptavidin-coated beads and cellulose. In one embodiment, the solid support material is streptavidin-coated beads. In another embodiment, the solid support material is part of the reaction vessel itself, such as a reaction wall.

[0085] One embodiment of the present invention provides the method disclosed above, wherein the multiple repeated copies of the template are attached to the support material in a continuous manner via a single attachment point. The multiple repeated copies of the template can be separated by a linker, for example as in Figure 3 The multiple repeated copies of the template may be direct repeats, ie they are not separated by linkers.

[0086] Another embodiment of the present invention provides the method disclosed above, wherein the property allowing separation of template from product is the molecular weight of the template.For example, repeated copies of the template sequence can be present on a single oligonucleotide, with or without a linker sequence.

[0087] Another embodiment of the present invention provides a method as disclosed above, wherein the template, or the template and support material is recycled for use in a future reaction, for example as described in detail below. Another embodiment of the present invention provides a method as disclosed above, wherein the reaction is carried out using a continuous or semi-continuous flow process, for example as described in Figure 4 or Figure 5 As shown in .

[0088] In one embodiment of the invention, described method is for large-scale preparation oligonucleotide, especially therapeutic oligonucleotide.In the context of the present invention, the large-scale preparation of oligonucleotide refers to the preparation of the scale being greater than or equal to 1 liter, for example, in 1 L or larger reactor, implement described method.Alternately or in addition, in the context of the present invention, the large-scale preparation of oligonucleotide refers to the product preparation at gram scale, especially the production being greater than or equal to 10 grams of products.In one embodiment of the invention, the amount of the oligonucleotide product of production is at gram scale.In one embodiment of the invention, the amount of the product of production is greater than or equal to: 1,2,3,4,5,6,7,8,9,10,20,30,40,50,60,70,80,90 or 100 grams. In one embodiment of the invention, the amount of the oligonucleotide product produced is greater than or equal to: 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950 grams. In one embodiment of the invention, the amount of the oligonucleotide product produced is 500 grams or larger. In one embodiment of the invention, the oligonucleotide product produced is on a kilogram scale. In one embodiment of the invention, the amount of the oligonucleotide product produced is 1 kg or more. In one embodiment of the invention, the amount of the oligonucleotide product produced is greater than or equal to: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 kg. In one embodiment of the invention, the amount of the oligonucleotide product produced is greater than or equal to: 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 kg.

[0089] In one embodiment of the present invention, the amount of the product produced is between 10 grams and 100 kg. In one embodiment of the present invention, the amount of the product produced is between 10 grams and 50 kg. In one embodiment of the present invention, the amount of the product produced is between 100 grams and 100 kg. In one embodiment of the present invention, the amount of the product produced is between 100 grams and 50 kg. In one embodiment of the present invention, the amount of the product produced is between 500 grams and 100 kg. In one embodiment of the present invention, the amount of the product produced is between 500 grams and 50 kg. In one embodiment of the present invention, the amount of the product produced is between 1 kg and 50 kg. In one embodiment of the present invention, the amount of the product produced is between 10 kg and 50 kg.

[0090] In one embodiment of the invention, oligonucleotide production occurs on a scale greater than or equal to 2, 3, 4, 5, 6, 7, 8, 9, 10 liters, for example, in a 2, 3, 4, 5, 6, 7, 8, 9, or 10 L reactor. In one embodiment of the invention, oligonucleotide production occurs on a scale greater than or equal to 20, 25, 30, 35, 40, 45, 50, 55, 60, 65 70, 75, 80, 85, 90, 95, 100 liters, for example, in a 20, 25, 30, 35, 40, 45, 50, 55, 60, 65 70, 75, 80, 85, 90, 95, 100 L reactor. In one embodiment of the invention, oligonucleotide preparation occurs on a scale greater than or equal to 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000 liters, for example in a 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000 L reactor.

[0091] In one embodiment of the invention, the reactor volume is about 10,000 L, about 5000 L, about 2000 L, about 1000 L, about 500 L, about 125 L, about 50 L, about 20 L, about 10 L, or about 5 L.

[0092] In one embodiment of the present invention, the reactor volume is between 5-10,000 L, between 10-5000 L, between 20-2000 L, or between 50-1000 L.

[0093] Compared to RNA or DNA based oligonucleotides, an oligonucleotide according to the invention may have at least one backbone modification, and / or at least one sugar modification and / or at least one base modification.

[0094] One embodiment of the present invention provides the method disclosed above, wherein the product comprises at least one modified nucleotide residue. In another embodiment, the modification is at the 2' position of the sugar moiety.

[0095] The oligonucleotides used in the methods of the present invention may include a sugar modification, ie, a modified form of the ribosyl moiety, such as a 2'-O-modified RNA such as a 2'-O-alkyl or 2'-O-(substituted)alkyl For example2'-O-methyl, 2'-O-(2-cyanoethyl), 2'-O-(2-methoxy)ethyl (2'-MOE), 2'-O-(2-thiomethyl)ethyl, 2'-O-butyryl, 2'-O-propargyl, 2'-O-allyl, 2'-O-(3-amino)propyl, 2'-O-(3-(dimethylamino)propyl), 2'-O-(2-amino)ethyl, 2'-O-(2-(dimethylamino)ethyl); 2'-deoxy (DNA); 2'-O-(haloalkoxy)methyl (Arai K. et al. . Bioorg. Med. Chem. 2011, 21, 6285) For example 2'-O-(2-chloroethoxy)methyl (MCEM), 2'-O-(2,2-dichloroethoxy)methyl (DCEM); 2'-O-alkoxycarbonyl For example 2'-O-[2-(methoxycarbonyl)ethyl] (MOCE), 2'-O-[2-(N-methylcarbamoyl)ethyl] (MCE), 2'-O-[2-(N, N-dimethylcarbamoyl)ethyl] (DCME); 2'-halogenated For example 2'-F, FANA (2'-F arabinosyl nucleic acid); carbasugar and azasugar modifications; 3'-O-alkyl For example 3'-O-methyl, 3'-O-butyryl, 3'-O-propargyl; and derivatives thereof.

[0096] In one embodiment of the invention, the sugar modification is selected from the group consisting of: 2'-fluoro (2'-F), 2'-O-methyl (2'-OMe), 2'-O-methoxyethyl (2'-MOE) and 2'-amino. In another embodiment, the modification is 2'-MOE.

[0097] Other sugar modifications include "bridged" or "bicyclic" nucleic acids (BNA), such as locked nucleic acids (LNA), xylo-LNA, α-L-LNA, β-D-LNA, cEt (2'-O, 4'-C constrained ethyl) LNA, cMOEt (2'-O, 4'-C constrained methoxyethyl) LNA, ethylene-bridged nucleic acids (ENA), tricyclic DNA; unlocked nucleic acids (UNA); cyclohexenyl nucleic acids (CeNA), altriol nucleic acids (ANA), hexitol nucleic acids (HNA), fluoro-HNA (F-HNA), pyranosyl-RNA (p-RNA), 3'-deoxypyranosyl-DNA (p-DNA); morpholinos (e.g., in PMO, PPMO, PMOPlus, PMO-X); and derivatives thereof.

[0098] The oligonucleotides used in the methods of the present invention may include other modifications, such as peptide-base nucleic acids (PNA), boron-modified PNA, pyrrolidine-oxy-peptide nucleic acids (POPNA), ethylene glycol or glycerol-based nucleic acids (GNA), threose-based nucleic acids (TNA), acyclic threonine-based nucleic acids (aTNA), oligonucleotides with integrated bases and backbones (ONIBs), pyrrolidine-amide oligonucleotides (POMs); and their derivatives.

[0099] In one embodiment of the present invention, the modified oligonucleotide comprises a phosphorodiamidate morpholino oligomer (PMO), a locked nucleic acid (LNA), a peptide nucleic acid (PNA), a bridged nucleic acid (BNA) such as ( S )-cEt-BNA or SPIEGELMER.

[0100] In another embodiment, the modification is in the nucleoside base. Base modifications include modified forms of natural purine and pyrimidine bases (e.g., adenine, uracil, guanine, cytosine, and thymine), such as inosine, hypoxanthine, orotic acid, agmatidine, lysidine, 2-thiopyrimidines (e.g., 2-thiouracil, 2-thiothymine), G-clamps and their derivatives, 5-substituted pyrimidines (e.g., 5-methylcytosine, 5-methyluracil, 5-halouracil, 5-propynyluracil, 5-propynylcytosine, 5-aminomethyluracil, 5-hydroxymethyluracil, 5-aminomethylcytosine, 5-hydroxymethylcytosine, Super T), 2,6-diaminopurine, 7-deazaguanine, 7-deazaadenine, 7-aza-2,6-diaminopurine, 8-aza-7-deazaguanine, 8-aza-7-deazaadenine, 8-aza-7-deaza-2,6-diaminopurine, Super G, Super A and N4-ethylcytosine or its derivatives; N 2 -cyclopentylguanine (cPent-G), N 2 -cyclopentyl-2-aminopurine (cPent-AP) and N 2-propyl-2-aminopurine (Pr-AP) or its derivatives; and a degenerate or universal base, such as 2,6-difluorotoluene, or a lack of a base, such as an abasic site (e.g., 1-deoxyribose, 1,2-dideoxyribose, 1-deoxy-2-O-methylribose; or a pyrrolidine derivative in which the ring oxygen has been replaced by nitrogen (azaribose)). Examples of derivatives of Super A, Super G, and Super T can be found in US6683173. cPent-G, cPent-AP, and Pr-AP have been shown to reduce the immunostimulatory effect when incorporated into siRNA (Peacock H. et al. J. Am. Chem. Soc. (2011), 133, 9200).

[0101] In one embodiment of the present invention, the nucleobase modification is selected from the group consisting of: 5-methylpyrimidine, 7-deazaguanosine, and abasic nucleotides. In one embodiment, the modification is 5-methylcytosine.

[0102] The oligonucleotides used in the methods of the invention may include backbone modifications, for example, a modified form of a phosphodiester present in RNA, such as phosphorothioate (PS), phosphorodithioate (PS2), phosphonoacetate (PACE), phosphonoacetamide (PACA), thiophosphonoacetate, thiophosphonoacetamide, phosphorothioate prodrugs, H-phosphonates, methyl phosphonates, methyl thiophosphonates, methyl phosphates, methyl thiophosphonates, ethyl phosphates, ethyl thiophosphonates, borophosphates, borophosphorothioates, methyl borophosphorothioates, methyl borophosphorothioates, methyl borophosphonates, methyl borophosphonates, and derivatives thereof. Another modification includes phosphoramidite, phosphoramidate, N3'→P5' phosphoramidate, phosphorodiamidate, thiophosphorodiamidate, sulfamate, dimethylene sulfoxide, sulfonate, triazole, oxalyl, carbamate, methyleneimino (MMI) and thioacetamido nucleic acid (TANA); and derivatives thereof.

[0103] In another embodiment, the modification is in the main chain and is selected from: phosphorothioate (PS), phosphoramidate (PA) and phosphorodiamidates. In one embodiment of the invention, the modified oligonucleotide is a phosphorodiamidates morpholino oligomer (PMO). PMO has a methylene morpholine ring main chain that connects containing phosphorodiamidates. In one embodiment of the invention, the product has a phosphorothioate (PS) main chain.

[0104] In one embodiment of the invention, the oligonucleotide comprises a combination of two or more modifications as disclosed above.It will be appreciated by those skilled in the art that there are many synthetic oligonucleotide derivatives.

[0105] In one embodiment of the invention, the product is a gapmer. In one embodiment of the invention, the 5' and 3' wings of the gapmer comprise or consist of 2'-MOE modified nucleotides. In one embodiment of the invention, the gap segment of the gapmer comprises or consists of nucleotides that contain a hydrogen at the 2' position of the sugar moiety (i.e., is DNA-like). In one embodiment of the invention, the 5' and 3' wings of the gapmer consist of 2'MOE modified nucleotides, and the gap segment of the gapmer consists of nucleotides that contain a hydrogen at the 2' position of the sugar moiety (i.e., deoxynucleotides). In one embodiment of the invention, the 5' and 3' wings of the gapmer consist of 2'MOE modified nucleotides, and the gap segment of the gapmer consists of nucleotides that contain a hydrogen at the 2' position of the sugar moiety (i.e., deoxynucleotides), and the linkages between all nucleotides are phosphorothioate bonds.

[0106] One embodiment of the present invention provides the method described above, wherein the product obtained is greater than 90% pure. In another embodiment, the product is greater than 95% pure. In another embodiment, the product is greater than 96% pure. In another embodiment, the product is greater than 97% pure. In another embodiment, the product is greater than 98% pure. In another embodiment, the product is greater than 99% pure. Use any suitable method, for example high performance liquid chromatography (HPLC) or mass spectrometry (MS), especially liquid chromatography-MS (LC-MS), HPLC-MS or capillary electrophoresis mass spectrometry (CEMS), the purity of the oligonucleotide can be determined.

[0107] Another embodiment of the present invention provides a method for producing a double-stranded oligonucleotide, wherein two complementary single-stranded oligonucleotides are produced by the method of any of the preceding embodiments and then mixed under conditions that allow annealing. In one embodiment, the product is siRNA.

[0108] One embodiment of the present invention provides oligonucleotides produced by the methods described above. In one embodiment of the present invention, the oligonucleotides produced are RNA. In one embodiment of the present invention, the oligonucleotides produced are DNA. In one embodiment of the present invention, the oligonucleotides produced comprise RNA and DNA. In another embodiment of the present invention, the oligonucleotides produced are modified oligonucleotides. In one embodiment of the present invention, the oligonucleotides produced are antisense oligonucleotides. In one embodiment of the present invention, the oligonucleotides produced are aptamers. In one embodiment of the present invention, the oligonucleotides produced are miRNA. In one embodiment of the present invention, the product is a therapeutic oligonucleotide.

[0109] The invention disclosed herein utilizes the binding properties of oligonucleotides to provide improved methods for their production. By providing the template oligonucleotide with 100% complementarity to the target sequence and controlling the reaction conditions so that the product can be released and isolated under specific conditions, a product with high purity can be obtained.

[0110] Release of product (or impurities) from template, i.e. denaturation of template duplex and separation of the product (or impurities) (or impurities)

[0111] Releasing the product from the template (or wherein the method comprises an additional step of impurity release, any impurity) requires destroying the Watson-Crick base pairing between the template oligonucleotide chain and the product (or impurity). The product (or impurity) can then be separated from the template. This can occur as two separate steps, or as a combined step.

[0112] If the method is carried out in a column reactor, the release and isolation of the product (or impurities) can occur in one step. Running in a buffer that changes pH or salt concentration or contains chemicals that disrupt base pairing (such as formamide or urea) will cause denaturation of the oligonucleotide chain and the product (or impurities) will be eluted in the buffer.

[0113] When implementing described method in other reaction vessels, the release of described product (or impurity) and separation can occur as two-step method.First, Watson-Crick base pair is destroyed to separate chain, then from described reaction vessel, take out described product (or impurity).When releasing and separating described product when two-step method is implemented, by changing buffer conditions (pH, salt) or mixing chemical destruction reagent (formamide, urea), can realize the destruction of Watson-Crick base pair.Alternatively, raising temperature will also cause the dissociation of two chains.Then can take out product (or impurity) from described reaction vessel by method, described method comprises separation based on molecular weight, separation based on charge, separation based on hydrophobicity, separation based on specific sequence or the combination of these methods.

[0114] When implementing described method in continuous or semi-continuous flow reactor, the release of described product (or impurity) and separation can be in one step or two steps.For example, by increasing temperature to cause the dissociation of two chains and in the same part of the reactor for raising the temperature, separate the chain of release based on molecular weight, it is possible to implement the release and separation of described product (or impurity) in one step.By increasing temperature in a part of reactor to cause the dissociation of two chains and in the different parts of reactor, separate the chain of release based on molecular weight, it is possible to implement the release and separation of described product (or impurity) in two steps.

[0115] Specific release and separation of impurities from the template, but retention of the product on the template

[0116] Impurities are produced when the wrong nucleotide is incorporated into the oligonucleotide chain during chain extension, or when the chain extension reaction is terminated prematurely. Impurities are also produced when the reaction includes a step of ligating segmented oligonucleotides and one or more ligation steps do not occur. Figure 6 The diagram shows the types of impurities that can be generated.

[0117] The property of Watson-Crick base pairing can be used to specifically release any impurities bound to the template before the product is released. Every double-stranded oligonucleotide will dissociate under specific conditions, and those conditions are different for sequences that do not have 100% complementarity compared to sequences with 100% complementarity. Determining such conditions is within the skill's remit.

[0118] A common way to denature oligonucleotides is to increase the temperature. The temperature at which half of the base pairs are dissociated (i.e., when 50% of the duplex is in a single-stranded state) is called the melting temperature (T). m The most reliable and accurate way to determine the melting temperature is empirically. However, this is cumbersome and usually not necessary. Several formulas can be used to calculate T m Value (NucleicAcids Research 1987, 15 (13): 5069-5083; PNAS 1986, 83 (11): 3746-3750; Biopolymers 1997, 44 (3): 217-239), and numerous melting temperature calculators can be found online, hosted by reagent suppliers and universities. It is known that for a given oligonucleotide sequence, a variant with all phosphorothioate bonds will melt at a lower temperature than a variant with all phosphodiester bonds. Increasing the number of phosphorothioate bonds in an oligonucleotide tends to reduce the T of the oligonucleotide for its intended target. m .

[0119] In order to specifically separate impurities from the reaction mixture, the melting temperature of the product: template duplex is first calculated. The reaction vessel is then heated to a first temperature, such as a temperature lower than the melting temperature of the product: template duplex, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 degrees Celsius lower than the melting temperature. This will cause the denaturation of oligonucleotides that are not products from the template (i.e., not 100% complementary to the template). These can then be removed from the reaction vessel using one of the methods disclosed above (e.g., separation based on molecular weight, separation based on charge, separation based on hydrophobicity, separation based on specific sequences, or a combination of these methods). The reaction vessel is then raised to a second, higher temperature, such as higher than the calculated melting temperature, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 degrees Celsius higher than the melting temperature, to cause the denaturation of the product from the template. The product can then be removed from the reaction vessel using one of the methods disclosed above (e.g., separation based on molecular weight, separation based on charge, separation based on hydrophobicity, separation based on specific sequences, or a combination of these methods).

[0120] When the destroying agent is a reagent or chemical destroying agent that causes changes in pH or salt concentration, a similar method can be used. The concentration of the destroying agent is increased until it is just below the concentration when the product can dissociate, so as to cause denaturation of the oligonucleotide of the product that is not from the template. These impurities can then be removed from the reaction vessel using one of the methods disclosed above. The concentration of the destroying agent is then increased to above the concentration that dissociates the product from the template. The product can then be taken out from the reaction vessel using one of the methods disclosed above.

[0121] The product obtained from the method (such as the method disclosed above) has high purity and does not require other purification steps. For example, the product obtained is greater than 95% pure.

[0122] Template performance

[0123] The template needs to have a performance that allows it to remain in the reaction vessel when the product is taken out to prevent it from becoming an impurity in the product. In other words, the template has a performance that allows it to be separated from the product. In one embodiment of the invention, this retention is achieved by coupling the template to a support material. This coupling produces a template-support complex with a high molecular weight, and therefore can remain in the reaction vessel when taking out impurities and product (e.g., by filtering). The template can be coupled to solid support materials such as polymer beads, fiber supports, membranes, streptavidin-coated beads and cellulose. The template can also be coupled to soluble support materials such as polyethylene glycol, soluble organic polymers, DNA, proteins, dendrimers, polysaccharides, oligosaccharides, and carbohydrates.

[0124] Each support material may have multiple points to which a template may be attached, and each attachment point may have multiple templates attached, for example Figure 3 The method shown in .

[0125] In the case where the template itself is not attached to a support material, it can be of high molecular weight, for example, it can be a molecule having multiple copies of the template, for example, Figure 3 The ways shown are separated by joints.

[0126] The ability to retain the template in the reaction vessel also allows for recycling of the template for future reactions, either by recovery or by use in a continuous or semi-continuous flow process.

[0127] Method for separating template from product (or impurities)

[0128] The properties of the templates disclosed above allow separation of the template and the product, or separation of the template bound to the product and impurities. Separation based on molecular weight, separation based on charge, separation based on hydrophobicity, separation based on specific sequence, or a combination of these methods can be used.

[0129] In the case where the template is attached to a solid support, separation of the template from the product or separation of impurities from the product bound to the template is achieved by washing the solid support under appropriate conditions. In the case where the template is coupled to a soluble support or itself consists of a repeating template sequence, separation of the template from the product or separation of the template-bound product from impurities can be achieved by means of separation based on molecular weight. This can be achieved by using techniques such as ultrafiltration or nanofiltration, in which the filter material is selected so that larger molecules are retained by the filter and smaller molecules pass through. In the case where a single separation step of the impurity-template-product complex or the separation of the product from the template is not effective enough, multiple continuous filtration steps can be used to increase the separation efficiency and thus produce a product that meets the desired purity.

[0130] It is desirable to provide methods for isolating such oligonucleotides that are efficient and applicable on an industrial production scale. "Therapeutic oligonucleotides: The state of the art in purification technologies" Sanghvi et al. Current Opinion in Drug Discovery (2004) Vol. 7 No. 8 reviews methods for oligonucleotide purification.

[0131] WO-A-01 / 55160 discloses the purification of oligonucleotides by forming an imine bond with a contaminant and then removing the imine-linked impurities by chromatography or other techniques. "Size Fractionation of DNA Fragments Ranging from 20 to 30,000 Base Pairs by Liquid / Liquid chromatography" Muller et al. Eur. J. Biochem (1982) 128-238 discloses the use of a solid column of microcrystalline cellulose on which a PEG / dextran phase has been deposited for the separation of nucleotide sequences. "Separation and identification of oligonucleotides by hydrophilic interaction chromatography." Easter et al. The Analyst (2010); 135(10) discloses the separation of oligonucleotides using a variant of HPLC that employs a solid silica gel support phase. "Fractionation of oligonucleotides of yeast soluble ribonucleic acids by countercurrent distribution" Doctor et al. Biochemistry (1965) 4(1) 49-54 discloses the use of dry solid columns packed with dry DEAE-cellulose. "Oligonucleotide composition of a yeast lysine transfer ribonucleic acid" Madison et al.; Biochemistry, 1974, 13(3) discloses the use of solid-phase chromatography for the separation of nucleotide sequences.

[0132] Liquid-liquid chromatography is a known separation method. "Countercurrent Chromatography The Support-Free Liquid Stationary Phase" Billardello, B.; Berthod, A; Wilson & Wilson's Comprehensive Analytical Chemistry 38; Berthod, A., ed.; Elsevier Science BV: Amsterdam (2002) pp. 177-200 provides a useful general description of liquid-liquid chromatography. Various liquid-liquid chromatography techniques are known. One technique is liquid-liquid countercurrent chromatography (referred to herein as "CCC"). Another known technique is centrifugal partition chromatography (referred to herein as "CPC").

[0133] The methods disclosed above and those described in WO 2013 / 030263 can be used to isolate the product oligonucleotide, for example from the template and / or impurities.

[0134] Oligonucleotides used as starting materials

[0135] The oligonucleotide that is used as the starting material of method for the present invention is described as " set " in this article and its definition is provided in the above.Described set is the non-homogeneous set of oligonucleotide.The oligonucleotide that forms set has been produced by other oligonucleotide production method, and therefore can contain high degree impurity.Therefore, when this oligonucleotide set was applied to method for the present invention, the ability of removing impurity specifically as described herein can cause purification step to occur.

[0136] The set can contain oligonucleotides that are intended to be the same length as the template oligonucleotide (although will contain impurities of varying lengths and residues mistakenly incorporated). The set can also be composed of segments of product oligonucleotides that are linked together when assembled on the template. Each segment will be a heterogeneous set containing impurities of varying lengths and residues mistakenly incorporated.

[0137] Ligase

[0138] In one aspect of the present invention, a ligase is provided. In one embodiment of the present invention, the ligase is an ATP-dependent ligase. ATP-dependent ligases are in the size range of 30 to >100 kDa. In one embodiment of the present invention, the ligase is an NAD-dependent ligase. NAD-dependent enzymes are highly homologous and are monomeric proteins of 70-80 kDa. In one embodiment of the present invention, the ligase is a thermostable ligase. A thermostable ligase can be derived from thermophilic bacteria.

[0139] In one embodiment of the present invention, the ligase is a modified ligase. For example, the modified ligase includes a modified T4 DNA ligase, a modified Enterobacteriaceae phage CC31 ligase, a modified Shigella phage Shf125875 ligase, and a modified Chlorella ligase.

[0140] In one embodiment, the wild-type T4 DNA ligase is modified at amino acid position 368 or amino acid position 371 of SEQ ID NO:3.

[0141] In one embodiment, the mutant ligase comprises or consists of SEQ ID NO: 3, wherein the amino acid at position 368 is R or K.

[0142] In one embodiment, the mutant ligase comprises or consists of SEQ ID NO: 3, wherein the amino acid at position 371 is any one of the following amino acids: L, K, Q, V, P, R.

[0143] In one embodiment, the corresponding residues disclosed above for T4 DNA ligase are mutated in any of Enterobacteriaceae phage CC31 ligase, Shigella phage Shf125875 ligase, and Chlorella vulgaris ligase. Conserved regions of DNA ligases are disclosed in Chem. Rev. 2006, 106, 687-699 and Nucleic Acids Research, 2000, Vol. 28, No. 21, 4051-4058. In one embodiment, the ligase is modified in the linker region.

[0144] In one embodiment of the present invention, the ligase comprises or consists of SEQ ID NO: 23 or a ligase having at least 90% sequence identity thereto, excluding the wild-type ligase For example Enterobacteriaceae phage CC31 ligase.

[0145] In one embodiment of the present invention, the ligase comprises or consists of any one of the following amino acid sequences: SEQ ID NO: 10-28.

[0146] In one embodiment of the invention, the ligase is immobilized, for example on beads.

[0147] In one aspect of the present invention, there is provided the use of a ligase comprising the amino acid sequence of SEQ ID NO: 6 or SEQ ID NO: 8 for linking a 5' segment (containing one or more modified sugar moieties) to a 3' segment, wherein all sugar moieties within the 3' segment are unmodified. In one embodiment of the present invention, there is provided the use of a ligase comprising the amino acid sequence of SEQ ID NO: 6 or SEQ ID NO: 8 for linking a 5' segment (containing one or more sugar moieties with a 2'-OMe modification) to a 3' segment, wherein all sugar moieties within the 3' segment are unmodified. In one embodiment, all sugar moieties in the 5' segment contain a 2'-OMe modification. In one embodiment, the 5' segment contains five sugar moieties with a 2'-OMe modification.

[0148] The present invention includes the following items:

[0149] 1. A method for producing a single-stranded oligonucleotide product, the method comprising:

[0150] a) providing a template oligonucleotide (I) complementary to the sequence of the product, said template having properties allowing its separation from said product;

[0151] b) providing a collection of oligonucleotides (II);

[0152] c) contacting (I) and (II) under conditions permitting annealing; and

[0153] d) changing the conditions to remove the product.

[0154] 2. The method according to item 1, comprising:

[0155] a) providing a template oligonucleotide (I) complementary to the sequence of the product, said template having properties allowing its separation from said product;

[0156] b) providing a collection of oligonucleotides (II);

[0157] c) contacting (I) and (II) under conditions allowing annealing;

[0158] d) modifying the conditions to remove impurities; and

[0159] e) changing the conditions to remove the product.

[0160] 3. The method according to item 1 or 2, comprising:

[0161] a) providing a template oligonucleotide (I) complementary to the sequence of the product, said template having properties allowing its separation from said product;

[0162] b) providing a pool of oligonucleotides (II) containing oligonucleotides that are segments of the product sequence;

[0163] c) contacting (I) and (II) under conditions allowing annealing;

[0164] d) ligating the segment oligonucleotides to form the product;

[0165] e) modifying the conditions to remove impurities; and

[0166] f) changing the conditions to remove the product.

[0167] 4. The method according to any preceding item, wherein the method occurs in a reaction vessel, and wherein changing the conditions to remove the product comprises the steps of separating the annealed oligonucleotide chains and removing the product from the reaction vessel.

[0168] 5. The method of any one of items 2-4, wherein the method occurs in a reaction vessel, and wherein changing the conditions to remove impurities comprises a step of separating the annealed oligonucleotide chains and a step of removing impurities from the reaction vessel.

[0169] 6. A method according to item 4 or 5, wherein the chain separation results from an increase in temperature.

[0170] 7. The method according to item 6, comprising two steps of increasing the temperature i) to separate the annealed impurities and ii) to separate the annealed product.

[0171] 8. The method of item 3-7, wherein the segment oligonucleotides are linked by enzymatic ligation.

[0172] 9. The method of claim 8, wherein the enzyme is a ligase.

[0173] 10. A method according to any one of items 3-9, wherein the segment is 3-15 nucleotides long.

[0174] 11. The method of any preceding item, wherein the product is 10-200 nucleotides long.

[0175] 12. The method of claim 11 , wherein the product is 20-25 nucleotides long.

[0176] 13. The method according to item 12, comprising three segment oligonucleotides: a 5' segment of 7 nucleotides in length, a central segment of 6 nucleotides in length and a 3' segment of 7 nucleotides in length.

[0177] 14. The method according to item 12, comprising three segment oligonucleotides: a 5' segment of 6 nucleotides in length, a central segment of 8 nucleotides in length and a 3' segment of 6 nucleotides in length.

[0178] 15. The method according to item 12, comprising three segment oligonucleotides: a 5' segment of 5 nucleotides in length, a central segment of 10 nucleotides in length, and a 3' segment of 5 nucleotides in length.

[0179] 16. The method according to any preceding item, wherein the property allowing separation of the template from the product is that the template is attached to a support material.

[0180] 17. The method according to item 16, wherein the support material is a soluble support material.

[0181] 18. The method of claim 17, wherein the support material is polyethylene glycol.

[0182] 19. The method of any one of items 11 to 18, wherein the plurality of repeated copies of the template are attached to the support material in a serial manner via a single point of attachment.

[0183] 20. A method according to any one of items 1 to 15, wherein the property allowing separation of the template from the product is the molecular weight of the template.

[0184] 21. The method of any preceding item, wherein the template, or the template and support material, is recycled for use in future reactions.

[0185] 22. The method according to any preceding item, wherein the reaction is carried out using a continuous flow process.

[0186] 23. The method of any preceding item, wherein the product contains at least one modified nucleotide residue.

[0187] 24. The method of claim 23, wherein at least one segment contains at least one modified nucleotide residue.

[0188] 25. A method according to item 23 or item 24, wherein the modification is at the 2' position of the sugar moiety.

[0189] 26. The method of claim 25, wherein the modification is selected from 2'-F, 2'-OMe, 2'-MOE and 2'-amino.

[0190] 27. The method of claim 24, wherein the oligonucleotide comprises a PMO, an LNA, a PNA, a BNA or a Spiegelmer™.

[0191] 28. A method according to item 23 or item 24, wherein the modification is in the nucleoside base.

[0192] 29. The method of claim 28, wherein the modification is selected from 5-methylpyrimidine, 7-deazaguanosine, and abasic nucleotides.

[0193] 30. A method according to item 23 or item 24, wherein the modification is in the main chain.

[0194] 31. The method of claim 30, wherein the modification is selected from phosphorothioate, phosphoramidate, and phosphorodiamidate.

[0195] 32. The method of any preceding item, wherein the product obtained is at least 90% pure.

[0196] 33. The method of claim 32, wherein the product is at least 95% pure.

[0197] 34. A method for producing a double-stranded oligonucleotide, wherein two complementary single-stranded oligonucleotides are produced by the method of any preceding item and then mixed under conditions that allow annealing.

[0198] 35. A method as claimed in any one of the preceding items, wherein the method is for producing a therapeutic oligonucleotide.

[0199] 36. An oligonucleotide produced by the method of any one of items 1 to 35.

[0200] 37. The oligonucleotide according to claim 36, wherein the oligonucleotide is a modified oligonucleotide produced by the method of any one of items 23-35.

[0201] 38. The oligonucleotide of claim 37, wherein the oligonucleotide is a gapmer.

[0202] 39. A ligase comprising the amino acid sequence set forth in SEQ ID NO: 3, wherein the amino acid at position 368 is substituted with R or K; and / or the amino acid at position 371 is substituted with L, K, Q, V, P or R.

[0203] 40. A ligase comprising the amino acid sequence described in any one of the following SEQ ID NOs: 10-28.

[0204] 41. Use of a ligase comprising the amino acid sequence depicted in SEQ ID NO: 6 or SEQ ID NO: 8 for ligating a 5' segment containing one or more sugar moieties having a 2'-OMe modification to a 3' segment, wherein all sugar moieties within the 3' segment are unmodified. Example

[0205] abbreviation

[0206] OMe O-methyl

[0207] MOE O-methoxyethyl (DNA backbone) or methoxyethyl (RNA backbone)

[0208] CBD cellulose binding domain

[0209] HPLC high-performance liquid chromatography

[0210] PBS Phosphate-buffered saline

[0211] HAA Hexylammonium Acetate

[0212] SDS PAGE Sodium dodecyl sulfate polyacrylamide gel electrophoresis

[0213] LCMS liquid chromatography mass spectrometry

[0214] PO phosphate diester

[0215] DTT dithiothreitol.

[0216] Example 1: Assembly and ligation of oligonucleotide (DNA) segments using wild-type T4 DNA ligase

[0217] 1.1 Initiation and control of chemical synthesis sequences

[0218] To demonstrate that multiple short oligonucleotides ("segments") could be assembled and ligated in the correct order on complementary template strands to produce the desired end product ("target"), the segment, target, and template sequences detailed in Table 1 were chemically synthesized using standard methods.

[0219] 1.2 HPLC analysis

[0220] HPLC analysis was performed using an Agilent ZORBAX Eclipse Plus XDB-C18 column (4.6 x 150 mm, 5 μm dp. Agilent P / N 993967-902) operating at 0.2 ml / min, while monitoring absorbance at 258 nm. The column was maintained at 60°C. A 20 µl sample was injected, and a gradient from 20 to 31% buffer B was run over 20 minutes, followed by an increase to 80% buffer B for 5 minutes.

[0221] Buffer A: 75 ml 1 M HAA, 300 ml isopropanol, 200 ml acetonitrile, 4425 ml water

[0222] Buffer B: 650 ml isopropanol, 350 ml acetonitrile.

[0223] Table 1

[0224] name sequence %HPLC purity Amount (mg) 5' segment 5'-GGC CAA-3' 100.0 21.6 Central section 5'-(p)ACC TCG GC-3' 96.9 58.1 3' segment 5'-(p)T TAC CT-3' 98.8 39.8 target 5'-GGC CAA ACC TCG GCT TAC CT-3' (SEQ ID NO:1) 98.4 101.7 Biotinylated template 5'-Biotin TT TAG GTA AGC CGA GGT TTG GCC-3' (SEQ ID NO: 2) 96.9 130.7

[0225] 1.3 Oligonucleotide assembly and ligation method using commercially available T4 DNA ligase (SEQ ID NO: 3)

[0226] The 5' segment, central segment, and 3' segment were assembled on the template: Each segment and template were dissolved in water at a concentration of 1 mg / ml and then mixed as follows.

[0227] 5' segment 2µl

[0228] Central section 2µl

[0229] 3' segment 2µl

[0230] 2µl biotinylated template

[0231] H2O 36µl.

[0232] The combined oligonucleotide solution was incubated at 94°C for 5 minutes and cooled to 37°C, then incubated at 37°C for an additional 5 minutes to allow the segments to anneal to the template. 2µl (equivalent to 2µg) of T4 DNA ligase (NEB) and 4µl of 1x T4 DNA Ligation Buffer (NEB) were then added, and the reaction (total reaction volume 50µl) was incubated at room temperature for 1 hour. Afterwards, 40µl of streptavidin-coated magnetic beads were added, and the suspension was incubated at room temperature for 10 minutes to allow the biotinylated template to bind to the streptavidin beads. The streptavidin beads were washed with 2 x 100µl PBS to remove unbound segments. The washes were analyzed by HPLC. The reaction mixture was then incubated at 94°C for 10 minutes to separate the bound ligation product (or any bound segments) from the template, then rapidly cooled on ice to 'melt' the DNA and halt reannealing of the oligonucleotide product (or segments) to the template. The ligation reaction was then analyzed by HPLC.

[0233] 1.4 Oligonucleotide Assembly and Ligation Method Using In-House T4 DNA Ligase Bead Slurry

[0234] 1.4.1 Bead slurry production

[0235] T4 ligase (SEQ ID NO:4) fused to a cellulose binding domain (CBD) at the N-terminus was produced using standard cloning, expression, and extraction methods. The amino acid sequence of this T4 ligase differs from the commercially available T4 ligase sequence (SEQ ID NO:3) in that the N-terminal methionine (M) has been replaced with glycine and serine (GS). This facilitates the production and expression of the CBD fusion protein. The CBD-T4 ligase fusion protein was expressed in BL21 A1 cells (INVITROGEN). The supernatant was harvested and added to 600 µl of PERLOZA 100 (PERLOZA) beads and shaken at 26°C for 1 hour. The PERLOZA cellulose beads were then collected and washed with 2 ml of buffer (50 mM Tris pH 8.0, 200 mM NaCl, 0.1% Tween 20, 10% glycerol), followed by 5 ml of PBS, and finally resuspended in 200 µl of PBS (10 mM PO4 3- , 137 mM NaCl, 2.7 mM KCl pH 7.4). For analysis of protein expression, 15 µl of PERLOZA bead slurry was mixed with 5 µl of SDS loading buffer and incubated at 80°C for 10 min before running on an SDS PAGE gradient gel (4–20%) according to standard protocols.

[0236] 1.4.2 Oligonucleotide Assembly and Ligation Using Bead Slurry

[0237] For T4 ligase bound to PERLOZA beads, modify the assembly and ligation protocol described in 1.3 above as follows. Reduce the amount of 36 µl of HO in the initial segment mixture to 8 µl. After annealing, replace 2 µl of commercially available T4 DNA ligase with 20 µl of PERLOZA bead slurry. Centrifuge the PERLOZA beads and remove the supernatant before adding the streptavidin magnetic beads. Add the streptavidin magnetic beads to the supernatant and incubate at room temperature for 10 minutes to allow the biotinylated template to bind to the streptavidin beads.

[0238] 1.5 Results and Conclusions

[0239] The product, template, and all three-segment oligonucleotides are clearly resolved in the control chromatogram.

[0240] HPLC analysis of the ligase reaction indicated that some unligated oligonucleotide segments remained, but commercially available T4 DNA ligase (NEB) was able to catalyze the ligation of the three segments to produce the desired product oligonucleotide ( Figure 7 PERLOZA bead-bound T4 DNA ligase appeared to be less efficient in ligating oligonucleotide segments, which appeared in both control wash samples and reaction samples ( Figure 8 However, it is difficult to ensure that the same amount of enzyme is added to the beads compared to commercially available enzymes, so a direct comparison of ligation efficiencies is not possible.

[0241] Example 2: Assembly and ligation of 2'-OMe ribose-modified oligonucleotide segments using wild-type T4 DNA ligase

[0242] 2.1 2'OMe at each nucleotide position in each segment

[0243] To determine whether T4 DNA ligase can ligate oligonucleotide segments with modifications at the 2' position of the ribose ring, oligonucleotide segments having the same sequence as in Example 1 were synthesized, except that the 2' position of the ribose ring was substituted with an OMe group and thymidine was replaced with uridine as shown below.

[0244] Table 2

[0245] name sequence %HPLC purity Amount (mg) 5' segment 2'-OMe 5'-GGC CAA-3' 21 97.8 Central segment 2'-OMe 5'-(p)ACC UCG GC-3' 15.5 97.7 3' segment 2'-OMe 5'-(p)U UAC CU-3' 21.2 98.1 Target 2'-OMe 5'-GGC CAA ACC UCG GCU UAC CU-3' (SEQ ID NO:5) 88 96.9

[0246] (p) = phosphate.

[0247] Assembly, ligation, and HPLC analysis were performed using the methods of Example 1 using commercially available NEB ligase and T4 ligase CBD fusions bound to PERLOZA beads. For experiments using commercially available T4 DNA ligase (NEB), the amount of water used in the reaction mixture was 26 µl instead of 36 µl, resulting in a final reaction volume of 40 µl. For experiments using homemade T4 DNA ligase bead slurry, the amount of water used in the reaction mixture was 23 µl, and the amount of beads used was 5 µl, resulting in a final reaction volume of 40 µl. Control experiments using unmodified DNA (instead of 2'-OMe DNA) were performed in parallel.

[0248] Results from control experiments were according to Example 1. For the 2'-OMe experiments, no product was detected using HPLC, indicating that T4 DNA ligase was unable to completely ligate 2'-OMe modified oligonucleotide segments, regardless of whether commercial T4 DNA ligase bound to PERLOZA beads or a homemade T4 DNA ligase CBD fusion was used.

[0249] 2.2 2′-OMe at each nucleotide position in a single segment

[0250] Reactions were set up as detailed in Table 3 using a 1 mg / ml solution of each oligonucleotide.

[0251] Table 3

[0252] Experiment 1 (no ligase control) Experiment 2 (Single 2'-OMe Segment - 3') Experiment 3 (Single 2'-OMe segment - 5') Experiment 4 (all 2'-OMe) Experiment 5 (all unmodified) Volume (µl) template template template template template 2 5' segment 5' segment 2'-OMe substituted 5' segment 2'-OMe substituted 5' segment 5' segment 2 3' segment 2'-OMe substituted 3' segment 3' segment 2'-OMe substituted 3' segment 3' segment 2 Central section Central section Central section 2'-OMe substituted central segment Central section 2 <![CDATA[H2O]]> <![CDATA[H2O]]> <![CDATA[H2O]]> <![CDATA[H2O]]> <![CDATA[H2O]]> Until a total of 40

[0253] The commercially available NEB ligase and homemade PERLOZA-conjugated T4 DNA ligase were used to assemble and ligate the DNA using the method of Example 1.

[0254] The reactions were incubated at 94°C for 5 minutes, followed by 5 minutes at 37°C to allow annealing. 4 µl of 1x NEB T4 DNA Ligation Buffer was added to each reaction along with 5 µl (approximately 2 µg) of homemade T4 DNA ligase or 2 µl (approximately 2 µg) of commercially available T4 DNA ligase (except for Experiment 1, which served as a no-ligase control), and the ligation reaction was allowed to proceed at room temperature for 2 hours. Streptavidin magnetic beads were then added to each reaction, and the reactions were heated to 94°C and then rapidly cooled on ice as described in Example 1 to separate the template from the starting material and product.

[0255] The processed reaction was split in half: one half was analyzed by HPLC as described for Example 1 (Section 1.2). The other half of the sample was reserved for mass spectrometry to confirm the HPLC results.

[0256] Ligation of the unmodified oligonucleotide segments (Experiment 5) proceeded as expected to produce full-length products. Figure 9 As shown in , a small amount of ligation was seen when the 5' segment was replaced by 2'-OMe (Experiment 3), but no ligation was seen when the 3' segment was replaced by 2'-OMe (Experiment 2). According to 2.1, no significant product was seen when all three segments were replaced by 2'-OMe (Experiment 4).

[0257] 2.3 Conclusion

[0258] Wild-type T4 DNA ligase is poor at ligating 2'-OMe substituted oligonucleotide segments, but is slightly less sensitive to modifications of the 5' oligonucleotide segment compared to the 3' segment.

[0259] Example 3: Assembly and ligation of 2'-OMe ribose-modified oligonucleotide segments using wild-type and mutant ligases

[0260] 3.1 Materials

[0261] Using standard cloning, expression, and extraction methods, wild-type Enterobacteriaceae phage ligase CC31 (SEQ ID NO: 6), wild-type Shigella phage Shf125875 ligase (SEQ ID NO: 8), and 10 mutant T4 ligases of SEQ ID NOs: 10-19 (each fused to the CBD at the N-terminus) were produced. As disclosed in 1.4.1, to produce and express CBD fusion proteins, the N-terminal methionine (M) was replaced in each case with glycine and serine (GS) (e.g., SEQ ID NO: 7 for Enterobacteriaceae phage ligase CC31 and SEQ ID NO: 9 for Shigella phage Shf125875 ligase).

[0262] The following oligonucleotides were synthesized by standard solid phase methods.

[0263] Table 4

[0264] name sequence %HPLC purity Amount (mg) 5' segment* 5'-(OMe)G(OMe)G(OMe)C(OMe)C(OMe)AA-3' 99.35 24.7 Central section 5'-(p)ACC TCG GC-3' 96.9 58.1 3' segment 5'-(p)TTA CCT-3' 97.88 29.5 Biotinylated template 5'-Biotin TT TAG GTA AGC CGA GGT TTG GCC-3' (SEQ ID NO: 2) 96.9 130.7

[0265] NB OMe indicates a 2' methoxy substitution on the ribose ring

[0266] (p) = phosphate

[0267] * Note that the first five nucleotides are 2'-OMe modified (GGCCA), but the final A is not.

[0268] 3.2 Oligonucleotide Assembly and Ligation Methods Using Ligase Bead Slurry

[0269] 3.2.1 Bead slurry production

[0270] Ligase fused to the CBD was bound to PERLOZA beads as described in 1.4 to produce a bead slurry.

[0271] 3.2.2 Oligonucleotide Assembly and Ligation Using Bead Slurry

[0272] Prepare ligation reactions in a 96-well plate to a final volume of 50 µL using the following components:

[0273] 2µL ~1mg / mL 5'(2'-OMe) segment

[0274] 2µL ~1mg / mL central section

[0275] 2µL ~1mg / mL 3' segment

[0276] 2µL ~1mg / mL template

[0277] 5µL NEB T4 DNA Ligase Buffer

[0278] 22µL H2O

[0279] 15 µL PERLOZA-bound bead slurry.

[0280] The reaction was incubated at room temperature for 15 minutes to allow the segments to anneal to the template before adding the PERLOZA bead slurry. The PERLOZA bead slurry was added and the reaction was incubated at room temperature for 1 hour. After the 1-hour incubation, the solution was transferred to an ACOPRE Padvance 350 filter plate (PN 8082), which was placed on top of an ABGENE superplate (Thermo Scientific, #AB-2800) and centrifuged at 4,000 rpm for 10 minutes to remove the PERLOZA bead slurry. The solution was then analyzed by HPLC using the method described in Example 1 (Part 1.2).

[0281] Each oligonucleotide assembly and ligation was repeated 6 times for each ligase.

[0282] 3.3 Results and Conclusions

[0283] Wild-type Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 6) and wild-type Shigella phage Shf125875 ligase (SEQ ID NO: 8) are able to ligate a 2'OMe-substituted 5' segment (containing five 2'OMe nucleobases and one deoxynucleobase) to a segment containing only unmodified DNA. Furthermore, although wild-type T4 DNA ligase (SEQ ID NOs: 3 and 4) has difficulty performing this reaction, as shown in Example 2 and confirmed again here, a number of mutations at positions 368 and 371 confer the ability to ligate a 2'-OMe-substituted 5' segment (containing five 2'OMe nucleobases and one deoxynucleobase) to a segment containing only unmodified DNA at the ligase (SEQ ID NOs: 10-19).

[0284] Example 4: Assembly and ligation of 2'MOE ribose-modified and 5-methylpyrimidine-modified oligonucleotide segments using mutant DNA ligase

[0285] 4.1 Materials

[0286] Modified oligonucleotide segments as described in Table 5 below were synthesized by standard solid phase based methods.

[0287] Table 5

[0288] Segment sequence Molecular weight Mass (mg) %purity Central section 5'-(p)dCdCdTdCdGdG-3' 2044.122 39 98.96 MOE 3'-segment 5'-(p)dCdTmTmAmCmCmT-3' 2699.679 52 97.79 MOE 5'-segment 5'-mGmGmCmCmAdAdA-3' 2644.771 48 99.13

[0289] (p) = phosphate, mX = MOE base, dX = DNA base

[0290] All 5-methylpyrimidines.

[0291] Mutant DNA ligases (SEQ ID NOs: 20-28) based on wild-type Enterobacteriaceae phage CC31 ligase, wild-type T4 ligase, and wild-type Shigella phage Shf125875 ligase were each fused to a cellulose binding domain (CBD) at the N-terminus using standard cloning, expression, and extraction methods. The CBD-fused ligases were bound to PERLOZA beads as described in 1.4 to produce a bead slurry. To release the ligases from the PERLOZA beads, 2 µl of TEV protease was added to the slurry and incubated overnight at 4°C. The cleaved proteins (now lacking the cellulose binding domain) were collected by centrifugation at 4000 rpm for 10 minutes.

[0292] 4.2 Methods

[0293] The reaction was set up as follows:

[0294] The final concentration of the central region is 20µM.

[0295] MOE 3' segment final 20µM

[0296] MOE 5' segment final 20µM

[0297] Template final 20µM

[0298] NEB T4 DNA Ligase Buffer 5µl

[0299] Mutant DNA ligase 15µl

[0300] HO to prepare a final reaction volume of 50 µl.

[0301] All components were mixed and vortexed before adding DNA ligase. The reaction was incubated at 35°C for 1 hour. After 1 hour, the reaction was stopped by heating at 95°C for 5 minutes in a PCR block.

[0302] Samples were analyzed by HPLC and LCMS to confirm product identity according to the HPLC protocol used in Example 1 (Section 1.2). A control of commercially available NEB T4 DNA ligase and a negative control (H2O substituted for any ligase) were included.

[0303] 4.3 Results and Conclusions

[0304] Figure 10HPLC traces for a control reaction and a reaction catalyzed by SEQ ID NO:23 (clone A4 - mutant Enterobacteriaceae phage CC31 ligase) are shown. Product and template co-elute using this HPLC method, so the product appears as an increase in the peak area of the product + template peak. In the mutant ligase trace (b), not only does the product + template peak increase, but two new peaks appear at 10.3 and 11.2 minutes. These peaks correspond to the ligation of the central segment with either the MOE 5' segment or the MOE 3' segment. Furthermore, the input segment peak is substantially smaller than the control, consistent with the increase in product and intermediate peaks. The NEB commercially available T4 ligase trace (a) shows a slight increase in the template + product peak area, with a concomitant decrease in a small amount of intermediate ligation product and the input oligonucleotide segment. However, the mutant ligase (SEQ ID NO:23) shows a substantially larger product + template peak area and a concomitant decrease in the peak area of the input oligonucleotide segment. Thus, for the 2'MOE substituted segment, the mutant ligase (SEQ ID NO: 23) was a much more efficient ligase than the commercial T4 DNA ligase. Similar improvements were demonstrated for the other mutant ligases (SEQ ID NOs: 20, 21, 24-28).

[0305] Example 5: Effects of different nucleotide pairings at the junction site

[0306] 5.1 Materials

[0307] Using standard cloning, expression, and extraction methods, a mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was produced with a cellulose binding domain (CBD) fused to the N-terminus. The extracted CBD-mutant Enterobacteriaceae phage CC31 ligase fusion protein was added to 25 ml of PERLOZA 100 (PERLOZA) cellulose beads and shaken at 20°C for 1 hour. The PERLOZA beads were then collected and washed with 250 ml of buffer (50 mM Tris pH 8.0, 200 mM NaCl, 0.1% Tween 20, 10% glycerol), followed by 250 ml of PBS, and finally resuspended in 10 ml of PBS (10 mM PO4 3-, 137 mM NaCl, 2.7 mM KCl pH 7.4). To analyze protein expression, 15 µl of PERLOZA bead slurry was mixed with 5 µl of SDS loading buffer and incubated at 80°C for 10 minutes, then run on an SDS PAGE gradient gel (4-20%) according to standard protocols. To release the ligase from the beads, 70 µl of TEV protease was added and incubated at 4°C overnight with shaking. The ligase was collected by washing the digested beads with 80 ml of PBS. The ligase was then concentrated to 1.2 ml using an Amicon 30 Kd MCO filter.

[0308] The following biotinylated DNA template oligonucleotides (Table 6) and DNA segment oligonucleotides (Table 7) were synthesized by standard solid phase methods. Note that the nucleotides in bold are the nucleotides present at the ligation site (i.e., those nucleotides that are ligated together in the ligation reaction - Table 7; and those nucleotides that are complementary to those ligated together via the ligation reaction - Table 6).

[0309] Table 6:

[0310]

[0311] Table 7

[0312]

[0313] (p) = phosphate.

[0314] NB Please note that unlike the previous examples, the ligation reaction in this example involves joining two segments together: the 5'-segment and the 3'-segment, ie, there is no central segment.

[0315] 5.2 Methods

[0316] The reaction was set up as follows:

[0317] Table 8

[0318]

[0319] For each 50 µL reaction:

[0320] 3' segment (1 mM stock, final 20µM) 1µl

[0321] 5' segment (1 mM stock, final 20µM) 1µl

[0322] Template (1 mM stock, final 20µM) 1µl

[0323] NEB DNA Ligase Buffer (for T4 Ligase) 5µl

[0324] Mutant CC31 DNA ligase (0.45 mM stock, final 90 µM) 10 µl

[0325] Add H2O to 50µl.

[0326] Each reaction mixture was incubated at 35° C. for 30 minutes and 1 hour. Each reaction was terminated by heating at 95° C. for 5 minutes and subjected to HPLC analysis.

[0327] 5.3 Results and Conclusions

[0328] In HPLC analysis, all reactions produced product peaks after 1 hour incubation. Therefore, the ligation method works for all combinations of nucleotides at the joint to be connected. Optimization to increase product yield is possible, but not necessary, as the results are clear and it is clear that the reaction works for all combinations of nucleotides at the joint to be connected.

[0329] Example 6: Effects of different modifications at the attachment site

[0330] 6.1 Materials

[0331] Mutant Enterobacteria phage CC31 ligase (SEQ ID NO: 23) was produced as described in 5.1, and Chlorella virus DNA ligase (SEQ ID NO: 29, commercially available as SplintR ligase, NEB) was purchased.

[0332] The following biotinylated template oligonucleotides and segment oligonucleotides were synthesized by standard solid phase methods (Table 9).

[0333] Table 9

[0334] name Sequence (the junction is modified in bold) template 5'-BiotinTTTAGGTAAGCCGAGGTTTGGCC-3' (SEQ ID NO:2) 5' segment (WT) 5'-GGCCAAA-3' 5' segment (Mo1) 5'-GGCCAA(OMe)A-3' 5' segment (Mo2) 5'-GGCCAA(F)A-3' Central segment (WT) 5'-(p)CCTCGG-3' Central section (Mo4A) 5'-(p)(OMe)CCTCGG-3' Central section (Mo5A) 5'-(p)(F)CCTCGG-3' Central section (Mo7) 5'(p)(Me)CCTCGG-3' 3' segment (WT) 5'-(p)CTTACCT-3' 3' segment (Mo8) 5'-(p)(Me)CTTACCT3'

[0335] OMe indicates a 2' methoxy substitution on the ribose ring

[0336] F indicates a 2' fluorine substitution on the ribose ring

[0337] All remaining sugar residues are deoxyribose residues

[0338] Me indicates 5-methylcytosine.

[0339] 6.2 Methods

[0340] The reaction was set up as follows:

[0341] Table 10

[0342]

[0343] For each 50 µL reaction:

[0344] 3'-segment (1 mM, stock, final 20µM) 1µl

[0345] Central fraction (1 mM, stock, final 20µM) 1µl

[0346] 5'-segment (1 mM, stock, final 20µM) 1µl

[0347] Template (1 mM, stock, final 20µM) 1µl

[0348] Add H2O to 50µl.

[0349] For mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23)

[0350] DNA ligase buffer (50 mM Tris-HCl, 1 mM DTT) 5 µl

[0351] Mutant CC31 DNA ligase (0.45 mM) 10 µl

[0352] MnCl2 (50mM) 5µl

[0353] ATP (10 mM) 10µl.

[0354] However, for Chlorella viral DNA ligase (SEQ ID NO: 29, commercially available as SplintR ligase, NEB)

[0355] NEB DNA Ligase Buffer (for Chlorella) 5µl

[0356] 2µl of Chlorella virus DNA ligase.

[0357] Each reaction mixture was incubated at 20° C. for 1 hour. Each reaction was terminated by heating at 95° C. for 10 minutes. HPLC analysis was performed using the method of Example 1.

[0358] 6.3 Results and Conclusions

[0359] In the HPLC analysis, all reactions produced product peaks. Therefore, the ligation method worked for all modification combinations tested at the joint to be connected. Optimization to increase product yield is possible, but not required, as the results are clear and it is clear that the reaction worked for all modification combinations tested at the joint to be connected.

[0360] Example 7: Ability to construct larger oligonucleotides using different numbers of segments

[0361] 7.1 Materials

[0362] The mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was produced as described in 5.1.

[0363] The following biotinylated template DNA oligonucleotides and DNA segment oligonucleotides (Table 11) were synthesized by standard solid phase methods.

[0364] Table 11

[0365] name sequence template 5'-biotinTTTGGTGCGAAGCAGAAGGTAAGCCGAGGTTTGGCC-3' (SEQ ID NO:47) 5' segment (1) 5'-GGCCAAA-3' Central section (2) 5'-(p)CCTCGG-3' Central segment / 3' segment (5) 5'-(p)TCTGCT-3' Central section (3) 5'-(p)CTTACCT-3' 3' segment (4) 5'-(p)TCGCACC-3'

[0366] (p) = phosphate,

[0367] 7.2 Methods

[0368] The reaction was set up as follows:

[0369] Table 12

[0370] 5' segment Central section 3' segment Total number of segments 1 2 and 3 5 4 1 2, 3 and 5 4 5

[0371] Reactions were run in phosphate buffered saline (pH = 7.04) in a total volume of 100µl and set up as follows:-

[0372] Template (final 20µM)

[0373] Each segment (final 20µM)

[0374] MgCl2 (final 10 mM)

[0375] ATP (final 100µM)

[0376] Mutant CC31 DNA ligase (final 25 µM).

[0377] Each reaction was incubated overnight at 28° C. and then terminated by heating for 1 minute at 94° C. The products were analyzed by HPLC mass spectrometry.

[0378] 7.3 Results and Conclusions

[0379] The reaction of using 4 sections produced the product of the complete connection of 27 base pairs in length.The reaction of using 5 sections produced the product of 33 base pairs in length.In two cases, the quality of the product observed is consistent with the expection of the sequence of expectation.In sum, it is obviously possible to assemble multiple sections to produce the oligonucleotide of desired length and the sequence defined by appropriate complementary template sequence.

[0380] Example 8: Assembly and ligation of 5-10-5 segments to form a gapmer, wherein the 5' and 3' segments comprise: (i) a 2'-OMe ribose sugar modification, (ii) a phosphorothioate linkage, or (iii) a 2'-OMe ribose sugar modification and a phosphorothioate linkage; and wherein the central segment is unmodified DNA.

[0381] 8.1 Materials

[0382] The mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was produced as described in 5.1. The following biotinylated template DNA oligonucleotides and segment oligonucleotides were synthesized by standard solid phase methods (Table 13).

[0383] Table 13

[0384] name sequence template 5'-biotinTTTGGTGCGAAGCAGACTGAGGC-3' (SEQ ID NO:30) 3' segment (3OMe) 5'-(p)(OMe)G(OMe)C(OMe)C(OMe)T(OMe)C-3' 3' segment (3PS+OMe) 5'-(p)(OMe)G*(OMe)C*(OMe)C*(OMe)T*(OMe)C-3' 3' segment (3PS) 5'-(p)G*C*C*T*C-3' Central section (D) 5'-(p)AGTCTGCTTC-3' 5' segment (5OMe) 5'-(OMe)G(OMe)C(OMe)A(OMe)C(OMe)C-3' 5' segment (5PS+OMe) 5'-(OMe)G*(OMe)C*(OMe)A*(OMe)C*(OMe)C-3' 5' segment (5PS) 5'-G*C*A*C*C-3'

[0385] OMe indicates a 2' methoxy substitution on the ribose ring

[0386] All remaining sugar residues are deoxyribose residues

[0387] *Phosphorothioate.

[0388] 8.2 Methods

[0389] The reaction was set up as follows:

[0390] Table 14

[0391] reaction 3' segment Central section 5' segment 1 3PS D 5PS 2 3OMe D 5OMe 3 3PS+OMe D 5PS+OMe

[0392] Reactions 1, 2 and 3 were each set up in phosphate buffered saline with the following components in a final volume of 100 µl:

[0393] 3' segment final 20µM

[0394] The final concentration of the central region is 20µM.

[0395] 5' segment final 20µM

[0396] Template final 20µM

[0397] MgCl2 final 10 mM

[0398] ATP final 50µM

[0399] The enzyme concentration was 25 µM.

[0400] Each reaction mixture was incubated overnight at 20° C. Each reaction was terminated by heating at 95° C. for 10 minutes and analyzed by HPLC mass spectrometry.

[0401] 8.3 Results and Conclusions

[0402] Product oligonucleotides corresponding to the successful ligation of all three fragments were produced in all three reactions. Thus, the mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was able to link the three segments together to form a 'gapmer' in which the 5' and 3' wings had phosphorothioate backbones, while the central region had a phosphodiester backbone, and all sugar residues in the gapmer were deoxyribose residues. The Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was also able to link the three segments together to form a 'gapmer' in which the 5' and 3' wings had 2'-methoxyribose (2'-OMe) residues, while the central region had deoxyribose residues, and all connections were phosphodiester bonds. Finally, Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) is able to join the three segments together to form a 'gapmer' in which the 5' and 3' wings have combined modifications (phosphorothioate backbone and 2'-methoxyribose residues) while the central region has deoxyribose residues and phosphodiester bonds.

[0403] Example 9: Assembly and ligation of segments containing locked nucleic acids (LNA)

[0404] 9.1 Materials

[0405] The mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was produced as described in 5.1. The mutant S. aureus NAD-dependent ligase (NAD-14) was produced as described in 13.1.

[0406] The following biotinylated template DNA oligonucleotides and segment oligonucleotides were synthesized by standard solid phase methods (Table 15).

[0407] Table 15

[0408] name sequence template 5'-biotinTTTGGTGCGAAGCAGACTGAGGC-3' (SEQ ID NO:30) 5'-segment 5'-GCCTCAG-3' LNA 5'-segment (oligo 1) 5'-GCCTCA(LNA)G-3' Central section 5'-(p)TCTGCT-3' LNA central segment (oligo 2) 5'-(p) (LNA)TCTGCT-3'

[0409] (p) = phosphate

[0410] LNA = Locked Nucleic Acid

[0411] 9.2 Methods

[0412] The reaction was set up as follows:

[0413] Reaction volume 100µl

[0414] Template final 20µM

[0415] Enzyme final 25µM

[0416] All oligonucleotides were concentrated to a final concentration of 20 µM.

[0417] Different enzymes (mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) or NAD-14), divalent cations (Mg 2+ or Mn 2+ ) and the combinations of oligonucleotide segments described in Table 16, and set up the reaction.

[0418] Table 16

[0419] enzymes SEQ ID NO:23 SEQ ID NO:23 NAD-14 NAD-14 Divalent cations <![CDATA[10 mM MgCl2]]> <![CDATA[10 mM MnCl2]]> <![CDATA[10 mM MgCl2]]> <![CDATA[10 mM MnCl2]]> cofactors 100µM ATP 100µM ATP 100µM NAD 100µM NAD buffer PBS, pH = 7.04 PBS, pH = 7.04 <![CDATA[50 mM KH2PO4, pH 7.5]]> <![CDATA[50 mM KH2PO4, pH 7.5]]> 5' segment + central segment product product product product Oligo 1 + central segment product product product product 5' segment + oligo 2 product product product product Oligo 1 + Oligo 2 No product No product No product No product

[0420] Each reaction mixture was incubated overnight at 28° C. Each reaction was terminated by heating at 94° C. for 1 minute and analyzed by HPLC mass spectrometry.

[0421] 9.3 Results and Conclusions

[0422] Product oligonucleotides were produced in control reactions (unmodified oligonucleotides only) in which a single locked nucleic acid was included in one segment at the junction junction, regardless of whether it was on the 3' or 5' side of the junction. No product was detected when the locked nucleic acid was included on both sides (Oligo 1 + Oligo 2). The data were similar for both enzymes and regardless of whether Mg2+ or Mn2+ was used.

[0423] Enzyme mutagenesis and / or selection screens can be performed to identify enzymes capable of ligating segments having locked nucleic acids on both the 3' and 5' sides of the junction.

[0424] Example 10: In the presence of Mg 2+ or Mn 2+ The three segments (7-6-7) were assembled and ligated using a variant of the Enterobacteriaceae phage CC31 ligase in the presence of α-glucose to form a gapmer, in which the 5' and 3' segments contained a 2'MOE ribose sugar modification and all linkages were phosphorothioate bonds.

[0425] 10.1 Phosphorothioate Bond Formation

[0426] To determine whether the mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was able to ligate modified oligonucleotide segments having a phosphorothioate backbone, a 2'MOE ribose sugar modification, and a 5-methylated pyrimidine base, reactions were performed using the oligonucleotide segments shown in Table 15. 2+ and Mn 2+ The reaction is performed in the presence of ions.

[0427] 10.2 Materials

[0428] Oligonucleotides were chemically synthesized using standard methods as follows:

[0429] Table 17

[0430] name sequence 5' segment 2'-MOE PS 5'-mG*mG*mC*mC*mA*dA*dA-3' Central section PS 5'-(p)*dC*dC*dT*dC*dG*dG -3' 3' segment 2'-MOE PS 5'-(p)*dC*dT*mU*mA*mC*mC*mU-3' Biotinylated template 5'-Biotin dT dT dT dA dG dG dT dA dA dG dC dC dG dA dG dG dT dT dT dG dG dC dC-3' (SEQ ID NO: 2)

[0431] (p)* = 5'-phosphorothioate, * = phosphorothioate bond, mX = MOE base, dX = DNA base

[0432] All segments and products have 5-methylpyrimidines (except template)

[0433] mT and m(Me)U are considered equivalent.

[0434] NB When hybridized to the biotinylated template shown in Table 17, the target 2'MOE PS molecule generated by ligating the segments in Table 17 is:

[0435] 5'-mG*mG*mC*mC*mA*dA*dA*dC*dC*dT*dC*dG*dG*dC*dT*mU*mA*mC*mC*mU-3'(SEQ ID NO:1)

[0436] Purified mutant Enterobacteriaceae phage CC31 ligase (SEQ ID NO: 23) was prepared as described in Example 5.1. HPLC analysis was performed.

[0437] 10.3 Oligonucleotide Assembly and Ligation Using Enterobacteriaceae Phage CC31 Ligase Variant (SEQ ID NO: 23) method

[0438] The reactants were prepared as follows:

[0439] MgCl2 reaction

[0440]

[0441] MnCl2 reaction

[0442]

[0443] The final reaction contained 20 µM of each segment and template, 5 mM MgCl2 or 5 mM MnCl2, 1 mM ATP, 50 mM Tris-HCl, 10 mM DTT (pH 7.5), and 4.9 µM ligase. An additional reaction without enzyme was prepared and used as a negative control. The reaction was incubated at 25°C for 16 hours and then quenched by heating to 95°C for 5 minutes. The precipitated protein was clarified by centrifugation, and the sample was analyzed by HPLC.

[0444] 10.4 Results and Conclusions

[0445] The oligonucleotide of the present invention is connected with the oligonucleotide of the template of the oligonucleotide of the template of the template.Product, template and segment oligonucleotide obviously decompose and do not observe connection in the control chromatogram.5mM MgCl is arranged in the presence of the ligase reaction that carries out causes the formation of the intermediate product that forms from the connection of 5 ' segment and central section, but does not detect full-length product.MnCl is arranged in the presence of the ligase reaction that carries out has produced full-length product and intermediate (5 ' segment+central section intermediate).Two kinds of ligase reactions show that unconnected oligonucleotide segment is retained.But, in order to maximize product yield, the optimization of scheme is possible.

[0446] Example 11: In the presence of natural Mg 2+ Three segments (7-6-7) were assembled and ligated using wild-type Chlorella viral DNA ligase in the presence of α-glucose to form a gapmer, in which the 5' and 3' segments contained 2'MOE ribose sugar modifications and all linkages were phosphorothioate bonds.

[0447] 11.1 Materials

[0448] To determine whether Chlorella viral DNA ligase (SEQ ID NO: 29, commercially available as SplintR ligase, NEB) was able to ligate modified oligonucleotide segments having a phosphorothioate backbone, a 2'MOE ribose sugar modification, and a 5-methylated pyrimidine base, reactions were performed using the oligonucleotide segments shown in Example 10.2, Table 17. Reactions were performed at 25°C, 30°C, and 37°C to investigate the effect of temperature on enzyme activity.

[0449] 11.2 Oligonucleotide Assembly and Ligation Method Using Commercially Available Chlorella Virus DNA Ligase (SEQ ID NO: 29) Law

[0450] Dissolve each oligonucleotide segment and template in nuclease-free water as detailed below:

[0451]

[0452] The reactants were prepared as follows:

[0453]

[0454] The final reaction contained 20 µM of each segment and template, 50 mM Tris-HCl, 10 mM MgCl2, 1 mM ATP, 10 mM DTT (pH 7.5), and 2.5 U / µl ligase. Reactions were incubated at 25°C, 30°C, and 37°C. Additional reactions without enzyme were prepared and used as negative controls. After a 16-hour incubation, the reaction was quenched by heating to 95°C for 10 minutes. The precipitated protein was clarified by centrifugation, and samples were analyzed by HPLC.

[0455] 11.3 Results and Conclusions

[0456] The product, template, and segment oligonucleotides were clearly resolved in the control chromatogram, and no ligation was observed. HPLC analysis of the ligase reaction showed that unligated oligonucleotide segments were retained, but Chlorella viral DNA ligase was able to successfully ligate the segments. Ligase activity increased with increasing temperature. At 25°C, Chlorella viral DNA ligase was able to successfully ligate the 5' segment and the central segment, but no full-length product was observed. At 30°C and 37°C, the full-length product was detected, in addition to intermediates formed from the 5' segment and the central segment.

[0457] Example 12: Screening a panel of 15 ATP and NAD ligases for activity in ligating three segments (7-6-7) to form a gapmer, wherein the 5' and 3' segments contain 2'MOE ribose sugar modifications and all linkages are phosphorothioate bonds

[0458] 12.1 Materials

[0459] The wild-type ATP- and NAD-dependent ligases described in Tables 18 and 19 were each fused to the CBD at the N-terminus. The genes were synthesized using standard cloning, expression, and extraction methods, cloned into pET28a, and expressed in E. coli BL21 (DE3).

[0460]

[0461]

[0462] The CBD-ligase fusion was conjugated to PERLOZA beads as described in 1.4 with the following modifications. CBD-ligase fusion proteins were cultured from a single colony of BL21(DE3) cells (NEB) and grown in 50 mL expression cultures. Cells were harvested by centrifugation, resuspended in 5-10 mL Tris-HCl (50 mM, pH 7.5), and lysed by sonication. The lysate was clarified by centrifugation, and 1 mL of PERLOZA 100 (PERLOZA) beads (50% slurry, pre-equilibrated with 50 mM Tris-HCl pH 7.5) was added to the supernatant, which was shaken at 20°C for 1 hour. The PERLOZA cellulose beads were then collected and washed with 30 ml of buffer (50 mM Tris pH 8.0, 200 mM NaCl, 0.1% Tween 20, 10% glycerol), followed by 10 ml of Tris-HCl (50 mM, pH 7.5), and finally resuspended in 1 mL of Tris-HCl (50 mM, pH 7.5). To analyze protein expression, 20 µl of PERLOZA bead slurry was mixed with 20 µl of SDS loading buffer and incubated at 95°C for 5 minutes before running on an SDS PAGE gradient gel (4-20%) according to standard protocols.

[0463] 12.2 2'MOE and Phosphorothioate-Modified Oligonucleotide Assembly and Ligation Methods

[0464] Modified oligonucleotide segments with phosphorothioate backbones, 2'MOE ribose sugar modifications, and 5-methylated pyrimidine bases as shown in Table 17 of Example 10.2 were used. Each oligonucleotide segment and template were dissolved in nuclease-free water as detailed below:

[0465]

[0466] Prepare ATP assay mix as follows:

[0467]

[0468] Prepare the NAD assay mixture as follows:

[0469]

[0470] Each immobilized protein (40 µl, 50% PERLOZA bead slurry) was transferred to a PCR tube. The beads were pelleted by centrifugation, and the supernatant was removed by aspiration. The assay mixture (40 µl) was added to each reaction (the final reaction contained 20 µM of each segment and template, 50 mM Tris-HCl, 10 mM MgCl2, 1 mM ATP or 100 µM NAD, 10 mM DTT (pH 7.5), and 40 µl of ligase on PERLOZA beads). A reaction containing no protein served as a negative control. The reaction was incubated at 30°C for 18 hours and then quenched by heating to 95°C for 10 minutes. The precipitated protein was clarified by centrifugation, and the samples were analyzed by HPLC.

[0471] 12.3 Results and Conclusions

[0472] The product, template, and segment oligonucleotides were clearly resolved in the control chromatogram and no ligation was observed. HPLC analysis of the ligase reactions showed that all proteins catalyzed successful ligation of the 5' segment and the central segment to form an intermediate product, but only some ligases catalyzed ligation of all three segments to produce the full-length product described in Table 20. The NAD-dependent ligase from Staphylococcus aureus (SaNAD, SEQ ID NO: 61) produced the most full-length product. Optimization to increase product yield is possible and within the skill of the artisan.

[0473] Table 20

[0474]

[0475] * Conversion was calculated from the HPLC peak area relative to the template, which was not consumed in the reaction and used as an internal standard. Conversion = product area / (template + product area) * 100.

[0476] Example 13: Semi-continuous ligation reaction

[0477] 13.1 Materials

[0478] A mutant S. aureus ligase (NAD-14) fused to the CBD at the N-terminus was produced using standard cloning, expression, and extraction methods. The CBD-NAD-14 mutant ligase was then bound to PERLOZA beads: 50 ml of protein lysate was added to 7.5 ml of PERLOZA beads, incubated at room temperature for 1 hour, and the beads were collected in a glass column (BioRad Econo-Column 10 cm length, 2.5 cm diameter #7372512). The beads were washed with 200 ml of buffer Y (50 mM Tris 8, 500 mM NaCl, 0.1% Tween 20, 10% glycerol), followed by 200 ml of buffer Z (50 mM Tris 8, 200 mM NaCl, 0.1% Tween 20, 10% glycerol) and 200 ml of PBS. The estimated concentration of mutant NAD-14 ligase on the beads was 69 µM ligase / ml beads.

[0479] The following template DNA oligonucleotides and segment oligonucleotides were synthesized by standard solid phase methods (Table 21).

[0480] Table 21

[0481] Segment sequence Molecular weight %HPLC purity Template 3 5'TTTGGTGCGAAGCAGACTGAGGC-3' (SEQ ID NO:30) Central section 5'-(p)dTdCdTdGdCdT-3' 1865.4 97.5 MOE 3'-segment 5'-(p)dTdCmGmCmAmCmC-3' 2546.7 98.0 MOE 5'-segment 5'-mGmCmCmUmCdAdG-3' 2492.7 98.3

[0482] (p) = phosphate, mX = MOE base, dX = DNA base

[0483] All 5-methylpyrimidines

[0484] All linkages are phosphodiester bonds.

[0485] A "tri-template hub'" (approximately 24 kDa) was produced, comprising a support material called "hub'" and three template sequences ( Figure 11 ). Each template copy is covalently linked to the "hub" at its own separate attachment point. The "three-template hub" molecule is a higher molecular weight than the target (product) oligonucleotide (100% complementary to the template sequence), thereby allowing it to be retained when separating impurities and products from the reaction mixture. It should be noted that in this particular case, the template sequence is SEQ ID NO: 30, and three copies are attached to the hub. In the following examples, three-template hubs were also produced, but with different template sequences. Therefore, as the template sequence varies between examples, the three-template hub also varies.

[0486] The following reaction mixture was prepared (total volume 5 ml):

[0487] 250µl 1 M KH2PO4, pH 7.5 (final 50 mM)

[0488] 108µl 0.07011 M central fragment (final 1.5 mM)

[0489] 137 µl 0.05481 M 3'-segment (final 1.5 mM)

[0490] 168µl 0.04461 M 5'-segment (final 1.5 mM)

[0491] 750µl 0.00387M Hub (template) (final 0.55 mM)

[0492] 350µl 50mM NAD + (Final 3.5 mM)

[0493] 1000µl 50mM MgCl2 (final 10 mM)

[0494] 2237µl nuclease-free H2O.

[0495] 13.2 Methods

[0496] As in Figure 12 A semi-continuous system was established as shown in .

[0497] 4 ml of PERLOZA beads and immobilized mutant NAD-14 ligase were loaded into a Pharmacia XK16 column (B). Using a column water compartment, the column temperature was maintained at 30°C using a water bath and a peristaltic pump (C). The beads were equilibrated by running 120 ml (30 x column volume) of 50 mM KH2PO4 buffer (at pH 7.5) at 1 ml / min for 120 minutes. Flow through the Pharmacia XK16 column was established using an AKTA probe pump A1 (A).

[0498] After post balance, 5 ml reaction mixture (mixed uniformly by vortex) is loaded onto post, collected in reservoir test tube (D) and passed through post recirculation using AKTA probe pump A1. Reaction mixture is recirculated 16 hours in the system at a flow velocity of 1 ml / min in continuous circulation mode. After 30 minutes, 60 minutes, 90 minutes, 4 hours, 5 hours, 6 hours, 7 hours, 14 hours and 16 hours, sample is collected for HPLC analysis.

[0499] 13.3 Results and Conclusions

[0500] Table 22

[0501] sample 5'-segment (%) Central section (%) 3'-segment (%) 5'+ central intermediate (%) product(%) 30min 7.80 11.00 24.00 39.40 4.40 60min 3.30 4.50 18.80 41.50 18.90 90 minutes 2.10 2.80 15.10 33.80 34.40 4 h 1.80 2.40 9.60 19.20 56.80 5 h 1.72 2.40 8.30 15.70 61.50 6 h 1.69 2.50 8.60 15.90 69.40 7 h 1.70 2.40 8.30 14.80 70.90 14 h 2.00 2.70 1.60 5.40 88.10 16 h 0.80 1.90 1.70 3.30 89.50

[0502] The percentage of each segment, intermediate, and product was expressed as peak area fraction relative to the triple-template hub peak area.

[0503] In summary, the semi-continuous flow reaction worked and after 16 hours, the reaction was almost complete.

[0504] Example 14: Separation of oligonucleotides of different sizes by filtration: a) Separation of a 20-mer oligonucleotide (SEQ ID NO: 1) and a hub comprising three non-complementary 20-mer oligonucleotides (SEQ ID NO: 30); (b) Separation of a 20-mer oligonucleotide (SEQ ID NO: 1) from segmented 6-mer and 8-mer oligonucleotides (see Table 1) and a hub comprising three complementary 20-mer oligonucleotides (SEQ ID NO: 2).

[0505] 14.1 Materials

[0506] All oligonucleotides used were synthesized by standard solid phase methods.

[0507] Using the three-template hub ( Figure 11 ).

[0508] As shown in Tables 23 and 24, a variety of filters with different molecular weight cutoffs and from different manufacturers were used.

[0509] 14.2 Methods

[0510] 14.2.1 Blind-End Filtration Setup and Protocol for Screening Polymer Membranes: (Protocol 1)

[0511] As in Figure 13 A blind-end filtration setup was set up as shown in , comprising a MET blind-end filtration unit placed in a water bath, which was placed on a magnetic stirrer and a hot plate. The pressure within the unit was provided by a nitrogen flow.

[0512] First, place the membrane sample to be tested (14 cm 2) are cut to appropriate size and placed in the cell. First, the membrane is regulated with HPLC-grade water (200 ml), and then regulated with PBS buffer (200 ml). The cell is then depressurized, the remaining PBS solution is removed and replaced with a solution containing oligonucleotide (40 ml oligonucleotide in PBS, 1 g / L concentration). The cell is placed on a hot stirring plate and the solution is heated to the desired temperature while stirring with magnetic agitation. Pressure is applied to the cell (reaching approximately 3.0 bar; record actual pressure in each case). Stop or continue stirring of the solution, and collect permeate solution (approximately 20 ml) and analyze by HPLC. Record flow. The system is then depressurized to allow sampling and HPLC analysis of the retentate solution. Then more PBS buffer (20 ml) is added to the filter cell and the previous operation is repeated 3 times. Finally, the membrane is washed with PBS buffer.

[0513] All samples were analyzed by HPLC without any dilution.

[0514] 14.2.2 Cross-Flow Filtration Setup and Protocol for Screening Polymer Membranes: (Protocol 2)

[0515] As in Figure 14 The cross-flow filtration setup is set up as shown in . A feed vessel (1) consisting of a conical flask contains the oligonucleotide solution to be purified. The solution is pumped to a cross-flow filtration unit (4) using an HPLC pump (2), while a heating plate (3) is used to maintain the temperature within the unit. A gear pump (6) is used to recirculate the solution within the unit. A pressure gauge (5) allows the pressure to be read during the experiment. A sample of the retentate solution is taken from a sampling valve (7), and a sample of the permeate solution is taken from a permeate collection vessel (8).

[0516] The sample of the film to be tested is first cut to an appropriate size and placed in the unit. The system is washed with PBS solution (100 ml). The temperature of the solution is adjusted to the desired set point. A solution (7.5 ml, 1 g / L) of the oligonucleotide product containing PBS is fed into the system. The PBS solution is then pumped into the system using an HPLC pump at a flow velocity (usually, 3 ml / min) that matches the flow velocity of the permeate solution. Pressure is recorded using a pressure gauge. Per 5 diafiltration volumes are sampled for HPLC analysis of the retentate solution. Each diafiltration volume is sampled for HPLC analysis of the permeate solution. After 20 diafiltration volumes, the experiment is stopped.

[0517] In the case of an experiment using a Snyder membrane (lot 120915R2) with a 5 kDa molecular weight cutoff and SEQ ID NO: 46 and SEQ ID NO: 30, the above method was modified as follows. First, a sample of the membrane to be tested was cut to an appropriate size and placed in the cell. The system was washed with a potassium phosphate solution (100 ml, 50 mM, pH 7.5). The temperature of the solution was adjusted to the desired set point. A solution containing the oligonucleotide product in potassium phosphate (approximately 1 g / L) was added to ethylenediaminetetraacetic acid (EDTA) (230 μL, 500 mM solution). The solution was then fed into the system. Potassium phosphate buffer was then pumped into the system using an HPLC pump at a flow rate (typically 4 ml / min) that matched the flow rate of the permeate solution. A pressure gauge was used to record the pressure. The retentate solution was sampled for HPLC analysis every 5 diafiltration volumes. The permeate solution was sampled for HPLC analysis every diafiltration volume. The experiment was stopped after 15 diafiltration volumes.

[0518] 14.3 Results

[0519] Table 23: Results of blind-end filtration experiments according to Scheme 1 (14.2.1)

[0520]

[0521] MWCO = molecular weight cutoff.

[0522] In experiments using a 10 kDa MWCO NADIR membrane at 60°C, clear separation between the product sequence (SEQ ID NO: 1) and the non-complementary three-template hub (comprising SEQ ID NO: 30) was demonstrated. Figure 15 Shown are a) the chromatogram of the retentate solution after 2 diafiltration volumes, which was retained in the filtration unit and contained mainly the tri-template hub; and b) the chromatogram of the product-rich permeate solution after 2 diafiltration volumes.

[0523]

[0524] In experiments using a 5 kDa MWCO Snyder membrane at 50°C and 3.0 bar pressure, clear separation between the segment sequence (see Table 1) and the complementary three-template hub (comprising SEQ ID NO: 2) and product (SEQ ID NO: 1) was demonstrated. Figure 16 Shown are a) chromatograms of the retentate solution after 20 diafiltration volumes, which contained primarily the tri-template hub and product; and b) chromatograms of the permeate after 20 diafiltration volumes, which contained primarily segmented oligonucleotides.

[0525] In experiments using a 5 kDa MWCO Snyder membrane at 80°C and 3.1 bar pressure, clear separation between the complementary three-template hub (comprising SEQ ID NO: 2) and the product (SEQ ID NO: 1) was demonstrated. Figure 17 Shown are a) the chromatogram of the retentate solution after 20 diafiltration volumes, which contains only the tri-template hub; and b) the chromatogram of the permeate solution after 2 diafiltration volumes, which contains only the product.

[0526] 14.4 Conclusion

[0527] Oligonucleotides of different lengths and molecular weights can be separated using filtration. As shown above, the type of membrane and conditions (such as temperature) can affect the separation level. For the set of oligonucleotides of given different lengths / molecular weights, suitable membranes and conditions can be selected to allow required separation. For example, we have confirmed that the segment oligonucleotides (short polymers of 6 and 8 nucleotide lengths) as shown in Table 1 can be separated from product oligonucleotides (20-aggregates with SEQ ID NO: 1) and three-template hubs (20-aggregates comprising 3 SEQ ID NO: 2, which are attached to a solid support), and the product oligonucleotides and three-template hubs can be separated from each other again.

[0528] Conclusion

[0529] We have demonstrated that it is possible to synthesize oligonucleotides, including oligonucleotides with many therapeutically relevant chemical modifications, in solution by assembling short oligonucleotide segments on complementary templates, ligating the segments together, and separating the product oligonucleotide from impurities and its complementary template in an efficient process that is scalable and suitable for large-scale production of therapeutic oligonucleotides.

[0530] By synthesizing oligonucleotides in solution, we have avoided the amplification constraints generated by solid phase methods. In using the inherent properties of DNA to specifically recognize complementary sequences and to bind complementary sequences with an affinity that reflects the fidelity of the complementary sequence and the length of the complementary sequence, we have been able to produce highly pure oligonucleotides without the need for chromatography. This improves the efficiency of the production method and the scalability of the method. By recovering the template in an unchanged state in the separation method, we can reuse the template for future synthesis rounds, and so have avoided the economic consequences of having to prepare 1 equivalent of template for each equivalent of formed product oligonucleotide.

[0531] Finally, although wild-type ligases are known to efficiently ligate normal DNA, we have demonstrated that modifications to DNA can lead to reduced ligation efficiency, and that multiple modifications to DNA are additive in their effects on reducing ligation efficiency, which in some cases can render the DNA ligase completely ineffective. We have demonstrated that ligation efficiency can be restored through appropriate mutagenesis and evolution of DNA ligases, and that appropriately modified DNA ligases are efficient catalysts for synthesizing oligonucleotides containing multiple modifications.

[0532] Sequence Listing

[0533] SEQ ID NO Sequence identifier 1 Example 1 Desired Product Oligonucleotide Sequence ("Target") 2 Examples 1-4, 6, 10 and 11 Template Oligonucleotide Sequences 3 Wild-type T4 DNA ligase protein sequence 4 Wild-type T4 DNA ligase protein sequence (when fused to CBD) 5 Example 2 Target sequence 6 Wild-type Enterobacteriaceae phage CC31 DNA ligase protein sequence 7 Wild-type Enterobacteriaceae phage CC31 DNA ligase protein sequence (when fused to CBD) 8 Wild-type Shigella phage Shf125875 DNA ligase protein sequence 9 Wild-type Shigella phage Shf125875 DNA ligase protein sequence (when fused to CBD) 10 Mutant ligase (T4 backbone) protein sequence 11 Mutant ligase (T4 backbone) protein sequence 12 Mutant ligase (T4 backbone) protein sequence 13 Mutant ligase (T4 backbone) protein sequence 14 Mutant ligase (T4 backbone) protein sequence 15 Mutant ligase (T4 backbone) protein sequence 16 Mutant ligase (T4 backbone) protein sequence 17 Mutant ligase (T4 backbone) protein sequence 18 Mutant ligase (T4 backbone) protein sequence 19 Mutant ligase (T4 backbone) protein sequence 20 Mutant ligase (T4 backbone) protein sequence 21 Mutant ligase (T4 backbone) protein sequence 22 Mutant ligase (T4 backbone) protein sequence 23 Mutant ligase (Enterobacterial phage CC31 backbone - clone A4) protein sequence 24 Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence 25 Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence 26 Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence 27 Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence 28 Mutant ligase (Shigella phage Shf125875 backbone) protein sequence 29 Wild-type Chlorella vulgaris ligase protein sequence 30 Examples 5, 8 and 9 template oligonucleotide sequences 31 Example 5 Template Oligonucleotide Sequence 32 Example 5 Template Oligonucleotide Sequence 33 Example 5 Template Oligonucleotide Sequence 34 Example 5 Template Oligonucleotide Sequence 35 Example 5 Template Oligonucleotide Sequence 36 Example 5 Template Oligonucleotide Sequence 37 Example 5 Template Oligonucleotide Sequence 38 Example 5 Template Oligonucleotide Sequence 39 Example 5 Template Oligonucleotide Sequence 40 Example 5 Template Oligonucleotide Sequence 41 Example 5 Template Oligonucleotide Sequence 42 Example 5 Template Oligonucleotide Sequence 43 Example 5 Template Oligonucleotide Sequence 44 Example 5 Template Oligonucleotide Sequence 45 Example 5 Template Oligonucleotide Sequence 46 Example 14 "20-mer" oligonucleotide sequence 47 Example 7 Template Oligonucleotide Sequence 48 Paramecium chlororaphis virus NE-JV-4 ligase 49 Paramecium viridis virus NYs1 ligase 50 Paramecium chlororaphis virus NE-JV-1 ligase 51 Acanthocystis turfacea Chlorella virus Canal-1 ligase 52 Acanthocystis turfacea Chlorella virus Br0604L ligase 53 Acanthocystis turfacea Chlorella virus NE-JV-2 ligase 54 Acanthocystis turfacea Chlorella virus TN603.4.2 ligase 55 Acanthocystis turfacea Chlorella virus GM0701.1 ligase 56 Synechococcus phage S-CRM01 ligase 57 Marine sediment metagenome ligase 58 Mycobacterium tuberculosis (strain ATCC 25618 / H37Rv) ligase 59 Enterococcus faecalis (strain ATCC 700802 / V583) ligase 60 Haemophilus influenzae (strain ATCC 51907 / DSM 11121 / KW20 / Rd) ligase 61 Staphylococcus aureus ligase 62 Streptococcus pneumoniae (strain P1031) ligase

[0534]

[0535]

[0536]

[0537]

[0538]

[0539]

[0540]

[0541]

[0542]

[0543] . Sequence Listing <110> GlaxoSmithKline Intellectual Property Development Limited <120> Novel method for producing oligonucleotides <130> PB66127 <160> 62 <170> PatentIn version 3.5 <210> 1 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Example 1 Desired Product Oligonucleotide Sequence ("Target") <400> 1 ggccaaacct cggcttacct 20 <210> 2 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Examples 1-4, 6, 10 and 11 Template Oligonucleotide Sequences <400> 2 tttaggtaag ccgaggtttg gcc 23 <210> 3 <211> 487 <212> PRT <213> Enterobacteriaceae phage T4 <400> 3 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Val Ile Asp Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 4 <211> 488 <212> PRT <213> Artificial sequence <220> <223> Wild-type T4 DNA ligase protein sequence (when fused to CBD) <400> 4 Gly Ser Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr 1 5 10 15 Lys Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu 20 25 30 Lys Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile 35 40 45 Lys Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu 50 55 60 Thr Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg 65 70 75 80 Lys Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr 85 90 95 Asp Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg 100 105 110 Asp Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro 115 120 125 Gly Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu 130 135 140 Lys Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys 145 150 155 160 Ala Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp 165 170 175 Asp Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp 180 185 190 Leu Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile 195 200 205 His Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln 210 215 220 Val Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro 225 230 235 240 Glu Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr 245 250 255 Ala Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu 260 265 270 Lys Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu 275 280 285 Val Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg 290 295 300 Phe Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu 305 310 315 320 Ile Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr 325 330 335 Lys Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile 340 345 350 Asp Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys 355 360 365 Glu Val Ile Asp Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg 370 375 380 Lys Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly 385 390 395 400 Lys Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val 405 410 415 Lys Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr 420 425 430 Tyr Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser 435 440 445 Asp Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg 450 455 460 Leu Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly 465 470 475 480 Asp Phe His Glu Val Thr Gly Leu 485 <210> 5 <211> 20 <212> RNA <21三> Artificial sequence <220> <223> Target sequence of Example 2 <400> 5 ggccaaaccu cggcuuaccu 20 <210> 6 <211> 482 Note: There seems to be a small error in the translation of "<21三>", it should probably be "<213>". I've translated it as "Artificial sequence" based on the context. Please check if this is correct in your original text. <212> PRT <213> Enterobacteriophage CC31 <400> 6 Met Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu Asp [[ID=二十]] 50 55 60 Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr Gly 65 70 75 80 Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser Asp 85 90 95 Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg Cys 一百 [[ID=三十五]] Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile Pro 115 120 125 Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile Glu 130 135 140 It should be noted that there seems to be an incorrect "二十" in the translation of line 20 which should be "50". And "一百" in the translation of line 100 should be "100", and "三十五" in the translation of line 35 should be "35". These are likely just mistakes in the provided translation content and not in the original rules of translation. The correct translation should be as follows: <212> PRT <213> Enterobacteriophage CC31 <400> 6 Met Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu Asp 50 55 60 Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr Gly 65 70 75 80 Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser Asp 85 90 95 Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg Cys 100 105 110 Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile Pro 115 120 125 Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile Glu 130 135 140 Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala Asp Gly Ala 145 150 155 160 Arg Ala Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp Val Lys Ile 165 170 175 Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu Leu Lys Gln 180 185 190 Gln Leu Ile Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Gly Gly 195 200 205 Val Met Ile Asp Gly Glu Leu Val Tyr His Ala Ser Thr Leu Pro Ala 210 215 220 Gly Pro Leu Asp Asp Ile Phe Gly Asp Leu Pro Glu Leu Ser Lys Ala 225 230 235 240 Lys Glu Phe Lys Glu Glu Ser Arg Thr Met Ser Asn Gly Leu Ala Asn 245 250 255 Lys Ser Leu Lys Gly Thr Ile Ser Ala Lys Glu Ala Ala Gly Met Lys 260 265 270 Phe Gln Val Trp Asp Tyr Val Pro Leu Asp Val Val Tyr Ser Glu Gly 275 280 285 Lys Gln Ser Gly Phe Ala Tyr Asp Val Arg Phe Arg Ala Leu Glu Leu 290 295 300 Met Val Gln Gly Tyr Ser Gln Met Ile Leu Ile Glu Asn His Ile Val 305 310 315 320 His Asn Leu Asp Glu Ala Lys Val Ile Tyr Arg Lys Tyr Val Asp Glu 325 330 335 Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Gly Ala Phe Trp Glu Asn 340 345 350 Thr Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu Val Ile Thr Ile Asp 355 360 365 Leu Arg Ile Val Asp Ile Tyr Glu His Ser Lys Gln Pro Gly Lys Ala 370 375 380 Gly Gly Phe Tyr Leu Glu Ser Glu Cys Gly Leu Ile Lys Val Lys Ala 385 390 395 400 Gly Ser Gly Leu Lys Asp Lys Pro Gly Lys Asp Ala His Glu Leu Asp 405 410 415 Arg Thr Arg Ile Trp Glu Asn Lys Asn Asp Tyr Ile Gly Gly Val Leu 420 425 430 Glu Ser Glu Cys Asn Gly Trp Leu Ala Ala Glu Gly Arg Thr Asp Tyr 435 440 445 Val Lys Leu Phe Leu Pro Ile Ala Ile Lys Met Arg Arg Asp Lys Asp 450 455 460 Val Ala Asn Thr Phe Ala Asp Ile Trp Gly Asp Phe His Glu Val Thr 465 470 475 480 Gly Leu <210> 7 <211> 483 <212> PRT <213> Artificial sequence <220> <223> Wild-type Enterobacteriaceae phage CC31 DNA ligase protein sequence (when fused to CBD) <400> 7 Gly Ser Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr 1 5 10 15 Lys Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu 20 25 30 Lys Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile 35 40 45 Lys Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu 50 55 60 Asp Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr 65 70 75 80 Gly Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser 85 90 95 Asp Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg 100 105 110 Cys Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile 115 120 125 Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile 130 135 140 Glu Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala Asp Gly 145 150 155 160 Ala Arg Ala Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp Val Lys 165 170 175 Ile Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu Leu Lys 180 185 190 Gln Gln Leu Ile Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Gly 195 200 205 Gly Val Met Ile Asp Gly Glu Leu Val Tyr His Ala Ser Thr Leu Pro 210 215 220 Ala Gly Pro Leu Asp Asp Ile Phe Gly Asp Leu Pro Glu Leu Ser Lys 225 230 235 240 Ala Lys Glu Phe Lys Glu Glu Ser Arg Thr Met Ser Asn Gly Leu Ala 245 250 255 Asn Lys Ser Leu Lys Gly Thr Ile Ser Ala Lys Glu Ala Ala Gly Met 260 265 270 Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Asp Val Val Tyr Ser Glu 275 280 285 Gly Lys Gln Ser Gly Phe Ala Tyr Asp Val Arg Phe Arg Ala Leu Glu 290 295 300 Leu Met Val Gln Gly Tyr Ser Gln Met Ile Leu Ile Glu Asn His Ile 305 310 315 320 Val His Asn Leu Asp Glu Ala Lys Val Ile Tyr Arg Lys Tyr Val Asp 325 330 335 Glu Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Gly Ala Phe Trp Glu 340 345 350 Asn Thr Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu Val Ile Thr Ile 355 360 365 Asp Leu Arg Ile Val Asp Ile Tyr Glu His Ser Lys Gln Pro Gly Lys 370 375 380 Ala Gly Gly Phe Tyr Leu Glu Ser Glu Cys Gly Leu Ile Lys Val Lys 385 390 395 400 Ala Gly Ser Gly Leu Lys Asp Lys Pro Gly Lys Asp Ala His Glu Leu 405 410 415 Asp Arg Thr Arg Ile Trp Glu Asn Lys Asn Asp Tyr Ile Gly Gly Val 420 425 430 Leu Glu Ser Glu Cys Asn Gly Trp Leu Ala Ala Glu Gly Arg Thr Asp 435 440 445 Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Lys Met Arg Arg Asp Lys 450 455 460 Asp Val Ala Asn Thr Phe Ala Asp Ile Trp Gly Asp Phe His Glu Val 465 470 475 480 Thr Gly Leu <210> 8 <211> 497 <212> PRT <213> Shigella phage Shf125875 <400> 8 Met Ile Leu Asp Ile Leu Asn Gln Ile Ala Ala Ile Gly Ser Thr Lys 1 5 10 15 Thr Lys Gln Glu Ile Leu Lys Lys Asn Lys Asp Asn Lys Leu Leu Glu 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ala Arg Gly Ile Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Gly Pro Gly Glu Arg Ser Gln Ala Tyr Gly Leu Leu Glu 50 55 60 Leu Asp Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Lys Glu Leu Met Gly Tyr Ile Ala Asp 85 90 95 Gly Lys Pro Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Val Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Gln Leu Gln Pro Gln Met Leu Ala Ser Ala Tyr Asp Glu Lys 130 135 140 Leu Ile Thr Lys Asn Ile Lys Trp Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Asp Asp Gly Val Gln Phe 165 170 175 Phe Ser Arg Ala Gly Asn Glu Tyr His Gly Leu Thr Leu Leu Ala Asp 180 185 190 Glu Leu Met Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Asn Gly 195 200 205 Val Leu Ile Asp Gly Glu Leu Val Tyr His Ser Phe Asp Ile Lys Lys 210 215 220 Ala Val Ser Ser Gly Asn Asp Leu Ser Phe Leu Phe Gly Asp Asn Glu 225 230 235 240 Glu Ser Glu Glu Val Gln Val Ala Asp Arg Ser Thr Ser Asn Gly Leu 245 250 255 Ala Asn Lys Ser Leu Gln Gly Thr Ile Ser Pro Lys Glu Ala Glu Gly 260 265 270 Met Val Leu Gln Ala Trp Asp Tyr Val Pro Leu Asp Glu Val Tyr Ser 275 280 285 Asp Gly Lys Ile Lys Gly Gln Lys Tyr Asp Val Arg Phe Ala Ala Leu 290 295 300 Glu Asn Met Ala Glu Gly Phe Lys Arg Ile Glu Pro Ile Glu Asn Gln 305 310 315 320 Leu Val His Asn Leu Asp Glu Ala Lys Val Val Tyr Lys Lys Tyr Val 325 330 335 Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Arg Asp Ser Tyr Trp 340 345 350 Glu Asn Lys Arg Ser Lys Asn Leu Ile Lys Phe Lys Glu Val Ile Asp 355 360 365 Ile Ala Leu Glu Val Val Gly Tyr Tyr Glu His Ser Lys Asp Pro Asn 370 375 380 Lys Leu Gly Gly Val Glu Leu Val Ser Arg Cys Arg Arg Ile Thr Thr 385 390 395 400 Asp Cys Gly Ser Gly Phe Lys Asp Thr Thr His Lys Thr Val Asp Gly 405 410 415 Val Lys Val Leu Ile Pro Leu Asp Glu Arg His Asp Leu Asp Arg Glu 420 425 430 Arg Leu Met Ala Glu Ala Arg Glu Gly Lys Leu Ile Gly Arg Ile Ala 435 440 445 Asp Cys Glu Cys Asn Gly Trp Val His Ser Lys Gly Arg Glu Gly Thr 450 455 460 Val Gly Ile Phe Leu Pro Ile Ile Lys Gly Phe Arg Phe Asp Lys Thr 465 470 475 480 Glu Ala Asp Ser Phe Glu Asp Val Phe Gly Pro Trp Ser Gln Thr Gly 485 490 495 Leu <210> 9 <211> 498 <212> PRT <213> Artificial sequence <220> <223> Wild-type Shigella phage Shf125875 DNA ligase protein sequence (when fused to CBD) <400> 9 Gly Ser Ile Leu Asp Ile Leu Asn Gln Ile Ala Ala Ile Gly Ser Thr 1 5 10 15 Lys Thr Lys Gln Glu Ile Leu Lys Lys Asn Lys Asp Asn Lys Leu Leu 20 25 30 Glu Arg Val Tyr Arg Leu Thr Tyr Ala Arg Gly Ile Gln Tyr Tyr Ile 35 40 45 Lys Lys Trp Pro Gly Pro Gly Glu Arg Ser Gln Ala Tyr Gly Leu Leu 50 55 60 Glu Leu Asp Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg 65 70 75 80 Lys Leu Thr Gly Asn Ala Ala Ile Lys Glu Leu Met Gly Tyr Ile Ala 85 90 95 Asp Gly Lys Pro Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg 100 105 110 Asp Leu Glu Val Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro 115 120 125 Gly Leu Ile Gln Leu Gln Pro Gln Met Leu Ala Ser Ala Tyr Asp Glu 130 135 140 Lys Leu Ile Thr Lys Asn Ile Lys Trp Pro Ala Phe Ala Gln Leu Lys 145 150 155 160 Ala Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Asp Asp Gly Val Gln 165 170 175 Phe Phe Ser Arg Ala Gly Asn Glu Tyr His Gly Leu Thr Leu Leu Ala 180 185 190 Asp Glu Leu Met Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Asn 195 200 205 Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Ser Phe Asp Ile Lys 210 215 220 Lys Ala Val Ser Ser Gly Asn Asp Leu Ser Phe Leu Phe Gly Asp Asn 225 230 235 240 Glu Glu Ser Glu Glu Val Gln Val Ala Asp Arg Ser Thr Ser Asn Gly 245 250 255 Leu Ala Asn Lys Ser Leu Gln Gly Thr Ile Ser Pro Lys Glu Ala Glu 260 265 270 Gly Met Val Leu Gln Ala Trp Asp Tyr Val Pro Leu Asp Glu Val Tyr 275 280 285 Ser Asp Gly Lys Ile Lys Gly Gln Lys Tyr Asp Val Arg Phe Ala Ala 290 295 300 Leu Glu Asn Met Ala Glu Gly Phe Lys Arg Ile Glu Pro Ile Glu Asn 305 310 315 320 Gln Leu Val His Asn Leu Asp Glu Ala Lys Val Val Tyr Lys Lys Tyr 325 330 335 Val Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Arg Asp Ser Tyr 340 345 350 Trp Glu Asn Lys Arg Ser Lys Asn Leu Ile Lys Phe Lys Glu Val Ile 355 360 365 Asp Ile Ala Leu Glu Val Val Gly Tyr Tyr Glu His Ser Lys Asp Pro 370 375 380 Asn Lys Leu Gly Gly Val Glu Leu Val Ser Arg Cys Arg Arg Ile Thr 385 390 395 400 Thr Asp Cys Gly Ser Gly Phe Lys Asp Thr Thr His Lys Thr Val Asp 405 410 415 Gly Val Lys Val Leu Ile Pro Leu Asp Glu Arg His Asp Leu Asp Arg 420 425 430 Glu Arg Leu Met Ala Glu Ala Arg Glu Gly Lys Leu Ile Gly Arg Ile 435 440 445 Ala Asp Cys Glu Cys Asn Gly Trp Val His Ser Lys Gly Arg Glu Gly 450 455 460 Thr Val Gly Ile Phe Leu Pro Ile Ile Lys Gly Phe Arg Phe Asp Lys 465 470 475 480 Thr Glu Ala Asp Ser Phe Glu Asp Val Phe Gly Pro Trp Ser Gln Thr 485 490 495 Gly Leu <210> 10 <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 10 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Arg 355 360 365 Val Ile Asp Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 11 <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 11 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Gly 355 360 365 Val Ile Asp Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 12 <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 12 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Lys 355 360 365 Val Ile Asp Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 45​​​​​​​​​​​​​​​​​​​​​​​​ 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Val Ile Leu Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 14 <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 14 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Val Ile Gln Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 15 <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 15 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Val Ile Gln Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 16 <211> 487 <212> PRT <213> Artificial Sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 16 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Val Ile Val Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 17 <211> 487 <212> PRT <213> Synthetic Sequence <220> <223> Mutant Ligase (T4 backbone) Protein Sequence <400> 17 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Val Ile Arg Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 18 <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 18 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Ala Ile Asp Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 19 <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 19 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Light Light Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Glu 355 360 365 Lys Ile Asp Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 20 <211> 487 <212> PRT <213> Artificial Sequence <220> <223> Mutant Ligase (T4 Backbone) Protein Sequence <400> 20 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Arg 355 360 365 Val Ile Val Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> 21 <211> 487 <212> PRT <213> Artificial Sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> 21 Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Lys 355 360 365 Val Ile Glu Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> twenty two <211> 487 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (T4 backbone) protein sequence <400> twenty two Met Ile Leu Lys Ile Leu Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Gln Lys Gln Ala Ile Leu Glu Lys Asn Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ser Arg Gly Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Lys Pro Gly Ile Ala Thr Gln Ser Phe Gly Met Leu Thr 50 55 60 Leu Thr Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Glu Glu Leu Thr Gly Tyr Ile Thr Asp 85 90 95 Gly Lys Lys Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Cys Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Pro Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys 130 135 140 Gly Ile Asn Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp 165 170 175 Val Arg Leu Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu 180 185 190 Leu Lys Glu Glu Leu Ile Lys Met Thr Ala Glu Ala Arg Gln Ile His 195 200 205 Pro Glu Gly Val Leu Ile Asp Gly Glu Leu Val Tyr His Glu Gln Val 210 215 220 Lys Lys Glu Pro Glu Gly Leu Asp Phe Leu Phe Asp Ala Tyr Pro Glu 225 230 235 240 Asn Ser Lys Ala Lys Glu Phe Ala Glu Val Ala Glu Ser Arg Thr Ala 245 250 255 Ser Asn Gly Ile Ala Asn Lys Ser Leu Lys Gly Thr Ile Ser Glu Lys 260 265 270 Glu Ala Gln Cys Met Lys Phe Gln Val Trp Asp Tyr Val Pro Leu Val 275 280 285 Glu Ile Tyr Ser Leu Pro Ala Phe Arg Leu Lys Tyr Asp Val Arg Phe 290 295 300 Ser Lys Leu Glu Gln Met Thr Ser Gly Tyr Asp Lys Val Ile Leu Ile 305 310 315 320 Glu Asn Gln Val Val Asn Asn Leu Asp Glu Ala Lys Val Ile Tyr Lys 325 330 335 Lys Tyr Ile Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Asp 340 345 350 Gly Leu Trp Glu Asn Ala Arg Ser Lys Asn Leu Tyr Lys Phe Lys Lys 355 360 365 Val Ile His Val Asp Leu Lys Ile Val Gly Ile Tyr Pro His Arg Lys 370 375 380 Asp Pro Thr Lys Ala Gly Gly Phe Ile Leu Glu Ser Glu Cys Gly Lys 385 390 395 400 Ile Lys Val Asn Ala Gly Ser Gly Leu Lys Asp Lys Ala Gly Val Lys 405 410 415 Ser His Glu Leu Asp Arg Thr Arg Ile Met Glu Asn Gln Asn Tyr Tyr 420 425 430 Ile Gly Lys Ile Leu Glu Cys Glu Cys Asn Gly Trp Leu Lys Ser Asp 435 440 445 Gly Arg Thr Asp Tyr Val Lys Leu Phe Leu Pro Ile Ala Ile Arg Leu 450 455 460 Arg Glu Asp Lys Thr Lys Ala Asn Thr Phe Glu Asp Val Phe Gly Asp 465 470 475 480 Phe His Glu Val Thr Gly Leu 485 <210> twenty three <211> 482 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (Enterobacterial phage CC31 backbone-clone A4) protein sequence <400> twenty three Met Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu Asp 50 55 60 Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr Gly 65 70 75 80 Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser Asp 85 90 95 Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg Cys 100 105 110 Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile Pro 115 120 125 Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile Glu 130 135 140 Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala Asp Gly Ala 145 150 155 160 Arg Ala Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp Val Lys Ile 165 170 175 Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu Leu Lys Gln 180 185 190 Gln Leu Ile Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Gly Gly 195 200 205 Val Met Ile Asp Gly Glu Leu Val Tyr His Ala Ser Thr Leu Pro Ala 210 215 220 Gly Pro Leu Asp Asp Ile Phe Gly Asp Leu Pro Glu Leu Ser Lys Ala 225 230 235 240 Lys Glu Phe Lys Glu Glu Ser Arg Thr Met Ser Asn Gly Leu Ala Asn 245 250 255 Lys Ser Leu Lys Gly Thr Ile Ser Ala Lys Glu Ala Ala Gly Met Lys 260 265 270 Phe Gln Val Trp Asp Tyr Val Pro Leu Asp Val Val Tyr Ser Glu Gly 275 280 285 Lys Gln Ser Gly Phe Ala Tyr Asp Val Arg Phe Arg Ala Leu Glu Leu 290 295 300 Met Val Gln Gly Tyr Ser Gln Met Ile Leu Ile Glu Asn His Ile Val 305 310 315 320 His Asn Leu Asp Glu Ala Lys Val Ile Tyr Arg Lys Tyr Val Asp Glu 325 330 335 Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Gly Ala Phe Trp Glu Asn 340 345 350 Thr Arg Ser Lys Asn Leu Tyr Lys Phe Lys Arg Val Ile Val Ile Asp 355 360 365 Leu Arg Ile Val Asp Ile Tyr Glu His Ser Lys Gln Pro Gly Lys Ala 370 375 380 Gly Gly Phe Tyr Leu Glu Ser Glu Cys Gly Leu Ile Lys Val Lys Ala 385 390 395 400 Gly Ser Gly Leu Lys Asp Lys Pro Gly Lys Asp Ala His Glu Leu Asp 405 410 415 Arg Thr Arg Ile Trp Glu Asn Lys Asn Asp Tyr Ile Gly Gly Val Leu 420 425 430 Glu Ser Glu Cys Asn Gly Trp Leu Ala Ala Glu Gly Arg Thr Asp Tyr 435 440 445 Val Lys Leu Phe Leu Pro Ile Ala Ile Lys Met Arg Arg Asp Lys Asp 450 455 460 Val Ala Asn Thr Phe Ala Asp Ile Trp Gly Asp Phe His Glu Val Thr 465 470 475 480 Gly Leu <210> twenty four <211> 482 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence <400> twenty four Met Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu Asp 50 55 60 Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr Gly 65 70 75 80 Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser Asp 85 90 95 Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg Cys 100 105 110 Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile Pro 115 120 125 Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile Glu 130 135 140 Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala Asp Gly Ala 145 150 155 160 Arg Ala Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp Val Lys Ile 165 170 175 Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu Leu Lys Gln 180 185 190 Gln Leu Ile Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Gly Gly 195 200 205 Val Met Ile Asp Gly Glu Leu Val Tyr His Ala Ser Thr Leu Pro Ala 210 215 220 Gly Pro Leu Asp Asp Ile Phe Gly Asp Leu Pro Glu Leu Ser Lys Ala 225 230 235 240 Lys Glu Phe Lys Glu Glu Ser Arg Thr Met Ser Asn Gly Leu Ala Asn 245 250 255 Lys Ser Leu Lys Gly Thr Ile Ser Ala Lys Glu Ala Ala Gly Met Lys 260 265 270 Phe Gln Val Trp Asp Tyr Val Pro Leu Asp Val Val Tyr Ser Glu Gly 275 280 285 Lys Gln Ser Gly Phe Ala Tyr Asp Val Arg Phe Arg Ala Leu Glu Leu 290 295 300 Met Val Gln Gly Tyr Ser Gln Met Ile Leu Ile Glu Asn His Ile Val 305 310 315 320 His Asn Leu Asp Glu Ala Lys Val Ile Tyr Arg Lys Tyr Val Asp Glu 325 330 335 Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Gly Ala Phe Trp Glu Asn 340 345 350 Thr Arg Ser Lys Asn Leu Tyr Lys Phe Lys Lys Val Ile Lys Ile Asp 355 360 365 Leu Arg Ile Val Asp Ile Tyr Glu His Ser Lys Gln Pro Gly Lys Ala 370 375 380 Gly Gly Phe Tyr Leu Glu Ser Glu Cys Gly Leu Ile Lys Val Lys Ala 385 390 395 400 Gly Ser Gly Leu Lys Asp Lys Pro Gly Lys Asp Ala His Glu Leu Asp 405 410 415 Arg Thr Arg Ile Trp Glu Asn Lys Asn Asp Tyr Ile Gly Gly Val Leu 420 425 430 Glu Ser Glu Cys Asn Gly Trp Leu Ala Ala Glu Gly Arg Thr Asp Tyr 435 440 445 Val Lys Leu Phe Leu Pro Ile Ala Ile Lys Met Arg Arg Asp Lys Asp 450 455 460 Val Ala Asn Thr Phe Ala Asp Ile Trp Gly Asp Phe His Glu Val Thr 465 470 475 480 Gly Leu <210> 25 <211> 482 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence <400> 25 Met Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu Asp 50 55 60 Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr Gly 65 70 75 80 Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser Asp 85 90 95 Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg Cys 100 105 110 Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile Pro 115 120 125 Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile Glu 130 135 140 Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala Asp Gly Ala 145 150 155 160 Arg Ala Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp Val Lys Ile 165 170 175 Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu Leu Lys Gln 180 185 190 Gln Leu Ile Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Gly Gly 195 200 205 Val Met Ile Asp Gly Glu Leu Val Tyr His Ala Ser Thr Leu Pro Ala 210 215 220 Gly Pro Leu Asp Asp Ile Phe Gly Asp Leu Pro Glu Leu Ser Lys Ala 225 230 235 240 Lys Glu Phe Lys Glu Glu Ser Arg Thr Met Ser Asn Gly Leu Ala Asn 245 250 255 Lys Ser Leu Lys Gly Thr Ile Ser Ala Lys Glu Ala Ala Gly Met Lys 260 265 270 Phe Gln Val Trp Asp Tyr Val Pro Leu Asp Val Val Tyr Ser Glu Gly 275 280 285 Lys Gln Ser Gly Phe Ala Tyr Asp Val Arg Phe Arg Ala Leu Glu Leu 290 295 300 Met Val Gln Gly Tyr Ser Gln Met Ile Leu Ile Glu Asn His Ile Val 305 310 315 320 His Asn Leu Asp Glu Ala Lys Val Ile Tyr Arg Lys Tyr Val Asp Glu 325 330 335 Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Gly Ala Phe Trp Glu Asn 340 345 350 Thr Arg Ser Lys Asn Leu Tyr Lys Phe Lys Gly Val Ile Phe Ile Asp 355 360 365 Leu Arg Ile Val Asp Ile Tyr Glu His Ser Lys Gln Pro Gly Lys Ala 370 375 380 Gly Gly Phe Tyr Leu Glu Ser Glu Cys Gly Leu Ile Lys Val Lys Ala 385 390 395 400 Gly Ser Gly Leu Lys Asp Lys Pro Gly Lys Asp Ala His Glu Leu Asp 405 410 415 Arg Thr Arg Ile Trp Glu Asn Lys Asn Asp Tyr Ile Gly Gly Val Leu 420 425 430 Glu Ser Glu Cys Asn Gly Trp Leu Ala Ala Glu Gly Arg Thr Asp Tyr 435 440 445 Val Lys Leu Phe Leu Pro Ile Ala Ile Lys Met Arg Arg Asp Lys Asp 450 455 460 Val Ala Asn Thr Phe Ala Asp Ile Trp Gly Asp Phe His Glu Val Thr 465 470 475 480 Gly Leu <210> 26 <211> 482 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence <400> 26 Met Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu Asp 50 55 60 Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr Gly 65 70 75 80 Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser Asp 85 90 95 Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg Cys 100 105 110 Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile Pro 115 120 125 Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile Glu 130 135 140 Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala Asp Gly Ala 145 150 155 160 Arg Ala Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp Val Lys Ile 165 170 175 Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu Leu Lys Gln 180 185 190 Gln Leu Ile Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Gly Gly 195 200 205 Val Met Ile Asp Gly Glu Leu Val Tyr His Ala Ser Thr Leu Pro Ala 210 215 220 Gly Pro Leu Asp Asp Ile Phe Gly Asp Leu Pro Glu Leu Ser Lys Ala 225 230 235 240 Lys Glu Phe Lys Glu Glu Ser Arg Thr Met Ser Asn Gly Leu Ala Asn 245 250 255 Lys Ser Leu Lys Gly Thr Ile Ser Ala Lys Glu Ala Ala Gly Met Lys 260 265 270 Phe Gln Val Trp Asp Tyr Val Pro Leu Asp Val Val Tyr Ser Glu Gly 275 280 285 Lys Gln Ser Gly Phe Ala Tyr Asp Val Arg Phe Arg Ala Leu Glu Leu 290 295 300 Met Val Gln Gly Tyr Ser Gln Met Ile Leu Ile Glu Asn His Ile Val 305 310 315 320 His Asn Leu Asp Glu Ala Lys Val Ile Tyr Arg Lys Tyr Val Asp Glu 325 330 335 Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Gly Ala Phe Trp Glu Asn 340 345 350 Thr Arg Ser Lys Asn Leu Tyr Lys Phe Lys Gly Val Ile Leu Ile Asp 355 360 365 Leu Arg Ile Val Asp Ile Tyr Glu His Ser Lys Gln Pro Gly Lys Ala 370 375 380 Gly Gly Phe Tyr Leu Glu Ser Glu Cys Gly Leu Ile Lys Val Lys Ala 385 390 395 400 Gly Ser Gly Leu Lys Asp Lys Pro Gly Lys Asp Ala His Glu Leu Asp 405 410 415 Arg Thr Arg Ile Trp Glu Asn Lys Asn Asp Tyr Ile Gly Gly Val Leu 420 425 430 Glu Ser Glu Cys Asn Gly Trp Leu Ala Ala Glu Gly Arg Thr Asp Tyr 435 440 445 Val Lys Leu Phe Leu Pro Ile Ala Ile Lys Met Arg Arg Asp Lys Asp 450 455 460 Val Ala Asn Thr Phe Ala Asp Ile Trp Gly Asp Phe His Glu Val Thr 465 470 475 480 Gly Leu <210> 27 <211> 482 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (Enterobacterial phage CC31 backbone) protein sequence <400> 27 Met Ile Leu Asp Ile Ile Asn Glu Ile Ala Ser Ile Gly Ser Thr Lys 1 5 10 15 Glu Lys Glu Ala Ile Ile Arg Arg His Lys Asp Asn Glu Leu Leu Lys 20 25 30 Arg Val Phe Arg Met Thr Tyr Asp Gly Lys Leu Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Asp Thr Arg Pro Lys Gly Asp Ile His Leu Thr Leu Glu Asp 50 55 60 Met Leu Tyr Leu Leu Glu Glu Lys Leu Ala Lys Arg Val Val Thr Gly 65 70 75 80 Asn Ala Ala Lys Glu Lys Leu Glu Ile Ala Leu Ser Gln Thr Ser Asp 85 90 95 Ala Asp Ala Glu Val Val Lys Lys Val Leu Leu Arg Asp Leu Arg Cys 100 105 110 Gly Ala Ser Arg Ser Ile Ala Asn Lys Val Trp Lys Asn Leu Ile Pro 115 120 125 Glu Gln Pro Gln Met Leu Ala Ser Ser Tyr Asp Glu Lys Gly Ile Glu 130 135 140 Lys Asn Ile Lys Phe Pro Ala Phe Ala Gln Leu Lys Ala Asp Gly Ala 145 150 155 160 Arg Ala Phe Ala Glu Val Arg Gly Asp Glu Leu Asp Asp Val Lys Ile 165 170 175 Leu Ser Arg Ala Gly Asn Glu Tyr Leu Gly Leu Asp Leu Leu Lys Gln 180 185 190 Gln Leu Ile Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Gly Gly 195 200 205 Val Met Ile Asp Gly Glu Leu Val Tyr His Ala Ser Thr Leu Pro Ala 210 215 220 Gly Pro Leu Asp Asp Ile Phe Gly Asp Leu Pro Glu Leu Ser Lys Ala 225 230 235 240 Lys Glu Phe Lys Glu Glu Ser Arg Thr Met Ser Asn Gly Leu Ala Asn 245 250 255 Lys Ser Leu Lys Gly Thr Ile Ser Ala Lys Glu Ala Ala Gly Met Lys 260 265 270 Phe Gln Val Trp Asp Tyr Val Pro Leu Asp Val Val Tyr Ser Glu Gly 275 280 285 Lys Gln Ser Gly Phe Ala Tyr Asp Val Arg Phe Arg Ala Leu Glu Leu 290 295 300 Met Val Gln Gly Tyr Ser Gln Met Ile Leu Ile Glu Asn His Ile Val 305 310 315 320 His Asn Leu Asp Glu Ala Lys Val Ile Tyr Arg Lys Tyr Val Asp Glu 325 330 335 Gly Leu Glu Gly Ile Ile Leu Lys Asn Ile Gly Ala Phe Trp Glu Asn 340 345 350 Thr Arg Ser Lys Asn Leu Tyr Lys Phe Lys Arg Val Ile Phe Ile Asp 355 360 365 Leu Arg Ile Val Asp Ile Tyr Glu His Ser Lys Gln Pro Gly Lys Ala 370 375 380 Gly Gly Phe Tyr Leu Glu Ser Glu Cys Gly Leu Ile Lys Val Lys Ala 385 390 395 400 Gly Ser Gly Leu Lys Asp Lys Pro Gly Lys Asp Ala His Glu Leu Asp 405 410 415 Arg Thr Arg Ile Trp Glu Asn Lys Asn Asp Tyr Ile Gly Gly Val Leu 420 425 430 Glu Ser Glu Cys Asn Gly Trp Leu Ala Ala Glu Gly Arg Thr Asp Tyr 435 440 445 Val Lys Leu Phe Leu Pro Ile Ala Ile Lys Met Arg Arg Asp Lys Asp 450 455 460 Val Ala Asn Thr Phe Ala Asp Ile Trp Gly Asp Phe His Glu Val Thr 465 470 475 480 Gly Leu <210> 28 <211> 497 <212> PRT <213> Artificial sequence <220> <223> Mutant ligase (Shigella phage Shf125875 backbone) protein sequence <400> 28 Met Ile Leu Asp Ile Leu Asn Gln Ile Ala Ala Ile Gly Ser Thr Lys 1 5 10 15 Thr Lys Gln Glu Ile Leu Lys Lys Asn Lys Asp Asn Lys Leu Leu Glu 20 25 30 Arg Val Tyr Arg Leu Thr Tyr Ala Arg Gly Ile Gln Tyr Tyr Ile Lys 35 40 45 Lys Trp Pro Gly Pro Gly Glu Arg Ser Gln Ala Tyr Gly Leu Leu Glu 50 55 60 Leu Asp Asp Met Leu Asp Phe Ile Glu Phe Thr Leu Ala Thr Arg Lys 65 70 75 80 Leu Thr Gly Asn Ala Ala Ile Lys Glu Leu Met Gly Tyr Ile Ala Asp 85 90 95 Gly Lys Pro Asp Asp Val Glu Val Leu Arg Arg Val Met Met Arg Asp 100 105 110 Leu Glu Val Gly Ala Ser Val Ser Ile Ala Asn Lys Val Trp Pro Gly 115 120 125 Leu Ile Gln Leu Gln Pro Gln Met Leu Ala Ser Ala Tyr Asp Glu Lys 130 135 140 Leu Ile Thr Lys Asn Ile Lys Trp Pro Ala Phe Ala Gln Leu Lys Ala 145 150 155 160 Asp Gly Ala Arg Cys Phe Ala Glu Val Arg Asp Asp Gly Val Gln Phe 165 170 175 Phe Ser Arg Ala Gly Asn Glu Tyr His Gly Leu Thr Leu Leu Ala Asp 180 185 190 Glu Leu Met Glu Met Thr Lys Glu Ala Arg Glu Arg His Pro Asn Gly 195 200 205 Val Leu Ile Asp Gly Glu Leu Val Tyr His Ser Phe Asp Ile Lys Lys 210 215 220 Ala Val Ser Ser Gly Asn Asp Leu Ser Phe Leu Phe Gly Asp Asn Glu 225 230 235 240 Glu Ser Glu Glu Val Gln Val Ala Asp Arg Ser Thr Ser Asn Gly Leu 245 250 255 Ala Asn Lys Ser Leu Gln Gly Thr Ile Ser Pro Lys Glu Ala Glu Gly 260 265 270 Met Val Leu Gln Ala Trp Asp Tyr Val Pro Leu Asp Glu Val Tyr Ser 275 280 285 Asp Gly Lys Ile Lys Gly Gln Lys Tyr Asp Val Arg Phe Ala Ala Leu 290 295 300 Glu Asn Met Ala Glu Gly Phe Lys Arg Ile Glu Pro Ile Glu Asn Gln 305 310 315 320 Leu Val His Asn Leu Asp Glu Ala Lys Val Val Tyr Lys Lys Tyr Val 325 330 335 Asp Gln Gly Leu Glu Gly Ile Ile Leu Lys Asn Arg Asp Ser Tyr Trp 340 345 350 Glu Asn Lys Arg Ser Lys Asn Leu Ile Lys Phe Lys Arg Val Ile Val 355 360 365 Ile Ala Leu Glu Val Val Gly Tyr Tyr Glu His Ser Lys Asp Pro Asn 370 375 380 Lys Leu Gly Gly Val Glu Leu Val Ser Arg Cys Arg Arg Ile Thr Thr 385 390 395 400 Asp Cys Gly Ser Gly Phe Lys Asp Thr Thr His Lys Thr Val Asp Gly 405 410 415 Val Lys Val Leu Ile Pro Leu Asp Glu Arg His Asp Leu Asp Arg Glu 420 425 430 Arg Leu Met Ala Glu Ala Arg Glu Gly Lys Leu Ile Gly Arg Ile Ala 435 440 445 Asp Cys Glu Cys Asn Gly Trp Val His Ser Lys Gly Arg Glu Gly Thr 450 455 460 Val Gly Ile Phe Leu Pro Ile Ile Lys Gly Phe Arg Phe Asp Lys Thr 465 470 475 480 Glu Ala Asp Ser Phe Glu Asp Val Phe Gly Pro Trp Ser Gln Thr Gly 485 490 495 Leu <210> 29 <211> 298 <212> PRT <213> Paramecium bursaria chlorella virus PBCV-1 <400> 29 Met Ala Ile Thr Lys Pro Leu Leu Ala Ala Thr Leu Glu Asn Ile Glu 1 5 10 15 Asp Val Gln Phe Pro Cys Leu Ala Thr Pro Lys Ile Asp Gly Ile Arg 20 25 30 Ser Val Lys Gln Thr Gln Met Leu Ser Arg Thr Phe Lys Pro Ile Arg 35 40 45 Asn Ser Val Met Asn Arg Leu Leu Thr Glu Leu Leu Pro Glu Gly Ser 50 55 60 Asp Gly Glu Ile Ser Ile Glu Gly Ala Thr Phe Gln Asp Thr Thr Ser 65 70 75 80 Ala Val Met Thr Gly His Lys Met Tyr Asn Ala Lys Phe Ser Tyr Tyr 85 90 95 Trp Phe Asp Tyr Val Thr Asp Asp Pro Leu Lys Lys Tyr Ile Asp Arg 100 105 110 Val Glu Asp Met Lys Asn Tyr Ile Thr Val His Pro His Ile Leu Glu 115 120 125 His Ala Gln Val Lys Ile Ile Pro Leu Ile Pro Val Glu Ile Asn Asn 130 135 140 Ile Thr Glu Leu Leu Gln Tyr Glu Arg Asp Val Leu Ser Lys Gly Phe 145 150 155 160 Glu Gly Val Met Ile Arg Lys Pro Asp Gly Lys Tyr Lys Phe Gly Arg 165 170 175 Ser Thr Leu Lys Glu Gly Ile Leu Leu Lys Met Lys Gln Phe Lys Asp 180 185 190 Ala Glu Ala Thr Ile Ile Ser Met Thr Ala Leu Phe Lys Asn Thr Asn 195 200 205 Thr Lys Thr Lys Asp Asn Phe Gly Tyr Ser Lys Arg Ser Thr His Lys 210 215 220 Ser Gly Lys Val Glu Glu Asp Val Met Gly Ser Ile Glu Val Asp Tyr 225 230 235 240 Asp Gly Val Val Phe Ser Ile Gly Thr Gly Phe Asp Ala Asp Gln Arg 245 250 255 Arg Asp Phe Trp Gln Asn Lys Glu Ser Tyr Ile Gly Lys Met Val Lys 260 265 270 Phe Lys Tyr Phe Glu Met Gly Ser Lys Asp Cys Pro Arg Phe Pro Val 275 280 285 Phe Ile Gly Ile Arg His Glu Glu Asp Arg 290 295 <210> 30 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Examples 5, 8 and 9 template oligonucleotide sequences <400> 30 tttggtgcga agcagactga ggc 23 <210> 31 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 31 tttggtgcga agcagagtga ggc 23 <210> 32 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 32 tttggtgcga agcagattga ggc 23 <210> 33 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 33 tttggtgcga agcagaatga ggc 23 <210> 34 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 34 tttggtgcga agcagtctga ggc 23 <210> 35 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 35 tttggtgcga agcagtgtga ggc 23 <210> 36 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 36 tttggtgcga agcagtttga ggc 23 <210> 37 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 37 tttggtgcga agcagtatga ggc 23 <210> 38 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 38 tttggtgcga agcagcctga ggc 23 <210> 39 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 39 tttggtgcga agcagcgtga ggc 23 <210> 40 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 40 tttggtgcga agcagcttga ggc 23 <210> 41 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 41 tttggtgcga agcagcatga ggc 23 <210> 42 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 42 tttggtgcga agcaggctga ggc 23 <210> 43 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 43 tttggtgcga agcagggtga ggc 23 <210> 44 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 44 tttggtgcga agcaggttga ggc 23 <210> 45 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Example 5 Template Oligonucleotide Sequence <400> 45 tttggtgcga agcaggatga ggc 23 <210> 46 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Example 14 "20-mer" oligonucleotide sequence <400> 46 gccucagtct gcttcgcacc 20 <210> 47 <211> 36 <212> DNA <213> Artificial sequence <220> <223> Example 7 Template Oligonucleotide Sequence <400> 47 tttggtgcga agcagaaggt aagccgaggt ttggcc 36 <210> 48 <211> 298 <212> PRT <213> Paramecium chlororaphis virus NE-JV-4 <400> 48 Met Ala Ile Thr Lys Pro Leu Leu Ala Ala Thr Leu Glu Asn Ile Glu 1 5 10 15 Asp Val Gln Phe Pro Cys Leu Ala Thr Pro Lys Ile Asp Gly Ile Arg 20 25 30 Ser Val Lys Gln Thr Gln Met Leu Ser Arg Thr Phe Lys Pro Ile Arg 35 40 45 Asn Ser Val Met Asn Arg Leu Leu Thr Glu Leu Leu Pro Glu Gly Ser 50 55 60 Asp Gly Glu Ile Ser Ile Glu Gly Ala Thr Phe Gln Asp Thr Thr Ser 65 70 75 80 Ala Val Met Thr Gly His Lys Met Tyr Asn Ala Lys Phe Ser Tyr Tyr 85 90 95 Trp Phe Asp Tyr Val Thr Asp Asp Pro Leu Lys Lys Tyr Ser Asp Arg 100 105 110 Val Glu Asp Met Lys Asn Tyr Ile Thr Ala His Pro His Ile Leu Asp 115 120 125 His Glu Gln Val Lys Ile Ile Pro Leu Ile Pro Val Glu Ile Asn Asn 130 135 140 Ile Thr Glu Leu Leu Gln Tyr Glu Arg Asp Val Leu Ser Lys Gly Phe 145 150 155 160 Glu Gly Val Met Ile Arg Lys Pro Asp Gly Lys Tyr Lys Phe Gly Arg 165 170 175 Ser Thr Leu Lys Glu Gly Ile Leu Leu Lys Met Lys Gln Phe Lys Asp 180 185 190 Ala Glu Ala Thr Ile Ile Ser Met Thr Ala Leu Phe Lys Asn Thr Asn 195 200 205 Thr Lys Thr Lys Asp Asn Phe Gly Tyr Ser Lys Arg Ser Thr His Lys 210 215 220 Asn Gly Lys Val Glu Glu Asp Val Met Gly Ser Ile Glu Val Asp Tyr 225 230 235 240 Asp Gly Val Val Phe Ser Ile Gly Thr Gly Phe Asp Ala Asp Gln Arg 245 250 255 Arg Asp Phe Trp Gln Asn Lys Glu Ser Tyr Ile Gly Lys Met Val Lys 260 265 270 Phe Lys Tyr Phe Glu Met Gly Ser Lys Asp Cys Pro Arg Phe Pro Val 275 280 285 Phe Ile Gly Ile Arg His Glu Glu Asp His 290 295 <210> 49 <211> 298 <212> PRT <213> Paramecium bursaria Chlorella virus NYs1 <400> 49 Met Thr Ile Ala Lys Pro Leu Leu Ala Ala Thr Leu Glu Asn Leu Asp 1 5 10 15 Asp Val Lys Phe Pro Cys Leu Val Thr Pro Lys Ile Asp Gly Ile Arg 20 25 30 Ser Leu Lys Gln Gln His Met Leu Ser Arg Thr Phe Lys Pro Ile Arg 35 40 45 Asn Ser Val Met Asn Lys Leu Leu Ser Glu Leu Leu Pro Glu Gly Ala 50 55 60 Asp Gly Glu Ile Cys Ile Glu Asp Ser Thr Phe Gln Ala Thr Thr Ser 65 70 75 80 Ala Val Met Thr Gly His Lys Val Tyr Asp Glu Lys Phe Ser Tyr Tyr 85 90 95 Trp Phe Asp Tyr Val Val Asp Asp Pro Leu Lys Ser Tyr Thr Asp Arg 100 105 110 Val Asn Asp Met Lys Lys Tyr Val Asp Asp His Pro His Ile Leu Glu 115 120 125 His Glu Gln Val Lys Ile Ile Pro Leu Ile Pro Val Glu Ile Asn Asn 130 135 140 Ile Asp Glu Leu Ser Gln Tyr Glu Arg Asp Val Leu Ala Lys Gly Phe 145 150 155 160 Glu Gly Val Met Ile Arg Arg Pro Asp Gly Lys Tyr Lys Phe Gly Arg 165 170 175 Ser Thr Leu Lys Glu Gly Ile Leu Leu Lys Met Lys Gln Phe Lys Asp 180 185 190 Ala Glu Ala Thr Ile Ile Ser Met Ser Pro Arg Leu Lys Asn Thr Asn 195 200 205 Ala Lys Ser Lys Asp Asn Leu Gly Tyr Ser Lys Arg Ser Thr His Lys 210 215 220 Ser Gly Lys Val Glu Glu Glu Thr Met Gly Ser Ile Glu Val Asp Tyr 225 230 235 240 Asp Gly Val Val Phe Ser Ile Gly Thr Gly Phe Asp Asp Glu Gln Arg 245 250 255 Lys His Phe Trp Glu Asn Lys Asp Ser Tyr Ile Gly Lys Leu Leu Lys 260 265 270 Phe Lys Tyr Phe Glu Met Gly Ser Lys Asp Ala Pro Arg Phe Pro Val 275 280 285 Phe Ile Gly Ile Arg His Glu Glu Asp Cys 290 295 <210> 50 <211> 301 <212> PRT <213> Paramecium bursaria Chlorella virus NE-JV-1 <400> 50 Met Thr Ala Ile Gln Lys Pro Leu Leu Ala Ala Ser Phe Lys Lys Leu 1 5 10 15 Thr Val Ala Asp Val Lys Tyr Pro Val Phe Ala Thr Pro Lys Leu Asp 20 25 30 Gly Ile Arg Ala Leu Lys Ile Asp Gly Ala Phe Val Ser Arg Thr Phe 35 40 45 Lys Pro Ile Arg Asn Arg Ala Ile Ala Asp Ala Leu Gln Asp Leu Leu 50 55 60 Pro Asn Gly Ser Asp Gly Glu Ile Leu Ser Gly Ser Thr Phe Gln Asp 65 70 75 80 Ala Ser Ser Ala Val Met Thr Ala Lys Ala Gly Ile Gly Ala Asn Thr 85 90 95 Ile Phe Tyr Trp Phe Asp Tyr Val Lys Asp Asp Pro Asn Lys Pro Tyr 100 105 110 Leu Asp Arg Met Thr Asp Met Glu Asn Tyr Leu Lys Glu Arg Pro Glu 115 120 125 Ile Leu Asn Asp Asp Arg Ile Lys Ile Val Pro Leu Ile Pro Lys Lys 130 135 140 Ile Glu Thr Lys Asp Glu Leu Asp Thr Phe Glu Lys Ile Cys Leu Asp 145 150 155 160 Gln Gly Phe Glu Gly Val Met Ile Arg Ser Gly Ala Gly Lys Tyr Lys 165 170 175 Phe Gly Arg Ser Thr Glu Lys Glu Gly Ile Leu Ile Lys Ile Lys Gln 180 185 190 Phe Glu Asp Asp Glu Ala Val Val Ile Gly Phe Thr Pro Met Gln Thr 195 200 205 Asn Thr Asn Asp Lys Ser Met Asn Glu Leu Gly Asp Met Lys Arg Ser 210 215 220 Ser His Lys Asp Gly Lys Val Asn Leu Asp Thr Leu Gly Ala Leu Glu 225 230 235 240 Val Asp Trp Asn Gly Ile Thr Phe Ser Ile Gly Thr Gly Phe Asp His 245 250 255 Ala Leu Arg Asp Lys Leu Trp Ser Glu Arg Asp Lys Leu Ile Gly Lys 260 265 270 Ile Val Lys Phe Lys Tyr Phe Ala Gln Gly Val Lys Thr Ala Pro Arg 275 280 285 Phe Pro Val Phe Ile Gly Phe Arg Asp Pro Asp Asp Met 290 295 300 <210> 51 <211> 300 <212> PRT <213> Acanthocystis turfacea Chlorella virus Canal-1 <400> 51 Met Ala Ile Gln Lys Pro Leu Leu Ala Ala Ser Leu Lys Lys Met Ser 1 5 10 15 Val Gly Asp Leu Thr Phe Pro Val Phe Ala Thr Pro Lys Leu Asp Gly 20 25 30 Ile Arg Ala Leu Lys Val Gly Gly Thr Ile Val Ser Arg Thr Phe Lys 35 40 45 Pro Val Arg Asn Ser Ala Ile Ser Glu Val Leu Ala Ser Ile Leu Pro 50 55 60 Asp Gly Ser Asp Gly Glu Ile Leu Ser Gly Lys Thr Phe Gln Glu Ser 65 70 75 80 Thr Ser Thr Val Met Thr Ala Asp Ala Gly Leu Gly Ser Gly Thr Met 85 90 95 Phe Phe Trp Phe Asp Tyr Val Lys Asp Asp Pro Asn Lys Gly Tyr Leu 100 105 110 Asp Arg Ile Ala Asp Met Lys Ser Phe Thr Asp Arg His Pro Glu Ile 115 120 125 Leu Lys Asp Lys Arg Val Thr Ile Val Pro Leu Phe Pro Lys Lys Ile 130 135 140 Asp Thr Thr Glu Glu Leu His Glu Phe Glu Lys Trp Cys Leu Asp Gln 145 150 155 160 Gly Phe Glu Gly Val Met Val Arg Asn Ala Gly Gly Lys Tyr Lys Phe 165 170 175 Gly Arg Ser Thr Glu Lys Glu Gln Ile Leu Val Lys Ile Lys Gln Phe 180 185 190 Glu Asp Asp Glu Ala Val Val Ile Gly Val Ser Ala Leu Gln Thr Asn 195 200 205 Thr Asn Asp Lys Lys Leu Asn Gln Leu Gly Glu Met Arg Arg Thr Ser 210 215 220 His Gln Asp Gly Lys Val Glu Leu Glu Met Leu Gly Ala Leu Asp Val 225 230 235 240 Asp Trp Asn Gly Ile Arg Phe Ser Ile Gly Thr Gly Phe Asp Arg Asp 245 250 255 Thr Arg Val Asp Leu Trp Lys Arg Arg Glu Gly Val Ile Gly Lys Ile 260 265 270 Val Lys Phe Lys Tyr Phe Ser Gln Gly Ile Lys Thr Ala Pro Arg Phe 275 280 285 Pro Val Phe Leu Gly Phe Arg Asp Lys Asp Asp Met 290 295 300 <210> 52 <211> 300 <212> PRT <213> Acanthocystis turfacea Chlorella virus Br0604L <400> 52 Met Ala Ile Gln Lys Pro Leu Leu Ala Ala Ser Leu Lys Lys Leu Ser 1 5 10 15 Val Asp Asp Leu Thr Phe Pro Val Tyr Ala Thr Pro Lys Leu Asp Gly 20 25 30 Ile Arg Ala Leu Lys Ile Asp Gly Thr Leu Val Ser Arg Thr Phe Lys 35 40 45 Pro Ile Arg Asn Thr Thr Ile Ser Lys Val Leu Thr Ser Leu Leu Pro 50 55 60 Asp Gly Ser Asp Gly Glu Ile Leu Ser Gly Lys Thr Phe Gln Asp Ser 65 70 75 80 Thr Ser Thr Val Met Ser Ala Asp Ala Gly Ile Gly Ser Gly Thr Thr 85 90 95 Phe Phe Trp Phe Asp Tyr Val Lys Asp Asp Pro Asn Lys Gly Tyr Leu 100 105 110 Asp Arg Ile Ala Asp Ile Lys Lys Phe Ile Asp Cys Arg Pro Glu Ile 115 120 125 Leu Lys Asp Ser Arg Val Ile Ile Val Pro Leu Phe Pro Lys Lys Ile 130 135 140 Asp Thr Ala Glu Glu Leu Asn Val Phe Glu Lys Trp Cys Leu Asp Gln 145 150 155 160 Gly Phe Glu Gly Val Met Val Arg Asn Ala Gly Gly Lys Tyr Lys Phe 165 170 175 Gly Arg Ser Thr Glu Lys Glu Gln Ile Leu Val Lys Ile Lys Gln Phe 180 185 190 Glu Asp Asp Glu Ala Val Val Ile Gly Val Ser Ala Leu Gln Thr Asn 195 200 205 Thr Asn Asp Lys Lys Val Asn Glu Leu Gly Glu Met Arg Arg Thr Ser 210 215 220 His Gln Asp Gly Lys Val Asp Leu Asp Met Leu Gly Ala Leu Asp Val 225 230 235 240 Asp Trp Asn Gly Ile Arg Phe Gly Ile Gly Thr Gly Phe Asp Lys Asp 245 250 255 Thr Arg Glu Asp Leu Trp Lys Arg Arg Asp Ser Ile Ile Gly Lys Ile 260 265 270 Val Lys Phe Lys Tyr Phe Ser Gln Gly Val Lys Thr Ala Pro Arg Phe 275 280 285 Pro Val Phe Leu Gly Phe Arg Asp Lys Asn Asp Met 290 295 300 <210> 53 <211> 300 <212> PRT <213> Acanthocystis turfacea Chlorella virus NE-JV-2 <400> 53 Met Ala Ile Gln Lys Pro Leu Leu Ala Ala Ser Leu Lys Lys Leu Ser 1 5 10 15 Val Asp Asp Leu Thr Phe Pro Val Tyr Ala Thr Pro Lys Leu Asp Gly 20 25 30 Ile Arg Ala Leu Lys Ile Asp Gly Thr Ile Val Ser Arg Thr Phe Lys 35 40 45 Pro Ile Arg Asn Thr Thr Ile Ser Asn Val Leu Met Ser Leu Leu Pro 50 55 60 Asp Gly Ser Asp Gly Glu Ile Leu Ser Gly Lys Thr Phe Gln Asp Ser 65 70 75 80 Thr Ser Thr Val Met Ser Ala Asp Ala Gly Ile Gly Ser Gly Thr Thr 85 90 95 Phe Phe Trp Phe Asp Tyr Val Lys Asp Asp Pro Asp Lys Gly Tyr Leu 100 105 110 Asp Arg Ile Ala Asp Met Lys Lys Phe Val Asp Ser His Pro Glu Ile 115 120 125 Leu Lys Asp Arg Arg Val Thr Ile Val Pro Leu Ile Pro Lys Lys Ile 130 135 140 Asp Thr Val Glu Glu Leu Asn Val Phe Glu Gln Trp Cys Leu Asp Gln 145 150 155 160 Gly Phe Glu Gly Val Met Val Arg Asn Ala Gly Gly Lys Tyr Lys Phe 165 170 175 Gly Arg Ser Thr Glu Lys Glu Gln Ile Leu Val Lys Ile Lys Gln Phe 180 185 190 Glu Asp Asp Glu Ala Val Val Ile Gly Val Ser Ala Leu Gln Thr Asn 195 200 205 Val Asn Asp Lys Lys Met Asn Glu Leu Gly Asp Met Arg Arg Thr Ser 210 215 220 His Lys Asp Gly Lys Ile Asp Leu Glu Met Leu Gly Ala Leu Asp Val 225 230 235 240 Glu Trp Asn Gly Ile Arg Phe Gly Ile Gly Thr Gly Phe Asp Lys Asp 245 250 255 Thr Arg Glu Asp Leu Trp Lys Lys Arg Asp Ser Ile Ile Gly Lys Val 260 265 270 Val Lys Phe Lys Tyr Phe Ser Gln Gly Ile Lys Thr Ala Pro Arg Phe 275 280 285 Pro Val Phe Leu Gly Phe Arg Asp Glu Asn Asp Met 290 295 300 <210> 54 <211> 300 <212> PRT <213> Acanthocystis turfacea Chlorella virus TN603.4.2 <400> 54 Met Ala Ile Gln Lys Pro Leu Leu Ala Ala Ser Leu Lys Lys Met Ser 1 5 10 15 Val Asp Asn Leu Thr Phe Pro Val Tyr Ala Thr Pro Lys Leu Asp Gly 20 25 30 Ile Arg Ala Leu Lys Ile Asp Gly Thr Leu Val Ser Arg Thr Phe Lys 35 40 45 Pro Ile Arg Asn Thr Thr Ile Ser Lys Val Leu Ala Ser Leu Leu Pro 50 55 60 Asp Gly Ser Asp Gly Glu Ile Leu Ser Gly Lys Thr Phe Gln Asp Ser 65 70 75 80 Thr Ser Thr Val Met Thr Thr Asp Ala Gly Ile Gly Ser Asp Thr Thr 85 90 95 Phe Phe Trp Phe Asp Tyr Val Lys Asp Asp Pro Asp Lys Gly Tyr Leu 100 105 110 Asp Arg Ile Ala Asp Met Lys Thr Phe Val Asp Gln His Pro Glu Ile 115 120 125 Leu Lys Asp Ser Cys Val Thr Ile Val Pro Leu Phe Pro Lys Lys Ile 130 135 140 Asp Thr Pro Glu Glu Leu His Val Phe Glu Lys Trp Cys Leu Asp Gln 145 150 155 160 Gly Phe Glu Gly Val Met Val Arg Thr Ala Gly Gly Lys Tyr Lys Phe 165 170 175 Gly Arg Ser Thr Glu Lys Glu Gln Ile Leu Val Lys Ile Lys Gln Phe 180 185 190 Glu Asp Asp Glu Ala Val Val Ile Gly Val Ser Ala Leu Gln Thr Asn 195 200 205 Thr Asn Asp Lys Lys Leu Asn Gln Leu Gly Glu Met Arg Arg Thr Ser 210 215 220 His Gln Asp Gly Lys Val Asp Leu Asp Met Leu Gly Ala Leu Asp Val 225 230 235 240 Asp Trp Asn Gly Ile Arg Phe Ser Ile Gly Thr Gly Phe Asp Lys Asp 245 250 255 Thr Arg Glu Asp Leu Trp Lys Gln Arg Asp Ser Ile Val Gly Lys Val 260 265 270 Val Lys Phe Lys Tyr Phe Ser Gln Gly Ile Lys Thr Ala Pro Arg Phe 275 280 285 Pro Val Phe Leu Gly Phe Arg Asp Glu Asn Asp Met 290 295 300 <210> 55 <211> 300 <212> PRT <213> Acanthocystis turfacea Chlorella virus GM0701.1 <400> 55 Met Ala Ile Gln Lys Pro Leu Leu Ala Ala Ser Leu Lys Lys Met Ser 1 5 10 15 Val Asp Asp Leu Thr Phe Pro Val Tyr Thr Thr Pro Lys Leu Asp Gly 20 25 30 Ile Arg Ala Leu Lys Ile Asp Gly Thr Leu Val Ser Arg Thr Phe Lys 35 40 45 Pro Val Arg Asn Ser Ala Ile Ser Glu Val Leu Ala Ser Leu Leu Pro 50 55 60 Asp Gly Ser Asp Gly Glu Ile Leu Ser Gly Lys Thr Phe Gln Asp Ser 65 70 75 80 Thr Ser Thr Val Met Thr Thr Asp Ala Gly Ile Gly Ser Asp Thr Thr 85 90 95 Phe Phe Trp Phe Asp Tyr Val Lys Asp Asp Pro Asn Lys Gly Tyr Leu 100 105 110 Asp Arg Ile Ala Asp Met Lys Thr Phe Ile Asp Gln His Pro Glu Met 115 120 125 Leu Lys Asp Asn His Val Thr Ile Val Pro Leu Ile Pro Lys Lys Ile 130 135 140 Asp Thr Val Glu Glu Leu Asn Ile Phe Glu Lys Trp Cys Leu Asp Gln 145 150 155 160 Gly Phe Glu Gly Val Met Val Arg Asn Ala Gly Gly Lys Tyr Lys Phe 165 170 175 Gly Arg Ser Thr Glu Lys Glu Gln Ile Leu Val Lys Ile Lys Gln Phe 180 185 190 Glu Asp Asp Glu Ala Val Val Ile Gly Val Ser Ala Leu Gln Thr Asn 195 200 205 Thr Asn Asp Lys Lys Leu Asn Gln Leu Gly Glu Met Arg Arg Thr Ser 210 215 220 His Gln Asp Gly Lys Ile Asp Leu Glu Met Leu Gly Ala Leu Asp Val 225 230 235 240 Asp Trp Asn Gly Ile Arg Phe Ser Ile Gly Thr Gly Phe Asp Arg Asp 245 250 255 Thr Arg Val Asp Leu Trp Lys Arg Arg Asp Gly Ile Val Gly Arg Thr 260 265 270 Ile Lys Phe Lys Tyr Phe Gly Gln Gly Ile Lys Thr Ala Pro Arg Phe 275 280 285 Pro Val Phe Leu Gly Phe Arg Asp Lys Asp Asp Met 290 295 300 <210> 56 <211> 287 <212> PRT <213> Synechococcus phage S-CRM01 <400> 56 Met Leu Ala Gly Asn Phe Asp Pro Lys Lys Ala Lys Phe Pro Tyr Cys 1 5 10 15 Ala Thr Pro Lys Ile Asp Gly Ile Arg Phe Leu Met Val Asn Gly Arg 20 25 30 Ala Leu Ser Arg Thr Phe Lys Pro Ile Arg Asn Glu Tyr Ile Gln Lys 35 40 45 Leu Leu Ser Lys His Leu Pro Asp Gly Ile Asp Gly Glu Leu Thr Cys 50 55 60 Gly Asp Thr Phe Gln Ser Ser Thr Ser Ala Ile Met Arg Ile Ala Gly 65 70 75 80 Glu Pro Asp Phe Lys Ala Trp Ile Phe Asp Tyr Val Asp Pro Asp Ser 85 90 95 Thr Ser Ile Leu Pro Phe Ile Glu Arg Phe Asp Gln Ile Ser Asp Ile 100 105 110 Ile Tyr Asn Gly Pro Ile Pro Phe Lys His Gln Val Leu Gly Gln Ser 115 120 125 Ile Leu Tyr Asn Ile Asp Asp Leu Asn Arg Tyr Glu Glu Ala Cys Leu 130 135 140 Asn Glu Gly Tyr Glu Gly Val Met Leu Arg Asp Pro Tyr Gly Thr Tyr 145 150 155 160 Lys Phe Gly Arg Ser Ser Thr Asn Glu Gly Ile Leu Leu Lys Val Lys 165 170 175 Arg Phe Glu Asp Ala Glu Ala Thr Val Ile Arg Ile Asp Glu Lys Met 180 185 190 Ser Asn Gln Asn Ile Ala Glu Lys Asp Asn Phe Gly Arg Thr Lys Arg 195 200 205 Ser Ser Cys Leu Asp Gly Met Val Pro Met Glu Thr Thr Gly Ala Leu 210 215 220 Phe Val Arg Asn Ser Asp Gly Leu Glu Phe Ser Ile Gly Ser Gly Leu 225 230 235 240 Asn Asp Glu Met Arg Asp Glu Ile Trp Lys Asn Lys Ser Ser Tyr Ile 245 250 255 Gly Lys Leu Val Lys Tyr Lys Tyr Phe Pro Gln Gly Val Lys Asp Leu 260 265 270 Pro Arg His Pro Val Phe Leu Gly Phe Arg Asp Pro Asp Asp Met 275 280 285 <210> 57 <211> 322 <212> PRT <213> Marine sediment metagenome <400> 57 Met Asp Ala His Glu Leu Met Lys Leu Asn Glu Tyr Ala Glu Arg Gln 1 5 10 15 Asn Gln Lys Gln Lys Lys Gln Ile Thr Lys Pro Met Leu Ala Ala Ser 20 25 30 Leu Lys Asp Ile Thr Gln Leu Asp Tyr Ser Lys Gly Tyr Leu Ala Thr 35 40 45 Gln Lys Leu Asp Gly Ile Arg Ala Leu Met Ile Asp Gly Lys Leu Val 50 55 60 Ser Arg Thr Phe Lys Pro Ile Arg Asn Asn His Ile Arg Glu Met Leu 65 70 75 80 Glu Asp Val Leu Pro Asp Gly Ala Asp Gly Glu Ile Val Cys Pro Gly 85 90 95 Ala Phe Gln Ala Thr Ser Ser Gly Val Met Ser Ala Asn Gly Glu Pro 100 105 110 Glu Phe Ile Tyr Tyr Met Phe Asp Tyr Val Lys Asp Asp Ile Thr Lys 115 120 125 Glu Tyr Trp Arg Arg Thr Gln Asp Met Val Gln Trp Leu Ile Asn Gln 130 135 140 Gly Pro Thr Arg Thr Pro Gly Leu Ser Lys Leu Lys Leu Leu Val Pro 145 150 155 160 Thr Leu Ile Lys Asn Tyr Asp His Leu Lys Thr Tyr Glu Thr Glu Cys 165 170 175 Ile Asp Lys Gly Phe Glu Gly Val Ile Leu Arg Thr Pro Asp Ser Pro 180 185 190 Tyr Lys Cys Gly Arg Ser Thr Ala Lys Gln Glu Trp Leu Leu Lys Leu 195 200 205 Lys Arg Phe Ala Asp Asp Glu Ala Val Val Ile Gly Phe Thr Glu Lys 210 215 220 Met His Asn Asp Asn Glu Ala Thr Lys Asp Lys Phe Gly His Thr Val 225 230 235 240 Arg Ser Ser His Lys Glu Asn Lys Arg Pro Ala Gly Thr Leu Gly Ser 245 250 255 Leu Ile Val Arg Asp Ile Lys Thr Glu Ile Glu Phe Glu Ile Gly Thr 260 265 270 Gly Phe Asp Asp Glu Leu Arg Gln Lys Ile Trp Asp Ala Arg Pro Glu 275 280 285 Trp Asp Gly Leu Cys Val Lys Tyr Lys His Phe Ala Ile Ser Gly Val 290 295 300 Lys Glu Lys Pro Arg Phe Pro Ser Phe Ile Gly Val Arg Asp Val Glu 305 310 315 320 Asp Met <210> 58 <211> 691 <212> PRT <213> Mycobacterium tuberculosis (strain ATCC 25618 / H37Rv) <400> 58 Met Ser Ser Pro Asp Ala Asp Gln Thr Ala Pro Glu Val Leu Arg Gln 1 5 10 15 Trp Gln Ala Leu Ala Glu Glu Val Arg Glu His Gln Phe Arg Tyr Tyr 20 25 30 Val Arg Asp Ala Pro Ile Ile Ser Asp Ala Glu Phe Asp Glu Leu Leu 35 40 45 Arg Arg Leu Glu Ala Leu Glu Glu Gln His Pro Glu Leu Arg Thr Pro 50 55 60 Asp Ser Pro Thr Gln Leu Val Gly Gly Ala Gly Phe Ala Thr Asp Phe 65 70 75 80 Glu Pro Val Asp His Leu Glu Arg Met Leu Ser Leu Asp Asn Ala Phe 85 90 95 Thr Ala Asp Glu Leu Ala Ala Trp Ala Gly Arg Ile His Ala Glu Val 100 105 110 Gly Asp Ala Ala His Tyr Leu Cys Glu Leu Lys Ile Asp Gly Val Ala 115 120 125 Leu Ser Leu Val Tyr Arg Glu Gly Arg Leu Thr Arg Ala Ser Thr Arg 130 135 140 Gly Asp Gly Arg Thr Gly Glu Asp Val Thr Leu Asn Ala Arg Thr Ile 145 150 155 160 Ala Asp Val Pro Glu Arg Leu Thr Pro Gly Asp Asp Tyr Pro Val Pro 165 170 175 Glu Val Leu Glu Val Arg Gly Glu Val Phe Phe Arg Leu Asp Asp Phe 180 185 190 Gln Ala Leu Asn Ala Ser Leu Val Glu Glu Gly Lys Ala Pro Phe Ala 195 200 205 Asn Pro Arg Asn Ser Ala Ala Gly Ser Leu Arg Gln Lys Asp Pro Ala 210 215 220 Val Thr Ala Arg Arg Arg Leu Arg Met Ile Cys His Gly Leu Gly His 225 230 235 240 Val Glu Gly Phe Arg Pro Ala Thr Leu His Gln Ala Tyr Leu Ala Leu 245 250 255 Arg Ala Trp Gly Leu Pro Val Ser Glu His Thr Thr Leu Ala Thr Asp 260 265 270 Leu Ala Gly Val Arg Glu Arg Ile Asp Tyr Trp Gly Glu His Arg His 275 280 285 Glu Val Asp His Glu Ile Asp Gly Val Val Val Lys Val Asp Glu Val 290 295 300 Ala Leu Gln Arg Arg Leu Gly Ser Thr Ser Arg Ala Pro Arg Trp Ala 305 310 315 320 Ile Ala Tyr Lys Tyr Pro Pro Glu Glu Ala Gln Thr Lys Leu Leu Asp 325 330 335 Ile Arg Val Asn Val Gly Arg Thr Gly Arg Ile Thr Pro Phe Ala Phe 340 345 350 Met Thr Pro Val Lys Val Ala Gly Ser Thr Val Gly Gln Ala Thr Leu 355 360 365 His Asn Ala Ser Glu Ile Lys Arg Lys Gly Val Leu Ile Gly Asp Thr 370 375 380 Val Val Ile Arg Lys Ala Gly Asp Val Ile Pro Glu Val Leu Gly Pro 385 390 395 400 Val Val Glu Leu Arg Asp Gly Ser Glu Arg Glu Phe Ile Met Pro Thr 405 410 415 Thr Cys Pro Glu Cys Gly Ser Pro Leu Ala Pro Glu Lys Glu Gly Asp 420 425 430 Ala Asp Ile Arg Cys Pro Asn Ala Arg Gly Cys Pro Gly Gln Leu Arg 435 440 445 Glu Arg Val Phe His Val Ala Ser Arg Asn Gly Leu Asp Ile Glu Val 450 455 460 Leu Gly Tyr Glu Ala Gly Val Ala Leu Leu Gln Ala Lys Val Ile Ala 465 470 475 480 Asp Glu Gly Glu Leu Phe Ala Leu Thr Glu Arg Asp Leu Leu Arg Thr 485 490 495 Asp Leu Phe Arg Thr Lys Ala Gly Glu Leu Ser Ala Asn Gly Lys Arg 500 505 510 Leu Leu Val Asn Leu Asp Lys Ala Lys Ala Ala Pro Leu Trp Arg Val 515 520 525 Leu Val Ala Leu Ser Ile Arg His Val Gly Pro Thr Ala Ala Arg Ala 530 535 540 Leu Ala Thr Glu Phe Gly Ser Leu Asp Ala Ile Ala Ala Ala Ser Thr 545 550 555 560 Asp Gln Leu Ala Ala Val Glu Gly Val Gly Pro Thr Ile Ala Ala Ala 565 570 575 Val Thr Glu Trp Phe Ala Val Asp Trp His Arg Glu Ile Val Asp Lys 580 585 590 Trp Arg Ala Ala Gly Val Arg Met Val Asp Glu Arg Asp Glu Ser Val 595 600 605 Pro Arg Thr Leu Ala Gly Leu Thr Ile Val Val Thr Gly Ser Leu Thr 610 615 620 Gly Phe Ser Arg Asp Asp Ala Lys Glu Ala Ile Val Ala Arg Gly Gly 625 630 635 640 Lys Ala Ala Gly Ser Val Ser Lys Lys Thr Asn Tyr Val Val Ala Gly 645 650 655 Asp Ser Pro Gly Ser Lys Tyr Asp Lys Ala Val Glu Leu Gly Val Pro 660 665 670 Ile Leu Asp Glu Asp Gly Phe Arg Arg Leu Leu Ala Asp Gly Pro Ala 675 680 685 Ser Arg Thr 690 <210> 59 <211> 676 <212> PRT <213> Enterococcus faecalis (strain ATCC 700802 / V583) <400> 59 Met Glu Gln Gln Pro Leu Thr Leu Thr Ala Ala Thr Thr Arg Ala Gln 1 5 10 15 Glu Leu Arg Lys Gln Leu Asn Gln Tyr Ser His Glu Tyr Tyr Val Lys 20 25 30 Asp Gln Pro Ser Val Glu Asp Tyr Val Tyr Asp Arg Leu Tyr Lys Glu 35 40 45 Leu Val Asp Ile Glu Thr Glu Phe Pro Asp Leu Ile Thr Pro Asp Ser 50 55 60 Pro Thr Gln Arg Val Gly Gly Lys Val Leu Ser Gly Phe Glu Lys Ala 65 70 75 80 Pro His Asp Ile Pro Met Tyr Ser Leu Asn Asp Gly Phe Ser Lys Glu 85 90 95 Asp Ile Phe Ala Phe Asp Glu Arg Val Arg Lys Ala Ile Gly Lys Pro 100 105 110 Val Ala Tyr Cys Cys Glu Leu Lys Ile Asp Gly Leu Ala Ile Ser Leu 115 120 125 Arg Tyr Glu Asn Gly Val Phe Val Arg Gly Ala Thr Arg Gly Asp Gly 130 135 140 Thr Val Gly Glu Asn Ile Thr Glu Asn Leu Arg Thr Val Arg Ser Val 145 150 155 160 Pro Met Arg Leu Thr Glu Pro Ile Ser Val Glu Val Arg Gly Glu Cys 165 170 175 Tyr Met Pro Lys Gln Ser Phe Val Ala Leu Asn Glu Glu Arg Glu Glu 180 185 190 Asn Gly Gln Asp Ile Phe Ala Asn Pro Arg Asn Ala Ala Ala Gly Ser 195 200 205 Leu Arg Gln Leu Asp Thr Lys Ile Val Ala Lys Arg Asn Leu Asn Thr 210 215 220 Phe Leu Tyr Thr Val Ala Asp Phe Gly Pro Met Lys Ala Lys Thr Gln 225 230 235 240 Phe Glu Ala Leu Glu Glu Leu Ser Ala Ile Gly Phe Arg Thr Asn Pro 245 250 255 Glu Arg Gln Leu Cys Gln Ser Ile Asp Glu Val Trp Ala Tyr Ile Glu 260 265 270 Glu Tyr His Glu Lys Arg Ser Thr Leu Pro Tyr Glu Ile Asp Gly Ile 275 280 285 Val Ile Lys Val Asn Glu Phe Ala Leu Gln Asp Glu Leu Gly Phe Thr 290 295 300 Val Lys Ala Pro Arg Trp Ala Ile Ala Tyr Lys Phe Pro Pro Glu Glu 305 310 315 320 Ala Glu Thr Val Val Glu Asp Ile Glu Trp Thr Ile Gly Arg Thr Gly 325 330 335 Val Val Thr Pro Thr Ala Val Met Ala Pro Val Arg Val Ala Gly Thr 340 345 350 Thr Val Ser Arg Ala Ser Leu His Asn Ala Asp Phe Ile Gln Met Lys 355 360 365 Asp Ile Arg Leu Asn Asp His Val Ile Ile Tyr Lys Ala Gly Asp Ile 370 375 380 Ile Pro Glu Val Ala Gln Val Leu Val Glu Lys Arg Ala Ala Asp Ser 385 390 395 400 Gln Pro Tyr Glu Met Pro Thr His Cys Pro Ile Cys His Ser Glu Leu 405 410 415 Val His Leu Asp Glu Glu Val Ala Leu Arg Cys Ile Asn Pro Lys Cys 420 425 430 Pro Ala Gln Ile Lys Glu Gly Leu Asn His Phe Val Ser Arg Asn Ala 435 440 445 Met Asn Ile Asp Gly Leu Gly Pro Arg Val Leu Ala Gln Met Tyr Asp 450 455 460 Lys Gly Leu Val Lys Asp Val Ala Asp Leu Tyr Phe Leu Thr Glu Glu 465 470 475 480 Gln Leu Met Thr Leu Asp Lys Ile Lys Glu Lys Ser Ala Asn Asn Ile 485 490 495 Tyr Thr Ala Ile Gln Gly Ser Lys Glu Asn Ser Val Glu Arg Leu Ile 500 505 510 Phe Gly Leu Gly Ile Arg His Val Gly Ala Lys Ala Ala Lys Ile Leu 515 520 525 Ala Glu His Phe Gly Asp Leu Pro Thr Leu Ser Arg Ala Thr Ala Glu 530 535 540 Glu Ile Val Ala Leu Asp Ser Ile Gly Glu Thr Ile Ala Asp Ser Val 545 550 555 560 Val Thr Tyr Phe Glu Asn Glu Glu Val His Glu Leu Met Ala Glu Leu 565 570 575 Glu Lys Ala Gln Val Asn Leu Thr Tyr Lys Gly Leu Arg Thr Glu Gln 580 585 590 Leu Ala Glu Val Glu Ser Pro Phe Lys Asp Lys Thr Val Val Leu Thr 595 600 605 Gly Lys Leu Ala Gln Tyr Thr Arg Glu Glu Ala Lys Glu Lys Ile Glu 610 615 620 Asn Leu Gly Gly Lys Val Thr Gly Ser Val Ser Lys Lys Thr Asp Ile 625 630 635 640 Val Val Ala Gly Glu Asp Ala Gly Ser Lys Leu Thr Lys Ala Glu Ser 645 650 655 Leu Gly Val Thr Val Trp Asn Glu Gln Glu Met Val Asp Ala Leu Asp 660 665 670 Ala Ser His Phe 675 <210> 60 <211> 670 <212> PRT <213> Haemophilus influenzae (strain ATCC 51907 / DSM 11121 / KW20 / Rd) <400> 60 Met Thr Asn Ile Gln Thr Gln Leu Asp Asn Leu Arg Lys Thr Leu Arg 1 5 10 15 Gln Tyr Glu Tyr Glu Tyr His Val Leu Asp Asn Pro Ser Val Pro Asp 20 25 30 Ser Glu Tyr Asp Arg Leu Phe His Gln Leu Lys Ala Leu Glu Leu Glu 35 40 45 His Pro Glu Phe Leu Thr Ser Asp Ser Pro Thr Gln Arg Val Gly Ala 50 55 60 Lys Pro Leu Ser Gly Phe Ser Gln Ile Arg His Glu Ile Pro Met Leu 65 70 75 80 Ser Leu Asp Asn Ala Phe Ser Asp Ala Glu Phe Asn Ala Phe Val Lys 85 90 95 Arg Ile Glu Asp Arg Leu Ile Leu Leu Pro Lys Pro Leu Thr Phe Cys 100 105 110 Cys Glu Pro Lys Leu Asp Gly Leu Ala Val Ser Ile Leu Tyr Val Asn 115 120 125 Gly Glu Leu Thr Gln Ala Ala Thr Arg Gly Asp Gly Thr Thr Gly Glu 130 135 140 Asp Ile Thr Ala Asn Ile Arg Thr Ile Arg Asn Val Pro Leu Gln Leu 145 150 155 160 Leu Thr Asp Asn Pro Pro Ala Arg Leu Glu Val Arg Gly Glu Val Phe 165 170 175 Met Pro His Ala Gly Phe Glu Arg Leu Asn Lys Tyr Ala Leu Glu His 180 185 190 Asn Glu Lys Thr Phe Ala Asn Pro Arg Asn Ala Ala Ala Gly Ser Leu 195 200 205 Arg Gln Leu Asp Pro Asn Ile Thr Ser Lys Arg Pro Leu Val Leu Asn 210 215 220 Ala Tyr Gly Ile Gly Ile Ala Glu Gly Val Asp Leu Pro Thr Thr His 225 230 235 240 Tyr Ala Arg Leu Gln Trp Leu Lys Ser Ile Gly Ile Pro Val Asn Pro 245 250 255 Glu Ile Arg Leu Cys Asn Gly Ala Asp Glu Val Leu Gly Phe Tyr Arg 260 265 270 Asp Ile Gln Asn Lys Arg Ser Ser Leu Gly Tyr Asp Ile Asp Gly Thr 275 280 285 Val Leu Lys Ile Asn Asp Ile Ala Leu Gln Asn Glu Leu Gly Phe Ile 290 295 300 Ser Lys Ala Pro Arg Trp Ala Ile Ala Tyr Lys Phe Pro Ala Gln Glu 305 310 315 320 Glu Leu Thr Leu Leu Asn Asp Val Glu Phe Gln Val Gly Arg Thr Gly 325 330 335 Ala Ile Thr Pro Val Ala Lys Leu Glu Pro Val Phe Val Ala Gly Val 340 345 350 Thr Val Ser Asn Ala Thr Leu His Asn Gly Asp Glu Ile Glu Arg Leu 355 360 365 Asn Ile Ala Ile Gly Asp Thr Val Val Ile Arg Arg Ala Gly Asp Val 370 375 380 Ile Pro Gln Ile Ile Gly Val Leu His Glu Arg Arg Pro Asp Asn Ala 385 390 395 400 Lys Pro Ile Ile Phe Pro Thr Asn Cys Pro Val Cys Asp Ser Gln Ile 405 410 415 Ile Arg Ile Glu Gly Glu Ala Val Ala Arg Cys Thr Gly Gly Leu Phe 420 425 430 Cys Ala Ala Gln Arg Lys Glu Ala Leu Lys His Phe Val Ser Arg Lys 435 440 445 Ala Met Asp Ile Asp Gly Val Gly Gly Lys Leu Ile Glu Gln Leu Val 450 455 460 Asp Arg Glu Leu Ile His Thr Pro Ala Asp Leu Phe Lys Leu Asp Leu 465 470 475 480 Thr Thr Leu Thr Arg Leu Glu Arg Met Gly Ala Lys Ser Ala Glu Asn 485 490 495 Ala Leu Asn Ser Leu Glu Asn Ala Lys Ser Thr Thr Leu Ala Arg Phe 500 505 510 Ile Phe Ala Leu Gly Ile Arg Glu Val Gly Glu Ala Thr Ala Leu Asn 515 520 525 Leu Ala Asn His Phe Lys Thr Leu Asp Ala Leu Lys Asp Ala Asn Leu 530 535 540 Glu Glu Leu Gln Gln Val Pro Asp Val Gly Glu Val Val Ala Asn Arg 545 550 555 560 Ile Phe Ile Phe Trp Arg Glu Ala His Asn Val Ala Val Val Glu Asp 565 570 575 Leu Ile Ala Gln Gly Val His Trp Glu Thr Val Glu Val Lys Glu Ala 580 585 590 Ser Glu Asn Leu Phe Lys Asp Lys Thr Val Val Leu Thr Gly Thr Leu 595 600 605 Thr Gln Met Gly Arg Asn Glu Ala Lys Ala Leu Leu Gln Gln Leu Gly 610 615 620 Ala Lys Val Ser Gly Ser Val Ser Ser Lys Thr Asp Phe Val Ile Ala 625 630 635 640 Gly Asp Ala Ala Gly Ser Lys Leu Ala Lys Ala Gln Glu Leu Asn Ile 645 650 655 Thr Val Leu Thr Glu Glu Glu Phe Leu Ala Gln Ile Thr Arg 660 665 670 <210> 61 <211> 667 <212> PRT <213> Staphylococcus aureus <400> 61 Met Ala Asp Leu Ser Ser Arg Val Asn Glu Leu His Asp Leu Leu Asn 1 5 10 15 Gln Tyr Ser Tyr Glu Tyr Tyr Val Glu Asp Asn Pro Ser Val Pro Asp 20 25 30 Ser Glu Tyr Asp Lys Leu Leu His Glu Leu Ile Lys Ile Glu Glu Glu 35 40 45 His Pro Glu Tyr Lys Thr Val Asp Ser Pro Thr Val Arg Val Gly Gly 50 55 60 Glu Ala Gln Ala Ser Phe Asn Lys Val Asn His Asp Thr Pro Met Leu 65 70 75 80 Ser Leu Gly Asn Ala Phe Asn Glu Asp Asp Leu Arg Lys Phe Asp Gln 85 90 95 Arg Ile Arg Glu Gln Ile Gly Asn Val Glu Tyr Met Cys Glu Leu Lys 100 105 110 Ile Asp Gly Leu Ala Val Ser Leu Lys Tyr Val Asp Gly Tyr Phe Val 115 120 125 Gln Gly Leu Thr Arg Gly Asp Gly Thr Thr Gly Glu Asp Ile Thr Glu 130 135 140 Asn Leu Lys Thr Ile His Ala Ile Pro Leu Lys Met Lys Glu Pro Leu 145 150 155 160 Asn Val Glu Val Arg Gly Glu Ala Tyr Met Pro Arg Arg Ser Phe Leu 165 170 175 Arg Leu Asn Glu Glu Lys Glu Lys Asn Asp Glu Gln Leu Phe Ala Asn 180 185 190 Pro Arg Asn Ala Ala Ala Gly Ser Leu Arg Gln Leu Asp Ser Lys Leu 195 200 205 Thr Ala Lys Arg Lys Leu Ser Val Phe Ile Tyr Ser Val Asn Asp Phe 210 215 220 Thr Asp Phe Asn Ala Arg Ser Gln Ser Glu Ala Leu Asp Glu Leu Asp 225 230 235 240 Lys Leu Gly Phe Thr Thr Asn Lys Asn Arg Ala Arg Val Asn Asn Ile 245 250 255 Asp Gly Val Leu Glu Tyr Ile Glu Lys Trp Thr Ser Gln Arg Glu Ser 260 265 270 Leu Pro Tyr Asp Ile Asp Gly Ile Val Ile Lys Val Asn Asp Leu Asp 275 280 285 Gln Gln Asp Glu Met Gly Phe Thr Gln Lys Ser Pro Arg Trp Ala Ile 290 295 300 Ala Tyr Lys Phe Pro Ala Glu Glu Val Val Thr Lys Leu Leu Asp Ile 305 310 315 320 Glu Leu Ser Ile Gly Arg Thr Gly Val Val Thr Pro Thr Ala Ile Leu 325 330 335 Glu Pro Val Lys Val Ala Gly Thr Thr Val Ser Arg Ala Ser Leu His 340 345 350 Asn Glu Asp Leu Ile His Asp Arg Asp Ile Arg Ile Gly Asp Ser Val 355 360 365 Val Val Lys Lys Ala Gly Asp Ile Ile Pro Glu Val Val Arg Ser Ile 370 375 380 Pro Glu Arg Arg Pro Glu Asp Ala Val Thr Tyr His Met Pro Thr His 385 390 395 400 Cys Pro Ser Cys Gly His Glu Leu Val Arg Ile Glu Gly Glu Val Ala 405 410 415 Leu Arg Cys Ile Asn Pro Lys Cys Gln Ala Gln Leu Val Glu Gly Leu 420 425 430 Ile His Phe Val Ser Arg Gln Ala Met Asn Ile Asp Gly Leu Gly Thr 435 440 445 Lys Ile Ile Gln Gln Leu Tyr Gln Ser Glu Leu Ile Lys Asp Val Ala 450 455 460 Asp Ile Phe Tyr Leu Thr Glu Glu Asp Leu Leu Pro Leu Asp Arg Met 465 470 475 480 Gly Gln Lys Lys Val Asp Asn Leu Leu Ala Ala Ile Gln Gln Ala Lys 485 490 495 Asp Asn Ser Leu Glu Asn Leu Leu Phe Gly Leu Gly Ile Arg His Leu 500 505 510 Gly Val Lys Ala Ser Gln Val Leu Ala Glu Lys Tyr Glu Thr Ile Asp 515 520 525 Arg Leu Leu Thr Val Thr Glu Ala Glu Leu Val Glu Ile His Asp Ile 530 535 540 Gly Asp Lys Val Ala Gln Ser Val Val Thr Tyr Leu Glu Asn Glu Asp 545 550 555 560 Ile Arg Ala Leu Ile Gln Lys Leu Lys Asp Lys His Val Asn Met Ile 565 570 575 Tyr Lys Gly Ile Lys Thr Ser Asp Ile Glu Gly His Pro Glu Phe Ser 580 585 590 Gly Lys Thr Ile Val Leu Thr Gly Lys Leu His Gln Met Thr Arg Asn 595 600 605 Glu Ala Ser Lys Trp Leu Ala Ser Gln Gly Ala Lys Val Thr Ser Ser 610 615 620 Val Thr Lys Asn Thr Asp Val Val Ile Ala Gly Glu Asp Ala Gly Ser 625 630 635 640 Lys Leu Thr Lys Ala Gln Ser Leu Gly Ile Glu Ile Trp Thr Glu Gln 645 650 655 Gln Phe Val Asp Lys Gln Asn Glu Leu Asn Ser 660 665 <210> 62 <211> 652 <212> PRT <213> Streptococcus pneumoniae (strain P1031) <400> 62 Met Asn Lys Arg Met Asn Glu Leu Val Ala Leu Leu Asn Arg Tyr Ala 1 5 10 15 Thr Glu Tyr Tyr Thr Ser Asp Asn Pro Ser Val Ser Asp Ser Glu Tyr 20 25 30 Asp Arg Leu Tyr Arg Glu Leu Val Glu Leu Glu Thr Ala Tyr Pro Glu 35 40 45 Gln Val Leu Ala Asp Ser Pro Thr His Arg Val Gly Gly Lys Val Leu 50 55 60 Asp Gly Phe Glu Lys Tyr Ser His Gln Tyr Pro Leu Tyr Ser Leu Gln 65 70 75 80 Asp Ala Phe Ser Arg Glu Glu Leu Asp Ala Phe Asp Ala Arg Val Arg 85 90 95 Lys Glu Val Ala His Pro Thr Tyr Ile Cys Glu Leu Lys Ile Asp Gly 100 105 110 Leu Ser Ile Ser Leu Thr Tyr Glu Lys Gly Ile Leu Val Ala Gly Val 115 120 125 Thr Arg Gly Asp Gly Ser Ile Gly Glu Asn Ile Thr Glu Asn Leu Lys 130 135 140 Arg Val Lys Asp Ile Pro Leu Thr Leu Pro Glu Glu Leu Asp Ile Thr 145 150 155 160 Val Arg Gly Glu Cys Tyr Met Pro Arg Ala Ser Phe Asp Gln Val Asn 165 170 175 Gln Ala Arg Gln Glu Asn Gly Glu Pro Glu Phe Ala Asn Pro Arg Asn 180 185 190 Ala Ala Ala Gly Thr Leu Arg Gln Leu Asp Thr Ala Val Val Ala Lys 195 200 205 Arg Asn Leu Ala Thr Phe Leu Tyr Gln Glu Ala Ser Pro Ser Thr Arg 210 215 220 Asp Ser Gln Glu Lys Gly Leu Lys Tyr Leu Glu Gln Leu Gly Phe Val 225 230 235 240 Val Asn Pro Lys Arg Ile Leu Ala Glu Asn Ile Asp Glu Ile Trp Asn 245 250 255 Phe Ile Gln Glu Val Gly Gln Glu Arg Glu Asn Leu Pro Tyr Asp Ile 260 265 270 Asp Gly Val Val Ile Lys Val Asn Asp Leu Ala Ser Gln Glu Glu Leu 275 280 285 Gly Phe Thr Val Lys Ala Pro Lys Trp Ala Val Ala Tyr Lys Phe Pro 290 295 300 Ala Glu Glu Lys Glu Ala Gln Leu Leu Ser Val Asp Trp Thr Val Gly 305 310 315 320 Arg Thr Gly Val Val Thr Pro Thr Ala Asn Leu Thr Pro Val Gln Leu 325 330 335 Ala Gly Thr Thr Val Ser Arg Ala Thr Leu His Asn Val Asp Tyr Ile 340 345 350 Ala Glu Lys Asp Ile Arg Lys Asp Asp Thr Val Ile Val Tyr Lys Ala 355 360 365 Gly Asp Ile Ile Pro Ala Val Leu Arg Val Val Glu Ser Lys Arg Val 370 375 380 Ser Glu Glu Lys Leu Asp Ile Pro Thr Asn Cys Pro Ser Cys Asn Ser 385 390 395 400 Asp Leu Leu His Phe Glu Asp Glu Val Ala Leu Arg Cys Ile Asn Pro 405 410 415 Arg Cys Pro Ala Gln Ile Met Glu Gly Leu Ile His Phe Ala Ser Arg 420 425 430 Asp Ala Met Asn Ile Thr Gly Leu Gly Pro Ser Ile Val Glu Lys Leu 435 440 445 Phe Ala Ala Asn Leu Val Lys Asp Val Ala Asp Ile Tyr Arg Leu Gln 450 455 460 Glu Glu Asp Phe Leu Leu Leu Glu Gly Val Lys Glu Lys Ser Ala Ala 465 470 475 480 Lys Leu Tyr Gln Ala Ile Gln Ala Ser Lys Glu Asn Ser Ala Glu Lys 485 490 495 Leu Leu Phe Gly Leu Gly Ile Arg His Val Gly Ser Lys Ala Ser Gln 500 505 510 Leu Leu Leu Gln Tyr Phe His Ser Ile Glu Asn Leu Tyr Gln Ala Asp 515 520 525 Ser Glu Glu Val Ala Ser Ile Glu Ser Leu Gly Gly Val Ile Ala Lys 530 535 540 Ser Leu Gln Thr Tyr Phe Ala Thr Glu Gly Ser Glu Ile Leu Leu Arg 545 550 555 560 Glu Leu Lys Glu Thr Gly Val Asn Leu Asp Tyr Lys Gly Gln Thr Val 565 570 575 Val Ala Asp Ala Ala Leu Ser Gly Leu Thr Val Val Leu Thr Gly Lys 580 585 590 Leu Glu Arg Leu Lys Arg Ser Glu Ala Lys Ser Lys Leu Glu Ser Leu 595 600 605 Gly Ala Lys Val Thr Gly Ser Val Ser Lys Lys Thr Asp Leu Val Val 610 615 620 Val Gly Ala Asp Ala Gly Ser Lys Leu Gln Lys Ala Gln Glu Leu Gly 625 630 635 640 Ile Gln Val Arg Asp Glu Ala Trp Leu Glu Ser Leu 645 650

Claims

1. A method for purifying a single-stranded oligonucleotide product from impurities, wherein the single-stranded oligonucleotide has at least one modified nucleotide residue, the modification selected from the group consisting of: modification of the 2' position of the sugar moiety of the nucleotide residue, modification of the nucleobase of the nucleotide residue, and modification of the backbone of the nucleotide residue, the method comprising: a) providing a template oligonucleotide (I) complementary to the sequence of the single-stranded oligonucleotide product, said template having properties allowing it to be separated from the single-stranded oligonucleotide product; b) providing a pool of oligonucleotides (II) comprising the single-stranded oligonucleotide product and impurities; c) contacting (I) and (II) under conditions allowing annealing; d) changing the conditions to separate any impurities, including denaturing the annealed template and impurity oligonucleotide chains and separating the impurities; e) changing the conditions to separate the single-stranded oligonucleotide product, comprising denaturing the annealed template and the single-stranded oligonucleotide product and separating the single-stranded oligonucleotide product; and f) recycling the template.

2. The method of claim 1, wherein the denaturation is caused by increasing the temperature, changing the pH or changing the salt concentration of the buffer solution.

3. The method according to claim 1 or 2, comprising two steps of increasing the temperature to: i) denature any annealed impurities and ii) denature the annealed single-stranded oligonucleotide products.

4. The method of claim 1 or 2, wherein the single-stranded oligonucleotide product is 10-200 nucleotides long. The method of claim 4 , wherein the single-stranded oligonucleotide product is 20-30 nucleotides long. The method of claim 4 , wherein the single-stranded oligonucleotide product is 20-25 nucleotides in length.

7. The method according to claim 1 or 2, wherein the property allowing separation of the template from the single-stranded oligonucleotide product is that the template is attached to a support material. The method according to claim 7 , wherein the support material is a soluble support material.

9. The method of claim 8, wherein the support material is selected from the group consisting of polyethylene glycol, soluble organic polymers, DNA, proteins, dendrimers, polysaccharides, oligosaccharides and carbohydrates.

10. The method of claim 7, wherein the support material is an insoluble support material.

11. The method according to claim 10, wherein the insoluble support material is selected from the group consisting of glass beads, polymer beads, fibrous supports, membranes, streptavidin-coated beads, cellulose; or is part of the reaction vessel itself or the reaction wall.

12. The method of claim 7, wherein multiple replicate copies of the template are attached to the support material in a serial manner via a single point of attachment.

13. The method of claim 1 or 2, wherein the property that allows separation of the template from the single-stranded oligonucleotide product is the molecular weight of the template.

14. The method of claim 1 or 2, wherein the modification is selected from 2'-F, 2'-OMe, 2'-MOE and 2'-amino, or wherein the oligonucleotide comprises a PMO, LNA, PNA or BNA.

15. The method of claim 1 or 2, wherein the modification is selected from 5-methylpyrimidine, 7-deazaguanosine, and abasic nucleotides.

16. The method of claim 1 or 2, wherein the modification is selected from phosphorothioate, phosphoramidate, and phosphorodiamidate.

17. The method of claim 1 or 2, wherein the resulting single-stranded oligonucleotide product is at least 90% pure.

18. The method of claim 17, wherein the resulting single-stranded oligonucleotide product is at least 95% pure.

19. The method of claim 17, wherein the resulting single-stranded oligonucleotide product is at least 98% pure.

20. A method as claimed in claim 1 or 2, wherein the method is used to purify a therapeutic single stranded oligonucleotide product.

21. A method as claimed in claim 1 or 2, wherein the single stranded oligonucleotide product is purified on a gram or kilogram scale, and / or the method is carried out in a reactor of 1 L or larger.

22. The process as claimed in claim 1 or 2, wherein the reaction is carried out using a continuous or semi-continuous flow process.

Citation Information

Patent Citations

  • Tm leveling methods

    US6683173B2

  • Purification of oligonucleotides

    WO2001055160A1

  • Process for separation of oligonucleotide of interest from a mixture

    WO2013030263A1

  • Single-stranded nucleic acid template-mediated recombination and nucleic acid fragment isolation

    WO2001064864A2

  • Synthesis of single-stranded DNA

    WO2009097673A1