Methods for modifying glycoproteins with beta-(1,4)-n-acetylgalactosaminyltransferase or mutants thereof
By using the enzymatic reaction of β-(1,4)-N-acetylgalactosamine transferase or its mutant with sugar derivative nucleotides, the problem of low glycoprotein modification efficiency in the prior art has been solved, achieving efficient and selective glycoprotein modification, which is suitable for glycoprotein labeling and detection in vitro and in vivo.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SYNAFFIX BV
- Filing Date
- 2015-08-04
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies struggle to effectively utilize β-(1,4)-N-acetylgalactosamine transferase to modify glycoproteins, particularly through non-natural GalNAc derivatives such as 2-keto or 2-azidoacetyl derivatives, resulting in low glycoprotein modification efficiency.
β-(1,4)-N-acetylgalactosamine transferase or its mutants are used in the presence of glycoproteins to contact the sugar derivative nucleotide Su(A)-Nuc, and the sugar derivative is linked to a glycan containing a terminal GlcNAc moiety through an enzymatic reaction to form a specifically modified glycoprotein.
It achieves efficient and selective modification of glycoproteins, improving the efficiency and specificity of glycoprotein modification, and is suitable for in vitro and in vivo glycoprotein labeling and detection.
Smart Images

Figure CN107109454B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for enzymatically modifying glycoproteins. More specifically, this invention relates to a method for modifying glycoproteins with sugar derivative nucleotides using β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof, and to β-(1,4)-N-acetylgalactosamine transferase mutants. Background Technology
[0002] Glycosyltransferases constitute a superfamily of enzymes involved in the synthesis of complex carbohydrates present on glycoproteins and glycolipids. The fundamental function of glycosyltransferases is to transfer the glycosyl moiety of nucleotide derivatives to specific sugar receptors. β-1,4-galactosyltransferase (β4Gal-T) (EC 2.4.1.38) forms a subfamily of the glycosyltransferase superfamily, which includes at least seven members, Gal-T1 to Gal-T7, and catalyzes the transfer of galactose (Gal) from UDP-Gal to different sugar receptors. The shared motif generated by galactosyltransferases at the terminal GlcNAc residue is the lactosamine sequence Galβ4GlcNAc-R (LacNAc or LN), which is subsequently modified in various ways by adding other sugars and sulfate groups. The most common and important sugar structure of membrane glycoconjugates is poly-N-acetyllactosamine (poly-LN), which is linked to proteins (or lipids) and plays a crucial role in cell communication, adhesion, and signal transduction, and is an important molecule in the regulation of immune responses.
[0003] Another common terminal motif present in vertebrate and invertebrate glycoconjugates is the GalNAcβ4GlcNAc-R (LaCdiNAc or LDN) sequence. The LDN motif is present in mammalian pituitary glycoprotein hormones, where the terminal GalNAc residue is 4-O-sulfated and serves as a recognition marker for endothelial cell Man / S4GGnM receptor clearance. However, non-pituitary mammalian glycoproteins also contain the LDN determinant. Furthermore, LDN and LDN sequence modifications are common antigenic determinants in many parasitic nematodes and trematodes. LDN biosynthesis involves the transfer of GalNAc to the terminal GlcNAc, a process performed by highly specific GalNAc transferases. For example, Miller et al. reported in J. Biol. Chem. 2008, 283, p. 1985 (included in this article by reference) that two closely related β1,4-N-acetylgalactosamine transferases—β4GalNAc-T3 and β4GalNAc-T4—cause protein-specific addition of β1,4-linked GalNAc to Asn-linked oligosaccharides on many glycoproteins, including the glycoprotein luteinizing hormone (LH) and carbonic anhydrase-6 (CA6) .
[0004] β-(1,4)-acetylgalactosamine transferase (β-(1,4)-GalNAcT) has been identified in a range of organisms, including humans, *Caenorhabditis elegans* (Kawar et al., J. Biol. Chem. 2002, 277, 34924, incorporated herein by reference), *Drosophila melanogaster* (Hoskins et al., Science 2007, 316, 1625, incorporated herein by reference), and *Trichoplusia ni* (Vadaie et al., J. Biol. Chem. 2004, 279, 33501, incorporated herein by reference).
[0005] Finally, besides GalT and GalNAcT, which are involved in N-glycoprotein modification, an unrelated class of enzymes called UDP-N-acetylgalactosamine:polypeptide N-acetylgalactosamine transferases (also known as ppGalNAcT) are responsible for the biosynthesis of mucin-type linkages (GalNAc--1-O-SeR / ThR). These enzymes transfer GalNAc from the sugar donor UDP-GalNAc to serine and threonine residues, forming the α-terminal isomers typical of O-glycoproteins. Although the catalytic function of ppGalNAcT appears simple, computer analysis estimates that there are only 24 unique ppGalNAcT human genes. Because O-linked glycosylation proceeds stepwise, adding GalNAc to serine or threonine represents the first key step in mucin biosynthesis. Despite this apparent simplicity, multiple members of the ppGalNAcT family appear to be essential for the complete glycosylation of their protein substrates.
[0006] It has been demonstrated that β-1,4-galactosyltransferase 1 (β4Gal-T1), in addition to transferring its native substrate UDP-Gal, is capable of transferring a range of non-natural galactose derivatives to the receptor GlcNAc substrate. Specifically, as reported by Ramakrishnan et al., J. Biol. Chem. 2002, 23, 20833 (included hereby cited), a mutation from TyR289 residue to Leu289 in bovine β4Gal-T1 creates a cavity in the enzyme's catalytic pocket that can accommodate UDP-Gal molecules, such as 2-keto-Gal, carrying a chemical stem at C2. This mutant enzyme β4GalT (Y289L) has been used in vitro to detect the presence of O-GlcNAc residues on proteins or terminal GlcNAc moieties on cell surface glycans in normal and malignant tumor tissues through a two-step process involving the first transfer of the non-natural galactose moiety followed by the attachment of the oxime to the C-2 stem.
[0007] For example, Khidekel et al., J. Am. Chem. Soc. 2003, 125, 16162 (included in this paper by reference), disclosed the chemoselective mounting of a non-natural ketone functional group to an O-GlcNAc modified protein with β4GalT (Y289L). The ketone moiety acts as a unique marker to “label” O-GlcNAc glycosylated proteins using oxime linkage with biotin. Once biotinylated, the glycoconjugate can be readily detected by chemiluminescence using streptavidin conjugated with horseradish peroxidase (HRP).
[0008] For example, WO 2007 / 095506, WO 2008 / 029281 (both from Invitrogen Corporation), WO2014 / 065661 (SynAffix BV), and Clark et al. J. Am. Chem. Soc. 2008, 130, 11576 (all incorporated herein by reference) reported similar methods that achieved similar success using β4GalT (Y289L) and an azide acetyl variant of galactosamine.
[0009] Recently, the mutant β4GalT (Y289L) has also been used in the preparation of antibody heavy chain glycans for site-selective radiolabeling, as reported by Zeglis et al. in Bioconj. Chem. 2013, 24, 1057 (included in this article by reference). In particular, the integration of azide-modified N-acetylgalactosamine monosaccharide (GalNAz) into antibody glycans allows for the use of [the product / method / applied] after click chemistry introduction with a suitable chelating agent. 89 Zr is used for controlled labeling.
[0010] Ramakrishnan et al. described the loss of Mn in the double mutant β4GalT (Y289L, M344H) in Biochemistry 2004, 43, 12513 (included in this paper by reference). 2+ 98% of the activity is dependent on Mg, but in Mg 2+ It exhibits 25-30% activity in the presence of Mn, including the ability to transfer C-2 modified galactose substrates. The double mutants β4GalT (Y289L, M344H) were found to be usable for in vitro galactosylation assays because 5-10 mM Mn is known to be effective. 2+ Typical requirements for cells have potential cytotoxic effects.
[0011] Mercer et al., Bioconj. Chem. 2013, 24, 144 (included in this paper by reference), described the use of Mg 2+In the presence of this enzyme, the double mutant Y289L-M344H-β4Gal-T1 transfers GalNAc and its analog sugar to the receptor GlcNAc.
[0012] Attempts to use wild-type β-(1,4)-N-acetylgalactosamine transferase (also referred to in this article as β-(1,4)-GalNAcT) for the transfer of C-2 modified GalNAc have so far achieved modest success.
[0013] Bertozzi et al., in ACS Chem. Biol. 2009, 4, 1068 (included in this paper by reference), applied bioorthogonal chemistry reporting techniques for molecular imaging of mucin-type O-glycans in live *C. elegans*. Treatment of the worms with an azide-glycosidic variant of N-acetylgalactosamine (GalNAz) enabled the in vivo incorporation of this non-native sugar. Although metabolic integration of GalNAz into the glycoprotein was observed, digestion of *C. elegans* lysates with chondroitinase ABC and peptidyl N-glycosidase F (PNGase F), followed by Staudinger linkage using a phospho-Flag tag, and subsequent Western blotting of the glycoprotein using an α-Flag antibody, indicated that most of the GalNAz residues on the glycoprotein were located in other types of glycans besides N-glycans. Furthermore, no detectable binding was observed between the azide-tagged glycoprotein and the N-glycan-specific lectin concanavalin A (ConA), consistent with the hypothesis that most labeled glycans are O-linked rather than N-linked. Based on these observations, it can be concluded that GalNAz does not metabolically integrate into N-GlcNAc-linked proteins in this organism.
[0014] Recently, Burnham-Marusich et al. reached similar conclusions in PLOS One 2012, 7, e49020 (included in this paper by reference), where they also observed a lack of signal reduction upon PNGAse treatment—indicating that GalNAz does not significantly integrate into N-glycoproteins. Burnham-Marusich et al. described a study using Cu(I)-catalyzed azide-alkyne cyclization addition reactions of terminal alkyne probes with azide-tagged glycoproteins to detect metabolically labeled glycoproteins. The results showed that most GalNAz tags integrate into glycans insensitive to pNGase F and are therefore not N-glycoproteins.
[0015] The high substrate specificity of β-(1,4)-GalNAcT for UDP-GalNAc becomes apparent from its poor recognition of UDP-GlcNAc, UDP-Glc, and UDP-Gal, with only 0.7%, 0.2%, and 1% of the transferase activity remaining for UDP-GlcNAc, UDP-Glc, and UDP-Gal, respectively, as reported by Kawar et al., J. Biol. Chem. 2002, 277, 34924 (included in this paper by reference).
[0016] In summary, it is not surprising that there are no reports of in vitro methods for modifying glycoproteins using GalNAc transferases derived from non-natural GalNAc derivatives (such as 2-keto or 2-azidoacetyl derivatives).
[0017] Taron et al., Carbohydr. Res. 2012, 362, 62 (included in this paper by reference), describe the in vivo metabolic integration of GalNAz in GPI anchors. Summary of the Invention
[0018] This invention relates to a method for modifying a glycoprotein, the method comprising the steps of: contacting the glycoprotein with a sugar derivative nucleotide Su(A)-Nuc in the presence of β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof, wherein the glycoprotein comprises a glycan containing a terminal GlcNAc moiety, wherein:
[0019] (i) Polysaccharides containing terminal GlcNAc moieties are shown in formula (1) or (2):
[0020]
[0021] in:
[0022] b is 0 or 1;
[0023] d is 0 or 1;
[0024] e is 0 or 1; and
[0025] G is a monosaccharide, or a straight-chain or branched oligosaccharide containing 2 to 20 sugar moieties; and
[0026] (ii) The sugar derivative nucleotide Su(A)-Nuc is shown in formula (3):
[0027]
[0028] in:
[0029] a is 0 or 1;
[0030] Nuc stands for nucleotide;
[0031] U is [C(R)] 1 )2] n or [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q Where n is an integer from 0 to 24; o is an integer from 0 to 12; q and p are independently 0, 1, or 2; R 1 Independently selected from H, F, Cl, Br, I and optionally substituted C1-C 24 alkyl;
[0032] T is C3-C 12 (Hetero)arylene, wherein the (hetero)arylene is optionally substituted; and
[0033] A is selected from:
[0034] (a)-N3
[0035] (b)-C(O)R 3
[0036] Where R 3 C1-C is an optional substitute 24 alkyl;
[0037] (c)-C(O)R 4
[0038] Where R 4 Hydrogen or optionally substituted C1-C 24 alkyl;
[0039] (d)-SH
[0040] (e)-SC(O)R 8
[0041] Where R 8 C1-C is an optional substitute 24 alkyl;
[0042] (f)-SC(V)OR 8
[0043] Where V is O or S, R 8 C1-C is an optional substitute 24 alkyl;
[0044] (g)-X
[0045] Where X is selected from F, Cl, Br and I;
[0046] (h)-OS(O)2R 5
[0047] Where R 5 Selected from C1-C 24 Alkyl, C6-C 24 Aryl, C7-C 24 alkylaryl and C7-C 24 Arylalkyl, wherein the alkyl, aryl, alkylaryl and arylalkyl groups are optionally substituted;
[0048] (i)R 11
[0049] Where R 11 C2-C is an optional substitute 24 alkyl;
[0050] (j)R 12
[0051] Where R 12 The terminal C2-C is optionally substituted. 24 alkenyl; and
[0052] (k)R 13
[0053] Where R 13 For optional substitution of terminal C3-C 24 Allenyl.
[0054] In another aspect, the present invention relates to β-(1,4)-N-acetylgalactosamine transferase mutants suitable for the methods of the present invention. Attached Figure Description
[0055] exist Figure 1 Examples of glycoproteins that can be modified by the methods of the present invention are shown, the glycoproteins comprising a glycan containing a terminal GlcNAc moiety.
[0056] exist Figure 2 An embodiment of a method for modifying glycoproteins is shown, wherein the glycoprotein is an antibody. In this embodiment, the sugar derivative Su(A)-Nuc is linked to the terminal GlcNAc moiety of the antibody glycan by β-(1,4)-N-acetylgalactosamine transferase to form a modified antibody.
[0057] Figure 3 Different glycoforms of antibody glycans G0, G1, G2, G0F, G1F, and G2F are shown.
[0058] Figure 4A method is shown for providing a glycoprotein containing a glycan of formula (27) by treating a mixture of glycoforms G0, G1, G2, G0F, G1F, and G2F with sialidase and galactosidase, and a method for providing a glycoprotein containing a glycan of formula (1) by treating a mixture of glycoforms G0, G1, G2, G0F, G1F, and G2F with endoglycosidase. Incubation of the glycoprotein containing the glycan of formula (27) or (1) with the azide-modified UDP-GalNAc derivative UDP-GalNAz yields azide-modified glycoprotein (33) or (32), respectively.
[0059] Figure 5 SDS-PAGE of a series of β-(1,4)-GalNAc-T (crude product after transient expression in CHO) is shown.
[0060] Figure 6 Non-reducing SDS-PAGE of a series of β-(1,4)-CeGalNAc-T mutants is shown.
[0061] Figure 7 The activity profiles of a series of different β-(1,4)-GalNAcT mutants compared to the β-(1,4)-GalT(Y289L) mutant for the transfer of UDP-F2-GalNAz to GlcNAc are shown, as determined by the R&D System Glycosyltransferase Activity Kit.
[0062] Figure 8 The activity curves of a series of different β-(1,4)-CeGalNAcT mutants Y257L, Y257M and Y257A for transferring UDP-F2-GalNAz to GlcNAc are shown, as determined by the R&D System Glycosyltransferase Activity Kit. Invention Details
[0064] definition
[0065] The verb "comprising" and its variations, as used in this specification and claims, are used in their non-limiting sense to mean that items following the word are included, but not excluding items not specifically mentioned.
[0066] Furthermore, mentioning an element by the indefinite article "a" or "an" does not preclude the possibility that there may be more than one element, unless the context clearly requires that there be one and only one of that element. Therefore, the indefinite article "a" or "an" usually means "at least one".
[0067] Unsubstituted alkyl groups have the general formula C n H 2n+1It can be straight-chain or branched. Unsubstituted alkyl groups may also contain cyclic moieties, thus having the corresponding general formula C1. n H 2n-1 Optionally, the alkyl group may be substituted with one or more substituents as further detailed herein. Examples of alkyl groups include methyl, ethyl, propyl, 2-propyl, tert-butyl, 1-hexyl, 1-dodecyl, etc.
[0068] The aryl group comprises 6 to 12 carbon atoms and can include monocyclic and bicyclic structures. Optionally, the aryl group can be substituted by one or more substituents, which are further detailed herein. Examples of aryl groups are phenyl and naphthyl.
[0069] Arylalkyl and alkylaryl groups comprise at least seven carbon atoms and may include monocyclic and bicyclic structures. Optionally, arylalkyl and alkylaryl groups may be substituted with one or more substituents as further detailed herein. Arylalkyl groups are, for example, benzyl. Alkylaryl groups are, for example, 4-tert-butylphenyl.
[0070] A heteroaryl group comprises at least two carbon atoms (i.e., at least C2) and one or more heteroatoms, N, O, P, or S. The heteroaryl group may have a monocyclic or bicyclic structure. Optionally, the heteroaryl group may be substituted with one or more substituents as further detailed herein. Examples of suitable heteroaryl groups include pyridyl, quinolinyl, pyrimidinyl, pyrazinyl, pyrazolyl, imidazoleyl, thiazolyl, pyrroleyl, furanyl, triazolyl, benzofuranyl, indolyl, purinyl, benzoxazolyl, thiophene, phosphoryl, and oxazolyl.
[0071] Heteroarylalkyl and alkylheteroaryl groups contain at least three carbon atoms (i.e., at least C3) and may include monocyclic and bicyclic structures. Optionally, the heteroaryl group may be substituted by one or more substituents as further detailed herein.
[0072] When an aryl group is represented as (hetero)aryl, this designation means that it includes both aryl and heteroaryl groups. Similarly, alkyl(hetero)aryl means that it includes both alkylaryl and alkylheteroaryl groups, and (hetero)arylalkyl means that it includes both arylalkyl and heteroarylalkyl groups. Therefore, C2-C 24 (Miscellaneous) aryl groups are interpreted as including C2-C 24 heteroaryl and C6-C 24 Aryl. Similarly, C3-C 24 Alkyl (hetero)aryl means including C7-C 24 Alkyl aryl and C3-C 24 Alkyl heteroaryl, and C3-C 24 (Hetero)arylalkyl means including C7-C 24 Arylalkyl and C3-C 24 Heteroarylalkyl.
[0073] Unless otherwise stated, alkyl, alkenyl, olefin, alkyne, (hetero)aryl, (hetero)arylalkyl, alkyl(hetero)aryl, alkylene, alkenylene, cycloalkylene, (hetero)aryl, alkyl(hetero)aryl, (hetero)arylalkylene, alkenyl, alkyne, cycloalkyl, alkoxy, alkenoxy, (hetero)aryloxy, alkyneoxy, and cycloalkoxy may be substituted by one or more substituents independently selected from the following: C1-C 12 Alkyl, C2-C 12 alkenyl, C2-C 12 alkynyl group, C3-C 12 cycloalkyl, C5-C 12 Cycloalkenyl, C8-C 12 Cycloalkynyl, C1-C 12 Alkoxy, C2-C 12 Alkenyl group, C2-C 12 Acryloxy group, C3-C 12 Cycloalkoxy, halogen, amino, oxo, and silyl groups, wherein the silyl group can be derived from formula (R 2 )3Si- represents, where R 2 Independently selected from C1-C 12 Alkyl, C2-C 12 alkenyl, C2-C 12 alkynyl group, C3-C 12 cycloalkyl, C1-C 12 Alkoxy, C2-C 12 Alkenyl group, C2-C 12 Acryloxy groups and C3-C 12 Cycloalkoxy, wherein alkyl, alkenyl, alkynyl, cycloalkyl, alkoxy, alkenyloxy, alkynyloxy and cycloalkoxy are optionally substituted, and alkyl, alkoxy, cycloalkyl and cycloalkoxy are optionally separated by one or more heteroatoms selected from O, N and S.
[0074] The alkynyl group contains a carbon-carbon triple bond. An unsubstituted alkynyl group containing one triple bond has the general formula C0. n H 2n-3 A terminal alkynyl group is an alkynyl group in which the triple bond is located at the end of the carbon chain. Optionally, the alkynyl group is substituted by one or more substituents, which are further detailed herein, and / or spaced by heteroatoms selected from oxygen, nitrogen, and sulfur. Examples of alkynyl groups include ethynyl, propynyl, butynyl, octyynyl, etc.
[0075] Cycloynyl groups are cyclic ynyl groups. Unsubstituted cyclic ynyl groups containing a triple bond have the general formula C1, C2, and C3. n H 2n-5 Optionally, the cycloynyl group is substituted with one or more substituents as further detailed herein. An example of a cycloynyl group is a cyclooctyynyl group.
[0076] A heterocyclic alkynyl group is a cyclic alkynyl group spaced by heteroatoms selected from oxygen, nitrogen, and sulfur. Optionally, the heterocyclic alkynyl group is substituted by one or more substituents, which are further described in detail herein. An example of a heterocyclic alkynyl group is an aza-heterocyclic octyrylyl group.
[0077] (Hetero)aryl includes aryl and heteroaryl. Alkyl(hetero)aryl includes alkylaryl and alkylheteroaryl. (Hetero)arylalkyl includes arylalkyl and heteroarylalkyl. (Hetero)ynyl includes ynyl and heteroynyl. (Hetero)cycloynyl includes cycloynyl and heterocycloynyl.
[0078] In this article, heterocyclic alkyne compounds are defined as compounds containing a heterocyclic alkyne group.
[0079] Several compounds disclosed in this specification and claims can be described as fused (hetero)cyclic alkynes, i.e., compounds in which the second ring structure is fused (i.e., cyclic) with a (hetero)cyclic alkyne group. For example, in fused (hetero)cyclic octyne compounds, a cycloalkyl group (e.g., cyclopropyl) or an aromatic hydrocarbon (e.g., benzene) can cyclically form a ring with the (hetero)cyclic octyne group. The triple bond of the (hetero)cyclic octyne group in a fused (hetero)cyclic octyne compound can be located at any of the three possible positions, i.e., at positions 2, 3, or 4 of the cyclooctyne moiety (according to IUPAC Nomenclature of Organic Chemistry, Rule A31.2). The description of any fused (hetero)cyclic octyne compound in this specification and claims means including all three separate regioisomers of the cyclooctyne moiety.
[0080] The general term "sugar" in this document refers to monosaccharides such as glucose (Glc), galactose (Gal), mannose (Man), and fucose (Fuc). The term "sugar derivative" in this document refers to derivatives of monosaccharides, i.e., monosaccharides containing substituents and / or functional groups. Examples of sugar derivatives include amino sugars and sugar acids, such as glucosamine (GlcNH2), galactosamine (GalNH2), N-acetylglucosamine (GlcNAc), N-acetylgalactosamine (GalNAc), sialic acid (Sia) (also known as N-acetylneuraminic acid (NeuNAc)), and N-acetylmuramic acid (MurNAc), glucuronic acid (GlCA), and iduronic acid (IdoA).
[0081] The term "nucleotide" as used in this article is used in its usual scientific sense. A nucleotide is a molecule consisting of a nucleobase, a pentose sugar (ribose or 2-deoxyribose), and one, two, or three phosphate groups. Without phosphate groups, the nucleobase and sugar form a nucleoside. Therefore, a nucleotide can also be called a nucleoside monophosphate, nucleoside diphosphate, or nucleoside triphosphate. The nucleobase can be adenine, guanine, cytosine, uracil, or thymine. Examples of nucleotides include uridine diphosphate (UDP), guanosine diphosphate (GDP), thymidine diphosphate (TDP), cytidine diphosphate (CDP), and cytidine monophosphate (CMP).
[0082] The term "protein" is used in its usual scientific sense in this article. In this article, a polypeptide containing approximately 10 or more amino acids is considered a protein. Proteins can contain naturally occurring amino acids, but also include non-natural amino acids.
[0083] The term "glycoprotein" as used herein is used in its usual scientific sense and refers to a protein containing one or more monosaccharide or oligosaccharide chains ("glycans") covalently bonded to the protein. Glycans can be attached to the hydroxyl groups of a protein (O-linked glycans), for example, to the hydroxyl groups of serine, threonine, tyrosine, hydroxylysine, or hydroxyproline; or to the amide functional groups of a protein (N-glycoproteins), such as asparagine or arginine; or to the carbon groups of a protein (C-glycoproteins), such as tryptophan. Glycoproteins can contain more than one glycan, can contain combinations of one or more monosaccharides and one or more oligosaccharides, and can contain combinations of N-linked, O-linked, and C-linked glycans. It is estimated that more than 50% of all proteins have some form of glycosylation and are therefore considered glycoproteins. Examples of glycoproteins include PSMA (prostate-specific membrane antigen), CAL (Candida antarcticis lipase), gp41, gp120, EPO (erythropoietin), antifreeze proteins, and antibodies.
[0084] The term "glycan" as used herein is used in its usual scientific sense and refers to a monosaccharide or oligosaccharide chain linked to a protein. Therefore, the term glycan refers to the carbohydrate portion of a glycoprotein. A glycan is linked to a protein via the C-1 carbon of a sugar, which may be unsubstituted (monosaccharide) or may be further substituted at one or more of its hydroxyl groups (oligosaccharide). Naturally occurring glycans typically contain 1 to approximately 10 sugar moieties. However, when a longer sugar chain is linked to a protein, that sugar chain is also considered a glycan herein.
[0085] The glycans of glycoproteins can be monosaccharides. Typically, the monosaccharide glycans of glycoproteins consist of a single N-acetylglucosamine (GlcNAc), glucose (Glc), mannose (Man), or fucose (Fuc) covalently linked to the protein.
[0086] Glycans can also be oligosaccharides. The oligosaccharide chains of glycoproteins can be linear or branched. In oligosaccharides, the sugar directly linked to the protein is called the core sugar. In oligosaccharides, the sugar not directly linked to the protein but linked to at least two other sugars is called the internal sugar. In oligosaccharides, the sugar not directly linked to the protein but linked to a single other sugar, i.e., the sugar without other sugar substituents at one or more of its other hydroxyl groups, is called the terminal sugar. To avoid ambiguity, multiple terminal sugars can exist in the oligosaccharides of glycoproteins, but only one core sugar is present.
[0087] Glycans can be O-linked, N-linked, or C-linked. In O-linked glycans, the monosaccharide or oligosaccharide is typically bonded to the O atom of the protein's amino acid via the hydroxyl group of serine (Ser) or threonine (Thr). In N-linked glycans, the monosaccharide or oligosaccharide is bonded to the protein via the N atom of the protein's amino acid, typically via the amide nitrogen in the asparagine (Asn) or arginine (Arg) side chain. In C-linked glycans, the monosaccharide or oligosaccharide is bonded to the C atom of the protein's amino acid, typically via the C atom of tryptophan (Trp).
[0088] The end of an oligosaccharide that is directly linked to a protein is called the reducing end of the glycan. The other end of the oligosaccharide is called the non-reducing end of the glycan.
[0089] For O-linked glycans, a wide variety of chains exist. Naturally occurring O-linked glycans are typically characterized by an α-O-GalNAc moiety linked by serine or threonine, further substituted with another GalNAc, galactose, GlcNAc, sialic acid, and / or fucose, preferably galactose, GlcNAc, sialic acid, and / or fucose. The hydroxylated amino acid with glycan substitutions can be part of any amino acid sequence in the protein.
[0090] For N-linked glycans, a wide variety of chains exist. Naturally occurring N-linked glycans are typically characterized by an asparagine-linked β-N-GlcNAc moiety, which is further substituted with β-GlcNAc at its 4-OH position, then with β-Man at the 4-OH position, and finally with α-Man at its 3-OH and 6-OH positions, yielding the glycan Man3GlcNAc2. The core GlcNAc moiety can be further substituted with α-Fuc at its 6-OH position. Man3GlcNAc2 is a common oligosaccharide scaffold for almost all N-linked glycoproteins and can carry a variety of other substituents, including but not limited to Man, GlcNAc, Gal, and sialic acid. The asparagine substituted with glycan on its side chain is typically part of the sequence Asn-X-Ser / Thr, where X is any amino acid except proline, and Ser / Thr is serine or threonine.
[0091] The term "antibody" as used herein is used in its usual scientific sense. An antibody is a protein produced by the immune system that recognizes and binds to a specific antigen. An antibody is an example of a glycoprotein. The term "antibody" as used herein is used in its broadest sense and specifically includes monoclonal antibodies, polyclonal antibodies, dimers, multimers, multispecific antibodies (e.g., bispecific antibodies), antibody fragments, and double-chain and single-chain antibodies. The term "antibody" as used herein also refers to human antibodies, humanized antibodies, chimeric antibodies, and antibodies that specifically bind to cancer antigens. The term "antibody" refers to whole antibodies, but also includes antibody fragments such as antibody Fab fragments, F(ab')2, Fv or Fc fragments derived from cleaved antibodies, scFv-Fc fragments, microantibodies, bispecific antibodies, or scFv. Furthermore, the term includes genetically engineered antibodies and antibody derivatives. Antibodies, antibody fragments, and genetically engineered antibodies can be obtained by methods known in the art. Suitable commercially available antibodies mainly include abciximab, rituximab, baliximab, palilizumab, infliximab, trastuzumab, alenzumab, adalimumab, tosimomab-I131, cetuximab, ibrituximab tiuxetan, omalizumab, bevacizumab, natezumab, ranibizumab, panitumumab, ikuzumab, certolizumab pegol, golimumab, cananumab, caputoxumab, ustekinumab, tocilizumab, ofamumab, deshumab, belimumab, ipilimumab, and brentuximab.
[0092] Sameness / Similarity
[0093] In the context of this invention, a protein or protein fragment is represented by an amino acid sequence.
[0094] It should be understood that each protein or protein fragment or peptide or derived peptide or polypeptide identified herein by a given sequence identification number (SEQ ID NO) is not limited to the specific sequence disclosed. "Sequence identity" as used herein is defined as the relationship between two or more amino acid (peptide or protein) sequences determined by comparing sequences. In the art, "identity" also means the degree of sequence similarity between amino acid sequences that, depending on the circumstances, can be determined by matching strings of such sequences. Unless otherwise stated herein, identity or similarity with a given SEQ ID NO means identity or similarity based on the full length of the sequence (i.e., over its entire length or as a whole).
[0095] The “similarity” between two amino acid sequences is determined by comparing the amino acid sequence of one polypeptide and its conserved amino acid substitutions with the sequence of the second polypeptide. "Identity" and "similarity" can be readily calculated by known methods, including but not limited to those described in: Computational Molecular Biology, Lesk, AM (ed.), Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW (ed.), Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG (ed.), Humana Publishing, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heine, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991; and Carillo, H. and Lipman, D., SIAM J. Applied Math., 48:1073 (1988).
[0096] The preferred method for determining identity is designed to give the maximum match between two or more sequences being tested. Methods for determining identity and similarity are encoded in publicly available computer programs. Preferred computer program methods for determining identity and similarity between two sequences include, for example, the GCG package (Devereux, J. et al., Nucleic Acids Research 12(1):387(1984)), BestFit, BLASTP, BLASTN, and FASTA (Altschul, SF et al., J. Mol. Biol. 215:403-410(1990)). The BLAST X program is publicly available from NCBI and other sources (BLAST Anual, Altschul, S. et al., NCBI NLM NIH Bethesda, MD20894; Altschul, S. et al., J. Mol. Biol. 215:403-410(1990)). The well-known Smith-Waterman algorithm can also be used to determine identity.
[0097] Preferred parameters for peptide sequence comparison include the following: Algorithm: Needleman and Wunsch, J. Mol. Biol. 48:443-453 (1970); Comparison matrix: from Hentikoff and Hentikoff's BLOSSUM 62, Proc. Natl. Acad. Sci. USA. 89:10915-10919 (1992); Gap penalty: 12; and Gap length penalty: 4. A useful program with these parameters is publicly available as "as" Sequence U program from the Genetics Computer Group, Madison, WI. The above parameters are the default parameters for amino acid comparison (and there is no penalty for terminal gaps).
[0098] Optionally, in determining the degree of amino acid similarity, those skilled in the art may also consider so-called “conservative” amino acid substitutions, which will be clear to them. Conservative amino acid substitution refers to the interchangeability of residues with similar side chains. For example, the amino acid group with aliphatic side chains is glycine, alanine, valine, leucine, and isoleucine; the amino acid group with aliphatic-hydroxy side chains is serine and threonine; the amino acid group with amide-containing side chains is asparagine and glutamine; the amino acid group with aromatic side chains is phenylalanine, tyrosine, and tryptophan; the amino acid group with basic side chains is lysine, arginine, and histidine; and the amino acid group with sulfur-containing side chains is cysteine and methionine. Preferred conserved amino acid substituents are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine. Substitutional variants of the amino acid sequences disclosed herein are those in which at least one residue in the disclosed sequence has been removed and a different residue has been inserted at its position. Preferably, the amino acid changes are conserved. The preferred conservative substitutions for each naturally occurring amino acid are as follows: Ala to Ser; Arg to Lys; Asn to Gln or His; Asp to Glu; Cys to Ser or Ala; Gln to Asn; Glu to Asp; Gly to Pro; His to Asn or Gln; Ile to Leu or Val; Leu to Ile or Val; Lys to Arg; Gln to Glu; Met to Leu or Ile; Phe to Met, Leu, or Tyr; Ser to Thr; Thr to Ser; Trp to Tyr or His; Tyr to Trp or Phe; and Val to Ile or Leu.
[0099] Methods for modifying glycoproteins
[0100] This invention relates to an in vitro method for modifying glycoproteins to obtain modified glycoproteins, the method using β-(1,4)-N-acetylgalactosamine transferase. Preferably, the method is an in vitro method. Specifically, this invention relates to a method for modifying glycoproteins, the method comprising the steps of: contacting a glycoprotein with a sugar derivative nucleotide Su(A)-Nuc in the presence of β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof, particularly in the presence of β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof, wherein the glycoprotein comprises a glycan containing a terminal GlcNAc moiety, wherein:
[0101] (i) Polysaccharides containing terminal GlcNAc moieties are shown in formula (1) or (2):
[0102]
[0103] in:
[0104] b is 0 or 1;
[0105] d is 0 or 1;
[0106] e is 0 or 1; and
[0107] G is a monosaccharide, or a straight-chain or branched oligosaccharide containing 2 to 20 sugar moieties; and
[0108] (ii) The sugar derivative nucleotide Su(A)-Nuc is shown in formula (3):
[0109]
[0110] in:
[0111] a is 0 or 1;
[0112] Nuc stands for nucleotide;
[0113] U is [C(R)] 1 )2] n or [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q Where n is an integer from 0 to 24; o is an integer from 0 to 12; q and p are independently 0, 1, or 2; R 1 Independently selected from H, F, Cl, Br, I and optionally substituted C1-C 24 Alkyl; T is C3-C 12 (Hetero)arylene, wherein the (hetero)arylene is optionally substituted; and
[0114] A is selected from:
[0115] (a)-N3
[0116] (b)-C(O)R 3
[0117] Where R 3 C1-C is an optional substitute 24 alkyl;
[0118] (c)-C(O)R 4
[0119] Where R 4 Hydrogen or optionally substituted C1-C 24 alkyl;
[0120] (d)-SH
[0121] (e)-SC(O)R 8
[0122] Where R 8 C1-C is an optional substitute 24 alkyl;
[0123] (f)-SC(V)OR 8
[0124] Where V is O or S, R 8 C1-C is an optional substitute 24 alkyl;
[0125] (g)-X
[0126] Where X is selected from F, Cl, Br and I;
[0127] (h)-OS(O)2R 5
[0128] Where R 5 Selected from C1-C 24 Alkyl, C6-C 24 Aryl, C7-C 24 alkylaryl and C7-C 24 Arylalkyl, wherein the alkyl, aryl, alkylaryl and arylalkyl groups are optionally substituted;
[0129] (i)R 11
[0130] Where R 11 C2-C is an optional substitute 24 alkyl;
[0131] (j)R 12
[0132] Where R 12 The terminal C2-C is optionally substituted. 24 alkenyl; and
[0133] (k)R 13
[0134] Where R 13 For optional substitution of terminal C3-C 24 Allenyl.
[0135] As described above, the method of the present invention for modifying glycoproteins provides modified glycoproteins. Hereinafter, a modified glycoprotein is defined as a glycoprotein comprising a glycan of formula (4) or (5):
[0136]
[0137] in:
[0138] b, d, e, and G are as defined above; and
[0139] Su(A) is a sugar derivative of formula (6):
[0140]
[0141] in:
[0142] a, U, A, and T are defined as above.
[0143] In the modified glycoproteoglycans of formulas (4) and (5), the C1 of the sugar derivative Su(A) is linked to the C4 of the GlcNAc moiety via a β-1,4-O-glycosidic bond.
[0144] The method for modifying glycoproteins may further include the step of providing a glycoprotein comprising a glycan containing a terminal GlcNAc moiety. Therefore, the present invention also relates to a method for modifying glycoproteins, comprising the following steps:
[0145] (1) Provides a glycoprotein comprising a glycan containing a terminal GlcNAc moiety, wherein the glycan containing the terminal GlcNAc moiety is as defined in formula (1) or (2) above; and
[0146] (2) In the presence of β-(1,4)-N-acetylgalactosamine transferase or its mutant, the glycoprotein is contacted with a sugar derivative nucleotide Su(A)-Nuc, wherein Su(A)-Nuc is as defined in formula (3) above.
[0147] The following describes in more detail β-(1,4)-N-acetylgalactosamine transferase, glycoproteins containing glycans with a terminal GlcNAc moiety, sugar derivative nucleotide Su(A)-Nuc, and modified glycoproteins, and preferred embodiments thereof.
[0148] enzymes
[0149] The method of the present invention comprises the step of contacting a glycoprotein containing a glycan with a terminal GlcNAc moiety with a sugar derivative nucleotide Su(A)-Nuc in the presence of β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof, particularly in the presence of β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof. In a second aspect, the present invention relates to mutants of β-(1,4)-N-acetylgalactosamine transferase as described herein, which are specifically designed for carrying out the method of the present invention. The mutants of β-(1,4)-N-acetylgalactosamine transferase are derived from naturally occurring β-(1,4)-N-acetylgalactosamine transferase. In this document, β-(1,4)-N-acetylgalactosamine transferase is also referred to as (1,4)-GalNAcT enzyme, or β-(1,4)-GalNAcT or GalNAcT. The term “β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof” refers to a glycosyltransferase that is or is derived from β-(1,4)-N-acetylgalactosamine transferase.
[0150] β-(1,4)-N-acetylgalactosamine transferase (β-(1,4)-GalNAcT) is known in the art. Typically, β-(1,4)-GalNAcT is an enzyme that catalyzes the transfer of N-acetylgalactosamine (GalNAc) from uridine diphosphate-GalNAc (UDP-GalNAc, also known as GalNAc-UDP) to the terminal GlcNAc moiety of a glycoprotein glycan, wherein the C1 of the GalNAc moiety is linked to the C4 of the GlcNAc moiety via a β-1,4-O-glycosidic bond. As described in more detail below, b is the GlcNAc moiety in the glycan of formula (1), i.e., the GlcNAc moiety in the glycan composed of fucosylated GlcNAc is also considered herein to be the terminal GlcNAc moiety.
[0151] In the method of the present invention, β-(1,4)-GalNAcT enzyme or a mutant thereof catalyzes the transfer of sugar derivative Su(A) from sugar derivative nucleotide Su(A)-Nuc to the terminal GlcNAc moiety of a glycoprotein glycan, wherein Su(A) is as shown in formula (6), Su(A)-Nuc is as shown in formula (3), and the glycan containing the terminal GlcNAc moiety is as shown in formula (1) or (2), as described above. In this method, C1 of the Su(A) moiety is linked to C4 of the GlcNAc moiety via a β-1,4-O-glycosidic bond.
[0152] Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention is or is derived from an invertebrate β-(1,4)-GalNAcT enzyme, that is, it is or is derived from β-(1,4)-GalNAcT from an invertebrate species. The β-(1,4)-GalNAcT enzyme can be or can be derived from any invertebrate β-(1,4)-GalNAcT enzyme known to those skilled in the art. Preferably, the β-(1,4)-GalNAcT enzyme is or is derived from a β-(1,4)-GalNAcT enzyme from the phylum Nematoda, preferably from the class Chromadorea or Seminentea, or from the phylum Arthropoda, preferably from the class Insecta. Preferably, the β-(1,4)-GalNAcT enzyme is derived from or is derived from *Caenorhabditis elegans*, *Caenorhabditis remanei*, *Caenorhabditis briggsae*, *Ascaris suum*, *Trichoplusiani*, *Drosophila melanogaster*, *Wuchereria bancrofti*, *Loa loa*, *Cerapachys biroi*, *Zootermopsis nevadensis*, *Camponotus floridanus*, *Crassostrea gigas*, or *Danaus plexippus*, and more preferably from *Caenorhabditis elegans*, *Ascaris suum*, *Trichoplusiani*, *Drosophila melanogaster*, *Wuchereria bancrofti*, *Loa loa*, *Cerapachys biroi*, *Zootermopsis nevadensis*, *Camponotus floridanus*, *Crassostrea gigas*, or *Danaus plexippus*. More preferably, the β-(1,4)-GalNAcT enzyme is or is derived from a β-(1,4)-GalNAcT enzyme derived from *C. elegans*, *Ascaris suis*, or *Spodoptera litura*. In other preferred embodiments, the β-(1,4)-GalNAcT enzyme is or is derived from a β-(1,4)-GalNAcT enzyme derived from *Ascaris suis*. In another preferred embodiment, the β-(1,4)-GalNAcT enzyme is or is derived from a β-(1,4)-GalNAcT enzyme derived from *Spodoptera litura*. In another preferred embodiment, the β-(1,4)-GalNAcT enzyme is or is derived from a β-(1,4)-GalNAcT enzyme derived from *C. elegans*.
[0153] In this article, *C. elegans* is also referred to as Ce, *Ascaris suis* as As, *Ophiocortis nigra* as Tn, and *Drosophila melanogaster* as Dm.
[0154] Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with the sequence selected from SEQ ID NO:2-5, i.e. SEQ ID NO:2, 3, 4, or 5.
[0155] Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention is or is derived from any naturally occurring or wild-type β-(1,4)-GalNAcT enzyme selected from the following: *C. elegans* β-(1,4)-GalNAcT (SEQ ID NO:2), named CeGalNAcT herein; *Ascaris suis* β-(1,4)-GalNAcT (SEQ ID NO:3), named TnGalNAcT herein; *D. melanogaster* β-(1,4)-GalNAcT (SEQ ID NO:4), named DmGalNAcT herein; *Caenorhabditis briggsae* β-(1,4)-GalNAcT (SEQ ID NO:5), named DmGalNAcT herein; *Caenorhabditis briggsae* β-(1,4)-GalNAcT (SEQ ID NO:15); ... NO:16), Wu Ce Nematode β-(1,4)-GalNAcT (SEQ ID NO:17), Loa filaria β-(1,4)-GalNAcT (SEQ ID NO:18), Biscuit β-(1,4)-GalNAcT (SEQ ID NO:19), Dampwood Termite β-(1,4)-GalNAcT (SEQ ID NO:20), Florida Carpenter's Ant β-(1,4)-GalNAcT (SEQ ID NO:21), Oyster β-(1,4)-GalNAcT (SEQ ID NO:22), and Monarch Butterfly β-(1,4)-GalNAcT (SEQ ID NO:23).
[0156] Further preferred are β-(1,4)-GalNAcT enzymes, which are β-(1,4)-GalNAcT enzymes derived from invertebrate species, namely, nematodes, preferably Chromadorea, preferably Rhabditida, preferably Rhabditidae, and preferably Caenorhabditis. Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with the sequences of SEQ ID NO:2, 15, and 16. Most preferably, the invertebrate species is *C. elegans*. Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with SEQ ID NO:2.
[0157] In another preferred embodiment, the β-(1,4)-GalNAcT enzyme used in the method of the present invention is a β-(1,4)-GalNAcT enzyme that is or is derived from a β-(1,4)-GalNAcT enzyme from an invertebrate species, wherein the invertebrate is a phylum Nematodea, preferably a class Secementea, preferably an order Ascaridida, preferably a family Ascarididae, preferably a genus Ascaris. More preferably, the invertebrate species is *Ascaris suis*. Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with the sequence of SEQ ID NO:3.
[0158] In another preferred embodiment, the β-(1,4)-GalNAcT enzyme used in the method of the present invention is or is derived from a β-(1,4)-GalNAcT enzyme from an invertebrate species, wherein the invertebrate species is an arthropod, preferably an insect, preferably a lepidoptera, preferably a noctuid moth, and preferably a genus *Trichoplusia*. More preferably, the invertebrate species is *Trichoplusia*. *Trichoplusia* is sometimes also called *Phytometra brassicae*, *Plusia innata*, or cabbage looper. Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with the sequence of SEQ ID NO:4.
[0159] In another preferred embodiment, the β-(1,4)-GalNAcT enzyme used in the method of the present invention is or is derived from a β-(1,4)-GalNAcT enzyme from an invertebrate species, wherein the invertebrate species is arthropoda, preferably insect, preferably diptera, preferably Drosophilidae, preferably Drosophila. More preferably, the invertebrate species is Drosophila melanogaster. Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with the sequence of SEQ ID NO:5.
[0160] In this document, "derived from" should be understood as having an amino acid sequence altered from naturally occurring β-(1,4)-GalNAcT enzymes by substitution, insertion, deletion, or addition of one or more amino acids, preferably 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, or more. In this document, β-(1,4)-GalNAcT enzymes derived from β-(1,4)-GalNAcT enzymes are also referred to as derived β-(1,4)-GalNAcT enzymes, modified β-(1,4)-GalNAcT enzymes, β-(1,4)-GalNAcT mutant enzymes, or β-(1,4)-GalNAcT mutants.
[0161] Preferably, the derived β-(1,4)-GalNAcT enzyme is modified by adding additional N- or C-terminal amino acids or chemical moieties, or by deleting N- or C-terminal amino acids, to increase stability, solubility, activity, and / or ease of purification.
[0162] Preferably, the β-(1,4)-GalNAcT enzyme is modified by deleting the N-terminal cytoplasmic domain and the transmembrane domain, which is referred to herein as the truncated enzyme.
[0163] For example, CeGalNAcT(30-383) should be understood herein as a truncated Caenorhabditis elegans β-(1,4)-GalNAcT enzyme consisting of the amino acid sequence represented by the amino acids at positions 30-383 of SEQ ID NO:2. It is known in the art that the deletion of these domains produces enzymes exhibiting increased solubility in aqueous solutions.
[0164] Similarly, AsGalNAcT(30-383) should be understood herein as a truncated β-(1,4)-GalNAcT enzyme of Ascaris suis, consisting of the amino acid sequence represented by amino acids at positions 30-383 of SEQ ID NO:3; TnGalNAcT(33-421) should be understood herein as a truncated β-(1,4)-GalNAcT enzyme of Arnebia pulcherrima, consisting of the amino acid sequence represented by amino acids at positions 33-421 of SEQ ID NO:4; and DmGalNAcT(47-403) should be understood herein as a truncated β-(1,4)-GalNAcT enzyme of Drosophila melanogaster, consisting of the amino acid sequence represented by amino acids at positions 47-403 of SEQ ID NO:5. Preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with any of the sequences in SEQ ID NO: 6-9. More preferably, the β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity with the sequence SEQ ID NO:8.
[0165] β-(1,4)-GalNAcT enzymes in which one or more amino acids have been substituted, added, or deleted are also referred to herein as mutant β-(1,4)-GalNAcT enzymes or derived β-(1,4)-GalNAcT enzymes. Preferably, β-(1,4)-GalNAcT enzymes are modified by deleting the N-terminal cytoplasmic domain and the transmembrane domain, and mutated by substituting one or more amino acids. In this document, the substitution of one or more amino acids is also referred to as mutation. Enzymes containing one or more substituted amino acids are also referred to as mutant enzymes.
[0166] In the method of the present invention, when the glycosyltransferase is derived from *C. elegans* β-(1,4)-GalNAcT enzyme or a truncated β-(1,4)-GalNAcT enzyme, the enzyme preferably further comprises one or more mutations. Preferred mutations include the substitution of isoleucine (Ile, also known as I) at position 257 by leucine (Leu, also known as L), methionine (Met, also known as M), or alanine (Ala, also known as A). Preferred mutations also include the substitution of methionine (Met, also known as M) at position 312 by histidine (His, also known as H). Therefore, when the glycosyltransferase is derived from CeGalNAcT or CeGalNAcT(30-383), the enzyme preferably comprises the I257L, I257M, or I257A mutation, and / or the M312H mutation.
[0167] It should be noted that the amino acid positions numbered in this document are based on the amino acid positions in the wild-type β-(1,4)-GalNAcT enzyme. When the β-(1,4)-GalNAcT enzyme is, for example, a truncated enzyme, the numbers used herein indicate, for example, the position of the amino acid substitution corresponding to the amino acid position number in the corresponding wild-type β-(1,4)-GalNAcT enzyme.
[0168] As an example, in wild-type CeGalNAcT (SEQ ID NO:2), isoleucine (Ile, I) is present at amino acid position 257. In CeGalNAcT (I257L), the isoleucine amino acid at position 257 is replaced by the leucine amino acid (Leu, L). As stated above, CeGalNAcT (30-383) should be understood herein as a truncated CeGalNAcT enzyme composed of the amino acid sequence represented by amino acids at positions 30-383 of SEQ ID NO:2, while CeGalNAcT (30-383) itself is represented by SEQ ID NO:6. In CeGalNAcT (30-383; I257L), the number “” in I257L indicates that it is the I amino acid at position 257 in the corresponding wild-type CeGalNAcT (i.e., the number 257 of SEQ ID NO:2 replaced by the L amino acid). The isoleucine amino acid at position 257 of SEQ ID NO:2 is represented by the isoleucine amino acid at position 228 of SEQ ID NO:6.
[0169] Preferred truncated Caenorhabditis elegans β-(1,4)-GalNAcT mutant enzymes include CeGalNAcT(30-383; I257L) (SEQ ID NO:10), CeGalNAcT(30-383; I257M) (SEQ ID NO:11), CeGalNAcT(30-383; I257A) (SEQ ID NO:12) and CeGalNAcT(30-383; M312H) (SEQ ID NO:13).
[0170] In the method of the present invention, when the glycosyltransferase is derived from the *Spodoptera litura* β-(1,4)-GalNAcT enzyme or a truncated *Spodoptera litura* β-(1,4)-GalNAcT enzyme, the enzyme preferably further comprises one or more mutations. Preferred mutations include the substitution of tryptophan (Trp, also known as W) at position 336 by phenylalanine (Phe, also known as F), histidine (His, also known as H), or valine (Val, also known as V). Thus, when the glycosyltransferase is derived from TnGalNAcT or TnGalNAcT(33-421), the enzyme preferably comprises the W336F, W336H, or W336V mutation. Preferred mutations in TnGalNAcT or TnGalNAcT(33-421) also include the substitution of glutamic acid (Glu, also known as E) at position 339 by alanine (Ala, also known as A), aspartic acid (Asp, also known as D), or serine (Ser, also known as S). Therefore, when the glycosyltransferase is derived from TnGalNAcT or TnGalNAcT(33-421), the enzyme preferably contains the E339A, E339D, or E339S mutation. Another preferred mutation of TnGalNAcT or TnGalNAcT(33-421) includes the substitution of leucine (Leu, also known as L) at position 302 with either alanine (Ala, also known as A) or glycine (Gly, also known as G). Other preferred mutations include the substitution of isoleucine (Ile, also known as I) at position 299 with either methionine (Met, also known as M), alanine (Ala, also known as A), or glycine (Gly, also known as G). Another preferred mutation includes the substitution of isoleucine (Ile, also known as I) at position 311 with methionine (Met, also known as M). The most preferred mutant of TnGalNAcT or TnGalNAcT(33-421) contains the L302A mutation.
[0171] Glycosyltransferases derived from TnGalNAcT or TnGalNAcT(33-421) may contain more than one mutation, such as the mutations at positions 336 and 339 as described above. In one embodiment, glycosyltransferases derived from TnGalNAcT or TnGalNAcT(33-421) contain mutations of W336F, W336H, or W336V and mutations of E339A, E339G, E339D, or E339S.
[0172] In a preferred embodiment of the method of the present invention, the glycosyltransferase that is or is derived from the β-(1,4)-GalNAcT enzyme is selected from the following β-(1,4)-GalNAcT enzymes of the moth *Spodoptera litura*: TnGalNAcT(33-421; W336F) (SEQ ID NO:25), TnGalNAcT(33-421; W336H) (SEQ ID NO:26), TnGalNAcT(33-421; W336V) (SEQ ID NO:27), TnGalNAcT(33-421; E339A) (SEQ ID NO:28), TnGalNAcT(33-421; E339D) (SEQ ID NO:30); TnGalNAcT(33-421; E339S) (SEQ ID NO:31); TnGalNAcT(33-421; L302A) (SEQ ID NO:25). NO: 29); TnGalNAcT (33-421; L302G) (SEQ ID NO: 35); TnGalNAcT (33-421; I299M) (SEQ ID NO: 36); TnGalNAcT (33-421; I299A) (SEQ ID NO:37); TnGalNAcT (33-421; I299G) (SEQ ID NO:38); and TnGalNAcT (33-421; I311M) (SEQ ID NO:39);
[0173] The glycosyltransferases derived from or used in the methods of this invention may also contain more than one mutation in the β-(1,4)-GalNAcT enzyme of the moth, such as TnGalNAcT(33-421; W336H, E339A) (SEQ ID NO:32), TnGalNAcT(33-421; W336H, E339D) (SEQ ID NO:33), and TnGalNAcT(33-421; W336H, E339S) (SEQ ID NO:34).
[0174] In the method of the present invention, when the glycosyltransferase is derived from β-(1,4)-GalNAcT enzyme of Ascaris suis or a truncated β-(1,4)-GalNAcT enzyme of Ascaris suis, the enzyme preferably further comprises one or more mutations. Preferred mutations include substitution of tryptophan (Trp, also known as W) with histidine (His, also known as H) at position 282, and / or substitution of glutamic acid (Glu, also known as E) with aspartic acid (Asp, also known as D) at position 285, and / or substitution of phenylalanine (Phe, also known as F) with alanine (Ala, also known as A) at position 248, and / or substitution of phenylalanine (Phe, also known as F) with glycine (Gly, also known as G) at position 248, and / or substitution of valine (Val, also known as V) with methionine (Met, also known as M) at position 245. Therefore, when the glycosyltransferase is derived from AsGalNAcT or AsGalNAcT(30-383), the enzyme preferably contains the W282H mutation and / or the E285D mutation.
[0175] In another preferred embodiment of the method of the present invention, the glycosyltransferase that is or is derived from the β-(1,4)-GalNAcT enzyme is selected from the following swine ascarid β-(1,4)-GalNAcT: AsGalNAcT(30-383; F248A) (SEQ ID NO:40), AsGalNAcT(30-383; F248G) (SEQ ID NO:41) and AsGalNAcT(30-383; V245M) (SEQ ID NO:42).
[0176] In a preferred embodiment of the method of the present invention, the glycosyltransferase that is or is derived from the β-(1,4)-GalNAcT enzyme is selected from the following swine ascarid β-(1,4)-GalNAcT: AsGalNAcT(30-383; W282H) (SEQ ID NO:46) and AsGalNAcT(30-383; E285D) (SEQ ID NO:47).
[0177] In a preferred embodiment, the β-(1,4)-GalNAcT enzyme as defined herein comprises a sequence encoding an easily purified tag. Preferably, the tag is selected from, but is not limited to, FLAG tags, His tags, HA tags, Myc tags, SUMO tags, GST tags, MBP tags, or CBP tags, more preferably a 6xHis tag. Preferably, the tag is covalently linked to the C-terminus of the β-(1,4)-GalNAcT enzyme. In another preferred embodiment, the tag is covalently linked to the N-terminus of the β-(1,4)-GalNAcT enzyme.
[0178] When the β-(1,4)-GalNAcT enzyme is derived from Caenorhabditis elegans β-(1,4)-GalNAcT, the His-labeled β-(1,4)-GalNAcT enzyme is preferably linked to the C-terminus of the β-(1,4)-GalNAcT enzyme, denoted as CeGalNAcT(30-383)-His6 (SEQ ID NO:14).
[0179] In a preferred embodiment of the method of the present invention, when the β-(1,4)-GalNAcT enzyme is or is derived from the β-(1,4)-GalNAcT of the white-spotted armyworm, the His-labeled β-(1,4)-GalNAcT enzyme is or is derived from His6-TnGalNAcT(33-421) (SEQ ID NO:49).
[0180] In another preferred embodiment of the method of the present invention, when the β-(1,4)-GalNAcT enzyme is or is derived from β-(1,4)-GalNAcT of the moth *Spodoptera litura*, the His-labeled β-(1,4)-GalNAcT enzyme is or is derived from His6-TnGalNAcT(33-421; W336F) (SEQ ID NO:50), His6-TnGalNAcT(33-421; W336H) (SEQ ID NO:51), His6-TnGalNAcT(33-421; W336V) (SEQ ID NO:52), His6-TnGalNAcT(33-421; 339A) (SEQ ID NO:53), and His6-TnGalNAcT(33-421; E339D) (SEQ ID NO:50). NO: 55), His6-TnGalNAcT (33-421; E339S) (SEQ ID NO: 56), His6-TnGalNAcT (33-421; L302A) (SEQ ID NO: 43), His6-TnGalNAcT (33-421; L302G) (SEQ ID NO: 44), His6-TnGalNAcT (33-421; I299M) (SEQ ID NO: 45), His6-TnGalNAcT (33-421; I299A) (SEQ ID NO: 48), His6-TnGalNAcT (33-421; I299G) (SEQ ID NO:54), His6-TnGalNAcT (33-421; I311M) (SEQ ID NO:60).
[0181] In another preferred embodiment of the method of the present invention, when the β-(1,4)-GalNAcT enzyme is or is derived from β-(1,4)-GalNAcT of Ascaris suis, the His-labeled β-(1,4)-GalNAcT enzyme is or is derived from His6-AsGalNAcT(30-383) (SEQ ID NO:71).
[0182] In another preferred embodiment of the method of the present invention, when the β-(1,4)-GalNAcT enzyme is or is derived from β-(1,4)-GalNAcT of Ascaris suis, the His-labeled β-(1,4)-GalNAcT enzyme is or is derived from His6-AsGalNAcT(30-383; W282H) (SEQ ID NO:72) or His6-AsGalNAcT(30-383; E285D) (SEQ ID NO:73).
[0183] In a preferred embodiment, the β-(1,4)-N-acetylgalactosamine transferase used in the method of the present invention is or is derived from a sequence selected from SEQ ID NO:2-23.
[0184] As stated above, the term "derived from" includes, for example, truncated enzymes, mutant enzymes, and enzymes containing tags that facilitate purification, modifications which are described in more detail above. The term "derived from" also includes enzymes containing combinations of modifications described in more detail above.
[0185] In another preferred embodiment, the β-(1,4)-N-acetylgalactosamine transferase used in the method of the present invention has at least 50% identity with a sequence selected from SEQ ID NO:2-23. More preferably, the β-(1,4)-N-acetylgalactosamine transferase used in the method of the present invention has a sequence selected from SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22 and SEQ ID NO:13. The sequence of NO:23 has at least 50% sequence identity, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity.
[0186] In another preferred embodiment of the method, the β-(1,4)-N-acetylgalactosamine transferase is or is derived from wild-type β-(1,4)-GalNAcT, preferably invertebrate β-(1,4)-GalNAcT. In another preferred embodiment of the method, the glycosyltransferase is or is derived from invertebrate β-(1,4)-GalNAcT. In another preferred embodiment, the glycosyltransferase is or is derived from *C. elegans* β-(1,4)-GalNAcT (CeGalNAcT), *Ascaris suis* β-(1,4)-GalNAcT (AsGalNAcT), or *Spodoptera litura* β-(1,4)-GalNAcT (TnGalNAcT). β-(1,4)-GalNAcT that is or is derived from (CeGalNAcT), (AsGalNAcT), or (TnGalNAcT) is as described in more detail above. In this embodiment, the β-(1,4)-N-acetylgalactosamine transferase used in the method is particularly preferred to be or derived from sequences selected from SEQ ID NO:2-9, i.e., sequences selected from SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, and SEQ ID NO:9. More preferably, the β-(1,4)-N-acetylgalactosamine transferase is or derived from sequences selected from SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:7, and SEQ ID NO:8, even more preferably from sequences selected from SEQ ID NO:6, SEQ ID NO:7, and SEQ ID NO:8, even more preferably from sequences selected from SEQ ID NO:7 and SEQ ID NO:8. Most preferably, the β-(1,4)-N-acetylgalactosamine transferase used in the method is or derived from SEQ ID NO:8.
[0187] In another particularly preferred embodiment, the β-(1,4)-N-acetylgalactosamine transferase used in the method has at least 50% sequence identity with a sequence selected from SEQ ID NO:2-9, i.e., selected from SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8 and SEQ ID NO:9, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity. More preferably, the β-(1,4)-N-acetylgalactosamine transferase used in the method has at least 50% sequence identity with sequences selected from SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:7 and SEQ ID NO:8, more preferably selected from SEQ ID NO:6, SEQ ID NO:7 and SEQ ID NO:8, and even more preferably selected from SEQ ID NO:7 and SEQ ID NO:8, having even more than 50% sequence identity, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or preferably 100% sequence identity. Most preferably, the β-(1,4)-N-acetylgalactosamine transferase used in the method has at least 50% sequence identity with SEQ ID NO:8, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity.
[0188] In another particularly preferred embodiment of the method of the present invention, the glycosyltransferase is or is derived from Caenorhabditis elegans β-(1,4)-GalNAcT (CeGalNAcT). In another particularly preferred embodiment, CeGalNAcT is or is derived from SEQ ID NO:2 or SEQ ID NO:6.
[0189] In another particularly preferred embodiment, CeGalNAcT used in the method is or is derived from SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13 or SEQ ID NO:14.
[0190] In another particularly preferred embodiment, the CeGalNAcT used in the method has at least 50% sequence identity with SEQ ID NO:2 or SEQ ID NO:6, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity. In another particularly preferred embodiment, the CeGalNAcT used in the method has at least 50% sequence identity with sequences SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13 or SEQ ID NO:14, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity.
[0191] In another particularly preferred embodiment of the method of the present invention, the glycosyltransferase is or is derived from *Spodoptera litura* β-(1,4)-GalNAcT (TnGalNAcT). In another preferred embodiment of the method, TnGalNAcT is or is derived from SEQ ID NO:4 or SEQ ID NO:8. In yet another preferred embodiment, the TnGalNAcT used in the method is or is derived from a sequence having at least 50% sequence identity with SEQ ID NO:4 or SEQ ID NO:8, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity. In another preferred embodiment, the TnGalNAcT used in the method is or is derived from a sequence selected from SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, and SEQ ID NO:60.In another preferred embodiment, the TnGalNAcT used in the method is selected from SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59 and SEQ ID NO:50. The sequence of NO:60 has at least 50% sequence identity, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity.
[0192] In another particularly preferred embodiment of the method of the present invention, the glycosyltransferase is or is derived from β-(1,4)-GalNAcT (AsGalNAcT) of *Ascaris suis*. In this embodiment, it is further preferred that AsGalNAcT is or is derived from SEQ ID NO:3 or SEQ ID NO:7. In another preferred embodiment, the AsGalNAcT used in the method has at least 50% sequence identity with SEQ ID NO:3 or SEQ ID NO:7, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity. In another preferred embodiment, the AsGalNAcT used in the method is or is derived from a sequence selected from SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:71, SEQ ID NO:72, and SEQ ID NO:73. In yet another preferred embodiment of the method, the AsGalNAcT used in the method has at least 50% sequence identity with a sequence selected from SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:71, SEQ ID NO:72, and SEQ ID NO:73, preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or preferably 100% sequence identity.
[0193] Preferably, the derived or wild-type β-(1,4)-GalNAcT enzyme used in the methods of the present invention has UDP-F2-GalNAz transfer activity. UDP-F2-GalNAz is a sugar derivative nucleotide of formula (18), described in more detail below. The UDP-F2-GalNAz transfer activity is preferably assessed by the method exemplified herein, i.e., by the method of the R&D Systems Glycosyltransferase Activity Kit (http: / / www.rndsystems.com / product_detail_objectname_glycotransferase_assay_principle.aspx, product number EA001).
[0194] In short, the glycosyltransferase kit determines the activity of specific glycosyltransferases via a coupling assay that detects the release of UDP during the transfer of donor sugars (from sugar-UDP nucleotides) to recipient sugars. More specifically, UDP released during sugar transfer is hydrolyzed by a specific enzyme (CD39L3 / rectum nucleoside triphosphate-bisphosphate hydrolase-3, also known as NTPDase-3), thereby producing UMP and an equivalent of phosphate (Pi). This phosphate is then detected by malachite green, which is also added to the mixture. The green color is proportional to the amount of inorganic phosphate released, and the absorbance of the color at 620 nm is measured as a direct measure of glycosyltransferase activity. In this case, the transfer of F2-GalNAz from the UDP-substrate to GlcNAc on the protein is accompanied by the release of UDP, which is hydrolyzed by CD39L3 and the resulting phosphate. Preferably, the derived or wild-type β-(1,4)-GalNAcT enzyme used in the method of the present invention has at least 30%, 33%, 50%, 75%, 100%, 150%, 200%, or more preferably at least 300% UDP-F2-GalNAz transfer activity compared to the UDP-F2-GalNAz transfer activity of the β-(1,4)-galactosyltransferase mutant bovine (Bos taurus) GalT-Y289L (SEQ ID NO:1). The transfer activity is assessed using the R&D Systems glycosyltransferase activity kit, under the conditions described in detail in Example 18, and using equal amounts of the enzyme to be tested and the β-(1,4)-galactosyltransferase mutant bovine (Bos taurus) GalT-Y289L (SEQ ID NO:1).
[0195] The mutant of β-(1,4)-N-acetylgalactosamine transferase of the second aspect of the present invention is preferably selected from SEQ ID NO:1, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:50, SEQ ID NO:51, SEQ ID One of the sequences NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:72 and SEQ ID NO:73 has at least 50% sequence identity, more preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and most preferably 100% sequence identity.
[0196] Glycoproteins and modified glycoproteins
[0197] In the method of the present invention, the glycoprotein to be modified comprises a glycan containing a terminal GlcNAc moiety, i.e., a GlcNAc moiety present at the non-reducing end of the glycan. The glycan comprises one or more sugar moieties and may be linear or branched. A glycan containing a terminal GlcNAc moiety is shown in formula (1) or (2):
[0198]
[0199] in:
[0200] b is 0 or 1;
[0201] d is 0 or 1;
[0202] e is 0 or 1; and
[0203] G is a monosaccharide, or a straight-chain or branched oligosaccharide containing 2 to 20 sugar moieties.
[0204] The glycoprotein to be modified may contain more than one glycan containing a terminal GlcNAc moiety. The glycoprotein may also contain other glycans that do not contain a terminal GlcNAc moiety.
[0205] The core GlcNAc moiety (i.e., the GlcNAc moiety linked to the protein) is optionally fucosylated (b is 0 or 1). When the core GlcNAc moiety is fucosylated, the fucose most typically links α-1,6 to C6 of the GlcNAc moiety.
[0206] It should be noted that the GlcNAc portion of the glycan with b=1 in equation (1), that is, the GlcNAc portion of the glycan composed of fucosylated GlcNAc, is also considered as the terminal GlcNAc portion in this paper.
[0207] In one embodiment, the glycan containing the terminal GlcNAc moiety consists of a single GlcNAc moiety, and the glycan is a glycan in formula (1) where b is 0. In another embodiment, the glycan consists of a fucoidylated GlcNAc moiety, and the glycan is a glycan in formula (1) where b is 1.
[0208] In another embodiment, the glycan is a glycan of formula (2), wherein the core GlcNAc (if present) is optionally fucoidylated (b is 0 or 1). In the glycan of formula (2), G represents a monosaccharide or a linear or branched oligosaccharide containing 2 to 20, preferably 2 to 12, more preferably 2 to 10, even more preferably 2, 3, 4, 5, 6, 7 or 8, most preferably 2, 3, 4, 5 or 6 sugar moieties. When G is a branched oligosaccharide, G may contain one or more terminal GlcNAc moieties. Thus, the glycan of formula (2) may contain more than one terminal GlcNAc moieties. In glycan (2), preferably e is 1 when d is 0, and d is 1 when e is 0. More preferably, d is 1 in glycan (2), even more preferably d is 1 and e is 1.
[0209] Sugar moieties that may be present in polysaccharides are known to those skilled in the art, including, for example, glucose (Glc), galactose (Gal), mannose (Man), fucose (Fuc), N-acetylglucosamine (GlcNAc), N-acetylglucosamine (GalNAc), N-acetylneuraminic acid (NeuNAc), or sialic acid and xyl.
[0210] In a preferred embodiment of the method of the present invention, the glycan containing the terminal GlcNAc moiety is as shown in Formula (1), as defined above. In another preferred embodiment, the glycan containing the terminal GlcNAc moiety is as shown in Formula (2). When the glycan containing the terminal GlcNAc moiety is as shown in Formula (2), it is further preferred that the glycan shown in Formula (2) is a glycan shown in Formula (26), (27), (28), (29) or (30):
[0211]
[0212] Where b is 0 or 1.
[0213] In a preferred embodiment of the method of the present invention, the glycan containing the terminal GlcNAc portion is a glycan as shown in formula (1), (26), (27), (28), (29), or (30), more preferably an N-linked glycan as shown in formula (1), (26), (27), (28), (29), or (30). In other preferred embodiments, the glycan containing the terminal GlcNAc portion is a glycan as shown in formula (1), (26), or (27), more preferably an N-linked glycan as shown in formula (1), (26), or (27). Most preferably, the glycan containing the terminal GlcNAc portion is a glycan as shown in formula (1) or (27), more preferably an N-linked glycan as shown in formula (1) or (27).
[0214] Glycoproteins containing glycans with terminal GlcNAc moieties are preferably as shown in formula (7) or (8):
[0215]
[0216] in:
[0217] b, d, e, and G, and their preferred implementations, are defined above;
[0218] y is an integer from 1 to 24; and
[0219] Pr is a protein.
[0220] In the method of the present invention, the glycoprotein to be modified comprises one or more glycans containing a terminal GlcNAc moiety (y is 1 to 24). Preferably, y is an integer from 1 to 12, more preferably an integer from 1 to 10. More preferably, y is 1, 2, 3, 4, 5, 6, 7 or 8, and even more preferably, y is 1, 2, 3, 4, 5 or 6. Even more preferably, y is 1, 2, 3 or 4. As described above, the glycoprotein may also comprise one or more glycans without a terminal GlcNAc moiety.
[0221] When the glycoprotein to be modified in the method of the present invention is as shown in formula (7) or (8), it is also preferred that the glycan containing the terminal GlcNAc portion is as shown in formula (1), (26), (27), (28), (29) or (30), more preferably as shown in formula (1), (26) or (27), and most preferably as shown in formula (1) or (27), preferably an N-linked glycan, as described above. Most preferably, the glycan containing the terminal GlcNAc portion is an N-linked glycan as shown in formula (1) or (27).
[0222] In a preferred embodiment of the method of the present invention, the glycoprotein containing a glycan with a terminal GlcNAc moiety is an antibody, more preferably an antibody as shown in formula (7) or (8), wherein the protein (Pr) is an antibody (Ab). Furthermore, when the glycoprotein to be modified is an antibody, the glycan containing a terminal GlcNAc moiety is preferably as shown in formula (1), (26), (27), (28), (29), or (30), more preferably as shown in formula (1), (26), or (27), and even more preferably as shown in formula (1) or (27), as described above. In this embodiment, it is further preferred that the glycan containing a terminal GlcNAc moiety is an N-linked glycan as shown in formula (1), (26), (27), (28), (29), or (30), more preferably an N-linked glycan as shown in formula (1), (26), or (27), and most preferably an N-linked glycan as shown in formula (1) or (27).
[0223] When the glycoprotein to be modified is an antibody, y is preferably 1, 2, 3, 4, 5, 6, 7 or 8, more preferably y is 1, 2, 4, 6 or 8, even more preferably y is 1, 2 or 4, and most preferably y is 1 or 2.
[0224] As defined above, the antibody can be a whole antibody or an antibody fragment. When the antibody is a whole antibody, it preferably contains one or more, more preferably one terminal non-reducing GlcNAc glycan on each heavy chain. Therefore, the whole antibody preferably contains two or more, more preferably two, four, six or eight of the glycans, more preferably two or four, and most preferably two glycans. In other words, when the antibody is a whole antibody, y is preferably 2, 4, 6 or 8, more preferably y is 2 or 4, and most preferably y is 2. When the antibody is an antibody fragment, y is preferably 1, 2, 3 or 4, more preferably y is 1 or 2.
[0225] In a preferred embodiment, the antibody is a monoclonal antibody (mAb). Preferably, the antibody is selected from IgA, IgD, IgE, IgG, and IgM antibodies. More preferably, the antibody is an IgG1, IgG2, IgG3, or IgG4 antibody, and most preferably, the antibody is an IgG1 antibody.
[0226] In the method of the present invention, a glycoprotein mixture comprising fucosylated and non-fucosylated glycans can be used as a starting glycoprotein. The mixture may, for example, comprise a glycoprotein containing one or more fucosylated (b=1) glycans (1) and / or (2) and / or one or more non-fucosylated (b=0) glycans (1) and / or (2). Therefore, removing fucose from the fucosylated glycans prior to the method of the present invention is not necessary but optional.
[0227] Glycoproteins containing glycans with terminal GlcNAc moieties are also referred to herein as “terminal non-reducing GlcNAc proteins”, and glycans containing terminal GlcNAc moieties are also referred to herein as “terminal non-reducing GlcNAc glycans”. It should be noted that the term “terminal non-reducing GlcNAc protein” includes proteins of formula (7) wherein b is 1, and the term “terminal non-reducing GlcNAc glycan” includes glycans of formula (1) wherein b is 1.
[0228] Terminally non-reducing GlcNAc proteins may comprise one or more linear and / or one or more branched terminally non-reducing GlcNAc glycans. The glycan is bonded to the protein via a C1 bond to a core sugar moiety, which is preferably a core GlcNAc moiety. Therefore, when the terminally non-reducing GlcNAc glycan bonded to the protein is a glycan as shown in Formula (2), d is preferably 1. More preferably, when the glycan is as shown in Formula (2), d is 1 and e is 1.
[0229] In a preferred embodiment, the C1 of the core sugar moiety of the terminal non-reducing GlcNAc glycan is bonded to the protein via an N-glycosidic bond, which is bonded to a nitrogen atom of an amino acid residue in the protein, more preferably to an amide nitrogen atom in the side chain of an asparagine (Asn) or arginine (Arg) amino acid. However, the C1 of the core sugar moiety of the non-reducing GlcNAc glycan can also be bonded to the protein via an O-glycosidic bond, which is bonded to an oxygen atom of an amino acid residue in the protein, more preferably to an oxygen atom in the side chain of a serine (Ser) or threonine (Thr) amino acid. In this embodiment, the core sugar moiety of the glycan is preferably an O-GlcNAc moiety or an O-GalNAc moiety, with an O-GlcNAc moiety being preferred. The C1 of the core sugar moiety of the non-reducing GlcNAc glycan can also be bonded to the protein via a C-glycosidic bond, which is bonded to a carbon atom on the protein, for example, to tryptophan (Trp). As mentioned above, glycoproteins can contain more than one glycan and can contain combinations of N-linked, O-linked and / or C-linked glycoproteins.
[0230] Terminal non-reducing GlcNAc glycans can be present at the native glycosylation sites of proteins, but can also be introduced into different sites of proteins.
[0231] When the glycoprotein is an antibody, it is preferable that the glycan containing the terminal GlcNAc portion is linked to a conserved N-glycosylation site at the asparagine position (typically at N297) in region 290-305 of the Fc fragment.
[0232] Several examples of terminally non-reducing GlcNAc proteins that can be modified in the method of the present invention are shown below. Figure 1 middle. Figure 1 (A) illustrates a glycoprotein comprising a single, optionally fucosylated GlcNAc moiety. This GlcNAc glycan may be linked to the protein, for example, via an N-glycosidic bond or an O-glycosidic bond. Figure 1 The glycoprotein in (A) can be obtained, for example, by conventional expression followed by trimming with an endoglucosidase or a combination of endoglucosidases. Figure 1 (B) illustrates a glycoprotein comprising a branched oligosaccharide glycan, wherein one of the branches contains a terminal GlcNAc moiety (this glycan is also known as GnM5). The core GlcNAc moiety may optionally be fucoidylated. Figure 1 The glycoprotein in (B) can be obtained, for example, by expressing the glycoprotein in a mammalian system in the presence of squalene or by expressing it in an engineered host organism (e.g., LeC1CHO or Pichia pastoris). Figure 1 (C) shows an antibody comprising branched oligosaccharides, wherein the core GlcNAc moiety is optionally fucosylated, and wherein all branches contain terminal GlcNAc moieties. Figure 1 The glycoproteins in (C) can be obtained, for example, by trimming a conventional mixture of antibody glycoforms (G0, G1, G2, G0F, G1F, and G2F) under the combined action of sialidase and galactosidase.
[0233] exist Figure 2 An embodiment of a method for modifying glycoproteins, wherein the glycoprotein is an antibody, is illustrated. In this embodiment, a sugar derivative Su(A) is transferred from Su(A)-Nuc to the terminal GlcNAc moiety of an antibody glycan using β-(1,4)-N-acetylgalactosamine transferase to form a modified antibody.
[0234] As described above, the method for modifying glycoproteins according to the present invention may further include the step of providing a glycoprotein comprising a glycan containing a terminal GlcNAc moiety. Therefore, the present invention also relates to a method for modifying glycoproteins, which includes the following steps:
[0235] (1) Provides a glycoprotein comprising a glycan containing a terminal GlcNAc moiety, wherein the glycan containing the terminal GlcNAc moiety is as shown in formula (1) or (2), as defined above; and
[0236] (2) In the presence of β-(1,4)-N-acetylgalactosamine transferase or a mutant thereof, the glycoprotein is contacted with a sugar derivative nucleotide Su(A)-Nuc, wherein Su(A)-Nuc is as shown in formula (3) as defined above.
[0237] When the glycoprotein to be modified, for example in the method of the present invention, comprises a polysaccharide of formula (1), in step (1) of the method, the glycoprotein to be modified can be provided by a method comprising the following steps: pruning the glycoprotein comprising oligosaccharide polysaccharides under the action of a suitable enzyme, preferably an endoglucosidase.
[0238] In a large number of glycans, the second GlcNAc residue is bonded to the GlcNac residue that is directly bonded to the glycoprotein, as well as... Figure 1 As shown in (B) and (C), the glycan in which the second GlcNAc residue is bonded to a GlcNAc residue directly bonded to the glycoprotein can be trimmed to obtain a glycoprotein containing the glycan of formula (1). The trimming occurs between the two GlcNAc residues.
[0239] A “suitable enzyme” is defined as an enzyme, and therefore the glycan to be pruned is the substrate. The preferred type of enzyme to be used in step (1) of this specific embodiment of the method of the invention depends on the specific glycan being pruned. In a preferred embodiment of this specific embodiment of the method of the invention, the enzyme in step (1) of this specific embodiment of the method is selected from endoglucosidases.
[0240] Endoglycosidases can cleave internal glycosidic bonds in glycan structures, which offers advantages for reconstructive and synthetic processes. For example, when endoglycosidases cleave at predictable sites within conserved glycan regions, they can be used to facilitate the homogenization of heteroglycan populations. In this regard, the most important class of glycosidases includes endo-β-N-acetylglucosidases (EC 3.2.1.96, commonly referred to as Endo and ENGase), a class of hydrolases that remove N-glycans from glycoproteins by hydrolyzing the β-1,4-glycosidic bonds in the N,N′-diacetylchitobiose core (a review by Wong et al., Chem. Rev. 2011, 111, 4259, incorporated herein by reference), leaving a single-core N-linked GlcNAc residue. Endo-β-N-acetylglucosidases (ENGases) have been found to be widely distributed in nature as common chemoenzyme variants, including Endo D, which is specific to oligomannose; Endo A and Endo H, which are specific to high-mannose; the Endo F isotype, ranging from high-mannose to biantennary complexes; and Endo M, which can cleave most N-glycan structures (high-mannose / complex / hybrid) except for fucoidan, with significantly higher hydrolytic activity for high-mannose oligosaccharides than for complex and hybrid oligosaccharides. These ENGases exhibit specificity for distal N-glycan structures rather than for proteins that display them, making them suitable for cleaving most N-linked glycans from glycoproteins under native conditions.
[0241] Endoglycosidases F1, F2, and F3 are best suited for the deglycosylation of native proteins. The linkage specificity of endo F1, F2, and F3 indicates a general strategy for protein deglycosylation that can remove all classes of N-linked oligosaccharides without denaturing the protein. Diantennary and triantennary structures can be immediately removed by endoglycosidases F2 and F3, respectively. Oligomannoses and hybrid structures can be removed by Endo F1.
[0242] The unique feature of Endo F3 is its sensitivity to the peptide bond state and core fucosylation state of oligosaccharides during cleavage. Endoglycosidase F3 cleaves asparagine-linked diantennary and triantennary complex oligosaccharides. It cleaves unfucosylated diantennary and triantennary structures at a slow rate, but only when peptides are linked. Core fucosylated diantennary structures are efficient substrates for Endo F3, exhibiting up to 400-fold activity. It has no activity against oligomannoses and hybrid molecules. See, for example, Tarentino et al., Glycobiology 1995, 5, 599, incorporated herein by reference.
[0243] Endo S is a glycoside endonuclease secreted by *Streptococcus pyogenes* and belongs to the glycoside hydrolase family 18, as described by Collin et al. (EMBO J. 2001, 20, 3046, incorporated herein by reference). However, compared to the aforementioned ENGases, Endo S exhibits more specificity, specifically cleaving only the conserved N-glycan in the Fc domain of human IgG (no other substrates have been identified to date), suggesting that protein-protein interactions between the enzyme and IgG provide this specificity.
[0244] Endo S49, also known as Endo S2, is described in WO 2013 / 037824 (Genovis AB) and is incorporated herein by reference. Endo S49 was isolated from Streptococcus poyogenes NZ131 and is a homolog of Endo S. Endo S49 exhibits specific endoglucosidase activity against natural IgG and cleaves more types of Fc glycans than Endo S.
[0245] In a preferred embodiment, the enzyme in step (1) of this embodiment is end-β-N-acetylglucosidase. In other preferred embodiments, end-β-N-acetylglucosidase is selected from Endo S, Endo S49, Endo F1, Endo F2, Endo F3, Endo H, Endo M and Endo A or combinations thereof.
[0246] When the glycan to be trimmed is a complex double-antennae structure, the end-β-N-acetylglucosidase is preferably selected from EndoS, Endo S49, Endo F1, Endo F2 and Endo F3 or combinations thereof.
[0247] When the glycoprotein is an antibody and the oligosaccharide to be pruned is a complex biantennary structure (i.e., ... Figure 1 (as shown in (C)), and when it is present at the conserved N-glycosylation site of IgG at N297, the end-β-N-acetylglucosidase is preferably selected from Endo S, Endo S49, Endo F1, Endo F2 and Endo F3 or a combination thereof, more preferably selected from Endo S and Endo S49 or a combination thereof.
[0248] When the glycoprotein is an antibody and the glycan to be pruned is a complex biantennary structure, and it does not have the conserved N-glycosylation site of IgG at N297, the end-β-N-acetylglucosidase is preferably selected from Endo F1, Endo F2, and Endo F3 or a combination thereof.
[0249] When the polysaccharide to be pruned is high in mannose, the end-β-N-acetylglucosidase is preferably selected from EndoH, EndoM, EndoA and EndoF1.
[0250] Therefore, when the glycoprotein to be modified in the method of the present invention contains a glycan of formula (1), in step (1) of the method, the glycoprotein to be modified is preferably provided by a method comprising the following steps: trimming the glycan of the glycoprotein containing the oligosaccharide glycan by the action of an end-β-N-acetylglucosidase to provide a glycoprotein containing a glycan of formula (1).
[0251] In other preferred embodiments, the endo-β-N-acetylglucosidase is selected from Endo S, Endo S49, Endo F1, Endo F2, Endo F3, Endo H, Endo M, Endo A, and any combination thereof. More preferably, the endo-β-N-acetylglucosidase is selected from Endo S, Endo S49, Endo H, Endo F1, Endo F2, Endo F3, and any combination thereof. Most preferably, the endo-β-N-acetylglucosidase is Endo S or Endo S49.
[0252] A method for providing a glycoprotein containing a glycan of formula (1) by treating a mixture of glycoforms G0, G1, G2, G0F, G1F, and G2F with an endoglucosidase is shown in the figure. Figure 4 middle. Figure 4 The treatment of glycoforms G0, G1, G2, G0F, G1F, and G2F with endoglycosidases is shown. Figure 2 The glycoprotein (in this case, the antibody) of the mixture of UDP-GalNAz and N-azidoacetylgalactosamine (GalNAz) is then transferred from UDP-GalNAz using the β-(1,4)-GalNAcT enzyme to produce the modified antibody of formula (32).
[0253] When the glycoprotein to be modified, for example, in the method of the present invention comprises a glycan of formula (26), the glycoprotein comprising the optional fucosylated glycan of formula (26) can be provided in a variety of ways (also referred to as "fucosylated glycan"). In this embodiment, the glycoprotein is preferably provided by hybridization of N-glycoprotein expression in the presence of squalene, as described, for example, in Kanda et al., Glycobiology 2006, 17, 104–118 (incorporated herein by reference), and subsequently treated with sialidase / galactosidase if necessary. Alternative methods include genetic engineering of the host organism. For example, LeC1CHO is a knockout CHO cell line lacking the gene expressing Mns-II. Therefore, the biosynthesis of N-glycan inevitably stops at the GnM5 stage of the glycan (which can be isolated and purified from the supernatant). Larger-scale methods require engineering host organisms that are abnormally programmed to produce hybrid or complex N-glycans, such as yeast or insect cells. However, it has been well demonstrated that these non-mammalian host cells (e.g., Glycoswitch) TM It can also be used to selectively express a single glycoform of a specific N-glycoprotein, including GnM5 glycans and M5 glycans.
[0254] Therefore, when the glycoprotein to be modified in the method of the present invention comprises a glycan of formula (26), in step (1) of the method, the glycoprotein comprising the optionally fucosylated glycan of formula (26) is preferably provided by a method comprising expressing the glycoprotein in a host organism in the presence of squalene. Preferably, the host organism is a mammalian cell line, such as HEK293 or NSO or CHO cell lines. The resulting glycoprotein can be obtained as a mixture comprising the following proteins: a glycan of formula (26) (also called GnM5), a glycan called GalGnM5, a sialylated glycan called SiaGalGnM5, and / or mixtures thereof. The non-reducing sialic acid and / or galactose moiety (if present) can be removed by treating the glycoprotein with sialic acidase (to remove the sialic acid moiety) and / or β-galactosidase (to remove the galactose moiety) to obtain the glycoprotein comprising the glycan of formula (26). Preferably, the treatment with sialidase and β-galactosidase occurs in a single step of (1b). In this embodiment, it is further preferred that in step (1) of the method, the glycoprotein to be modified is provided by a method comprising the following steps:
[0255] (1a) Expression of glycoproteins in a host organism in the presence of sorghumin; and
[0256] (1b) The obtained glycoprotein is treated with sialidase and / or β-galactosidase to obtain a glycoprotein containing the polysaccharide of formula (26).
[0257] When the glycoprotein to be modified in the method of the present invention comprises a glycan of formula (27), in step (1) of the method, the glycoprotein to be modified can be provided, for example, by a method comprising the following steps: treating a mixture of glycoforms G0, G1, G2, G0F, G1F, and G2Fy of the glycoprotein with sialidase and galactosidase. Figure 3 The image shows glycoforms of antibodies containing biantennary glycans, namely G0, G1, G2, G0F, G1F, and G2F.
[0258] Figure 4 A method is shown for providing a glycoprotein (in this case, an antibody) containing a polysaccharide of formula (27): a mixture of glycoforms G0, G1, G2, G0F, G1F, and G2F is treated with sialidase and galactosidase, and then N-azidoacetylgalactosamine (GalNAz) is transferred from UDP-GalNAz using β-(1,4)-GalNAcT enzyme to obtain a modified antibody of formula (33).
[0259] Sugar derivative nucleotide Su(A)-Nuc
[0260] In the method for modifying glycoproteins of the present invention, a glycoprotein comprising a glycan of formula (1) or (2) is contacted with a sugar derivative nucleotide Su(A)-Nuc by (mutant) β-(1,4)-acetylgalactosamine transferase. The sugar derivative nucleotide Su(A)-Nuc is shown in formula (3):
[0261]
[0262] in:
[0263] a is 0 or 1;
[0264] Nuc stands for nucleotide;
[0265] U is [C(R)] 1 )2] n or [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q Where n is an integer from 0 to 24; o is an integer from 0 to 12; p and q are independently 0, 1, or 2; R 1 Independently selected from H, F, Cl, Br, I and optionally substituted C1-C 24 alkyl;
[0266] T is C3-C 12 (Hetero)arylene, wherein the (hetero)arylene is optionally substituted; and
[0267] A is selected from:
[0268] (a)-N3
[0269] (b)-C(O)R 3
[0270] Where R 3 C1-C is an optional substitute 24 alkyl;
[0271] (c)-C(O)R 4
[0272] Where R 4 Hydrogen or optionally substituted C1-C 24 alkyl;
[0273] (d)-SH
[0274] (e)-SC(O)R 8
[0275] Where R 8 C1-C is an optional substitute 24 alkyl;
[0276] (f)-SC(V)OR 8
[0277] Where V is O or S, R 8 C1-C is an optional substitute 24 alkyl;
[0278] (g)-X
[0279] Where X is selected from F, Cl, Br and I;
[0280] (h)-OS(O)2R 5
[0281] Where R 5 Selected from C1-C 24 Alkyl, C6-C 24 Aryl, C7-C 24 alkylaryl and C7-C 24 Arylalkyl, wherein the alkyl, aryl, alkylaryl and arylalkyl groups are optionally substituted;
[0282] (i)R 11
[0283] Where R 11 C2-C is an optional substitute 24 alkyl;
[0284] (j)R 12
[0285] Where R 12 The terminal C2-C is optionally substituted. 24 alkenyl; and
[0286] (k)R 13
[0287] Where R 13 For optional substitution of terminal C3-C 24 Allenyl.
[0288] Nuc is defined herein as a nucleotide. Nuc is preferably selected from nucleoside monophosphate and nucleoside diphosphate, more preferably from uridine diphosphate (UDP), guanosine diphosphate (GDP), thymidine diphosphate (TDP), cytidine diphosphate (CDP), and cytidine monophosphate (CMP), and even more preferably from uridine diphosphate (UDP), guanosine diphosphate (GDP), and cytidine diphosphate (CDP). Most preferably, Nuc is uridine diphosphate (UDP). Therefore, in a preferred embodiment, Su(A)-Nuc(3) is Su(A)-UDP(31):
[0289]
[0290] U, a, T, and A are defined as above.
[0291] In one implementation, A is azide-N3.
[0292] In another embodiment, A is a ketone-C(O)R 3 , where R 3 It is an optional substitution of C1-C 24 Alkyl groups, preferably optionally substituted C1-C 12 Alkyl, more preferably optional substituted C1-C6 alkyl. Even more preferably, R 3 It is methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl, most preferably, R 3 It is a methyl group.
[0293] In another embodiment, A is an alkynyl group. In a preferred embodiment, the alkynyl group is a (hetero)cyclic alkynyl group, preferably a (hetero)cyclic octyrylyl group. In another preferred embodiment, the alkynyl group is -C≡CR 4 , where R 4 It is hydrogen or an optional substituted C1-C 24 Alkyl groups, preferably hydrogen-based or optionally substituted C1-C 12 Alkyl, more preferably hydrogen-containing or optionally substituted C1-C6 alkyl. Most preferably, R 4It is hydrogen, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. In this embodiment, it is further preferred that the alkynyl group is a terminal alkynyl group, i.e., R. 4 Hydrogen is preferred.
[0294] In another implementation, A is a thiol-SH group.
[0295] In another embodiment, A is a mercapto-SC(O)R 8 The precursor, in which R 8 It is an optional substitution of C1-C 24 Alkyl group. Preferably, R 8 It is an optional substitution of C1-C 12 Alkyl, more preferably R 8 It is an optional substituted C1-C6 alkyl group, or even more preferably R. 8 It is methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. Most preferably, R 8 It is a methyl group. In the method of the present invention for modifying glycoproteins, a sugar derivative nucleotide in which A is a thiol precursor can be used. In this process, the thiol precursor is converted into a thiol group.
[0296] In another implementation, A is -SC(V)OR 8 Where V is O or S, R 8 It is an optional substitution of C1-C 24 Alkyl group. In a preferred embodiment, A is -SC(O)OR 8 In another preferred embodiment, A is -SC(S)OR 8 . When A is -SC(O)OR 8 And when A is -SC(S)OR 8 At that time, R 8 The preferred option is the optional substitution of C1-C. 12 Alkyl, more preferably R 8 It is an optional substituted C1-C6 alkyl group, or even more preferably R. 8 It is methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. Most preferably, R 8 It is a methyl group.
[0297] In another embodiment, A is a halogen X. X is selected from F, Cl, Br, and I, preferably from Cl, Br, and I, more preferably from Cl and Br. Most preferably, X is Cl.
[0298] In another embodiment, A is sulfonyloxy-OS(O)2R 5 , where R 5 Selected from C1-C 24 Alkyl, C6-C 24Aryl, C7-C 24 alkylaryl and C7-C 24 Arylalkyl, alkyl, aryl, alkylaryl, and arylalkyl are optionally substituted. Preferably, R 5 It is C1-C 12 Alkyl, C6-C 12 Aryl, C7-C 12 alkylaryl or C7-C 12 Arylalkyl. More preferably, R 5 Selected from -CH3, -C2H5, C3 straight-chain or branched alkyl, C4 straight-chain or branched alkyl, C6-C 10 Aryl and C7 alkyl aryl. Even more preferably, R 5 It is methyl, ethyl, phenyl or p-tolyl. Most preferably, the sulfonyloxy group is methanesulfonate (-OS(O)2CH3), benzenesulfonate (-OS(O)2(C6H5)) or toluenesulfonate (-OS(O)2(C6H4CH3)).
[0299] In another implementation, A is R 11 , where R 11 It is an optional substitution of C2-C 24 Alkyl groups, preferably optionally substituted C2-C 12 Alkyl, more preferably optional substituted C2-C6 alkyl. Even more preferably, R 11 It is ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl, most preferably, R 11 It is an ethyl group.
[0300] In another implementation, A is R 12 , where R 12 It is the optional substituted terminal C2-C 24 Alkenyl. The term "terminal alkenyl" in this text refers to an alkenyl group in which the carbon-carbon double bond is located at the end of the alkenyl group. Therefore, the terminal C2-C... 24 The alkenyl group terminates at the C=CH2 moiety. Preferably, R 12 It is the optional substituted terminal C2-C 12 Alkenyl, more preferably a terminal C2-C6 alkenyl with optional substituted groups. More preferably, the terminal alkenyl is a straight-chain alkenyl, preferably an unsubstituted straight-chain alkenyl. Even more preferably R 12 Selected from -C(H)=CH2, -CH2-C(H)=CH2, -CH2-CH2-C(H)=CH2, -CH2-CH2-CH2-C(H)=CH2 and -CH2-CH2-CH2-CH2-C(H)=CH2. Even more preferably, R 12Selected from -C(H)=CH2, -CH2-C(H)=CH2 and -CH2-CH2-C(H)=CH2. Most preferably, R 12 It is -C(H)=CH2.
[0301] In another implementation, A is R 13 , where R 13 It is an optional substituted terminal C3-C 24 Propylene group. In this document, the term "terminal propadienyl" refers to a propadienyl group in which the C=C=C portion is located at the end of the propadienyl group. Therefore, terminal C3-C... 24 The alkenyl group terminates with a -C=C=CH2 moiety. Preferably, R 13 It is an optional substituted terminal C3-C 12 Alkenyl, more preferably a terminal C3-C6 alkenyl with optional substituted groups. More preferably, the terminal propadienyl group is a straight-chain propadienyl group, preferably an unsubstituted straight-chain propadienyl group. Even more preferably, R 13 Selected from -C(H)=C=CH2, -CH2-C(H)=C=CH2, -CH2-CH2-C(H)=C=CH2, and -CH2-CH2-CH2-C(H)=C=CH2. Even more preferably, R 13 Selected from -C(H)=C=CH2 and -CH2-C(H)=C=CH2. Most preferably, R 13 It is -C(H)=C=CH2. When A is R 13 In particular, it is preferred that in Su(A)-Nuc(3), neither U nor T exists, that is, it is particularly preferred that a is 0, when U is [C(R 1 )2] n When n is 0, then n is 0, and when U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q When , then o, p, and q are all 0.
[0302] In Su(A)-Nuc(3), T is C3-C 12 (Hetero)aryl, wherein the (hetero)aryl group is optionally substituted. In a preferred embodiment, T is absent (a is 0). In another preferred embodiment, T is present (a is 1). When a is 1, the (hetero)aryl group T in (3) is substituted with A, wherein A is as defined above.
[0303] (Miscellaneous) aryl T optionally further substituents R by one or more substituents 2 Replace, where R 2Independently selected from halogens (-F, -Cl, -Br, -I, preferably -F, -Cl, -Br), -CN, -NO2, -C(O)R 9 -C(O)OR 9 -C(O)N(R) 10 2. C1-C 12 Alkyl, C2-C 12 alkenyl, C2-C 12 alkynyl group, C3-C 12 cycloalkyl, C5-C 12 Cycloalkenyl, C8-C 12 Cycloalkynyl, C1-C 12 Alkoxy, C2-C 12 Alkenyl group, C2-C 12 Acryloxy group, C3-C 12 Cycloalkoxy, amino (preferably -N(R) 10 )2) Oxide groups and -Si(R) 7 )3 groups, wherein alkyl, alkenyl, alkynyl, cycloalkyl, cycloalkenyl, cycloalkynyl, alkoxy, alkenyloxy, alkynyloxy, and cycloalkoxy are optionally spaced by one or more heteroatoms selected from O, N, and S, wherein R 7 Independently selected from C1-C 12 Alkyl, C2-C 12 alkenyl, C2-C 12 alkynyl group, C3-C 12 cycloalkyl, C1-C 12 Alkoxy, C2-C 12 Alkenyl group, C2-C 12 Acryloxy groups and C3-C 12 Cycloalkoxy, wherein alkyl, alkenyl, alkynyl, cycloalkyl, alkoxy, alkenyloxy, alkynyloxy, and cycloalkoxy are optionally substituted, wherein R 9 It is C1-C 12 Alkyl, wherein R 10 Independently selected from hydrogen and C1-C 12 Alkyl group. Preferably, R 9 It is a C1-C6 alkyl group, or more preferably a C1-C4 alkyl group, with methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl being the most preferred. Preferably, R... 10 It is hydrogen or C1-C6 alkyl, more preferably hydrogen or C1-C4 alkyl, and most preferably R. 10 It is hydrogen, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl.
[0304] When R 2 It is -Si(R) 7 When )3 groups are used, R is preferred. 7 Independently, it is C1-C 12Alkyl, more preferably independently C1-C6 alkyl, even more preferably independently C1-C4 alkyl, most preferably R 7 It is independently methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl.
[0305] Preferably, R 2 (When present) Independently selected from -F, -Cl, -Br, -I, -CN, -NO2, -C(O)R 9 -C(O)OR 9 -C(O)N(R) 10 2. C1-C 12 Alkyl, C1-C 12 Alkoxy, amino (-N(R) 10 )2) Oxide groups and -Si(R) 7 )3 groups, of which R 7 R 9 R 10 and R 7 R 9 R 10 The preferred implementation scheme is defined above.
[0306] More preferably, R 2 (When present) Independently selected from -F, -Cl, -Br, -CN, -NO2, -C(O)R 9 -C(O)OR 9 -C(O)N(R) 10 )2. C1-C6 alkyl, C1-C6 alkoxy, amino, oxo groups and -Si(R) 7 )3 groups, of which R 7 R 9 R 10 and R 7 R 9 R 10 The preferred implementation scheme is defined above.
[0307] Even more preferably, R 2 (When present) Independently selected from -F, -Cl, -Br, -CN, -NO2, -C(O)R 9 -C(O)OR 9 -C(O)N(R) 10 2. C1-C4 alkyl and C1-C4 alkoxy, wherein R 9 and R 10 and R 9 and R 10 The preferred implementation scheme is defined above.
[0308] Even more preferably, R 2(When present) independently selected from -F, -Cl, -Br, -CN, -NO2, methyl, methoxy, ethyl, ethoxy, n-propyl, n-propoxy, isopropyl, isopropoxy, n-butyl, n-butoxy, sec-butyl, sec-butoxy, tert-butyl, and tert-butoxy. Most preferably, R 2 (When present) Independently selected from -F, -Cl, -Br, -CN, -NO2, methyl and methoxy.
[0309] In a preferred embodiment, the (hetero)arylene in (3) is unsubstituted. In another preferred embodiment, the (hetero)arylene in (3) comprises one or more substituents R. 2 , where R 2 and R 2 The preferred implementation scheme is defined above.
[0310] The term “(hetero)aryl” in this document refers to aryl and heteroaryl. The term “(hetero)aryl” in this document refers to monocyclic (hetero)aryl and bicyclic (hetero)aryl. The (hetero)aryl in Su(A)-Nuc(3) can be any aryl or any heteroaryl.
[0311] In a preferred embodiment of the method of the present invention, the (hetero)aryl T in (3) is selected from phenylene, naphthylene, anthraceneylene, pyrrolylene, pyrroliumylene, furanylene, thiophenylene (i.e., thiofuranylene), pyrazolylene, imidazolyl, pyrimidiniumylene, and imidazolyl. oxazoliumylene, isoxazolyl, isoxazolyl, oxazoliumylene, isoxazolyl, isoxazolyl, 1,2,3-triazolyl, 1,3,4-triazolyl, diazolyl, 1-oxa-2,3-diazolyl, 1-oxa-2,4-diazolyl, 1-oxa-2,5-diazolyl, 1-oxa-3,4-diazolyl, 1-thia-2,3-diazolyl, 1-thia-2,4-diazolyl Azolyl, 1-thia-2,5-diazolyl, 1-thia-3,4-diazolyl, tetrazolyl, pyridinyl, pyridazinyl, pyrimidinyl, pyridiniumylene, pyrimidiniumylene, benzofuranyl, benzothiophenyl, benzoimidazolyl, indazoleyl, benzotriazolyl, pyrrolo[2,3-b]pyridinyl, pyrrolo[2,3] -c]pyridyl, pyrrolo[3,2-c]pyridyl, pyrrolo[3,2-b]pyridyl, imidazo[4,5-b]pyridyl, imidazo[4,5-c]pyridyl, pyrazolo[4,3-d]pyridyl, pyrazolo[4,3-c]pyridyl, pyrazolo[3,4-c]pyridyl, pyrazolo[3,4-b]pyridyl, isoindolyl, indolinyl, purineyl, dihydroindolyl (indolininylene) (group), imidazo[1,2-a]pyridyl, imidazo[1,5-a]pyridyl, pyrazolo[1,5-a]pyridyl, pyrrolo[1,2-b]pyridazinyl, imidazo[1,2-c]pyrimidinyl, quinolineyl, isoquinolineyl, cinnolineyl, quinolineyl, quinoxalinyl, phthalazinyl, 1,6-naphthyl, 1,7-naphthyl, 1,8-naphthyl, 1,5-naphthyl, 2,6-naphthyl, 2 ,7-naphthylene, pyrido[3,2-d]pyrimidinylene, pyrido[4,3-d]pyrimidinylene, pyrido[3,4-d]pyrimidinylene, pyrido[2,3-d]pyrimidinylene, pyrido[2,3-b]pyrazinylene, pyrido[3,4-b]pyrazinylene, pyrimido[5,4-d]pyrimidinylene, pyrazino[2,3-b]pyrazinylene, and pyrimido[4,5-d]pyrimidinylene, all groups optionally with one or more substituents R 2 Replace, where R2 and R 2 The preferred implementation scheme is defined above.
[0312] In another preferred embodiment, the (hetero)arylene T is selected from phenylene, pyridinyl, pyridinium, pyrimidinyl, pyrimidinium, pyrazinyl, pyridadiazinyl, pyrroloyl, pyrrolounyl, furanyl, thiofuranylyl (i.e., thiofuranylene), diazonyl, quinolinyl, imidazolyl, pyrimidinium, imidazolyl, oxazolyl, and oxazolium, and all groups are optionally replaced by one or more substituents R. 2 Replace, where R 2 and R 2 The preferred implementation scheme is defined above.
[0313] Even more preferably, the (hetero)arylene T is selected from phenylene, pyridinyl, pyridiniumyl, pyrimidinyl, pyrimidiniumyl, imidazolyl, pyrimidiniumyl, imidazolyl, pyrroleyl, furanyl, and thiopheneyl, and all groups are optionally replaced by one or more substituents R. 2 Replace, where R 2 and R 2 The preferred implementation scheme is defined above.
[0314] Most preferably, the (hetero)aryl T is selected from phenylene, imidazolyl, imidazolyl, pyridinium, pyridinium, and pyridinium, and all groups are optionally replaced by one or more substituents R. 2 Replace, where R 2 and R 2 The preferred implementation scheme is defined above.
[0315] In Su(A)-Nuc(3), U is [C(R 1 )2] n or [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q Where n is an integer from 0 to 24; o is an integer from 0 to 12; p and q are independently 0, 1, or 2; R 1 Independently selected from H, F, Cl, Br, I and optionally substituted C1-C 24 Alkyl group. Preferably, U is [C(R 1 )2] n .
[0316] In a preferred embodiment of the method of the present invention, U does not exist, that is, n, p, o and q are all 0.
[0317] In another preferred embodiment of the method of the present invention, U exists, that is, n, p, o, and q are not all 0. Therefore, in this embodiment, when U is [C(R 1 )2] n When n is an integer from 1 to 24, when U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q When , o is an integer from 0 to 12 and / or p is 1 or 2 and / or q is 1 or 2. In other words, when U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q At that time, at least one of o, p, and q is not 0.
[0318] When U is [C(R) 1 )2] n When n is an integer from 0 to 24. In a preferred embodiment, n is an integer from 1 to 24, preferably from 1 to 12. In this embodiment, n is more preferably 1, 2, 3, 4, 5, 6, 7 or 8, even more preferably 1, 2, 3, 4, 5 or 6, even more preferably 1, 2, 3 or 4, even more preferably 1, 2 or 3, even more preferably 1 or 2, and most preferably n is 1. In another preferred embodiment, n is 0. Particularly preferred is n being 0 or 1.
[0319] When U is [C(R) 1 )2] n And when n is 1 or greater, R 1 Independently selected from H, F, Cl, Br, I and optionally substituted C1-C 24 Alkyl groups, preferably selected from H, F, Cl, Br, I, and optionally substituted C1-C groups. 12 Alkyl groups, more preferably selected from H, F, Cl, Br, I, and optionally substituted C1-C6 alkyl groups. Even more preferably, R 1 Independently selected from H, F, Cl, Br, I, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. Even more preferably, R 1 Independently selected from H, F, Cl and methyl, most preferably, R 1Selected independently from H and F.
[0320] When U is [C(R) 1 )2] n And when n is 1 or 2, -[C(R) in Su(A)-Nuc 1 )2] n Preferred examples of the - portion include -(CH2)-, -(CF2)-, -(CCl2)-, -(CBr2)-, -(CMe2)-, -(CH2CH2)-, -(CH2CF2)-, -(CH2CCl2)-, -(CH2CBr2)-, -(CH2CI2)-, -(CH2CMe2)-, -(CF2CF2)-, -(CCl2CCl2)-, -(CBr2CBr2)- and -(CMe2CMe2)-, more preferably -(CH2)-, -(CF2)-, -(CH2CH2)-, -(CH2CF2)- and -(CF2CF2)-.
[0321] When U is [C(R) 1 )2] n And when n is 3 or greater, -[C(R) in Su(A)-Nuc 1 )2] n -Preferred examples include-(C) n H 2n )-、-(C n F 2n )-、-(C n Cl 2n )-、-(C n Br 2n )-、-(C (n-1) H 2(n-1) CF2)-、-(C (n-1) H 2(n-1) CCl2)-、-(C (n-1) H 2(n-1) CBr2)- and -(C (n-1) H 2(n-1) CMe2)-, such as -(C3H6)-, -(C3F6)-, -(C3Cl6)-, -(C3Br6)-, -(CH2CH2CF2)-, -(CH2CH2CCl2)-, -(CH2CH2CBr2)-, and -(C4H8)-. More preferred examples include -(C n H 2n )-、-(C n F 2n )-, for example -(C3H6)-, -(C4H8)-, -(C3F6)- and -(C4F8)-.
[0322] When U is [C(R)1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q When o is an integer from 0 to 12, p and q are independently 0, 1, or 2. Preferably, o is an integer from 1 to 10; more preferably, o is 1, 2, 3, 4, 5, 6, 7, or 8; even more preferably, o is 1, 2, 3, 4, 5, or 6; even more preferably, o is 1, 2, 3, or 4; even more preferably, o is 1, 2, or 3; even more preferably, o is 1 or 2; most preferably, o is 1. In another preferred embodiment, o is 0. Particularly preferred is o being 0, 1, or 2. When o is 0, it is further preferred that when p is 0, q is 1 or 2, and when q is 0, p is 1 or 2.
[0323] When U is [C(R) 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q And when o and / or p and / or q are 1 or greater, R 1 Independently selected from H, F, Cl, Br, I and optionally substituted C1-C 24 Alkyl groups, preferably selected from H, F, Cl, Br, I, and optionally substituted C1-C groups. 12 Alkyl groups, more preferably selected from H, F, Cl, Br, I, and optionally substituted C1-C6 alkyl groups. Even more preferably, R 1 Independently selected from H, F, Cl, Br, I, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. Even more preferably, R 1 Independently selected from H, F, Cl, and methyl. Most preferably, R 1 For H.
[0324] When U is [C(R) 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q At that time, -[C(R) in Su(A)-Nuc 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R1 )2] q Preferred examples include -CH2-O-, -(CH2)2-O-, -O-CH2-, -O-(CH2)2-, and -CH2-O-(CH2CH2O). o -、-(CH2)2-O-(CH2CH2O) o -、-O-(CH2CH2O) o -、-O-(CH2CH2O) o -CH2-, -O-(CH2CH2O) o -(CH2)2-、-CH2-O-(CH2CH2O) o -CH2-, -CH2-O-(CH2CH2O) o -(CH2)2-, -(CH2)2-O-(CH2CH2O) o -CH2- and -(CH2)2-O-(CH2CH2O) o -(CH2)2-, wherein o is 1, 2, 3, 4, 5 or 6, preferably o is 1, 2, 3 or 4, more preferably o is 1 or 2, and most preferably o is 1.
[0325] In the method of the present invention, it is preferable that a, n, o, p, and q are not all 0. Therefore, when U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q And when o, p, and q are 0, a is preferably not 0, and / or when U is [C(R 1 )2] n Furthermore, when n is 0, a is preferably not 0. In other words, when U does not exist (i.e. when o, p, and q are 0 and n is 0), a is preferably not 0.
[0326] In a preferred embodiment of the method of the present invention, a is 0, and U is [C(R 1 )2] n n is an integer from 1 to 24. In this embodiment, the sugar derivative nucleotide Su(A)-Nuc is as shown in formula (9), as defined below, where U is [C(R 1 )2] n In this embodiment, it is further preferred that a is 0 and n is 1 to 12, more preferably a is 0 and n is 1, 2, 3, 4, 5, 6, 7 or 8, even more preferably a is 0 and n is 1, 2, 3, 4, 5 or 6, even more preferably a is 0 and n is 1, 2, 3 or 4, even more preferably a is 0 and n is 1 or 2, and most preferably a is 0 and n is 1.
[0327] [C(R 1 )2] n The preferred examples are described in more detail above.
[0328] In another preferred embodiment of the method of the present invention, a is 0, and U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q Furthermore, p, o, and q are not all 0, i.e., o is an integer from 1 to 12 and / or p is 1 or 2 and / or q is 1 or 2. In this embodiment, the sugar derivative nucleotide Su(A)-Nuc is as shown in formula (9), as defined below, where U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q In this embodiment, it is further preferred that a is 0 and o is 1 to 12, more preferably a is 0 and o is 1 to 10, even more preferably a is 0 and o is 1, 2, 1, 2, 3, 4, 5, 6, 7 or 8, even more preferably a is 0 and o is 1, 2, 3, 4, 5 or 6, even more preferably a is 0 and o is 1, 2, 3 or 4, even more preferably a is 0 and o is 1 or 2, and most preferably a is 0 and o is 1. Also in this embodiment, p and q are independently 0, 1 or 2. [C(R] 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q The preferred examples are described in more detail above.
[0329] In another preferred embodiment, a is 1, and U is [C(R 1 )2] n n is an integer from 1 to 24. In this embodiment, it is further preferred that a is 1 and n is 1 to 12, more preferably a is 1 and n is 1, 2, 3, 4, 5, 6, 7 or 8, even more preferably a is 1 and n is 1, 2, 3, 4, 5 or 6, even more preferably a is 1 and n is 1, 2, 3 or 4, even more preferably a is 1 and n is 1 or 2, and most preferably a is 1 and n is 1. [C(R] 1 )2] nThe preferred examples are described in more detail above.
[0330] In another preferred embodiment, a is 1, and U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q , o is an integer from 1 to 12, and p and q are independently 0, 1, or 2. In this embodiment, it is further preferred that a is 1, and U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q And o is 1 to 10, more preferably a is 1, U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q And o is 1, 2, 3, 4, 5, 6, 7 or 8, or even more preferably a is 1, U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q Furthermore, o can be 1, 2, 3, 4, 5, or 6, and even more preferably a can be 1, and U can be [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q Furthermore, o can be 1, 2, 3, or 4, and even more preferably a can be 1, and U can be [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q And o is 1 or 2, the optimal choice is a is 1, and U is [C(R 1 )2] p -O-[C(R 1)2C(R 1 )2O] o -[C(R 1 )2] q And o is 1. In this implementation, p and q are independently 0, 1, or 2. [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q The preferred examples are described in more detail above.
[0331] In a preferred embodiment of the method of the present invention, a is 0 and U exists, that is, a is 0 and when U is [C(R 1 )2] n When n is an integer from 1 to 24, or a is 0 and when U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q When o is an integer from 1 to 12 and / or p is 1 or 2 and / or q is 1 or 2, in other words, in this implementation, when U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q When, at least one of o, p, and q is not 0 and a is 0.
[0332] In another preferred embodiment of the method of the present invention, a is 1 and U does not exist. When U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q And when p, o, and q are all 0, or when U is [C(R 1 )2] n And when n is 0, U does not exist. In this embodiment, the sugar derivative nucleotide Su(A)-Nuc is as shown in formula (10), as defined below.
[0333] In a preferred embodiment of the method of the present invention, the sugar derivative nucleotide Su(A)-Nuc is as shown in formula (9) or (10):
[0334]
[0335] Nuc, A, T, and U, and their preferred embodiments, are as defined above.
[0336] When Su(A)-Nuc is as shown in equation (9), U may or may not exist. As mentioned above, U is [C(R 1 )2] n or [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q U does not exist when n, p, o, and q are all 0.
[0337] When Su(A)-Nuc is as shown in equation (9) and U is absent, A is preferably selected from R. 12 and R 13 R 12 and R 13 The preferred embodiments are described in more detail above. When U is absent in (9), R is particularly preferred. 12 It is -C(H)=CH2, R 13 It is –C(H)=C=CH2. Nuc is preferably UDP. In a particularly preferred embodiment, Su(A)-Nuc is as shown in equation (9), where U is absent and A is R. 13 , where R 13 For -C(H)=C=CH2. Therefore, in a particularly preferred embodiment, Su(A)-Nuc is as shown in equation (44):
[0338]
[0339] When Su(A)-Nuc is as shown in equation (9) and U exists, U is preferably [C(R 1 )2] n In this embodiment, n is preferably an integer from 1 to 24, and as described above, it is further preferred that n is from 1 to 12, more preferably that n is 1, 2, 3, 4, 5, 6, 7 or 8, even more preferably that n is 1, 2, 3, 4, 5 or 6, even more preferably that n is 1, 2, 3 or 4, even more preferably that n is 1 or 2, and most preferably that n is 1.
[0340] Similarly, when Su(A)-Nuc is as shown in equation (9) and when U is [C(R 1 )2] n At that time, in (9) -[C(R) 1)2] n Preferred examples of the - part include: when n is 1 or 2, -(CH2)-, -(CF2)-, -(CCl2)-, -(CBr2)-, -(CMe2)-, -(CH2CH2)-, -(CH2CF2)-, -(CH2CCl2)-, -(CH2CBr2)-, -(CH2CI2)-, -(CH2CMe2)-, (CF2CF2)-, -(CCl2CCl2)-, -(CBr2CBr2)- and -(CMe2CMe2)-; when n is 3 or greater, -(C n H 2n )-、-(C n F 2n )-、-(C n Cl 2n )-、-(C n Br 2n )-、-(C (n-1) H 2(n-1) CF2)-、-(C (n-1) H 2(n-1) CCl2)-、-(C (n-1) H 2(n-1) CBr2)- and -(C (n-1) H 2(n-1) CMe2)-, such as -(C3H6)-, -(C3F6)-, -(C3Cl6)-, -(C3Br6)-, -(CH2CH2CF2)-, -(CH2CH2CCl2)-, -(CH2CH2CBr2)-, -(C4H8)-, -(C4F8)-, -(C4Cl8)- and -(C4Br8)-.
[0341] In a particularly preferred embodiment, when U is [C(R 1 )2] n When n is 1 or 2, R 1 It is H or F. Therefore, in a particularly preferred embodiment, U is [C(R 1 )2] n And in (9) -[C(R) 1 )2] n - Some are -(CH2)-, -(CF2)-, -(CH2CH2)-, -(CF2CF2)- or -(CH2CF2)-.
[0342] When Su(A)-Nuc is as shown in equation (9), when U is [C(R 1 )2] n When and when U is [C(R 1 )2] p -O-[C(R 1 )2C(R1 )2O] o -[C(R 1 )2] q In all cases, A is preferably selected from -N3 and -C≡CR. 4 -SH, -SC(O)R 8 -SC(V)OR 8 -X and -OS(O)2R 5 V and R 4 R 5 R 8 The preferred embodiments thereof are as defined above. X is F, Cl, Br or I. When A is X, X is preferably Cl or Br, and most preferably X is Cl. More preferably, A is selected from -N3, -SH, -SC(O)CH3 and -X, wherein X is preferably Cl or Br, and more preferably Cl. Further preferably U is [C(R 1 )2] n .
[0343] Similarly, when Su(A)-Nuc is as shown in equation (9), when U is [C(R 1 )2] n When and when U is [C(R 1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q In all cases, Nuc is preferred as UDP. Further preferred is U as [C(R 1 )2] n .
[0344] Several particularly preferred sugar derivative nucleotides of formula (9) are shown below. Therefore, in a preferred embodiment of the method of the present invention, Su(A)-Nuc is as shown in formulas (17), (18), (19), (20), (21), or (22):
[0345]
[0346]
[0347] in:
[0348] X is F, Cl, Br, or I;
[0349] V is either O or S;
[0350] R 8 C1-C is an optional substitute 24 alkyl;
[0351] r is 0 or 1;
[0352] s is an integer from 1 to 10.
[0353] t is an integer from 1 to 10; and
[0354] u is an integer from 0 to 10.
[0355] In (19), X is selected from F, Cl, Br and I, preferably X is Cl or Br, and more preferably X is Cl.
[0356] In (20), when r is 0 and when r is 1, s is preferably 1, 2, 3, 4, 5, or 6. More preferably, when r is 0 and when r is 1, s is 1, 2, 3, or 4; even more preferably, when r is 0 and when r is 1, s is 1, 2, or 3; and even more preferably, when r is 0 and when r is 1, s is 2 or 3. In particular, when s is 2 or 3, the modified glycoprotein obtained by this embodiment of the method of the present invention yields a particularly stable maleimide conjugate via subsequent coupling with maleimide. When r is 0 and when r is 1, R is preferably 1, 2, 3, 4, 5, or 6. 8 It is an optional substitution of C1-C 12 Alkyl groups, more preferably, when r is 0 and when r is 1, R 8 All are optionally substituted C1-C6 alkyl groups, and even more preferably, when r is 0 and when r is 1, R 8 All are methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. Most preferably, when r is 0 and when r is 1, R... 8 All are methyl groups. In (20), it is particularly preferred that R is methyl when r is 0 and when r is 1. 8 All are methyl groups and s is 1, 2, 3 or 4, more particularly preferred when r is 0 and when r is 1, R 8 All are methyl groups and s is 1, 2, or 3. The optimal values are when r is 0 and when r is 1. 8 All are methyl groups and s is 1. Further preferred is r is 0.
[0357] In (21), t is preferably 1, 2, 3, 4, 5, or 6. More preferably t is 1, 2, 3, or 4, even more preferably t is 1, 2, or 3, and most preferably t is 2 or 3. In particular, when t is 2 or 3, the modified glycoprotein obtained by this embodiment of the method of the present invention yields a particularly stable maleimide conjugate via subsequent coupling with maleimide.
[0358] In (22), u is preferably 1, 2, 3, 4, 5 or 6. More preferably u is 1, 2, 3 or 4, and most preferably u is 1 or 2.
[0359] When Su(A)-Nuc is as shown in equation (9) and U is [C(R1 )2] p -O-[C(R 1 )2C(R 1 )2O] o -[C(R 1 )2] q At that time, several preferred sugar derivative nucleotides of formula (9) are shown below. Therefore, in a preferred embodiment of the method of the present invention, Su(A)-Nuc is as shown in formulas (36), (37), (38), (39), (40), (41), (42) or (43):
[0360]
[0361]
[0362] Where p and q are independently 0, 1 or 2.
[0363] When Su(A)-Nuc is as in equations (36), (37), (38), (39), (40), (41), (42) or (43), it is more preferable that p and q are 1 or 2.
[0364] In another embodiment of the method of the present invention, the sugar derivative nucleotide Su(A)-Nuc(3) is as shown in formula (45), (46) or (47):
[0365]
[0366] In this embodiment, A comprises an optionally substituted cyclopropenyl group or an optionally substituted cyclopropene group. More preferably, A comprises an optionally substituted C3-C 12 Cyclopropene or optionally substituted C3-C 12 Cyclopropene group, or even more preferably, optionally substituted C3-C8 cyclopropene group. Even more preferably, A comprises optionally substituted C4-C... 12 Cyclopropene or optionally substituted C4-C 12 Cyclopropene, or even more preferably, optionally substituted C4-C8 cyclopropene or optionally substituted C4-C8 cyclopropene.
[0367] When Su(A)-Nuc is as shown in formula (10), it is also preferred that Nuc is UDP. In (10), preferred embodiments of the (hetero)aryl T and T are described in more detail above with respect to (3). Also in (10), T is optionally replaced by one or more independently selected substituents R. 2 Replace. In (10) R 2 Its preferred implementation scheme is described in more detail above for (3).
[0368] When Su(A)-Nuc is as shown in equation (10), A is preferably selected from -N3, -C≡CR. 4 -SH, -SC(O)R 8 -SC(V)OR 8 and -OS(O)2R 5 V and R 4 R 5 R 8 Its preferred embodiments are as defined above. More preferably, A is selected from -N3, -SH and -SC(O)R. 8 .
[0369] Several particularly preferred sugar derivative nucleotides of formula (10) are shown below. Therefore, in a preferred embodiment of the method of the present invention, Su(A)-Nuc is as shown in formulas (11), (12), (13), (14), (15), (16), (34), or (35):
[0370]
[0371] in:
[0372] Nuc, A and R 2 Its preferred implementation scheme is as defined above for (3);
[0373] m is 0, 1, 2, 3, or 4; and
[0374] R 6 Selected from H and optionally substituted C1-C 24 alkyl.
[0375] In (11), m is 0, 1, 2, 3 or 4; in (12), m is 0, 1, 2 or 3; in (13), m is 0, 1, 2 or 3; in (14), m is 0, 1 or 2; in (15), m is 0, 1 or 2; in (16), m is 0 or 1.
[0376] In (11), (12), (13), (14), (15), or (16), R is preferred. 2 (When present) Independently selected from -F, -Cl, -Br, -CN, -NO2, -C(O)R 9 -C(O)OR 9 -C(O)N(R) 10 2. C1-C4 alkyl and C1-C4 alkoxy, wherein R 9 For C1-C 12 Alkyl, wherein R 10 Independently selected from hydrogen and C1-C 12 Alkyl group. Preferably, R9 It is a C1-C6 alkyl group, or more preferably a C1-C4 alkyl group, with methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl being the most preferred. Preferably, R... 10 It is hydrogen or C1-C6 alkyl, more preferably hydrogen or C1–C4 alkyl, and most preferably R. 10 It is hydrogen, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. More preferably, R 2 (When present) independently selected from -F, -Cl, -Br, CN, -NO2, methyl, methoxy, ethyl, ethoxy, n-propyl, n-propoxy, isopropyl, isopropoxy, n-butyl, n-butoxy, sec-butyl, sec-butoxy, tert-butyl, and tert-butoxy. Even more preferably, R 2 (When present) independently selected from -F, -Cl, -Br, -CN, -NO2, methyl, and methoxy. Most preferably, R 2 (When present) selected from F and Cl. In other preferred embodiments, m is 1 or 2. In another preferred embodiment, m is 0.
[0377] In (11), (12), (13), (14), (15) or (16), A is preferably selected from -N3, -C≡CR 4 -SH, -SC(O)R 8 -SC(V)OR 8 and -OS(O)2R 5 V and R 4 R 5 R 8 Its preferred embodiments are as defined above. More preferably, A is selected from -N3, -SH and -SC(O)R. 8 Nuc is further preferred as a UDP port.
[0378] In (13), R is preferred. 6 H or optionally substituted C1-C 12 Alkyl, more preferably R 6 H or optionally substituted C1-C6 alkyl, or even more preferably R 6 The most preferred compounds are H, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl. 6 It is H or methyl.
[0379] In a particularly preferred embodiment of Su(A)-Nuc(10), A is -N3, Nuc is UDP, m is 0, 1, 2, 3 or 4, and R 2 (When m is 1, 2, 3, or 4) is X, where X is F, Cl, Br, or I. In this embodiment, it is further preferred that X is F or Cl.
[0380] In a particularly preferred embodiment of the method of the present invention, Su(A)-Nuc is as shown in formula (23), (24) or (25):
[0381]
[0382] In (24) and (25), X can be the same or different. Preferably, X is the same. More preferably, X is F, Cl or Br, more preferably F or Cl. In a particularly preferred embodiment, X is F in (24) and (25). In another particularly preferred embodiment, X is Cl in (24) and (25).
[0383] The method for preparing the modified glycoprotein of the present invention is preferably carried out in a solution containing a suitable concentration of Mn. 2+ or Mg 2+ The ion is transported in a suitable buffer solution (e.g., phosphate, buffered saline (e.g., phosphate-buffered saline, tris-buffered saline), citrate, HEPES, tris, and glycine). Suitable buffers are known in the art. Preferably, the buffer is phosphate-buffered saline (PBS) or tris buffer.
[0384] The method is preferably carried out at a temperature of about 4 to about 50°C, more preferably about 10 to about 45°C, even more preferably about 15 to about 40°C, and most preferably about 20 to about 37°C.
[0385] The method is preferably carried out at a pH of about 5 to about 9, more preferably about 5.5 to about 8.5, and more preferably about 6 to about 8. Most preferably, the method is carried out at a pH of about 7 to about 8.
[0386] In one specific embodiment, when the method of the present invention is performed using a specific mutant β-(1,4)-GalNAcT, for example CeGalNAcT(M312H), the method can also be carried out in a Mg-containing environment. 2+ Ions rather than Mn 2+ The reaction is carried out in a suitable buffer solution of ions.
[0387] The method of the present invention has several advantages. First, the method uses the β-(1,4)-GalNAcT enzyme. The enzyme can be wild-type β-(1,4)-GalNAcT or a mutant thereof. Many β-(1,4)-GalNAcTs from different organisms are available in nature. Furthermore, many β-(1,4)-GalNAcTs can be readily obtained from CHO by transient expression followed by purification using a simple cation exchange column. In this way, enzymes with a purity of at least 75% are typically obtained.
[0388] Furthermore, the β-(1,4)-GalNAcT transferase activity for transferring non-natural UDP-GalNAc derivatives (such as the sugar derivative Su(A) as described in more detail above) can be higher, or even significantly higher, than the transferase activity of the β-(1,4)-galactosyltransferase GalT(Y289L) mutant known in the art.
[0389] For example, it has been determined that several β-(1,4)-GalNAcTs and their mutants exhibit significantly higher transferase activities for several sugar derivatives Su(A) in the methods of this invention. In particular, CeGalNAcT was found to exhibit 10-fold higher activity than GalT(Y289L) for transferring azide-based GalNAc derivatives to glycoproteins containing terminal GlcNAc. Figure 7 It is clearly visible.
[0390] Figure 7 A series of different β-(1,4)-GalNAcT mutants compared to the GalT(Y289L) mutant are shown, demonstrating their activity profiles for transferring F2-GalNAz from the sugar derivative Su(A) of formula (18) to GlcNAc, determined using the R&D System Glycosyltransferase Activity Kit as described in more detail above. Figure 7 It is clear that, for the same process, the transferase activities of CeGalNAcT, CeGalNAcT-His, and TnGalNAcT are significantly higher than those of GalT(Y289L).
[0391] Similarly, Figure 8 Activity profiles of a series of different CeGalNAcT mutants, namely CeGalNAcT(Y257L), CeGalNAcT(Y257M) and CeGalNAcT(Y257A), for transferring F2-GalNAz from UDP-F2-GalNAz(18) to GlcNAc are shown, as determined by the R&D System Glycosyltransferase Activity Kit. Figure 8 It is clearly shown that, for the same process, the transferase activities of mutants β-(1,4)-GalNAcT CeGalNAcT(Y257L), CeGalNAcT(Y257M), and CeGalNAcT(Y257A) are also significantly higher than those of GalT(Y289L). Example
[0392] Example 1. Selection and design of GalNAc transferase
[0393] Five specific sequences were selected for initial evaluation, specifically Uniprot registry numbers: Q9GUM2 (Caenorhabditis elegans; SEQ ID NO:2 in this paper), U1MEV9 (Ascaris suis; SEQ ID NO:3 in this paper), Q6J4T9 (Butterfly moth; SEQ ID NO:4 in this paper), Q7KN92 (Drosophila melanogaster; SEQ ID NO:5 in this paper), and Q6L9W6 (Homo sapiens).
[0394] The following peptides were constructed based on the predicted deletion of cytoplasmic and transmembrane domains. These peptides contain the predicted:
[0395] Caenorhabditis elegans (CeGalNAcT[30-383] represented by SEQ ID NO: 6)
[0396] KIPSLYENLTIGSSTLIADVDAMEAVLGNTASTSDDLLDTWNSTFSPISEVNQTSFMEDIRPILFPDNQTLQFCNQTPPHLVGPIRVFLDEPDFKTLEKIYPDTHAGGHGMPKDCVARHRVAIIVPYRDREAHLRIMLHNLHSLLAKQQLDYAIFIVEQVANQTFNRGKLMNVGYDV ASRLYPWQCFIFHDVDLLPEDDRNLYTCPIQPRHMSVAIDKFNYKLPYSAIFGGISALTKDHLKKINGFSNDFWGWGGEDDDLATRTSMAGLKVSRYPTQIARYKMIKHSTEATNPVNKCRYKIMGQTKRRWTRDGLSNLKYKLVNLELKPLYTRAVVDLLEKDCRRELRRDFPTCF
[0397] Ascaris suis (represented by SEQ ID NO: 7 AsGalNAcT[30-383])
[0398] DYSFWSPAFIISAPKTLTTLQPFSQSTSTNDLAVSALESVEFSMLDNSSILHASDNWTNDELVMRAQNENLQLCPMTPPALVGPIKVWMDAPSFAELERLYPFLEPGGHGMPTACRARHRVAIVVPYRDRESHLRTFLHNLHSLLTKQQLDYAIFVVEQTANETFNRAKLMNVGYAEAIRLYDWRCFIFHDVDLLPEDDRNLYSCPDEPRHMSVAVDKFNYKLPYGSIFGGISALTREQFEGINGFSNDYWGWGGEDDDLSTRVTLAGYKISRYPAEIARYKMIKHNSEKKNPVNRCRYKLMSATKSRWRNDGLSSLSYDLISLGRLPLYTHIKVDLLEKQSRRYLRTHGFPTC
[0399] Tobacco budworm (TnGalNAcT[33 - 421] represented by SEQ ID NO: 8)
[0400] SPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTELELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWGWGGEDDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0401] Drosophila melanogaster (DmGalNAcT[47 - 403] represented by SEQ ID NO: 9)
[0402] HKYAHIYGNASSDGAGGSEASRLPASPLALSKDRERDQELNGGPNSTIRTVLATANFTSIPQDLTRFLLGTKKFLPPRQKSTSALLANCTDPDPRDGGPITPNTTLESLDVIEAELGPLLRPGGAFEPENCNAQHHVAIVVPFRDRYAHLLLFLRNIHPFLMKQRIAYRIFIVEQTNGKPFNRAAMMNIGYLEALKLYQWDCFIFHDVDLLPLDDRNLYNCPRQPRHMSVAIDTLNFRLPYRSIFGGVSAMTREHFQAVNGFSNSFFGWGGEDDDMSNRLKHANLFISRYPVNIARYKMLKHQKEKANPKRYENLQNGMSKIEQDGINSIKYSIYSIKQFPTFTWYLAELKNSERKS
[0403] Homo sapiens (HuGalNAcT [57-998] represented by SEQ ID NO: 24)
[0404] RYGSWRELAKALASRNIPAVDPHLQFYHPQRLSLEDHDIDQGVSSNSSYLKWNKPVPWLSEFRGRANLHVFEDWCGSSIQQLRRNLHFPLYPHIRTTLRKLAVSPKWTNYGLRIFGYLHPFTDGKIQFAIAADDNAEFWLSLDDQVSGLQLLASVGKTGKEWTAPGEFGKFRSQISKPVSLSASHRYYFEVLHKQNEEGTDHVEVAWRRNDPGAKFTTIIDSLSLSLFTNETFLQMDEVGHIPQTAASHYDSSNALPRDEQPPADMLRPDPRDTLYRVPLIPKSHLRHVLPDCPYKPSYLVDGLPLQRYQGLRFVHLSFVYPNDYTRLSHMETHNKCFYQENAYYQDRFSFQEYIKIDQPEKQGLEQPGFEENLLEESQYGEVAEETPASNNQNARMLEGRQTPASTLEQDATDYRLRSLRKLLAQPREGLLAPFSKRNSTASFPGRTSHIPVQQPEKRKQKPSPEPSQDSPHSDKWPPGHPVKNLPQMRGPRPRPAGDSPRKTQWLNQVESYIAEQRRGDRMRPQAPGRGWHGEEEVVAAAGQEGQVEGEEEGEEEEEEEDMSEVFEYVPVFDPVVNWDQTFSARNLDFQALRTDWIDLSCNTSGNLLLPEQEALEVTRVFLKKLNQRSRGRYQLQRIVNVEKRQDQLRGGRYLLELELLEQGQRVVRLSEYVSARGWQGIDPAGGEEVEARNLQGLVWDPHNRRRQVLNTRAQEPKLCWPQGFSWSHRAVVHFVVPVKNQARWVQQFIKDMENLFQVTGDPHFNIVITDYSSEDMDVEMALKRSKLRSYQYVKLSGNFERSAGLQAGIDLVKDPHSIIFLCDLHIHFPAGVIDAIRKHCVEGKMAFAPMVMRLHCGATPQWPEGYWEVNGFGLLGIYKSDLDRIGGMNTKEFRDRWGGEDWELLDRILQGLDVERLSLRNFFHHFHSKRGMWSRRQMKTL
[0405] In addition, peptide variants containing an N-terminal His-tag were constructed as AsGalNAcT(30-383): (His6-AsGalNAcT(30-383) represented by SEQ ID NO: 71) and TnGalNAcT(33-421) (His6-TnGalNAcT(33-421) represented by SEQ ID NO: 49).
[0406] Example 2. Design of the I257 Caenorhabditis elegans GalNAcT mutant
[0407] Based on the sequence alignment of CeGalNAcT and GalT, by isoleucine 257 Three active site mutants were designed by mutating to leucine, methionine, or alanine (underlined).
[0408] CeGalNacT(30-383; I257L) represented by SEQ ID NO: 10
[0409] KIPSLYENLTIGSSTLIADVDAMEAVLGNTASTSDDLLDTWNSTFSPISEVNQTSFMEDIRPILFPDNQTLQFCNQTPPHLVGPIRVFLDEPDFKTLEKIYPDTHAGGHGMPK DCVARHRVAIIVPYRDREAHLRIMLHNLHSLLAKQQLDYAIFIVEQVANQTFNRGKLMNVGYDVASRLYPWQCFIFHDVDLLPEDDRNLYTCPIQPRHMSVAIDKFNYKLPYSA L FGGISALTKDHLKKINGFSNDFWGWGGEDDDLATRTSMAGLKVSRYPTQLARYKMIKHSTEATNPVNKCRYKIMGQTKRRWTRDGLSNLKYKLVNLELKPLYTRAVVDLLEKDCRRELRRDFPTCF
[0410] CeGalNAcT(30-383; I257M) represented by SEQ ID NO: 11
[0411] KIPSLYENLTIGSSTLIADVDAMEAVLGNTASTSDDLLDTWNSTFSPISEVNQTSFMEDIRPILFPDNQTLQFCNQTPPHLVGPIRVFLDEPDFKTLEKIYPDTHAGGHGMPK DCVARHRVAIIVPYRDREAHLRIMLHNLHSLLAKQQLDYAIFIVEQVANQTFNRGKLMNVGYDVASRLYPWQCFIFHDVDLLPEDDRNLYTCPIQPRHMSVAIDKFNYKLPYSA M FGGISALTKDHLKKINGFSNDFWGWGGEDDDLATRTSMAGLKVSRYPTQIARYKMIKHSTEATNPVNKCRYKIMGQTKRRWTRDGLSNLKYKLVNLELKPLYTRAVVDLLEKDCRRELRRDFPTCF
[0412] CeGalNacT(30-383; 1257A) represented by SEQ ID NO: 12
[0413] KIPSLYENLTIGSSTLIADVDAMEAVLGNTASTSDDLLDTWNSTFSPISEVNQTSFMEDIRPILFPDNQTLQFCNQTPPHLVGPIRVFLDEPDFKTLEKIYPDTHAGGHGMPK DCVARHRVAIIVPYRDREAHLRIMLHNLHSLLAKQQLDYAIFIVEQVANQTFNRGKLMNVGYDVASRLYPWQCFIFHDVTLLPEDDRNLYTCPIQPRHMSVAIDKFNYKLPYSA A FGGISALTKDHLKKINGFSNDFWGWGGEDDDLATRTSMAGLKVSRYPTQLARYKMIKHSTEATNPVNKCRYKIMGQTKRRWTRDGLSNLKYKLVNLELKPLYTRAVVDLLEKDCRRELRRDFPTCF
[0414] Example 3. Design of the GalNAcT mutant of Caenorhabditis elegans M312
[0415] By methionine 312 The CeGalNAcT mutant was designed by mutating histidine.
[0416] CeGalNacT(30-383; M312H) represented by SEQ ID NO: 13
[0417] KIPSLYENLTIGSSTLIADVDAMEAVLGNTASTSDDLLDTWNSTFSPISEVNQTSFMEDIRPILFPDNQTLQFCNQTPPHLVGPIRVFLDEPDFKTLEKIYPDTHAGGHGMPKDCVARHRVAIIVPYRDREAHLRIMLHNL HSLLAKQQLDYAIFIVEQVANQTFNRGKLMNVGYDVASRLYPWQCFIFHDVDLLPEDDRNLYTCPIQPRHMSVAIDKFNYKLPYSAIFGGISALTKDHLKKINGFSNDFWGWGGEDDDLATRTSMAGLKVSRYPTQIARYK H IKHSTEATNPVNKCRYKIMGQTKRRWTRDGLSNLKYKLVNLELKPLYTRAVVDLLEKDCRRELRRDFPTCF
[0418] Example 4. Design of CeGalNAcT with C-terminal His6-tag
[0419] Design CeGalNAcT (CeGalNAcT-His6) with a C-terminal His6-label.
[0420] CeGalNAcT(30-383)-His6, represented by SEQ ID NO: 14
[0421] KIPSLYENLTIGSSTLIADVDAMEAVLGNTASTSDDLLDTWNSTFSPISEVNQTSFMEDIRPILFPDNQTLQFCNQTPPHLVGPIRVFLDEPDFKTLEKIYPDTHAGGHGMPKDCVARHRVAIIVPYRDREAHLRIMLHNLHSLLAKQQLDYAIFIVEQVANQTFNRGKLMNVGYDVASR LYPWQCFIFHDVDLLPEDDRNLYTCPIQPRHMSVAIDKFNYKLPYSAIFGGISALTKDHLKKINGFSNDFWGWGGEDDDLATRTSMAGLKVSRYPTQIARYKMIKHSTEATNPVNKCRYKIMGQTKRRWTRDGLSNLKYKLVNLELKPLYTRAVVDLLEKDCRRELRRDFPTCFHHHHHH
[0422] Example 5. Transient expression of enzyme in CHO
[0423] The protein was transiently expressed in CHO K1 cells using Evitria (Zurich, Switzerland) at a volume of 20 mL. All GalNAcT proteins except HuGalNAcT were successfully expressed, such as... Figure 5 and Figure 6 As shown.
[0424] Figure 5 SDS-PAGE of a series of β-(1,4)-GalNAcT (crude product after transient expression in CHO) is shown. Lanes 2-6: reduced enzyme; lanes 7-11: unreduced same enzyme. Lane 1: marker, Biorad UltraPrecision Protein Standard; MW from top to bottom: 250 kDa, 150 kDa, 100 kDa, 75 kDa, 50 kDa, 37 kDa, 25 kDa, 20 kDa. Lanes 2+7: TnGalNAc. Lanes 3+8: DmGalNAcT. Lanes 4+9: AsGalNAcT. Lanes 5+10: CeGalNAcT with a C-terminal His6-tag. Lanes 6+11: huGalNAcT. The visible protein bands of 50-55 kDa are the desired GalNAcT; all variants except huGalNAcT were successfully expressed.
[0425] Figure 6A series of β-(1,4)-CeGalNAcT mutants were shown using non-reducing SDS-PAGE. Lane 1: Markers: GE rainbow molecular weight markers; MW from top to bottom: 225kDa, 150kDa, 102kDa, 76kDa, 52kDa, 38kDa, 33kDa, 24kDa, 17kDa. Lane 2: CeGalNAcT (M312H). Lane 3: CeGalNAcT (M312H). Lane 4: CeGalNAcT (I257A). Lane 5: CeGalNAcT (I257A). Lane 6: CeGalNAcT (I257L). Lane 7: CeGalNAcT (I257L). Lane 8: CeGalNAcT (1257M). Lane 9: CeGalNAcT (I257L). Lane 10: CeGalNAcT with a His6 tag. The visible protein bands at ~52 kDa are monomeric CeGalNAcT species, and the visible bands at ~102 kDa are dimer CeGalNAcT species. Protein bands at higher MW (>225 kDa) were not characterized.
[0426] Example 6. Purification protocol for non-His-labeled GalNAcT protein
[0427] The purification protocol is based on cation exchange on an SP column (GE Healthcare), followed by size exclusion.
[0428] In a typical purification experiment, the supernatant from CHO containing expressed GalNAcT is dialyzed against 20 mM Tris buffer (pH 7.5). The supernatant (typically 25 mL) is filtered through a 0.45 μM filter and then purified using a cation exchange column (SP column, 5 mL, GE Healthcare), equilibrated with 20 mM Tris buffer (pH 7.5) before use. Purification is performed on an AKTA Prime chromatography system equipped with an external fraction collector. The sample is loaded from system pump A. Unbound protein is eluted from the column by washing the column with 10 column volumes (CV) of 20 mM Tris buffer (pH 7.5). The retained protein is eluted with elution buffer (20 mM Tris, 1 NaCl, pH 7.5; 10 mL). The collected fractions are analyzed by SDS-PAGE on a polyacrylamide gel (12%), and the fractions containing the target protein are combined and concentrated to a volume of 0.5 mL using spin filtration. Next, the protein was purified using an AKTA purifier system (UNICORN v6.3) on a preparative Superdex size exclusion column. This purification step resulted in the identification and separation of the dimer and the monomer fraction of the target protein. Both fractions were analyzed by SDS-PAGE and stored at -80°C before further use.
[0429] Example 7. Purification of CeGalNAcT-His6
[0430] In a typical purification assay, the CHO supernatant was filtered through a 0.45 μm pore size filter and applied to a Ni-NTA column (GE Healthcare, 5 mL) equilibrated with buffer A (20 mM Tris buffer, 20 mM imidazole, 500 mM NaCl, pH 7.5) before use. Before filtration, imidazole was added to the CHO supernatant to a final concentration of 20 mM to minimize non-specific binding to the column. The column was first washed with buffer A (50 mL). Retained proteins were eluted with buffer B (20 mM Tris, 500 mM NaCl, 250 mM imidazole, pH 7.5, 10 mL). Fractions were analyzed by SDS-PAGE on a polyacrylamide gel (12%), and fractions containing purified target proteins were pooled and the buffer was exchanged for 20 mM Tris (pH 7.5) by overnight dialysis at 4 °C. The purified protein was stored at -80 °C before further use. Note: Additional SEC purification (as described above) is required to identify the monomer and dimer CeGalNAcT-His6 species.
[0431] Example 8. Synthesis of ethyl 2-azido-2,2-difluoroacetate
[0432] Sodium azide (365 mg, 5.62 mmol) was added to an anhydrous DMSO (5 mL) solution of ethyl 2-bromo-2,2-difluoroacetate (950 mg, 4.68 mmol). After stirring overnight at room temperature, the reaction mixture was poured into water (150 mL). The layers were separated, and dichloromethane was added to the organic layer, which was then dried over sodium sulfate. After filtration, the solvent was removed under reduced pressure (300 mbar) at 35 °C to give the crude product ethyl 2-azido-2,2-difluoroacetate (250 mg, 1.51 mmol, 32%).
[0433] 1 H-NMR (300MHz, CDCl3): δ4.41 (q, J=7.2hz, 2H), 1.38 (t, J=6.9Hz, 3H).
[0434] Example 9. Synthesis of α-2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose-1-phosphate
[0435] The process for preparing 2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose 1-phosphate from D-galactosamine is described in Linhardt et al., J. Org. Chem. 2012, 77, 1449-1456 (included in this paper by reference).
[0436] 1 H-NMR (300MHz, CD3OD): δ5.69 (dd, J=7.2, 3.3Hz, 1H), 5.43-5.42 (m, 1H), 5.35 (dd, J=11.1, 3.3Hz, 1H), 4.53 (t, J=7. 2Hz,1H),4.21-4.13(m,1H),4.07-4.00(m,1H),3.82(dt,J=10.8,2.7Hz,1H),2.12(s,3H),2.00(s,3H),1.99(s,3H).
[0437] C 12 H 17 N3O 11 P(MH + The calculated LRMS(ESI-) value is 410.06, and the measured value is 410.00.
[0438] Example 10. Synthesis of α-2-amino-3,4,6-tri-O-acetyl-D-galactose-1-phosphate
[0439] Pd / C (20 mg) was added to a MeOH (3 mL) solution of α-2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose (105 mg, 0.255 mmol). The reaction was stirred for 2 hours under a hydrogen atmosphere and filtered through diatomaceous earth. The filter was rinsed with MeOH (10 mL) and the filtrate was concentrated under vacuum to give free amine (94 mg, 0.244 mmol, 96%).
[0440] 1 H-NMR(300MHz,D2O): δ5.87-5.76(m,1H),5.44(br s,1H),5.30-5.20(m,1H),4.55(t,J=6.3Hz,1H),4.28-4.00(m,3H),2.11(s,3H),2.03(s,3H),2.00(s,3H).
[0441] C 12 H 19 NO 11 P(MH + The calculated LRMS(ESI-) value is 384.07, and the measured value is 384.10.
[0442] Example 11. Synthesis of α-(2'-azido-2',2'-difluoroacetamido)-3,4,6-tri-O-acetyl-D-galactose-1-phosphate
[0443] Anhydrous DMF (3 mL) solution of α-2-amino-3,4,6-tri-O-acetyl-D-galactose-1-phosphate (94 mg, 0.244 mmol) was mixed with ethyl difluoroazidoacetamide (48 mg, 0.293 mmol) and Et3N (68 μL, 0.488 mmol). The reaction was stirred for 6 hours and then concentrated under vacuum to give the crude product. Rapid chromatography (100:0-50:50 EtOAC:MeOH) yielded α-(2'-azido-2',2'-difluoroacetamido)-3,4,6-tri-O-acetyl-D-galactose-1-phosphate (63 mg, 0.125 mmol, 51%).
[0444] Example 12. Synthesis of UDP-α-(2'-azido-2',2'-difluoroacetamido)-3,4,6-tri-O-acetyl-D-galactose
[0445] According to Baisch et al., Bioorg. Med. Chem., 1997, 5, 383-391 (included in this paper by reference), α-(2'-azido-2',2'-difluoroacetamido)-3,4,6-tri-O-acetyl-D-galactose-1-phosphate was coupled to UMP.
[0446] Therefore, the solution of D-uridine-5'-monophosphate disodium salt (98 mg, 0.266 mmol) in H2O (1 mL) was prepared using a DOWEX 50Wx8 (H) solution. + The mixture was treated with (type) for 40 minutes and filtered. The filtrate was stirred vigorously at room temperature while tributylamine (63 μL, 0.266 mmol) was added dropwise. After further stirring for 30 minutes, the reaction mixture was lyophilized and then further dried under vacuum with P2O5 for 5 h.
[0447] The obtained uridine-5'-tributylammonium monophosphate was dissolved in anhydrous DMF (15 mL) under an argon atmosphere. Carbonyl diimidazole (35 mg, 0.219 mmol) was added, and the reaction mixture was stirred at room temperature for 30 min. Next, anhydrous MeOH (4.63 μL) was added and stirred for 15 min to remove excess carbonyl diimidazole. The remaining MeOH was removed under high vacuum for 15 min. Subsequently, N-methylimidazole HCl salt (61 mg, 0.52 mmol) was added to the reaction mixture, and the resulting compound (63 mg, 0.125 mmol) was dissolved in anhydrous DMF (15 mL) and added dropwise to the reaction mixture. The reaction was stirred overnight at room temperature and then concentrated under vacuum. The consumption of the imidazole-UMP intermediate was monitored by MS. UDP-α-(2'-azido-2',2'-difluoroacetamido)-3,4,6-tri-O-acetyl-D-galactose was obtained by rapid chromatography (7:2:1-5:2:1EtOAC:MeOH:H2O).
[0448] 1 H-NMR (300MHz, D2O): δ7.87(d,J=8.1Hz,1H),5.913-5.85(m,2H),5.67(dd,J=6.6,2.7Hz,1H),5.56-5.50(m,1H),5 .47-5.43(m,1H),5.31-5.25(m,2H),4.61-4.43(m,2H),4.31-4.05(m,5H),2.16(s,3H),2.02(s,3H),1.94(s,3H).
[0449] C 23 H 29 F2N6O 20 P2(MH + The calculated LRMS(ESI-) value is 809.09, and the measured value is 809.1.
[0450] 1H-NMR (300MHz, CD3OD): δ5.64(m,1H),5.47(d,J=2.4Hz,1H),5.35(dd,J=11.4,3.0Hz,1H),4. 58-4.48(m,2H),4.25-4.15(m,1H),4.09-4.00(m,1H),2.14(s,3H),2.00(s,3H),1.93(s,3H).
[0451] C 14 H 18 F2N4O 12 P(MH + The calculated LRMS(ESI-) value is 503.06, and the measured value is 503.0.
[0452] Example 13. Synthesis of α-UDP-2-(2'-azido-2',2'-difluoroacetamido)-2-deoxy-D-galactose (UDP-F2-GalNAz, 18)
[0453] According to Kiso et al., Glycoconj.J., 2006, 23, 565 (included in this paper by reference), the deacetylation of UDP-α-(2'-azido-2',2'-difluoroacetamido)-3,4,6-tri-O-acetyl-D-galactose was carried out.
[0454] Therefore, UDP-α-(2'-azido-2',2'-difluoroacetamido)-3,4,6-tri-O-acetyl-D-galactose was dissolved in H2O (1 mL), and triethylamine (1 mL) and MeOH (2.4 mL) were added. The reaction mixture was stirred for 2 h, and then concentrated under vacuum. α-UDP-2-(2'-azido-2',2'-difluoroacetamido)-2-deoxy-D-galactose (18) was obtained by rapid chromatography (7:2:1-5:2:1EtOAC:MeOH:H2O).
[0455] 1 H-NMR (300MHz, D2O): δ7.86 (d, J=8.1Hz, 1H), 5.91-5.85 (m, 2H), 5.54 (dd, J=6.6, 3.6Hz, 1H), 4.31-3.95 (m, 9H), 3.74-3.62 (m, 2H).
[0456] C 17 H 23 F2N6O 17 P2(MH + The calculated LRMS(ESI-) value is 683.06, and the measured value is 683.10.
[0457] Example 14. Synthesis of α-UDP-2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose
[0458] According to Baisch et al., Bioorg. Med. Chem., 1997, 5, 383-391, α-2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose-1-phosphate prepared as in Example 9 was linked to UMP.
[0459] Therefore, the solution of D-uridine-5'-monophosphate disodium salt (1.49 g, 4.05 mmol) in H2O (15 mL) was prepared using a DOWEX 50Wx8 (H) solution. + The mixture was treated with (type) for 30 minutes and filtered. The filtrate was stirred vigorously at room temperature while tributylamine (0.966 mL, 4.05 mmol) was added dropwise. After stirring for another 30 minutes, the reaction mixture was lyophilized and then further dried under vacuum with P2O5 for 5 h.
[0460] The obtained uridine-5'-tributylammonium monophosphate was dissolved in anhydrous DMF (25 mL) under an argon atmosphere. Carbonyl diimidazole (1.38 g, 8.51 mmol) was added, and the reaction mixture was stirred at room temperature for 30 min. Next, anhydrous MeOH (180 μL) was added and stirred for 15 min to remove excess carbonyl diimidazole. The remaining MeOH was removed under high vacuum for 15 min. The resulting compound (2.0 g, 4.86 mmol) was dissolved in anhydrous DMF (25 mL) and added dropwise to the reaction mixture. The reaction was stirred at room temperature for 2 days, then concentrated under vacuum. The consumption of the imidazole-UMP intermediate was monitored by MS. Rapid chromatography (7:2:1-5:2:1EtOAC:MeOH:H2O) yielded α-UDP-2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose (1.08 g, 1.51 mmol, 37%).
[0461] 1 H-NMR (300MHz, D2O): δ7.96(d,J=8.0Hz,1H),5.98-5.94(m,2H),5.81-5.79(m,1H),5.70(dd,J=7.1,3.3Hz,1H),5.49(dd,J=15.2,2.6Hz,1H ),5.30(ddd,J=18.5,11.0,3.2Hz,2H),4.57(q,J=6.0Hz,2H),4.35-4.16(m,9H),4.07-3.95(m,2H),2.17(s,3H),2.08(s,3H),2.07(s,3H).
[0462] C 21 H 29 N5O 19 P2(MH + The calculated LRMS(ESI-) value is 716.09, and the measured value is 716.3.
[0463] Example 15. Synthesis of α-UDP-2-azido-2-deoxy-D-galactose
[0464] According to Kiso et al., Glycoconj.J., 2006, 23, 565, the deacetylation of α-UDP-2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose prepared as in Example 14 was performed.
[0465] Therefore, α-UDP-2-azido-2-deoxy-3,4,6-tri-O-acetyl-D-galactose (222 mg, 0.309 mmol) was dissolved in H2O (2.5 mL), and triethylamine (2.5 mL) and MeOH (6 mL) were added. The reaction mixture was stirred for 3 hours and then concentrated under vacuum to obtain the crude product α-UDP-2-azido-2-deoxy-D-galactose. 1 H-NMR (300MHz, D2O): δ7.99 (d, J=8.2Hz, 1H), 6.02-5.98 (m, 2H), 5.73 (dd, J=7.4, 3.4Hz, 1H), 4 .42-4.37(m,2H),4.30-4.18(m,4H),4.14-4.04(m,2H),3.80-3.70(m,2H),3.65-3.58(m,1H).
[0466] C 15 H 23 N5O 16 P2(MH + The calculated LRMS(ESI-) value is 590.05, and the measured value is 590.2.
[0467] Example 16. Synthesis of α-UDP-D-galactosamine (Gal-NH2)
[0468] Lindlar catalyst (50 mg) was added to a 1:1 (4 mL) H₂O:MeOH solution of α-UDP-2-azido-2-deoxy-D-galactose prepared as in Example 15. The reaction was stirred for 5 h under a hydrogen atmosphere and filtered through diatomaceous earth. The filter was rinsed with 10 mL H₂O, and the filtrate was concentrated under vacuum to give α-UDP-D-galactosamine (UDP-GalNH₂) (169 mg, 0.286 mmol, 92% yield in two steps). 1H-NMR (300MHz, D2O): δ7.93(d,J=8.1Hz,1H),5.99-5.90(m,2H),5.76-5.69(m,1H),4.39-4.34(m,2 H),4.31-4.17(m,5H),4.05-4.01(m,1H),3.94-3.86(m,1H),3.82-3.70(m,3H),3.30-3.16(m,1H). C 15 H 25 N3O 16 P2(MH + The calculated LRMS(ESI-) value is 564.06, and the measured value is 564.10.
[0469] Example 17. Synthesis of α-UDP-N-(4'-azido-3',5'-difluorobenzoyl)-D-galactosamine (24, X=F)
[0470] According to Rademan et al., Angew. Chem. Int. Ed., 2012, 51, 9441-9447 (included in this paper by reference), the process for preparing 4-azido-3,5-difluorobenzoic acid succinimide ester is described.
[0471] Therefore, dicyclohexylcarbodiimide (1.1 equivalents) and N-hydroxysuccinimide (1.2 equivalents) were added to a solution of 4-azido-3,5-difluorobenzoic acid, and the resulting suspension was stirred overnight, followed by vacuum filtration. The filtrate was concentrated and dissolved in EtOAC, then washed with saturated NaHCO3 and brine. The organic layer was dried over Na2SO4, filtered, and concentrated under vacuum, and the crude product was used in the next reaction.
[0472] 1 H-NMR (300MHz, CDCl3): δ7.74-7.66 (m, 2H), 2.91 (s, 4H).
[0473] Next, UDP-GalNH2 (30 mg, 0.0531 mmol) prepared in Example 16 was dissolved in 0.1 M NaHCO3 (0.2 M), and N-hydroxysuccinimide ester of 4-azido-3,5-difluorobenzoic acid (31 mg, 0.106 mmol, 2 equivalents) dissolved in DMF (0.2 M) was added. The reaction was stirred overnight at room temperature and concentrated under vacuum. Rapid chromatography (7:2:1-5:2:1 EtOAC:MeOH:H2O) gave product 24 (X=F) (8 mg, 0.0107 mmol, 20%).
[0474] 1H-NMR (300MHz, D2O): δ7.73(d,J=8.4Hz,1H),7.52-7.31(m,2H),5.87-5.71(m,2 H),5.65-5.57(m,1H),5.47-5.33(m,1H),4.43-3.96(m,8H),3.76-3.60(m,2H).
[0475] C 22 H 25 F2N6O 17 P2(MH + LRMS (ESI) - Calculated value: 745.07, measured value: 744.9.
[0476] Example 18. Determination of specific enzyme activity
[0477] The specific activity of the enzyme was determined by a coupled glycosyltransferase process as described by Wu et al. in Glycobiology 2010, 21, 723-733 (included hereby by reference), and is available as a kit from R&D systems: (http: / / www.rndsystems.com / Products / EA001 (accessed July 31, 2014)).
[0478] In short, the glycosyltransferase reaction was performed for 20 min at room temperature in 96-well plates in 50 μL of reaction buffer (25 mM Tris, 150 mM NaCl, 5 mM MgCl2, and 5 mM MnCl2, pH 7.5). To determine the kinetic parameters of the glycosyltransferase, multiple reactions were performed simultaneously with various amounts of enzyme (0.01, 0.02, 0.04, 0.06, 0.08, 0.10, and 0.12 μg enzyme) in the presence of all other components (GlcNAc: 20 mM, UDP-F2-GalNAz: 500 μM), including coupled phosphatase 1 (ENTPD3 / CD39L3) (2.5 μL in a 20 ng / μL solution). Wells containing all components except the enzyme served as blank controls. The reaction was initiated by adding substrate and phosphatase to the enzyme and terminated by adding 30 μL of malachite reagent A and 100 μL of water to each well. Color development was achieved by adding 30 μL of malachite reagent B to each well, followed by gentle mixing and incubation at room temperature for 20 min. After color development, the plate was read at 620 nm using a multi-well plate reader. A phosphate standard curve was also prepared to determine the conversion factor between absorbance and inorganic phosphate content. The preparation of UDP-F2-GalNAz is described in Example 13.
[0479] The specific activity of GalNAcT described herein is compared with that of GalT(Y289L), an enzyme known for transferring UDP-GalNAz as previously disclosed in WO 2007 / 095506 and WO 2008 / 029281 (both from Invitrogen Corporation).
[0480] Data collected from the determination of the specific activities of several enzymes are shown in Tables 1 and 2.
[0481] Table 1: Data from specific activity assays of CeGalNAcT, CeGalNAcT-His, TnGalNAcT and GalT(Y289L).
[0482]
[0483]
[0484] Table 2: Data from specific activity assays of CeGalNAcT (I257L), CeGalNAcT (I257A), and CeGalNAcT (I257M).
[0485]
[0486] The conversion rate (in pmol / min / μg) plotted relative to the amount of enzyme is shown in the graph. Figure 7 (Regarding Table 1) and Figure 8 (Referring to Table 2). Specific activity was calculated from these curves using linear regression.
[0487] Example 19. Activity determination of CeGalNAcT (M312H)
[0488] The activity of CeGalNAcT(M312H) was also measured using the same process as described in Example 14, however, Mg was used in this case. 2+ Replace Mn 2+ Under these conditions, a specific activity of 15 pmol / min / μg was measured.
[0489] Example 20. Trimming IgG polysaccharides with Endo S (General Procedure)
[0490] IgG glycan trimming was performed using Endo S (obtained from Genovis, Lund, Sweden) from *Streptococcus pyogenes*. IgG (10 mg / mL) was incubated with Endo S (40 U / mL) at 37°C for approximately 16 hours in 25 mM Tris pH 8.0. The deglycosylated IgG was concentrated and washed with 10 mM MnCl2 and 25 mM Tris-HCl pH 8.0 using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore).
[0491] Example 21. Trimming trastuzumab
[0492] The above-described trimming protocol was applied to trastuzumab. The trimmed antibody was analyzed on a JEOL AccuToF equipped with an Agilent 1100 HPLC. The sample (2 μL) was reduced with DTT (2 μL) over 10 min, followed by dilution with H2O (40 μL) and then injected. After deconvolution of the peaks, the mass spectrometry showed one peak in the light chain and two peaks in the heavy chain. The two peaks in the heavy chain belonged to a major product (49496 Da, 90% of total heavy chain) generated by trastuzumab with core GlcNAc (Fuc) substitution and a minor product (49351 Da, ±10% of total heavy chain) generated by trastuzumab with core GlCNac substitution.
[0493] This is an example of a glycoprotein containing a polysaccharide of formula (1).
[0494] Example 22. Expression of trastuzumab in the presence of sorghum extract
[0495] Trastuzumab was transiently expressed in CHO K1 cells by Evitria (Zurich, Switzerland) in the presence of 10 or 25 μg / mL sorghum extract (commercially available from Sigma-Aldrich), purified using protein A agarose gel, and analyzed by mass spectrometry. Two concentrations of sorghum extract produced three major heavy chain products of trastuzumab, corresponding to the trastuzumab heavy chains replaced by GlcNAc-Man5-GlcNAc-GlcNAc(Fuc)- (c = d = 0, 50712 Da, ±20% of total heavy chain products), Gal-GlcNAc-Man5-GlcNAc-GlcNAc(Fuc)- (c = 1, d = 0, 50874 Da, ±35% of total heavy chain products), and Sial-Gal-GlcNAc-Man5-GlcNAc-GlcNAc(Fuc)- (c = d = 1, 51164 Da, ±35% of total heavy chain products).
[0496] Example 23. Pruning with sialidase / galactosidase to obtain trast-Man5GlcNAc
[0497] Trastuzumab (10 mg / mL) transiently expressed in the presence of stilbene as described in Example 18 was incubated with neuraminidase (0.5 mU / mg IgG) from Vibrio cholerae (commercially available from Sigma-Aldrich) in 100 mM sodium acetate at pH 6.0 and 2 mM CaCl2 for 16 hours, which resulted in the complete removal of sialic acid (the two major heavy chain products, 50712 and 50874 Da, which correspond to approximately 20% and 70% of the total heavy chain products, respectively). When the same reaction was performed in the presence of β(1,4)-galactosidase (3 mU / mg IgG) from Streptococcus pneumoniae (commercially available from Calbiochem), a single major heavy chain product was observed, corresponding to trastuzumab with a heavy chain substituted with GlcNAc-Man5-GlcNAc-GlcNAc(Fuc) (c = d = 0.50712 Da, ±% of total heavy chain product, and minor heavy chain product of 50700–50900 Da).
[0498] This is an example of a glycoprotein containing a polysaccharide of formula (26).
[0499] Example 24. Trimming with galactosidase to obtain trast-Man3GlcNAc2
[0500] Trastuzumab (10 mg / mL) in 50 mM sodium phosphate at pH 6.0 and β(1,4)-galactosidase (3 mU / mg IgG) from Streptococcus pneumoniae (commercially available from Calbiochem) were stirred at 37 °C for 16 h, followed by MS analysis. A single major heavy chain product was observed, corresponding to trastuzumab with a heavy chain substituted with GlcNAc2-Man3-GlcNAc-GlcNAc(Fuc) (c = d = 0, 50592 Da).
[0501] This is an example of a glycoprotein containing a polysaccharide of formula (27).
[0502] General protocol for mass spectrometry analysis of IgG
[0503] A solution of approximately 70 μL of 50 μg (modified) IgG, 1 M Tris-HCl pH 8.0, 1 mM EDTA, and 30 mM DTT was incubated at 37 °C for 20 min to reduce disulfide bonds, thereby enabling analysis of the light and heavy chains. If present, azide functional groups were reduced to amines under these conditions. The reduced sample was washed twice with milliQ using an Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore) and concentrated to 10 μM (modified) IgG. The reduced IgG was analyzed by electrospray ionization time-of-flight (ESI-TOF) on a JEOLAccuTOF. Deconvolutional spectra were obtained using Magtran software.
[0504] Glycosyl transfer of galactose derivatives (e.g., azide) using GalNAcT (general procedure)
[0505] Galactose derivatives (e.g., azido-containing sugars) were enzymatically introduced onto IgG using GalNAc transferase or its mutants. Deglycosylated IgG (prepared as described above, 10 mg / mL) was incubated with modified UDP-galactose derivatives (e.g., azido-modified sugar-UDP derivatives) (0.4 mM) and GalNAcT (1 mg / mL) at 30°C for 16 h in 10 mM MnCl2 and 25 mM Tris-HCl pH 8.0. Functionalized IgG (e.g., azido-functionalized IgG) was incubated with protein A agarose (40 μL / mg IgG) at 4°C for 2 h. The protein A agarose was washed three times with PBS, and the IgG was eluted with 100 mM glycine-HCl pH 2.7. The eluted IgG was neutralized with 1M Tris-HCl pH 8.0, concentrated, and washed with PBS using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore) to a concentration of 15-20 mg / mL.
[0506] Example 25. Trastuzumab (GalNAz) 2
[0507] Trimmed trastuzumab was subjected to a glycosyltransfer protocol using UDP-N-azidoacetylgalactosamine (UDP-GalNAz) and CeGalNAcT. After protein A affinity purification, small samples were reduced with DTT and subsequently analyzed by MS, revealing the formation of a major product (49713 Da, 90% of total heavy chain) resulting from the transfer of GalNAz to the core GlcNAc (Fuc)-substituted trastuzumab, and a minor product (49566 Da, ±10% of total heavy chain) resulting from the transfer of GalNAz to the core GlcNAc-substituted trastuzumab.
[0508] Example 26. Trastuzumab (F2-GalNAz) 2
[0509] The pruned trastuzumab was subjected to a glycosyltransfer protocol using UDP-N-azidodifluoroacetylgalactosamine (UDP-F2-GalNAz) and GalNAcT (GalNAcT selected from CeGalNAcT (or its mutants as described above), AsGalNAcT, TnGalNAcT, or DmGalNAcT). After protein A affinity purification, small samples were reduced with DTT and subsequently analyzed by MS, which showed the formation of a major heavy chain product (49865 Da, approximately 90% of the total heavy chain) resulting from the transfer of F2-GalNAz to the core GlcNAc (Fuc)-substituted trastuzumab (which reacted with DTT during sample preparation).
[0510] Example 27. Trastuzumab-(F2-GalNBAz)2
[0511] Trimmed trastuzumab (10 mg / mL, 6.6 nmol) (obtained by treating trastuzumab as in Formula 1 with Endo S) was incubated overnight at 30 °C with UDP-F2GalNBAz (24, X = F, 7 mM) and CeGalNAcT (2 mg / mL) in 10 mM MnCl2 and 25 mM Tris-HCl at pH 8.0. Mass spectrometry analysis of the reduced sample showed the formation of a major product (49815 Da, approximately 90% of the total heavy chain), which was generated by the transfer of F2-GalNBAz to the core GlcNAc (Fuc)-substituted trastuzumab heavy chain.
[0512] Example 28. Trastuzumab-(Man5GlcNAc-GalNAz)2
[0513] Trastuzumab substituted with GlcNAc-Man5-GlcNAc-GlcNAc(Fuc) (as in Formula 26, obtained by transient expression of trastuzumab in the presence of squalene and pruned with neuraminidase and galactosidase as described in Example 19) was subjected to a glycosyltransfer protocol using UDP-GalNAz and GalNAcT. The pruned antibody was incubated with UDP-GalNAz (0.5 mM) (commercially available from Glycohub, Inc.) and CeGalNAcT (0.1 mg / mL) at 30°C for 16 h in 10 mM MnCl2 and 25 mM Tris-HCl pH 8.0, which resulted in complete conversion to GalNAz-GlcNAc-Man5-GlcNAc-GlcNAc(Fuc)-substituted trastuzumab (50929 Da major heavy chain product, ±90% of total heavy chain product).
[0514] Example 29. Trastuzumab-(Man3(GlcNAc-GalNAz)2)2
[0515] Trastuzumab obtained after galactosidase trimming (as in Formula 27, the preparation of which is described in Example 24) was subjected to a glycosyltransfer protocol using UDP-GalNAz and GalNAcT. The trimmed antibody was incubated with UDP-GalNAz (0.5 mM) (commercially available from Glycohub, Inc.) and CeGalNAcT (0.1 mg / mL) at 22°C for 16 h in 10 mM MnCl2 and 25 mM Tris-HCl pH 8.0, which resulted in complete conversion to (GalNAz-GlcNAc)2-Man3-GlcNAc-GlcNAc(Fuc)-substituted trastuzumab (51027 Da major heavy chain product, ±90% of total heavy chain product).
[0516] Example 30. Using CeGalNAcT(M312H) and Mg 2+ Trastuzumab (F2-GalNAz)2
[0517] In the presence of 10 mM MgCl2, trimmed trastuzumab was subjected to a glycosyltransfer protocol using UDP-N-azidodifluoroacetylgalactosamine (UDP-F2-GalNAz) and CeGalNAcT (M312H), as well as GalNAcT (1 mg / mL). After protein A affinity purification, small samples were reduced with DTT and subsequently analyzed by MS, revealing the formation of a major heavy chain product (49873 Da, approximately 25% of the total heavy chain, the remainder being the remaining trimmed antibody), which was generated by the transfer of F2-GalNAz to the core GlcNAc (Fuc)-substituted trastuzumab (which had already reacted with DTT during sample preparation).
[0518] Synthesize additional sugar derivative nucleotides Su(A)-Nuc(3)
[0519] The additional sugar derivative nucleotide Su(A)-Nuc of formula (3) is prepared according to the procedures disclosed, for example, WO 2014 / 065661 (SynAffix B.V.); Pouilly et al., ACS Chem. Biol. 2012, 7, 753 and Guan et al., Chem. Eur. J. 2010, 16, 13343 (all incorporated herein by reference).
[0520] Example 31. Synthesize UDP-GalNAcSAC((20)), where V is 0, r is 0 and R 8 (for CH3)
[0521] UDP-D-galactosamine (Example 16) (45 mg, 0.0796 mmol) was dissolved in a pH 7 buffer (0.5 Mk2HPO4) (2 mL). N-succinimide-S-acetylthioacetate (37 mg, 0.159 mmol) and DMF (2 mL) were added, and the reaction was stirred overnight at room temperature. An additional 36 mg of N-succinimide-S-acetylthioacetate was added, and after 3 hours, the reaction was concentrated under vacuum. Rapid chromatography (7:2:1–5:2:1 EtOAC:MeOH:H2O) yielded UDP-GalNAcSAc (28 mg, 0.041 mmol, 52%).
[0522] 1 H-NMR (300MHz, D2O): δ7.84 (d, J = 8.1Hz, 1H), 5.90-5.82 (m, 2H), 5.48-5.41 (m, 1H), 4. 29-4.22(m,2H),4.20-4.00(m,5H),3.98-3.82(m,2H),3.79-3.59(m,4H),2.30(s,3H). C 19 H 29 N3O 18 P2S(MH + The calculated LRMS(ESI-) value is 680.06, and the measured value is 680.1.
[0523] Example 32. Synthesis of UDP-GalNAcCl ((19), where X is Cl)
[0524] UDP-D-galactosamine (Example 16) (42 mg, 0.074 mmol) was dissolved in 0.1 M NaHCO3 (1 mL), and N-(chloroacetoxy)succinimide (29 mg, 0.149 mmol) (prepared according to Hosztafi et al., Helv. Chim. Acta, 1996, 79, 133-136) and DMF (1 mL) were added. The reaction was stirred overnight at room temperature, and another 10 mg of N-(chloroacetoxy)succinimide was added and stirring was continued overnight. The reaction was concentrated under vacuum and purified by rapid chromatography (7:2:1-5:2:1 EtOAC:MeOH:H2O) to give UDP-GalNAcCl (25 mg, 0.039 mmol, 53%).
[0525] 1 H-NMR (300MHz, D2O): δ7.84 (d, J = 8.1Hz, 1H), 5.89-5.84 (m, 2H), 5.53-5.46 ( m,1H),4.33-4.00(m,9H),3.99-3.88(m,2H),3.77-3.59(m,2H),1.83(s,1H).
[0526] C 17 H 26 ClN3O 17 P2(MH + The calculated LRMS (ESI-) values were 640.03 (100%) and 642.03 (32%), while the measured values were 640.1 (100%) and 642.2 (35%).
[0527] Example 33. Synthesis of UDP-GalNAcBr ((19), where X is Br)
[0528] UDP-D-galactosamine (Example 16) (42 mg, 0.088 mmol) was dissolved in 0.1 M NaHCO3 (3 mL), and N-(bromoacetoxy)succinimide (63 mg, 0.265 mmol) (prepared according to Hosztafi et al., Helv. Chim. Acta, 1996, 79, 133-136) and DMF (2 mL) were added. The reaction was stirred overnight at room temperature and concentrated under vacuum. The compound was purified by rapid chromatography (7:2:1-5:2:1 EtOAC:MeOH:H2O) to give UDP-GalNAcBr (28 mg, 0.048 mmol, 65%).
[0529] 1H-NMR (300MHz, D2O): δ7.86 (d, J = 3.2Hz, 1H), 5.97-5.84 (m, 2H), 5.54-5.46 (m, 1 H),4.33-4.04(m,6H),3.99-3.85(m,2H),3.79-3.60(m,2H),2.75-2.68(m,3H).
[0530] C 17 H 26 BrN3O 17 P2(MH + The calculated LRMS(ESI-) values were 683.98 (100%) and 685.98 (98%), while the measured values were 687.1 (100%), 688.0 (92%), 686.0 (85%), and 689.0 (72%).
[0531] Example 34. Design of GalNAcT mutants of the white-spotted armyworm and the pig ascarid.
[0532] Mutants of TnGalNAcT and AsGalNAcT were designed based on the crystal structure of bovine β(1,4)-Gal-T1 complexed with UDP-N-acetyl-galactosamine (PDB entry 1OQM) and the β(1,4)-Gal-T1 (Y289L) mutant reported by Qasba et al. (J. Biol. Chem. 2002, 277: 20833-20839, incorporated herein by reference). Mutants of TnGalNAcT and AsGalNAcT were designed based on sequence alignment of TnGalNAcT and AsGalNAcT with bovine β(1,4)Gal-T1. The corresponding amino acid residues between these proteins are shown in Table 3.
[0533] Table 3. Number of corresponding amino acids in different GalNAcT / GalT types
[0534] TnGalNAcT AsGalNAcT Bovine β(1,4)-Gal-T1 I311 I257 Y289 W336 W282 W314 E339 E285 E317
[0535] Example 35. Site-directed mutagenesis of His6-TnGalNAcT(33-421) mutants
[0536] Obtain the pET15b vector containing the codon-optimized sequence (encoding residues 33-421 of TnGalNAcT (represented by SEQ ID NO:8) between the NdeI-BamHI sites) from Genscript to generate His6-TnGalNAcT(33-421) (represented by SEQ ID NO:49). The TnGalNaCT mutant gene was amplified from the above constructs by linear amplification PCR using a set of overlapping primers. The overlapping primer sets used for each mutant are shown in Table 4. For constructing His6-TnGalNAcT(33-421; W336F) (represented by SEQ ID NO:50), the DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:79 and SEQ ID NO:80. For constructing His6-TnGalNAcT(33-421; W336H) (represented by SEQ ID NO:51), the DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:81 and SEQ ID NO:82. To construct His6-TnGalNAcT(33-421; W336V) (represented by SEQ ID NO:52), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:83 and SEQ ID NO:84. To construct His6-TnGalNAcT(33-421; E339A) (represented by SEQ ID NO:53), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:85 and SEQ ID NO:86. To construct His-TnGalNAcT(33-421; E339D) (represented by SEQ ID NO:55), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:88 and SEQ ID NO:89. To construct His6-TnGalNAcT(33-421; I299M) (represented by SEQ ID NO:45), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:61 and SEQ ID NO:62. To construct His6-TnGalNAcT(33-421; I299A) (represented by SEQ ID NO:48), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:63 and SEQ ID NO:64. To construct His6-TnGalNAcT(33-421; I299G) (represented by SEQ ID NO:54), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:65 and SEQ ID NO:66.To construct His6-TnGalNAcT(33-421; L302A) (represented by SEQ ID NO:43), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:67 and SEQ ID NO:68. To construct His6-TnGalNAcT(33-421; L302G) (represented by SEQ ID NO:44), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:69 and SEQ ID NO:70. To construct His6-TnGalNAcT(33-421; I311M) (represented by SEQ ID NO:60), a DNA fragment was amplified using a primer pair defined herein as SEQ ID NO:74 and SEQ ID NO:87. After PCR amplification, the reaction mixture was treated with DpnI to digest the template DNA, and then transformed into NEB 10-β competent cells (obtained from New En gland Biolabs). DNA was isolated and the sequences were confirmed by sequence analysis of the following mutants: His6-TnGalNAcT(33-421; W336F) (represented by SEQ ID NO:50), His6-TnGalNAcT(33-421; W336V) (represented by SEQ ID NO:52), His6-TnGalNAcT(33-421; E339A) (represented by SEQ ID NO:53), His6-TnGalNAcT(33-421; I299M) (represented by SEQ ID NO:45), His6-TnGalNAcT(33-421; I299A) (represented by SEQ ID NO:48), His6-TnGalNAcT(33-421; I299G) (represented by SEQ ID NO:54), and His6-TnGalNAcT(33-421; L302A) (represented by SEQ ID NO:50). (represented by NO:43), His6-TnGalNAcT(33-421; L302G) (represented by SEQ ID NO:44), and His6-TnGalNAcT(33-421; I311M) (represented by SEQ ID NO:60).
[0537] Table 4. Sequence identifiers of the primers used. Codons corresponding to the mutant amino acids are shown in bold.
[0538]
[0539]
[0540] Example 36. His6-TnGalNAcT (33-421), His6-TnGalNAcT (33-421; W336F), His6-TnGalNAcT (33-421; W3 36V), His6-TnGalNAcT (33-421; I299M), His6-TnGalNAcT (33-421; I299A), His6-TnGalNAcT (33-42 1. Expression and refolding of His6-TnGalNAcT(33-421; L302A), His6-TnGalNAcT(33-421; L302G), His6-TnGalNAcT(33-421; I311M), His6-TnGalNAcT(33-421; E339D) and His6-TnGalNAcT(33-421; E339A) in Escherichia coli
[0541] His6-TnGalNAcT(33-421); His6-TnGalNAcT(33-421; W336F); His6-TnGalNAcT(33-421; W336V); His6-TnGalNAcT(33-421; I299M); His6-TnGalNAcT(33-421; I299A); His6 The following proteins were identified: His6-TnGalNAcT (33-421; I299G), His6-TnGalNAcT (33-421; L302A), His6-TnGalNAcT (33-421; L302G), His6-TnGalNAcT (33-421; I311M), His6-TnGalNAcT (33-421; E339D), and His6-TnGalNAcT (33-421; E339A). Expression, inclusion body isolation, and refolding were performed according to the procedure reported by Qasba et al. (Prot. Expr. Pur. 2003, 30, 219-76229, incorporated herein by reference). After refolding, insoluble proteins were removed by centrifugation (14000 × g for 10 min at 4 °C), followed by filtration through a 0.45 μM filter. Soluble proteins were purified and concentrated using a HisTrap HP 5mL column (GE Healthcare). The column was first washed with buffer A (5mM Tris buffer, 20mM imidazole, 500mM NaCl, pH 7.5). Retained proteins were eluted with buffer B (20mM Tris, 500mM NaCl, 500mM imidazole, pH 7.5, 10mL). Fractions were analyzed by SDS-PAGE on a 12% polyacrylamide gel. Fractions containing the purified target protein were combined and dialyzed overnight at 4°C, replacing the buffer with 20mM Tris pH 7.5 and 500mM NaCl. The purified protein was concentrated to at least 2 mg / mL using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore) and stored at -80°C before further use.
[0542] Example 37. Design of the GalNAcT mutant of W336 from the moth *Hemiberlesia lataniae*
[0543] Based on the sequence alignment of TnGalNAcT and GalT, by using tryptophan... 336 Three active site mutants were designed by mutating to phenylalanine, histidine, or valine (underlined).
[0544] T. ni (His6-TnGalNAcT[33-421; W336F] represented by SEQ ID NO: 50)
[0545] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTELELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWG F GGEDDDMSYRLKKINYHIARYKMSIARYAMIDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0546] T. ni (His6-TnGalNAcT[33-421; W336H] represented by SEQ ID NO: 51)
[0547] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTELELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWG HGGEDDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0548] The white-striped armyworm (His6-TnGalNAcT[33-421; W336V] represented by SEQ ID NO: 52)
[0549] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAPAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTEL ELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWG V GGEDDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0550] Example 38. Design of the GalNAcT mutant of L302 in the moth *Arbuscular simonii*
[0551] Based on the sequence alignment of TnGalNAcT and GalT, by leucine 302 Two active site mutants were designed by mutating to glycine or alanine (underlined).
[0552] The white-striped armyworm (His6-TnGalNAcT[33-421; L302G] represented by SEQ ID NO: 44)
[0553] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTELELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDK G HFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWGWGGEDDDMSYRLKKINYHLARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0557] Based on the sequence alignment of TnGalNAcT and GalT, by glutamic acid 339 Two active site mutants were designed by mutating to glycine, alanine, or aspartic acid (underlined).
[0558] The white-spotted moth (His6-TnGalNAcT[33-421; E339A] represented by SEQ ID NO: 53)
[0559] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTELE LEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWGWGG A DDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0560] The white-striped armyworm (His6-TnGalNAcT[33-421; E339D] represented by SEQ ID NO: 55)
[0561] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTELE LEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWGWGG D DDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0562] Example 40. Design of the GalNAcT mutant I299M of the white-spotted armyworm.
[0563] Based on the sequence alignment of TnGalNAcT and GalT, by isoleucine 299 Three active site mutants were designed by mutating to methionine, alanine, or glycine (underlined).
[0564] The white-striped armyworm (His6-TnGalNAcT[33-421; I299M] represented by SEQ ID NO: 45)
[0565] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLP LCDSMPPDLGPITLNKTELELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSAS MDKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWGWGGEDDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0566] Trichoplusia ni (His6-TnGalNAcT[33-421; I299A] represented by SEQ ID NO: 48)
[0567] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKTELELEWVEKKFPEVEWGGRYSPPNCTARHRVAHVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASADKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWGWGGEDDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0568] Trichoplusia ni (His6-TnGalNAcT[33-421; I299G] represented by SEQ ID NO: 54)
[0569] MGSSHHHHHHSSGLVPRGSHMSPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLP LCDSMPPDLGPITLNKTELELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSAS G DKLHFKLPYEDIFGGVSAMTLEQFTRVNGFSNKYWGWGGEDDDMSYRLKKINYHIARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0570] Example 41. Design of the GalNAcT mutant I311M of the white-spotted armyworm.
[0571] By using isoleucine 311 The TnGalNAcT mutant was designed by mutating to methionine.
[0572] Powdered Noctuid moth (represented by SEQ ID NO: 60, TnGalNAcT[33-421; I311M])
[0573] SPLRTYLYTPLYNATQPTLRNVERLAANWPKKIPSNYIEDSEEYSIKNISLSNHTTRASVVHPPSSITETASKLDKNMTIQDGAFAMISPTPLLITKLMDSIKSYVTTEDGVKKAEAVVTLPLCDSMPPDLGPITLNKT ELELEWVEKKFPEVEWGGRYSPPNCTARHRVAIIVPYRDRQQHLAIFLNHMHPFLMKQQIEYGIFIVEQEGNKDFNRAKLMNVGFVESQKLVAEGWQCFVFHDIDLLPLDTRNLYSCPRQPRHMSASIDKLHFKLPYED MFGGVSAMTLEQFTRVNGFSNKYWGWGGEDDDMSYRLKKINYHLARYKMSIARYAMLDHKKSTPNPKRYQLLSQTSKTFQKDGLSTLEYELVQVVQYHLYTHILVNIDERS
[0574] Example 42. Synthesis of UDP-GalNPropN3 ((9), where U is [CH2]2 and A = N3)
[0575] UDP-GalNPropN3, a sugar derivative nucleotide of formula (3), was prepared according to the procedures disclosed in, for example, WO 2014 / 065661 (SynAffix BV); Pouilly et al., ACS Chem. Biol. 2012, 7, 753 and Guan et al., Chem. Eur. J. 2010, 16, 13343 (all incorporated herein by reference).
[0576] Example 43. Synthesize UDP-GalNButN3 ((9), where U is [CH2]3 and A = N3)
[0577] UDP-GalNButN3, a sugar derivative nucleotide of formula (3), was prepared according to the procedures disclosed in, for example, WO 2014 / 065661 (SynAffix BV); Pouilly et al., ACS Chem. Biol. 2012, 7, 753 and Guan et al., Chem. Eur. J. 2010, 16, 13343 (all incorporated herein by reference).
[0578] Example 44. Synthesis of UDP-GalNProSH ((21), where t=2)
[0579] UDP-GalNProSH, a sugar derivative nucleotide of formula (3), is prepared according to the procedures disclosed in, for example, WO 2015 / 057063 (SynAffix BV); Pouilly et al., ACS Chem. Biol. 2012, 7, 753 and Guan et al., Chem. Eur. J. 2010, 16, 13343 (all incorporated herein by reference).
[0580] Example 45. Synthesis of UDP-GalNB z N3((23), where X is H)
[0581] UDP-GalNBzN3, a sugar derivative nucleotide of formula (3), was prepared according to the procedures disclosed in, for example, WO 2015 / 112013 (SynAffix BV); Pouilly et al., ACS Chem. Biol. 2012, 7, 753 and Guan et al., Chem. Eur. J. 2010, 16, 13343 (all incorporated herein by reference).
[0582] Example 46. Synthesize UDP-GalNPyrN3((12), where A is N3, R 2 (for H)
[0583] UDP-GalNPyrN3, a sugar derivative nucleotide of formula (3), was prepared according to the procedures disclosed in, for example, WO 2015 / 112013 (SynAffix BV); Pouilly et al., ACS Chem. Biol. 2012, 7, 753 and Guan et al., Chem. Eur. J. 2010, 16, 13343 (all incorporated herein by reference).
[0584] Glycosyl transfer of galactose derivatives (e.g., sucrase) using GalNAcT (general procedure)
[0585] Galactose derivatives (e.g., sulfur-containing sugar 21) were enzymatically introduced onto IgG using GalNAc transferase or its mutants. Deglycosylated IgG (prepared as described above, 10 mg / mL) was incubated with modified UDP-galactose derivatives (e.g., thio-modified sugar-UDP derivatives) (2 mM) and GalNAcT (0.2 mg / mL) at 30°C for 16 h in 10 mM MnCl2 and 50 mM Tris-HCl pH 6.0. Functionalized IgG (e.g., thio-functionalized trastuzumab) was incubated with protein A agarose (40 μL / mg IgG) at room temperature for 1 h. The protein A agarose was washed three times with TBS (pH 6.0), and the IgG was eluted with 100 mM glycine-HCl pH 2.5. The eluted IgG was neutralized with 1M Tris-HCl pH 7.0, concentrated, and washed with Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore) with 50mM Tris-HCl pH 6.0 to a concentration of 15-20 mg / mL.
[0586] Example 47. Trastuzumab (GalNProSH)2 was prepared by transferring the glycosyl group of UDP-GalNProSH ((21), where t=2) to deglycosylated trastuzumab using CeGalNAcT.
[0587] The trimmed trastuzumab was subjected to glycosyltransferase using UDP-GalNProSH (2 mM) and CeGalNAcT (1 mg / mL). After incubation overnight, small samples were transferred to a fabricator. TM Spectroscopic analysis was performed after digestion (50 U, in 10 μL PBS pH 6.6), followed by washing with MiliQ using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore). Complete conversion of the starting material to two products was observed. The major product (24387 Da, expected mass 24388) corresponded to deglycosylated trastuzumab + GalNProSH (trastuzumab-(GalNProSH)2), while the minor product (25037 Da, expected mass 25038) corresponded to deglycosylated trastuzumab + GalNProS-UDPGalNProS disulfide. The ratio of the two products was approximately 60:40.
[0588] Example 48. Trastuzumab (GalNProSH)2 was prepared by transferring the glycosyl group of UDP-GalNProSH ((21), where t=2) to deglycosylated trastuzumab using TnGalNAcT.
[0589] The trimmed trastuzumab was subjected to glycosyltransferase using UDP-GalNProSH (2 mM) and TnGalNAcT (0.2 mg / mL). After incubation overnight, small samples were transferred to Fabricator. TM Spectroscopic analysis was performed after digestion (50 U, in 10 μL PBS pH 6.6), followed by washing with MiliQ using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore). Complete conversion of the starting material to two products was observed. The major product (24387 Da, expected mass 24388) corresponded to deglycosylated trastuzumab + GalNProSH (trastuzumab-(GalNProSH)2), while the minor product (25037 Da, expected mass 25038) corresponded to deglycosylated trastuzumab + GalNProS-UDPGalNProS disulfide. The ratio of the two products was approximately 60:40.
[0590] Example 49. Trastuzumab (GalNProSH)2 was prepared by transferring the glycosyl group of UDP-GalNProSH ((21), where t=2) to deglycosylated trastuzumab using AsGalNAcT.
[0591] Trimmed trastuzumab was subjected to a glycosyltransfer protocol using UDP-GalNProSH (2 mM) and AsGalNAcT (0.2 mg / mL). After overnight incubation, small samples were reduced with DTT and subsequently analyzed by MS, which showed the formation of a major product (49755 Da, 95% of total heavy chain) resulting from the transfer of GalNProSH to the core GlcNAc (Fuc)-substituted trastuzumab.
[0592] Example 50. By using TnGalNAcT to convert UDP-GalNPyrN3((12), where A is N3, R 2 Trastuzumab (GalPyrN3) was prepared by converting H) glycosyltransferase to deglycosylated trastuzumab.
[0593] According to the general scheme of glycosyl transfer, in UDP-GalNPyrN3((12), where A is N3, R 2 Trimmed trastuzumab was treated with TnGalNAcT (0.5 mg / mL) in the presence of H (4 mM). After incubation overnight, small samples were processed using Fabricator. TM Digestion (50 U, in 10 μL PBS pH 6.6) was followed by spectral analysis and washing with MiliQ using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore). Approximately 50% conversion was observed. The resulting product (24445 Da, expected mass 24445) was generated by the transfer of GalNPyrN3 to the core GlcNAc (Fuc)-substituted trastuzumab.
[0594] Example 51. Trastuzumab (F2-GalNAz) was prepared by transferring the glycosyl group of UDP-F2-GalNAz (18) to deglycosylated trastuzumab using TnGalNAcT.
[0595] Following the general protocol for glycosyltransferase, the trimmed trastuzumab was treated with TnGalNAcT (wt), TnGalNAcT (W336V), TnGalNAcT (W336F), TnGalNAcT (E339A), or TnGalNAcT (L302A) at a concentration of 0.25 mg / mL in the presence of UDP-F2-GalNAz (18, 1 mM). After overnight incubation, small samples were processed using Fabricator.TM Spectroscopic analysis was performed after digestion (50 U, in 10 μL PBS pH 6.6), followed by washing with MiliQ using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore). MS analysis showed the formation of a major product (MW = 24420 Da) of trastuzumab substituted by F2-GalNAz transfer to the core GlcNAc (Fuc) and a minor product (MW = 24273 Da) of trastuzumab substituted by F2-GalNAz transfer to the core GlcNAc. The cumulative conversion rates of the major and minor products observed for TnGalNAcT (wt), TnGalNAcT (W336V), TnGalNAcT (W336F), TnGalNAcT (E339A), or TnGalNAcT (L302A) were 95%, 27%, 24%, 5%, and 95%, respectively.
[0596] Example 52. Trastuzumab (GalBzN3) was prepared by transferring the glycosyl group of UDP-GalNBzN3 ((23), where X is H) to deglycosylated trastuzumab using TnGalNAcT.
[0597] Following a general protocol for glycosyltransferase, the trimmed trastuzumab was treated with TnGalNAcT (0.7 mg / mL) in the presence of UDP-GalNBzN3 (4 mM) ((23), where X is H). After overnight incubation, small samples were processed using Fabricator. TM Digestion (50 U, in 10 μL PBS pH 6.6) was followed by spectral analysis and washing with MiliQ using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore). Approximately 70% conversion was observed. The resulting product (24444 Da, expected mass 24443) was generated by the transfer of GalNBzN3 to the core GlcNAc (Fuc)-substituted trastuzumab.
[0598] Example 53. Trastuzumab (F2-GalNBzN3)2 was prepared by transferring the glycosyl group of UDP-F2-GalNBzN3(24) to deglycosylated trastuzumab using TnGalNAcT.
[0599] Following the general protocol for glycosyltransferase, trimmed trastuzumab was treated with TnGalNAcT (wt), TnGalNAcT (W336H), TnGalNAcT (I299M), TnGalNAcT (L302A), or TnGalNAcT (L302G) at a concentration of 0.5 mg / mL in the presence of UDP-F2-GalNBzN3 (24, 1 mM). After overnight incubation, small samples were processed using a fabricator. TM Digestion (50 U, in 10 μL PBS pH 6.6) followed by washing with MiliQ using Amicon Ultra-0.5, Ultracel-10 Membrane (Millipore). MS analysis showed the formation of a major product (MW = 24479 Da) of trastuzumab substituted with F2-GalNBzN3 transferred to the core GlcNAc (Fuc) and a minor product (MW = 24332 Da) of trastuzumab substituted with F2-GalNBzN3 transferred to the core GlcNAc. The cumulative conversion rates of the major and minor products observed for TnGalNAcT(wt), TnGalNAcT(W336H), TnGalNAcT(I299M), TnGalNAcT(L302A), or TnGalNAcT(L302G) were 14%, 62%, 26%, 37%, and 14%, respectively.
[0600] Example 54. Trastuzumab (GalNProN3)2 was prepared by transferring the glycosyl group of UDP-GalNPropN3 ((31), where U is [CH2]2, A = N3) to deglycosylated trastuzumab using AsGalNAcT.
[0601] Following a general protocol for glycosyltransfer, trimmed trastuzumab was treated with AsGalNAcT at a concentration of 0.2 mg / mL in the presence of UDP-GalNPropN3 ((31), where U is [CH2]2, A = N3, 0.7 mM). After overnight incubation, small samples were reduced with DTT and subsequently analyzed by MS, revealing the formation of a major product (49759 Da, 70% of total heavy chain) resulting from the transfer of trastuzumab from GalNPropN3 to the core GlcNAc (Fuc).
[0602] Example 55. Trastuzumab (GalNButN3)2 was prepared by transferring the glycosyl group of UDP-GalNButN3 ((31), where U is [CH2]3 and A = N3) to deglycosylated trastuzumab using AsGalNAcT.
[0603] Following a general glycosyltransfer protocol, trimmed trastuzumab was treated with AsGalNAcT at a concentration of 0.2 mg / mL in the presence of UDP-GalNButN3 ((31), where U is [CH2]3, A = N3, 0.7 mM). After overnight incubation, small samples were reduced with DTT and subsequently analyzed by MS, revealing the formation of a major product (49772 Da, 50% of total heavy chain) resulting from the transfer of GalNButN3 to the core GlcNAc (Fuc)-substituted trastuzumab. sequence list <110> Sina Fox Corporation <120> Methods of modifying glycoproteins with β-(1,4)-N-acetylgalactosamine transferase or its mutants <130> P6052334PCT <150> EP14179713.4 <151> 2014-08-04 <160> 89 <170> PatentIn version 3.3 <210> 1 <211> 402 <212> PRT <213> Artificial sequence <220> <223> Bovine GalT Y289L mutant <400> 1 Met Lys Phe Arg Glu Pro Leu Leu Gly Gly Ser Ala Ala Met Pro Gly 1 5 10 15 Ala Ser Leu Gln Arg Ala Cys Arg Leu Leu Val Ala Val Cys Ala Leu 20 25 30 His Leu Gly Val Thr Leu Val Tyr Tyr Leu Ala Gly Arg Asp Leu Arg 35 40 45 Arg Leu Pro Gln Leu Val Gly Val His Pro Pro Leu Gln Gly Ser Ser 50 55 60 His Gly Ala Ala Ala Ile Gly Gln Pro Ser Gly Glu Leu Arg Leu Arg 65 70 75 80 Gly Val Ala Pro Pro Pro Pro Leu Gln Asn Ser Ser Lys Pro Arg Ser 85 90 95 Arg Ala Pro Ser Asn Leu Asp Ala Tyr Ser His Pro Gly Pro Gly Pro 100 105 110 Gly Pro Gly Ser Asn Leu Thr Ser Ala Pro Val Pro Ser Thr Thr Thr 115 120 125 Arg Ser Leu Thr Ala Cys Pro Glu Glu Ser Pro Leu Leu Val Gly Pro 130 135 140 Met Leu Ile Glu Phe Asn Ile Pro Val Asp Leu Lys Leu Ile Glu Gln 145 150 155 160 Gln Asn Pro Lys Val Lys Leu Gly Gly Arg Tyr Thr Pro Met Asp Cys 165 170 175 Ile Ser Pro His Lys Val Ala Ile Ile Ile Leu Phe Arg Asn Arg Gln 180 185 190 Glu His Leu Lys Tyr Trp Leu Tyr Tyr Leu His Pro Met Val Gln Arg 195 200 205 Gln Gln Leu Asp Tyr Gly Ile Tyr Val Ile Asn Gln Ala Gly Glu Ser 210 215 220 Met Phe Asn Arg Ala Lys Leu Leu Asn Val Gly Phe Lys Glu Ala Leu 225 230 235 240 Lys Asp Tyr Asp Tyr Asn Cys Phe Val Phe Ser Asp Val Asp Leu Ile 245 250 255 Pro Met Asn Asp His Asn Thr Tyr Arg Cys Phe Ser Gln Pro Arg His 260 265 270 Ile Ser Val Ala Met Asp Lys Phe Gly Phe Ser Leu Pro Tyr Val Gln 275 280 285 Leu Phe Gly Gly Val Ser Ala Leu Ser Lys Gln Gln Phe Leu Ser Ile 290 295 300 Asn Gly Phe Pro Asn Asn Tyr Trp Gly Trp Gly Gly Glu Asp Asp Asp 305 310 315 320 Ile Tyr Asn Arg Leu Ala Phe Arg Gly Met Ser Val Ser Arg Pro Asn 325 330 335 Ala Val Ile Gly Lys Cys Arg Met Ile Arg His Ser Arg Asp Lys Lys 340 345 350 Asn Glu Pro Asn Pro Gln Arg Phe Asp Arg Ile Ala His Thr Lys Glu 355 360 365 Thr Met Leu Ser Asp Gly Leu Asn Ser Leu Thr Tyr Met Val Leu Glu 370 375 380 Val Gln Arg Tyr Pro Leu Tyr Thr Lys Ile Thr Val Asp Ile Gly Thr 385 390 395 400 Pro Serum <210> 2 <211> 383 <212> PRT <213> Caenorhabditis elegans <400> 2 Met Ala Phe Arg His Leu Ala Val Ala Arg Leu Lys Ser Leu Leu Val 1 5 10 15 Leu Cys Ala Val Leu Leu Leu Val His Ala Met Ile Tyr Lys Ile Pro 20 25 30 Ser Leu Tyr Glu Asn Leu Thr Ile Gly Ser Ser Thr Leu Ile Ala Asp 35 40 45 Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser Thr Ser Asp 50 55 60 Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile Ser Glu Val 65 70 75 80 Asn Gln Thr Ser Phe Met Glu Asp Ile Arg Pro Ile Leu Phe Pro Asp 85 90 95 Asn Gln Thr Leu Gln Phe Cys Asn Gln Thr Pro Pro His Leu Val Gly 100 105 110 Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Lys Thr Leu Glu Lys 115 120 125 Ile Tyr Pro Asp Thr His Ala Gly Gly His Gly Met Pro Lys Asp Cys 130 135 140 Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg Asp Arg Glu 145 150 155 160 Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu Leu Ala Lys 165 170 175 Gln Gln Leu Asp Tyr Ala Ile Phe Ile Val Glu Gln Val Ala Asn Gln 180 185 190 Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp Val Ala Ser 195 200 205 Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val Asp Leu Leu 210 215 220 Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln Pro Arg His 225 230 235 240 Met Ser Val Ala Ile Asp Lys Phe Asn Tyr Lys Leu Pro Tyr Ser Ala 245 250 255 Ile Phe Gly Gly Ile Ser Ala Leu Thr Lys Asp His Leu Lys Lys Ile 260 265 270 Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu Asp Asp Asp 275 280 285 Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser Arg Tyr Pro 290 295 300 Thr Gln Ile Ala Arg Tyr Lys Met Ile Lys His Ser Thr Glu Ala Thr 305 310 315 320 Asn Pro Val Asn Lys Cys Arg Tyr Lys Ile Met Gly Gln Thr Lys Arg 325 330 335 Arg Trp Thr Arg Asp Gly Leu Ser Asn Leu Lys Tyr Lys Leu Val Asn 340 345 350 Leu Glu Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp Leu Leu Glu 355 360 365 Lys Asp Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr Cys Phe 370 375 380 <210> 3 <211> 383 <212> PRT <213> Ascaris suum <400> 3 Met Asn Ser Lys Leu Lys Leu Val Ile Val Leu Thr Leu Cys Val Ala 1 5 10 15 Ile Ile His Phe Leu Leu Ser Asp Cys Pro Ile Ser Pro Asp Tyr Ser 20 25 30 Phe Trp Ser Pro Ala Phe Ile Ile Ser Ala Pro Lys Thr Leu Thr Thr 35 40 45 Leu Gln Pro Phe Ser Gln Ser Thr Ser Thr Asn Asp Leu Ala Val Ser 50 55 60 Ala Leu Glu Ser Val Glu Phe Ser Met Leu Asp Asn Ser Ser Ile Leu 65 70 75 80 His Ala Ser Asp Asn Trp Thr Asn Asp Glu Leu Val Met Arg Ala Gln 85 90 95 Asn Glu Asn Leu Gln Leu Cys Pro Met Thr Pro Pro Ala Leu Val Gly 100 105 110 Pro Ile Lys Val Trp Met Asp Ala Pro Ser Phe Ala Glu Leu Glu Arg 115 120 125 Leu Tyr Pro Phe Leu Glu Pro Gly Gly His Gly Met Pro Thr Ala Cys 130 135 140 Arg Ala Arg His Arg Val Ala Ile Val Val Pro Tyr Arg Asp Arg Glu 145 150 155 160 Ser His Leu Arg Thr Phe Leu His Asn Leu His Ser Leu Leu Thr Lys 165 170 175 Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Thr Ala Asn Glu 180 185 190 Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Ala Glu Ala Ile 195 200 205 Arg Leu Tyr Asp Trp Arg Cys Phe Ile Phe His Asp Val Asp Leu Leu 210 215 220 Pro Glu Asp Asp Arg Asn Leu Tyr Ser Cys Pro Asp Glu Pro Arg His 225 230 235 240 Met Ser Val Ala Val Asp Lys Phe Asn Tyr Lys Leu Pro Tyr Gly Ser 245 250 255 Ile Phe Gly Gly Ile Ser Ala Leu Thr Arg Glu Gln Phe Glu Gly Ile 260 265 270 Asn Gly Phe Ser Asn Asp Tyr Trp Gly Trp Gly Gly Glu Asp Asp Asp 275 280 285 Leu Ser Thr Arg Val Thr Leu Ala Gly Tyr Lys Ile Ser Arg Tyr Pro 290 295 300 Ala Glu Ile Ala Arg Tyr Lys Met Ile Lys His Asn Ser Glu Lys Lys 305 310 315 320 Asn Pro Val Asn Arg Cys Arg Tyr Lys Leu Met Ser Ala Thr Lys Ser 325 330 335 Arg Trp Arg Asn Asp Gly Leu Ser Ser Leu Ser Tyr Asp Leu Ile Ser 340 345 350 Leu Gly Arg Leu Pro Leu Tyr Thr His Ile Lys Val Asp Leu Leu Glu 355 360 365 Lys Gln Ser Arg Arg Tyr Leu Arg Thr His Gly Phe Pro Thr Cys 370 375 380 <210> 4 <211> 421 <212> PRT <213> Trichoplusia ni <400> 4 Met Gly Gly Arg Ala Thr Arg Ala Leu Arg Leu Leu Leu Leu Leu Val 1 5 10 15 Leu Ala Leu Ala Ala Val Glu Tyr Leu Phe Gly Ser Ile Leu Asp Ala 20 25 30 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 35 40 45 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 50 55 60 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 65 70 75 80 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 85 90 95 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 100 105 110 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 115 120 125 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 130 135 140 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 145 150 155 160 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 165 170 175 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 180 185 190 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 195 200 205 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 210 215 220 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 225 230 235 240 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 245 250 255 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 260 265 270 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 275 280 285 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 290 295 300 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 305 310 315 320 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 325 330 335 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 340 345 350 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 355 360 365 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 370 375 380 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 385 390 395 400 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 405 410 415 Ile Asp Glu Arg Ser 420 <210> 5 <211> 403 <212> PRT <213> Drosophila melanogaster <400> 5 Met Tyr Leu Phe Thr Lys Ala Asn Leu Ile Arg Phe Leu Ala Gly Ala 1 5 10 15 Ile Cys Leu Leu Leu Val Leu Asn Phe Val Gly Phe Arg Ser Asp Gly 20 25 30 Gly Ser Ala Thr Ser Leu Ser Lys Leu Ser Ile Arg Arg Val His Lys 35 40 45 Tyr Ala His Ile Tyr Gly Asn Ala Ser Ser Asp Gly Ala Gly Gly Ser 50 55 60 Glu Ala Ser Arg Leu Pro Ala Ser Pro Leu Ala Leu Ser Lys Asp Arg 65 70 75 80 Glu Arg Asp Gln Glu Leu Asn Gly Gly Pro Asn Ser Thr Ile Arg Thr 85 90 95 Val Ile Ala Thr Ala Asn Phe Thr Ser Ile Pro Gln Asp Leu Thr Arg 100 105 110 Phe Leu Leu Gly Thr Lys Lys Phe Leu Pro Pro Arg Gln Lys Ser Thr 115 120 125 Ser Ala Leu Leu Ala Asn Cys Thr Asp Pro Asp Pro Arg Asp Gly Gly 130 135 140 Pro Ile Thr Pro Asn Thr Thr Leu Glu Ser Leu Asp Val Ile Glu Ala 145 150 155 160 Glu Leu Gly Pro Leu Leu Arg Pro Gly Gly Ala Phe Glu Pro Glu Asn 165 170 175 Cys Asn Ala Gln His His Val Ala Ile Val Val Pro Phe Arg Asp Arg 180 185 190 Tyr Ala His Leu Leu Leu Phe Leu Arg Asn Ile His Pro Phe Leu Met 195 200 205 Lys Gln Arg Ile Ala Tyr Arg Ile Phe Ile Val Glu Gln Thr Asn Gly 210 215 220 Lys Pro Phe Asn Arg Ala Ala Met Met Asn Ile Gly Tyr Leu Glu Ala 225 230 235 240 Leu Lys Leu Tyr Gln Trp Asp Cys Phe Ile Phe His Asp Val Asp Leu 245 250 255 Leu Pro Leu Asp Asp Arg Asn Leu Tyr Asn Cys Pro Arg Gln Pro Arg 260 265 270 His Met Ser Val Ala Ile Asp Thr Leu Asn Phe Arg Leu Pro Tyr Arg 275 280 285 Ser Ile Phe Gly Gly Val Ser Ala Met Thr Arg Glu His Phe Gln Ala 290 295 300 Val Asn Gly Phe Ser Asn Ser Phe Phe Gly Trp Gly Gly Glu Asp Asp 305 310 315 320 Asp Met Ser Asn Arg Leu Lys His Ala Asn Leu Phe Ile Ser Arg Tyr 325 330 335 Pro Val Asn Ile Ala Arg Tyr Lys Met Leu Lys His Gln Lys Glu Lys 340 345 350 Ala Asn Pro Lys Arg Tyr Glu Asn Leu Gln Asn Gly Met Ser Lys Ile 355 360 365 Glu Gln Asp Gly Ile Asn Ser Ile Lys Tyr Ser Ile Tyr Ser Ile Lys 370 375 380 Gln Phe Pro Thr Phe Thr Trp Tyr Leu Ala Glu Leu Lys Asn Ser Glu 385 390 395 400 Arg Lys Ser <210> 6 <211> 354 <212> PRT <213> Artificial sequence <220> <223> CeGalNAcT(30‑383) <400> 6 Lys Ile Pro Ser Leu Tyr Glu Asn Leu Thr Ile Gly Ser Ser Thr Leu 1 5 10 15 Ile Ala Asp Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser 20 25 30 Thr Ser Asp Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile 35 40 45 Ser Glu Val Asn Gln Thr Ser Phe Met Glu Asp Ile Arg Pro Ile Leu 50 55 60 Phe Pro Asp Asn Gln Thr Leu Gln Phe Cys Asn Gln Thr Pro Pro His 65 70 75 80 Leu Val Gly Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Lys Thr 85 90 95 Leu Glu Lys Ile Tyr Pro Asp Thr His Ala Gly Gly His Gly Met Pro 100 105 110 Lys Asp Cys Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu 130 135 140 Leu Ala Lys Gln Gln Leu Asp Tyr Ala Ile Phe Ile Val Glu Gln Val 145 150 155 160 Ala Asn Gln Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp 165 170 175 Val Ala Ser Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln 195 200 205 Pro Arg His Met Ser Val Ala Ile Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Ser Ala Ile Phe Gly Gly Ile Ser Ala Leu Thr Lys Asp His Leu 225 230 235 240 Lys Lys Ile Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser 260 265 270 Arg Tyr Pro Thr Gln Ile Ala Arg Tyr Lys Met Ile Lys His Ser Thr 275 280 285 Glu Ala Thr Asn Pro Val Asn Lys Cys Arg Tyr Lys Ile Met Gly Gln 290 295 300 Thr Lys Arg Arg Trp Thr Arg Asp Gly Leu Ser Asn Leu Lys Tyr Lys 305 310 315 320 Leu Val Asn Leu Glu Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp 325 330 335 Leu Leu Glu Lys Asp Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr 340 345 350 CysPhe <210> 7 <211> 354 <212> PRT <213> artificial sequence <220> <223> AsGalNAcT (30‑383) <400> 7 Asp Tyr Ser Phe Trp Ser Pro Ala Phe Ile Ile Ser Ala Pro Lys Thr 1 5 10 15 Leu Thr Thr Leu Gln Pro Phe Ser Gln Ser Thr Ser Thr Asn Asp Leu 20 25 30 Ala Val Ser Ala Leu Glu Ser Val Glu Phe Ser Met Leu Asp Asn Ser 35 40 45 Ser Ile Leu His Ala Ser Asp Asn Trp Thr Asn Asp Glu Leu Val Met 50 55 60 Arg Ala Gln Asn Glu Asn Leu Gln Leu Cys Pro Met Thr Pro Pro Ala 65 70 75 80 Leu Val Gly Pro Ile Lys Val Trp Met Asp Ala Pro Ser Phe Ala Glu 85 90 95 Leu Glu Arg Leu Tyr Pro Phe Leu Glu Pro Gly Gly His Gly Met Pro 100 105 110 Thr Ala Cys Arg Ala Arg His Arg Val Ala Ile Val Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ser His Leu Arg Thr Phe Leu His Asn Leu His Ser Leu 130 135 140 Leu Thr Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Thr 145 150 155 160 Ala Asn Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Ala 165 170 175 Glu Ala Ile Arg Leu Tyr Asp Trp Arg Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Ser Cys Pro Asp Glu 195 200 205 Pro Arg His Met Ser Val Ala Val Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Gly Ser Ile Phe Gly Gly Ile Ser Ala Leu Thr Arg Glu Gln Phe 225 230 235 240 Glu Gly Ile Asn Gly Phe Ser Asn Asp Tyr Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ser Thr Arg Val Thr Leu Ala Gly Tyr Lys Ile Ser 260 265 270 Arg Tyr Pro Ala Glu Ile Ala Arg Tyr Lys Met Ile Lys His Asn Ser 275 280 285 Glu Lys Lys Asn Pro Val Asn Arg Cys Arg Tyr Lys Leu Met Ser Ala 290 295 300 Thr Lys Ser Arg Trp Arg Asn Asp Gly Leu Ser Ser Leu Ser Tyr Asp 305 310 315 320 Leu Ile Ser Leu Gly Arg Leu Pro Leu Tyr Thr His Ile Lys Val Asp 325 330 335 Leu Leu Glu Lys Gln Ser Arg Arg Tyr Leu Arg Thr His Gly Phe Pro 340 345 350 Thr Cys <210> 8 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421) <400> 8 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 9 <211> 357 <212> PRT <213> Artificial Sequence [[ID=十七]]<220> <223> DmGalNAcT (47‑403) <400> 9 His Lys Tyr Ala His Ile Tyr Gly Asn Ala Ser Ser Asp Gly Ala Gly 1 5 10 15 Gly Ser Glu Ala Ser Arg Leu Pro Ala Ser Pro Leu Ala Leu Ser Lys 20 25 30 Asp Arg Glu Arg Asp Gln Glu Leu Asn Gly Gly Pro Asn Ser Thr Ile 35 40 45 Arg Thr Val Ile Ala Thr Ala Asn Phe Thr Ser Ile Pro Gln Asp Leu 50 55 60 Thr Arg Phe Leu Leu Gly Thr Lys Lys Phe Leu Pro Pro Arg Gln Lys 65 70 75 80 Ser Thr Ser Ala Leu Leu Ala Asn Cys Thr Asp Pro Asp Pro Arg Asp<000254*9*>85 90 95 Gly Gly Pro Ile Thr Pro Asn Thr Thr Leu Glu Ser Leu Asp Val Ile 100 105 110 It should be noted that in the above translation, there is an error in line 17 where "十七" is used instead of the correct translation. It should be "<220>". This is a mistake in the original text provided for translation. The correct translation should be: Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 9 <211> 357 <212> PRT <213> Artificial Sequence <220> <223> DmGalNAcT (47‑403) <400> 9 His Lys Tyr Ala His Ile Tyr Gly Asn Ala Ser Ser Asp Gly Ala Gly 1 5 10 15 Gly Ser Glu Ala Ser Arg Leu Pro Ala Ser Pro Leu Ala Leu Ser Lys 20 25 30 <000254*2*> Asp Arg Glu Arg Asp Gln Glu Leu Asn Gly Gly Pro Asn Ser Thr Ile Arg Thr Val Ile Ala Thr Ala Asn Phe Thr Ser Ile Pro Gln Asp Leu 50 55 60 Thr Arg Phe Leu Leu Gly Thr Lys Lys Phe Leu Pro Pro Arg Gln Lys 65 70 75 80 Ser Thr Ser Ala Leu Leu Ala Asn Cys Thr Asp Pro Asp Pro Arg Asp 85 90 95 Gly Gly Pro Ile Thr Pro Asn Thr Thr Leu Glu Ser Leu Asp Val Ile 100 105 110 Glu Ala Glu Leu Gly Pro Leu Leu Arg Pro Gly Gly Ala Phe Glu Pro 115 120 125 Glu Asn Cys Asn Ala Gln His His Val Ala Ile Val Val Pro Phe Arg 130 135 140 Asp Arg Tyr Ala His Leu Leu Leu Phe Leu Arg Asn Ile His Pro Phe 145 150 155 160 Leu Met Lys Gln Arg Ile Ala Tyr Arg Ile Phe Ile Val Glu Gln Thr 165 170 175 Asn Gly Lys Pro Phe Asn Arg Ala Ala Met Met Asn Ile Gly Tyr Leu 180 185 190 Glu Ala Leu Lys Leu Tyr Gln Trp Asp Cys Phe Ile Phe His Asp Val 195 200 205 Asp Leu Leu Pro Leu Asp Asp Arg Asn Leu Tyr Asn Cys Pro Arg Gln 210 215 220 Pro Arg His Met Ser Val Ala Ile Asp Thr Leu Asn Phe Arg Leu Pro 225 230 235 240 Tyr Arg Ser Ile Phe Gly Gly Val Ser Ala Met Thr Arg Glu His Phe 245 250 255 Gln Ala Val Asn Gly Phe Ser Asn Ser Phe Phe Gly Trp Gly Gly Glu 260 265 270 Asp Asp Asp Met Ser Asn Arg Leu Lys His Ala Asn Leu Phe Ile Ser 275 280 285 Arg Tyr Pro Val Asn Ile Ala Arg Tyr Lys Met Leu Lys His Gln Lys 290 295 300 Glu Lys Ala Asn Pro Lys Arg Tyr Glu Asn Leu Gln Asn Gly Met Ser 305 310 315 320 Lys Ile Glu Gln Asp Gly Ile Asn Ser Ile Lys Tyr Ser Ile Tyr Ser 325 330 335 Ile Lys Gln Phe Pro Thr Phe Thr Trp Tyr Leu Ala Glu Leu Lys Asn 340 345 350 Ser Glu Arg Lys Ser 355 <210> 10 <211> 354 <212> PRT <213> artificial sequence <220> <223> CeGalNacT(30‑383; I257L) <400> 10 Lys Ile Pro Ser Leu Tyr Glu Asn Leu Thr Ile Gly Ser Ser Thr Leu 1 5 10 15 Ile Ala Asp Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser 20 25 30 Thr Ser Asp Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile 35 40 45 Ser Glu Val Asn Gln Thr Ser Phe Met Glu Asp Ile Arg Pro Ile Leu 50 55 60 Phe Pro Asp Asn Gln Thr Leu Gln Phe Cys Asn Gln Thr Pro Pro His 65 70 75 80 Leu Val Gly Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Lys Thr 85 90 95 Leu Glu Lys Ile Tyr Pro Asp Thr His Ala Gly Gly His Gly Met Pro 100 105 110 Lys Asp Cys Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu 130 135 140 Leu Ala Lys Gln Gln Leu Asp Tyr Ala Ile Phe Ile Val Glu Gln Val 145 150 155 160 Ala Asn Gln Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp 165 170 175 Val Ala Ser Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln 195 200 205 Pro Arg His Met Ser Val Ala Ile Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Ser Ala Leu Phe Gly Gly Ile Ser Ala Leu Thr Lys Asp His Leu 225 230 235 240 Lys Lys Ile Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser 260 265 270 Arg Tyr Pro Thr Gln Ile Ala Arg Tyr Lys Met Ile Lys His Ser Thr 275 280 285 Glu Ala Thr Asn Pro Val Asn Lys Cys Arg Tyr Lys Ile Met Gly Gln 290 295 300 Thr Lys Arg Arg Trp Thr Arg Asp Gly Leu Ser Asn Leu Lys Tyr Lys 305 310 315 320 Leu Val Asn Leu Glu Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp 325 330 335 Leu Leu Glu Lys Asp Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr 340 345 350 Cys Phe <210> 11 <211> 354 <212> PRT <213> Artificial Sequence <220> <223> CeGalNAcT(30‑383; I257M) <400> 11 Lys Ile Pro Ser Leu Tyr Glu Asn Leu Thr Ile Gly Ser Ser Thr Leu 1 5 10 15 Ile Ala Asp Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser 20 25 30 Thr Ser Asp Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile 35 40 45 Ser Glu Val Asn Gln Thr Ser Phe Met Glu Asp Ile Arg Pro Ile Leu 50 55 60 Phe Pro Asp Asn Gln Thr Leu Gln Phe Cys Asn Gln Thr Pro Pro His 65 70 75 80 Leu Val Gly Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Lys Thr 85 90 95 Leu Glu Lys Ile Tyr Pro Asp Thr His Ala Gly Gly His Gly Met Pro 100 105 110 Lys Asp Cys Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu 130 135 140 Leu Ala Lys Gln Gln Leu Asp Tyr Ala Ile Phe Ile Val Glu Gln Val 145 150 155 160 Ala Asn Gln Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp 165 170 175 Val Ala Ser Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln 195 200 205 Pro Arg His Met Ser Val Ala Ile Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Ser Ala Met Phe Gly Gly Ile Ser Ala Leu Thr Lys Asp His Leu 225 230 235 240 Lys Lys Ile Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser 260 265 270 Arg Tyr Pro Thr Gln Ile Ala Arg Tyr Lys Met Ile Lys His Ser Thr 275 280 285 Glu Ala Thr Asn Pro Val Asn Lys Cys Arg Tyr Lys Ile Met Gly Gln 290 295 300 Thr Lys Arg Arg Trp Thr Arg Asp Gly Leu Ser Asn Leu Lys Tyr Lys 305 310 315 320 Leu Val Asn Leu Glu Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp 325 330 335 Leu Leu Glu Lys Asp Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr 340 345 350 CysPhe <210> 12 <211> 354 <212> PRT <213> artificial sequence <220> <223> CeGalNacT(30‑383; I257A) <400> 12 Lys Ile Pro Ser Leu Tyr Glu Asn Leu Thr Ile Gly Ser Ser Thr Leu 1 5 10 15 Ile Ala Asp Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser 20 25 30 Thr Ser Asp Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile 35 40 45 Ser Glu Val Asn Gln Thr Ser Phe Met Glu Asp Ile Arg Pro Ile Leu 50 55 60 Phe Pro Asp Asn Gln Thr Leu Gln Phe Cys Asn Gln Thr Pro Pro His 65 70 75 80 Leu Val Gly Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Lys Thr 85 90 95 Leu Glu Lys Ile Tyr Pro Asp Thr His Ala Gly Gly His Gly Met Pro 100 105 110 Lys Asp Cys Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu 130 135 140 Leu Ala Lys Gln Gln Leu Asp Tyr Ala Ile Phe Ile Val Glu Gln Val 145 150 155 160 Ala Asn Gln Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp 165 170 175 Val Ala Ser Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln 195 200 205 Pro Arg His Met Ser Val Ala Ile Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Ser Ala Ala Phe Gly Gly Ile Ser Ala Leu Thr Lys Asp His Leu 225 230 235 240 Lys Lys Ile Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser 260 265 270 Arg Tyr Pro Thr Gln Ile Ala Arg Tyr Lys Met Ile Lys His Ser Thr 275 280 285 Glu Ala Thr Asn Pro Val Asn Lys Cys Arg Tyr Lys Ile Met Gly Gln 290 295 300 Thr Lys Arg Arg Trp Thr Arg Asp Gly Leu Ser Asn Leu Lys Tyr Lys 305 310 315 320 Leu Val Asn Leu Glu Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp 325 330 335 Leu Leu Glu Lys Asp Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr 340 345 350 Cys Phe <210> 13 <2·11> 354 <212> PRT <213> Artificial sequence <220> <223> CeGalNacT(30‑383; M312H) <400> 13 Lys Ile Pro Ser Leu Tyr Glu Asn Leu Thr Ile Gly Ser Ser Thr Leu 1 5 10 15 Ile Ala Asp Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser 20 25 30 Thr Ser Asp Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile 35 40 45 Ser Glu Val Asn Gln Thr Ser Phe Met Glu Asp Ile Arg Pro Ile Leu 50 55 60 Phe Pro Asp Asn Gln Thr Leu Gln Phe Cys Asn Gln Thr Pro Pro His 65 70 75 80 Leu Val Gly Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Lys Thr 85 90 95 Leu Glu Lys Ile Tyr Pro Asp Thr His Ala Gly Gly His Gly Met Pro 100 105 110 Lys Asp Cys Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu 130 135 140 Leu Ala Lys Gln Gln Leu Asp Tyr Ala Ile Phe Ile Val Glu Gln Val 145 150 155 160 Ala Asn Gln Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp 165 170 175 Val Ala Ser Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln 195 200 205 Pro Arg His Met Ser Val Ala Ile Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Ser Ala Ile Phe Gly Gly Ile Ser Ala Leu Thr Lys Asp His Leu 225 230 235 240 Lys Lys Ile Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser 260 265 270 Arg Tyr Pro Thr Gln Ile Ala Arg Tyr Lys His Ile Lys His Ser Thr 275 280 285 Glu Ala Thr Asn Pro Val Asn Lys Cys Arg Tyr Lys Ile Met Gly Gln 290 295 300 Thr Lys Arg Arg Trp Thr Arg Asp Gly Leu Ser Asn Leu Lys Tyr Lys 305 310 315 320 Leu Val Asn Leu Glu Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp 325 330 335 Leu Leu Glu Lys Asp Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr 340 345 350 CysPhe <210> 14 <211> 360 <212> PRT <213> artificial sequence <220> <223> CeGalNAcT(30‑383)‑His <400> 14 Lys Ile Pro Ser Leu Tyr Glu Asn Leu Thr Ile Gly Ser Ser Thr Leu 1 5 10 15 Ile Ala Asp Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser 20 25 30 Thr Ser Asp Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile 35 40 45 Ser Glu Val Asn Gln Thr Ser Phe Met Glu Asp Ile Arg Pro Ile Leu 50 55 60 Phe Pro Asp Asn Gln Thr Leu Gln Phe Cys Asn Gln Thr Pro Pro His 65 70 75 80 Leu Val Gly Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Lys Thr 85 90 95 Leu Glu Lys Ile Tyr Pro Asp Thr His Ala Gly Gly His Gly Met Pro 100 105 110 Lys Asp Cys Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu 130 135 140 Leu Ala Lys Gln Gln Leu Asp Tyr Ala Ile Phe Ile Val Glu Gln Val 145 150 155 160 Ala Asn Gln Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp 165 170 175 Val Ala Ser Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln 195 200 205 Pro Arg His Met Ser Val Ala Ile Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Ser Ala Ile Phe Gly Gly Ile Ser Ala Leu Thr Lys Asp His Leu 225 230 235 240 Lys Lys Ile Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser 260 265 270 Arg Tyr Pro Thr Gln With Only Arg Tyr Lys With Lys Ser Thr 275 280 285 Glu Only Thr Asn Pro Val Asn Lys Cys Arg Tyr Ile Met Gly Gln 290,295,300 Thr Lys Arg Arg Trp Thr Arg Asp Gly Leu Ser Asn Leu Lys Tyr Lys 305 310 315 320 You Val Asn You Glu You Lys You You Tyr Thr Arg Ala Val Val Asp 325 330 335 Leu Leu Glu Lys Asp Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr 340 345 350 How Phe His His His His His His 355,360 <210> 15 <211> 383 <212> PRT <213> Caenorhabditis remains (Caenorhabditis remains) <400> 15 Met Ala Leu Arg His Leu Ala Val Ala Lys Leu Lys Thr Phe Phe Val 1 5 10 15 Leu Cys Wing Leu Leu Leu Val His Thr Met Ile Tyr Lys Wing Pro 20 25 30 Ser Leu Tyr Glu Asn Phe Ser Ile Gly Ser Ser Thr Leu Ile Ala Asp 35 40 45 Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser Thr Ser Tyr 50 55 60 Asp Leu Leu Asp Thr Trp Asn Ser Thr Phe Ser Pro Ile Ser Glu Val 65 70 75 80 Asn Gln Thr Ser Phe Leu Glu Asp Val Arg Pro Ile Leu Phe Thr Asp 85 90 95 Asn Gln Thr Lys Pro Phe Cys Asn Gln Thr Pro Pro His Leu Val Gly 100 105 110 Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Ala Thr Leu Glu Lys 115 120 125 Ile Tyr Pro Asp Val His Thr Gly Gly His Gly Ile Pro Asp Glu Cys 130 135 140 Ile Ala Arg His Arg Val Ala Val Ile Val Pro Tyr Arg Asp Arg Glu 145 150 155 160 Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu Leu Ala Lys 165 170 175 Gln Gln Leu Asp Tyr Ala Ile Ile Val Val Glu Gln Ile Val Asn Gln 180 185 190 Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp Val Ala Ser 195 200 205 Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val Asp Leu Leu 210 215 220 Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln Pro Arg His 225 230 235 240 Met Ser Val Ala Ile Asp Lys Phe Asp Tyr Lys Leu Pro Tyr Ser Thr 245 250 255 Ile Phe Gly Gly Ile Ser Ala Leu Thr Gln Glu His Val Lys Lys Ile 260 265 270 Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu Asp Asp Asp 275 280 285 Leu Ala Thr Arg Thr Ser Met Ala Gly Leu Lys Val Ser Arg Tyr Pro 290 295 300 Ala Gln Ile Ala Arg Tyr Lys Met Ile Lys His Ser Thr Glu Ala Thr 305 310 315 320 Asn Pro Val Asn Lys Cys Arg Tyr Lys Ile Met Gly Gln Thr Lys Arg 325 330 335 Arg Trp Thr Arg Asp Gly Leu Ser Ser Leu Lys Tyr Lys Leu Val Lys 340 345 350 Leu Asp Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp Leu Leu Glu 355 360 365 Lys Asp Cys Arg Arg Glu Leu Arg Lys Asp Phe Pro Thr Cys Phe 370 375 380 <210> 16 <211> 384 <212> PRT <213> Caenorhabditis briggsae <400> 16 Met Ala Phe Arg His Leu Ala Ser Ala Lys Leu Lys Thr Phe Phe Val 1 5 10 15 Leu Cys Ala Ala Leu Leu Leu Val His Ala Met Ile Tyr Lys Val Pro 20 25 30 Ser Leu Tyr Glu Asn Phe Ser Ile Gly Ser Ser Thr Leu Ile Ala Asp 35 40 45 Val Asp Ala Met Glu Ala Val Leu Gly Asn Thr Ala Ser Thr Ser Asp 50 55 60 Asp Pro Phe Asp Val Trp Asn Ser Thr Phe Ser Pro Ile Ser Glu Val 65 70 75 80 Asn Gln Thr Ala Phe Met Glu Asp Ile Arg Pro Ile Leu Phe Gly Asp 85 90 95 Ala Asn Glu Thr Arg Pro His Cys Asn Gln Thr Pro Pro His Leu Val 100 105 110 Gly Pro Ile Arg Val Phe Leu Asp Glu Pro Asp Phe Ala Thr Leu Glu 115 120 125 Lys Ile Tyr Pro Glu Thr His Pro Gly Gly His Gly Ile Pro Thr Glu 130 135 140 Cys Val Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg Asp Arg 145 150 155 160 Glu Ala His Leu Arg Ile Met Leu His Asn Leu His Ser Leu Leu Ala 165 170 175 Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Val Ala Asn 180 185 190 Gln Thr Phe Asn Arg Gly Lys Leu Met Asn Val Gly Tyr Asp Val Ala 195 200 205 Ser Arg Leu Tyr Pro Trp Gln Cys Phe Ile Phe His Asp Val Asp Leu 210 215 220 Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys Pro Ile Gln Pro Arg 225 230 235 240 His Met Ser Val Ala Ile Asp Lys Phe His Tyr Lys Leu Pro Tyr Ser 245 250 255 Gly Gly Served Only Thr Gln Glu His Val Lys Ala 260 265 270 Ile Asn Gly Phe Ser Asn Asp Phe Trp Gly Trp Gly Gly Glu Asp Asp 275 280 285 Asp Leu Alpha Thr Arg Thr Served Gln Ala Gly Leu Lys Val Served Arg Tyr 290,295,300 Pro Ala Gln Ile Ala Arg Tyr Lys Met Ile Lys His Serving Thr Glu Ala 305 310 315 320 Thr Asn Pro Val Asn Lys Cys Arg Tyr Lys With Gly Gln Thr Lys 325 330 335 Arg Arg Trp Lys Thr Asp Gly Leu Ser Ser Leu Lys Tyr Lys Leu Val 340 345 350 Lys Leu Glu Leu Lys Pro Leu Tyr Thr Arg Ala Val Val Asp Leu Leu 355 360 365 Glu Lys Glu Cys Arg Arg Glu Leu Arg Arg Asp Phe Pro Thr Cys Phe 370 375 380 <210> 17 <211> 464 <212> PRT <213> Bankruptcy (Wucheraria bankruptcy) <400> 17 Met Pro Ala Ala Gly Arg Phe Val Ile Ile Leu Leu Ile Phe Gly Ala 1 5 10 15 Ala Ala His Ile Phe Leu Gly Gly Gly Leu Ser Phe Ile Ser Asp Tyr 20 25 30 His Ile Trp Arg Pro Val Val Glu Ser Ser Arg Gln Glu Ile Val Leu 35 40 45 Val His Asn Ile Asp Asn Asn Ser Asp Gln Asn Ala Glu Lys Ile Ile 50 55 60 Ser Asn Asn Glu Thr Lys Phe His Leu Thr Ser Ala Thr Pro Ile Asp 65 70 75 80 Asn Leu Val Ser Ile His Ser Asn Phe Tyr Glu Leu Phe Ile Asn Gly 85 90 95 Leu Arg Phe Gly Lys Leu Thr Thr Val Tyr Pro Ile Ile Asn Gln Ser 100 105 110 Ile Asn Asn Gly Ser Thr Thr Asp Lys Ser Thr Glu Thr Tyr Ala Glu 115 120 125 Ser Val Tyr Phe Leu Lys Thr Asp Gly Asn Ile His Ser Asn Thr Leu 130 135 140 Leu Ser Thr Ile Thr Asp Ala Gln Ser Thr Arg Gln Leu Phe Gly Asn 145 150 155 160 Glu Thr Leu Ser Ala Cys Asn Val Ile Pro Ser Phe Gln Met Met His 165 170 175 Gln Asn Leu Ser Leu Val Asn Cys Pro Val Thr Pro Pro Gly Leu Val 180 185 190 Gly Pro Ile Lys Val Trp Tyr Asp Glu Pro Thr Phe Glu Glu Ile Glu 195 200 205 Arg Leu Asn Pro Asn Leu Glu Ala Gly Gly His Gly Lys Pro Glu Asn 210 215 220 Cys Leu Ser Arg His Arg Val Ala Val Ile Val Pro Tyr Arg Asp Arg 225 230 235 240 Glu Ala His Leu Arg Ile Leu Leu His Asn Leu His Ser Leu Leu Thr 245 250 255 Lys Gln Gln Leu Asp Tyr Gly Ile Phe Val Ile Glu Gln His Glu Asn 260 265 270 Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Val Glu Ala 275 280 285 Leu Lys Leu Tyr Asp Trp Gln Cys Phe Val Phe His Asp Val Asp Leu 290 295 300 Leu Ala Glu Asp Asp Arg Asn Ile Tyr Ser Cys Pro Asp Gln Pro Arg 305 310 315 320 His Met Ser Val Ala Val Asn Lys Phe Lys Tyr Lys Leu Pro Tyr Gly 325 330 335 Ser Ile Phe Gly Gly Val Ser Ala Ile Arg Thr Glu Gln Phe Ala Thr 340 345 350 Leu Asn Gly Phe Ser Asn Ser Tyr Trp Gly Trp Gly Gly Glu Asp Asp 355 360 365 Asp Leu Ser Met Arg Val Thr Ser Ala Gly Tyr Lys Ile Met Arg Tyr 370 375 380 Pro Ser Glu Ile Ala Arg Tyr Gln Met Val Gln His Lys Ser Glu Met 385 390 395 400 Lys Asn Pro Ile Asn Arg Cys Arg Tyr Asp Leu Leu Ala Lys Thr Lys 405 410 415 Val Arg Gln Gln Thr Asp Gly Ile Ser Ser Leu Lys Tyr Glu Cys Tyr 420 425 430 Asp Leu Gln Phe Phe Thr Leu Phe Thr His Ile Lys Val Lys Leu Phe 435 440 445 Glu Gln Glu Ser Lys Ala Gln Leu Arg Glu Glu Gly Phe Lys Arg Cys 450 455 460 <210> 18 <211> 291 <212> PRT <213> Loa loa <400> 18 Met Glu Arg Gln Asn Leu Ser Leu Val Asp Cys Pro Ile Ile Pro Pro 1 5 10 15 Gly Leu Val Gly Pro Ile Lys Val Trp Tyr Asp Glu Pro Thr Phe Glu 20 25 30 Glu Ile Glu Arg Leu Asn Pro Tyr Leu Glu Leu Gly Gly His Gly Lys 35 40 45 Pro Gly Ser Cys Leu Ser Arg His Arg Val Ala Ile Ile Val Pro Tyr 50 55 60 Arg Asp Arg Glu Ala His Leu Arg Ile Leu Leu His Asn Leu His Ser 65 70 75 80 Leu Leu Thr Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Ile Glu Gln 85 90 95 His Glu Asn Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr 100 105 110 Thr Glu Ala Met Lys Leu Tyr Asp Trp Gln Cys Phe Ile Phe His Asp 115 120 125 Val Asp Leu Leu Ala Glu Asp Asp Arg Asn Ile Tyr Ser Cys Pro Asp 130 135 140 Gln Pro Arg His Met Ser Val Ala Ile Asn Lys Phe Lys Tyr Arg Leu 145 150 155 160 Pro Tyr Gly Ser Ile Phe Gly Gly Val Ser Ala Ile Arg Thr Glu Gln 165 170 175 Phe Leu Lys Met Asn Gly Phe Ser Asn Ser Tyr Trp Gly Trp Gly Gly 180 185 190 Glu Asp Asp Asp Served With Arg Val Thr Served With Gly Tyr Lys Ile 195 200 205 Met Arg Tyr Pro Leu Glu Ile Ala Arg Tyr Gln Met Val Lys His Glu 210 215 220 Ser Glu Thr Lys Asn Pro Ile Asn Arg Cys Arg Tyr Asp Leu Leu Ala 225 230 235 240 Lys Thr Lys Val Arg Gln Gln Met Asp Gly Ile Ser Leu Lys Tyr 245 250 255 Glu Cys Tyr Asp Leu His Phe Leu Pro Leu Phe Thr His Ile Lys Val 260 265 270 Lys Leu Phe Glu Gln Glu Ser Lys Ala Gln Leu Arg Glu Glu Gly Phe 275 280 285 Lys Lys Cys 290 <210> 19 <211> 296 <212> PRT <213> Cerapachys biroi <400> 19 Met Pro Ile Arg Asn Leu Ala Gly Asn Gly Gly Thr Ala Arg Glu Leu 1 5 10 15 Pro Val Ala Asn Thr Thr Ser Asn Ala Thr Ile Pro Arg Cys Pro Leu 20 25 30 Ile Pro Pro Asn Leu Val Gly Pro Val Ala Val Ser Lys Ser Pro Pro 35 40 45 Pro Leu Ser Glu Met Glu Arg Ser Phe Val Glu Val Lys Ala Gly Gly 50 55 60 Lys Gly Arg Pro Ala Asp Cys Val Ala Arg His Arg Val Ala Ile Ile 65 70 75 80 Ile Pro Phe Arg Asp Arg Pro Gln His Leu Gln Thr Leu Leu Tyr Asn 85 90 95 Leu His Pro Ile Leu Leu Arg Gln Gln Ile Asp Tyr Gln Ile Phe Val 100 105 110 Ile Glu Gln Glu Gly Thr Gly Thr Phe Asn Arg Ala Met Leu Met Asn 115 120 125 Val Gly Tyr Val Glu Ala Leu Lys Glu Arg Ile Phe Asp Cys Phe Ile 130 135 140 Phe His Asp Val Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr 145 150 155 160 Cys Pro Glu Gln Pro Arg His Met Ser Val Ala Val Asp Lys Phe Lys 165 170 175 Tyr Arg Leu Pro Tyr Ala Asp Leu Phe Gly Gly Val Ser Ala Met Ser 180 185 190 Arg Glu His Phe Gln Leu Val Asn Gly Phe Ser Asn Val Phe Trp Gly 195 200 205 Trp Gly Gly Glu Asp Asp Asp Met Ala Asn Arg Ile Lys Ala His Gly 210 215 220 Leu His Ile Ser Arg Tyr Pro Ala Asn Val Ala Arg Tyr Lys Met Leu 225 230 235 240 Thr His Lys Lys Glu Lys Ala Asn Pro Lys Arg Tyr Glu Phe Leu Lys 245 250 255 Thr Gly Lys Lys Arg Phe Ser Thr Asp Gly Leu Ala Asn Leu Gln Tyr 6]260 265 270 Glu Leu Cys Asp Lys Arg Lys Pro Lys Leu Tyr Thr Trp Leu Leu Val 275 280 285 Arg Leu Thr Pro Pro Gln Pro Ser 290 295 <210> 20 <211> 422 <212> PRT <213> Dampwood termite (Zootermopsis nevadensis) <400> 20 Met Arg Cys Arg Cys Leu Ser Ala Trp Ser Arg Ile Thr Gln His Val 1 5 10 15 Pro Arg Gln Pro Cys Leu His Ile His Ser His Leu Cys Lys Val Val 20 25 30 Ile Val Leu Ala Val Leu Ile Ala Leu Gln Phe Leu Leu Thr Thr Ile 35 40 45 Phe Glu Ala Arg Gln Ile Glu Pro Leu Phe Thr Val Asn Phe Thr Tyr 50 55 60 Ser Gly Arg Arg Ser Arg Trp Gly Leu Ile Ser His Ser Arg Gly Leu 65 70 75 80 Leu Ser Pro Ser His Asn Ser Ser Phe Asn Gly Ser Met Arg Val Ser 85 90 95 Val Glu Arg Thr Leu Ser Pro Val Glu Asn Ile Ser Gly Glu Thr Lys 100 105 110 Asn Leu Ser Phe Leu His Thr His Glu Asn Ala Val Arg Asn Ala Ser 115 120 125 Ser Leu Val Leu Asn Ile Ser Leu Pro Ser Asp Leu Asn Pro Thr Thr 130 135 140 Ser Pro Ser Leu Thr Val Pro Phe Thr Gly Lys Ser Leu Cys Pro Pro 145 150 155 160 Ile Pro Pro Asn Leu Asn Gly Pro Ile Lys Val Leu Lys Asp Ser Pro 165 170 175 Ser Leu Glu Glu Leu Glu Lys Met Phe Pro Leu Leu Glu Pro Gly Gly 180 185 190 His Tyr His Pro Glu Glu Cys Gln Ala Arg Asp Arg Val Ala Ile Ile 195 200 205 Val Pro Tyr Arg Asp Arg Ala Glu His Leu Ser Thr Phe Leu Leu Asn 210 215 220 Leu His Pro Leu Leu Gln Arg Gln Gln Leu Asp Tyr Gly Met Phe Val 225 230 235 240 Ile Glu Gln Gly Gly Asp Gly Pro Phe Asn Arg Ala Met Leu Met Asn 245 250 255 Val Gly Phe Val Glu Ala Leu Lys Leu Tyr Ser Tyr Asp Cys Phe Ile 260 265 270 Phe His Asp Val Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr 275 280 285 Cys Pro Glu Gln Pro Arg His Met Ser Val Ala Val Asp Val Leu Lys 290 295 300 Tyr Lys Leu Pro Tyr Gln Ala Ile Phe Gly Gly Val Ser Ala Met Thr 305 310 315 320 Lys Thr Gln Phe Gln Lys Val Asn Gly Phe Ser Asn Leu Phe Trp Gly 325 330 335 Trp Gly Gly Glu Asp Asp Asp Met Ser Asn Arg Val Arg His His Gly 340 345 350 Tyr His Ile Ser Arg Tyr Pro Ala Asn Ile Ala Arg Tyr Lys Met Leu 355 360 365 Ala His Arg Lys Gln His Ala Asn Pro Lys Arg Tyr Glu Phe Leu Asn 370 375 380[[ID=十六]] [[ID=十七]]Thr Gly Arg Lys Arg Phe Lys Thr Asp Gly Leu Ser Asn Leu Gln Tyr[[ID=十八]] [[ID=十九]]385 390 395 400[[ID=二十]] [[ID=二十一]]Asp Arg Lys Glu Leu Asn Leu Gly Lys Leu Tyr Thr Arg Val Leu Val[[ID=二十二]] [[ID=二十三]]405 410 415[[ID=二十四]] [[ID=二十五]]Glu Leu Ala Thr Pro Ser[[ID=二十六]] [[ID=二十七]]420[[ID=二十八]] [[ID=二十九]]<210> 21[[ID=三十]]<(0003160)>[[ID=三十一]]<211> 295[[ID=三十二]] [[ID=三十三]]<212> PRT[[ID=三十四]] [[ID=三十五]]<213> Camponotus floridanus[[ID=三十六]] [[ID=三十七]]<400> 21[[ID=三十八]] [[ID=三十九]]Met Pro Thr Arg Asn Leu Val Gly Gly Gly Thr Ala Arg Glu Leu Pro[[ID=四十]] [[ID=四十一]]1 5 10 15[[ID=四十二]] [[ID=四十三]]Val Ala Asn Ala Thr Asn Asn Thr Thr Met Pro Arg Cys Pro Leu Ile[[ID=四十四]] [[ID=四十五]]20 25 30[[ID=四十六]] It should be noted that in the above translation, for the 7-digit tags like , , etc., they are kept exactly as in the original text as required. If there are any specific requirements or corrections regarding the biological terms or other content, further adjustments can be made accordingly.Pro Pro Asn Leu Val Gly Pro Met Val Val Ser Lys Ser Pro Pro Pro 35 40 45 Leu Ser Glu Met Glu Arg Ser Phe Val Glu Val Asn Ala Gly Gly Arg 50 55 60 Gly Arg Pro Ala Asp Cys Val Ala Arg His Arg Val Ala Ile Ile Ile 65 70 75 80 Pro Phe Arg Asp Arg Pro Gln His Leu Gln Thr Leu Leu Tyr Asn Leu 85 90 95 His Pro Ile Leu Leu Arg Gln Gln Ile Glu Tyr Gln Ile Phe Val Ile 100 105 110 Glu Gln Glu Gly Thr Gly Ala Phe Asn Arg Ala Met Leu Met Asn Val 115 120 125 Gly Tyr Val Glu Ala Leu Lys Glu Arg Thr Phe Asp Cys Phe Ile Phe 130 135 140 His Asp Val Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Thr Cys 145 150 155 160 Pro Glu Gln Pro Arg His Met Ser Val Ala Val Asp Lys Phe Lys Tyr 165 170 175 Arg Leu Pro Tyr Thr Asp Leu Phe Gly Gly Val Ser Ala Met Ser Arg 180 185 190 Glu His Phe Gln Leu Val Asn Gly Phe Ser Asn Val Phe Trp Gly Trp 195 200 205 Gly Gly Glu Asp Asp Asp Met Ala Asn Arg Ile Lys Ala His Gly Leu 210 215 220 His Ile Ser Arg Tyr Pro Ala Asn Val Ala Arg Tyr Lys Met Leu Thr 225 230 235 240 His Lys Lys Glu Lys Ala Asn Pro Lys Arg Tyr Glu Phe Leu Lys Thr 245 250 255 Gly Lys Lys Arg Phe Ser Thr Asp Gly Leu Ala Asn Leu Gln Tyr Glu 260 265 270 Leu Ser Asp Lys Arg Lys Pro Lys Leu Tyr Thr Trp Leu Leu Val Arg 275 280 285 Leu Thr Pro Pro Gln Pro Ser 290 295 <210> 22 <211> 310 <212> PRT <213> Pacific oyster (Crassostrea gigas) <400> 22 Met Asp Arg Gly Cys Lys Pro Met Arg Val Cys Ser Ser Ser Pro Ser 1 5 10 15 Asp Leu Val Gly Ser Leu Ala Thr Tyr Lys Glu Ala Pro Ser Tyr Lys 20 25 30 Glu Met Ile Lys Ile Tyr Pro Leu Val Arg Pro Gly Gly Leu Tyr Thr 35 40 45 Pro Pro Asp Cys Ile Ala Arg Glu Arg Val Ala Ile Ile Ile Pro Phe 50 55 60 Arg Asp Arg Glu Glu His Leu Arg Ile Leu Leu His Asn Leu His Pro 65 70 75 80 Met Leu Gln Arg Gln Gln Leu Asp Tyr Gly Ile Tyr Val Val Glu Gln 85 90 95 Glu Asn Gly Thr Gln Phe Asn Arg Ala Met Leu Met Asn Ile Gly Tyr 100 105 110 Ala Glu Ser Ile Lys Leu Tyr Asn Tyr Thr Cys Phe Ile Phe His Asp 115 120 125 Val Asp Leu Ile Pro Glu Asn Asp Arg Ile Met Tyr Asp Cys Arg Asp 130 135 140 Ser Pro Arg His Leu Ser Ser Ala Val Asp Lys Phe Lys Tyr Lys Leu 145 150 155 160 Pro Tyr Pro Gln Leu Phe Gly Gly Val Thr Ala Ile Lys Arg Ala His 165 170 175 Phe Glu Lys Val Asn Gly His Ser Asn Lys Phe Phe Gly Trp Gly Gly 180 185 190 Glu Asp Asp Asp Met Phe Arg Arg Leu Val Asn Asn Gly Phe Lys Ile 195 200 205 Ser Arg Tyr Gln Ala Ser Leu Ser Lys Tyr Lys Met Ile Lys His Leu 210 215 220 His Asp Ala Gly Asn Lys Ala Asn Lys Arg Arg His Leu Ile Lys 225 230 235 240 Thr Gly Lys Gly Arg Tyr Arg Arg Asp Gly Ile Asn Asn Leu His Tyr 245 250 255 Lys Lys Leu Gly Ile Glu Tyr Gln Tyr Leu His Thr Arg Ile Leu Val 260 265 270 There Is No Glue Thr Lys Will Met Thr Will Be Leu Leu Tyr Met Tyr 275 280 285 Ser Ser Thr Thr Val Tyr Ile Ile Val Asn Ile Thr Thr Ile Tyr Cys 290,295,300 Lys Ser Arg Asn Ile Arg 305 310 <210> 23 <211> 338 <212> PRT <213> Danaus plexippus <400> 23 Met Ala Lys Lys Leu Leu Thr Gln Gly Thr Glu Ser Val Thr Asn Tyr 1 5 10 15 Thr His Thr Thr Asn Ser Ser Asn Lys Asn Pro Ala Lys Glu Thr Phe 20 25 30 Asn Met Thr Lys Pro Asn Leu Ser Asp Asp Thr Ser Thr Pro Leu Leu 35 40 45 Ile Thr Lys Ile Met Glu Ser Ile Lys Asn Leu Val Thr Thr Glu Glu 50 55 60 Asp Phe Arg Asp Glu Pro Ser Leu Pro Leu Cys Asp Glu Met Pro Pro 65 70 75 80 Asp Leu Gly Pro Ile Ser Val Asn Lys Thr Glu Ile Glu Leu Asp Trp 85 90 95 Val Glu Lys Arg Tyr Pro Glu Val Arg Ser Gly Gly Ile Tyr Ser Ser 100 105 110 Ser Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr Arg 115 120 125 Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro Phe 130 135 140 Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Tyr Ile Ile Glu Gln Glu 145 150 155 160 Gly Thr Ser Glu Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe Val 165 170 175 Glu Ser Gln Arg Gln Arg Ser Trp Gln Cys Phe Ile Phe His Asp Ile 180 185 190 Asp Leu Leu Pro Leu Asp Ser Arg Asn Met Tyr Ser Cys Pro Lys Gln 195 200 205 Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu Asn Phe Arg Leu Pro 210 215 220 Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu Glu Gln Phe 225 230 235 240 Thr Lys Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Met Phe Tyr Arg Leu Lys Lys Met Asn Tyr His Ile Ala 260 265 270 Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp His Lys Lys 275 280 285 Ser Ala Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln Thr Ser Lys 290 295 300 Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu Val Ile Lys 305 310 315 320 Val Thr Ala Asn His Leu Tyr Thr His Ile Leu Val Asn Ile Asp Glu 325 330 335 Arg Ser <210> 24 <211> 941 <212> PRT <213> Artificial sequence <220> <223> HuGalNAcT (57-998) <400> 24 Arg Tyr Gly Ser Trp Arg Glu Leu Ala Lys Ala Leu Ala Ser Arg Asn 1 5 10 15 Ile Pro Ala Val Asp Pro His Leu Gln Phe Tyr His Pro Gln Arg Leu 20 25 30 Ser Leu Glu Asp His Asp Ile Asp Gln Gly Val Ser Ser Asn Ser Ser 35 40 45 Tyr Leu Lys Trp Asn Lys Pro Val Pro Trp Leu Ser Glu Phe Arg Gly 50 55 60 Arg Ala Asn Leu His Val Phe Glu Asp Trp Cys Gly Ser Ser Ile Gln 65 70 75 80 Gln Leu Arg Arg Asn Leu His Phe Pro Leu Tyr Pro His Ile Arg Thr 85 90 95 Thr Leu Arg Lys Leu Ala Val Ser Pro Lys Trp Thr Asn Tyr Gly Leu 100 105 110 Arg Ile Phe Gly Tyr Leu His Pro Phe Thr Asp Gly Lys Ile Gln Phe 115 120 125 Ala Ile Ala Ala Asp Asp Asn Ala Glu Phe Trp Leu Ser Leu Asp Asp 130 135 140 Gln Val Ser Gly Leu Gln Leu Leu Ala Ser Val Gly Lys Thr Gly Lys 145 150 155 160 Glu Trp Thr Ala Pro Gly Glu Phe Gly Lys Phe Arg Ser Gln Ile Ser 165 170 175 Lys Pro Val Ser Leu Ser Ala Ser His Arg Tyr Tyr Phe Glu Val Leu 180 185 190 His Lys Gln Asn Glu Glu Gly Thr Asp His Val Glu Val Ala Trp Arg 195 200 205 Arg Asn Asp Pro Gly Ala Lys Phe Thr Ile Ile Asp Ser Leu Ser Leu 210 215 220 Ser Leu Phe Thr Asn Glu Thr Phe Leu Gln Met Asp Glu Val Gly His 225 230 235 240 Ile Pro Gln Thr Ala Ala Ser His Val Asp Ser Ser Asn Ala Leu Pro 245 250 255 Arg Asp Glu Gln Pro Pro Ala Asp Met Leu Arg Pro Asp Pro Arg Asp 260 265 270 Thr Leu Tyr Arg Val Pro Leu Ile Pro Lys Ser His Leu Arg His Val 275 280 285 Leu Pro Asp Cys Pro Tyr Lys Pro Ser Tyr Leu Val Asp Gly Leu Pro 290 295 300 Leu Gln Arg Tyr Gln Gly Leu Arg Phe Val His Leu Ser Phe Val Tyr 305 310 315 320 Pro Asn Asp Tyr Thr Arg Leu Ser His Met Glu Thr His Asn Lys Cys 325 330 335 Phe Tyr Gln Glu Asn Ala Tyr Tyr Gln Asp Arg Phe Ser Phe Gln Glu 340 345 350 Tyr Ile Lys Ile Asp Gln Pro Glu Lys Gln Gly Leu Glu Gln Pro Gly 355 360 365 Phe Glu Glu Asn Leu Leu Glu Glu Ser Gln Tyr Gly Glu Val Ala Glu 370 375 380 Glu Thr Pro Ala Ser Asn Asn Gln Asn Ala Arg Met Leu Glu Gly Arg 385 390 395 400 Gln Thr Pro Ala Ser Thr Leu Glu Gln Asp Ala Thr Asp Tyr Arg Leu 405 410 415 Arg Ser Leu Arg Lys Leu Leu Ala Gln Pro Arg Glu Gly Leu Leu Ala 420 425 430 Pro Phe Ser Lys Arg Asn Ser Thr Ala Ser Phe Pro Gly Arg Thr Ser 435 440 445 His Ile Pro Val Gln Gln Pro Glu Lys Arg Lys Gln Lys Pro Ser Pro 450 455 460 Glu Pro Ser Gln Asp Ser Pro His Ser Asp Lys Trp Pro Pro Gly His 465 470 475 480 Pro Val Lys Asn Leu Pro Gln Met Arg Gly Pro Arg Pro Arg Pro Ala 485 490 495 Gly Asp Ser Pro Arg Lys Thr Gln Trp Leu Asn Gln Val Glu Ser Tyr 500 505 510 Ile Ala Glu Gln Arg Arg Gly Asp Arg Met Arg Pro Gln Ala Pro Gly 515 520 525 Arg Gly Trp His Gly Glu Glu Glu Val Val Ala Ala Ala Gly Gln Glu 530 535 540 Gly Gln Val Glu Gly Glu Glu Glu Gly Glu Glu Glu Glu Glu Glu Glu 545 550 555 560 Asp Met Ser Glu Val Phe Glu Tyr Val Pro Val Phe Asp Pro Val Val 565 570 575 Asn Trp Asp Gln Thr Phe Ser Ala Arg Asn Leu Asp Phe Gln Ala Leu 580 585 590 Arg Thr Asp Trp Ile Asp Leu Ser Cys Asn Thr Ser Gly Asn Leu Leu 595 600 605 Leu Pro Glu Gln Glu Ala Leu Glu Val Thr Arg Val Phe Leu Lys Lys 610 615 620 Leu Asn Gln Arg Ser Arg Gly Arg Tyr Gln Leu Gln Arg Ile Val Asn 625 630 635 640 Val Glu Lys Arg Gln Asp Gln Leu Arg Gly Gly Arg Tyr Leu Leu Glu 645 650 655 Leu Glu Leu Leu Glu Gln Gly Gln Arg Val Val Arg Leu Ser Glu Tyr 660 665 670 Val Ser Ala Arg Gly Trp Gln Gly Ile Asp Pro Ala Gly Gly Glu Glu 675 680 685 Val Glu Ala Arg Asn Leu Gln Gly Leu Val Trp Asp Pro His Asn Arg 690 695 700 Arg Arg Gln Val Leu Asn Thr Arg Ala Gln Glu Pro Lys Leu Cys Trp 705 710 715 720 Pro Gln Gly Phe Ser Trp Ser His Arg Ala Val Val His Phe Val Val 725 730 735 Pro Val Lys Asn Gln Ala Arg Trp Val Gln Gln Phe Ile Lys Asp Met 740 745 750 Glu Asn Leu Phe Gln Val Thr Gly Asp Pro His Phe Asn Ile Val Ile 755 760 765 Thr Asp Tyr Ser Ser Glu Asp Met Asp Val Glu Met Ala Leu Lys Arg 770 775 780 Ser Lys Leu Arg Ser Tyr Gln Tyr Val Lys Leu Ser Gly Asn Phe Glu 785 790 795 800 Arg Ser Ala Gly Leu Gln Ala Gly Ile Asp Leu Val Lys Asp Pro His 805 810 815 Ser Ile Ile Phe Leu Cys Asp Leu His Ile His Phe Pro Ala Gly Val 820 825 830 Ile Asp Ala Ile Arg Lys His Cys Val Glu Gly Lys Met Ala Phe Ala 835 840 845 Pro Met Val Met Arg Leu His Cys Gly Ala Thr Pro Gln Trp Pro Glu 850 855 860 Gly Tyr Trp Glu Val Asn Gly Phe Gly Leu Leu Gly Ile Tyr Lys Ser 865 870 875 880 Asp Leu Asp Arg Ile Gly Gly Met Asn Thr Lys Glu Phe Arg Asp Arg 885 890 895 Trp Gly Gly Glu Asp Trp Glu Leu Leu Asp Arg Ile Leu Gln Gly Leu 900 905 910 Asp Val Glu Arg Leu Ser Leu Arg Asn Phe Phe His His Phe His Ser 915 920 925 Lys Arg Gly Met Trp Ser Arg Arg Gln Met Lys Thr Leu 930 935 940 <210> 25 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; W336F) <400> 25 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Phe 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 26 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; W336H) <400> 26 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly His 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 27 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; W336V) <400> 27 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Val 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 28 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; E339A) <400> 28 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Ala Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 29 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; L302A) <400> 29 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Ala His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 30 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; E339D) <400> 30 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Asp Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 31 <211> 389 <212> PRT <213> Artificial sequence <220> <223> TnGalNAcT(33-421; E339S) <400> 31 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Ser Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 32 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; W336H,E339A) <400> 32 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly His 290 295 300 Gly Gly Ala Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 33 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; W336H,E339D) <400> 33 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly His 290 295 300 Gly Gly Asp Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 34 <211> 389 <212> PRT <213> Artificial sequence <220> <223> TnGalNAcT(33‑421; W336H,E339S) <400> 34 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly His 290 295 300 Gly Gly Ser Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 35 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; L302G) <400> 35 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Gly His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 36 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; I299M) <400> 36 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Met Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 37 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; I299A) <400> 37 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Light Light Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ala Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 38 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; I299G) <400> 38 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Gly Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 39 <211> 389 <212> PRT <213> artificial sequence <220> <223> TnGalNAcT(33‑421; I311M) <400> 39 Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu Tyr Asn Ala Thr Gln 1 5 10 15 Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala Asn Trp Pro Lys Lys 20 25 30 Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu Tyr Ser Ile Lys Asn 35 40 45 Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser Val Val His Pro Pro 50 55 60 Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp Lys Asn Met Thr Ile 65 70 75 80 Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr Pro Leu Leu Ile Thr 85 90 95 Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr Thr Glu Asp Gly Val 100 105 110 Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu Cys Asp Ser Met Pro 115 120 125 Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr Glu Leu Glu Leu Glu 130 135 140 Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp Gly Gly Arg Tyr Ser 145 150 155 160 Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala Ile Ile Val Pro Tyr 165 170 175 Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu Asn His Met His Pro 180 185 190 Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile Phe Ile Val Glu Gln 195 200 205 Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu Met Asn Val Gly Phe 210 215 220 Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp Gln Cys Phe Val Phe 225 230 235 240 His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg Asn Leu Tyr Ser Cys 245 250 255 Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile Asp Lys Leu His Phe 260 265 270 Lys Leu Pro Tyr Glu Asp Met Phe Gly Gly Val Ser Ala Met Thr Leu 275 280 285 Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn Lys Tyr Trp Gly Trp 290 295 300 Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu Lys Lys Ile Asn Tyr 305 310 315 320 His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg Tyr Ala Met Leu Asp 325 330 335 His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr Gln Leu Leu Ser Gln 340 345 350 Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser Thr Leu Glu Tyr Glu 355 360 365 Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr His Ile Leu Val Asn 370 375 380 Ile Asp Glu Arg Ser 385 <210> 40 <211> 354 <212> PRT <213> artificial sequence <220> <223> AsGalNAcT(30‑383; F248A) <400> 40 Asp Tyr Ser Phe Trp Ser Pro Ala Phe Ile Ile Ser Ala Pro Lys Thr 1 5 10 15 Leu Thr Thr Leu Gln Pro Phe Ser Gln Ser Thr Ser Thr Asn Asp Leu 20 25 30 Ala Val Ser Ala Leu Glu Ser Val Glu Phe Ser Met Leu Asp Asn Ser 35 40 45 Ser Ile Leu His Ala Ser Asp Asn Trp Thr Asn Asp Glu Leu Val Met 50 55 60 Arg Ala Gln Asn Glu Asn Leu Gln Leu Cys Pro Met Thr Pro Pro Ala 65 70 75 80 Leu Val Gly Pro Ile Lys Val Trp Met Asp Ala Pro Ser Phe Ala Glu 85 90 95 Leu Glu Arg Leu Tyr Pro Phe Leu Glu Pro Gly Gly His Gly Met Pro 100 105 110 Thr Ala Cys Arg Ala Arg His Arg Val Ala Ile Val Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ser His Leu Arg Thr Phe Leu His Asn Leu His Ser Leu 130 135 140 Leu Thr Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Thr 145 150 155 160 Ala Asn Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Ala 165 170 175 Glu Ala Ile Arg Leu Tyr Asp Trp Arg Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Ser Cys Pro Asp Glu 195 200 205 Pro Arg His Met Ser Val Ala Val Asp Lys Ala Asn Tyr Lys Leu Pro 210 215 220 Tyr Gly Ser Ile Phe Gly Gly Ile Ser Ala Leu Thr Arg Glu Gln Phe 225 230 235 240 Glu Gly Ile Asn Gly Phe Ser Asn Asp Tyr Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ser Thr Arg Val Thr Leu Ala Gly Tyr Lys Ile Ser 260 265 270 Arg Tyr Pro Ala Glu Ile Ala Arg Tyr Lys Met Ile Lys His Asn Ser 275 280 285 Glu Lys Lys Asn Pro Val Asn Arg Cys Arg Tyr Lys Leu Met Ser Ala 290 295 300 Thr Lys Ser Arg Trp Arg Asn Asp Gly Leu Ser Ser Leu Ser Tyr Asp 305 310 315 320 Leu Ile Ser Leu Gly Arg Leu Pro Leu Tyr Thr His Ile Lys Val Asp 325 330 335 Leu Leu Glu Lys Gln Ser Arg Arg Tyr Leu Arg Thr His Gly Phe Pro 340 345 350 Thr Cys <210> 41 <211> 354 <212> PRT <213> artificial sequence <220> <223> AsGalNAcT(30‑383; F248G) <400> 41 Asp Tyr Ser Phe Trp Ser Pro Ala Phe Ile Ile Ser Ala Pro Lys Thr 1 5 10 15 Leu Thr Thr Leu Gln Pro Phe Ser Gln Ser Thr Ser Thr Asn Asp Leu 20 25 30 Ala Val Ser Ala Leu Glu Ser Val Glu Phe Ser Met Leu Asp Asn Ser 35 40 45 Ser Ile Leu His Ala Ser Asp Asn Trp Thr Asn Asp Glu Leu Val Met 50 55 60 Arg Ala Gln Asn Glu Asn Leu Gln Leu Cys Pro Met Thr Pro Pro Ala 65 70 75 80 Leu Val Gly Pro Ile Lys Val Trp Met Asp Ala Pro Ser Phe Ala Glu 85 90 95 Leu Glu Arg Leu Tyr Pro Phe Leu Glu Pro Gly Gly His Gly Met Pro 100 105 110 Thr Ala Cys Arg Ala Arg His Arg Val Ala Ile Val Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ser His Leu Arg Thr Phe Leu His Asn Leu His Ser Leu 130 135 140 Leu Thr Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Thr 145 150 155 160 Ala Asn Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Ala 165 170 175 Glu Ala Ile Arg Leu Tyr Asp Trp Arg Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Ser Cys Pro Asp Glu 195 200 205 Pro Arg His Met Ser Val Ala Val Asp Lys Gly Asn Tyr Lys Leu Pro 210 215 220 Tyr Gly Ser Ile Phe Gly Gly Ile Ser Ala Leu Thr Arg Glu Gln Phe 225 230 235 240 Glu Gly Ile Asn Gly Phe Ser Asn Asp Tyr Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ser Thr Arg Val Thr Leu Ala Gly Tyr Lys Ile Ser 260 265 270 Arg Tyr Pro Ala Glu Ile Ala Arg Tyr Lys Met Ile Lys His Asn Ser 275 280 285 Glu Lys Lys Asn Pro Val Asn Arg Cys Arg Tyr Lys Leu Met Ser Ala 290 295 300 Thr Lys Ser Arg Trp Arg Asn Asp Gly Leu Ser Ser Leu Ser Tyr Asp 305 310 315 320 Leu Ile Ser Leu Gly Arg Leu Pro Leu Tyr Thr His Ile Lys Val Asp 325 330 335 Leu Leu Glu Lys Gln Ser Arg Arg Tyr Leu Arg Thr His Gly Phe Pro 340 345 350 Thr Cys <210> 42 <211> 354 <212> PRT <213> Artificial Sequence <220> <223> AsGalNAcT(30‑383; V245M) <400> 42 Asp Tyr Ser Phe Trp Ser Pro Ala Phe Ile Ile Ser Ala Pro Lys Thr 1 5 10 15 Leu Thr Thr Leu Gln Pro Phe Ser Gln Ser Thr Ser Thr Asn Asp Leu 20 25 30 Ala Val Ser Ala Leu Glu Ser Val Glu Phe Ser Met Leu Asp Asn Ser 35 40 45 Ser Ile Leu His Ala Ser Asp Asn Trp Thr Asn Asp Glu Leu Val Met 50 55 60 Arg Ala Gln Asn Glu Asn Leu Gln Leu Cys Pro Met Thr Pro Pro Ala 65 70 75 80 Leu Val Gly Pro Ile Lys Val Trp Met Asp Ala Pro Ser Phe Ala Glu 85 90 95 Leu Glu Arg Leu Tyr Pro Phe Leu Glu Pro Gly Gly His Gly Met Pro 100 105 110 Thr Ala Cys Arg Ala Arg His Arg Val Ala Ile Val Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ser His Leu Arg Thr Phe Leu His Asn Leu His Ser Leu 130 135 140 Leu Thr Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Thr 145 150 155 160 Ala Asn Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Ala 165 170 175 Glu Ala Ile Arg Leu Tyr Asp Trp Arg Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Ser Cys Pro Asp Glu 195 200 205 Pro Arg His Met Ser Val Ala Met Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Gly Ser Ile Phe Gly Gly Ile Ser Ala Leu Thr Arg Glu Gln Phe 225 230 235 240 Glu Gly Ile Asn Gly Phe Ser Asn Asp Tyr Trp Gly Trp Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ser Thr Arg Val Thr Leu Ala Gly Tyr Lys Ile Ser 260 265 270 Arg Tyr Pro Ala Glu Ile Ala Arg Tyr Lys Met Ile Lys His Asn Ser 275 280 285 Glu Lys Lys Asn Pro Val Asn Arg Cys Arg Tyr Lys Leu Met Ser Ala 290 295 300 Thr Lys Ser Arg Trp Arg Asn Asp Gly Leu Ser Ser Leu Ser Tyr Asp 305 310 315 320 Leu Ile Ser Leu Gly Arg Leu Pro Leu Tyr Thr His Ile Lys Val Asp 325 330 335 Leu Leu Glu Lys Gln Ser Arg Arg Tyr Leu Arg Thr His Gly Phe Pro 340 345 350 Thr Cys <210> 43 <211> 410 <212> PRT <213> Artificial sequence <220> <223> His6-TnGalNAcT(33-421; L302A) <400> 43 Met Gly Ser Ser His His His His His His Ser Ser Gly Leu Val Pro 1 5 10 15 Arg Gly Ser His Met Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu 20 25 30 Tyr Asn Ala Thr Gln Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala 35 40 45 Asn Trp Pro Lys Lys Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu 50 55 60 Tyr Ser Ile Lys Asn Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser 65 70 75 80 Val Val His Pro Pro Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp 85 90 95 Lys Asn Met Thr Ile Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr 100 105 110 Pro Leu Leu Ile Thr Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr 115 120 125 Thr Glu Asp Gly Val Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu 130 135 140 Cys Asp Ser Met Pro Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr 145 150 155 160 Glu Leu Glu Leu Glu Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp 165 170 175 Gly Gly Arg Tyr Ser Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala 180 185 190 Ile Ile Val Pro Tyr Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu 195 200 205 Asn His Met His Pro Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile 210 215 220 Phe Ile Val Glu Gln Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu 225 230 235 240 Met Asn Val Gly Phe Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp 245 250 255 Gln Cys Phe Val Phe His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg 260 265 270 Asn Leu Tyr Ser Cys Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile 275 280 285 Asp Lys Ala His Phe Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val 290 295 300 Ser Ala Met Thr Leu Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn 305 310 315 320 Lys Tyr Trp Gly Trp Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu 325 330 335 Lys Lys Ile Asn Tyr His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg 340 345 350 Tyr Ala Met Leu Asp His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr 355 360 365 Gln Leu Leu Ser Gln Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser 370 375 380 Thr Leu Glu Tyr Glu Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr 385 390 395 400 His Ile Leu Val Asn Ile Asp Glu Arg Ser 405 410 <210> 44 <211> 410 <212> PRT <213> Artificial sequence <220> <223> His6‑TnGalNAcT(33‑421; L302G) <400> 44 Met Gly Ser Ser His His His His His His Ser Ser Gly Leu Val Pro 1 5 10 15 Arg Gly Ser His Met Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu 2 '0 25 30 Tyr Asn Ala Thr Gln Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala 35 40 45 Asn Trp Pro Lys Lys Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu 50 55 60 Tyr Ser Ile Lys Asn Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser 65 70 75 80 Val Val His Pro Pro Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp 85 90 95 Lys Asn Met Thr Ile Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr 100 105 110 Pro Leu Leu Ile Thr Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr 115 120 125 Thr Glu Asp Gly Val Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu 130 135 140 Cys Asp Ser Met Pro Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr 145 150 155 160 Glu Leu Glu Leu Glu Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp 165 170 175 Gly Gly Arg Tyr Ser Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala 180 185 190 Ile Ile Val Pro Tyr Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu 195 200 205 Asn His Met His Pro Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile 210 215 220 Phe Ile Val Glu Gln Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu 225 230 235 240 Met Asn Val Gly Phe Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp 245 250 255 Gln Cys Phe Val Phe His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg 260 265 270 Asn Leu Tyr Ser Cys Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile 275 280 285 Asp Lys Gly His Phe Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val 290 295 300 Ser Ala Met Thr Leu Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn 305 310 315 320 Lys Tyr Trp Gly Trp Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu 325 330 335 Lys Lys Ile Asn Tyr His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg 340 345 350 Tyr Ala Met Leu Asp His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr 355 360 365 Gln Leu Leu Ser Gln Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser 370 375 380 Thr Leu Glu Tyr Glu Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr 385 390 395 400 His Ile Leu Val Asn Ile Asp Glu Arg Ser 405 410 <210> 45 <211> 410 <212> PRT <213> Artificial Sequence <220> <223> His6-TnGalNAcT(33-421; I299M) <400> 45 Met Gly Ser Ser His His His His His His Ser Ser Gly Leu Val Pro 1 5 10 15 Arg Gly Ser His Met Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu 20 25 30 Tyr Asn Ala Thr Gln Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala 35 40 45 Asn Trp Pro Lys Lys Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu 50 55 60 Tyr Ser Ile Lys Asn Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser 65 70 75 80 Val Val His Pro Pro Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp 85 90 95 Lys Asn Met Thr Ile Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr 100 105 110 Pro Leu Leu Ile Thr Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr 115 120 125 Thr Glu Asp Gly Val Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu 130 135 140 Cys Asp Ser Met Pro Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr 145 150 155 160 Glu Leu Glu Leu Glu Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp 165 170 175 Gly Gly Arg Tyr Ser Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala 180 185 190 Ile Ile Val Pro Tyr Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu 195 200 205 Asn His Met His Pro Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile 210 215 220 Phe Ile Val Glu Gln Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu 225 230 235 240 Met Asn Val Gly Phe Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp 245 250 255 Gln Cys Phe Val Phe His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg 260 265 270 Asn Leu Tyr Ser Cys Pro Arg Gln Pro Arg His Met Ser Ala Ser Met 275 280 285 Asp Lys Leu His Phe Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val 290 295 300 Ser Ala Met Thr Leu Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn 305 310 315 320 Lys Tyr Trp Gly Trp Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu 325 330 335 Lys Lys Ile Asn Tyr His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg 340 345 350 Tyr Ala Met Leu Asp His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr 355 360 365 Gln Leu Leu Ser Gln Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser 370 375 380 Thr Leu Glu Tyr Glu Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr 385 390 395 400 His Ile Leu Val Asn Ile Asp Glu Arg Ser 405 410 <210> 46 <211> 354 <212> PRT <213> artificial sequence <220> <223> AsGalNAcT(30‑383; W282H) <400> 46 Asp Tyr Ser Phe Trp Ser Pro Ala Phe Ile Ile Ser Ala Pro Lys Thr 1 5 10 15 Leu Thr Thr Leu Gln Pro Phe Ser Gln Ser Thr Ser Thr Asn Asp Leu 20 25 30 Ala Val Ser Ala Leu Glu Ser Val Glu Phe Ser Met Leu Asp Asn Ser 35 40 45 Ser Ile Leu His Ala Ser Asp Asn Trp Thr Asn Asp Glu Leu Val Met 50 55 60 Arg Ala Gln Asn Glu Asn Leu Gln Leu Cys Pro Met Thr Pro Pro Ala 65 70 75 80 Leu Val Gly Pro Ile Lys Val Trp Met Asp Ala Pro Ser Phe Ala Glu 85 90 95 Leu Glu Arg Leu Tyr Pro Phe Leu Glu Pro Gly Gly His Gly Met Pro 100 105 110 Thr Ala Cys Arg Ala Arg His Arg Val Ala Ile Val Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ser His Leu Arg Thr Phe Leu His Asn Leu His Ser Leu 130 135 140 Leu Thr Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Thr 145 150 155 160 Ala Asn Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Ala 165 170 175 Glu Ala Ile Arg Leu Tyr Asp Trp Arg Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Ser Cys Pro Asp Glu 195 200 205 Pro Arg His Met Ser Val Ala Val Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Gly Ser Ile Phe Gly Gly Ile Ser Ala Leu Thr Arg Glu Gln Phe 225 230 235 240 Glu Gly Ile Asn Gly Phe Ser Asn Asp Tyr Trp Gly His Gly Gly Glu 245 250 255 Asp Asp Asp Leu Ser Thr Arg Val Thr Leu Ala Gly Tyr Lys Ile Ser 260 265 270 Arg Tyr Pro Ala Glu Ile Ala Arg Tyr Lys Met Ile Lys His Asn Ser 275 280 285 Glu Lys Lys Asn Pro Val Asn Arg Cys Arg Tyr Lys Leu Met Ser Ala 290 295 300 Thr Lys Ser Arg Trp Arg Asn Asp Gly Leu Ser Ser Leu Ser Tyr Asp 305 310 315 320 Leu Ile Ser Leu Gly Arg Leu Pro Leu Tyr Thr His Ile Lys Val Asp 325 330 335 Leu Leu Glu Lys Gln Ser Arg Arg Tyr Leu Arg Thr His Gly Phe Pro 340 345 350 Thr Cys <210> 47 <211> 354 <212> PRT <213> artificial sequence <220> <223> AsGalNAcT(30‑383; E285D) <400> 47 Asp Tyr Ser Phe Trp Ser Pro Ala Phe Ile Ile Ser Ala Pro Lys Thr 1 5 10 15 Leu Thr Thr Leu Gln Pro Phe Ser Gln Ser Thr Ser Thr Asn Asp Leu 20 25 30 Ala Val Ser Ala Leu Glu Ser Val Glu Phe Ser Met Leu Asp Asn Ser 35 40 45 Ser Ile Leu His Ala Ser Asp Asn Trp Thr Asn Asp Glu Leu Val Met 50 55 60 Arg Ala Gln Asn Glu Asn Leu Gln Leu Cys Pro Met Thr Pro Pro Ala 65 70 75 80 Leu Val Gly Pro Ile Lys Val Trp Met Asp Ala Pro Ser Phe Ala Glu 85 90 95 Leu Glu Arg Leu Tyr Pro Phe Leu Glu Pro Gly Gly His Gly Met Pro 100 105 110 Thr Ala Cys Arg Ala Arg His Arg Val Ala Ile Val Val Pro Tyr Arg 115 120 125 Asp Arg Glu Ser His Leu Arg Thr Phe Leu His Asn Leu His Ser Leu 130 135 140 Leu Thr Lys Gln Gln Leu Asp Tyr Ala Ile Phe Val Val Glu Gln Thr 145 150 155 160 Ala Asn Glu Thr Phe Asn Arg Ala Lys Leu Met Asn Val Gly Tyr Ala 165 170 175 Glu Ala Ile Arg Leu Tyr Asp Trp Arg Cys Phe Ile Phe His Asp Val 180 185 190 Asp Leu Leu Pro Glu Asp Asp Arg Asn Leu Tyr Ser Cys Pro Asp Glu 195 200 205 Pro Arg His Met Ser Val Ala Val Asp Lys Phe Asn Tyr Lys Leu Pro 210 215 220 Tyr Gly Ser Ile Phe Gly Gly Ile Ser Ala Leu Thr Arg Glu Gln Phe 225 230 235 240 Glu Gly Ile Asn Gly Phe Ser Asn Asp Tyr Trp Gly Trp Gly Gly Asp 245 250 255 Asp Asp Asp Leu Ser Thr Arg Val Thr Leu Ala Gly Tyr Lys Ile Ser 260 265 270 Arg Tyr Pro Ala Glu Ile Ala Arg Tyr Lys Met Ile Lys His Asn Ser 275 280 285 Glu Lys Lys Asn Pro Val Asn Arg Cys Arg Tyr Lys Leu Met Ser Ala 290 295 300 Thr Lys Ser Arg Trp Arg Asn Asp Gly Leu Ser Ser Leu Ser Tyr Asp 305 310 315 320 Leu Ile Ser Leu Gly Arg Leu Pro Leu Tyr Thr His Ile Lys Val Asp 325 330 335 Leu Leu Glu Lys Gln Ser Arg Arg Tyr Leu Arg Thr His Gly Phe Pro 340 345 350 Thr Cys <210> 48 <211> 410 <212> PRT <213> artificial sequence <220> <223> His6‑TnGalNAcT(33‑421; I299A) <400> 48 Met Gly Ser Ser His His His His His His Ser Ser Gly Leu Val Pro 1 5 10 15 Arg Gly Ser His Met Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu 20 25 30 Tyr Asn Ala Thr Gln Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala 35 40 45 Asn Trp Pro Lys Lys Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu 50 55 60 Tyr Ser Ile Lys Asn Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser 65 70 75 80 Val Val His Pro Pro Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp 85 90 95 Lys Asn Met Thr Ile Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr 100 105 110 Pro Leu Leu Ile Thr Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr 115 120 125 Thr Glu Asp Gly Val Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu 130 135 140 Cys Asp Ser Met Pro Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr 145 150 155 160 Glu Leu Glu Leu Glu Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp 165 170 175 Gly Gly Arg Tyr Ser Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala 180 185 190 Ile Ile Val Pro Tyr Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu 195 200 205 Asn His Met His Pro Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile 210 215 220 Phe Ile Val Glu Gln Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu 225 230 235 240 Met Asn Val Gly Phe Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp 245 250 255 Gln Cys Phe Val Phe His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg 260 265 270 Asn Leu Tyr Ser Cys Pro Arg Gln Pro Arg His Met Ser Ala Ser Ala 275 280 285 Asp Lys Leu His Phe Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val 290 295 300 Ser Ala Met Thr Leu Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn 305 310 315 320 Lys Tyr Trp Gly Trp Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu 325 330 335 Lys Lys Ile Asn Tyr His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg 340 345 350 Tyr Ala Met Leu Asp His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr 355 360 365 Gln Leu Leu Ser Gln Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser 370 375 380 Thr Leu Glu Tyr Glu Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr 385 390 395 400 His Ile Leu Val Asn Ile Asp Glu Arg Ser 405 410 <210> 49 <211> 410 <212> PRT <213> Artificial Sequence <220> <223> His6-TnGalNAcT(33-421) <400> 49 Met Gly Ser Ser His His His His His His Ser Ser Gly Leu Val Pro 1 5 10 15 Arg Gly Ser His Met Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu 20 25 30 Tyr Asn Ala Thr Gln Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala 35 40 45 Asn Trp Pro Lys Lys Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu 50 55 60 Tyr Ser Ile Lys Asn Ile Ser Leu Ser Asn His Thr Thr Arg Ala Ser 65 70 75 80 Val Val His Pro Pro Ser Ser Ile Thr Glu Thr Ala Ser Lys Leu Asp 85 90 95 Lys Asn Met Thr Ile Gln Asp Gly Ala Phe Ala Met Ile Ser Pro Thr 100 105 110 Pro Leu Leu Ile Thr Lys Leu Met Asp Ser Ile Lys Ser Tyr Val Thr 115 120 125 Thr Glu Asp Gly Val Lys Lys Ala Glu Ala Val Val Thr Leu Pro Leu 130 135 140 Cys Asp Ser Met Pro Pro Asp Leu Gly Pro Ile Thr Leu Asn Lys Thr 145 150 155 160 Glu Leu Glu Leu Glu Trp Val Glu Lys Lys Phe Pro Glu Val Glu Trp 165 170 175 Gly Gly Arg Tyr Ser Pro Pro Asn Cys Thr Ala Arg His Arg Val Ala 180 185 190 Ile Ile Val Pro Tyr Arg Asp Arg Gln Gln His Leu Ala Ile Phe Leu 195 200 205 Asn His Met His Pro Phe Leu Met Lys Gln Gln Ile Glu Tyr Gly Ile 210 215 220 Phe Ile Val Glu Gln Glu Gly Asn Lys Asp Phe Asn Arg Ala Lys Leu 225 230 235 240 Met Asn Val Gly Phe Val Glu Ser Gln Lys Leu Val Ala Glu Gly Trp 245 250 255 Gln Cys Phe Val Phe His Asp Ile Asp Leu Leu Pro Leu Asp Thr Arg 260 265 270 Asn Leu Tyr Ser Cys Pro Arg Gln Pro Arg His Met Ser Ala Ser Ile 275 280 285 Asp Lys Leu His Phe Lys Leu Pro Tyr Glu Asp Ile Phe Gly Gly Val 290 295 300 Ser Ala Met Thr Leu Glu Gln Phe Thr Arg Val Asn Gly Phe Ser Asn 305 310 315 320 Lys Tyr Trp Gly Trp Gly Gly Glu Asp Asp Asp Met Ser Tyr Arg Leu 325 330 335 Lys Lys Ile Asn Tyr His Ile Ala Arg Tyr Lys Met Ser Ile Ala Arg 340 345 350 Tyr Ala Met Leu Asp His Lys Lys Ser Thr Pro Asn Pro Lys Arg Tyr 355 360 365 Gln Leu Leu Ser Gln Thr Ser Lys Thr Phe Gln Lys Asp Gly Leu Ser 370 375 380 Thr Leu Glu Tyr Glu Leu Val Gln Val Val Gln Tyr His Leu Tyr Thr 385 390 395 400 His Ile Leu Val Asn Ile Asp Glu Arg Ser 405 410 <210> 50 <211> 410 <212> PRT <213> Artificial sequence <220> <223> His6-TnGalNAcT(33-421; W336F) <400> 50 Met Gly Ser Ser His His His His His His Ser Ser Gly Leu Val Pro 1 5 10 15 Arg Gly Ser His Met Ser Pro Leu Arg Thr Tyr Leu Tyr Thr Pro Leu 20 25 30 Tyr Asn Ala Thr Gln Pro Thr Leu Arg Asn Val Glu Arg Leu Ala Ala 35 40 45 Asn Trp Pro Lys Lys Ile Pro Ser Asn Tyr Ile Glu Asp Ser Glu Glu 50 55 60 Tyr Ser Ile Lys Asn Ile Ser Leu Ser Asn His ...
Claims
1. A method for modifying glycoproteins, the method comprising the following steps: In the presence of β-(1,4)-N-acetylgalactosamine transferase, a glycoprotein containing a glycan with a terminal GlcNAc moiety is contacted with a sugar derivative nucleotide Su(A)-Nuc, wherein: (i) The polysaccharide containing the terminal GlcNAc moiety is as shown in formula (1) or (2): in: b is 0 or 1; d is 0 or 1; e is 0 or 1; and G is a monosaccharide, or a straight-chain or branched oligosaccharide containing 2 to 20 sugar moieties; and (ii) The sugar derivative nucleotide Su(A)-Nuc is as shown in formula (3): in: a is 0 or 1; Nuc stands for nucleotide; U is [C(R)] 1 )2] n Where n is 0, 1, 2, or 3; R 1 Independently selected from H, F, Cl, Br, and I; T is a (hetero)arylene, wherein the (hetero)arylene is selected from phenylene, pyridinyl, pyridiniumyl, pyrimidinyl, pyrimidiniumyl, imidazolyl, imidazolyl, pyrrololyl, furanyl, and thiopheneyl, wherein the (hetero)arylene is optionally replaced by one or more substituents R. 2 Replace, where R 2 Independently selected from -F, -Cl, -Br, -CN, -NO2, methyl, and methoxy; and A is selected from: (a) -N3 (b) -C(O)R 3 Where R 3 It can be methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl; (c) -C≡C-R 4 Where R 4 It can be hydrogen, methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl; (d) -SH (e) -SC(O)R 8 Where R 8 It can be methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl; (f) -SC(V)OR 8 Where V is O or S, R 8 It can be methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl; (g) -X Where X is selected from F, Cl, Br and I; (h) -OS(O)2R 5 Where R 5 Selected from methyl, ethyl, phenyl, or p-tolyl; (i) R 11 Where R 11 It is ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl; (j) R 12 Where R 12 Selected from -C(H)=CH2, -CH2-C(H)=CH2 and -CH2-CH2-C(H)=CH2; and (k) R 13 Where R 13 Selected from -C(H)=C=CH2 and CH2-C(H)=C=CH2; and (iii) wherein the β-(1,4)-N-acetylgalactosamine transferase is a sequence selected from SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 and 14, or a sequence selected from SEQ ID NO: 25, 26, 27, 28, 29, 35, 36, 43, 44, 45, 49, 50, 51, 52, 53 and 71.
2. The method of claim 1, wherein the β-(1,4)-N-acetylgalactosamine transferase is derived from Caenorhabditis elegans, Ascaris lumbricoides, Spodoptera litura, and Drosophila melanogaster.
3. The method of claim 1, wherein n is 1, 2 or 3.
4. The method of claim 1, wherein n is 0.
5. The method of claim 1, wherein the sugar derivative nucleotide Su(A)-Nuc is as shown in formula (9) or (10): Nuc, A, U and T are as defined in claim 1.
6. The method of claim 1, wherein the sugar derivative nucleotide is as shown in formula (11), (12), (13), (14), (15), (16), (34), or (35): in: A as defined in claim 1; and m is an integer from 0 to 4; R 2 Independently selected from -F, -Cl, -Br, -CN, -NO2, methyl, and methoxy; and R 6 It is H or methyl.
7. The method of claim 1, wherein the nucleotide is UDP.
8. The method of claim 1, wherein the sugar derivative nucleotide is as shown in formula (17), (18), (19), (20), (21) or (22): in: X is F, Cl, Br, or I; R 8 It can be methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, or tert-butyl; V is either O or S; r is 0 or 1; s is 2 or 3; t is 1, 2, or 3; and u is 1 or 2.
9. The method of claim 1, wherein the sugar derivative nucleotide is as shown in formula (23), (24) or (25): Where X is F, Cl, Br or I.
10. The method of claim 8 or 9, wherein the sugar derivative nucleotide is as shown in formula (17), (18), (19), (23) or (24), wherein (17), (18) and (19) are as defined in claim 8, wherein (23) and (24) are as defined in claim 9, and wherein X is Cl or F.
11. The method of claim 1, wherein the sugar derivative nucleotide Su(A)-Nuc is Su(A)-UDP as shown in formula (31): U, T, A, and a are as defined in claim 1.
12. The method of claim 1, wherein the polysaccharide containing the terminal GlcNAc moiety is as shown in formula (1), (26), or (27): in: b is as defined in claim 1.
13. The method of claim 1, wherein the glycoprotein comprising a glycan containing a terminal GlcNAc moiety is as shown in formula (7) or (8): in: b, d, e, and G are as defined in claim 1; y is an integer from 1 to 24; and Pr is a protein.
14. The method of claim 1, wherein the glycoprotein comprising a glycan containing a terminal GlcNAc moiety is an antibody.
Citation Information
Patent Citations
Oligosaccharide modification and labeling of proteins
WO2007095506A1
Labeling and detection of post translationally modified proteins
WO2008029281A2
Endoglycosidase from streptococcus pyogenes and methods using it
WO2013037824A1
Modified glycoprotein, protein-conjugate and process for the preparation thereof
WO2015057063A1
Process for the attachment of a galnac moiety comprising a (hetero)aryl group to a glcnac moiety, and product obtained thereby
WO2015112013A1