Polypeptides for bioconjugation
Engineered microbial transglutaminase enzymes address the challenge of heterogeneous bioconjugates by enabling site-specific conjugation at lysine or glutamine residues, enhancing the efficiency and specificity of antibody-drug conjugates.
Patent Information
- Application Number
- PCT/US2025/020148
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-25
AI Technical Summary
Existing methods for bioconjugation of compounds to biomolecules, such as antibody-drug conjugates, result in heterogeneous conjugates due to the ubiquity of lysine and cysteine residues, complicating characterization and therapeutic use, and current attempts at controlled bioconjugation have not achieved efficient and widely applicable yields.
Engineered microbial transglutaminase enzymes with specific amino acid sequences, capable of catalyzing site-specific conjugation between biomolecules and compounds, particularly at lysine or glutamine residues, to form targeted biomolecule-compound conjugates.
The engineered enzymes enable high specificity and efficiency in bioconjugation, resulting in predominantly conjugated compounds at defined positions on biomolecules, improving the characterization and therapeutic potential of antibody-drug conjugates.
Smart Images

Figure IMGF000008_0001 
Figure IMGF000008_0002 
Figure IMGF000036_0001
Abstract
Description
POLYPEPTIDES FOR BIOCONJUGATIONCROSS-REERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 568,183, filed March 21, 2024, the disclosure of which is incorporated herein by reference in its entirety.REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY
[0002] The sequence listing of the present application is submitted electronically as an XML formatted sequence listing created on January 15, 2025, with the file name “25874.xml” and having a size of 116 kb. This sequence listing submitted electronically is part of the specification and is herein incorporated by reference in its entirety.FIELD OF THE INVENTION
[0003] The present disclosure relates generally to polypeptides that can be used for bioconjugation. More specifically, the present disclosure relates to polypeptides that are engineered microbial transglutaminase enzymes.BACKGROUND OF THE INVENTION
[0004] Bioconjugation of compounds to biomolecules has significant therapeutic applications, for example in synthesis of antibody-drug conjugates (ADCs). ADCs are being tested in various clinical trials, and they have the potential to deliver cytotoxic agents to specific locations within the body due to the selectivity of the antibodies in the ADC.
[0005] Common approaches of conjugating biomolecules to compounds rely on various lysine or cysteine residues of the biomolecules. However, since these residues are ubiquitous, such approaches yield heterogeneous conjugates that complicate their characterization and therapeutic use.
[0006] While there are some attempts to achieve more controlled bioconjugation, such as the genetic encoding of non-standard amino acids, such attempts have not yet reached efficient and widely applicable yields, such as for synthesis of ADCs. Therefore, there is a need in the field for approaches that allow specific conjugation of compounds to biomolecules.SUMMARY OF THE INVENTION
[0007] Disclosed are polypeptides that can be used for bioconjugation. For example, the polypeptides can be used for antibody-drug conjugate synthesis. The disclosed polypeptides are engineered microbial transglutaminase enzymes (mTG).
[0008] In some aspects, the invention comprises polypeptides comprising an amino acid sequence that has at least 90% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 91% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, polypeptides comprise an amino acid sequence that has at least 92% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 93% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 94% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 95% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 96% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 97% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 98% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the polypeptides comprise an amino acid sequence that has at least 99% identity with the amino acid sequence of SEQ ID NO: 8.
[0009] In some embodiments of these aspects of the invention, the polypeptide comprises the amino acid sequence of SEQ ID NO: 9. In some embodiments, the polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 1 to 7. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 2. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 3. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 4. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 5. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 6. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 7.
[0010] In some embodiments, the polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 31 to 37. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 31. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 32. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 33. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 34. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 35. In some embodiments, the polypeptide comprises the amino acidsequence of SEQ ID NO: 36. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 37.
[0011] In some embodiments, the polypeptide consists essentially of the amino acid sequence of any one of SEQ ID NOs: 31 to 37. In some embodiments, the invention relates to a polypeptide that consists of the amino acid sequence of any one of SEQ ID NOs: 31 to 37. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 31. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 32. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 33. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 34. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 35. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 36. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 37. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 31. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 32. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 33. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 34. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 35. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 36. In some embodiments, the polypeptide consists of the amino acid sequence of SEQ ID NO: 37.
[0012] In some embodiments of these aspects of the invention, the polypeptides have transglutaminase activity. In some embodiments, the polypeptide is capable of catalyzing a reaction between a lysine residue of a biomolecule and a compound that comprises a primary amide group, to thereby make a biomolecule-compound conjugate. In some embodiments, the biomolecule is an antibody having a lysine at position 222 based on the Eu numbering scheme.
[0013] In some aspects, a polypeptide of the invention comprises an amino acid sequence that has at least 90% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 91% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 92% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 93% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 94% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 95% identity with the amino acid sequence ofSEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 96% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 97% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 98% identity with the amino acid sequence of SEQ ID NO: 28. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 99% identity with the amino acid sequence of SEQ ID NO: 28.
[0014] In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 29.
[0015] In some embodiments, the polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 11 to 23. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 11. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 12. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 13. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 14. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 15. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 16. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 17. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 18. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 19. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 20. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 21. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 22. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 23.
[0016] In some embodiments, the polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 41 to 53. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 41. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 42. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 43. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 44. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 45. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 46. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 47. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 48. In some embodiments, the polypeptide comprises the amino acidsequence of SEQ ID NO: 49. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 50. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 51. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 52. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 53.
[0017] In some embodiments, the polypeptide consists essentially of the amino acid sequence of any one of SEQ ID NOs: 41 to 53. In some embodiments, the polypeptide consists of the amino acid sequence of any one of SEQ ID NOs: 41 to 53. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO: 41. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO: 42. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO: 43. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO: 44. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO: 45. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO: 46. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO:47. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO:48. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO:49. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO:50. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO:51. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO:52. In some embodiments, the polypeptide consists of or consists essentially of the amino acid sequence of SEQ ID NO:53.
[0018] In some embodiments, the polypeptide has transglutaminase activity. In some embodiments, the polypeptide comprises an amino acid sequence that has at least 90% identity with the amino acid sequence of SEQ ID NO:28 and has transglutaminase activity. In some embodiments, the polypeptide is capable of catalyzing a reaction between a glutamine residue of a biomolecule and a compound that comprises a primary amine group, to thereby make a biomolecule-compound conjugate. In some embodiments, the biomolecule is an antibody having a glutamine at position 295 based on the Eu numbering scheme.
[0019] Also described herein are compositions comprising one or more of the disclosed polypeptides and a buffering agent. Also described herein are kits comprising any of the polypeptides or disclosed compositions. Also described herein are isolated nucleic acids thatencode any of the disclosed polypeptides. In some embodiments, the isolated nucleic acids are DNAs. Also described herein are expression vectors comprising any of the disclosed nucleic acids. In some embodiments, said nucleic acids are operably linked to one or more control sequences suitable for directing expression of the encoded polypeptide in a host cell, wherein the one or more control sequences comprise a promoter. Also described herein are host cells comprising any of the disclosed expression vectors. Also described herein are methods of producing any of the disclosed polypeptides comprising cultivating a disclosed host cell in a medium under conditions suitable for expression of the polypeptide by the host cell; and optionally isolating the polypeptide from the medium.
[0020] Also described herein are methods of making a biomolecule-compound conjugate comprising reacting a biomolecule with a compound in the presence of a polypeptide of the invention to thereby make the biomolecule-compound conjugate. In some embodiments of the methods of the invention, the biomolecule comprises a lysine molecule. In some embodiments of the methods of the invention, the compound is conjugated to a lysine side chain of the biomolecule. In some embodiments of the methods of the invention, the lysine is at position K222, per the Eu numbering scheme and wherein the compound is predominantly conjugated to position K222 of the antibody, per the Eu numbering scheme. In some embodiments of the methods of the invention, more than 90% of the biomolecule-compound conjugates made by the method comprise a conjugation at position K222 of the antibody, per the Eu numbering scheme. In some embodiments of the methods of the invention, the biomolecule comprises a polypeptide. In some embodiments of the methods of the invention, biomolecule is an antibody. In some embodiments of the methods of the invention, the antibody is glycosylated at position N297, per the Eu numbering scheme.
[0021] In some embodiments of the methods of the invention, the biomolecule comprises, consists, or consists essentially of the amino acid sequence set forth in any one of SEQ ID NOs: 1-9 and 31-37. In some embodiments of the methods of the invention, the polypeptide comprises, consists, or consists essentially of the amino acid sequence set forth in any one of SEQ ID Nos: 28-29, 11-23 and 41-53.
[0022] In some embodiments of the methods of the invention, the compound comprises a drug. In some embodiments of the methods of the invention, the compound further comprises a linker, and wherein the linker comprises the primary amide group. In some embodiments of the methods of the invention, the linker comprises a spacer cleavable by an enzyme endogenous to a cell. In some embodiments of the methods of the invention, the compound is selected from Biotin K, Lys-Ala-Ala-PABC-MMAE,
[0023] Also described herein are methods of making a biomolecule-compound conjugate comprising reacting a biomolecule with a compound in the presence of a polypeptide of the invention to thereby make the biomolecule-compound conjugate, wherein the biomolecule comprises a glutamine molecule. In some embodiments of the methods of the invention, the compound is conjugated to a glutamine side chain of the biomolecule. In some embodiments of the methods of the invention, the glutamine is at position Q295, per the Eu numbering scheme and wherein the compound is predominantly conjugated to position Q295 of the biomolecule, per the Eu numbering scheme. In some embodiments of the methods of the invention, more than 90% of the biomolecule-compound conjugates made by the method comprise a conjugation at position Q295 of the biomolecule, per the Eu numbering scheme. In some embodiments of the methods of the invention, the biomolecule comprises a polypeptide. In some embodiments of the methods of the invention, biomolecule is an antibody. In some embodiments of the methods of the invention, the antibody is glycosylated at position N297, per the Eu numbering scheme.
[0024] In some embodiments of the methods of the invention, the biomolecule comprises, consists, or consists essentially of the amino acid sequence set forth in any one of SEQ ID NOs: l-9 and 31-37. In some embodiments of the methods of the invention, the biomolecule comprises, consists, or consists essentially of the amino acid sequence set forth in any one of SEQ ID NOs: 11-23, 28-29 and 41-53.
[0025] In some embodiments of the methods of the invention, the compound comprises a drug. In some embodiments of the methods of the invention, the compound further comprises a linker, and wherein the linker comprises the primary amine group. In some embodiments of the methods of the invention, the linker comprises a spacer cleavable by an enzyme endogenous to a cell. In some embodiments of the methods of the invention, the compound is selected from Biotin (Biotin K, Biotin Q, Biotin A), Lys-Ala-Ala-PABC-MMAE,BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIGURE 1 shows a scheme for conjugation of antibodies at their glutamine residue with a Biotin K Probe at 5mg scale.
[0027] FIGURE 2 shows a scheme for conjugation of antibodies at their glutamine residue with a Lys-Ala-Ala-PABC-MMAE Probe at 0.5mg scale.
[0028] FIGURE 3 shows a scheme for conjugation of antibodies at their glutamine residue with an Azide K Probe at Img scale.
[0029] FIGURE 4 shows a scheme for conjugation of antibodies at their lysine residue with a Biotin Q Probe.
[0030] FIGURE 5 shows a scheme for dual conjugation of antibodies at their glutamine and lysine residues with Biotin K and Q Probes at 2mg scale.
[0031] FIGURE 6 shows a scheme for dual conjugation of antibodies at their glutamine and lysine residues with Biotin K and Q Probes at 2mg scale.DETAILED DESCRIPTION OF THE INVENTION
[0032] Microbial transglutaminases naturally catalyze the cross-linkage of glutamine and lysine residues within proteins. The disclosed examples demonstrate that engineered transglutaminase enzymes are capable of catalyzing site-specific conjugation of molecules bearing lysine and glutamine residues and analogs that are linked with various cargoes. In this context, cargoes may consist of cleavable or non-cleavable linker elements that may include polyethylglycol spacers or peptides that are linked with cytotoxic molecules or other biologically active elements.Exemplary reactions along with corresponding enzyme sequences are disclosed. In some embodiments, fully glycosylated native mAbs are used as substrates for the conjugation. Some of the developed enzymes disclosed herein are capable of selectively modifying a specific lysine residue on a native mAb.DEFINITIONS
[0033] So that the invention may be more readily understood, certain technical and scientific terms are specifically defined below. Unless specifically defined elsewhere in this document, all other technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs.
[0034] As used herein, including the appended claims, the singular forms of words such as “a,” “an,” and “the,” include their corresponding plural references unless the context clearly dictates otherwise.
[0035] As used herein, “antibody” refers to an immunoglobulin, including recombinantly produced forms and includes any form of antibody that exhibits the desired biological activity. Thus, it is used in the broadest sense and specifically covers, but is not limited to, monoclonal antibodies (including full length monoclonal antibodies), polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), humanized antibodies, fully human antibodies, biparatopic antibodies, and chimeric antibodies.
[0036] The term “antibody” refers, in one embodiment, to a conventional antibody, which is a protein tetramer comprising two heavy chains (HCs) and two light chains (LCs) inter-connected by disulfide bonds, or an antigen binding portion thereof. In such an embodiment, each heavy chain is comprised of a heavy chain variable region or domain (abbreviated herein as VH) and a heavy chain constant region or domain. In certain naturally occurring IgG, IgD and IgA antibodies, the heavy chain constant region is comprised of three domains, CHI, CH2, and CH3. In certain naturally occurring antibodies, each light chain is comprised of a light chain variable region or domain (abbreviated herein as VL) and a light chain constant region or domain. The light chain constant region is comprised of one domain, CL. The human VH includes six family members: VH1, VH2, VH3, VH4, VH5, and VH6 and the human VL family includes 16 family members: VKI, VK2, VK3, VK4, VK5, VK6, V 1, VX2, VX3, VX4, VX5, VX6, VX7, VX8, VX9, and VX10. Each of these family members can be further divided into particular subtypes.
[0037] The VH and VL regions can be further subdivided into regions of hypervariability, termed complementarity determining regions (CDR), interspersed with regions that are more conserved, termed framework regions (FR). Each VH and VL is composed of three CDRs and four FRs, arranged from amino-terminus to carboxy -terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. The CDRs form a binding domain that interacts with an antigen.
[0038] The constant domains or regions of the antibodies may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (e.g., effector cells) and the first component (Clq) of the classical complement system. Typically, the numbering of the amino acids in the heavy chain constant domain begins with number 118, which is in accordance with the Eu numbering scheme. The Eu numbering scheme is based upon the amino acid sequence of human IgGl (Eu), which has a constant domain that begins at amino acid position 118 of the amino acid sequence of the IgGl described in Edelman et al., Proc. Natl.Acad. Sci. USA. 63: 78-85 (1969), and is shown for the IgGl, IgG2, IgG3, and IgG4 constant domains in Beranger, et al., Id. Unless otherwise noted, amino acid residues within an antibody are referred to herein based on the Eu numbering scheme, e.g., “an antibody having a lysine at position 22” means an antibody that has a lysine at position 222 per the Eu numbering scheme.
[0039] A conventional antibody tetramer includes two identical pairs of polypeptide chains, each pair having one “light” (about 25 kDa) and one “heavy” chain (about 50-70 kDa). The amino-terminal portion of each chain includes a variable region of about 100 to 110 or more amino acids primarily responsible for antigen recognition. The carboxy-terminal portion of the heavy chain may define a constant region primarily responsible for effector function. Typically, human light chains are classified as kappa and lambda light chains. Furthermore, human heavy chains are typically classified as mu, delta, gamma, alpha, or epsilon, and define the antibody’s isotype as IgM, IgD, IgG, IgA, and IgE, respectively. Within light and heavy chains, the variable and constant regions are joined by a “J” region of about 12 or more amino acids, with the heavy chain also including a “D” region of about 10 more amino acids. See generally, Fundamental Immunology Ch. 7 (Paul, W ., ed., 2nd ed. Raven Press, N.Y. (1989)).
[0040] As used herein, an “Fc domain” or “Fc region” each refer to the fragment crystallizable region of an antibody. The Fc domain comprises two heavy chain fragments comprising the CH2 and CH3 domains of an antibody. The two heavy chain fragments are held together by two or more disulfide bonds and by hydrophobic interactions of the CH3 domains. The Fc domain may be fused at the N-terminus or the C-terminus to a heterologous protein.
[0041] As used herein, “isolated” antibodies are at least partially free of other biological molecules from the cells or cell cultures in which they are produced. Such biological molecules include nucleic acids, proteins, lipids, or other material such as cellular debris and growth medium. An isolated antibody may further be at least partially free of expression system components such as biological molecules from a host cell or of the growth medium thereof. Generally, the term “isolated” is not intended to refer to a complete absence of such biological molecules or to an absence of water, buffers, or salts or to components of a pharmaceutical formulation that includes the antibodies.
[0042] As used herein, a “monoclonal antibody” refers to a population of substantially homogeneous antibodies, i.e., the antibody molecules comprising the population are identical in amino acid sequence except for possible naturally occurring mutations that may be present in minor amounts. In contrast, conventional (polyclonal) antibody preparations typically include a multitude of different antibodies having different amino acid sequences in their variable domains that are often specific for different epitopes. The modifier “monoclonal” indicates the characterof the antibody as being obtained from a substantially homogeneous population of antibodies; it is not to be construed as requiring production of the antibody by any particular method. For example, antibodies to be used in accordance with certain embodiments of the invention may be made by the hybridoma method first described by Kohler et al. (1975) Nature 256: 495, or they may be made by recombinant DNA methods (see, e.g., U.S. Pat. No. 4,816,567). Antibodies may also be isolated from phage antibody libraries using the techniques described in Clackson et al. (1991) Nature 352: 624-628 and Marks et al. (1991) J. Mol. Biol. 222: 581-597. See also Presta (2005) J. Allergy Clin. Immunol. 116:731.
[0043] As used herein, “conservatively modified variants” or “conservative substitution” refers to substitutions of amino acids with other amino acids having similar characteristics (e.g., charge, side-chain size, hydrophobicity / hydrophilicity, backbone conformation and rigidity, etc.), such that the changes can frequently be made without altering the biological activity of the protein. Those of skill in this art recognize that, in general, single amino acid substitutions in non- essential regions of a polypeptide do not substantially alter biological activity (see, e.g., Watson et al. (1987) Molecular Biology of the Gene, The Benjamin / Cummings Pub. Co., p. 224 (4th Ed.)). In addition, substitutions of structurally or functionally similar amino acids are less likely to disrupt biological activity. Exemplary conservative substitutions are set forth in Table A below.Table A
[0044] As used herein, “mutations” include substitutions (e.g., conservative substitutions).Mutations also include deletions and insertions (e.g., appearing as gaps in a sequence alignment).
[0045] As used herein, “isolated nucleic acid molecule” means a DNA or RNA of genomic, mRNA, cDNA, or synthetic origin or some combination thereof which is not associated with allor a portion of a polynucleotide in which the isolated polynucleotide is found in nature, or is linked to a polynucleotide to which it is not linked in nature. For purposes of this disclosure, it should be understood that “a nucleic acid molecule comprising” a particular nucleotide sequence does not encompass intact chromosomes. Isolated nucleic acid molecules “comprising” specified nucleic acid sequences may include, in addition to the specified sequences, coding sequences for up to ten or even up to twenty or more other proteins or portions or fragments thereof, or may include operably linked regulatory sequences that control expression of the coding region of the recited nucleic acid sequences, and / or may include vector sequences.
[0046] Any carbon or heteroatom with unsatisfied valences in the text, schemes, examples and tables herein is assumed to have sufficient hydrogen atom(s) to satisfy the valences. Any one or more of these hydrogen atoms can be deuterium.
[0047] Compounds herein may contain one or more stereogenic centers and can occur as racemates, racemic mixtures, single enantiomers, diastereomeric mixtures, and individual diastereomers. Additional asymmetric centers may be present depending upon the nature of the various substituents on the molecule. Each such asymmetric center will independently produce two optical isomers, and all possible optical isomers and diastereomers in mixtures and as pure or partially purified compounds are included within the disclosure. Any formulas, structures, or names of compounds described herein that do not specify a particular stereochemistry are meant to encompass any and all existing isomers as described above and mixtures thereof in any proportion. When stereochemistry is specified, the disclosure is meant to encompass that particular isomer in pure form or as part of a mixture with other isomers in any proportion.
[0048] Diastereomeric mixtures can be separated into their individual diastereomers on the basis of their physical chemical differences by methods well known to those skilled in the art, such as, for example, by chromatography and / or fractional crystallization. Enantiomers can be separated by converting the enantiomeric mixture into a diastereomeric mixture by reaction with an appropriate optically active compound (e.g., chiral auxiliary such as a chiral alcohol or Mosher’s acid chloride), separating the diastereomers and converting (e.g., hydrolyzing) the individual diastereomers to the corresponding pure enantiomers. Enantiomers can also be separated by use of chiral HPLC column.
[0049] All stereoisomers (for example, geometric isomers, optical isomers, and the like) of disclosed compounds (including those of the salts and solvates of compounds as well as the salts, solvates, and esters of prodrugs), such as those that may exist due to asymmetric carbons on various substituents, including enantiomeric forms (which may exist even in the absence of asymmetric carbons), rotameric forms, atropisomers, and diastereomeric forms, are contemplatedwithin the scope of this disclosure. Individual stereoisomers of compounds may, for example, be substantially free of other isomers, or may be admixed, for example, as racemates or with all other, or other selected, stereoisomers. The chiral centers can have the S or R configuration as defined by the IUPAC 1974 Recommendations.
[0050] The present disclosure further includes compounds and synthetic intermediates in all their isolated forms. For example, the above-identified compounds are intended to encompass all forms of the compounds such as, any solvates, hydrates, stereoisomers, and tautomers thereof.
[0051] Compounds can form salts that are also within the scope of this disclosure. Reference to a compound herein is understood to include reference to salts thereof, unless otherwise indicated. The term “salt(s),” as employed herein, denotes acidic salts formed with inorganic and / or organic acids, as well as basic salts formed with inorganic and / or organic bases. In addition, when a compound contains both a basic moiety, such as, but not limited to a pyridine or imidazole, and an acidic moiety, such as, but not limited to a carboxylic acid, zwitterions (“inner salts”) may be formed and are included within the term “salt(s)” as used herein. Pharmaceutically acceptable (i.e., non-toxic, physiologically acceptable) salts are preferred, although other salts are also useful. Salts of the compounds may be formed, for example, by reacting a compound with an amount of acid or base, such as an equivalent amount, in a medium such as one in which the salt precipitates or in an aqueous medium followed by lyophilization.
[0052] Exemplary acid addition salts include acetates, ascorbates, benzoates, benzenesulfonates, bisulfates, borates, butyrates, citrates, camphorates, camphorsulfonates, fumarates, hydrochlorides, hydrobromides, hydroiodides, lactates, maleates, methanesulfonates, naphthalenesulfonates, nitrates, oxalates, phosphates, propionates, salicylates, succinates, sulfates, tartarates, thiocyanates, toluenesulfonates (also known as tosylates,), and the like. Additionally, acids that are generally considered suitable for the formation of pharmaceutically useful salts from basic pharmaceutical compounds are discussed, for example, by P. Stahl et al., Camille G. (eds.) Handbook of Pharmaceutical Salts: Properties, Selection and Use (2002) Zurich: Wiley-VCH; S. Berge et al., J. Pharm. Sci. (1977) 66(1) 1-19; P. Gould, International J. of Pharmaceutics (1986) 33 201-217; Anderson et al., The Practice of Medicinal Chemistry (1996), Academic Press, New York; and in The Orange Book (Food & Drug Administration, Washington, D.C.). These disclosures are incorporated herein by reference thereto.
[0053] Exemplary basic salts include ammonium salts, alkali metal salts such as sodium, lithium, and potassium salts, alkaline earth metal salts such as calcium and magnesium salts, salts with organic bases (for example, organic amines) such as dicyclohexylamines, t-butyl amines, and salts with amino acids such as arginine, lysine, and the like. Basic nitrogen-containinggroups may be quarternized with agents such as lower alkyl halides (e.g., methyl, ethyl, and butyl chlorides, bromides and iodides), dialkyl sulfates (e.g., dimethyl, diethyl, and dibutyl sulfates), long chain halides (e.g., decyl, lauryl, and stearyl chlorides, bromides, and iodides), aralkyl halides (e.g., benzyl and phenethyl bromides), and others.
[0054] “Protein,” “polypeptide,” and “peptide” are used interchangeably herein to denote a polymer of at least two amino acids covalently linked by an amide bond, regardless of length or post-translational modification (e.g., glycosylation or phosphorylation, lipidation, myristoylation, ubiquitination, etc.). Included within this definition are d- and 1-amino acids, and mixtures of d- and 1-amino acids, as well as polymers comprising d- and 1-amino acids, and mixtures of d- and 1- amino acids. Proteins, polypeptides, and peptides may include a tag, such as a histidine tag, which should not be included when determining percentage of sequence identity.
[0055] “Amino acid” or “residue” as used in context of the polypeptides disclosed herein refers to the specific monomer at a sequence position. Amino acids are referred to herein by either their commonly known three-letter symbols or by the one-letter symbols recommended by IUPAC- IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single letter codes.
[0056] The abbreviations used for the genetically encoded amino acids are conventional and are as follows: alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartate (Asp or D), cysteine (Cys or C), glutamate (Glu or E), glutamine (Gin or Q), histidine (His or H), isoleucine (He or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Vai or V).
[0057] The abbreviations used for the genetically encoding nucleosides are conventional and are as follows: adenosine (A); guanosine (G); cytidine (C); thymidine (T); and uridine (U). Unless specifically delineated, the abbreviated nucleosides may be either ribonucleosides or 2'- deoxyribonucleosides. The nucleosides may be specified as being either ribonucleosides or 2'- deoxyribonucleosides on an individual basis or on an aggregate basis. When nucleic acid sequences are presented as a string of one-letter abbreviations, the sequences are presented in the 5' to 3' direction in accordance with common convention, and the phosphates are not indicated.
[0058] As used herein, “polynucleotide” and “nucleic acid’ refer to two or more nucleotides that are covalently linked together. The polynucleotide may be wholly comprised of ribonucleotides (i.e., RNA), wholly comprised of 2' deoxyribonucleotides (i.e., DNA), or comprised of mixtures of ribo- and 2' deoxyribonucleotides. While the nucleosides will typically be linked together via standard phosphodiester linkages, the polynucleotides may include one ormore non-standard linkages. The polynucleotide may be single-stranded or double-stranded, or the polynucleotide may include both single-stranded regions and double-stranded regions. Moreover, while a polynucleotide will typically be composed of the naturally occurring encoding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), it may include one or more modified and / or synthetic nucleobases, such as, for example, inosine, xanthine, hypoxanthine, etc. In some embodiments, such modified or synthetic nucleobases are nucleobases encoding amino acid sequences.
[0059] As used herein, “isolated polypeptide” refers to a polypeptide that is substantially separated from other contaminants that naturally accompany it (e.g., protein, lipids, and polynucleotides). The term embraces polypeptides that have been removed or purified from their naturally occurring environment or expression system (e.g., within a host cell or via in vitro synthesis). The recombinant polypeptides may be present within a cell, present in the cellular medium, or prepared in various forms, such as lysates or isolated preparations. As such, in some embodiments, the recombinant polypeptides can be an isolated polypeptide.
[0060] As used herein, “substantially pure polypeptide” or “purified protein” refers to a composition in which the polypeptide species is the predominant species present (i.e., on a molar or weight basis it is more abundant than any other individual macromolecular species in the composition), and is generally a substantially purified composition when the object species comprises at least about 50 percent of the macromolecular species present by mole or % weight. However, in some embodiments, an enzyme comprising composition comprises enzymes that are less than 50% pure (e.g., about 10%, about 20%, about 30%, about 40%, or about 50%).Generally, a substantially pure enzyme or polypeptide composition comprises about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more of all macromolecular species by mole or % weight present in the composition. In some embodiments, the object species is purified to essential homogeneity (i.e., contaminant species cannot be detected in the composition by conventional detection methods) wherein the composition consists essentially of a single macromolecular species. Solvent species, small molecules (<500 Daltons), and elemental ion species are not considered macromolecular species. In some embodiments, the isolated recombinant polypeptides are substantially pure polypeptide compositions.
[0061] As used herein, a “vector” is a DNA construct for introducing a DNA sequence into a cell. In some embodiments, the vector is an expression vector that is operably linked to a suitable control sequence capable of effecting the expression in a suitable host of the polypeptide encoded in the DNA sequence. In some embodiments, an “expression vector” has a promoter sequenceoperably linked to the DNA sequence (e.g., transgene) to drive expression in a host cell, and in some embodiments, also comprises a transcription terminator sequence.
[0062] As used herein, the term “expression” includes any step involved in the production of the polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from a cell.
[0063] As used herein, the term “produces” refers to the production of proteins and / or other compounds by cells. It is intended that the term encompass any step involved in the production of polypeptides including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from a cell.
[0064] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, signal peptide, terminator sequence, etc.) is “heterologous” to another sequence with which it is operably linked if the two sequences are not associated in nature. For example, a “heterologous polynucleotide” is any polynucleotide that is introduced into a host cell by laboratory techniques, and the term includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into a host cell.
[0065] As used herein, the terms “host cell” and “host strain” refer to suitable hosts for expression vectors comprising DNA provided herein (e.g., the polynucleotides encoding the variants). In some embodiments, the host cells are prokaryotic or eukaryotic cells that have been transformed or transfected with vectors constructed using recombinant DNA techniques as known in the art.
[0066] “Percentage of sequence identity,” “percent identity,” and “percent identical” are used herein to refer to comparisons between polynucleotide sequences or polypeptide sequences, and are determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps or truncations) as compared to the reference sequence for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which either the identical nucleic acid base or amino acid residue occurs in both sequences or a nucleic acid base or amino acid residue is aligned with a gap to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity. Determination of optimal alignment and percent sequence identity is performed using the BLAST and BLAST 2.0 algorithms (see e.g., Altschul et al.,1990, J. Mol. Biol. 215: 403-410; and Altschul et al., 1977, Nucleic Acids Res. 3389-3402). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website.
[0067] Briefly, the BLAST analyses involve first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as, the neighborhood word score threshold (Altschul etal., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, M = 5, N = -4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc. Natl. Acad. Sci. USA 89: 10915).
[0068] Numerous other algorithms are available that function similarly to BLAST in providing percent identity for two sequences. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math. 2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443, by the search for similarity method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Software Package), or by visual inspection (see generally, Current Protocols in Molecular Biology, F. M. Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). Additionally, determination of sequence alignment and percent sequence identity can employ the BESTFIT or GAP programs in the GCG Wisconsin Software package (Accelrys, Madison WI), using default parameters provided.
[0069] “ Substantial identity” refers to a polynucleotide or polypeptide sequence that has at least 80 percent sequence identity, preferably at least 85 percent sequence identity, more preferably at least 89 percent sequence identity, more preferably at least 95 percent sequence identity, and even more preferably at least 99 percent sequence identity as compared to a reference sequence over a comparison window of at least 20 residue positions, frequently over a window of at least 30-50 residues, wherein the percentage of sequence identity is calculated by comparing the reference sequence to a sequence that includes deletions or additions which total 20 percent or less of the reference sequence over the window of comparison. In specific embodiments applied to polypeptides, the term “substantial identity” means that two polypeptide sequences, when optimally aligned, such as by the programs GAP or BESTFIT using default gap weights, share at least 80 percent sequence identity, preferably at least 89 percent sequence identity, more preferably at least 95 percent sequence identity or more (e.g., 99 percent sequence identity). Preferably, residue positions which are not identical differ by conservative amino acid substitutions.
[0070] “Corresponding to”, “reference to” or “relative to” when used in the context of the numbering of a given amino acid or polynucleotide sequence refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue number or residue position of a given polymer is designated with respect to the reference sequence rather than by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, although the gaps are present, the numbering of the residue in the given amino acid or polynucleotide sequence is made with respect to the reference sequence to which it has been aligned.
[0071] The term “DAR” or “Drug Antibody Ratio,” as used in the Examples herein, refers to the average number of drug moieties attached to an antibody in a composition comprising more than one Antibody-Drug Conjugate molecule, wherein DAR is represented as a real number from 0 to 10. Accordingly, in one embodiment, for an Antibody -Drug Conjugate of the present disclosure, the DAR is an integer from 0 to 10, 0 to 9, 0 to 8, from 0 to 7, from 0 to 6, from 0 to 5, from 0 to 4, from 0 to 3, from 0 to 2, and from 0 to 1. In additional embodiments, for an Antibody-Drug Conjugate of the present disclosure, the DAR is an integer from 1 to 8, 1 to 4, 2 to 5, 3 to 6, 4 to 7, 5 to 8, and 6 to 8. In other embodiments, for an Antibody-Drug Conjugates of the present disclosure, the DAR is an integer from 1 to 3, 2 to 4, 3 to 5, 4 to 6, 5 to 7, and 6 to 8. In further embodiments, for an Antibody-Drug Conjugate of the present disclosure, the DAR isan integer from 1 to 2, 2 to 3, 3 to 4, 4 to 5, 5 to 6, 6 to 7, and 7 to 8. Likewise, for a composition comprising an Antibody-Drug Conjugate of the present disclosure, the DAR for the composition is an average of the DAR for all of the Antibody-Drug Conjugate molecules present in said composition. As such, for a composition comprising an Antibody-Drug Conjugate of the present disclosure, the DAR of the composition is a decimal from 0 to 8, from 0 to 7, from 0 to 6, from 0 to 5, from 0 to 4, from 0 to 3, from 0 to 2, and from 0 to 1. In additional embodiments, for a composition comprising an Antibody-Drug Conjugate of the present disclosure, the DAR of the composition is a decimal from 1 to 4, 2 to 5, 3 to 6, 4 to 7, 5 to 8, and 6 to 8. In other embodiments, for a composition comprising an Antibody-Drug Conjugate of the present disclosure, the DAR of the composition is a decimal from 1 to 3, 2 to 4, 3 to 5, 4 to 6, 5 to 7, and 6 to 8. In further embodiments, for a composition comprising an Antibody-Drug Conjugate of the present disclosure, the DAR of the composition is a decimal from 1 to 2, 2 to 3, 3 to 4, 4 to 5, 5 to 6, 6 to 7, and 7 to 8. The term “composition” as used above, is understood to encompass pharmaceutical compositions.Transglutaminases
[0072] Transglutaminases catalyze reactions between carboxamide groups (e.g., of glutamine residues) and amino groups (e.g., of lysine residues) to form an isopeptide bond. Thus, transglutaminase activity can be used to cross-link different molecules to each other.
[0073] In some aspects, the disclosed polypeptides have transglutaminase activity and can be used to conjugate a compound (e.g., having a primary amide group) specifically to a lysine moiety of an antibody. In some embodiments of the invention, polypeptides include those having the amino acid sequences of any one of SEQ ID NOs: 1-7 and 31-37. SEQ ID NOs: 31-37 differ from SEQ ID NOs: 1-7 by having a hexahistidine tag and a propeptide domain. In some embodiments, a polypeptide of the invention has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 conservative substitutions relative to any one of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide of the invention has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 conservative substitutions as compared to the amino acid sequence of any one of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide of the invention has 1 conservative substitution as compared to the amino acid sequence of any one of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide of the invention has 2 conservative substitutions, as compared to the amino acid sequence of any one of SEQ ID NOs: 1-7 and 31- 37. In some embodiments, a polypeptides has 3 conservative substitutions, relative to the amino acid sequences of any one of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptidehas 4 conservative substitution, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 5 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 6 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 7 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 8 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 9 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 10 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 11 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 12 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 13 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 14 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 15 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 16 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 17 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 18 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 19 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. In some embodiments, a polypeptide has 20 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 1-7 or 31-37. The invention further relates to any of the above described polypeptides, wherein the polypeptide has transglutaminase activity and can be used to conjugate a compound (e.g., having an amide group) specifically to lysine at position 222 of an antibody.
[0074] In some embodiments, the polypeptide comprises conservative substitutions, wherein all of the conservative substitutions are outside of the active site residues of the transglutaminase. In certain embodiments, polypeptides of the invention are enzymes that were evolved for K222modification using probes with primary amide functionality as shown in FIGURE. 4. One example of a compound that can be used for testing this reaction is 3-(4-(3-azidopropyl)-3,6- dioxopiperazin-2-yl) propenamide.
[0075] In some aspects, the disclosed polypeptides have transglutaminase activity and can be used to conjugate a compound (e.g., having a primary amine group) specifically to a glutamine moiety of an antibody. In some embodiments of the invention, the polypeptides that can be used to catalyze such a reaction include those having the amino acid sequences set forth in any one of SEQ ID NOs: 11-23 or 41-53. SEQ ID NOs: 41-53 differ from SEQ ID NOs: 11-23 by having a hexahistidine tag and a propeptide domain. In some embodiments of the invention, a polypeptide include those having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 conservative substitutions relative to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide of the invention has 1 conservative substitution as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 and 41-53. In some embodiments, a polypeptides of the invention has 2 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptides has 3 conservative substitutions, as compared to the any one of the amino acid sequences of SEQ ID NOs: 11-23 and 41-53. In some embodiments, a polypeptide has 4 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 5 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 6 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 7 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 8 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 9 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 10 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 11 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 12 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 13 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 14 conservative substitutions, ascompared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 15 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 16 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 17 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 18 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 19 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. In some embodiments, a polypeptide has 20 conservative substitutions, as compared to any one of the amino acid sequences of SEQ ID NOs: 11-23 or 41-53. The invention further relates to any of the above described polypeptides, wherein the polypeptide has transglutaminase activity and can be used to conjugate a compound (e.g., having a primary amine group) specifically to glutamine at position 295 of an antibody.
[0076] In some embodiments, the polypeptide comprises conservative substitutions, wherein all of the conservative substitutions are outside of the active site residues of the transglutaminase. In certain embodiments, polypeptides of the invention are enzymes that were evolved for mAb Q295 modification using probes with primary amine functionality as shown in FIGURE 1. One example of a compound that can be used for testing this reaction is 3-(4-aminobutyl)-l-(3- azidopropyl)piperazine-2, 5-dione.
[0077] In another aspect, the present disclosure provides polynucleotides encoding the polypeptides (i.e., enzymes) disclosed herein. The polynucleotides may be operatively linked to one or more heterologous regulatory sequences that control gene expression to create a recombinant polynucleotide capable of expressing the polypeptide. Expression constructs containing a heterologous polynucleotide encoding the transglutaminase can be introduced into appropriate host cells to express the corresponding transglutaminase polypeptide.
[0078] A “triplet” codon of four possible nucleotide bases can exist in over 60 variant forms. Because these codons provide the message for only 20 different amino acids (as well as transcription initiation and termination), some amino acids can be coded for by more than one codon, a phenomenon known as codon redundancy. Thus, having identified a particular amino acid sequence, those skilled in the art could make numerous different nucleic acid sequences that encode a polypeptide having a specific amino acid sequence of the invention. In this regard, the present disclosure specifically contemplates any polynucleotide that encodes any one of the polypeptides disclosed herein.
[0079] The invention also relates to nucleic acid sequences that encode a transglutaminase of the invention. In some embodiments of this aspect of the invention, the nucleic acid sequence encodes an enzyme capable of conjugating a compound to position K222 of an antibody, wherein the nucleic acid sequence is at least 90% identical to any one of SEQ ID NOs: 71-77. In some embodiments, the nucleic acid sequence is at least 92% identical to any one of SEQ ID NOs: 71- 77. In some embodiments, the nucleic acid sequence is at least 94% identical to any one of SEQ ID NOs: 71-77. In some embodiments, the nucleic acid sequence is at least 96% identical to any one of SEQ ID NOs: 71-77. In some embodiments, the nucleic acid sequence is at least 98% identical to any one of SEQ ID NOs: 71-77. In some embodiments, the nucleic acid molecule comprises a sequence of nucleic acids as set forth in any one of SEQ ID NOs: 71-77. In some embodiments, the nucleic acid molecule consists essentially of a sequence of nucleic acids as set forth in any one of SEQ ID NOs: 71-77. In some embodiments, the nucleic acid molecule consists of a sequence of nucleic acids as set forth in any one of SEQ ID NOs: 71-77.
[0080] The invention further relates to nucleic acid sequences that encode a transglutaminase that is capable of conjugating to position Q295 of an antibody, wherein the nucleic acid sequence is at least 90% identical to any one of SEQ ID NOs: 81-93. In some embodiments, the nucleic acid sequence is at least 92% identical to any one of SEQ ID NOs: 81-93. In some embodiments, the nucleic acid sequence is at least 94% identical to any one of SEQ ID NOs: 81-93. In some embodiments, the nucleic acid sequence is at least 96% identical to any one of SEQ ID NOs: 81- 93. In some embodiments, the nucleic acid sequence is at least 98% identical to any one of SEQ ID NOs: 81-93. In some embodiments, the nucleic acid molecule comprises a sequence of nucleotides as set forth in any one of SEQ ID NOs: 81-93. In some embodiments, the nucleic acid molecule consists essentially of a sequence of nucleotides as set forth in any one of SEQ ID NOs: 81-93. In some embodiments, the nucleic acid molecule consists of a sequence of nucleotides as set forth in any one of SEQ ID NOs: 81-93.
[0081] As understood by one of skill in the art, synthetic genes which have been designed to include a projected host cell's preferred codons provide an optimal form of foreign genetic material for practice of recombinant protein expression. Thus, one aspect of this invention is a nucleic acid molecule encoding a transglutaminase described herein that is codon-optimized for high-level expression in a host cell of choice. Thus, in various embodiments, the codons of the nucleic acid molecules of the invention are selected to align with the host cell in which the protein is being produced. For example, preferred codons used in bacteria are used to express the nucleic acid molecule in bacterial host cells; preferred codons used in yeast are used to expressthe nucleic acid molecule in yeast host cells; and preferred codons used in mammals are used to express the nucleic acid molecule in mammalian host cells.
[0082] In certain embodiments, less than 100% of the codons in the nucleic acid molecules of the invention are replaced to optimize the codon usage. Consequently, codon optimized polynucleotides encoding a transglutaminase enzyme of the invention may contain preferred codons at about 40%, 50%, 60%, 70%, 80%, 90% or greater than 90% of codon positions of the full-length coding region.
[0083] In various embodiments, an isolated polynucleotide encoding an improved transglutaminase polypeptide may be manipulated in a variety of ways to provide for optimal expression of the polypeptide. Manipulation of the isolated polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector. Techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well known in the art. Guidance is provided in Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, 3rd Ed., Cold Spring Harbor Laboratory Press; and Current Protocols in Molecular Biology, Ausubel. F. ed., Greene Pub. Associates, 1998, updates to 2006.
[0084] In some embodiments, an isolated polynucleotide encoding any of the transglutaminase polypeptides herein is manipulated in a variety of ways to facilitate expression of the transglutaminase polypeptide. In some embodiments, the invention relates to an expression vector comprising a nucleic acid sequence that encodes a transglutaminase of the invention and one or more control sequences, wherein the one or more control sequences is present to regulate the expression of the transglutaminase polynucleotides and / or polypeptides. Manipulation of the isolated polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector utilized. Techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well known in the art. In some embodiments, the control sequences include among others, promoters, leader sequences, polyadenylation sequences, propeptide sequences, signal peptide sequences, and transcription terminators. In some embodiments, suitable promoters are selected based on the host cells selection. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure, include, but are not limited to, promoters obtained from the E. coli lac operon, Streptomyces coelicolor agarase gene (dagA), Bacillus subtilis levansucrase gene (sacB), Bacillus licheniformis alpha-amylase gene (amyL), Bacillus stearothermophilus maltogenic amylase gene (amyM), Bacillus amyloliquefaciens alpha-amylase gene (amyQ), Bacillus licheniformis penicillinase gene (penP), Bacillus subtilis xylA and xylB genes, and prokaryotic beta-lactamase gene (See e.g., Villa-Kamaroff et al., Proc. Natl Acad. Sci.USA 75: 3727-3731 (1978]), as well as the tac promoter (See e.g., DeBoer et al., Proc. Natl Acad. Sci. USA 80: 21-25 (1983]). Exemplary promoters for filamentous fungal host cells, include, but are not limited to, promoters obtained from the genes for Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral alpha-amylase, Aspergillus niger acid stable alpha-amylase, Aspergillus niger o Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (See e.g, WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the genes for Aspergillus niger neutral alpha-amylase and Aspergillus oryzae triose phosphate isomerase), and mutant, truncated, and hybrid promoters thereof. Exemplary yeast cell promoters can be from the genes can be from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GALI), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3 -phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3 -phosphoglycerate kinase. Other useful promoters for yeast host cells are known in the art (See e.g., Romanos et al., Yeast 8:423-488 (1992)).
[0085] In some embodiments, the control sequence is also a suitable transcription terminator sequence (i.e., a sequence recognized by a host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' terminus of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable terminator that is functional in the host cell of choice finds use in the invention. Exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger alpha-glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the genes for Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3 -phosphate dehydrogenase. Other useful terminators for yeast host cells are known in the art (See e.g., Romanos et al., supra).
[0086] In some embodiments, the control sequence is also a suitable leader sequence (i.e., a non-translated region of an mRNA that is important for translation by the host cell). In some embodiments, the leader sequence is operably linked to the 5' terminus of the nucleic acid sequence encoding the transglutaminase. Any suitable leader sequence that is functional in the host cell of choice find use in the invention. Exemplary leaders for filamentous fungal host cells are obtained from the genes for Aspergillus oryzae TAKA amylase, and Aspergillus nidulans triose phosphate isomerase. Suitable leaders for yeast host cells are obtained from the genes forSaccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3 -phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0087] In some embodiments, the control sequence is also a polyadenylation sequence (i.e., a sequence operably linked to the 3' terminus of the nucleic acid sequence and which, when transcribed, is recognized by the host cell as a signal to add polyadenosine residues to transcribed mRNA). Any suitable polyadenylation sequence that is functional in the host cell of choice may be used in the invention. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsinlike protease, and Aspergillus niger alpha-glucosidase. Useful polyadenylation sequences for yeast host cells are known (See e.g., Guo and Sherman, Mol. Cell. Biol., 15:5983-5990 (1995)).
[0088] In some embodiments, the control sequence is also a signal peptide (i.e., a coding region that codes for an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the cell’s secretory pathway). In some embodiments, the 5' end of the coding sequence of the nucleic acid sequence inherently contains a signal peptide coding region naturally linked in translation reading frame with the segment of the coding region that encodes the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence contains a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of a host cell of choice finds use for expression of the engineered polypeptide(s). Effective signal peptide coding regions for bacterial host cells are the signal peptide coding regions include, but are not limited to, those obtained from the genes for Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are known in the art (See e.g., Simonen and Palva, Microbiol. Rev., 57: 109-137 (1993)). In some embodiments, effective signal peptide coding regions for filamentous fungal host cells include, but are not limited to, the signal peptide coding regions obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase. Useful signal peptides for yeast host cells include, but are not limited to, those from the genes for Saccharomyces cerevisiae alpha-factor and Saccharomyces cerevisiae invertase.
[0089] In some embodiments, regulatory sequences are also utilized. These sequences facilitate the regulation of the expression of the polypeptide relative to the growth of the host cell. Examples of regulatory systems are those that cause the expression of the gene to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, lac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or GALI system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, Aspergillus niger glucoamylase promoter, and Aspergillus oryzae glucoamylase promoter.
[0090] In another aspect, the invention is directed to a recombinant expression vector comprising a polynucleotide encoding a transglutaminase polypeptide as described herein (e.g., a polynucleotide encoding any one of SEQ ID NO’s 1-7, 11-23, 31-37 and 41-53), and one or more expression regulating regions such as a promoter and a terminator, a replication origin, etc., depending on the type of hosts into which they are to be introduced. In some embodiments, the various nucleic acid and control sequences described herein are joined together to produce recombinant expression vectors that include one or more convenient restriction sites to allow for insertion or substitution of the nucleic acid sequence encoding the enzyme polypeptide at such sites. Alternatively, in some embodiments, a nucleic acid sequence of the invention is expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate vector for expression. In some embodiments, the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression.
[0091] The recombinant expression vector may be any suitable vector (e.g., a plasmid or virus), that can be conveniently subjected to recombinant DNA procedures and bring about the expression of the enzyme polynucleotide sequence. The choice of the vector typically depends on the compatibility of the vector with the host cell into which the vector is to be introduced. The vectors may be linear or closed circular plasmids.
[0092] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extra-chromosomal entity, the replication of which is independent of chromosomal replication, such as a plasmid, an extra-chromosomal element, a minichromosome, or an artificial chromosome). The vector may contain any means for assuring self-replication. In some alternative embodiments, the vector is one in which, when introduced into the host cell, it is integrated into the genome and replicated together with the chromosome(s) into which it has been integrated. Furthermore, in some embodiments, a single vector or plasmid, or two or morevectors or plasmids which together contain the total DNA to be introduced into the genome of the host cell, and / or a transposon is utilized.
[0093] In some embodiments, the expression vector contains one or more selectable markers, which permit easy selection of transformed cells. A “selectable marker” is a gene, the product of which provides biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like. Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis. or markers, which confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase; e.g., from A. nidulans or A. orzyae), argB (ornithine carbamoyltransferases), bar (phosphinothricin acetyltransferase; e.g, from S. hygroscopicus), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase; e.g, from A. nidulans or A. orzyae), sC (sulfate adenyltransferase), and trpC (anthranilate synthase), as well as equivalents thereof.
[0094] In another aspect, the invention provides a host cell comprising at least one polynucleotide encoding at least one transglutaminase of the invention, the polynucleotide(s) being operatively linked to one or more control sequences for expression of the at least one transglutaminase in the host cell. Host cells suitable for use in expressing the polypeptides encoded by the expression vectors of the invention are well known in the art and include but are not limited to, bacterial cells, such as E. coli, Vibrio fluvialis. Streptomyces and Salmonella typhimurium cells; fungal cells, such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178));. Exemplary host cells also include various Escherichia coli strains (e.g., W3110 (AfhuA) and BL21). Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis, or markers, which confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, and or tetracycline resistance.
[0095] In some embodiments, an expression vector of the invention contains an element(s) that permits integration of the vector into the host cell’s genome or autonomous replication of the vector in the cell independent of the genome. In some embodiments, the vectors rely on the nucleic acid sequence encoding the polypeptide or any other element of the vector for integration of the vector into the genome by homologous or nonhomologous recombination.
[0096] In some alternative embodiments, the expression vectors contain additional nucleic acid sequences for directing integration by homologous recombination into the genome of the hostcell. The additional nucleic acid sequences enable the vector to be integrated into the host cell genome at a precise location(s) in the chromosome(s). To increase the likelihood of integration at a precise location, the integrational elements may contain a number of nucleotides, such as 100 to 10,000 base pairs, 400 to 10,000 base pairs, or 800 to 10,000 base pairs, which are highly homologous with the corresponding target sequence to enhance the probability of homologous recombination. The integrational elements may be any sequence that is homologous with the target sequence in the genome of the host cell. Furthermore, the integrational elements may be non-encoding or encoding nucleic acid sequences. On the other hand, the vector may be integrated into the genome of the host cell by non-homologous recombination.
[0097] For autonomous replication, the vector may further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. Examples of bacterial origins of replication are Pl 5 A ori or the origins of replication of plasmids pBR322, pUC19, pACYC177 (which contains the P15A ori), or pACYC184 (which contains the P15A ori) permitting replication in E. coli, and pUB 110, pE194, or pTA1060 permitting replication in Bacillus. Examples of origins of replication for use in a yeast host cell are the 2 micron origin of replication, ARS1, ARS4, the combination of ARS1 and CEN3, and the combination of ARS4 and CEN6. The origin of replication may be one having a mutation which makes its functioning temperature-sensitive in the host cell (See e.g., Ehrlich, Proc. Natl. Acad. Sci. USA 75: 1433 (1978)).
[0098] In some embodiments, more than one copy of a nucleic acid sequence of the invention is inserted into the host cell to increase production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the nucleic acid sequence where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the nucleic acid sequence, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.
[0099] Many expression vectors useful in the invention are commercially available. Suitable commercial expression vectors include, but are not limited to, the Novagen® pET E. coli T7 expression vectors (Millipore Sigma) and the p3xFLAGTM™ expression vectors (Sigma- Aldrich Chemicals). Other suitable expression vectors include, but are not limited to, pBluescriptll SK(-) and pBK-CMV (Stratagene), and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen) or pPoly (See e.g., Lathe et al., Gene 57: 193-201 (1987)).
[0100] Thus, in some embodiments, a vector comprising a sequence encoding at least one variant transglutaminase of the invention is transformed into a host cell in order to allow propagation of the vector and expression of the variant transglutaminase(s). In some embodiments, the transformed host cell described above is cultured in a suitable nutrient medium under conditions permitting the expression of the variant transglutaminase(s). Any suitable medium useful for culturing the host cells finds use in the invention, including, but not limited to minimal or complex media containing appropriate supplements. In some embodiments, host cells are grown in HTP media. Suitable media are available from various commercial suppliers or may be prepared according to published recipes (e.g., in catalogues of the American Type Culture Collection).
[0101] In another aspect, the present disclosure provides a host cell comprising a polynucleotide encoding an improved transglutaminase polypeptide of the invention, the polynucleotide being operatively linked to one or more control sequences for expression of the transglutaminase enzyme in the host cell. Host cells for use in expressing the transglutaminase polypeptides of the invention are well known in the art and include but are not limited to, bacterial cells, such as E. coli, B. sublilis. B. licheniformis, B. megaterium, B. stearothermophilus, B. amyloliquefciciens. Lactobacillus kejir. Lactobacillus brevis, Lactobacillus minor, Streptomyces and Salmonella typhimurium cells; fungal cells, such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178)).Appropriate culture mediums and growth conditions for the above-described host cells are well known in the art.
[0102] Polynucleotides for expression of the transglutaminases may be introduced into cells by various methods known in the art. Techniques include among others, electroporation, biolistic particle bombardment, liposome mediated transfection, calcium chloride transfection, and protoplast fusion. Various methods for introducing polynucleotides into cells will be apparent to the skilled artisan.
[0103] In some embodiments of the invention, the host cell is a filamentous fungus of any suitable genus and species, including, but not limited to Achlya, Acremonium, Aspergillus, Aureobasidium, Bjerkandera, Ceriporiopsis, Cephalosporium, Chrysosporium, Cochlioboluy Corynascus, Cryphonectria, Cryptococcus, Coprinus, Coriolus, Diplodia, Endothis, Fusarium, Gibberella, Gliocladium, Humicola, Hypocrea, Myceliophthora, Mucor, Neurospora, Penicillium, Podospora, Phlebia, Piromyces, Pyricularia, Rhizomucor, Rhizopus, Schizophyllum, Scytalidium, Sporotrichum, Talaromyces, Thermoascus, Thielavia, Trametes, Tolypocladium,Trichoderma, Verticillium, and / or Volvariella, and / or teleomorphs, or anamorphs, and synonyms, basionyms, or taxonomic equivalents thereof.
[0104] In some embodiments of the invention, the host cell is a yeast cell, including but not limited to cells of Candida, Hansenula, Saccharomyces, Schizosaccharomyces, Pichia, Kluyveromyces, or Yarrowia species. In some embodiments of the invention, the yeast cell is Hansenula polymorpha, Saccharomyces cerevisiae, Saccharomyces car Isber gensis, Saccharomyces diastaticus, Saccharomyces norbensis, Saccharomyces kluyveri, Schizosaccharomyces pombe, Pichia pastoris, Pichia fmlandica, Pichia trehalophila, Pichia kodamae, Pichia membranaefaciens, Pichia opuntiae, Pichia thermotolerans, Pichia salictaria, Pichia quercuum, Pichia pijperi, Pichia stipitis, Pichia methanolica, Pichia angusta, Kluyveromyces lactis, Candida albicans, or Yarrowia lipolytica.
[0105] In some other embodiments, the host cell is a prokaryotic cell. Suitable prokaryotic cells include, but are not limited to, Gram-positive, Gram-negative and Gram-variable bacterial cells. Any suitable bacterial organism finds use in the invention, including but not limited to Agrobacterium, Alicyclobacillus, Anabaena, Anacystis, Acinetobacter, Acidothermus, Arthrobacter, Azobacter, Bacillus, Bifidobacterium, Brevibacterium, Butyrivibrio, Buchnera, Campestris, Camplyobacter, Clostridium, Corynebacterium, Chromatium, Coprococcus, Escherichia, Enterococcus, Enterobacter, Erwinia, Fusobacterium, Faecalibacterium, Francisella, Flavobacterium, Geobacillus, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Lactococcus, Ilyobacter, Micrococcus, Microbacterium, Mesorhizobium, Methylobacterium, Methylobacterium, Mycobacterium, Neisseria, Pantoea, Pseudomonas, Prochlorococcus, Rhodobacter, Rhodopseudomonas, Rhodopseudomonas, Roseburia, Rhodospirillum, Rhodococcus, Scenedesmus, Streptomyces, Streptococcus, Synecoccus, Saccharomonospora, Staphylococcus, Serratia, Salmonella, Shigella, Thermoanaerobacterium, Tropheryma, Tularensis, Temecula, Thermosynechococcus, Thermococcus, Ureaplasma, Xanthomonas, Xylella, Yersinia and Zymomonas. In some embodiments, the host cell is a species of Agrobacterium, Acinetobacter, Azobacter, Bacillus, Bifidobacterium, Buchnera, Geobacillus, Campylobacter, Clostridium, Corynebacterium, Escherichia, Enterococcus, Erwinia, Flavobacterium, Lactobacillus, Lactococcus, Pantoea, Pseudomonas, Staphylococcus, Salmonella, Streptococcus, Streptomyces, or Zymomonas. In some embodiments, the bacterial host strain is non-pathogenic to humans. In some embodiments the bacterial host strain is an industrial strain. Numerous bacterial industrial strains are known and suitable in the invention. In some embodiments of the invention, the bacterial host cell is n Agrobacterium species (e.g., A. radiobacter, A. rhizogenes, and A. rubi). In some embodiments of the invention, the bacterialhost cell is an Arthrobacter species (e.g, A. aurescens, A. citreus, A. globiformis, A. hydrocarboglutamicus, A. mysorens, A. nicolianae, A. paraffmeus, A. protophonniae , A. roseoparqffmus, A. sulf r eus, and A. ureafaciens). In some embodiments of the invention, the bacterial host cell is a Bacillus species (e.g., B. thuringensis, B. anthracis, B. megaterium, B. subtilis, B. lentus, B. circulans, B. pumilus, B. lautus, B.coagulans, B. brevis, B.firmus, B. alkaophius, B. licheniformis, B. clausii, B. stearothermophilus, B. halodurans, and B. amyloliquefaciens . In some embodiments, the host cell is an industrial Bacillus strain including but not limited to B. subtilis, B. pumilus, B. licheniformis, B. megaterium, B. clausii, B. stearothermophilus, or B. amyloliquefaciens. In some embodiments, the Bacillus host cells are B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, and / or B. amyloliquefaciens. In some embodiments, the bacterial host cell is a Clostridium species (e.g., C. acetobutylicum, C. tetani E88, C. lituseburense, C. saccharobutylicum, C. perfringens, and C. beijerinckii). In some embodiments, the bacterial host cell is a Corynebacterium species e.g., C. glutamicum and C. acetoacidophilum). In some embodiments the bacterial host cell is an Escherichia species e.g., E. coli). In some embodiments, the host cell is Escherichia coli W3110. In some embodiments the host is Escherichia coli BL21 or BL21(DE3). In some embodiments, the bacterial host cell is n Erwinia species e.g., E. uredovora, E. carotovora, E. ananas, E. herbicola, E. punctata, and E. terreus). In some embodiments, the bacterial host cell is Pantoea species e.g., P. citrea, and P. agglomerans). In some embodiments the bacterial host cell is a Pseudomonas species e.g., P. putida, P. aeruginosa, P. mevalonii, and P. sp. D-01 10). In some embodiments, the bacterial host cell is a Streptococcus species e.g., S. equisimiles, S. pyogenes, and S. uberis). In some embodiments, the bacterial host cell is a Streptomyces species e.g., S. ambofaciens, S. achromogenes, S. avermitilis, S. coelicolor, S. aureofaciens, S. aureus, S. fungicidicus, S. griseus, and S. lividans). In some embodiments, the bacterial host cell is a Zymomonas species e.g., Z. mobilis, and Z. lipolytica).
[0106] Many prokaryotic and eukaryotic strains that find use in the invention are readily available to the public from a number of culture collections such as American Type Culture Collection (ATCC), Deutsche Sammlung von Mikroorganismen und Zellkulturen GmbH (DSM), Centraalbureau Voor Schimmel cultures (CBS), and Agricultural Research Service Patent Culture Collection, Northern Regional Research Center (NRRL).
[0107] In some embodiments, host cells are genetically modified to have characteristics that improve protein secretion, protein stability and / or other properties desirable for expression and / or secretion of a protein. Genetic modification can be achieved by genetic engineering techniques and / or classical microbiological techniques e.g., chemical or UV mutagenesis and subsequentselection). Indeed, in some embodiments, combinations of recombinant modification and classical selection techniques are used to produce the host cells. Using recombinant technology, nucleic acid molecules can be introduced, deleted, inhibited or modified, in a manner that results in increased yields of transglutaminase variant(s) within the host cell and / or in the culture medium. In one genetic engineering approach, homologous recombination is used to induce targeted gene modifications by specifically targeting a gene in vivo to suppress expression of the encoded protein. In alternative approaches, siRNA, antisense and / or ribozyme technology find use in inhibiting gene expression. A variety of methods are known in the art for reducing expression of protein in cells, including, but not limited to deletion of all or part of the gene encoding the protein and site-specific mutagenesis to disrupt expression or activity of the gene product. (See e.g., Chaveroche et al., Nucl. Acids Res., 28:22 e97 (2000)); Cho et al., Molec. Plant Microbe Interact., 19:7-15 (2006); Maruyama and Kitamoto, Biotechnol. Lett., 30: 1811- 1817 (2008); Takahashi et al., Mol. Gen. Genom., 272: 344-352 (2004); and You et al., Arch. Microbiol. ,191 :615-622 (2009), all of which are incorporated by reference herein). Random mutagenesis, followed by screening for desired mutations also finds use (See e.g., Combier et al., FEMS Microbiol. Lett., 220: 141-8 (2003); and Firon et al., Eukary. Cell 2:247-55 (2003), both of which are incorporated by reference).
[0108] Introduction of a vector or DNA construct into a host cell can be accomplished using any suitable method known in the art, including but not limited to calcium phosphate transfection, DEAE-dextran mediated transfection, PEG-mediated transformation, electroporation, or other common techniques known in the art.
[0109] In some embodiments, the engineered host cells (i.e., “recombinant host cells”) of the invention are cultured in conventional nutrient media modified as appropriate for activating promoters, selecting transformants, or amplifying the transglutaminase polynucleotide. Culture conditions, such as temperature, pH and the like, are those previously used with the host cell selected for expression, and are well-known to those skilled in the art. As noted, many standard references and texts are available for the culture and production of cells, including cells of bacterial, plant, animal (especially mammalian) and archebacterial origin.
[0110] In some embodiments, cells expressing a transglutaminase of the invention are grown under batch or continuous fermentations conditions. Classical “batch fermentation” is a closed system, wherein the compositions of the medium are set at the beginning of the fermentation and is not subject to artificial alternations during the fermentation. A variation of the batch system is a “fed-batch fermentation” that also finds use in the invention. In this variation, the substrate is added in increments as the fermentation progresses. Fed-batch systems are useful when cataboliterepression is likely to inhibit the metabolism of the cells and where it is desirable to have limited amounts of substrate in the medium. Batch and fed-batch fermentations are common and well known in the art. “Continuous fermentation” is an open system where a defined fermentation medium is added continuously to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the cultures at a constant high density where cells are primarily in log phase growth. Continuous fermentation systems strive to maintain steady state growth conditions. Methods for modulating nutrients and growth factors for continuous fermentation processes as well as techniques for maximizing the rate of product formation are well known in the art of industrial microbiology. [OHl] In some embodiments of the invention, cell-free transcription and translation systems find use in producing the transglutaminase(s). Several systems are commercially available, and the methods are well-known to those skilled in the art.Biomolecules
[0112] The disclosed transglutaminases can be used for conjugation reactions involving various biomolecules that have primary amine or primary amide groups. In some aspects, the biomolecules are peptides or polypeptides. In some embodiments, the primary amine groups are those of lysine side chains. In some embodiments, the primary amide groups are those of glutamine side chains.
[0113] In some aspects, the biomolecules are antibodies. In some embodiments, the antibodies are of an IgG isotype. In some embodiments, the antibodies comprise IgGl constant domains. In some embodiments, the antibodies comprise IgG2 constant domains. In some embodiments, the antibodies comprise IgG3 constant domains. In some embodiments, the antibodies comprise IgG4 constant domains.
[0114] In some embodiments, the antibodies are glycosylated at position N297 of the antibody. In some embodiments, the glycosylation comprises GOF, GIF, or G2F species.
[0115] In some embodiments, the antibodies comprise constant domains that are, individually or collectively, at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the corresponding sequences of SEQ ID NO: 61 or SEQ ID NO: 62. SEQ ID NO: 61 is the heavy chain of Her2 IgG and SEQ ID NO: 62 is the light chain Her2 IgG. In some embodiments, the antibodies comprise constant domains that are, individually or collectively, at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical tothe corresponding sequences of SEQ ID NO: 65 or SEQ ID NO: 66. SEQ ID NO: 65 is the heavy chain of WT-Trop2 and SEQ ID NO: 66 is the light chain TW-Trop2.Compounds
[0116] As noted before, some of the enzymes were evolved for antibody K222 modification using probes with primary amide functionality as shown in FIGURE 4. The following are some examples of other conjugation substrates (i.e., compounds) that reacted successfully with these enzymes: Biotin (Biotin K, Biotin Q, Biotin A), Lys-Ala-Ala-PABC-MMAE,
[0117] Many other compounds having such primary amide groups can be used in these reactions. In some embodiments, for synthesizing ADCs, compounds that have a linker and a drug (e.g., a small molecule drug) are used. The linker portions of these compounds include the primary amide group at one end for conjugation. At the other end, the linkers are bonded to the drug. In some embodiments, the linkers further include a spacer (e.g., at the end that bonds to the drug), which may be cleaved once the ADC reaches its target.
[0118] As noted before, some of the enzymes were evolved for antibody Q295 modification using probes with primary amine functionality as shown in FIGURE 1. The following are some examples of other tested conjugation substrates (i.e., compounds) that reacted successfully with t
[0119] Many other compounds having such primary amine groups can be used in these reactions. In some embodiments, for synthesizing ADCs, compounds that have a linker and a drug (e.g., a small molecule drug) are used. The linker portions of these compounds include the primary amine group at one end for conjugation. At the other end, the linkers are bonded to thedrug. In some embodiments, the linkers further include a spacer (e.g., at the end that bonds to the drug), which may be cleaved once the ADC reaches its target.
[0120] In some embodiments, the compounds comprise a linker without a drug. In such embodiments, the drug can then be conjugated to the linker in a separate reaction.Methods and Uses
[0121] The disclosed polypeptides can be used for making (i.e., synthesizing) biomoleculecompound conjugates.
[0122] For example, the polypeptides comprising a sequence of any one of SEQ ID NOs: 1-7 and 31-37 can be used to catalyze a reaction between a lysine (e.g., K222 of an antibody) of a biomolecule and a compound that comprises a primary amide group, to thereby make a biomolecule-compound conjugate. In some embodiments, the compound comprises a drug. In some embodiments, the compound further comprises a linker, wherein the linker comprises said primary amide group. In some embodiments, the linker comprises a spacer cleavable by an enzyme endogenous to a cell.
[0123] In some embodiments, the biomolecule comprises a polypeptide. In some embodiments, the biomolecule is an antibody. In some embodiments, the antibody is glycosylated at position N297, per the Eu numbering scheme. In some embodiments, the compound is conjugated to a lysine side chain of the biomolecule. In some embodiments, the compound is predominantly conjugated to K222 of an antibody, per the Eu numbering scheme. In some embodiments, more than 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the biomolecule-compound conjugates made by the method comprise a conjugation at K222 of the antibody, per the Eu numbering scheme.
[0124] For example, a polypeptide comprising an amino acid sequence of any one of SEQ ID NOs: 11-23 and 41-53 can be used to catalyze a reaction between a glutamine (e.g., Q295 of an antibody) of a biomolecule and a compound that comprises a primary amine group, to thereby make a biomolecule-compound conjugate. In some embodiments, the compound comprises a drug. In some embodiments, the compound further comprises a linker, wherein the linker comprises said primary amine group. In some embodiments, the linker comprises a spacer cleavable by an enzyme endogenous to a cell.
[0125] In some embodiments, the biomolecule comprises a polypeptide. In some embodiments, the biomolecule is an antibody. In some embodiments, the antibody is glycosylated at position N297 of the antibody. In some embodiments, the compound is conjugated to a glutamine side chain of the biomolecule. In some embodiments, the compound is predominantly conjugated toQ295 of an antibody, per the Eu numbering scheme. In some embodiments, more than 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the biomolecule-compound conjugates made by the method comprise a conjugation at Q295 of the antibody, per the Eu numbering scheme.
[0126] During the course of the reduction reactions, the pH of the reaction mixture may change. The pH of the reaction mixture may be maintained at a desired pH or within a desired pH range by the addition of an acid or a base during the course of the reaction. Alternatively, the pH may be controlled by using an aqueous solvent that comprises a buffer. Suitable buffers to maintain desired pH ranges are known in the art and include, for example, phosphate buffer, triethanolamine buffer, and the like. Combinations of buffering and acid or base addition may also be used.
[0127] In some embodiments, both K222 and Q295 of an antibody can be conjugated to suitable compounds, for example by sequentially carrying out each conjugation reaction (e.g., as demonstrated in Example 5).EXAMPLES
[0128] The following examples are provided to promote a further understanding of the invention.
[0129] The meanings of the abbreviations in Examples are shown below.DMSO = dimethyl sulfoxideMOPS = (3-(N-morpholino) propanesulfonic acid)DAR = drug to antibody ratioHRMS = High-resolution mass spectrometry Ni-NTA = nickel-IMAC resin PNGase = Peptide:N-glycosidaseExample 1 : Glutamine conjugation with Biotin K Probe at 5mg scale
[0130] 3 lOpl of 0. IM MOPS buffer pH 7.5 was added to a low protein binding Eppendorf tube. 500pl of a lOmg / ml solution of Her2 IgG (5 mg, 0.034 pmol) was added followed by 150pl of a 4mg / ml solution of Biotin K ((S)-2-acetamido-6-amino-N-(13-oxo-17-((3aS,4S,6aR)-2- oxohexahydro-lH-thieno[3,4-d]imidazol-4-yl)-3,6,9-trioxa-12-azaheptadecyl)hexanamide) (0.6 mg, 1.019 pmol) in DMSO. The HER2 antibody had a heavy chain comprising the sequence of amino acids set forth in SEQ ID NO: 61 and a light chain comprising the sequence of amino acids set forth in SEQ ID NO: 62. 40pl of a 1 Img / ml aqueous solution of SEQ ID NO: 50 (mTGBiol_Rdl3BB )(0.4 mg, 0.0103 pmol) was added and the reaction was put in a shaker and incubated at 37°C and 800 RPM for 18hrs. The final pH was pH 7.0. At the end of the reaction, mTG Biol_Rdl3BB was removed by treatment with Ni-NTA resin in well plate format. The resulting solution was passed through a desalting plate to remove DMSO and excess Biotin K. The reaction for Example 1 is also shown in FIGURE 1. Drug Antibody Ratio (DAR) was calculated with Agilent DAR Calculator (vl.2) using peak area from deconvoluted mass spectrum to determine the number of Biotin probe conjugated to mAb. Analysis of the modified mAb via UPLC-QTOF showed an average DAR of 2.2 with 95% DAR 2 species on single conjugation site Q298.
[0131] Glutamine conjugation with Biotin K Probe was performed with a WT-Trop2 antibody having a heavy chain comprising the sequence of amino acids set forth in of SEQ ID NO: 65 and a light chain comprising the sequence of amino acids set forth in of SEQ ID NO: 66. Analysis of the modified mAb via UPLC-QTOF showed an average DAR of 1.9 with 90% DAR 2 species on single conjugation site Q298.Table 1 : Glutamine conjugation using mTG Biol variantsExample 2: Glutamine conjugation with Lys-Ala-Ala-PABC-MMAE Probe at 0.5mg scale
[0132] 31 pl of 0. IM MOPS buffer pH 7.5 was added to a low protein binding Eppendorf tube. 50pl of a lOmg / ml solution of HerlgGl (0.5 mg, 0.004 pmol) was added followed by a solution of Lys-Ala-Ala-PABC-MMAE (0.2 mg, 0.2 pmol) in 15pl DMSO. 4pl of a l lmg / ml aq. solution of SEQ ID NO: 50 (mTG Biol_Rdl3BB ) (0.02 mg, 0.001 pmol) in buffer was added and the reaction incubated at 37C and 800rpm for 18-36hrs. The final pH was pH 7.0. At the end of the reaction, mTG Biol_Rdl3BB was removed by treatment with Ni-NTA resin in well plate format. The resulting solution was passed through a desalting plate to remove DMSO and excess probe. The reaction for Example 2 is also show in FIGURE 2. Analysis of the modified mAb via UPLC- QTOF showed and average DAR of 0.6 with 21% DAR 2 species on conjugation site Q298.
[0133] Glutamine conjugation with Lys-Ala-Ala-PABC-MMAE Probe was also performed with WT-Trop an antibody having a heavy chain of SEQ ID NO: 65 and a light chain of SEQ IDNO: 66. Analysis of the modified mAb via UPLC-QTOF showed and average DAR of 0.3 with undetectable amount of DAR 2 species on conjugation site Q299 for WT Trop-2 mAb.Table 2: Glutamine conjugation using mTG Biol variants and a Lys-Ala-Ala-PABC-MMAE ProbeExample 3: Glutamine conjugation with Azide K Probe at Img scale
[0134] 22pl of 0. IM MOPS buffer pH 7.5 was added to a low protein binding Eppendorf tube. lOOpl of a lOmg / ml solution of HerlgGl (1 mg, 0.007 pmol) was added followed by Azide K (0.12 mg, 0.4 pmol) in 40pl buffer. Then, 30pl DMSO were added. 8 pl of a 1 Img / ml solution of SEQ ID NO: 50 (mTG Biol_Rdl3BB ) (0.08 mg, 0.02 pmol) was added and the reaction incubated at 37°C and 800rpm for 18hrs. The final pH was pH 7.0. At the end of the reaction, mTG Biol_Rdl3BB was removed by treatment with a Ni-NTA resin in well plate format. The reaction for Example 3 is shown in FIGURE 3 The resulting solution was passed through a desalting plate removing DMSO and excess Azide K. Analysis of the modified mAb via UPLC- QTOF showed an average DAR of 1.9 with 94% DAR 2 species on single conjugation site Q298.Table 3: Glutamine conjugation using mTG Biol variants and an azide probeExample 4: Lysine conjugation at 2mg scale
[0135] 124pl of 0. IM MOPS buffer pH 7.5 was added to a low protein binding Eppendorf tube. 200pl of a lOmg / ml solution of HerlgGl (2 mg, 0.014 pmol) was added followed by 60pl of a 4mg / ml solution of Biotin Q (0.24 mg, 0.41 pmol) in DMSO. 16pl of a l lmg / ml aq. solution of SEQ ID NO: 41 (mTG Bio2_Rd4BB) (0.16 mg, 0.004 pmol) was added and the reaction incubated at 37°C and 800rpm for 18hrs. The final pH was pH 7.0. At the end of the reaction, mTG Bio2_Rd4BB was removed by treatment with a Ni-NTA resin in well plate format. The resulting solution was passed through a desalting plate removing DMSO and excess Biotin Q. The reaction for Example 4 is shown in FIGURE 4. Analysis of the modified mAb via UPLC-QTOF showed and average DAR of 1.7 with 65% DAR 2 species on predominantly conjugation site K225.Table 4: Lysine conjugation using mTG Bio2 variants and a Biotin Q Probe
[0136] 124pl of 0. IM MOPS buffer pH 7.5 was added to a low protein binding Eppendorf tube. 200pl of a lOmg / ml solution of HerlgGl (2 mg, 0.014 pmol) was added followed by Biotin K (0.24 mg, 0.4 pmol) in 60pl DMSO. 16 pl of a 1 Img / ml aq. solution of SEQ ID NO: 50 (mTG Biol_Rdl3BB ) (0.16 mg, 0.004 pmol) was added and the reaction was put in a shaker and incubated at 37°C and 800RPM for 18hrs. The final pH was pH 7.0. At the end of the reaction, SEQ ID NO: 50 (mTG Biol_Rdl3BB ) was removed by treatment with Ni-NTA resin in well plate format. The resulting solution was passed through a desalting plate removing DMSO and excess Biotin K. Analysis of the modified mAb via UPLC-QTOF showed complete modification of the Her IgG mAb with 90% DAR 2 species on conjugation site Q298. The purified reaction mixture was concentrated to 200pL using centrifugation tubes with 30kD MW cut-off. Then, 124pl of 0. IM MOPS buffer pH 7.5 was added followed by Biotin Q (0.24 mg, 0.41 pmol) in 60pl DMSO. 16pl of a 1 Img / ml aq. solution of mTG Bio2_Rd4BB (0.16 mg, 0.004 pmol) was added and the reaction incubated at 37°C and 800rpm for 18hrs. The final pH was pH 7.0. At the end of the reaction, mTG Bio2_Rd4BB was removed by treatment with a Ni-NTA resin in well plate format. The resulting solution was passed through a desalting plate removing DMSO and excess Biotin Q. The reaction for Example 5 is shown in FIGURE 5. Analysis of the modified mAb via UPLC-QTOF showed 18% dual conjugation (DAR 2 + 2).Table 5 Step 1 : Glutamine conjugation with Biotin KTable 6 Step 2: Lysine conjugation of biotinylated mAb from step 1 with Biotin QExample 6: Analytical assay and sample preparationDesalting
[0137] Crude reaction mixtures were first de-salted and then treated with enzyme (PNGase) before injecting to HRMS (Agilent IMQTOF 6560). A 40 MWCO Zeba 96-well spin desalting plate was used according to manufacturer protocol and equilibrate using 0. IM MOPS buffer pH 7.5. After the equilibration step, 100 pL of reaction sample was added to each well and centrifuged for 2 min at 1000 x g. The eluent was then collected.Sample PreparationIntact mass deglycosylation
[0138] A 20 pl sample was denatured by mixing the samples with 4 l of 5 X non-reducing Rapid PNGase F buffer (P071 IS) and incubating the mixture at 75 °C for 5 min. The deglycosylation was performed by adding 1 pl Rapid PNGase F, followed by incubation at 50 °C for 10-15 min. The sample was diluted with water to 0.4-1 pg / pl for LC-MS analysis.Reduced Mass
[0139] A 20 pl (~20pg) protein sample was mixed with 4 pl of 5X Rapid PNGase F buffer (P0710S) and 1 pl Rapid PNGase F, followed by incubation at 50 °C for 10-15 min.Table 7: Reversed phase analytical assay on UPLC-HRMS (Agilent 6560 IMQTOF)Example 7: Evolution screening process for identifying improved mTG variants
[0140] E. coli cultures, each harboring a plasmid that encodes the transglutaminase enzymes described herein were diluted serially to dilutions of 10'4, 10'5, and 10'6using Luria-Bretani Broth (culture media for cells) as a diluent. 100 pL of the dilutions were each spread on a petri dish containing Luria-Bertani agar, supplemented with 30 pg / mL of Kanamycin and 1% (w / v) glucose. The plates were placed in a 37 °C incubator overnight.
[0141] 200 pL per well of Luria-Bertani Broth (culture media for cells) (500 mL Luria-Bertani + 30pg / mL Kanamycin +l%(w / v) glucose) was aliquoted into labeled 96-well shallow well plates. The shallow well plates were loaded into a plate stacker of a colony picker. Agar plates containing colonies sufficiently diluted so that the majority of colonies were isolated from one another (known as single colonies to those schooled in the art) were picked into unique wells of the shallow well plates. The colonies were allowed to grow overnight, in a shaker at 30°C and 200ROM, and 85% RH.
[0142] 390 pL of Terrific Broth (TB) growth media (commercially available from ThermoFisher Scientific (TB + 50 pg / mL Kanamycin) was aliquoted into labeled 96-well deep well subculture plates. 13 pL of the overnight growth culture was transferred from each well of the master shallow well plate into the corresponding labeled deep well subculture plates. The plates were sealed with breathable film, and the plates were shaken for 2-2.5 h at 250 rpm, at 30°C and 85%RH. After shaking, optical density (ODeoo, optical density at 600nm wavelength) of at least one plate was measured for growth. When the ODeoo of this plate was in the range of 0.4-0.8, the deep well plates were inducted with 40 pL per well of induction media (for final concentration of 1 mM Isopropyl P- d-1 -thiogalactopyranoside). The plates were resealed and incubated with shaking for 18-20 h at 400 rpm, at 18 °C, and 85% RH.Harvest, lysis, and dispase treatment
[0143] The plates were centrifuged at 4000 rpm for 15 minutes at 4 °C. The media was disposed, the plates were vacuum sealed, and put in the -80 °C freezer. Plates were removed and thawed at 25 °C for 30 min. Cells were lysed as follows: 200 pL 2 g / L lysozyme (7 mM sodium phosphate buffer, pH 7.0, 140 mM NaCl solution, and 2.7 mM KC1 solution) was added to each well. The plates were put in 25 °C shakers for 2 h at 800 rmp. The plates were centrifuged down at 4000 RPM for 15 minutes at 4 °C.67 pL of dispase solution (45.50 mL of activation buffer and 1.40 mL 10 mg / ml dispase solution) and 133 pL lysate were mixed together. The plates were sealed and put in the shaker at 37 °C at 350 RPM for 30 minutes.Conjugation reaction
[0144] The conjugation reaction is depicted in FIGURE 6. Dispase treated lysate (16 pL) was added to PBS buffer (64 pL) and then mixed. The dilute lysate (5 pL) was then added to a solution of Biotin A (1 mM), Herceptin (hlgGl) (100 mM pH 6.0 MES) with DMSO cosolvent (22.5% v / v). Reactions were then incubated at 37 °C, 250 RPM for 18h.
[0145] Following the conjugation reaction, 40 MWCO Zeba 96-well spin desalting plate was used to remove small-molecules. An ELISA assay was performed after diluting the eluents 10,000-fold via serial dilution. For the ELISA assay, each well of the pre-coated protein L ELISA plates were rinsed three times with 200 pL of Wash Buffer (PBS buffer with 0.1% w / v Tween 137 mM NaCl, 2.7 mM KC1, 10 mM Na2HPO4, 1.8 mM K2HPO4, Tween (R) 20 detergent 0.1% (w / v)). 100 pL of sample was added to each well and the plates were incubated for 60 minutes at room temperature, at 300 rpm. Each well was rinsed three times with 200 pL of wash buffer. 100 pL 1 :4000 diluted Strep-HRP in dilution buffer was added add the plates were incubated at room temperature for 45 minutes at300 rpm. Each well was rinsed five times with 200 pL of Wash Buffer (PB ST). 50 pL / well TMB (3,3',5,5'-Tetramethylbenzidine) substrate was added. 50 pL / well of 0.16 M H2SO4 was added and the plates were read at 450 nm. Enzyme variants were then ranked by the ELISA response in each well with highly active enzymes showing greater response in streptavidin ELISA assay.SEQUENCE LISTING
[0146] A sequence listing file with SEQ ID NOs 1 to 93 is separately filed on the same date as the filing of this application, which is part of the specification as stated in the section titled “reference to sequence listing submitted electronically.” The alternatively formatted listing below provides the sequences of SEQ ID NOs 31 to 37 and 41 to 53.<SEQ ID NO: 031; Bio2_R2BB; AA; synthetic construct MDNRLGEETKSYAETYRLTADDVAYINALNESATAASNASPSFRAPYSDDRVTPPAEPL DRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTW VNSGQYPTNHLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRARE VASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGG NHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARP GYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDF DRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH <SEQ ID NO: 032; Bio2_R3BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATAASNASPSFRAPYSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNHLAFASFDEDRFKNELKNGRPDSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 033; Bio2_R5BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATAASHASPSFRAPYSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNHLAFASFDEDRFKNELKNGRPDSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFKERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYVARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGRLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVRGWPGGGSGHHHHHH<SEQ ID NO: 034; Bio2_R6BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATAASHASPSFRAPYSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNHLAFASFDEDRFKNELKNGRPDSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFKERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYVARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPEGRLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVRGWPGGGSGHHHHHH<SEQ ID NO: 035; Bio2_R7BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATTASNASPSFRAPYSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNHLAFASFDEDRFKNELKNGRPDSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFKERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGRLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVRGWPGGGSGHHHHHH<SEQ ID NO: 036; Bio2_R8BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATTASNASPSFRAPYIDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNHLAFASFDEDRFKNELKNGRPDSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFKERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGRLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVRGWPGGGSGHHHHHH<SEQ ID NO: 037; Bio2_R9BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATTASNASPSFRAPYIDDRVTPPAEPLDRMPDPYEPVYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNHLAFASFDEDRFKNELKNGRPDSGETRAEFEGRVAKWSFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDADSPFYSALRNTPSFKERNGGNCVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGRLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVLGWPGGGSGHHHHHH<SEQ ID NO: 041; Biol_R4BB; AA; synthetic constructMDNGAGEETKSYAETYRLTADDVANINALNESATAASNAGPSFRAPDSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNRLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFKERNGGNHDPSRMKAVIYSKHFWSGQDDSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRNWSEGYSDFDRGIYVITFIPKSWNTAPDKVKQGWPGGGSGHHHHHH<SEQ ID NO: 042; Biol_R5BB; AA; synthetic constructMDNGAGEETKSYAETYRLTADDVAYINALNESATAASNASPSFRAPDSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNRLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDDSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRNWSEGYSDFDRGIYVITFIPKSWNTAPDKVKQGWPGGGSGHHHHHH<SEQ ID NO: 043; Biol_R6BB; AA; synthetic constructMDNGAGEETKSYAETYRLTADDVAYINALNESATAASNASPSFRAPDSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNRLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRNWSEGYSDFDRGIYVITFIPKSWNTAPDKVKQGWPGGGSGHHHHHH<SEQ ID NO: 044; Biol_R7BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATAASNASPSFRAPDSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNRLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 045; Biol_R8BB; AA; synthetic constructMDNGSGEETKSYAETYRLTADDVAYINALNESATAASNASPSFRAPDSDDRVTPPAEPLDRMPDPYTPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMDEEQREWLSYGCVGVTWVNSGQYPTNRLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDELKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDRSLSADKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWSEGYSDFDRGIYVITFIPKSWNTAPDKVIQGWPGGGSGHHHHHH<SEQ ID NO: 046; Biol_R9BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATAASNASPSFRAPYSDDRVTPPAEPLDRMPDPYRPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNHLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDRSLSRDKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 047; Biol RlOBB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATANSNASPSFRAPYSDDRVTPPAEPLDRMPDPYDPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNDLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDARSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDRSLSRDKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 048; Biol Rl IBB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATANSNASPSFRAPDIDDRVTPPAEPLDRMPDPYDPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNDLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDADSPFYSALRNTPSFEERNGGNHVPSRMKAVIYSKHFWSGQDRSLSLDKRKYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 049; Biol_R12BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATANSNASPSFRAPYIDDRVTPPAEPLDRMPDPYDPSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNDLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDADSPFYSALRNTPSFEEVNGGNHVPSRMKAVIYSKHFWSGQDRNLSLDKRMYGDPDAFRPDRETDQVDMSTYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 050; Biol_R13BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATANSNASPSFRAPYIDDRVTPPAEPLDRMPDPYDGSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNDLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDADSPFYSALRNTPSFEEVNGGNHVPSRMKAVIYSKHFWSGQDRNLSLDKRMYGDPDAFRPDRETDQVDMSHYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWREGYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 051; Biol_R14BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATANSNASPSFRAPYIDDRVTPPAEPLDRMPDPYDGSYGRAETVVNNYIRKWQQVYSHRDGRKQQMTEEQREWLSYGCVGVTWVNSGQYPTNDLAFASFDEDRFKNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDADSPFYSALRNTPSFEEVNGGNHVPSRMKAVIYSKHFWSGQDRNLSLDKRMYGDPDAFRPDRETDQVDMSHYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWRELYSDF DRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 052; Biol_R15BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATANSNASPSFRMPYVDDRVTPPAEPLDRMPDPYDGSYGRAETVVNNYIRKWQQVYSHRDGIKQQMTEEQREWLSYGCVGVTWVNSGQYPTNDLAFASFDEDRFLNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDADSNFYSALRNTPSFEEVNGGNHVPSRMKAVIYSKHFWSGQDRNLSLDKRMYGDPDAFRPDRETDQVDMSHYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESKFRRWRELYSDFDRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH<SEQ ID NO: 053; Biol_R16BB; AA; synthetic constructMDNRLGEETKSYAETYRLTADDVAYINALNESATANSNASPSFRMPYVDDRVTPPAEPLDRMPDPYDGSYGRAETVVNNYIRKWQQVYSHRDGIKQQMTEEQREWLSYGCVGVTWVNSGQYPTNDLAFASFDEDRFLNELKNGRPRSGETRAEFEGRVAKESFDEEKGFQRAREVASVMNRALENAHDESAYLDNLKKELANGNDALRNEDADSNFYSALRNTPSFEEVNGGNHVPSRMKAVIYSKHFWSGQDRNLSLDKRMYGDPDAFRPDRETDQVDMSHYRYRARPGYVNFDYGWFGAQTEADADKTVWTHGNHYHAPNGSLGAMHVYESNFRRWRELYSDF DRGIYVITFIPKSWNTAPDKVVQGWPGGGSGHHHHHH
Claims
CLAIMSWhat is claimed is:
1. A polypeptide comprising an amino acid sequence that has at least 90% identity with the amino acid sequence of SEQ ID NO: 8.
2. The polypeptide of claim 1, comprising the amino acid sequence of SEQ ID NO: 9.
3. The polypeptide of claim 1, comprising the amino acid sequence of any one of SEQ ID NOs: 1 to 7.
4. The polypeptide of claim 1, consisting of the amino acid sequence of any one of SEQ ID NOs: 31 to 37.
5. The polypeptide of claim 1, wherein the polypeptide has transglutaminase activity.
6. The polypeptide of claim 5, wherein the polypeptide is capable of catalyzing a reaction between a lysine residue of a biomolecule and a compound that comprises an amide group, to thereby make a biomolecule-compound conjugate.
7. The polypeptide of claim 6, wherein the biomolecule is an antibody having a lysine at position 222 based on the Eu numbering scheme.
8. A polypeptide comprising an amino acid sequence that has at least 90% identity with the amino acid sequence of SEQ ID NO: 28.
9. The polypeptide of claim 8, comprising the amino acid sequence of SEQ ID NO: 29.
10. The polypeptide of claim 8, comprising the amino acid sequence of any one of SEQ ID NOs: 11 to 23.
11. The polypeptide of claim 8, consisting of the amino acid sequence of any one of SEQ ID NOs: 41 to 53.
12. The polypeptide of any one of claim 8, wherein the polypeptide has transglutaminase activity.
13. The polypeptide of claim 12, wherein the polypeptide is capable of catalyzing a reaction between a glutamine residue of a biomolecule and a compound that comprises a primary amine group, to thereby make a biomolecule-compound conjugate.
14. The polypeptide of claim 13, wherein the biomolecule is an antibody having a glutamine at position 295 based on the Eu numbering scheme.
15. A composition comprising the polypeptide of any one of claims 1 to 14 and a buffering agent.
16. A kit comprising the composition of claim 15.
17. An isolated nucleic acid encoding the polypeptide of any one of claims 1 to 14.
18. The isolated nucleic acid of claim 17, wherein said nucleic acid is a DNA.
19. An expression vector comprising the nucleic acid of claim 17 or 18.
20. The expression vector of claim 19, wherein said nucleic acid is operably linked to one or more control sequences suitable for directing expression of the encoded polypeptide in a host cell, wherein the one or more control sequences comprise a promoter.
21. A host cell comprising the expression vector of claim 19 or 20.
22. A method for producing the polypeptide of any one of claims 1 to 14, comprising cultivating the host cell of claim 21 in a medium under conditions suitable for expression of the polypeptide by the host cell; and optionally isolating the polypeptide from the medium.
23. A method of making a biomolecule-compound conjugate, comprising reacting a biomolecule with a compound in the presence of the polypeptide of any one of claims 1 to 6 tothereby make the biomolecule-compound conjugate, wherein the compound comprises a primary amide group.
24. The method of claim 23, wherein the biomolecule comprises a lysine moiety.
25. The method of any one of claims 23 to 24, wherein the compound comprises a drug.
26. The method of any one of claims 23 to 25, wherein the compound further comprises a linker, and wherein the linker comprises the primary amide group.
27. The method of claim 26, wherein the linker comprises a spacer cleavable by an enzyme endogenous to a cell.
28. The method of claim 23, wherein the compound is selected from Biotin K, Lys-Ala-Ala- PABC-MMAE,29. The method of any one of claims 25 to 28, wherein the biomolecule comprises a polypeptide.
30. The method of claim 29, wherein the biomolecule is an antibody.
31. The method of claim 30, wherein the antibody is glycosylated at position N297, per the Eu numbering scheme.
32. The method of any one of claims 24 to 31, wherein the compound is conjugated to a lysine side chain of the biomolecule.
33. The method of claim 24 or 32, wherein the lysine is at position K222, per the Eu numbering scheme and wherein the compound is predominantly conjugated to position K222 of the antibody, per the Eu numbering scheme.
34. The method of claim 33, wherein more than 90% of the biomolecule-compound conjugates made by the method comprise a conjugation at position K222 of the antibody, per the Eu numbering scheme.
35. A method of making a biomolecule-compound conjugate, comprising reacting a biomolecule with a compound in the presence of the polypeptide of any one of claims 7 to 12 to thereby make the biomolecule-compound conjugate, wherein the compound comprises a primary amine group.
36. The method of claim 35, wherein the biomolecule comprises a glutamine moiety.
37. The method of any one of claims 35 to 36, wherein the compound comprises a drug.
38. The method of any one of claims 35 to 37, wherein the compound further comprises a linker, and wherein the linker comprises the primary amine group.
39. The method of claim 38, wherein the linker comprises a spacer cleavable by an enzyme endogenous to a cell.
40. The method of claim 35, wherein the compound is selected from the group consisting ofBiotin (Biotin K, Biotin Q, Biotin A), Lys-Ala-Ala-PABC-MMAE,41. The method of any one of claims 35 to 40, wherein the biomolecule comprises a polypeptide.
42. The method of claim 41, wherein the biomolecule is an antibody.
43. The method of claim 42, wherein the antibody is glycosylated at position N297, per the Eu numbering scheme.
44. The method of any one of claims 36 to 43, wherein the compound is conjugated to a glutamine side chain of the biomolecule.
45. The method of claim 36 or 44, wherein the glutamine is at position Q295 per the Eu numbering scheme and wherein the compound is predominantly conjugated to position Q295 of the antibody per the Eu numbering scheme.
46. The method of claim 45, wherein more than 90% of the biomolecule-compound conjugates made by the method comprise a conjugation at position Q295 of the antibody, per the Eu numbering scheme.
Citation Information
Patent Citations
Transglutaminase variants
US20200263150A1
C-terminal lysine conjugated immunoglobulins
US20220088212A1