Thermostable binding scaffolds
A protein scaffold with tailored framework and loop regions addresses the limitations of existing affinity chromatography by providing stable and selective protein purification under extreme conditions, enhancing process-scale applicability to diverse targets.
Patent Information
- Application Number
- JP2025530528
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-28
- Filing Date
- 2023-11-28
- Publication Date
- 2025-12-16
Smart Images

Figure 2025540724000001_ABST
Abstract
Description
[Technical Field]
[0001] Sequence Listing This application contains an electronically filed Sequence Listing in Extensible Markup Language (XML) format, which is incorporated by reference in its entirety. The XML copy, created on November 22, 2023, is named 51027-005WO2_Sequence_Listing_11_22_23.XML and is 526,678 bytes in size.
[0002] Statement Regarding Federally Sponsored Research This invention was made with government support under Grant No. 1 R43 GM143942-01 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]
[0003] Affinity chromatography (AC) using target-specific immobilized capture agents is an established method for protein purification. In this technique, capture agents, e.g., proteins, nucleic acids, or small molecules, are coupled to a solid support and can then be used to isolate proteins of interest from complex mixtures. The technique has been widely used on a laboratory scale for the single-step purification of a variety of target proteins, including enzymes, transcription factors, growth factors, and antibodies.
[0004] The use of protein-based capture agents for AC in industrial applications has not been widespread because currently available approaches are either incompatible with the temperatures, extreme pH, and solvents often required for process-scale purification or are useful for only a limited number of targets. The purification of kilogram quantities of antibodies using an AC resin based on Staphylococcal Protein A is an exception. The development of Protein A resins highlights the use of protein engineering to improve AC resin robustness and some remaining limitations. Early versions of the resins using wild-type Protein A captured antibodies with high selectivity and capacity from cell culture media feedstocks but gradually lost activity after multiple cycles of cleaning-in-place with sodium hydroxide. Mutagenesis of Protein A has yielded variants with increased resistance to sodium hydroxide treatment and higher binding capacity. Despite its widespread use, Protein A resins are limited to antibody purification.
[0005] For process-scale purification of non-antibody targets, the use of AC has been much less widespread than the use of Protein A for purifying antibodies. For example, non-protein ligand-based approaches, such as small-molecule substrate mimics, are effective but are limited to specific enzyme classes and difficult to use with general proteins of interest. Alternatively, specialized affinity resins, such as glutathione or nickel, require the addition of non-native tags, which introduces downstream complexity for proteins targeted for therapeutic use. Immuno-AC, using antibody- or nanobody-based capture agents, is the most generally applicable approach and has been widely used to purify a variety of proteins at the laboratory scale. However, immuno-AC has several limitations. Generally, conjugation of antibodies to resins often results in heterogeneous coupling due to a lack of precise control over the site of conjugation. Furthermore, chromatography must be performed under oxidizing conditions to preserve disulfide bonds, which are essential for maintaining antibody structure. Also, target elution typically requires low or high pH conditions that are incompatible with some target proteins. For process-scale applications, the main limitation of Immuno-AC resins is their sensitivity to sodium hydroxide solutions, which are favored for clean-in-place procedures. Because of these limitations, new capture agents are needed that can perform under the various extreme conditions necessary for robust target purification. Summary of the Invention [Means for solving the problem]
[0006] In one aspect, the invention features a protein scaffold comprising framework regions and loop regions. The protein scaffold has the structure: A-F1-L1-F2-L2-F3-L3-F4-L4-F5-L5-F6-L6-F7-L7-F8-L8-F9-B where F1 to F9 correspond to framework regions 1 to 9, respectively. L1 to L8 correspond to loop regions 1 to 8, respectively. A and B are each independently absent or contain at least one amino acid; F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 4; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5), or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 5; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 6; L3 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 7; L4 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 8; L5 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 9; L6 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 10; L7 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 11; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 includes the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 12.
[0007] As described herein, sequences having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof include, for example, sequences having one insertion, two insertions, one deletion, two deletions, one substitution mutation, two substitution mutations, one insertion and one deletion, one insertion and one substitution mutation, or one deletion and one substitution mutation.
[0008] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X1, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X2, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X3 relative to SEQ ID NO: 1, wherein: X is any amino acid except the amino acid at the equivalent position in SEQ ID NO: 1, X1 is any amino acid except R or S; X2 is any amino acid except P or K, X3 is any amino acid except R or K.
[0009] In some embodiments, F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 4; F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5), or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 5; F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 6; F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 7; F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS; F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 9; F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 10; F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 11; F9 comprises the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 12.
[0010] In some embodiments, F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4); F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5); F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6), F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7); F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8); F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9), F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10), F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11), F9 contains the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12).
[0011] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X relative to SEQ ID NO: 1, wherein X is any amino acid.
[0012] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of N807D, S809T, R812H, S813T, E814P, S815G, D818V, N822S, N825D, N832S, W836E, K857E, E858V, I859V, K860E, L861V, D862G, R865H, K870A, N871D, N880T, K881R, K883R, N890G, K897R, K901H, K908Q, E912D, S914D, and K922Q relative to SEQ ID NO:1.
[0013] In some embodiments, at least one mutation is K870X and / or N890X. In some embodiments, at least one mutation is K870A and / or N890G. In some embodiments, at least one mutation is K870A. In some embodiments, at least one mutation is N890G.
[0014] In some embodiments, the protein scaffold comprises at least 3 fewer lysines than SEQ ID NO: 1. For example, in some embodiments, the protein scaffold comprises at least 3, 4, 5, 6, 7, 8, 9, or 10 fewer lysines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises at least 6 fewer lysines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises 9 fewer lysines than SEQ ID NO: 1. In some embodiments, the protein scaffold does not comprise any lysines.
[0015] In some embodiments, the protein scaffold comprises at least 3 fewer asparagines than SEQ ID NO: 1. For example, in some embodiments, the protein scaffold comprises at least 3, 4, 5, 6, 7, or 8 fewer asparagines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises at least 5 fewer asparagines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises 7 fewer asparagines than SEQ ID NO: 1. In some embodiments, the protein scaffold does not comprise any asparagines.
[0016] In some embodiments, A and B are each independently absent or at least one amino acid. For example, each of A and B can be independently absent. In some embodiments, A and B are each independently at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 30, 400, 500, 600, 700, 800, 900, 1,000 or more amino acids. In some embodiments, A and B each independently represent 0 to 1,000 amino acids, e.g., 1 to 10 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids), 10 to 100 amino acids (e.g., 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids), or 100 to 1,000 amino acids (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 amino acids).
[0017] In some embodiments, A and B are each independently zero or between 1 and 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids).
[0018] In some embodiments, L1 to L8 are each independently 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids).
[0019] In some embodiments, L1 to L8 are each independently 1 to 10 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids). In some embodiments, L1 to L8 are each independently 3 to 10 amino acids. In some embodiments, L1 to L8 are each independently 3 to 8 amino acids.
[0020] In some embodiments, L1 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), In some embodiments, L1 is 0 to 5 amino acids (e.g., 1 to 5 amino acids, e.g., 0, 1, 2, 3, 4, or 5 amino acids).
[0021] In some embodiments, L2 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L2 is 1 to 16 amino acids (e.g., 4 to 16 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 amino acids).
[0022] In some embodiments, L3 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L3 is 6 amino acids.
[0023] In some embodiments, L4 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L4 is 0 to 5 amino acids (e.g., 1 to 5 amino acids, e.g., 0, 1, 2, 3, 4, or 5 amino acids).
[0024] In some embodiments, L5 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L5 is 5 amino acids.
[0025] In some embodiments, L6 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), hi some embodiments, L6 is 3 to 6 amino acids (e.g., 3, 4, 5, or 6 amino acids).
[0026] In some embodiments, L7 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), hi some embodiments, L7 is 4 or 5 amino acids.
[0027] In some embodiments, L8 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), hi some embodiments, L8 is 4 to 6 amino acids (e.g., 4, 5, or 6 amino acids).
[0028] In some embodiments, L1 is 4 amino acids. In some embodiments, L2 is 7 amino acids. In some embodiments, L8 is 5 amino acids. In some embodiments, L1 is 4 amino acids, L2 is 7 amino acids, and / or L8 is 5 amino acids. In some embodiments, L1 is 4 amino acids, L2 is 7 amino acids, and L8 is 5 amino acids.
[0029] In some embodiments, L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1-X4 are each independently any amino acid. In some embodiments, X2 is V.
[0030] In some embodiments, L2 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1-X7 are each independently any amino acid.
[0031] In some embodiments, L8 comprises the sequence X1X2X3X4X5 (SEQ ID NO: 15), where X1-X5 are each independently any amino acid.
[0032] In some embodiments, L4 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1-X7 are each independently any amino acid.
[0033] In some embodiments, L6 comprises the sequence X1X2X3X4X5X6 (SEQ ID NO: 16), where X1-X6 are each independently any amino acid.
[0034] In some embodiments, L8 comprises at least two amino acids. In some embodiments, L8 comprises at least one amino acid.
[0035] In some embodiments, L4 comprises the sequence of (G / D)-GGSS (SEQ ID NO: 17) or GDT, or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 17 or GDT.
[0036] In some embodiments, L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 18.
[0037] In some embodiments, L4 comprises the sequence (G / D)-GGSS (SEQ ID NO: 17) or GDT, and L6 comprises the sequence TGAPAG (SEQ ID NO: 18).
[0038] In some embodiments, L3 comprises the sequence (E / K / S)-(V / E)-(V / I / T)-(E / K / P / S)-(V / L)-(G / D) (SEQ ID NO: 19), or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 19.
[0039] In some embodiments, L5 comprises the sequence LD-(G / N)-(E / S)-S (SEQ ID NO: 20) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 20.
[0040] In some embodiments, L7 comprises at least one amino acid.
[0041] In some embodiments, L7 comprises the sequence of ETPI-(S / E)-A (SEQ ID NO: 21) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 21.
[0042] In some embodiments, L3 comprises the sequence (E / K / S)-(V / E)-(V / I / T)-(E / K / P / S)-(V / L)-(G / D) (SEQ ID NO: 19) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 19; L5 comprises the sequence LD-(G / N)-(E / S)-S (SEQ ID NO: 20) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 20; and L7 comprises the sequence ETPI-(S / E)-A (SEQ ID NO: 21) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 21.
[0043] In some embodiments, A comprises the sequence (D / N / H)-P. In some embodiments, A comprises the sequence DP.
[0044] In some embodiments, B comprises the sequence of DELE (SEQ ID NO: 35).
[0045] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 24; L3 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 25; L4 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F5 comprises the sequence of DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 26; L5 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F6 comprises the sequence of WTTIREYDH (SEQ ID NO: 27) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 27; L6 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 28; L7 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F8 comprises the sequence of QYIRLTNLE (SEQ ID NO: 29) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 30.
[0046] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22), L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 contains the sequence LTFSEFAIVS (SEQ ID NO: 30).
[0047] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 24; L3 comprises the sequence of EVVEVG (SEQ ID NO: 31) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 31; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 25; L4 comprises the sequence GGGSS (SEQ ID NO: 32) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 32; F5 comprises the sequence of DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 26; L5 comprises the sequence of LDGES (SEQ ID NO: 33) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 33; F6 comprises the sequence of WTTIREYDH (SEQ ID NO: 27) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 27; L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 18; F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 28; L7 comprises the sequence of ETPISA (SEQ ID NO: 34) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 34; F8 comprises the sequence of QYIRLTNLE (SEQ ID NO: 29) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 30.
[0048] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) relative to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) relative to SEQ ID NO: 24; L3 comprises the sequence of EVVEVG (SEQ ID NO: 31) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 31; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 25; L4 comprises the sequence GGGSS (SEQ ID NO: 32) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 32; F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 26; L5 comprises the sequence of LDGES (SEQ ID NO: 33) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 33; F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 27; L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 18; F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 28; L7 comprises the sequence of ETPISA (SEQ ID NO: 34) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) relative to SEQ ID NO: 34; F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30) or a sequence having a single amino acid insertion, deletion, or substitution mutation (eg, a single substitution mutation) compared to SEQ ID NO: 30.
[0049] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22), L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 contains the sequence LTFSEFAIVS (SEQ ID NO: 30).
[0050] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22), L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1 to X4 are each independently any amino acid; F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1 to X7 are each independently any amino acid; F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 comprises the sequence X1X2X3X4X5 (SEQ ID NO: 15), where X1 to X5 are each independently any amino acid; F9 contains the sequence LTFSEFAIVS (SEQ ID NO: 30).
[0051] In some embodiments, A comprises the sequence of DP, F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22), L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1 to X4 are each independently any amino acid; F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1 to X7 are each independently any amino acid; F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 comprises the sequence X1X2X3X4X5 (SEQ ID NO: 15), where X1 to X5 are each independently any amino acid; F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30), B contains the sequence of DELE (SEQ ID NO: 35).
[0052] In some embodiments, L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1, X3, and X4 are each independently any amino acid, and X2 is V.
[0053] Another aspect features a protein scaffold that includes a polypeptide having at least 80% (e.g., at least 85%, 90%, 95%, 97%, or 99%) sequence identity to SEQ ID NO:3. In some embodiments, the polypeptide includes the sequence of SEQ ID NO:3. In some embodiments, the polypeptide does not include the sequence of SEQ ID NO:1. In some embodiments, the polypeptide does not include the sequence of SEQ ID NO:2.
[0054] In some embodiments, the polypeptide comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X1, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X2, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X3 relative to SEQ ID NO: 1, wherein: X is any amino acid except the amino acid at the equivalent position in SEQ ID NO: 1, X1 is any amino acid except R or S; X2 is any amino acid except P or K, X3 is any amino acid except R or K.
[0055] In some embodiments of any of the above aspects, the protein scaffold further comprises a mutation that adds a cysteine residue. In some embodiments, the protein scaffold comprises a first mutation that adds a first cysteine residue and a second mutation that adds a second cysteine residue. In some embodiments, the first cysteine residue and the second cysteine residue form a disulfide bond under oxidizing conditions.
[0056] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of F806C, P808C, S845C, L855C, V858C, V861C, K878C, W879C, L884C, L888C, A904C, P905C, A906GC, G907C, I924C, L926C, N928C, L936C, I943C, L948C.
[0057] In some embodiments, the protein scaffold comprises at least two or more mutations selected from the group consisting of F806C, P808C, S845C, L855C, V858C, V861C, K878C, W879C, L884C, L888C, A904C, P905C, A906GC, G907C, I924C, L926C, N928C, L936C, I943C, L948C.
[0058] In some embodiments, the protein scaffold comprises a cysteine mutation pair selected from the group consisting of K878C and G907C, K878C and A904C, V861C and I943C, P905C and L855C, S845C and L936C, W879C and N928C, L884C and L926C, F806C and L948C, V858C and L888C, K878C and G907C, K878C and A906GC, S845C and N928C, K878C and A904C, P808C and I943C, V861C and I924C, P808C and V861C, and I943C and L855C.
[0059] In some embodiments, the cysteine mutation pair is selected from the group consisting of K878C and G907C, K878C and A904C, S845C and L936C, W879C and N928C, W879C and N928C, L884C and L926C, V858C and L888C, K878C and G907C, and K878C and A906GC (i.e., replacement of alanine 906 with glycine and cysteine).
[0060] In some embodiments of any of the above aspects, the protein scaffold further comprises a tag covalently attached to the scaffold.
[0061] In some embodiments, the tag is an affinity tag (eg, a polyhistidine tag, eg, 4, 5, 6, 7, 8, 9, or 10 histidines), an epitope tag, a covalent tag, or a protein tag.
[0062] In some embodiments, the tag is attached to the N-terminus or C-terminus of the scaffold.
[0063] In some embodiments, the scaffold is conjugated to a functional group, which in some embodiments comprises biotin, streptavidin or a derivative of streptavidin, a polyethylene glycol moiety, a fluorescent dye, an enzyme, a radioactive moiety, a lanthanide, or a lanthanide binding motif.
[0064] In some embodiments, the scaffold is conjugated to a lanthanide or a lanthanide binding motif, hi some embodiments, the lanthanide is terbium.
[0065] In some embodiments, the scaffold is conjugated to a radioactive moiety, hi some embodiments, the radioactive moiety is an alpha or beta emitter.
[0066] In some embodiments, the functional group is conjugated to a sulfhydryl group or a primary amine.
[0067] Another aspect features a polynucleotide encoding a protein scaffold described herein, e.g., any of the above embodiments. In some embodiments, the polynucleotide is a ribonucleotide. In some embodiments, the polynucleotide is a deoxyribonucleotide.
[0068] Another aspect features a vector that includes a polynucleotide described herein.
[0069] Another aspect features a cell containing a polynucleotide encoding a protein scaffold or a vector containing the polynucleotide.
[0070] Another aspect features a method of producing a protein scaffold described herein, e.g., in any of the above embodiments. The method includes: (a) providing a cell transformed with a polynucleotide encoding the protein scaffold or a vector comprising the polynucleotide; (b) culturing the transformed cell under conditions for expression of the polynucleotide, wherein culturing results in expression of the protein scaffold. The method may further include (c) isolating the protein scaffold or using the protein scaffold to bind to a target.
[0071] Another aspect features a particle including the protein scaffold of any of the above embodiments. In some embodiments, the particle is a magnetic particle.
[0072] Another aspect features a resin that includes a plurality of particles, eg, containing a protein scaffold.
[0073] Another aspect features a column (eg, a chromatography column) containing, for example, particles or resins conjugated to a scaffold.
[0074] Another aspect features a method for purifying a target molecule from a plurality of molecules, the method including: (a) providing a sample comprising a mixture of the target molecule and the plurality of molecules, (b) contacting the sample with a protein scaffold of any one of the above embodiments, wherein the scaffold specifically binds to the target molecule, and (c) separating the target molecule bound to the protein scaffold from the plurality of molecules.
[0075] In some embodiments, the separating step comprises immobilizing the protein scaffold.
[0076] In some embodiments, the protein scaffold is conjugated to a particle. In some embodiments, the particle comprises a magnetic bead. In some embodiments, the protein scaffold is conjugated to a resin or monolith comprising a plurality of particles.
[0077] definition The carbohydrate-binding module family 32 (CBM32) scaffold of SEQ ID NO: 1 is derived from a single protein domain of Clostridium perfringens hyaluronidase (NagH), a multi-domain enzyme consisting of 1627 amino acids. Amino acid residue 1 of SEQ ID NO: 1 corresponds to amino acid residue 807 of NagH, and amino acid residue 140 of SEQ ID NO: 1 corresponds to amino acid residue 946 of NagH. The amino acid positions and mutations described herein generally relate to the corresponding positions on full-length NagH, unless otherwise specified.
[0078] The term "constant region," as used herein, generally refers to a region of a binding scaffold that does not include the variable loop regions involved in target binding. For example, a constant region may include a framework region (e.g., F1-F9) or a loop region (e.g., L3-L7) that is not one of the three loops (e.g., L1, L2, and L8) that has been mutagenized for target binding. The constant region may have sequence variability.
[0079] The term "non-naturally occurring amino acid" as used herein means a non-proteinogenic amino acid. Examples of non-naturally occurring amino acids include D-amino acids; amino acids having an acetylaminomethyl group attached to the sulfur atom of cysteine; pegylated amino acids; amino acids of the formula NH2(CH2) nOmega amino acids with COOH (where n is 2 to 6), neutral nonpolar amino acids such as sarcosine, t-butylalanine, t-butylglycine, N-methylisoleucine, and norleucine; oxymethionine; phenylglycine; citrulline; methionine sulfoxide; cysteic acid; ornithine; diaminobutyric acid; 3-aminoalanine; 3-hydroxy-D-proline; 2,4-diaminobutyric acid; 2-aminopentanoic acid; 2-aminooctanoic acid, 2-carboxypiperazine; piperazine-2-carboxylic acid, 2-amino-4-phenylbutanoic acid; 3-(2-naphthyl)alanine, and hydroxyproline. Other amino acids include α-aminobutyric acid, α-amino-α-methylbutyrate, aminocyclopropane-carboxylate, aminoisobutyric acid, aminonorbornyl-carboxylate, L-cyclohexylalanine, cyclopentylalanine, LN-methylleucine, LN-methylmethionine, LN-methylnorvaline, LN-methylphenylalanine, LN-methylproline, LN-methylserine, LN-methyltryptophan, D-ornithine, LN-methylethylglycine, L-norleucine, α-methyl-aminoisobutyrate, α-methylcyclohexylalanine, D-α-methylalanine, D-α-methylarginine, D-α-methylasparagine, D-α-methylaspartate, D-α-methylcysteine, D-α-methylglutamine, D-α-methylhistidine, D -α-methylisoleucine, D-α-methylleucine, D-α-methyllysine, D-α-methylmethionine, D-α-methylornithine, D-α-methylphenylalanine, D-α-methylproline, D-α-methylserine, DN-methylserine, D-α-methylthreonine, D-α-methyltryptophan, D-α-methyltyrosine, D-α-methylvaline, DN-methylalanine, DN-methylarginine, DN-methylasparagine, DN-methylaspartate, DN-methylcysteine, DN-methylglutamine, DN-methylglutamate, DN-methylhistidine, DN-methylisoleucine, DN-methylleucine, DN-methyllysine, N-methylcyclohexylalanine, DN-methylornithine, N-methylglycine, N-methylaminoisobutyrate,N-(1-methylpropyl)glycine, N-(2-methylpropyl)glycine, DN-methyltryptophan, DN-methyltyrosine, DN-methylvaline, γ-aminobutyric acid, Lt-butylglycine, L-ethylglycine, L-homophenylalanine, L-α-methylarginine, L-α-methylaspartate, L-α-methylcysteine, L-α-methylglutamine, L-α-methylhistidine, L-α-methylisoleucine, L-α-methylleucine, L-α-methylmethionine, L-α-methylnorvaline, L-α-methylphenylalanine L-α-methylserine, L-α-methyltryptophan, L-α-methylvaline, N-(N-(2,2-diphenylethyl)carbamylmethylglycine, 1-carboxy-1-(2,2-diphenyl-ethylamino)cyclopropane, 4-hydroxyproline, ornithine, 2-aminobenzoyl (anthraniloyl), D-cyclohexylalanine, 4-phenyl-phenylalanine, L-citrulline, α-cyclohexylglycine, L-1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, L-thiazolidine-4-carboxylic acid, L-homotyrosine, L-2-furylalanine, L-histidine (3-methyl), N-(3-guanidinopropyl)glycine, O-methyl-L-tyrosine, O-glycan-serine, meta-tyrosine, nor-tyrosine, LN,N',N"-trimethyllysine, homolysine, norlysine, N-glycan asparagine, 7-hydroxy-1,2,3,4-tetrahydro-4-fluorophenylalanine, 4-methylphenylalanine, bis-(2-picolyl)amine, pentafluorophenylalanine, indoline-2-carboxylic acid, 2-aminobenzoic acid, 3-amino-2-naphthoic acid, asymmetric dimethylarginine, L-tetrahydroisoquinoline-1-carboxylic acid, D-tetrahydroisoquinoline-1-carboxylic acid, 1-amino-cyclohexaneacetic acid, D / L-allylglycine, 4-aminobenzoic acid, 1-amino-cyclobutanecarboxylic acid, 2 or 3 or 4-aminocyclohexanecarboxylic acid, 1-amino-1-cyclopentanecarboxylic acid, 1-aminoindan-1-carboxylic acid, 4-amino-pyrrolidine-2-carboxylic acid, 2-aminotetralin-2-carboxylic acid, azetidine-3-carboxylic acid,4-Benzyl-pyrrolidine-2-carboxylic acid, tert-butylglycine, b-(benzothiazolyl-2-yl)-alanine, b-cyclopropylalanine, 5,5-dimethyl-1,3-thiazolidine-4-carboxylic acid, (2R,4S)4-hydroxypiperidine-2-carboxylic acid, (2S,4S) and (2S,4R)-4-(2-naphthylmethoxy)-pyrrolidine-2-carboxylic acid, (2S,4S) and (2S,4R)-4-phenoxy-pyrrolidine-2-carboxylic acid, (2R,5S) and (2S,5R)-5-phenyl-pyrrolidine-2- Carboxylic acid, (2S,4S)-4-amino-1-benzoyl-pyrrolidine-2-carboxylic acid, t-butylalanine, (2S,5R)-5-phenyl-pyrrolidine-2-carboxylic acid, 1-aminomethyl-cyclohexane-acetic acid, 3,5-bis-(2-amino)ethoxy-benzoic acid, 3,5-diamino-benzoic acid, 2-methylamino-benzoic acid, N-methylanthranilic acid, LN-methylalanine, LN-methylarginine, LN-methylasparagine, LN-methylaspartic acid, LN-methylcysteine, LN-methylglutamine, LN-methyl L-glutamic acid, N-methylhistidine, N-methylisoleucine, N-methyllysine, N-methylnorleucine, N-methylornithine, N-methylthreonine, N-methyltyrosine, N-methylvaline, N-methyl-t-butylglycine, L-norvaline, α-methyl-γ-aminobutyrate, 4,4'-biphenylalanine, α-methylcyclopentylalanine, α-methyl-α-naphthylalanine, α-methylpenicillamine, N-(4-aminobutyl)glycine, N-(2-aminoethyl)glycine, N-(3-amino Propyl)glycine, N-amino-α-methylbutyrate, α-naphthylalanine, N-benzylglycine, N-(2-carbamylethyl)glycine, N-(carbamylmethyl)glycine, N-(2-carboxyethyl)glycine, N-(carboxymethyl)glycine, N-cyclobutylglycine, N-cyclodecylglycine, N-cycloheptylglycine, N-cyclohexylglycine, N-cyclodecylglycine, N-cyclododecylglycine, N-cyclooctylglycine, N-cyclopropylglycine, N-cycloundecylglycine,N-(2,2-diphenylethyl)glycine, N-(3,3-diphenylpropyl)glycine, N-(3-guanidinopropyl)glycine, N-(1-hydroxyethyl)glycine, N-(hydroxyethyl)glycine, N-(imidazolylethyl)glycine, N-(3-indolylethyl)glycine, N-methyl-γ-aminobutyrate, DN-methylmethionine, N-methylcyclopentylalanine, DN-methylphenylalanine, DN-methylproline, DN-methylthreonine, N-(1-methylethyl)glycine, N-methyl-naphthyl Alanine, N-methylpenicillamine, N-(p-hydroxyphenyl)glycine, N-(thiomethyl)glycine, penicillamine, L-α-methylalanine, L-α-methylasparagine, L-α-methyl-t-butylglycine, L-methylethylglycine, L-α-methylglutamate, L-α-methylhomophenylalanine, N-(2-methylthioethyl)glycine, L-α-methyllysine, L-α-methylnorleucine, L-α-methylornithine, L-α-methylproline, L-α-methylthreonine, L-α-methyltyrosine, L-N-methyl- Homophenylalanine, N-(N-(3,3-diphenylpropyl)carbamylmethylglycine, L-pyroglutamic acid, D-pyroglutamic acid, O-methyl-L-serine, O-methyl-L-homoserine, 5-hydroxylysine, α-carboxyglutamate, phenylglycine, L-pipecolic acid (homoproline), L-homoleucine, L-lysine (dimethyl), L-2-naphthylalanine, L-dimethyldopa or L-dimethoxy-phenylalanine, L-3-pyridylalanine, L-histidine (benzoyloxymethyl), N-cycloheptyl Glycine, L-diphenylalanine, O-methyl-L-homotyrosine, L-β-homolysine, O-glycan-threonine, ortho-tyrosine, LN,N'-dimethyllysine, L-homoarginine, neotryptophan, 3-benzothienylalanine, isoquinoline-3-carboxylic acid, diaminopropionic acid, homocysteine, 3,4-dimethoxyphenylalanine, 4-chlorophenylalanine, L-1,2,3,4-tetrahydronorharman-3-carboxylic acid, adamantylalanine, symmetric dimethylarginine, 3-carboxythiomorpholine,D-1,2,3,4-tetrahydronorharman-3-carboxylic acid, 3-aminobenzoic acid, 3-amino-1-carboxymethyl-pyridin-2-one, 1-amino-1-cyclohexanecarboxylic acid, 2-aminocyclopentanecarboxylic acid, 1-amino-1-cyclopropanecarboxylic acid, 2-aminoindan-2-carboxylic acid, 4-amino-tetrahydrothiopyran-4-carboxylic acid, azetidine-2-carboxylic acid, b-(benzothiazol-2-yl)-alanine, neopentylglycine, 2-carboxymethylpiperidine, b-cyclobutylalanine, allylglycine, diaminopropionic acid, homo-cyclohexylalanine, (2S,4R)-4 -hydroxypiperidine-2-carboxylic acid, octahydroindole-2-carboxylic acid, (2S,4R) and (2S,4R)-4-(2-naphthyl), pyrrolidine-2-carboxylic acid, nipecotic acid, (2S,4R) and (2S,4S)-4-(4-phenylbenzyl)pyrrolidine-2-carboxylic acid, (3S)-1-pyrrolidine-3-carboxylic acid, (2S,4S)-4-tritylmercapto-pyrrolidine-2-carboxylic acid, (2S,4S)-4-mercaptoproline, t-butylglycine, N,N-bis(3-aminopropyl)glycine, 1-amino-cyclohexane-1-carboxylic acid, N-mercaptoethylglycine, and selenocysteine. In some embodiments, the amino acid residue can be charged or polar. Charged amino acids include alanine, lysine, aspartic acid, or glutamic acid, or non-naturally occurring analogs thereof. Polar amino acids include glutamine, asparagine, histidine, serine, threonine, tyrosine, methionine, or tryptophan, or non-naturally occurring analogs thereof. In some embodiments, it is specifically contemplated that the terminal amino group in an amino acid may be an amide group or a carbamate group.
[0080] As used herein, the term "percent (%) identity" refers to the percentage of amino acid residues in a candidate sequence, e.g., a protein scaffold, that are identical to the amino acid residues of a reference sequence, e.g., a wild-type CBM32 polypeptide, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent identity (i.e., gaps can be introduced into one or both of the candidate and reference sequences for optimal alignment, and non-homologous sequences can be disregarded for comparison purposes). Alignment for purposes of determining percent identity can be achieved in a variety of ways within the skill of one in the art, for example, using publicly available computer software, such as BLAST, ALIGN, or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. In some embodiments, the percent amino acid sequence identity of a given candidate sequence with a given reference sequence, or relative to a given reference sequence (which can be translated as a given candidate sequence having or containing a certain percent amino acid sequence identity with or against a given reference sequence) is determined using the following: 100×(A / B ratio) where A is the number of amino acid residues scored as identical in the alignment of the candidate sequence and the reference sequence, and where B is the total number of amino acid residues in the reference sequence. In some embodiments where the length of the candidate sequence is not equal to the length of the reference sequence, the percent amino acid sequence identity of the candidate sequence relative to the reference sequence is not equal to the percent amino acid sequence identity of the reference sequence relative to the candidate sequence. [Brief explanation of the drawings]
[0081] [Figure 1]Figure 1 shows a schematic overview of the protein engineering campaign to generate a member of the nC-B class of nanoCLAMPs. The starting nanoCLAMP was the anti-SUMO clone SMT3-A1, a member of the nC-A class of nanoCLAMPs. SMT3A1 was mutated over seven rounds. At the end of each round, the performance of the clones combined with that round and different mutations was assessed by DSF and SEC. The final product was clone P2788, whose constant region served as the basis for the nC-B class of nanoCLAMPs. [Figure 2] Space-filling model of P2788, an example of the nC-B class of nanoCLAMPs, with mutations in P2788 mapped to the CBM-32-3 crystal structure, and the sequence of the constant region of the nC-A class of nanoCLAMPs. The constant region of clone P2788 is the basis for the nC-B class of nanoCLAMPs. Side chains at mutated positions are displayed in green, and side chains in the variable loops are displayed in red. Other side chains are displayed in light gray. Framework residues are displayed in dark gray. The alignment compares the constant regions of the nC-A and nC-B classes of nanoCLAMPs. Residues in bold black text represent positions in nC-A where at least one mutation was tested (top row). Mutations were tested at 58% of the positions in the constant region (72 of 124). Residues in bold text represent mutations in P2788 (bottom row). In P2788, 24% of the positions in the constant region are mutated to nC-A (30 of 124). [Figure 3] This model shows the crystal structure superposition of AlphaFold models of CBM32-2 (basis for the nC-A class of nanoCLAMPs) and P2788 (basis for the nC-B class). CBM32-2 (PDB accession 2W1Q) and P2788 were aligned with jFATCAT (rigid) on the RCSB server, resulting in a high TM score (0.95). Backbone deviations are evident and expected for loops with different amino acid sequences. [Figure 4]Graph showing differential scanning fluorimetry analysis of SMT3-A1 (nC-A class) and P2788 (nC-B class). Both clones show classic melting curve shapes with low initial fluorescence. The 30 mutations in P2788 increase its Tm by 24°C compared to SMT3-A1. [Figure 5] Figure 5A shows a gel and graphs demonstrating the protease resistance of nanoCLAMPs of the nC-A and nC-B classes. Figure 5A shows SDS-PAGE analysis of SMT3-A1 (nC-A class) and P2788 (which has the same variable loop as SMT3-A1 but the constant region of the nC-B class) after 16 hours of incubation with trypsin or chymotrypsin. Figures 5B-5D show SDS-PAGE analysis of time-course tryptic digests of SMT3-A1, P2788, and P2808. P2788 and P2808 have the same constant region (nC-B class) but different loops. P2788 and P2808 were resistant to trypsin digestion for over 20 hours. Figure 5E shows quantitative densitometry analysis of time-course stained gels. FIG. 5F shows SDS-PAGE analysis of members of the nC-A class (SMT3-A1) and nC-B class (P2788, P2808, P2809, and P2811) of nanoCLAMPs after 16 hours of trypsin digestion. [Figure 6] 1 is a set of size-exclusion chromatograms showing monodispersity and melting point analysis of nC-B class anti-SUMO nanoCLAMPs. Size-exclusion chromatography (left panel) and differential scanning fluorescence (right panel) of nanoCLAMPs P2808, P2809, and P2811. [Figure 7]This graph shows the dynamic binding capacity of SMT3-A1 resin (nC-A class) and P2808 resin (nC-B class). Breakthrough curves were generated by loading 0.2 mg / ml Sumo-GFP solution in PBS onto 0.6 ml of packed resin in a column (3 cm height × 5 mm ID) at a flow rate of 0.5 ml / min and measuring the fluorescence of the eluate. The percent fluorescence of the load was calculated by dividing the fluorescence of the eluate by the fluorescence of the load. The dynamic binding capacity (DBC) was calculated using the following formula: DBC = (Vx - Vdelay) × c / (Vresin). Vx is the volume of collected eluate, Vdelay is the elution volume of the load under nonbinding conditions, c is the concentration of target in the load, and Vresin is the volume of packed resin in the column. P2808 resin has a dynamic binding capacity of 10 mg / ml resin (240 nmol / ml resin). [Figure 8] These gels show the performance of P2808 resin in a single-step affinity chromatography purification of GFP-SUMO from spiked lysate in resin-limited and protein-limited scenarios. The same column was loaded at 33% (Figure 8A) or 58% (Figure 8B) above the dynamic binding capacity of the column. Escherichia coli (E. coli) lysate with spiked-in target protein (SUMO-GFP) was used as the load. After loading and washing, bound protein was eluted with 3 M imidazole, pH 8. Total protein was loaded on SDS-PAGE: Figure 8A: Lysate = 32 μg, spiked lysate = 34 μg, FT = 47 μg, eluate = 6 μg; Figure 8B: Lysate = 17 μg, spiked lysate = 17 μg, FT = 21 μg, wash = NA, eluate = 3 μg. The purification metrics for Figures 8A and 8B are tabulated in Table 6. [Figure 9]13 shows the effect of sodium hydroxide treatment on the binding capacity of nanoCLAMP capture agents of the nC-A and nC-B classes. The binding capacity of resins using capture agents of the nC-A class (SMT3-A1, P1519, P1533) and nC-B class (P2808, P2809, P2811) was determined after each of 22 cycles of purification of GFP-SUMO from spiked E. coli lysate, followed by washing, elution, and cleaning in place with 0.1 M NaOH (10 min contact time). The % of starting binding capacity was determined by dividing the fluorescence of the eluate by the fluorescence of the load. Selectivity was determined by analyzing the eluate on SDS-PAGE (FIG. 13). [Figure 10] These graphs and gels show the effect of organic solvents and autoclaving on the binding capacity of resins prepared using the nC-B class of nanoCLAMPs. Resins P2808, P2809, and P2811 (nC-B class) and SMT3-A1 (nC-A class) were incubated in 100% DMF for 2 hours (Figures 10A and 10B) or autoclaved (a 105-minute liquid-vapor cycle including a 30-minute exposure to 120 °C and 20 psi) (Figures 10C and 10D), then re-equilibrated in fresh buffer and tested in affinity chromatography purification of SUMO-GFP from spiked E. coli lysate. The % intact binding capacity was determined by dividing the fluorescence of the eluate by the fluorescence of the control (untreated) eluate. Specificity was determined by Coomassie staining of SDS-PAGE (Figures 10B and 10D). [Figure 11] Figure 1 shows the kinetic thermal stability of nC-A (SMT3-A1) and nC-B (P2808, P2809, and P2811) nanoCLAMPs. The nanoCLAMPs were heat-treated, cooled, and centrifuged. The supernatants were tested for binding activity by biolayer interferometry. The percent of initial response was measured as the amplitude of binding divided by the amplitude obtained with the control sample (held at 20°C during heat treatment). [Figure 12]1 is a gel showing the static binding capacity of P2808 resin. Affinity resin prepared with P2808 (nC-B class) was incubated with spiked E. coli lysate, washed, eluted with 3 M imidazole pH 8, buffer exchanged, and quantified by A280. [Figure 13] A set of gels showing the effect of sodium hydroxide treatment on the specificity of nC-A and nC-B class nanoCLAMP capture agents. nC-A (SMT3-A1, P1519, P1533) and nC-B (P2808, P2809, P2811) Sumo-conjugated nanoCLAMPs were covalently coupled to 6% cross-linked agarose resin and subsequently used to purify Sumo-GFP fusions from crude E. coli lysates. Each cycle consisted of loading a crude Sumo-GFP spiked sample, washing, elution with 3 M imidazole (collection), washing, 0.1 M NaOH regeneration (10 min contact time per cycle), and a 5 min refolding wash. Target protein in the eluate was quantified by fluorescence spectroscopy (Figure 9), and the percent yield was calculated by dividing the fluorescence by that of the initial eluate. The purity of the eluted target from each cycle was assessed by SDS-PAGE and stained with Coomassie. Cycle numbers are indicated for each lane. L = load, M = marker. The prominent band in the eluate of each gel is SUMO-GFP (42 kD). [Figure 14] Figure 14A shows a graph and gel examining the stability of resin P2808 (nC-B class) over 20 low-pH elution cycles. SUMO-GFP was spiked into crude E. coli lysate, loaded onto the P2808 resin, washed, and eluted with 0.1 M citrate, pH 2.5, followed by regeneration with 0.1 N NaOH for 1 minute of contact time per cycle and a 5-minute re-equilibration wash. Figure 14A shows the target protein in the eluate, which was quantified by densitometry on a Coomassie-stained SDS-PAGE gel in Figure 14B, since the fluorescence of the eluate was destroyed by the low pH. The percent yield was calculated by dividing the band density by the band density of the initial eluate. The purity of the eluted target from each cycle was assessed by SDS-PAGE in Figure 14B. [Figure 15]This graph shows nanoCLAMP stably binding terbium (Tb). SMT3-A1 (nC-A class), P2808 (nC-B class), and a negative control protein (recombinant SMT3) were incubated overnight with CaCl or TbCl, followed by buffer exchange to remove unbound metal. Buffer-exchanged proteins were analyzed 24 hours after buffer exchange by time-resolved fluorescence (Ex / Em: 350 nm / 544 nm) with a 200 μs delay. [Figure 16] Model showing the front, back, top, and bottom surfaces of the nC-B class of nanoCLAMPs. A, B, F1-9, and L1-L8 are mapped to the clone P2808 sequence and 3D modeled to show the location of each region of the scaffold. [Figure 17] A model showing the alignment of P2808, an example of the nC-B class, in which loop 1 has been replaced with a (G4S)3 sequence or removed. [Figure 18] 10 is a model showing the alignment of nC-B nanoCLAMP P2808, in which loop 2 has been replaced with a (G4S)3 sequence or removed. [Figure 19] 10 is a model showing the alignment of nC-B nanoCLAMP P2808, in which loop 4 has been replaced with a (G4S)3 sequence or removed. [Figure 20] 1 is a model showing the alignment of nC-B nanoCLAMP P2808, where loop 6 is replaced with a (G4S)3 sequence or GG. [Figure 21] 1 is a model showing nC-B nanoCLAMP P2808, in which loop 8 is replaced with GGGGG (SEQ ID NO: 36), GGGG (SEQ ID NO: 37), GGG, GG, or G, or is removed. [Figure 22] 10 is a model showing the alignment of nC-B nanoCLAMP P2808, in which loop 8 has been replaced with a (G4S)3 sequence or removed. [Figure 23]10 is a model showing the alignment of nC-B nanoCLAMP P2808, in which loop 3 has been replaced with a (G4S)3 sequence or removed. [Figure 24] 10 is a model showing the alignment of nC-B nanoCLAMP P2808, in which loop 5 has been replaced with a (G4S)3 sequence or removed. [Figure 25] 1 is a model showing nC-B nanoCLAMP P2808, in which loop 7 is replaced with GGGGG (SEQ ID NO: 36), GGGG (SEQ ID NO: 37), GGG, GG, or G, or is removed. [Figure 26] 1 is a model showing the alignment of nC-B nanoCLAMP P2808, where loop 8 is replaced with a (G4S)3 sequence or G. [Figure 27] Figure 1 shows the introduction of an artificial disulfide bond into clones P2808 and P2960. SDS-PAGE analysis of P2808 and P2960 variants mutated to contain adjacent cysteine pairs under oxidizing and reducing conditions. Purified proteins were treated with SDS sample buffer containing (first lane of each set) or lacking (second lane of each set) a reducing agent (DTT). In samples lacking DTT, the presence of a faster-migrating species indicates a disulfide bond, likely due to a more compact folding and smaller hydrodynamic radius. P2808 and P2960 do not contain Cys residues and, as expected, migrate at the same rate in oxidizing and reducing sample buffers. BSA, which contains 17 disulfide bonds, was included as a control for the activity of the reducing agent (DTT). [Figure 28] 1 is a graph showing that the artificial disulfide improves the thermal stability of P3015 by 9° C. The graph shows a differential scanning fluorescence (DSF) analysis of the melting points of reduced and oxidized P3015. DETAILED DESCRIPTION OF THE INVENTION
[0082] Immunoaffinity chromatography is an established laboratory-scale technique for the isolation of target proteins with high yield and purity. However, the properties of antibodies and nanobodies often make immunoaffinity chromatography incompatible with conditions typical of many industrial-scale processes. To overcome these limitations, the present invention features an antibody-mimetic scaffold called nanoCLAMP that can be used in process-scale affinity chromatography. The 16 kD antibody mimic is based on a bacterial cysteine-free β-sandwich protein with a structure similar to that of immunoglobulin variable domains. Like antibodies and other antibody mimics, first-generation nanoCLAMPs generally exhibited high selectivity and affinity, but also suffered from sensitivity to high temperatures, protease digestion, and alkaline inactivation. The present invention solves this problem by engineering multiple mutations in the nanoCLAMP scaffold to improve nanoCLAMP's general robustness and resistance to extreme conditions.
[0083] This mutated scaffold served as the basis for an improved class of nanoCLAMPs, termed the nC-B class. Using phage display, hundreds of nC-B capture agents were generated that recognized diverse targets. The resulting immunoaffinity capture agents typically had K values below 80 nM. d , T exceeding 70°C m and t in 0.1 mg / ml trypsin for more than 20 hours 1 / 2The nC-B capture agents also maintained their binding capacity and selectivity over 20 purification cycles, each of which included a 10-minute wash in place with 0.1 M NaOH. Affinity chromatography resins made with the nC-B capture agent supported efficient single-step purification from crude mixtures. Target proteins could be eluted with either 3 M imidazole, pH 8, or 0.1 M sodium citrate, pH 2.5. Furthermore, affinity chromatography resins using the nC-B capture agent remained functional after exposure to 100% DMF and autoclaving. The robust nanoCLAMP scaffold described herein enables the development of custom high-speed affinity chromatography resins compatible with the rigorous conditions of process-scale applications, which may be able to accommodate a wide variety of target substrates.
[0084] Protein Scaffolds The scaffold described herein is derived from the carbohydrate-binding module family 32 (CBM32) protein domain of Clostridium perfringens hyaluronidase (NagH), a multi-domain enzyme consisting of 1627 amino acids. Amino acid residue 1 of SEQ ID NO:1 corresponds to amino acid residue 807 of NagH, and amino acid residue 140 of SEQ ID NO:1 corresponds to amino acid residue 946 of NagH. The amino acid positions and mutations described herein generally relate to the corresponding positions in full-length NagH, unless otherwise indicated. The WT sequence of CBM32 is shown below:
[0085] CBM32 SEQ (SEQ ID NO: 1) [ka]
[0086] Previous studies have identified scaffolds in which three or five loop regions (L1, L2, and L8, or L1, L2, L4, L6, and L8) have been mutagenized to form binders to diverse protein targets in place of carbohydrates, a property not predicted for carbohydrate-binding modules. In some embodiments, the protein scaffold does not retain the carbohydrate-binding activity of, for example, the native CBM scaffold. Loop L1 corresponds to residues 817-820, loop L2 corresponds to residues 838-844, and loop L8 corresponds to residues 931-935. The original scaffold (nC-A) is shown below and contains a single M929L mutation compared to SEQ ID NO: 1. X represents a variable loop residue, and each X can independently be any residue.
[0087] nC-A scaffold sequence (SEQ ID NO: 2) [ka]
[0088] The current scaffold (nC-B) described herein is based on the exemplary scaffold of SEQ ID NO: 3 shown below: X represents variable loop residues, and each X can independently be any residue.
[0089] nC-B scaffold sequence full length (SEQ ID NO: 3) [ka]
[0090] The protein scaffolds described herein comprise (e.g., consist of) framework regions (F) and loop regions (L). A scaffold generally comprises: A-F1-L1-F2-L2-F3-L3-F4-L4-F5-L5-F6-L6-F7-L7-F8-L8-F9-B It has the following structure.
[0091] F1-F9 correspond to framework regions 1-9, and L1-L8 correspond to loop regions 1-8. Framework and loop regions were selected based on where beta strands turn into loops or where beta strands make sharp turns out of the plane of the beta sheet of the strands (see Figures 2 and 16). The N- and C-termini of the scaffold, A and B, can each independently be present (e.g., contain one or more amino acids) or absent.
[0092] F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 4; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5), or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 5; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 6; L3 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 7; L4 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 8; L5 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 9; L6 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 10; L7 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 11; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 includes the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 12.
[0093] As described herein, sequences having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof include, for example, sequences having one insertion, two insertions, one deletion, two deletions, one substitution mutation, two substitution mutations, one insertion and one deletion, one insertion and one substitution mutation, or one deletion and one substitution mutation.
[0094] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X relative to SEQ ID NO: 1, wherein X is any amino acid.
[0095] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X1, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X2, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X3 relative to SEQ ID NO: 1, wherein: X is any amino acid except the amino acid at the equivalent position in SEQ ID NO: 1, X1 is any amino acid except R or S; X2 is any amino acid except P or K, X3 is any amino acid except R or K.
[0096] In some embodiments, F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 4; F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5), or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 5; F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 6; F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 7; F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 8; F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 9; F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 10; F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 11; F9 comprises the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 12.
[0097] In some embodiments, F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4), F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5); F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6), F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7); F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8); F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9), F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10), F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11), F9 contains the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12).
[0098] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X relative to SEQ ID NO: 1, wherein X is any amino acid.
[0099] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of N807D, S809T, R812H, S813T, E814P, S815G, D818V, N822S, N825D, N832S, W836E, K857E, E858V, I859V, K860E, L861V, D862G, R865H, K870A, N871D, N880T, K881R, K883R, N890G, K897R, K901H, K908Q, E912D, S914D, and K922Q relative to SEQ ID NO:1.
[0100] In some embodiments, at least one mutation is K870X and / or N890X. In some embodiments, at least one mutation is K870A and / or N890G. In some embodiments, at least one mutation is K870A. In some embodiments, at least one mutation is N890G.
[0101] In some embodiments, the protein scaffold comprises at least 3 fewer lysines than SEQ ID NO: 1. For example, in some embodiments, the protein scaffold comprises at least 3, 4, 5, 6, 7, 8, 9, or 10 fewer lysines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises at least 6 fewer lysines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises 9 fewer lysines than SEQ ID NO: 1. In some embodiments, the protein scaffold does not comprise any lysines.
[0102] In some embodiments, the protein scaffold comprises at least 3 fewer asparagines than SEQ ID NO: 1. For example, in some embodiments, the protein scaffold comprises at least 3, 4, 5, 6, 7, or 8 fewer asparagines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises at least 5 fewer asparagines than SEQ ID NO: 1. In some embodiments, the protein scaffold comprises 7 fewer asparagines than SEQ ID NO: 1. In some embodiments, the protein scaffold does not comprise any asparagines.
[0103] In some embodiments, A and B are each independently absent or at least one amino acid. For example, each of A and B can be independently absent. In some embodiments, A and B are each independently at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 30, 400, 500, 600, 700, 800, 900, 1,000 or more amino acids. In some embodiments, A and B each independently represent 0 to 1,000 amino acids, e.g., 1 to 10 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids), 10 to 100 amino acids (e.g., 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids), or 100 to 1,000 amino acids (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 amino acids).
[0104] In some embodiments, A and B are each independently zero or between 1 and 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids).
[0105] In some embodiments, L1 to L8 are each independently 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids).
[0106] In some embodiments, L1 to L8 are each independently 1 to 10 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids). In some embodiments, L1 to L8 are each independently 3 to 10 amino acids. In some embodiments, L1 to L8 are each independently 3 to 8 amino acids.
[0107] In some embodiments, L1 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), In some embodiments, L1 is 0 to 5 amino acids (e.g., 1 to 5 amino acids, e.g., 0, 1, 2, 3, 4, or 5 amino acids).
[0108] In some embodiments, L2 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L2 is 1 to 16 amino acids (e.g., 4 to 16 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 amino acids).
[0109] In some embodiments, L3 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L3 is 6 amino acids.
[0110] In some embodiments, L4 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L4 is 0 to 5 amino acids (e.g., 1 to 5 amino acids, e.g., 0, 1, 2, 3, 4, or 5 amino acids).
[0111] In some embodiments, L5 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, L5 is 5 amino acids.
[0112] In some embodiments, L6 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), hi some embodiments, L6 is 3 to 6 amino acids (e.g., 3, 4, 5, or 6 amino acids).
[0113] In some embodiments, L7 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), hi some embodiments, L7 is 4 or 5 amino acids.
[0114] In some embodiments, L8 is 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids), hi some embodiments, L8 is 4 to 6 amino acids (e.g., 4, 5, or 6 amino acids).
[0115] In some embodiments, L1 is 4 amino acids. In some embodiments, L2 is 7 amino acids. In some embodiments, L8 is 5 amino acids. In some embodiments, L1 is 4 amino acids, L2 is 7 amino acids, and / or L8 is 5 amino acids. In some embodiments, L1 is 4 amino acids, L2 is 7 amino acids, and L8 is 5 amino acids.
[0116] In some embodiments, L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1-X4 are each independently any amino acid. In some embodiments, X2 is V.
[0117] In some embodiments, L2 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1-X7 are each independently any amino acid.
[0118] In some embodiments, L8 comprises the sequence X1X2X3X4X5 (SEQ ID NO: 15), where X1-X5 are each independently any amino acid.
[0119] In some embodiments, L4 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1-X7 are each independently any amino acid.
[0120] In some embodiments, L6 comprises the sequence X1X2X3X4X5X6 (SEQ ID NO: 16), where X1-X6 are each independently any amino acid. In some embodiments, L8 comprises at least two amino acids. In some embodiments, L8 comprises at least one amino acid.
[0121] In some embodiments, L4 comprises the sequence of (G / D)-GGSS (SEQ ID NO: 17) or GDT, or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 17 or GDT.
[0122] In some embodiments, L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 18.
[0123] In some embodiments, L4 comprises the sequence (G / D)-GGSS (SEQ ID NO: 17) or GDT, and L6 comprises the sequence TGAPAG (SEQ ID NO: 18).
[0124] In some embodiments, L3 comprises the sequence (E / K / S)-(V / E)-(V / I / T)-(E / K / P / S)-(V / L)-(G / D) (SEQ ID NO: 19), or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 19.
[0125] In some embodiments, L5 comprises the sequence LD-(G / N)-(E / S)-S (SEQ ID NO: 20) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 20.
[0126] In some embodiments, L7 comprises at least one amino acid.
[0127] In some embodiments, L7 comprises the sequence of ETPI-(S / E)-A (SEQ ID NO: 21) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 21.
[0128] In some embodiments, L3 comprises the sequence (E / K / S)-(V / E)-(V / I / T)-(E / K / P / S)-(V / L)-(G / D) (SEQ ID NO: 19) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 19; L5 comprises the sequence LD-(G / N)-(E / S)-S (SEQ ID NO: 20) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 20; and L7 comprises the sequence ETPI-(S / E)-A (SEQ ID NO: 21) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 21.
[0129] In some embodiments, A comprises the sequence (D / N / H)-P. In some embodiments, A comprises the sequence DP.
[0130] In some embodiments, B comprises the sequence of DELE (SEQ ID NO: 35).
[0131] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 24; L3 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 25; L4 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F5 comprises the sequence of DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 26; L5 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F6 comprises the sequence of WTTIREYDH (SEQ ID NO: 27) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 27; L6 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 28; L7 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F8 comprises the sequence of QYIRLTNLE (SEQ ID NO: 29) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 30.
[0132] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22), L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 contains the sequence LTFSEFAIVS (SEQ ID NO: 30).
[0133] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 24; L3 comprises the sequence of EVVEVG (SEQ ID NO: 31) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 31; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 25; L4 comprises the sequence GGGSS (SEQ ID NO: 32) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 32; F5 comprises the sequence of DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 26; L5 comprises the sequence of LDGES (SEQ ID NO: 33) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 33; F6 comprises the sequence of WTTIREYDH (SEQ ID NO: 27) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 27; L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 18; F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 28; L7 comprises the sequence of ETPISA (SEQ ID NO: 34) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 34; F8 comprises the sequence of QYIRLTNLE (SEQ ID NO: 29) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 30.
[0134] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) relative to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 24; L3 comprises the sequence of EVVEVG (SEQ ID NO: 31) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 31; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 25; L4 comprises the sequence GGGSS (SEQ ID NO: 32) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 32; F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 26; L5 comprises the sequence of LDGES (SEQ ID NO: 33) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 33; F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 27; L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having a single amino acid insertion, deletion or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 18; F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 28; L7 comprises the sequence of ETPISA (SEQ ID NO: 34) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 34; F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29) or a sequence having a single amino acid insertion, deletion, or substitution mutation (e.g., a single substitution mutation) compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30) or a sequence having a single amino acid insertion, deletion, or substitution mutation (eg, a single substitution mutation) compared to SEQ ID NO: 30.
[0135] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22), L1 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 is absent or comprises at least one amino acid (e.g., 1 to 20 amino acids, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids); F9 contains the sequence LTFSEFAIVS (SEQ ID NO: 30).
[0136] In some embodiments, F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22), L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1 to X4 are each independently any amino acid; F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1 to X7 are each independently any amino acid; F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 comprises the sequence X1X2X3X4X5 (SEQ ID NO: 15), where X1 to X5 are each independently any amino acid; F9 contains the sequence LTFSEFAIVS (SEQ ID NO: 30).
[0137] In some embodiments, A comprises the sequence of DP, F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22), L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1 to X4 are each independently any amino acid; F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 comprises the sequence X1X2X3X4X5X6X7 (SEQ ID NO: 14), where X1 to X7 are each independently any amino acid; F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 comprises the sequence X1X2X3X4X5 (SEQ ID NO: 15), where X1 to X5 are each independently any amino acid; F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30), B contains the sequence of DELE (SEQ ID NO: 35).
[0138] In some embodiments, L1 comprises the sequence X1X2X3X4 (SEQ ID NO: 13), where X1, X3, and X4 are each independently any amino acid, and X2 is V.
[0139] Another aspect features a protein scaffold that includes a polypeptide having at least 80% (e.g., at least 85%, 90%, 95%, 97%, or 99%) sequence identity to SEQ ID NO:3. In some embodiments, the polypeptide includes the sequence of SEQ ID NO:3. In some embodiments, the polypeptide does not include the sequence of SEQ ID NO:1. In some embodiments, the polypeptide does not include the sequence of SEQ ID NO:2.
[0140] Another aspect features a polypeptide having at least 85% (e.g., at least 90%, 95%, 97%, 99%, or 100%) sequence identity to a polypeptide of Table 9 or Table 10. In some embodiments, the polypeptide comprises a sequence set forth in Table 9 or Table 10.
[0141] In some embodiments, the polypeptide comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X1, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X2, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X3 relative to SEQ ID NO: 1, wherein: X is any amino acid except the amino acid at the equivalent position in SEQ ID NO: 1, X1 is any amino acid except R or S; X2 is any amino acid except P or K, X3 is any amino acid except R or K.
[0142] Those skilled in the art will understand that the protein scaffolds described herein containing nine framework regions (i.e., F1-F9) can be optimized or replaced according to established biophysical techniques. Accordingly, the present invention also features protein scaffolds containing seven of the nine framework regions or eight of the nine framework regions described herein. Based on detailed structural analyses of scaffolds known in the art (see, e.g., Ficko-Blean et al., J. Mol. Bio. 390:208-220, 2009) and PDB ID 2w1q, one skilled in the art can swap one or more beta strands of the core protein fold or portions thereof (e.g., more than two residues in any one of a given framework region, e.g., F1-F9) while maintaining the structural integrity of the entire scaffold. Phage libraries can be generated that express protein scaffolds, each of which has a loop that confers binding to a specific target, and in which the amino acids at each position in a specific beta strand are randomized. The phage library can then be selected for library members that are thermostable and maintain binding to the target by selecting the library for members that can withstand incubation above 55°C without aggregation and for members that can bind to the immobilized target. Isolated clones with these properties represent scaffolds with shuffled beta strands or portions thereof.
[0143] Thus, in some embodiments, the present invention also contemplates a protein scaffold having at least seven, e.g., at least eight, of the following framework regions, wherein seven of the nine framework regions or eight of the nine framework regions have the following sequence: F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 4; F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5), or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO: 5; F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 6; F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 7; F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 8; F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 9; F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 10; F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 11; F9 includes the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 12.
[0144] In other embodiments, the present invention also contemplates a protein scaffold having at least seven, e.g., at least eight, of the following framework regions, wherein seven of the nine framework regions or eight of the nine framework regions have the following sequence: F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 22; F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 23; F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 24; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 25; F5 comprises the sequence of DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 26; F6 comprises the sequence of WTTIREYDH (SEQ ID NO: 27) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 27; F7 comprises QDVIDEDF (SEQ ID NO:28) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) relative to SEQ ID NO:28; F8 comprises the sequence of QYIRLTNLE (SEQ ID NO: 29) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 29; F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof (e.g., one or two substitution mutations) compared to SEQ ID NO: 30.
[0145] In other embodiments, the present invention also contemplates a protein scaffold having at least seven, e.g., at least eight, of the following framework regions, wherein seven of the nine framework regions or eight of the nine framework regions have the following sequence: F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22), F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), F9 contains the sequence LTFSEFAIVS (SEQ ID NO: 30).
[0146] In another aspect, the invention features a protein scaffold that includes a polypeptide having at least 80% (e.g., at least 85%, 90%, 95%, 97%, or 99%) sequence identity to framework regions (F1-F9) over the region of alignment corresponding to F1-F9 of a reference sequence (e.g., SEQ ID NO: 3).
[0147] In some embodiments, the protein scaffold comprises one or more unnatural amino acids. In some embodiments, one or more of the framework regions comprises unnatural amino acids. In some embodiments, one or more of the loop regions comprises unnatural amino acids.
[0148] Cysteine mutations and disulfide bridges The protein scaffolds described herein may lack natural cysteine residues. Thus, the scaffolds can be mutagenized to introduce one or more cysteine residues into the scaffold (e.g., in one or more loop or framework regions). When two or more cysteine residues are introduced into the scaffold at adjacent sites, the two cysteine residues can form disulfide bridges, for example, under oxidizing conditions. In some embodiments, the disulfide bridges enhance the thermal stability of the protein scaffold.
[0149] In some embodiments, the protein scaffold comprises a mutation that adds a cysteine residue. In some embodiments, the protein scaffold comprises a first mutation that adds a first cysteine residue and a second mutation that adds a second cysteine residue. In some embodiments, the first cysteine residue and the second cysteine residue form a disulfide bond under oxidizing conditions.
[0150] In some embodiments, the protein scaffold comprises at least one mutation selected from the group consisting of F806C, P808C, S845C, L855C, V858C, V861C, K878C, W879C, L884C, L888C, A904C, P905C, A906GC, G907C, I924C, L926C, N928C, L936C, I943C, L948C.
[0151] In some embodiments, the protein scaffold comprises at least two or more mutations selected from the group consisting of F806C, P808C, S845C, L855C, V858C, V861C, K878C, W879C, L884C, L888C, A904C, P905C, A906GC, G907C, I924C, L926C, N928C, L936C, I943C, L948C.
[0152] In some embodiments, the protein scaffold comprises a cysteine mutation pair selected from the group consisting of K878C and G907C, K878C and A904C, V861C and I943C, P905C and L855C, S845C and L936C, W879C and N928C, L884C and L926C, F806C and L948C, V858C and L888C, K878C and G907C, K878C and A906GC, S845C and N928C, K878C and A904C, P808C and I943C, V861C and I924C, P808C and V861C, and I943C and L855C.
[0153] In some embodiments, the cysteine mutation pair is selected from the group consisting of K878C and G907C, K878C and A904C, S845C and L936C, W879C and N928C, W879C and N928C, L884C and L926C, V858C and L888C, K878C and G907C, and K878C and A906GC.
[0154] Tags and Functional Groups The protein scaffolds described herein may further comprise a tag. The tag may provide for ease of purification, detection, or binding of the protein scaffold. The tag may be covalently attached to the scaffold. In some embodiments, A and / or B of the scaffold are or comprise a tag.
[0155] In some embodiments, the tag is an affinity tag (e.g., a polyhistidine tag, e.g., 4, 5, 6, 7, 8, 9, or 10 histidines, e.g., a Gly-His tag, e.g., an AviTag, e.g., a calmodulin-tag, e.g., a polyglutamate tag, e.g., a polyarginine tag, e.g., an SBP-tag).
[0156] In some embodiments, the tag is an epitope tag (e.g., ALFA-tag, C-tag, iCapTag, E-tag, FLAG-tag, HA-tag, Myc-tag, NE-tag, Rho1D4-tag, S-tag, Softag1, Softag3, Spot-tag, T7-tag, TC-tag, Ty-tag, V6-tag, VSV-tag, or Xpress-tag).
[0157] In some embodiments, the tag is a covalent protein tag (eg, Isopeptag, SpyTag, SnoopTag, DogTag, or SdyTag).
[0158] In some embodiments, the tag is a protein tag, such as a biotin carboxyl carrier protein tag, a glutathione-S-transferase (GST) tag, a green fluorescent protein (GFP) tag, a HaloTag, a SNAP-tag, a CLIP-tag, a HUH-tag, a maltose binding protein tag, a Nus tag, a thioredoxin tag, an Fc tag, an engineered intrinsically disordered tag, a CRDSAT tag, a SpyCatcher, a SnoopCatcher, a DogCatcher, a SdyCatcher, or a SUMO-tag.
[0159] In some embodiments, scaffold A and / or B comprise an affinity tag, an epitope tag, a covalent peptide tag, or a protein tag.
[0160] In some embodiments, the tag is attached to the N-terminus or C-terminus of the scaffold.
[0161] In some embodiments, the scaffold is conjugated to a functional group, which in some embodiments comprises biotin, streptavidin or a derivative of streptavidin, a polyethylene glycol moiety, a fluorescent dye, an enzyme, a radioactive moiety, a lanthanide, or a lanthanide binding motif.
[0162] In some embodiments, the scaffold is conjugated to a lanthanide or a lanthanide binding motif, hi some embodiments, the lanthanide is terbium.
[0163] In some embodiments, the scaffold is conjugated to a radioactive moiety, hi some embodiments, the radioactive moiety is an alpha or beta emitter.
[0164] In some embodiments, the functional group is conjugated to a sulfhydryl group or a primary amine (eg, on a cysteine residue or lysine).
[0165] Polynucleotides, vectors, and cells The protein scaffolds described herein can be encoded by a polynucleotide. In some embodiments, the polynucleotide is a ribonucleotide. In some embodiments, the polynucleotide is a deoxyribonucleotide. Also contemplated herein are vectors comprising a polynucleotide encoding a protein scaffold.
[0166] Other embodiments feature a cell comprising a polynucleotide encoding a protein scaffold or a vector comprising the polynucleotide. The polynucleotide or vector can include expression elements configured to drive expression of the protein scaffold. The cell can be a prokaryotic cell (e.g., E. coli). The cell can be a eukaryotic cell. In some embodiments, the eukaryotic cell is a yeast cell (e.g., S. cerevisiae) or a mammalian cell (e.g., a Chinese hamster ovary (CHO) cell). In some embodiments, the protein scaffold is secreted by the cell. In some embodiments, the protein scaffold is expressed intracellularly. Such cells (e.g., E. coli) can be lysed to provide a lysate comprising the protein scaffold.
[0167] Also featured herein are methods of producing the protein scaffolds described herein. The methods include providing a cell transformed with a polynucleotide encoding the protein scaffold or a vector containing the polynucleotide, and culturing the transformed cell under conditions that allow expression of the polynucleotide. The culturing step results in expression of the protein scaffold. The methods can further include isolating the protein scaffold or using the protein scaffold to bind to a target.
[0168] Particles, resins, and columns The protein scaffolds described herein can be conjugated to particles. In some embodiments, the particles are magnetic particles. Also featured are, for example, resins or monoliths comprising a plurality of particles containing the protein scaffold. Also contemplated herein are, for example, columns (e.g., chromatography columns) containing particles or resins conjugated to the scaffold.
[0169] The scaffold and its method of use can use a surface linked to the protein scaffold, the surface configured to bind to its target. The resin surface refers to the portion of the support structure (e.g., substrate) that is accessible for contact with one or more target molecules. The shape, form, material, and modification of the resin surface can be selected from a variety of options depending on the application. In one embodiment, the resin surface is SEPHAROSE®. In one embodiment, the resin surface is agarose.
[0170] The surface of the resin can be substantially flat or planar. Alternatively, the surface of the resin can be rounded or contoured. Exemplary contours that may be included in the surface of the resin are wells, depressions, pillars, ridges, channels, etc.
[0171] In one embodiment, the surface of the resin is modified to contain channels, patterns, layers, or other configurations (e.g., patterned surfaces). The surface can be in the form of a bead, box, column, cylinder, disk, dish (e.g., glass dish, PETRI dish), fiber, film, filter, microtiter plate (e.g., 96-well microtiter plate), multi-blade stick, net, pellet, plate, ring, rod, roll, sheet, slide, stick, tray, tube, or vial. The surface can be a single individual entity (e.g., a single tube, a single bead), any number of multiple surface entities (e.g., a rack of 10 tubes, several beads), or a combination thereof (e.g., a tray containing multiple microtiter plates, a column filled with beads, a microtiter plate filled with beads).
[0172] In some embodiments, the surface may comprise a membrane-based resin matrix. In some embodiments, the resin surface comprises a porous or non-porous resin. Examples of porous resins include additional agarose-based resins (e.g., cyanogen bromide-activated SEPHAROSE® (GE), WORKBEADS™ 40 ACT, and WORKBEADS™ 40 / 10000 ACT (Bioworks)), methacrylates (such as Tosoh 650M derivatives), polystyrene divinylbenzene (LifeTech Poros media / GE Source media), Fractogel, polyacrylamide, silica, controlled pore glass, dextran derivatives, acrylamide derivatives, convective interaction media (Sartorius), additional polymers, and combinations thereof.
[0173] In some embodiments, the surface can include one or more pores, hi some embodiments, the pore size can be between 300 and 8,000 angstroms, for example, between 500 and 4,000 angstroms.
[0174] The resins described herein include a plurality of particles. Example particle sizes are 5 μm to 500 μm, 20 μm to 300 μm, and 50 μm to 200 μm. In some embodiments, the particle size can be 50 μm, 60 μm, 70 μm, 80 μm, 90 μm, 100 μm, 110 μm, 120 μm, 130 μm, 140 μm, 150 μm, 160 μm, 170 μm, 180 μm, 190 μm, or 200 μm.
[0175] The protein scaffold can be immobilized, coated, bonded, affixed, adhered, or attached to any of the surface forms described herein (e.g., beads, boxes, columns, cylinders, discs, dishes (e.g., glass dishes, PETRI dishes), fibers, films, filters, microtiter plates (e.g., 96-well microtiter plates), multi-blade sticks, nets, pellets, plates, rings, rods, rolls, sheets, slides, sticks, trays, tubes, or vials).
[0176] Purification method Featured herein are methods for purifying a target molecule, e.g., from a plurality of molecules, e.g., from a crude lysate. The method includes providing a sample containing a mixture of the target molecule and the plurality of molecules, and contacting the sample with a protein scaffold described herein. The protein scaffold may be pre-generated with a loop region specific for a desired target. The scaffold (e.g., the loop region of the scaffold) specifically binds to the target molecule. The method further includes separating the target molecule bound to the protein scaffold from the plurality of molecules. In some embodiments, the separating step includes immobilizing the protein scaffold. In some embodiments, the protein scaffold is conjugated to a particle. In some embodiments, the particle comprises a magnetic bead. In some embodiments, the protein scaffold is conjugated to a resin or monolith described herein.
[0177] The scaffolds, particles, resins, and columns described herein are suitable for single-step purification from crude mixtures. For example, target proteins can be eluted with polyols, imidazole (e.g., 3 M imidazole, e.g., pH 8), or sodium citrate (e.g., or 0.1 M sodium citrate, e.g., pH 2.5). The scaffolds, particles, resins, and columns can be washed with, for example, alkaline substances, such as NaOH, e.g., 0.1 M NaOH. Furthermore, resins or columns made using the protein scaffolds described herein can remain functional after exposure to dimethylformamide (DMF), e.g., 100% DMF, and / or autoclaving. These features allow for the reuse of the scaffolds without loss of target binding capacity.
[0178] [Example] Example 1: Protein Engineering of NanoCLAMP Antibody-Mimetic for Use as an Affinity Chromatography Capture Agent that is Resistant to High Temperature, Trypsin, Low pH, Organic Solvents, and Sodium Hydroxide Immunoaffinity chromatography is an established laboratory-scale technique for the isolation of target proteins with high yield and purity. However, the properties of antibodies and nanobodies often make immunoaffinity chromatography incompatible with conditions typical of many industrial-scale processes. To overcome these limitations, we optimized an antibody-mimetic called nanoCLAMP for use in process-scale affinity chromatography. The 16 kD antibody mimic is based on a bacterial cysteine-free β-sandwich protein with a structure similar to that of immunoglobulin variable domains. Like antibodies and other antibody mimics, first-generation nanoCLAMPs generally exhibited high selectivity and affinity, but also suffered from sensitivity to high temperatures, protease digestion, and alkaline inactivation. In this study, we address these limitations with a protein engineering campaign to improve the general robustness of nanoCLAMP. Over seven rounds of mutagenesis and screening, we examined 185 protein variants with at least one mutation in 58% (72 of 124) of the positions in the constant region of nanoCLAMP. The campaign yielded proteins with mutations at 30 of 124 positions in the constant region and dramatically improved resistance to extreme conditions. The mutant proteins served as the basis for an improved class of nanoCLAMPs, termed the nC-B class. Phage display was then used to generate several nC-B capture agents that recognized diverse targets. The resulting immunoaffinity capture agents typically had K values less than 80 nM. d , T exceeding 70°C m and t in 0.1 mg / ml trypsin for over 20 hours 1 / 2The nC-B capture agent also maintained its binding capacity and selectivity over 20 purification cycles, each of which included a 10-minute wash in place with 0.1 M NaOH. Affinity chromatography resins fabricated with the nC-B capture agent supported efficient single-step purification from crude mixtures. The target protein could be eluted with either 3 M imidazole, pH 8, or 0.1 M sodium citrate, pH 2.5. Furthermore, affinity chromatography resins using the nC-B capture agent remained functional after exposure to 100% DMF and autoclaving. The robust nC-B scaffold developed in this study will enable the development of custom high-speed affinity chromatography resins compatible with the demanding conditions of process-scale applications.
[0179] introduction Affinity chromatography (AC) using target-specific immobilized capture agents is an established method for protein purification. In this technique, capture agents, e.g., proteins, nucleic acids, or small molecules, are coupled to a solid support and can then be used to isolate proteins of interest from complex mixtures. The technique has been widely used on a laboratory scale for the single-step purification of a variety of target proteins, including enzymes, transcription factors, growth factors, and antibodies.
[0180] The use of protein-based capture agents for AC in industrial applications has not been widespread because currently available approaches are either incompatible with the temperatures, extreme pH, and solvents often required for process-scale purification or are useful for only a limited number of targets. The purification of kilogram quantities of antibodies using an AC resin based on Staphylococcus aureus protein A is an exception. The development of Protein A resins highlights the use of protein engineering to improve AC resin robustness and some remaining limitations. Early versions of the resins using wild-type Protein A captured antibodies from cell culture media feedstocks with high selectivity and capacity, but gradually lost activity after multiple cycles of cleaning-in-place with sodium hydroxide. Mutagenesis of Protein A has yielded variants with increased resistance to sodium hydroxide treatment and higher binding capacity. Despite its widespread use, Protein A resins are limited to antibody purification.
[0181] For process-scale purification of non-antibody targets, the use of AC has been much less widespread than the use of Protein A for purifying antibodies. For example, non-protein ligand-based approaches, such as small-molecule substrate mimics, are effective but are limited to specific enzyme classes and difficult to use with general proteins of interest. Alternatively, specialized affinity resins, such as glutathione or nickel, require the addition of non-native tags, which introduces downstream complications for proteins targeted for therapeutic use. Immuno-AC, using antibody- or nanobody-based capture agents, is the most generally applicable approach and has been widely used to purify a variety of proteins at the laboratory scale. However, immuno-AC has several limitations. Generally, conjugation of antibodies to resins often results in heterogeneous coupling due to a lack of precise control over the site of conjugation. Furthermore, chromatography must be performed under oxidizing conditions to preserve disulfide bonds, which are essential for maintaining antibody structure. Also, target elution typically requires low or high pH conditions that are incompatible with some target proteins. For process-scale applications, the main limitation of Immuno-AC resins is their sensitivity to sodium hydroxide solutions, which are preferred for clean-in-place procedures.
[0182] These limitations have motivated the development of AC resins based on antibody mimetics, proteins that can be engineered to bind specific antigens with high affinity and specificity, like antibodies, but that are not directly derived from animal immune systems. Examples of antibody mimetics include those based on protein A, gamma-β crystallin, ubiquitin, cystatin, lipocalin, ankyrin repeat motifs, SH3 domains, fibronectin, OB-fold domains, lamprey variable lymphocyte receptor, minibodies, miniproteins, and Kunitz domains. Most of these antibody mimetics use animal-sparing phage display for their isolation and can be produced by microbial cells. Many possess unique cysteines that support uniform, site-specific coupling to sulfhydryl-reactive supports. However, most current antibody mimetics have not been shown to allow elution near neutral pH or to be compatible with the sometimes stringent conditions required for process-scale procedures. The development of custom peptide and protein-based affinity chromatography platforms for the purification of non-antibody protein therapeutics has been supported by several proprietary platforms, although the availability and technical details of these platforms are limited (e.g., Avitide, LigaTrap Technologies, Astrea Bioseparations, and Navigo Proteins).
[0183] The present inventors set out to develop a widely available and generally applicable class of protein-based affinity capture reagents useful for industrial protein purification. Specifically, the inventors sought to develop an AC capture agent technology that would address a wide range of targets, allow single-step purification from crude mixtures, elute targets at near-neutral pH, and maintain functionality after exposure to high temperatures, organic solvents, proteases, and extremes of pH.
[0184] We previously developed antibody mimics based on the 16 kD second type 32 carbohydrate-binding module (NagH CpCBM32-2) of the hyaluronoglucosaminidase nagH from Clostridium perfringens (Suderman et al., Protein Expr. Purif. 134:114-124, 2017). This binding module is a monomeric β-sandwich domain with variable loops that mimic the complementarity-determining regions of immunoglobulin variable domains. We named these antibody mimics nanoCLAMPs (nanoclostridial antibody-mimicking proteins). C lostridial A Antibody M imetic P NanoCLAMP has the unusual and favorable general property of releasing bound target proteins in solutions of non-denaturing polyols and ammonium sulfate at neutral pH. We have demonstrated that the K ranges from 1-100 nM before affinity maturation and 10-1000 pM after affinity maturation. d NanoCLAMPs have been isolated that recognize a variety of target proteins (Suderman et al., 2017). Affinity chromatography media produced using nanoCLAMPs support single-step purification to near homogeneity as assessed by Coomassie staining. Practical binding capacities range from 5 to 200 nmol of target protein per ml of packed beads. While these first-generation nanoCLAMP resins have sufficient selectivity and capacity for laboratory-scale purification, they are hampered by their moderate thermostability (Tmax ranging from 45 to 60 °C). m ), sensitivity to protease digestion (t<1 h in 0.1 mg / ml trypsin) 1 / 2 ), and moderate alkaline resistance (50% loss of activity after 12 cycles of incubation with 0.1 M NaOH), making it less than optimal for process-scale purification.
[0185] We sought to improve the performance of first-generation nanoCLAMPs to improve their general utility for process-scale purification of targets using AC. Toward this goal, we undertook a multi-round protein engineering campaign to improve the first-generation nanoCLAMPs. In each round, we generated site-directed mutations at specific positions, assessed the effect of the mutations on thermostability and monodispersity, and then combined beneficial mutations to generate the basis for the next round. The end product of the seven-round, over 180-round mutation campaign was a clone with significantly improved performance. We used this clone as the basis for the new "nC-B" class of nanoCLAMPs. Here, we report on our campaign to develop the improved nC-B class, including the generation and characterization of an nC-B nanoCLAMP (SMT3) for the exemplary protein yeast SUMO, and the performance of the nC-B nanoCLAMP after repeated exposure to extreme conditions.
[0186] result An approach to improving nanoCLAMP performance using a multi-round campaign of site-directed mutagenesis. We aimed to improve the thermal, proteolytic, and alkaline stability of the first-generation nanoCLAMP by using consensus protein design, an approach to improving protein thermal stability. Our initial attempts to directly synthesize several versions of the consensus sequence failed, resulting only in aggregated or multimeric proteins. Therefore, we decided to take an incremental approach, creating single mutations and determining their effects, alone or in combination, to achieve an improved protein over several rounds. We focused our efforts on surface residues and loops, generally striving to remove lysines, arginines, and asparagines whenever possible. The intention behind substituting lysines and arginines was to reduce the number of potential sites for trypsin cleavage. The intention behind substituting asparagines was to increase the protein's resistance to alkaline solutions commonly used to sanitize industrial chromatography columns. Asparagine, under certain circumstances, is susceptible to deamidation, and its removal has been shown to reduce loss of protein binding activity in sodium hydroxide.
[0187] We refer to each nanoCLAMP as a "clone," i.e., a specific isolate with a unique binding loop. Our starting clone consisted of an anti-SMT3 nanoCLAMP, SMT3-A1 (Suderman et al., supra), which allows for the purification of SUMO fusion proteins from complex lysates in a single step. The target of SMT3-A1 is the SUMO tag, which is widely used to improve the solubility and yield of proteins produced in E. coli and can be cleaved to leave the native sequence. SMT3-A1 consists of residues 807-946 of the second type 32 carbohydrate-binding module of NagH from Clostridium perfringens mutated at positions 817-820, 838-844, and 931-935, with the loop selected to bind yeast SUMO. Throughout the paper, we number the nanoCLAMP amino acids based on the NagH sequence. To identify evolutionarily conserved amino acids, we generated a multiple sequence alignment of 20 nonredundant BLAST hits selected to encompass a range of similarities. The percent identity of orthologs ranged from 58% for Clostridium nigeriense to 43% for Coprobacillus sp. AF21-8LB (Table 1).
[0188] [Table 1]
[0189] Our workflow for the protein engineering campaign is summarized in Figure 1. We generated site-specific mutations and purified each individual mutant protein for further biophysical evaluation. For initial evaluation, we measured the melting temperatures of the resulting mutants using differential scanning fluorimetry (DSF). We then evaluated two parameters: melting temperature and initial fluorescence. The rationale for including low initial fluorescence as a progress criterion is that we have previously observed a rough correlation between high initial fluorescence and the presence of soluble aggregates and multimers detected by size-exclusion chromatography. Initial fluorescence in DSF can be caused by binding of fluorophores to hydrophobic patches exposed prior to unfolding.
[0190] After identifying beneficial mutations, we generated several constructs with various combinations of the singly beneficial mutations, and the effects were usually, but not always, additive. As a secondary screen, we confirmed that the starting clones for each subsequent round of mutagenesis were monodisperse by size-exclusion chromatography.
[0191] Optimized proteins resulting from mutagenesis campaigns. The results of seven rounds of mutagenesis and evaluation are summarized in Table 2. The complete list of mutations is listed in Table 3.
[0192] [Table 2]
[0193] [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4] [Table 3-5] [Table 3-6] [Table 3-7] [Table 3-8] [Table 3-9] [Table 3-10] [Table 3-11] [Table 3-12] [Table 3-13] [Table 3-14] [Table 3-15] [Table 3-16] [Table 3-17] [Table 3-18]
[0194] Rounds 1 to 4 are T m Rounds 5 and 6 focused on increasing T mRound 7 involved many reversals to previous rounds, focusing on removing potential protease cleavage sites while maintaining T m We focused on removing the remaining asparagines while maintaining the constant region constant. Overall, mutations were examined for 58% of the amino acids in the constant region (72 of 124 positions). The clone resulting from seven rounds of mutagenesis was designated P2788. Overall, P2788 contained 30 mutations, representing approximately 24% of the positions in the constant region, and included a three-residue C-terminal extension derived from CBM32-2. The resulting mutations were distributed throughout the primary sequence, as shown by alignment with the original sequence (Figure 2), and throughout the 3D structure when mapped to the CBM32-2 crystal structure (PDB accession 2W1Q). Although all mutants retained the same binding loop as the initial clone, we predicted and observed a gradual decline in target binding with increasing number of mutations, many of which were adjacent to the binding loop. We speculate that the loss of binding was caused by a shift in the conformation of the binding loop (data not shown). Because our intention was to improve the stability of nanoCLAMP and then isolate new binders, our workflow did not include screening for SUMO binding.
[0195] The number of lysines and arginines representing potential trypsin cleavage sites was reduced from 11 in the constant region of the starting protein (clone SMT3-A1) to five in the constant region of the resulting protein (clone P2788). Three of the remaining arginines (R881, R897, and R925) are predicted to be involved in salt bridges as identified by the ESBRI algorithm (Costantini et al., ESBRI: a web server for evaluating salt bridges in proteins. Bioinformation 3:137-138, 2008). For these residues, we were unable to identify any substitution mutations that did not destabilize the protein. For another position, K883, substitution with arginine was beneficial, but several additional substitutions all resulted in a more than 10°C decrease in melting point or high initial fluorescence in DSF. For the one remaining position, K878, which is universally conserved in the alignment, 10 of 10 substitution mutations resulted in proteins with high initial fluorescence in DSF. In the wild-type NagH CpCBM32-2 structure, K878 Nζ forms hydrogen bonds with the carbonyl oxygens of P905 and G907 as determined by the RING 2.0 algorithm (Piovesan et al., Nucleic Acids Res 44:W367-374, 2016). The number of asparagines was reduced from eight in the original clone (SMT3-A1) to one in the mutant clone. The remaining asparagine, N928, is universally conserved in the consensus alignment and is buried in the 3D structure. N928 Nδ 2 forms hydrogen bonds with the carbonyl oxygens of S845 and L846 as determined by the RING 2.0 algorithm. We chose not to attempt substitution mutagenesis with N928 due to the low likelihood of deamidation based on its sequence context and the potential difficulty of finding a substitution with beneficial effect.
[0196] We next predicted the 3D structure of P2788 using AlphaFold to assess the possibility of large-scale changes in the 3D structure (Jumper et al., Nature 596:583-589, 2021). A 3D alignment of the crystal structure of CBM32-2 and the predicted structure of P2788 was performed using the jFATCAT (rigid) algorithm (Figure 3). As expected with conservative substitutions, high similarity, and the use of templates by the AlphaFold algorithm, the predicted structure of the constant region of P2788 shows no large-scale deviations from the solved crystal structure of NagH CpCBM32-2. As expected, due to differences in amino acid sequence, the variable loops show deviations, especially for the longest 838-844 loop. The overall similarity between the solved CBM32-2 structure and the predicted structure of P2788 is high (TM score = 0.95), even with differences in loop sequence.
[0197] Compared to the starting protein, T m The temperature increased by 24°C from 52°C to 76°C (Fig. 4). We next examined the resistance of P2788 to digestion with trypsin. After 16 hours of digestion in 0.1 mg / ml trypsin, no full-length SMT3-A1 remained, as assessed by SDS-PAGE, whereas no obvious digestion of P2788 occurred (Fig. 5A). The time course of trypsin digestion revealed that t 1 / 2 The cleavage time was determined to increase from 3 h for SMT3-A1 to over 16 h for P2788 (Figures 5C–5E). While our protein engineering campaign deleted only a few surface-exposed chymotrypsin cleavage sites, we also examined resistance to chymotrypsin as a reflection of general stability. After 16 h of digestion with 0.1 mg / ml chymotrypsin, approximately one-fifth of P2788 appeared to remain full-length, while SMT3-A1 was completely digested (Figure 5A).
[0198] Next, we sought to determine whether a new class of nanoCLAMPs based on the constant region of P2788 could confer these properties to newly isolated clones. We designated the P2788-derived class as the "nC-B class" (nanoCLAMP-B, with the identifier "B" referring to the next variant of the original class of nanoCLAMPs). n ano C The first generation nanoCLAMPs, represented by SMT3-A1 and others, are referred to as "nC-A class" (with the "A" designating the first class of nanoCLAMPs). n ano C LAMP- A )). For clarity, the nC-A class of nanoCLAMPs encompasses the first published nanoCLAMP (Suderman et al., supra). Compared to NagH CpCBM32-2, the nC-A class of nanoCLAMPs has an M929L mutation that removes a methionine in the variable loop as well as an amino acid difference.
[0199] Phage display library panning for improved SUMO capture agents of the nC-B class of nanoCLAMPs. To confirm that the optimized constant region of the nC-B class can support the general isolation of high-affinity binders with improved thermal, proteolytic, and alkaline stability, we sought to isolate new nanoCLAMP binders containing the nC-B constant region and variable loop diversity. We first constructed a phage display library with randomized binding loops in the context of the nC-B constant region and panned the library for binders to yeast SUMO (SMT3). This library has the same three variable loops with randomized residues as the previous library from which SMT3-A1 was isolated, but uses the nC-B constant region instead of the nC-A constant region. Degenerate oligonucleotides constructed using phosphoramidite trimers were designed so that the variable regions encode all amino acids except cysteine (omitted to avoid heterogeneous coupling to multiple cysteines), methionine (omitted to avoid the risk of inactivating oxidation), and lysine and arginine (omitted to avoid the addition of trypsin cleavage sites). Because valine and another small hydrophobic amino acid, isoleucine, appeared in one-quarter of the nanoCLAMPs from previous screens, position 818 was kept constant at the wild-type amino acid valine. The resulting library consisted of 10 of the nC-B nanoCLAMPs. 10 It contained more than 10 variants.
[0200] After the third round of panning of this library, we randomly selected 96 clones and screened them for target binding by semELISA, which confirmed 93 positives. Of these, 40 were sequenced, identifying 18 unique nanoCLAMP-SUMO conjugates. These conjugates were subcloned into bacterial expression vectors, expressed, purified by immobilized metal affinity chromatography (IMAC), and confirmed to be >90% pure as estimated by SDS-PAGE (data not shown). We then evaluated the ability of purified nanoCLAMP to function as an affinity capture agent in a medium-throughput, small-scale depletion assay. In this assay, nanoCLAMP was conjugated to cross-linked agarose under denaturing conditions, refolded on the resin, and incubated with SUMO to detect A. 280 The amount of unbound SUMO was measured by
[0201] The purified nanoCLAMP was subsequently analyzed for monodispersity by size exclusion chromatography and for T by DSF. m Of the 18 clones screened, 7 (38%) were greater than 90% monomeric. Of these, 5 had melting temperatures greater than 73°C and 4 had melting temperatures greater than 99°C (Table 4). Initial results using the SUMO assay suggest that the nC-B constant region generally favors the isolation of clones with high monodispersity and thermostability.
[0202] [Table 4]
[0203] Characterization of binding affinity, thermal stability, and protease resistance of nC-B class nanoCLAMP capture agents. We selected three nC-B nanoCLAMPs (P2808, P2809, and P2811) for further characterization. These were chosen to provide a diverse sample of binding loops. For the three clones, loops 817-820 and loops 931-935 did not share any obvious similarity. No identity was found at any position in these loops, except for V818, whose identity was fixed in the library. For loop 838-844, clones P2808 and P2809 are identical at five of the seven positions, whereas clone 2811 shows no identity with either.
[0204] All three nanoCLAMPs were produced in E. coli at yields exceeding 150 mg / liter (shaken E. coli culture) and used in biophysical characterization experiments.
[0205] We first investigated the quaternary structure of nanoCLAMPs by size-exclusion chromatography to ensure that subsequent results could be interpreted without confounding avidity effects from higher-order multimers or aggregates. All three clones migrated as monodisperse monomers (Figure 6). To rank these nanoCLAMPs by affinity for SUMO, we measured their dissociation constants using biolayer interferometry, which ranged from 5 to 80 nM (Table 5). Next, we measured the melting points of nanoCLAMPs using DSF. P2808 had an apparent T m P2809 and P811 had flat DSF curves, with no apparent melting transition between 25 and 99 °C (Figure 6). This observation confirmed the T mThis suggests that the binding activity exceeds the quantifiable range for this assay. We corroborated this observation with a functional binding assay to assess kinetic thermostability. In this assay, we incubated samples of each nanoCLAMP at various temperatures, cooled the solution, centrifuged, and measured the binding activity remaining in the supernatant by biolayer interferometry. In this test, we measured the binding activity of T 50 T was defined as the temperature of a 5-minute heat challenge at which 50% of the binding activity was irreversibly lost. 50 The T ranged from approximately 85°C for clone P2808 to over 100°C for clones P2809 and P2811. This rank order and values are consistent with the T measured by DSF. m (Fig. 11). Both T exceed 100°C. m Clones P2809 and P2811, which have the α- and β-blocking groups, maintained greater than 90% activity after 5 minutes of incubation at boiling point. The maintenance of binding activity indicates that the nanoCLAMPs remained in solution after heat treatment and did not irreversibly aggregate or precipitate. Taken together with the DSF data, the kinetic thermal stability measurements suggest that the two nanoCLAMPs can remain folded up to, and in some cases even beyond, 99°C.
[0206] [Table 5]
[0207] Next, we characterized the trypsin resistance of the three clones. Clones P2808 and P2809 were highly resistant to digestion with 0.1 mg / ml trypsin, whereas clone P2811 was less resistant (Figure 5F). Figures 5B-5D show the time course of trypsin digests comparing nanoCLAMPs with various combinations of constant regions and variable loops to understand the contribution of each component to trypsin resistance. We tested the original clone SMT3-A1 (nC-A constant region and original variable loop), P2788 (nC-B constant region and original variable loop from SMT3-A1), and P2808 (nC-B constant region and newly isolated variable loop). Both P2788 and P2808 survived for over 20 hours in 0.1 mg / ml trypsin, compared to approximately 4 hours for the original SMT3-A1 clone. 1 / 2 These results, combined with those of P2809 and P2811, indicate that trypsin resistance depends on the sequences of both the variable loop and the constant region. The observation of trypsin resistance for three of the four isolated clones with diverse loop sequences and identical constant regions indicates that, at least in the first test case, the constant region of nC-B can be used generically to isolate trypsin-resistant clones.
[0208] Measurement of affinity chromatography performance parameters using nC-B class nanoCLAMP trapping agents. For brevity, we refer to affinity chromatography resins made with the nC-B class of nanoCLAMP capture agents as "nC-B affinity resins." As a test case for the utility of nC-B affinity resins, we used P2808 as a capture agent for more detailed studies. To generate the affinity resin, the P2808 protein was expressed, purified by IMAC under denaturing conditions, conjugated to sulfhydryl-reactive 6% cross-linked agarose resin, and subsequently refolded by rinsing with buffered saline. We prepared a test mixture of crude E. coli lysate spiked with a SUMO-GFP fusion protein to optimize purification conditions and evaluate the selectivity and binding capacity of the resin.
[0209] In pilot experiments, we found that polyol elution at neutral pH was still possible, similar to the original nanoCLAMP (Suderman et al., supra), although qualitatively at a lower rate and yield (data not shown). As an alternative, we tested elution of the target protein using molar concentrations of imidazole, which has been used successfully to disrupt protein-peptide and antibody-protein A interactions.
[0210] A buffer with 3 M imidazole at pH 8 rapidly and completely eluted SUMO from the P2808 resin. As shown in Figure 12, after a 1-hour incubation with excess target protein in spiked E. coli lysate, the bound target was washed and eluted with greater than 90% purity, as estimated by densitometric analysis of Coomassie-stained SDS-PAGE. The static binding capacity under these conditions was 11.5 mg / ml (277 nmol / ml) of resin. In subsequent experiments, we found that imidazole elution worked consistently with three different AC resins targeting SUMO, mCherry, and GFP, along with P2808, P2809, and P2811, as well as the nC-A class of capture agents. Observations suggest that imidazole elution is a general property of both classes of nanoCLAMPs.
[0211] Using the established elution conditions, we next measured the dynamic binding capacity of the P2808 resin using a 0.6 ml packed column (ID 0.5 cm x 3.06 cm) at a constant flow rate of 0.5 ml / min. We utilized the fluorescence of the SUMO-GFP target to measure the QB 10 , or the capacity at which the fluorescence of the eluate equals 10% of the fluorescence of the load, was determined (see Materials and Methods for calculations). The dynamic binding capacity of the P2808 resin under these conditions was 10 mg / ml resin (240 nmol / ml resin) (Figure 7). This capacity represents an approximately 70% increase over our original SMT3-A1 resin.
[0212] We next examined the efficiency of the SUMO affinity resin P2808 under flow conditions. Two versions were tested using a 0.6 ml packed column and a flow rate of 0.5 ml / min (linear flow rate = 153 cm / hr).
[0213] In the capacity-overload model, SUMO was spiked into E. coli lysate at a low concentration (SUMO-GFP = 0.76% by weight of total protein) and loaded at 42% of the column's dynamic capacity (36% of its static binding capacity). Purity was greater than 90%, with a yield of 90%, as assessed by densitometry of Coomassie-stained SDS-PAGE gels (Figure 8A).
[0214] In the capacity-limited version, SUMO-GFP was spiked into E. coli lysate at a high concentration (SUMO-GFP = 5.7% by weight of total protein) and loaded at 133% of the column's dynamic binding capacity (116% of its static binding capacity) (Figure 8B). Purity and yield were comparable to the overcapacity version, both estimated to be greater than 90%. Performance metrics for both purifications are summarized in Table 6.
[0215] [Table 6]
[0216] Compatibility of nC-B affinity resin with repeated cycles of NaOH washing. For large-scale production of industrial proteins or biologics, column reuse and cleaning-in-place reduce production costs and maintain consistent performance. Sodium hydroxide solution is commonly used in cleaning-in-place procedures, and therefore we tested the general compatibility of the new nanoCLAMP resin with sodium hydroxide treatment. We tested three resins made with the nC-A nanoCLAMP and three resins made with the nC-B nanoCLAMP. Each cycle consisted of loading SUMO-GFP target spiked into E. coli lysate, a wash step, elution with 3 M imidazole pH 8, a 10-minute contact time with 0.1 M NaOH, followed by a 5-minute re-equilibration step. The eluate was collected and analyzed for purity by SDS-PAGE, and the purified target was quantified by fluorescence spectroscopy (Figures 9 and 13). In the first few cycles, we were surprised to observe a 5- to 20-% improvement in binding capacity when using the nC-B resin. The increased binding may be due to the removal of inhibitors by NaOH. For all three nC-B resins, the binding capacity reached a plateau over 20 cycles and remained at or above 100% of the starting capacity. In contrast, we observed a steady decrease in binding capacity of 25% to 50% with all three nC-A resins. Collectively, this data indicates that the nC-B class of nanoCLAMPs can generally function as capture agents for resins capable of single-step affinity purification of targets to homogeneity. These resins are also compatible with clean-in-place protocols using 0.1 M sodium hydroxide for more than 20 cycles without loss of binding capacity or specificity. Furthermore, in practice, clean-in-place cycles are typically not performed after each run; therefore, the expected lifetime is likely to exceed 100 cycles, assuming disinfection every fifth run.
[0217] Compatibility of nC-B affinity resin with repeated cycles of low pH elution. Because elution with 3 M imidazole may not be optimal for some applications, we also tested elution with pH 2.5 citrate buffer. In these experiments, the resin was washed with NaOH for 1 min between cycles, and then re-equilibrated with buffered saline. NanoCLAMP maintained 100% of its binding capacity and specificity over more than 20 cycles of loading, elution, and regeneration (Figures 14A and 14B).
[0218] Compatibility of nC-B affinity resin with organic solvents and autoclaving. We next examined the ability of the nC-B resin and the original SMT3-A1 resin to withstand more extreme conditions. First, we examined the ability of the resin to restore its selective binding capacity after exposure to 100% DMF. We incubated nanoCLAMP resin in 100% DMF for 2 hours, re-equilibrated with buffered saline, and then measured its binding capacity. Of the four resins tested (original SMT3-A1 resin and P2808, P2809, and P2811 resins), all retained greater than 85% of their binding capacity after treatment with DMF (Figure 10A) and maintained their apparent selectivity as assessed qualitatively by SDS-PAGE (Figure 10B).
[0219] Because the nC-B resin was robust across the wide range of conditions tested so far, we sought to determine whether the resin could also retain binding and specificity after autoclaving. We autoclaved the resin for 105 minutes on a liquid-vapor cycle, including 30 minutes of exposure to 120 °C and 20 psi, followed by re-equilibration at room temperature and subsequent measurement of binding capacity. Resins made with the nC-A nanoCLAMP (SMT3-A1) did not bind any detectable target protein after autoclaving. In contrast, resins made with the nC-B nanoCLAMPs (P2808, P2809, and P2811) retained 25% to 45% of their binding capacity with comparable specificity (Figures 10C and 10D). The SMT3-A1 eluate was not analyzed in Figure 10D because there was not enough protein in the autoclaved sample to prepare standardized aliquots for comparison with the control. To our knowledge, the nC-B nanoCLAMP-based affinity chromatography resin is the first protein-based affinity resin shown to retain significant binding capacity and specificity after autoclaving.
[0220] Consideration We report a protein engineering campaign that yielded an improved class of nanoCLAMPs suitable as capture agents for process-scale affinity chromatography. Similar to Protein A resins and antibody-based immunoaffinity resins, resins engineered with the novel nC-B class of nanoCLAMPs support single-step purification from complex mixtures with high yields and fold-through purification. Similar to Protein A resins, but unlike antibody-based resins, nC-B resins are compatible with NaOH wash-in-place, can be produced by bacterial expression of the capture reagent, and lack cysteine residues. Unlike Protein A or antibody-based resins, nC-B resins are unique in that they have been shown to be resistant to boiling points, trypsin, and organic solvents. Key performance parameters of nC-B resins and supporting results are summarized in Table 7.
[0221] [Table 7]
[0222] This study provides a test case supporting the potential of nC-B nanoCLAMP resin to extend Protein A-like performance to a broad range of proteins beyond antibodies. The efficiency, low manufacturing cost, and reusability of nC-B resin have the potential to reduce the overall cost of manufacturing for process-scale purification. Furthermore, nanoCLAMP's general compatibility with high temperatures, organic solvents, and extreme pH may enable industrial applications requiring extreme conditions.
[0223] The improved stability and performance of nC-B nanoCLAMPs also support their use in applications beyond immunoaffinity chromatography. NanoCLAMPs have been successfully used in bioelectrical and electrochemical sensors. Conditions at biosensor surfaces represent a challenging environment that may be enabled by improving the stability of nC-B nanoCLAMPs. For example, conjugation of nanoCLAMPs to surfaces in DMF allows for the use of reaction conditions that are compatible with reagents that have low solubility in aqueous buffers.
[0224] The performance characteristics demonstrated for the exemplary nanoCLAMP in this study support further exploration of the potential of nanoCLAMP to replace other capture agents in a wide range of applications, particularly in procedures requiring stringent, selective, or reversible binding and exposure to extreme conditions.
[0225] Materials and Methods Cloning of SMT3-A1 mutants The nanoCLAMP SMT3-A1-containing plasmid pET(SMT3-A1) was mutated by inverse PCR (Ochman et al., Genetics 120:621-623, 1988) by amplifying the plasmid with forward and reverse primers containing the desired mutation with 15-bp overlapping 5' ends, purifying the amplicon, cloning both ends together using In-Fusion (In Fusion HD Cloning Kit, Takara), and transforming chemically competent NEc1 E. coli (a BL21(DE3) derivative carrying slyDD (His151-His196) from Nectagen, Inc.). The plasmid was purified using a Qiagen miniprep kit (Qiagen), and the mutation was verified by sequencing the purified plasmid using Sanger sequencing (Genewiz). A glycerol stock of the plasmid in NEc1 cells was prepared for inoculation of expression cultures. The construct for conjugation to Sulfolink resin (Thermo) encoded nanoCLAMP with an N-terminal 6-His tag and a 13-amino acid C-terminal GS-linker followed by Cys. The construct for expressing nanoCLAMP for biophysical characterization lacked the GS-linker and C-terminal Cys to avoid disulfide-induced dimerization issues.
[0226] pET(SMT3-A1) sequence (SEQ ID NO: 38) [ka] [ka] [ka]
[0227] Expression and purification of nanoCLAMP under denaturing conditions for conjugation to affinity chromatography resins (1 L scale) A glycerol stock of NEc1 cells harboring the nanoCLAMP expression vector (described above) was used to inoculate a 3 ml starter culture of 2xYT / 2% glucose (Glu) / 100 mg / ml carbenicillin (CB) and grown overnight at 37°C and 250 rpm. The overnight culture was diluted 1:100 into 300 ml of Novagen Overnight Express Instant TB Medium / 1% glycerol / CB and incubated for 24 hours at 30°C and 250 rpm. Cells were pelleted at 10 k × g for 10 minutes at 4°C and lysed and homogenized using a Polytron in 30 ml of 100 mM NaH2PO4, 10 mM Tris, 6 M GuHCl (QAB) pH 8.5, +1 mM TCEP (QAB-TCEP, pH 8.5). Insoluble material was pelleted at 15 k × g for 20 min at 15°C, and the clarified supernatant was applied to Ni Sepharose 6 Fast Flow (Cytivia) and incubated with rotation for 1 h to overnight. The beads were transferred to a column and washed with 3 CV of QAB-TCEP, pH 8.5, followed by 3 CV of QAB, pH 8.5. Protein was eluted with QAB, pH 8.5 + 250 mM imidazole and quantified by A280. Purity of the eluted protein was determined by SDS-PAGE on a 12% NuPAGE Bis-Tris gel and Coomassie staining with Gel-Code Blue (after removal of GuHCl by cold ethanol precipitation). Yields for nanoCLAMP were typically 150–300 mg / L culture, and purity was typically greater than 90%.
[0228] Purified modified nanoCLAMP in QAB, pH 8.5, was reduced with 2 mM TCEP for use after storage and conjugated to Sulfolink-crosslinked 6% beaded agarose (Thermo). Briefly, the resin was equilibrated with QAB, pH 8.5, plus 5 mM EDTA, and transferred to a column. NanoCLAMP was adjusted to 8 mg / ml in two volumes of Sulfolink resin and then incubated with the resin for 30 minutes at room temperature with rotation. The resin was allowed to settle for 15 minutes, and the column was drained onto the top of the resin bed. The column was washed with QAB, pH 8.5, then incubated with 50 mM L-Cys in QAB, pH 8.5, and quenched for 15 minutes with rotation. The column was allowed to settle, drained, and washed again with 6 M GuHCl, 20 mM Tris, (QCB), pH 8. Finally, the nanoCLAMP was refolded onto the resin by rinsing with 6 CV of 20 mM MOPS, 150 mM NaCl (MBS), pH 6.5 + 1 mM CaCl2.
[0229] Measurement of the static binding capacity of nanoCLAMP resin The spiked E. coli lysate was measured at an OD 600 Cells were prepared by pelleting the equivalent of 8 cultures, discarding the supernatant, and lysing the cells with 4 ml of BPER per gram of pellet at room temperature for 20 minutes while rotating. Insoluble material was removed by centrifugation, and the clarified lysate was adjusted to a total protein concentration of approximately 1.87 mg / ml in 20% BPER in PBS, pH 7.4. The target protein, SUMO-GFP (described above), was spiked into the lysate using a highly concentrated stock to a final concentration of 0.025–0.2 mg / ml depending on the application, so that the total protein concentration remained unchanged. The spiked lysate was then incubated with 10 ml of nanoCLAMP resin (loading volume) in a total volume of 1.4 ml at 4°C for 1 hour while rotating. The resin was pelleted by centrifugation and transferred to a small column. The resin was washed four times with 400 ml of PBS, pH 7.4, and then eluted with 3 M imidazole, pH 8. The eluate was buffer exchanged twice on a Zeba column (7 kD MWCO, Thermo) and analyzed using an iD5 plate reader.280 or quantified by fluorescence.
[0230] Expression and purification of nanoCLAMP for biophysical characterization A glycerol stock of NEc1 cells harboring the nanoCLAMP expression vector (described above) was used to inoculate a 3 ml starter culture of 2xYT / 2% glucose (Glu) / 100 mg / ml carbenicillin (CB) and grown overnight at 37°C and 250 rpm. The overnight culture was diluted 1:100 into 35 ml of Novagen Overnight Express Instant TB Medium / 1% glycerol / CB and incubated for 24 hours at 30°C and 250 rpm. Cells were pelleted and lysed with QCB, pH 8, and insoluble material was removed by centrifugation at 15 k×g for 20 minutes at 15°C. The clarified lysate was incubated with Ni Sepharose 6 Fast Flow (Cytivia) for >1 hour at room temperature with rotation before being transferred to a 2 ml column. The column was washed with 6 × 1 ml of QCB, pH 8, followed by refolding with 11 ml of 20 mM MOPS, 150 mM NaCl (MBS), 1 mM CaCl2, pH 8. NanoCLAMP was eluted with MBS, 1 mM CaCl2, 250 mM imidazole, pH 8, buffer exchanged to remove imidazole using a Zeba 7 MWCO desalting column, and standardized to 1 mg / ml in MBS, 1 mM CaCl2, pH 6.5.
[0231] Expression and purification of target proteins for panning and affinity chromatography To prepare target proteins for panning library NL-26, we prepared a biotinylated yeast SUMO construct (B-SUMO, P1068) in a pET expression vector and transformed it into BL21(DE3) E. coli harboring constitutively expressed biotin ligase BirA. An overnight starter culture was diluted 1:100 into 500 ml of Novagen Overnight Express Instant TB Medium / 1% glycerol / CB / CAM containing 5 mM biotin and incubated for 24 hours at 30°C and 250 rpm. After induction, cells were pelleted and the medium was discarded. The pellet was frozen at -80°C, thawed on ice, resuspended in MBS, pH 7.4, plus Pierce Protease Inhibitor Tablet Mini at 5 ml / g pellet, and sonicated on ice for 10 minutes at a 50% duty cycle. Biotin was added to 100 μM, and the lysed cells were incubated at 37°C for 30 minutes at 250 rpm to drive biotinylation to completion. The lysate was clarified by centrifugation at 30 k×g for 20 minutes at 4°C, and the supernatant was transferred to a 2.25 ml column packed with SMT3-A1 resin (Nectagen, Inc.) at 1 ml / min. The resin was washed with 25 ml of MBS, pH 7.4, and the protein was eluted with polyol elution buffer (PEB): 10 mM Tris, 1 mM EDTA, 0.75 M ammonium sulfate, 40% propylene glycol, pH 7.9. The protein was desalted 2x into 50 mM Tris, pH 8, and stored as a 50% glycerol stock at -20°C.
[0232] B-SUMO sequence (P1068) (SEQ ID NO: 39) [ka]
[0233] To facilitate quantification of the affinity chromatography target protein in the eluate, we prepared a SUMO-GFP fusion (P10126RDG-1) that we could track by fluorescence spectroscopy. Briefly, the pET expression construct was transformed into NEc1 E. coli (Nectagen, Inc.), and an overnight culture was diluted 1:100 into 1 L of Novagen Overnight Express Instant TB Medium / 1% glycerol / CB and grown at 30°C and 250 rpm for 24 hours. The cells were pelleted, the medium removed, and lysed with BPER with Universal Nuclease (Thermo). Insoluble material was removed by centrifugation, and the clarified supernatant was loaded onto a 5 ml Ni Sepharose 6 Fast Flow column (Cytivia) at 1.5 ml / min. The resin was washed with 100 ml of 50 mM NaH2PO4, 300 mM NaCl, 20 mM imidazole, pH 8, followed by elution with the same buffer but with 250 mM imidazole. Protein purity of both B-SUMO and SUMO-GFP was assessed by SDS-PAGE under reducing conditions using 12% NuPAGE, BisTris, MES running buffer and stained with GelCode Blue (Thermo).
[0234] SUMO-GFP fusion (P10126RDG-1) (SEQ ID NO: 40) [ka]
[0235] Analysis of monodispersity by size exclusion chromatography Purified nanoCLAMP was diluted to a final concentration of 0.18 mg / ml in MBS, 1 mM CaCl2, pH 6.5, centrifuged at 20 k × g for 2 minutes at 4°C, and the supernatant was transferred to a clean tube. The sample was loaded into a 125 μl sample loop and injected at a flow rate of 0.65 ml / min onto a Superdex 75 10 / 300GL column (GE Healthcare Life Sciences, Pittsburgh, PA) equilibrated in MBS, 1 mM CaCl2, pH 6.5. The column was calibrated using Bio-Rad Gel Filtration Standard according to the manufacturer's instructions.
[0236] Melting point determination by differential scanning fluorimetry. The melting point of purified nanoCLAMP was determined using the GloMelt Thermal Shift Protein Stability Kit (Biotium) according to the manufacturer's instructions. Briefly, purified nanoCLAMP was adjusted to 1 mg / ml in MBS, 1 mM CaCl2, pH 6.5, diluted in half with 2x GloMelt (Biotium), dispensed into a 386-well plate, and sealed with optical film. The plate was then heated in a Quantstudio 5 qPCR instrument using the SYBR Green reporter without a passive reference. The heating profile was 25°C for 2 minutes, a 0.05°C / s ramp to 99°C, and 99°C for 2 minutes. m is defined as the inflection point of the unfolding curve.
[0237] Determination of protease stability by digestion with trypsin and chymotrypsin Digestion was performed by incubating 0.25 mg / ml nanoCLAMP in a 20 μl reaction containing 0.1 mg / ml trypsin (Roche Cat. 11418475001) or chymotrypsin (Roche Cat. 11418467001) diluted with 1 mM HCl to a final HCl concentration of 0.1 mM in the reaction. CaCl2 was added to the reaction to 10 mM. The protein and remaining diluent buffer was MBS, pH 6.5. Reactions were incubated at 37°C for the indicated times, stopped by adding 2 ml of 10x Protease Arrest (G-Biosciences), and analyzed by SDS-PAGE (12% NuPAGE, Bis Tris in MES buffer) in SDS sample buffer with reducing agent and stained with GelCode Blue (Thermo). Densitometry was performed using GelAnalyzer software to measure the relative staining intensity of full-length bands.
[0238] Construction of phage display library NL-26 for the nC-B class The pCombX phagemid template p2799 (Table 8) contained the N- and C-terminal constant regions of the nC-B class, separated by a stuffer region containing HindIII and SpeI cleavage sites. This template was digested with HindIII and SpeI and gel-purified, and the plasmid region was amplified using degenerate primers 1957T R and 1960T F, which added the N- and C-terminal portions of nC-B and randomized loops L1 and L8, respectively. Primers listed with a T indicate that they were degenerate primers constructed using a phosphoramidite trimer mixture (Glen Research) of oligos (IDT) containing all amino acids except Cys, Met, Lys, and Arg. A short internal region of p2788 was amplified using primers 1958T F and 1959 R, which added and randomized loop L2. PCR was performed using ClonAmp HiFi PCR Mix according to the manufacturer's instructions (Takara Bio, Mountain View, CA). The reaction cycle consisted of 10 seconds at 98°C, 10 seconds at 65°C, and 30 seconds at 72°C, repeated 30 times. These two amplicons containing overlapping ends were gel-purified and cloned together by Gibson Assembly (described below) to generate the nC-B construct with three variable loops: loop L1 (three residues -817, 819, 820), loop L2 (seven residues, 838-844), and loop L8 (five residues, 931-935), for a total of 15 variable residues in the three loops.
[0239] [Table 8]
[0240] To clone the library components, 10 μg of the large amplicon and 7.86 μg of the short amplicon were combined in a 2 ml reaction containing 1000 μl of Gibson Assembly Master Mix (2x) (NEB), incubated at 50°C for 30 minutes, and then placed on ice. The ligated DNA was then purified and concentrated in a Nucleospin Gel and PCR Cleanup Kit (Machery Nagel) and eluted in 45 μl of EB. The DNA was then desalted with ddH2O on a VSWP 0.025 μm membrane (EMD Millipore) for 40 minutes with 20-minute water changes. The desalted DNA was then adjusted to 100 ng / μl with ddH2O and used to electroporate electrocompetent TG1 cells (Lucigen). Approximately 50 μl of DNA was added to 1.25 ml of ice-cold TG1 cells and mixed by pipetting up and down four times on ice, after which 25 μl aliquots were transferred to 50 electroporation cuvettes (with a 1 mm gap) on ice. Cells were electroporated, immediately quenched with 975 μl of recovery medium (Lucigen), pooled, and incubated for 1 hour at 37°C and 250 rpm. To titrate the library, 10 μl of the recovery culture was serially diluted in 2xYT, and 10 μl of each dilution was spotted onto 2xYT / glu / carb and incubated overnight at 30°C. The remaining library was spread onto 3 L of 2xYT / glu / carb and amplified overnight at 30°C and 250 rpm. The next day, the library was pelleted at 10 k×g for 10 minutes at 4°C, and the medium was discarded. The pellet was analyzed at OD in 2xYT / 2% glucose / 18% glycerol. 600 The solution was resuspended to a volume of 75 ml, aliquoted, and stored at -80°C.
[0241] Panning of nanoCLAMP library NL-26 (nC-B library) For the first round of panning, 2.7 L of 2xYT medium with 2% glucose and 100 mg / ml carbenicillin (2xYT / Glu / CB) was cultured at an OD 600Dilute the glycerol stock of the NL-26 library to an OD of approximately 0.1. 600 = 75) 3.6 ml was inoculated and the OD 600 The cells were grown at 37°C and 250 rpm until an MOI of 0.52 was reached. The library was infected by adding helper phage VCSM13 (Stratagene, Cat#200251) to the 750 ml culture at an MOI of 20 phage / cell and incubating at 37°C, 100 rpm for 30 minutes, followed by an additional 30 minutes at 250 rpm. The cells were pelleted at 7500 x g for 10 minutes, and the medium was discarded. The cells were resuspended in 1.2 L of 2xYT / CB, 70 μg / ml kanamycin (KAN) and incubated at 30°C, 250 rpm for 15 hours. The cells were combined, and 100 ml was centrifuged at 10k x g for 10 minutes. The phage-containing supernatant was transferred to a clean tube and precipitated by adding 37.5 ml of 5xPEG / NaCl (20% polyethylene glycol 6000 / 2.5 M NaCl) and incubated on ice for 25 minutes. Phage were pelleted at 13 k×g for 25 min and the supernatant was discarded. Phage were resuspended in 10 ml of 20 mM NaH2PO4, 150 mM NaCl, pH 7.4 (PBS) and then centrifuged at 15 k×g for 15 min to remove insoluble material. Phage were precipitated a second time by adding ¼ volume of 5×PEG / NaCl, incubated on ice for 5 min, and pelleted at 13 k×g for 10 min at 4°C. Phage pellets were resuspended in 3 ml of PBS and quantified by absorbance at 268 nm (5×10 12 For phage / ml of solution, A 268 =1).
[0242] Two sets of 100 μl of Dynabeads MyOne Streptavidin T1 (ThermoFisher Scientific) magnetic bead slurry were washed with 2 × 1 ml of PBS-T (PBS with 0.05% Tween 20), and after applying a magnet between washes to remove the supernatant, they were blocked by rotation in 1 ml of a 2% solution of milk powder in PBS with 0.05% Tween 20 (2% M-PBS-T) for 1 h at room temperature. To preclear phage from beads alone, 1 ml of phage was added to 2 × 10 13 The beads were prepared at a concentration of 10 phage / ml. The block was removed from the first set of beads, and the phage was added to the beads and incubated with rotation for 1 hour. A magnet was applied to remove the precleared phage and transfer it to a clean tube. This step was repeated twice to ensure that phage-bound beads were not carried over to the next step. The biotinylated target (B-SUMO) was added to the precleared phage to a final concentration of 100 nM and incubated with rotation for 1 hour. The block was removed from the second set of beads, and the phage / B-SUMO mixture was added to the beads to precipitate the biotinylated target and bound phage. The beads were washed eight times with 1 ml of PBS-T, vortexing between each wash and applying a magnet. The washed beads were eluted with 800 μl of 0.1 M glycine, pH 2.0, for 10 minutes with rotation, and the magnet was applied. The eluate was neutralized by transferring it to 72 μl of 2 M Tris base. The neutralized phage were then cultured at OD 600 This was added to 9 ml of XL1-blue E. coli grown to a pH of 0.435 and placed on ice. The cells were infected at 37°C for 45 minutes at 175 rpm, then spread onto 100 ml of 2xYT / Glu / CB and incubated overnight at 30°C at 250 rpm.
[0243] The overnight culture was analyzed at OD 600 The cells were centrifuged at 10 k × g for 10 min, followed by OD in 2 × YT / 18% glycerol. 600To prepare phage for the next round of panning, cells were resuspended in 5 ml of 2xYT / Glu / CB to a final concentration of 75 OD. 600 Inoculate 5 μl of a glycerol stock at OD 600 The cells were incubated at 37°C and 250 rpm until the β-dimer ratio reached 0.5. Cells were superinfected at a 20:1 ratio of phage:cells, mixed thoroughly, and incubated at 37°C for 30 minutes at 150 rpm, followed by 30 minutes at 250 rpm. The cells were pelleted at 5500 x g for 10 minutes, the glucose-containing medium was discarded, and the cells were resuspended in 10 ml of 2xYT / CB / KAN and incubated overnight at 30°C and 250 rpm.
[0244] Overnight phage preps were processed as described above. Phages were then diluted in 2% M-PBS-T. 268 Panning and preclearance were continued as described, except that the phage were prepared at pH 7.0 = 0.8 and the biotinylated target concentration was reduced 10-fold per round in the second and third rounds. Post-phage capture washes were also increased to 12 washes in the third round. In round 2, neutravidin-coated magnetic beads (Spherotech) were used instead of streptavidin beads to reduce enrichment for streptavidin binders.
[0245] Qualitative semELISA of individual clones after panning. At the end of the final panning round, individual colonies were plated onto 2xYT / Glu / CB agar plates after recovery of XL1-blue cells infected with eluted phages at 37°C for 45 minutes at 150 rpm. The next day, 95 colonies were inoculated into 400 μl of 2xYT / Glu / CB in a 96-deep-well culture plate and grown overnight at 37°C and 300 rpm to generate master plates, to which glycerol was added to 18% for storage at -80°C. To prepare induction plates for ELISA, 5 μl of each master plate culture was inoculated into 400 μl of fresh 2xYT / 0.1% glucose / CB medium and incubated at 37°C and 300 rpm for 2.75 hours. IPTG was then added to 0.5 mM, and the plates were incubated overnight at 30°C with shaking at 300 rpm. Because the phagemid contains an amber stop codon, even though XL1-blue is a suppressor strain, some nanoCLAMP protein without the pIII domain is produced, resulting in periplasmic localization of some nanoCLAMP, a small percentage of which is ultimately secreted into the medium. The medium can then be used directly for ELISA assays (soluble expression-based monoclonal enzyme-linked immunosorbent assays: semELISA). After overnight induction, the plates were centrifuged at 1200 × g for 10 minutes to pellet the cells. Streptavidin-coated microtiter plates (ThermoFisher) were rinsed three times with 200 μl of PBS and coated with 2 μg / ml of biotinylated target protein at 100 μl / well and incubated for 1 hour. For blank controls, the plates were incubated with 100 μl / well of PBS. The coating solution was removed, and the plates were blocked with 2% M-PBS-T. The block was removed, and 50 μl of 4% M-PBS-T was added to each well. At this point, 50 μl of each induced plate supernatant was transferred to the blank and protein-coated wells, mixed by pipetting 10 times, and incubated for 1 hour. The plates were washed four times with 200 μl of PBS-T, removing the plates between washes and tapping them on a paper towel.After washing, 75 μl of 1:2000 diluted anti-FLAG-HRP (Sigma A8592) in 4% M-PBS-T was added to each well and incubated for 1 hour. The anti-FLAG-HRP was discarded, and the plate was washed as before. The plate was developed by adding 75 μl of TMB Ultra substrate (ThermoFisher) and analyzed for positive signals compared to controls. Positive clones were then expanded from the master plate by inoculating 3 μl of the glycerol stock into 1 ml of 2xYT / 2% glucose / 100 μg / ml CB and incubated at 37°C and 250 rpm for at least 6 hours. The cells were then pelleted and the medium discarded. Plasmid DNA was prepared from the pellet using a Qiaprep Spin Miniprep Kit, and the sequence was determined by Sanger sequencing at Genewiz (South Plainfield, NJ). The nanoCLAMP insert from a unique positive clone was amplified and cloned into a pET expression vector as described above.
[0246] NanoCLAMP biolayer interferometry Kinetic analysis of the interaction between nanoCLAMP and biotinylated SUMO was performed with OctetRed96 using a SAX streptavidin-coated sensor chip. The chip was first transferred to buffer (MBS, 1 mM CaCl2, pH 6.5 + 1% BSA) for 300 seconds, then to 2 mg / ml B-SUMO in buffer for 180 seconds, followed by buffer for 300 seconds, then to at least four dilutions of nanoCLAMP in buffer (binding) for 200 seconds, followed by buffer for 500 seconds (dissociation). Cells were constantly vortexed at 1000 rpm at room temperature. Kinetics were fitted to a 1:1 model and global fit analysis was used to estimate the K. d was calculated.
[0247] Dynamic bonding ability of P2808 resin and SMT3-A1 resin A 0.6 ml packed volume of P2808 or SMT3-A1 resin was packed into a Tricorn 5 / 50 column (ID 5 mm x height 3.06 cm) and equilibrated with 5 CV at 0.5 ml / min in 20 mM NaH2PO4, 150 mM NaCl, pH 7.4 (PBS). A load of Sumo-GFP fusion protein (MW = 41,559 g / mol) diluted to a concentration of 0.2 mg / ml in PBS was delivered to the system in column bypass mode. Fluorescence of the eluate was measured at Ex / Em 485 / 535 nm to determine the total loaded fluorescence. The delay volume, V delay was measured for the 0.5 ml configuration. The load was then applied to the column and the measured volume V x where V x where V is the fluorescence of the eluate = the volume where the fluorescence is 10% of the total load. The dynamic binding capacity was then calculated in mg / (ml resin) as follows: DBC = (V x -V delay ) × c / (volume of resin).
[0248] Purification of Sumo-GFP from spiked E. coli lysate by affinity chromatography using P2808 resin Clarified E. coli lysates were prepared by dissolving pellets of NEc1 E. coli (a derivative of BL21(DE3) in which the C-terminal region of SlyD had been recombinantly knocked out; Nectagen, Inc.) with BPER (Thermo) and removing insoluble material by centrifugation at 15 k × g for 20 min at 4°C. The clarified supernatant was diluted to a total protein concentration of approximately 3.3 mg / ml with PBS, pH 7.4, such that BPER reagent was present at 20% vol / vol. The target protein, SUMO-GFP (MW = 41,559 g / mol), was spiked in to a final concentration of either 0.2 mg / ml or 0.025 mg / ml. The spiked lysates were loaded onto the column at 0.5 ml / min for the indicated times, washed with 20 CV of PBS, pH 7.4, and then eluted with 3 M imidazole, pH 8. Fractions containing the eluted target were pooled and desalted twice on a Zeba 7 MWCO column, and the protein was purified by A 280Imidazole removal was verified by examining the A280 of the elution buffer alone after 2x desalting. Spiked lysates, early wash fractions, and pooled eluates (after buffer exchange) were analyzed by NuPAGE SDS PAGE under reducing conditions, 12% Bis-Tris in MES running buffer, and stained with Gel Code Blue (Thermo).
[0249] Repeated AC purification cycles including cleaning in place (CIP) of the resin using 0.1M NaOH Repeated affinity chromatography purifications were performed on an FPLC using a small 50 μl (packed) column with a running buffer of 20 mM MOPS, 150 mM NaCl, 1 mM CaCl, pH 7.2. The load consisted of Sumo-GFP spiked into 0.1 mg / ml of clarified E. coli lysate (described above in Purification of Sumo-GFP from Spiked E. coli Lysate). The cycle consisted of a 2 ml equilibration in running buffer at 1 ml / min, a 0.5 ml load of spiked lysate at 0.5 ml / min, a 3 ml wash with running buffer at 0.5 ml / min, a 2 ml elution with 3 M imidazole, pH 8 (collection) at 0.5 ml / min, a 0.5 ml wash with running buffer at 0.5 ml / min, a 1.5 ml wash with NaOH at 1 ml / min followed by a 2 ml stationary wash at 0.2 ml / min (total contact time 10 min), and a final refolding step with 5 ml of running buffer at 1 ml / min. Target concentrations in the eluates were measured in duplicate by fluorescence spectroscopy at Ex / Em 485 / 535 nm on an iD5 plate reader (Molecular Dynamics). The eluates were analyzed by SDS-PAGE using NuPAGE gels as described above.
[0250] Repeated AC purification cycles with low pH elution and short NaOH clean-in-place Repeated affinity chromatography purifications were performed as described above, except that the column was eluted with 0.1 M citrate, pH 2.5, instead of 3 M imidazole, pH 8. Additionally, the 0.1 M NaOH wash-in-place step was shortened to 1 ml at 1 ml / min (1 min contact time per cycle). Because the eluted SUMO-GFP was denatured by the low pH elution, relative elution concentrations were compared using densitometry of the target bands on SDS-PAGE.
[0251] Determining the effect of autoclaving or DMF incubation on the binding capacity and specificity of SUMO-binding resins For each resin tested, three 10 μl aliquots (fill volume) of resin were loaded into 1.5 ml screw-cap tubes. To one, 1 ml of DMF was added, and the tubes were incubated at room temperature for 2 hours. To the other two, 100 μl of MBS, 1 mM CaCl2, pH 7.2 was added. One of these, with its cap slightly loosened, was autoclaved on a 30-minute liquid cycle, which sterilized it at approximately 120°C and 20 psi for 30 minutes, followed by a slow reduction in pressure and temperature over the next 90 minutes. The other set of resins served as a control and was left on ice. After 2 hours, all resins were cooled to room temperature and centrifuged at 1 k × g for 1 minute. The control and autoclaved resins were stored overnight at 4°C. The DMF-treated resins were rinsed three times with fresh MBS, 1 mM CaCl2, pH 7.2, and then stored at 4°C overnight. The next day, all three bead sets were rinsed with fresh buffer and then incubated with 1.3 ml of E. coli lysate (prepared as described above) spiked with 0.2 mg / ml SUMO-GFP for 1 hour at 4°C, rotating. The resin was loaded onto a small tared column, rinsed four times with 400 ml of PBS, pH 7.4, and eluted with 3 x 25 ml of 3 M imidazole, pH 8. Fluorescence of the eluate was read in duplicate on an iD5 plate reader as described above, and concentrations were determined by comparison to a standard curve of the target and compared with controls. Concentrations were normalized and analyzed by SDS-PAGE as described above to assess purity.
[0252] Example 2: Use of protein scaffolds to target diverse antigens. We selected protein scaffolds that were conjugated to diverse protein targets. A summary of the protein scaffolds and their cognate targets is shown in Table 9 below. Table 9 contains a subset of a much larger set of target-specific nanoCLAMPs, the majority of which possess loop lengths of 4, 7, and 5 residues for loops 1, 2, and 8, respectively, as designed in library NL-26 (see above). To demonstrate the ability of protein scaffolds to tolerate various loop lengths, we included only nanoCLAMPs in Table 9 that possess at least one loop with a length different from the designed length. In Table 10, we demonstrate the ability of scaffolds to support vast loop diversity by tabulating the amino acid sequences of nanoCLAMPs specific to several targets, showing in several cases the diversity of loop sequences for a single target.
[0253] [Table 9-1] [Table 9-2] [Table 9-3] [Table 9-4] [Table 9-5] [Table 9-6] [Table 9-7] [Table 9-8] [Table 9-9] [Table 9-10] [Table 9-11] [Table 9-12] [Table 9-13] [Table 9-14] [Table 9-15] [Table 9-16]
[0254] [Table 10-1] [Table 10-2] [Table 10-3] [Table 10-4] [Table 10-5] [Table 10-6] [Table 10-7] [Table 10-8]
[0255] [Example 3] Terbium binding Lanthanide binding to the protein scaffold was demonstrated by incubating the protein with terbium, removing unbound terbium by buffer exchange, and measuring time-resolved fluorescence. Proteins were prepared at 30 μM in 20 mM MOPS, 150 mM NaCl (MBS), pH 6.5, and buffer-exchanged to remove any unbound Ca. SMT3-A1 (nC-A), P2808 (nC-B), and a negative control protein (recombinant SMT3) were added to a 140 μl reaction in the same buffer, so their final concentration was 8.57 μM, and either CaCl2 or TbCl3 was added to 300 μM. The reaction was incubated at 4°C for 16 h and then buffer-exchanged against MBS, pH 6.5, using a Zeba 7 MWCO desalting column. After diluting the protein to 0.5 μM in MBS, pH 6.5, 200 μl was analyzed in duplicate on an iD5 (Molecular Devices) plate reader using time-resolved fluorescence with Ex / Em: 350 / 544 nm and a 200 μsec delay. As shown in Figure 15, P972 (nC-A) exhibited 9-fold higher fluorescence when incubated with terbium instead of calcium, and P2808 (nC-B) exhibited 23-fold higher fluorescence when incubated with terbium instead of calcium.
[0256] [Example 4] Loop length modeling We next undertook modeling experiments to determine the protein scaffold's ability to maintain its secondary and tertiary structure while varying the loop length of each of L1-L8. For each loop, plasticity was modeled using P2808 as follows: To explore longer lengths, each loop was replaced with a flexible 15-amino acid (G4S)3 linker (sequence: GGGGSGGGSGGGGS (SEQ ID NO: 41)). The protein fold was modeled in AlphaFold mmseq without relaxation. The top result was aligned with P2808 in the Swiss PDB Viewer with MagicFit function.
[0257] The amino acids of the loops, flanking N-terminus, and flanking C-terminus are shown in different shades. Structures were qualitatively assessed for maintenance of the overall beta-sheet structure. Structures that maintained the overall beta-sheet structure were considered to maintain the overall fold. To explore shorter loop lengths, each loop was deleted completely and modeled as above. If the complete deletion did not affect the overall fold, no further constructs were modeled. Complete deletions were aligned and qualitatively assessed in the Swiss PDB Viewer with the MagicFit function.
[0258] Deletions were considered to result in a disruption of the overall beta-sheet structure if the beta-strand secondary structure assignment was converted to a coil assignment, or if one or more beta-strands lost connectivity with adjacent beta-strands. If a complete deletion resulted in a disruption of the overall beta-sheet structure, a series of deletions was created, beginning with replacing each wild-type amino acid with G, followed by removing Gs one at a time. Constructs of a series of deletions with the shortest loop lengths that qualitatively maintained the fold were aligned and evaluated. Results from these modeling experiments are shown in Figures 17-26.
[0259] Our database of nanoCLAMP conjugates was searched for any clones whose loops deviated from the standard lengths of loop 1: 4 amino acids, loop 2: 7 amino acids, and loop 8: 5 amino acids. The resulting clones are listed in Table 9 above. Tables 11 and 12 below summarize the lengths observed for clones using the nC-B and nC-A scaffolds. Table 11 shows the top-face loops, and Table 12 shows the bottom-face loops.
[0260] The third column shows the observed length variation among orthologs. The observed variation roughly correlates with the modeling data.
[0261] [Table 11]
[0262] [Table 12]
[0263] Taken together, the modeling results, the observed loop lengths in isolated nanoCLAMPs, and the natural variation in loop lengths between species indicate that the length and sequence of each of L1–L8 can be varied independently without disrupting the core fold of the protein.
[0264] [Example 5] Introduction of artificial disulfides into the nC-B scaffold to improve stability The well-characterized SUMO conjugate P2808 was modeled in AlphaFold to rationally select adjacent residues on adjacent beta strands for substitution with Cys residues, with the goal of further stabilizing the protein by introducing disulfide bonds. Substitution mutations were chosen by visual inspection of the structure to identify residues whose side chains were located in the core of the protein, oriented toward each other, and whose alpha carbons were separated by approximately the same distance as would be observed with a native disulfide bond. AlphaFold modeling predicted that some of the selected substitution mutations would form disulfide bonds and some would not (Table 13). We cloned, expressed, and purified 14 mutants predicted to form disulfides (12 mutants of P2808 and two mutants of P2960, a P2808 mutant with X and Y loops DGGGSS871-876GDT and DHTGAP900-905SST from C. celatum). Proteins were immobilized on Ni Sepharose 6 Fast Flow beads under denaturing and reducing conditions and refolded in the presence of reduced and oxidized glutathione by gradually decreasing the concentration of denaturant to favor disulfide bond formation. Refolded proteins were then eluted from the resin and examined for the presence of disulfide bonds by mobility shift on SDS-PAGE under reducing versus oxidizing conditions. Because proteins that possess intramolecular disulfides remain smaller than proteins that do not, proteins with disulfides usually run faster on SDS-PAGE due to their smaller hydrodynamic radius. Thirteen of the 14 proteins predicted by AlphaFold to form disulfide bonds migrated faster on SDS-PAGE in sample buffer lacking reducing agent than in sample buffer containing reducing agent. This observation is consistent with the agent reducing disulfide bonds in these proteins (Figure 27). Proteins that appear to possess disulfide bonds also had bands that ran similarly to their reduced forms.This observation may indicate a mixed population of proteins with and without disulfides. These preparations also contained a small percentage of high-molecular-weight species, likely intermolecular disulfide-linked multimers, which largely disappeared upon reduction. The cysteine-free control proteins P2808 and P2960 migrated at the same rate in both reducing and nonreducing buffers. Furthermore, the unrelated disulfide-containing control protein, bovine serum albumin (BSA), showed the expected decrease in electrophoretic mobility in reducing buffer. Collectively, these results indicate that disulfide bonds can be successfully introduced into the nC-B scaffold in the 12 positive constructs in Table 13. Clones P3007, P3008, and P3009 do not contain lysines. The absence of lysines is expected to reduce susceptibility to trypsin and increase the specificity of labeling with amine-reactive reagents. These scaffolds have only a single primary amine located at the N-terminal alpha amino group and are predicted to be specifically modified at this position with amine-reactive reagents. Clones P3013 and P3014 do not contain asparagine, a common site for deamidation, and are therefore predicted to be less susceptible to deamidation.
[0265] [Table 13-1] [Table 13-2]
[0266] To ensure that thermostability was not negatively affected by the mutation, we measured the melting temperatures of the mutants and the parent using differential scanning fluorimetry (DSF). We determined that the T mIntermediate or better results were defined as a positive or negative effect on the melting point. Clones P3010, P3011, P3017, and P3021 had a detrimental effect on thermostability and were not pursued further. Of the 14 proteins tested, 10 had an intermediate or better effect on melting point, exhibiting thermostability at least as good as the parental construct (Table 13). One clone, P3015, which has a disulfide bond between residues 884 and 926, showed an 8°C improvement in thermostability under oxidizing versus reducing conditions (Figure 27).
[0267] Materials and Methods AlphaFold modeling of introduced disulfides The 2808 sequences were modeled in AlphaFold with pairs of cysteine substitutions. The version is mmseq and modeling was performed using relaxation.
[0268] Disulfide analysis by SDS PAGE NanoCLAMP was purified under denaturing conditions as described above, except for the modified refolding step. Briefly, nanoCLAMP was bound to Ni Sepharose 6 Fast Flow (Cytivia) in 6 M GuHCl, 20 mM Tris, pH 8 (QCB) + 5 mM TCEP. The resin was washed with 5 column volumes (CV) of QCB + 1 mM TCEP, followed by 5 CV of QCB (without TCEP). The protein was then gradually refolded by washing with 10 CV of QCB + 2 mM GSH / 1 mM GSSG, followed by dilution with 20 mM MOPS, 150 mM NaCl (MBS), 1 mM CaCl2, 2 mM GSH / 1 mM GSSG, pH 8, down to 4, 3, 2, 1, and finally 0 M GuHCl in 5 CV steps. Each 5 CV refolding step was incubated for 30 min. The refolded protein was then washed with 10 CV of MBS, pH 8, 1 mM CaCl2, and finally eluted with MBS, pH 8, 1 mM CaCl2, 250 mM imidazole. Protein was normalized to 1 mg / ml in MBS, 1 mM CaCl2, pH 6.5. Protein was diluted 10-fold into SDS sample buffer containing 50 mM DTT (reduced) and SDS sample buffer lacking reducing agent. Protein was heated to 95°C for 5 min, cooled, and 1 μg was separated on a 12% NuPAGE BisTris gel (Thermo) using MES running buffer. The gel was stained with GelCode Blue (Thermo).
[0269] Differential Scanning Fluorometry (DSF) DSC was performed as previously described, except that in the reduction case, TCEP was included in the DSC cocktail at 50 mM (final). The DSC program was run as described, and the T was measured at the inflection point of the curve.
[0270] [Example 6] Identification of framework variants that maintain binding activity Phage library NL-26 contains nanoCLAMP clones with the wild-type nC-B framework and mutations resulting from errors in gene synthesis, PCR, and phage propagation. The library was screened and analyzed to identify nanoCLAMPs that maintained their ability to bind to their intended targets and contained one or more mutations in the framework region. The analysis resulted in the identification of 105 nanoCLAMP variants, each recognizing one of the four target antigens and each containing one or more mutations in the framework region. The number of mutations identified in each framework and their locations are summarized in Tables 14 and 15. A listing of each enriched clone, target, and framework mutation is provided in Table 16. The identification of these variants in this non-exhaustive analysis indicates that each framework region can tolerate one or more mutations while maintaining the ability to be displayed on the phage surface and mediate binding to its target.
[0271] method The NL-26 phage library was panned against recombinant GFPMut2, human serum albumin, mCherry, and TEV protease as described above. Phages were enriched in two rounds, and approximately 200,000 clones from each round were sequenced by next-generation DNA sequencing using an Illumina MiSeq system. Sequencing reads were processed and clustered using PipeBio software to count similar sequences. Clones that met the following criteria were identified: 1) a normalized enrichment greater than twofold from round 1 to round 2 and 2) one or more mutations in the framework regions.
[0272] [Table 14]
[0273] [Table 15]
[0274] [Table 16-1] [Table 16-2] [Table 16-3] [Table 16-4] [Table 16-5]
[0275] Other embodiments While the invention has been described in relation to particular embodiments thereof, it will be understood that the invention is capable of further modifications, and this application is intended to cover any variations, uses, or adaptations of the invention which generally follow the principles of the invention and include such departures from the invention as come within known or customary practice in the art to which the invention pertains, and may be applied to the essential features hereinbefore described, and which comply with the scope of the appended claims.
[0276] Other embodiments are within the scope of the following claims.
Claims
1. structure: A-F1-L1-F2-L2-F3-L3-F4-L4-F5-L5-F6-L6-F7-L7-F8-L8-F9-B 1. A protein scaffold comprising: where: F1 to F9 correspond to framework regions 1 to 9. L1 to L8 correspond to loop regions 1 to 8, A and B are each independently absent or contain at least one amino acid; F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 4; L1 is absent or comprises at least one amino acid, F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 5; L2 is absent or comprises at least one amino acid, F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 6; L3 is absent or comprises at least one amino acid, F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 7; L4 is absent or comprises at least one amino acid, F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 8; L5 is absent or comprises at least one amino acid; F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 9; L6 is absent or comprises at least one amino acid, F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 10; L7 is absent or comprises at least one amino acid; F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 11; L8 is absent or comprises at least one amino acid; F9 comprises the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 12; The protein scaffold is N807X, S809X, R812X, S813X, E814X, S815X compared to SEQ ID NO: 1 1 , D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X 2 , L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X 3 wherein the at least one mutation is selected from the group consisting of: X is any amino acid except the amino acid at the equivalent position in SEQ ID NO: 1, X 1 is any amino acid except R or S, X 2 is any amino acid except P or K, X 3 is any amino acid except R or K.
2. F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 4; F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 5; F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 6; F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having one amino acid insertion, deletion, or substitution mutation compared to SEQ ID NO: 7; F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having one amino acid insertion, deletion, or substitution mutation compared to SEQ ID NO: 8; F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having one amino acid insertion, deletion, or substitution mutation compared to SEQ ID NO: 9; F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 10; F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 11; 2. The protein scaffold of claim 1, wherein F9 comprises the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having one amino acid insertion, deletion, or substitution mutation compared to SEQ ID NO:
12.
3. F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4); F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5); F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6), F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7); F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8), F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9), F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10), F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11), 3. The protein scaffold of claim 2, wherein F9 comprises the sequence LTFSEFA-(I / V)-VS (SEQ ID NO: 12).
4. 4. The protein scaffold of any one of claims 1 to 3, wherein the protein scaffold comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X compared to SEQ ID NO: 1, wherein X is any amino acid.
5. 5. The protein scaffold of claim 4, comprising at least one mutation selected from the group consisting of N807D, S809T, R812H, S813T, E814P, S815G, D818V, N822S, N825D, N832S, W836E, K857E, E858V, I859V, K860E, L861V, D862G, R865H, K870A, N871D, N880T, K881R, K883R, N890G, K897R, K901H, K908Q, E912D, S914D, and K922Q compared to SEQ ID NO:
1.
6. 5. The protein scaffold of claim 4, wherein at least one mutation is K870X and / or N890X.
7. 7. The protein scaffold of claim 6, wherein at least one mutation is K870A and / or N890G.
8. 8. The protein scaffold of any one of claims 1 to 7, comprising at least three fewer lysines compared to SEQ ID NO:
1.
9. 9. The protein scaffold of claim 8, comprising at least 6 fewer lysines compared to SEQ ID NO:
1.
10. 10. The protein scaffold of claim 9, comprising 9 fewer lysines compared to SEQ ID NO:
1.
11. 10. The protein scaffold of claim 9, which is lysine-free.
12. 12. The protein scaffold of any one of claims 1 to 11, comprising at least three fewer asparagines compared to SEQ ID NO:
1.
13. 13. The protein scaffold of claim 12, comprising at least 5 fewer asparagines compared to SEQ ID NO:
1.
14. 14. The protein scaffold of claim 13, comprising 7 fewer asparagines compared to SEQ ID NO:
1.
15. 14. The protein scaffold of claim 13, which is asparagine-free.
16. 16. The protein scaffold of any one of claims 1 to 15, wherein A and B are each independently absent or 1 to 20 amino acids.
17. 17. The protein scaffold of any one of claims 1 to 16, wherein L1 to L8 are each independently 1 to 20 amino acids.
18. 18. The protein scaffold of claim 17, wherein L1 to L8 are each independently 1 to 10 amino acids.
19. 19. The protein scaffold of claim 18, wherein L1 to L8 are each independently 3 to 10 amino acids.
20. 20. The protein scaffold of claim 19, wherein L1 to L8 are each independently 3 to 8 amino acids.
21. 21. The protein scaffold of claim 20, wherein L1 is 4 amino acids, L2 is 7 amino acids, and / or L8 is 5 amino acids.
22. L1 is X 1 X 2 X 3 X 4 (SEQ ID NO: 13), wherein X 1 ~X 4 22. The protein scaffold of claim 21, wherein each is independently any amino acid.
23. L2 is X 1 X 2 X 3 X 4 X 5 X 6 X 7 (SEQ ID NO: 14), wherein X 1 ~X 7 23. The protein scaffold of claim 21 or 22, wherein each independently is any amino acid.
24. L8 is X 1 X 2 X 3 X 4 X 5 (SEQ ID NO: 15), wherein X 1 ~X 5 24. The protein scaffold of any one of claims 21 to 23, wherein each is independently any amino acid.
25. L4 is X 1 X 2 X 3 X 4 X 5 X 6 X 7 (SEQ ID NO: 14), wherein X 1 ~X 7 24. The protein scaffold of any one of claims 1 to 23, wherein each is independently any amino acid.
26. L6 is X 1 X 2 X 3 X 4 X 5 X 6 (SEQ ID NO: 16), wherein X 1 ~X 6 26. The protein scaffold of any one of claims 1 to 25, wherein each is independently any amino acid.
27. 27. The protein scaffold of any one of claims 1 to 26, wherein L8 comprises at least two amino acids.
28. 28. The protein scaffold of any one of claims 1 to 27, wherein L4 comprises the sequence of (G / D)-GGSS (SEQ ID NO: 17) or GDT or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 17 or GDT.
29. 29. The protein scaffold of any one of claims 1 to 28, wherein L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO:
18.
30. L4 comprises the sequence (G / D)-GGSS (SEQ ID NO: 17) or GDT, and 30. The protein scaffold of any one of claims 1 to 29, wherein L6 comprises the sequence TGAPAG (SEQ ID NO: 18).
31. 31. The protein scaffold of any one of claims 1 to 30, wherein L3 comprises the sequence (E / K / S)-(V / E)-(V / I / T)-(E / K / P / S)-(V / L)-(G / D) (SEQ ID NO: 19) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO:
19.
32. 32. The protein scaffold of any one of claims 1 to 31, wherein L5 comprises the sequence LD-(G / N)-(E / S)-S (SEQ ID NO: 20) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO:
20.
33. 33. The protein scaffold of any one of claims 1 to 32, wherein L7 comprises at least one amino acid.
34. 34. The protein scaffold of any one of claims 1 to 33, wherein L7 comprises the sequence of ETPI-(S / E)-A (SEQ ID NO: 21) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO:
21.
35. L3 comprises the sequence (E / K / S)-(V / E)-(V / I / T)-(E / K / P / S)-(V / L)-(G / D) (SEQ ID NO: 19) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 19; L5 comprises the sequence LD-(G / N)-(E / S)-S (SEQ ID NO: 20) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 20; 35. The protein scaffold of any one of claims 32 to 34, wherein L7 comprises the sequence of ETPI-(S / E)-A (SEQ ID NO: 21) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO:
21.
36. 36. The protein scaffold of any one of claims 1 to 35, wherein A comprises the sequence (D / N / H)-P.
37. 37. The protein scaffold of claim 36, wherein A comprises a sequence of DP.
38. 38. The protein scaffold of any one of claims 1 to 37, wherein B comprises the sequence DELE (SEQ ID NO: 35).
39. F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid, F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid, F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 24; L3 comprises the sequence of EVVEVG (SEQ ID NO: 31) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 31; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 25; L4 comprises the sequence GGGSS (SEQ ID NO: 32) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 32; F5 comprises the sequence of DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 26; L5 comprises the sequence of LDGES (SEQ ID NO: 33) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 33; F6 comprises the sequence of WTTIREYDH (SEQ ID NO: 27) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 27; L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 18; F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 28; L7 comprises the sequence of ETPISA (SEQ ID NO: 34) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 34; F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid; 39. The protein scaffold of any one of claims 1 to 38, wherein F9 comprises the sequence of LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO:
30.
40. F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22) or a sequence having one amino acid insertion, deletion, or substitution mutation compared to SEQ ID NO: 22; L1 is absent or comprises at least one amino acid, F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 23; L2 is absent or comprises at least one amino acid, F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 24; L3 comprises the sequence EVVEVG (SEQ ID NO: 31) or a sequence having an insertion, deletion, or substitution mutation of one amino acid compared to SEQ ID NO: 31; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having an insertion, deletion or substitution mutation of one amino acid compared to SEQ ID NO: 25; L4 comprises the sequence GGGSS (SEQ ID NO: 32) or a sequence having an insertion, deletion, or substitution mutation of one amino acid compared to SEQ ID NO: 32; F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one amino acid insertion, deletion, or substitution mutation compared to SEQ ID NO: 26; L5 comprises the sequence of LDGES (SEQ ID NO: 33) or a sequence having an insertion, deletion, or substitution mutation of one amino acid compared to SEQ ID NO: 33; F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 27; L6 comprises the sequence of TGAPAG (SEQ ID NO: 18) or a sequence having an insertion, deletion, or substitution mutation of one amino acid compared to SEQ ID NO: 18; F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28) or a sequence having an insertion, deletion or substitution mutation of one amino acid compared to SEQ ID NO: 28; L7 comprises the sequence of ETPISA (SEQ ID NO: 34) or a sequence having an insertion, deletion or substitution mutation of one amino acid compared to SEQ ID NO: 34; F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29) or a sequence having one amino acid insertion, deletion or substitution mutation compared to SEQ ID NO: 29; L8 is absent or comprises at least one amino acid; 40. The protein scaffold of claim 39, wherein F9 comprises the sequence of LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one amino acid insertion, deletion, or substitution mutation compared to SEQ ID NO:
30.
41. F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22), L1 is absent or comprises at least one amino acid, F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 is absent or comprises at least one amino acid, F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8 is absent or comprises at least one amino acid; 41. The protein scaffold of claim 40, wherein F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30).
42. F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22), L1 is X 1 X 2 X 3 X 4 (SEQ ID NO: 13), wherein X 1 ~X 4 are each independently any amino acid, F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 is X 1 X 2 X 3 X 4 X 5 X 6 X 7 (SEQ ID NO: 14), wherein X 1 ~X 7 are each independently any amino acid, F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8, X 1 X 2 X 3 X 4 X 5 (SEQ ID NO: 15), wherein X 1 ~X 5 are each independently any amino acid, 42. The protein scaffold of claim 41, wherein F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30).
43. A comprises the sequence DP, F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22), L1 is X 1 X 2 X 3 X 4 (SEQ ID NO: 13), wherein X 1 ~X 4 are each independently any amino acid, F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), L2 is X 1 X 2 X 3 X 4 X 5 X 6 X 7 (SEQ ID NO: 14), wherein X 1 ~X 7 are each independently any amino acid, F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), L3 comprises the sequence EVVEVG (SEQ ID NO: 31), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), L4 comprises the sequence GGGSS (SEQ ID NO: 32), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), L5 comprises the sequence LDGES (SEQ ID NO: 33), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), L6 comprises the sequence TGAPAG (SEQ ID NO: 18), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), L7 comprises the sequence ETPISA (SEQ ID NO: 34), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), L8, X 1 X 2 X 3 X 4 X 5 (SEQ ID NO: 15), wherein X 1 ~X 5 are each independently any amino acid, F9 comprises the sequence LTFSEFAIVS (SEQ ID NO: 30), 43. The protein scaffold of claim 42, wherein B comprises the sequence DELE (SEQ ID NO: 35).
44. L1 is X 1 X 2 X 3 X 4 (SEQ ID NO: 13), 1 , X 3 , and X 4 are each independently any amino acid, and X 2 44. The protein scaffold of claim 42 or 43, wherein
45. A protein scaffold comprising a polypeptide having at least 80% sequence identity to SEQ ID NO:
3.
46. 46. The protein scaffold of claim 45, wherein the polypeptide has at least 85%, 90%, 95%, 97%, or 99% sequence identity to SEQ ID NO:
3.
47. The polypeptide has the following amino acids: N807X, S809X, R812X, S813X, E814X, S815X compared to SEQ ID NO: 1 1 , D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X 2 , L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X 3 wherein the at least one mutation is selected from the group consisting of: X is any amino acid except the amino acid at the equivalent position in SEQ ID NO: 1, X 1 is any amino acid except R or S, X 2 is any amino acid except P or K, X 3 is any amino acid except R or K.
48. 48. The protein scaffold of claim 47, wherein the protein scaffold comprises at least one mutation selected from the group consisting of N807X, S809X, R812X, S813X, E814X, S815X, D818X, N822X, N825X, N832X, W836X, K857X, E858X, I859X, K860X, L861X, D862X, R865X, K870X, N871X, N880X, K881X, K883X, N890X, K897X, K901X, K908X, E912X, S914X, and K922X compared to SEQ ID NO: 1, wherein X is any amino acid.
49. 49. The protein scaffold of claim 48, comprising at least one mutation selected from the group consisting of N807D, S809T, R812H, S813T, E814P, S815G, D818V, N822S, N825D, N832S, W836E, K857E, E858V, I859V, K860E, L861V, D862G, R865H, K870A, N871D, N880T, K881R, K883R, N890G, K897R, K901H, K908Q, E912D, S914D, and K922Q compared to SEQ ID NO:
1.
50. A protein scaffold comprising framework regions and loop regions, wherein the protein scaffold has the following structure: A-F1-L1-F2-L2-F3-L3-F4-L4-F5-L5-F6-L6-F7-L7-F8-L8-F9-B and comprising at least seven framework regions derived from where: F1 to F9 correspond to framework regions 1 to 9. L1-L8 each independently correspond to loop regions 1-8 that are absent or contain one or more amino acids; A and B are each independently absent or contain at least one amino acid; F1 comprises the sequence (T / S)-LI-(H / R)-(T / S)-(P / E)-(G / S)-W (SEQ ID NO: 4) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 4; F2 comprises the sequence G-(S / N / T)-E-(A / S)-(D / N / S / A)-LLDGDD-(S / N / T)-TGV-(E / W / A)-Y (SEQ ID NO: 5) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 5; F3 comprises the sequence S-(L / V)-AGEFIGLDLG (SEQ ID NO: 6) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 6; F4 comprises the sequence G-(I / V)-(H / R / Y / N)-FVIG-(A / K / R)-(D / N) (SEQ ID NO: 7) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 7; F5 comprises the sequence DKW-(T / N / S)-(R / K)-F-(R / K)-LEYS (SEQ ID NO: 8) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 8; F6 comprises the sequence WTTI-(R / K / H / Q)-EYD-(H / K / R / Q) (SEQ ID NO: 9) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 9; F7 comprises the sequence (Q / K)-DVI-(D / E)-E-(D / S)-F (SEQ ID NO: 10) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 10; F8 comprises the sequence (Q / K / R)-YIRLTNLE (SEQ ID NO: 11) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 11; A protein scaffold, wherein F9 comprises the sequence of LTFSEFA-(I / V)-VS (SEQ ID NO: 12) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO:
12.
51. F1 comprises the sequence of TLIHTPGW (SEQ ID NO: 22) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 22; F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 23; F3 comprises the sequence of SLAGEFIGLDLG (SEQ ID NO: 24) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 24; F4 comprises the sequence of GIHFVIGAD (SEQ ID NO: 25) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 25; F5 comprises the sequence of DKWTRFRLEYS (SEQ ID NO: 26) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 26; F6 comprises the sequence of WTTIREYDH (SEQ ID NO: 27) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 27; F7 comprises the sequence of QDVIDEDF (SEQ ID NO: 28) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or combinations thereof compared to SEQ ID NO: 28; F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO: 29; 51. The protein scaffold of claim 50, wherein F9 comprises at least seven of the framework regions comprising the sequence of LTFSEFAIVS (SEQ ID NO: 30) or a sequence having one or two amino acid insertions, deletions, substitution mutations, or a combination thereof compared to SEQ ID NO:
30.
52. F1 comprises the sequence TLIHTPGW (SEQ ID NO: 22), F2 comprises the sequence GSEADLLDGDDSTGVEY (SEQ ID NO: 23), F3 comprises the sequence SLAGEFIGLDLG (SEQ ID NO: 24), F4 comprises the sequence GIHFVIGAD (SEQ ID NO: 25), F5 comprises the sequence DKWTRFRLEYS (SEQ ID NO: 26), F6 comprises the sequence WTTIREYDH (SEQ ID NO: 27), F7 comprises the sequence QDVIDEDF (SEQ ID NO: 28), F8 comprises the sequence QYIRLTNLE (SEQ ID NO: 29), 52. The protein scaffold of claim 51, wherein F9 comprises at least seven of the framework regions comprising the sequence of LTFSEFAIVS (SEQ ID NO: 30).
53. 53. The protein scaffold of any one of claims 50 to 52, comprising at least eight of the framework regions F1 to F9.
54. 54. The protein scaffold of any one of claims 50 to 53, comprising at least 80%, 85%, 90%, 95%, 97%, or 99% sequence identity to framework regions F1 to F9 over one or more regions of alignment.
55. 55. The protein scaffold of any one of claims 1 to 54, further comprising a substitution mutation that adds a cysteine residue.
56. 56. The protein scaffold of claim 55, comprising a first substitution mutation that adds a first cysteine residue and a second substitution mutation that adds a second cysteine residue.
57. 57. The protein scaffold of claim 56, wherein the first cysteine residue and the second cysteine residue form a disulfide bond under oxidizing conditions.
58. 58. The protein scaffold of any one of claims 55 to 57, comprising at least one mutation selected from the group consisting of: F806C, P808C, S845C, L855C, V858C, V861C, K878C, W879C, L884C, L888C, A904C, P905C, A906GC, G907C, I924C, L926C, N928C, L936C, I943C, L948C.
59. 59. The protein scaffold of claim 58, comprising at least two or more mutations selected from the group consisting of: F806C, P808C, S845C, L855C, V858C, V861C, K878C, W879C, L884C, L888C, A904C, P905C, A906GC, G907C, I924C, L926C, N928C, L936C, I943C, L948C.
60. 60. The protein scaffold of claim 59, comprising a cysteine mutation pair selected from the group consisting of K878C and G907C, K878C and A904C, V861C and I943C, P905C and L855C, S845C and L936C, W879C and N928C, L884C and L926C, F806C and L948C, V858C and L888C, K878C and G907C, K878C and A906GC, S845C and N928C, K878C and A904C, P808C and I943C, V861C and I924C, P808C and V861C, and I943C and L855C.
61. 61. The protein scaffold of claim 60, wherein the cysteine mutation pairs are selected from the group consisting of K878C and G907C, K878C and A904C, S845C and L936C, W879C and N928C, W879C and N928C, L884C and L926C, V858C and L888C, K878C and G907C, and K878C and A906GC.
62. 62. The protein scaffold of any one of claims 1 to 61, further comprising a tag covalently attached to the scaffold.
63. 63. The protein scaffold of claim 62, wherein the tag is an affinity tag.
64. 64. The protein scaffold of claim 62 or 63, wherein the tag is attached to the N-terminus or C-terminus of the scaffold.
65. 65. The protein scaffold of any one of claims 1 to 64, wherein the scaffold is conjugated to a functional group.
66. 66. The protein scaffold of claim 65, wherein the functional group comprises biotin, streptavidin or a derivative of streptavidin, a polyethylene glycol moiety, a fluorescent dye, an enzyme, a radioactive moiety, a lanthanide, or a lanthanide binding motif.
67. 67. The protein scaffold of claim 66, wherein the lanthanide is terbium.
68. 67. The protein scaffold of claim 66, wherein the radioactive moiety is an alpha or beta emitter.
69. 69. The protein scaffold of any one of claims 65 to 68, wherein the functional group is conjugated to a sulfhydryl group or a primary amine.
70. 70. A polynucleotide encoding the protein scaffold of any one of claims 1 to 69.
71. 71. The polynucleotide of claim 70, which is a ribonucleotide.
72. 71. The polynucleotide of claim 70, which is a deoxyribonucleotide.
73. 73. A vector comprising the polynucleotide of claim 71 or 72.
74. 74. A cell comprising a polynucleotide according to any one of claims 70 to 72 or a vector according to claim 73.
75. 70. A method of producing a protein scaffold according to any one of claims 1 to 69, comprising: (a) providing a cell transformed with a polynucleotide according to any one of claims 70 to 72 or a vector according to claim 73; (b) culturing the transformed cells under conditions that allow expression of the polynucleotide, wherein culturing results in expression of the protein scaffold; and (c) isolating the protein scaffold A method comprising:
76. 70. A particle comprising the protein scaffold of any one of claims 1 to 69.
77. 77. The particle of claim 76, which is a magnetic particle.
78. 78. A resin comprising a plurality of particles according to claim 76 or 77.
79. 79. A column comprising the resin of claim 78.
80. 1. A method for purifying a target molecule from a plurality of molecules, comprising: (a) providing a sample containing a target molecule and a mixture of molecules; (b) contacting the sample with a protein scaffold according to any one of claims 1 to 69, wherein the scaffold specifically binds to the target molecule; and (c) separating the target molecule bound to the protein scaffold from the plurality of molecules. A method comprising:
81. 81. The method of Claim 80, wherein the separating step comprises immobilizing the protein scaffold.
82. 82. The method of claim 81, wherein the protein scaffold is conjugated to a particle.
83. 83. The method of claim 82, wherein the particles comprise magnetic beads.
84. 84. The method of claim 82 or 83, wherein the protein scaffold is conjugated to a resin or monolith comprising a plurality of particles.