A split intein and a method for preparing a recombinant polypeptide using the same

Novel side sequences for split inteins address the instability and immune response issues of double-specific antibodies, enabling efficient production of stable, long-lasting antibodies with reduced immune reactions.

CN114450406BActive Publication Date: 2025-07-15WUHAN YZY BIOPHARMA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080063323.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-09
Filing Date
2020-09-09
Publication Date
2025-07-15
Estimated Expiration
2040-09-09

AI Technical Summary

Technical Problem

The existing bispecific antibodies have short half-life in vivo, strong immune response, poor stability, and traditional modification methods are difficult to retain the Fc fragment effect function, resulting in limited application.

Method used

By performing amino acid mutations on the flanking sequence of the fractured inteppeptide, new flanking sequence pairs are screened for preparation of recombinant polypeptides, especially bispecific antibodies, avoiding the introduction of free thiol groups, improving splicing efficiency, and expressing them in mammalian cells, retaining the natural antibody structure.

Benefits of technology

The prepared bispecific antibodies have stable structure, low immunogenicity, long half-life, and have the Fc domain function of natural antibodies, avoid heavy chain light chain mismatch problems, and are suitable for humanized and whole-human sequence antibodies, reducing immune response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GPA0000318395260000141
    Figure GPA0000318395260000141
  • Figure GPA0000318395260000151
    Figure GPA0000318395260000151
  • Figure GPA0000318395260000161
    Figure GPA0000318395260000161
Patent Text Reader

Abstract

Flanking sequence pair for split intein, wherein the flanking sequence pair includes: flanking sequence a and flanking sequence b; the flanking sequence a is located at the N-terminus of the N-terminal protein splicing region (In) of the split intein and is between the N-terminal exon (En) and In; the flanking sequence b is located at the C-terminus of the C-terminal protein splicing region (Ic) of the split intein and is between Ic and the C-terminal exon (Ec); the split intein is NpuDnaE.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to split inteins containing novel flanking sequence pairs, recombinant polypeptides using the same, and the application of the inteins in the preparation of antibodies, especially bispecific antibodies. The present invention also relates to a screening method for the split inteins containing novel flanking sequence pairs. Background Art

[0002] Protein trans-splicing refers to a protein splicing reaction mediated by split inteins. In this type of splicing process, first, the N-terminal fragment or N-terminal protein splicing region (In) and the C-terminal fragment or C-terminal protein splicing region (Ic) of the split intein recognize and bind to each other non-covalently. Once bound and correctly folded, the split intein that reconstructs the active center completes the protein splicing reaction according to the typical protein splicing pathway, connecting the two flanking exteins (Saleh.L., Chemical Record. 6 (2006) 183-193).

[0003] In the technology of preparing recombinant proteins, the gene encoding the precursor protein can be split into two open reading frames, and the protein trans-splicing reaction can be catalyzed by a split intein including two parts: the N-terminal protein splicing region (N’fragment of intein, abbreviated as In) and the C-terminal protein splicing region (C’fragment of intein, abbreviated as Ic), so as to connect the two separated exteins (En, Ec) constituting the precursor protein with peptide bonds to obtain a recombinant protein (Ozawa.T., Nat Biotechbol. 21 (2003) 287-93).

[0004] A bispecific antibody refers to an antibody molecule that can simultaneously recognize two antigens or two epitopes. For example, bispecific or multispecific antibodies that can bind more than two antigens are known in the art and can be obtained in eukaryotic expression systems or prokaryotic expression systems by methods such as cell fusion, chemical modification, and gene recombination.

[0005] Currently, a wide variety of recombinant bispecific antibody formats have been developed. For example, tetravalent bispecific antibodies by fusing, for example, IgG antibody formats and single-chain domains (see, e.g., Coloma, M.J., et al., Nature Biotech. 15 (1997) 159-163; WO 2001077342; and Morrison, S., L., Nature Biotech. 25 (2007) 1233-1234). However, such antibodies, due to their large difference from the native antibody structure, can cause strong immune responses and short half-lives after entering the body.

[0006] In addition, several other novel formats capable of binding two or more antigens have also been developed, such as: small molecule antibodies like minibodies, several single-chain formats (scFv bis-scFv), etc. In these small molecule antibodies, the antibody central structure (IgA, IgD, IgE, IgG or IgM) is no longer maintained (Holliger, P., et al., Nature Biotech. 23 (2005) 1126-1136; Fischer, N, and Leger, O., Pathobiology 74 (2007) 3-14; Shen, J., et al., J. Immunol. Methods. 318 (2007) 65-74; Wu, C., et al., Nature Biotech. 25 (2007) 1290-1297).

[0007] This modification of connecting the core binding regions of antibodies with other antibody core binding regions through a linker has obvious advantages over bispecific antibodies. However, there are also problems in its application as a drug, which greatly limits its development as a drug.

[0008] In fact, in terms of immunogenicity, these foreign proteins may cause immune responses against the linker itself or the protein containing the linker, and even cause an immune storm. In addition, due to the flexible nature of these linkers, they tend to undergo protein degradation, easily leading to poor antibody stability, easy aggregation, shortened half-life and further enhanced immunogenicity. For example, Blinatumomab of Amgen has a half-life of only 1.25 hours in the blood, resulting in the need for continuous administration through an infusion pump for 24 hours, which greatly limits its application (Bargou, R and Leo.E., Science. 321 (2008) 974-7).

[0009] In addition, in the modification of bispecific antibodies, it is desirable to retain the effector functions of the Fc fragment of the antibody: for example, CDC (complement-dependent cytotoxicity), or ADCC (antibody-dependent cell-mediated cytotoxicity), and to extend the half-life of the antibody binding to the FcRn (Fc receptor) on the inner wall of blood vessels. These functions must be mediated by the Fc region. Therefore, it is necessary to retain the Fc region in the modified bispecific antibody.

[0010] Therefore, it is necessary to develop bispecific antibodies with structures extremely similar to those of naturally occurring antibodies (such as IgA, IgD, IgE, IgG, IgM). Furthermore, humanized bispecific antibodies and fully human bispecific antibodies with minimal differences from human antibody sequences are required.

[0011] Currently, the trans-splicing mechanism of the Npu-PCC73102 DnaE (abbreviated as NpuDnaE) intein has been attempted to prepare bispecific antibodies. When using the trans-splicing mechanism of the intein to prepare bispecific antibodies, there is no linker peptide in the splicing product, but there are the following problems: in the bispecific antibodies obtained in this way, the free sulfhydryl groups introduced by the Ic flanking sequence cannot be avoided, resulting in a high risk of misfolding and instability of such bispecific antibodies, and there are also problems with splicing efficiency (Han L, Zong H, et al., Naturally split intein Npu DnaE mediated rapid generation of bispecific IgG antibodies, Methods,. Vol 154, 2019 Feb 1;154:32-37).

[0012] The protein splicing efficiency mediated by split inteins is directly related to the intein sequence and flanking sequences of the intein.

[0013] More than 600 split inteins are listed in the NEB database (http: / / inteins.com / ). Commonly used ones are, for example, NpuDnaE, whose In flanking sequence is AEY (En-AEY-In), and the Ic flanking sequence is CFNGT (Ic-CFNGT-Ec). After splicing, En-AEY-In and Ic-CFNGT-Ec form a protein in the form of En-AEYCFNGT-Ec, which will retain a cysteine, resulting in free sulfhydryl groups in the splicing product and greatly increasing the risk of misfolding and instability of the product.

[0014] To avoid having free sulfhydryl groups in the splicing products, it is necessary to improve the existing flanking sequence pairs of split inteins, and novel flanking sequences are required: novel flanking sequence pairs that maintain the good splicing efficiency of the intein and do not contain cysteine residues.

[0015] It has been reported in the literature that by performing amino acid mutations on the flanking sequences of NpuDnaE and screening, the In flanking sequence was obtained as MGG (En-MGG-In), and the Ic flanking sequence was SVY (Ic-SVY-Ec). When using an intein with such flanking sequences for trans-splicing, En-MGG-In and Ic-SVY-Ec are spliced to obtain En-MGGSVY-Ec (Cheriyan M., et al., J Mol Biol. 2014 Dec 12;426(24):4018-4029), thus achieving no free sulfhydryl groups in the final product.

[0016] In addition, after performing amino acid mutations on the existing flanking sequence pairs of split inteins, it will affect the efficiency of their trans-splicing. Therefore, a screening method is needed to screen for inteins containing novel flanking sequence pairs, which have excellent trans-splicing efficiency and do not introduce free sulfhydryl groups at the interface in the splicing products. Further, a split intein suitable for preparing antibodies, especially bispecific antibodies, is needed, which has excellent trans-splicing efficiency and does not introduce free sulfhydryl groups at the interface in the cleavage products and contains novel flanking sequence pairs. Summary of the Invention

[0017] Through the diligent research of the inventors, the present invention obtained a class of split inteins with novel flanking sequence pairs by performing regular amino acid mutations on the existing flanking sequence pairs of inteins and screening for flanking sequence pairs with excellent trans-splicing efficiency. They have flanking sequences without cysteine residues, do not introduce free sulfhydryl groups at the interface in the cleavage products, have excellent trans-splicing efficiency, and are particularly suitable for preparing antibodies (especially bispecific antibodies).

[0018] Using the split intein of the present invention, polypeptide fragments from different proteins can be spliced together to form a recombinant fusion polypeptide protein with high splicing efficiency under relatively mild conditions (such as normal temperature, physiological salt concentration, neutral pH, etc.).

[0019] In addition, based on the screening of the above-mentioned split inteins, the present inventors have established a method for preparing recombinant polypeptides, especially bispecific antibodies, using split inteins. The bispecific antibodies prepared by the method for preparing bispecific antibodies according to the present invention do not have non-natural domains, and their structures are extremely similar to the structures of natural antibodies (IgA, IgD, IgE, IgG or IgM), and have an Fc domain. The structure of the bispecific antibody is intact and has good stability, and can retain or remove CDC (complement-dependent cytotoxicity) or ADCC (antibody-dependent cell cytotoxicity) or ADCP (antibody-dependent cell phagocytosis) or FcRn (Fc receptor) binding activity according to different IgG subtypes.

[0020] The bispecific antibodies prepared by the method of the present invention have the following advantages: the bispecific antibodies have a long in vivo half-life and low immunogenicity; no linker of any form is introduced, the stability of the antibody molecule is improved, and the immune response in vivo is reduced.

[0021] The bispecific antibodies prepared by the method of the present invention can be prepared using a mammalian cell expression system, thereby having glycosylation modifications consistent with wild-type IgG, obtaining better biological functions, being more stable, and having a long in vivo half-life; by using the in vitro splicing method performed by inteins, the problems of heavy chain mispairing and light chain mispairing that are extremely likely to occur in traditional methods can be completely avoided.

[0022] The method for preparing bispecific antibodies of the present invention can also be used to produce humanized bispecific antibodies and fully human sequence bispecific antibodies. The sequences of such antibodies prepared by the method of the present invention are closer to human antibodies, and the occurrence of immune responses can be effectively reduced.

[0023] The method for preparing bispecific antibodies of the present invention is a construction method for general bispecific antibodies, which is not restricted by antibody subtypes (IgG, IgA, IgM, IgD, IgE, and light chain κ and λ types), and does not require designing different mutations according to specific targets, and can be used to construct any bispecific antibody.

[0024] The present invention provides the following technical solutions.

[0025] 1. A pair of flanking sequences for split inteins, wherein,

[0026] the pair of flanking sequences includes: flanking sequence a and flanking sequence b; the flanking sequence a is located at the N-terminal of the N-terminal protein splicing region (In) of the split intein and is between the N-terminal exon peptide (En) and In; the flanking sequence b is located at the C-terminal of the C-terminal protein splicing region (Ic) of the split intein and is between Ic and the C-terminal exon peptide (Ec);

[0027] The cleavage-type intein is NpuDnaE,

[0028] The flanking sequence a is A -3 A -2 A -1 , and the flanking sequence b is B1B2B3, where:

[0029] A -3 is X or absent; A -2 is selected from D, F, G, L, N, S or W; A -1 is selected from G, A, K, Q, R, W, T or S;

[0030] B1 is S; B2 is E; B3 is X or absent, or preferably T, I, A, D, E, F, H, L, M, S, V, W or Y;

[0031] Preferably,

[0032] the flanking sequence a is GG, SG, XGG, XSG, GA, GK, GQ, GR, GW, GT, GS, XGA, XGK, XGQ, XGR, XGW, XGT, XGS, DG, FG, LG, NG, WG and the flanking sequence b is SE or SEX,

[0033] wherein the X is any one amino acid selected from: G, A, V, L, M, I, S, T, P, N, Q, F, Y, W, K, R, H, D, E, C.

[0034] 2. The flanking sequence pair for a cleavage-type intein according to 1 above, wherein the cleavage-type intein and the flanking sequence pair are used together for trans-splicing,

[0035] wherein,

[0036] the NpuDnaE is composed of In with the sequence of SEQ ID NO: 31 and Ic with the sequence of SEQ ID NO: 32,

[0037] Preferably, the flanking sequence a is GG or SG, and the flanking sequence b is SET or SEI or SES or SEH; or the flanking sequence a is GA, GK, GQ, GR, GW, GT, GS, and the flanking sequence b is SET or SEI or SES or SEH; or the flanking sequence a is DG, FG, LG, NG, WG and the flanking sequence b is SET or SEI or SES or SEH.

[0038] 3. A recombinant polypeptide obtained by trans-splicing using the flanking sequence pair for a cleavage-type intein according to 1 or 2 above.

[0039] 4. The recombinant polypeptide according to item 3 above, wherein the recombinant polypeptide is obtained by trans-splicing of component A and component B;

[0040] In component A, the N-terminus of the flanking sequence a is connected to the C-terminus of the N-terminal exon peptide (En), and the C-terminus of the flanking sequence a is connected to the In, and optionally a tag protein is connected to the C-terminus of the In;

[0041] In component B, the C-terminus of the flanking sequence b is connected to the N-terminus of the C-terminal exon peptide (Ec), and the N-terminus of the flanking sequence b is connected to the Ic, and optionally a tag protein is connected to the N-terminus of the Ic;

[0042] Wherein, the coding sequences of the N-terminal exon peptide (En) and the C-terminal exon peptide (Ec) are respectively from the N-terminal part and the C-terminal part of the same protein,

[0043] Preferably, the tag protein is selected from SEQ ID NO: 24, 25, 26, 27, 28, 29 or 30.

[0044] 5. The recombinant polypeptide according to item 3 above, wherein the recombinant polypeptide is obtained by trans-splicing of component A and component B;

[0045] In component A, the N-terminus of the flanking sequence a is connected to the C-terminus of the N-terminal exon peptide (En), and the C-terminus of the flanking sequence a is connected to the In, and optionally a tag protein is connected to the C-terminus of the In;

[0046] In component B, the C-terminus of the flanking sequence b is connected to the N-terminus of the C-terminal exon peptide (Ec), and the N-terminus of the flanking sequence b is connected to the Ic, and optionally a tag protein is connected to the N-terminus of the Ic;

[0047] Wherein, the coding sequences of the N-terminal exon peptide (En) and the C-terminal exon peptide (Ec) are from different proteins.

[0048] 6. The recombinant polypeptide according to item 4 or 5 above, which is a fluorescent protein, protease, signal peptide, antimicrobial peptide, antibody, or a polypeptide with biological toxicity.

[0049] 7. The recombinant polypeptide according to item 4 or 5 above, wherein one or more of the same protein or the different proteins are antibodies.

[0050] 8. The recombinant polypeptide according to item 7 above, wherein the antibody is of the natural immunoglobulin IgG, IgM, IgA, IgD, or IgE class, or immunoglobulin subclass: IgG1, IgG2, IgG3, IgG4, IgG5, or different classes of light chains: κ, λ; or single-domain antibody; or

[0051] The antibody is a full-length antibody or a functional fragment of an antibody.

[0052] 9. The recombinant polypeptide according to item 8 above, wherein the functional fragment of the antibody is selected from one or more of: the variable region of the heavy chain of the antibody VH, the variable region of the light chain of the antibody VL, the fragment Fc of the constant region of the heavy chain of the antibody, the constant region 1 CH1 of the heavy chain of the antibody, the constant region 2 CH2 of the heavy chain of the antibody, the constant region 3 CH3 of the heavy chain of the antibody, the constant region CL of the light chain of the antibody, or the variable region VHH of the single-domain antibody.

[0053] 10. The recombinant polypeptide according to item 7 above, wherein one or more of the same protein or the different proteins are specific for antigen or epitope A.

[0054] The antigen A includes: tumor cell surface antigen, immune cell surface antigen, cytokine, cytokine receptor, transcription factor, membrane protein, actin, virus, bacterium, endotoxin, FIXa, FX, CD3, SLAMF7, CD38, BCMA, CD20, CD16, CEA, PD-L1, PD-1, CTLA-4, TIGIT, LAG-3, VEGF, B7-H3, Claudin18.2, TGF-β, Her2, IL-10, Siglec-15, Ras, C-myc, and the epitope A is the immunogenic epitope of the antigen A.

[0055] 11. The recombinant polypeptide according to item 10 above, wherein one or more of the same protein or the different proteins are specific for an antigen or epitope B different from antigen or epitope A.

[0056] The antigen B includes: tumor cell surface antigen, immune cell surface antigen, cytokine, cytokine receptor, transcription factor, membrane protein, actin, virus, bacterium, endotoxin, FIXa, FX, CD3, SLAMF7, CD38, BCMA, CD20, CD16, CEA, PD-L1, PD-1, CTLA-4, TIGIT, LAG-3, VEGF, B7-H3, Claudin18.2, TGF-β, Her2, IL-10, Siglec-15, Ras, C-myc, and the epitope B is the immunogenic epitope of the antigen B.

[0057] 12. The recombinant polypeptide according to item 11 above, which is a bispecific antibody that can bind antigen or epitopes A and B simultaneously, preferably a humanized bispecific antibody or a bispecific antibody with a fully human sequence.

[0058] 13. The recombinant polypeptide according to any one of items 7 to 11 above, wherein

[0059] Component A comprises: the light chain of an antibody, the VH+CH1 chain of an antibody with In fused to the C-terminus, or the variable region VHHa of a single-domain antibody with In fused to the C-terminus, optionally with a tag protein linked to the C-terminus of In.

[0060] Component B comprises: the light chain of an antibody, the complete heavy chain of an antibody, and the Fc chain with Ic fused to the N-terminus, or the variable region VHHb of a single-domain antibody with Ic fused to the N-terminus, optionally with a tag protein linked to the N-terminus of Ic. VHHa and VHHb may be the same or different.

[0061] 14. The recombinant polypeptide according to any one of items 3 to 13 above, wherein,

[0062] The tag protein is selected from: Fc, His-tag, Strep-tag, Flag, HA, or maltose-binding protein MBP.

[0063] 15. A composition comprising the recombinant polypeptide according to any one of items 3 to 14 above.

[0064] 16. A composition which, in addition to the recombinant polypeptide according to any one of items 3 to 14 above, further comprises a carrier.

[0065] 17. The composition according to item 16 above, which is a pharmaceutical composition, and the carrier is a pharmaceutically acceptable carrier.

[0066] 18. A carrier linked to the recombinant polypeptide according to any one of items 3 to 14 above, preferably for purification purposes including chromatography.

[0067] 19. A kit comprising the recombinant polypeptide according to any one of items 3 to 14 above, which is used to detect the presence of antigen or epitope A and / or antigen or epitope B in a sample. Preferably, the recombinant polypeptide is in a state preserved in a liquid or freeze-dried powder, and may optionally exist alone or be in a state of being linked, complexed, associated, or chelated and immobilized on a carrier.

[0068] 20. An expression vector for preparing the expression vector of the recombinant polypeptide according to any one of items 3 to 14 above.

[0069] 21. A method for preparing a recombinant polypeptide, which includes:

[0070] (1) Providing component A and component B, wherein component A includes a flanking sequence a, an N-terminal exon peptide En, and In, the N-terminus of the flanking sequence a is connected to the C-terminus of the N-terminal exon peptide En, and the C-terminus of the flanking sequence a is connected to In, optionally with a tag protein further linked to the C-terminus of In;

[0071] Component B includes a flanking sequence b, a C-terminal exopeptide Ec, and Ic. The C-terminus of the flanking sequence b is connected to the N-terminus of the C-terminal exopeptide Ec, and the N-terminus of the flanking sequence b is connected to Ic. Optionally, a tag protein is connected to the N-terminus of Ic;

[0072] Among them, the flanking sequence a and the flanking sequence b are as described in 1 or 2 above. The coding sequences of the N-terminal exopeptide En and the C-terminal exopeptide Ec are from the same protein or different proteins; and

[0073] (2) Perform in vitro trans-splicing on Component A and Component B to obtain a recombinant polypeptide;

[0074] Preferably, in step (1), it includes expressing Component A and Component B in a cell containing nucleic acid sequences encoding Component A and Component B; preferably, the N-terminal exopeptide En and the C-terminal exopeptide Ec can be different domains of an antibody.

[0075] 22. The method for preparing the recombinant polypeptide described in 21 above, which further includes:

[0076] A first purification step of chromatography on Component A and Component B before trans-splicing;

[0077] A second purification step of chromatography on the recombinant polypeptide obtained by trans-splicing;

[0078] Preferably, the chromatography method in the first purification step is selected from protein A, protein G, nickel column, Strep-Tactin affinity chromatography, anti-Flag antibody affinity chromatography, anti-HA antibody affinity chromatography, or cross-linked starch affinity chromatography, and

[0079] Preferably, the chromatography method in the second purification step is an affinity chromatography method corresponding to the tag protein to remove unspliced components, or remove unspliced components by ion exchange, hydrophobicity, or molecular sieve.

[0080] 23. According to the method for preparing the recombinant polypeptide described in 21 above, wherein the recombinant polypeptide is a bispecific antibody, and the coding sequences of the bispecific antibody belong to two different antibodies P and antibody R respectively;

[0081] 1) Split antibody P into En P and Ec P , and design the sequences of Component A and Component B; split antibody R into En R and Ec R , and design Component A' and Component B'; wherein,

[0082] Component A includes a flanking sequence a, En P and In. The N-terminus of the flanking sequence a is connected to the EnP is connected to the C-terminus, and the C-terminus of the flank sequence a is connected to the In, and optionally a tag protein is further connected to the C-terminus of the In; Component B includes flank sequence b, Ec P and Ic, the C-terminus of the flank sequence b is connected to the N-terminus of Ec P is connected to the N-terminus, and the N-terminus of the flank sequence b is connected to the Ic, and optionally a tag protein is connected to the N-terminus of the Ic;

[0083] Component A' includes flank sequence a, En R and In, the N-terminus of the flank sequence a is connected to the C-terminus of the Ra, and the C-terminus of the flank sequence a is connected to the In, and optionally a tag protein is further connected to the C-terminus of the In; Component B' includes flank sequence b, Ec R and Ic, the C-terminus of the flank sequence b is connected to the N-terminus of Ec R is connected to the N-terminus, and the N-terminus of the flank sequence b is connected to the Ic, and optionally a tag protein is connected to the N-terminus of the Ic;

[0084] 2) Perform trans-splicing on the Component A and Component B', and / or on the Component A' and Component B to obtain the bispecific antibody.

[0085] 24. A screening method for the flank sequence pairs of split inteins, the method comprising:

[0086] 1) Split the amino acid sequence of protein P;

[0087] 2) The flank sequence a is a combination of 2 to 3 amino acids independently designed, denoted as flank sequences a1 to an, and the flank sequence b is a combination of 2 to 3 amino acids independently designed, denoted as flank sequences b1 to bn; wherein, the amino acids are any one of the amino acids selected from G, A, V, L, M, I, S, T, P, N, Q, F, Y, W, K, R, H, D, E, C;

[0088] 3) For split inteins, use the flank sequences a1 to an and b1 - bn designed in 2) to design the expression sequences of Components A1 to An and Components B1 to Bn containing the sequences split from protein P;

[0089] 4) Connect the expression sequences to vectors respectively, perform co-transfection of Components A and B in one-to-one correspondence and in-cell trans-splicing to obtain splicing products F1 to Fn;

[0090] 5) Detect the splicing products F1 to Fn, and select the flank sequence pairs with a splicing efficiency exceeding 20%;

[0091] 6) Analyze the selected flanking sequence pairs in 5), and eliminate the flanking sequences in the flanking sequences that can cause free sulfhydryl groups to be generated after splicing, so as to optimize the flanking sequence pairs selected in 5);

[0092] 7) Repeat the steps 1) to 5) above, and select flanking sequence pairs 1 to m whose splicing efficiency is in the top 20% among all candidate sequence pairs, and there are no free sulfhydryl groups in the recombinant polypeptide as the splicing product,

[0093] where n is 2 or 3, and m is a positive integer.

[0094] 25. The method for screening the flanking sequence pairs of the split intein according to 24 above, the method further includes:

[0095] 1) Split a protein R different from protein P;

[0096] 2) Using the flanking sequence pairs 1 to m, design the expression sequences of components A'1 to A'm and components B'1 to B'm;

[0097] 3) Connect the expression sequences with a vector, perform transfection, expression and purification to obtain components A'1 to A'm and components B'1 to B'm,

[0098] 4) Respectively perform in vitro trans-splicing on components A1 to Am and components B'1 to B'm obtained using the flanking sequence pairs 1 to m, and / or components A'1 to A'm and components B1 to Bm in a one-to-one correspondence, detect the spliced product protein, and select multiple flanking sequence pairs with a splicing efficiency exceeding 50%.

[0099] 26. A method for preparing a recombinant polypeptide, characterized in that it uses the flanking sequence pairs for split inteins described in 1 or 2 above for trans-splicing.

[0100] 27. The use of the flanking sequence pairs for split inteins described in 1 or 2 above, characterized in that it is used for preparing a recombinant polypeptide, preferably for performing trans-splicing together with a split intein.

[0101] The advantages of the recombinant polypeptide (such as bispecific antibody) mediated by the flanking sequence pairs for split inteins of the present invention include (1) no free sulfhydryl groups; (2) high-throughput and high-efficiency preparation; (3) the target product and impurities are easy to distinguish and identify.

[0102] Definition

[0103] It should be noted that an indefinite quantity of an entity should refer to one or more of that entity; for example, "bispecific antibody" should be understood to represent one or more bispecific antibodies. Similarly, the terms "one or more" and "at least one" with an indefinite quantity limitation can be used interchangeably herein.

[0104] As used herein, the term "polypeptide" includes both the singular "polypeptide" and the plural "polypeptides", and also refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). Polypeptides can be derived from natural biological sources or produced by recombinant techniques, not necessarily translated from a specified nucleic acid sequence, and can be produced in any manner, including chemical synthesis.

[0105] As used herein, the term "recombinant", when referring to a polypeptide or polynucleotide, refers to a form of polypeptide or polynucleotide that does not exist in nature, and a non-limiting example of which can be achieved by combining polynucleotides or polypeptides that do not normally occur together.

[0106] "Homology" or "identity" or "similarity" refers to the degree of sequence similarity between two peptide chain molecules or between two nucleic acid molecules. When there are the same bases or amino acids at the positions in the sequences being compared, the molecules at those positions are homologous. The degree of homology among multiple sequences is a function of the number of paired or homologous sites shared by these sequences. An "unrelated" or "non-homologous" sequence has less than 40% homology, but preferably less than 25% homology, with one of the sequences of the present invention.

[0107] A polynucleotide or polynucleotide region (or polypeptide or polypeptide region) has a certain percentage (e.g., 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99%) of "sequence identity" with another sequence means that when aligned, that percentage of bases (or amino acids) is the same when comparing these two sequences.

[0108] Biologically equivalent polynucleotides are polynucleotides having the above-mentioned specific percentage of homology and encoding polypeptides with the same or similar biological activities.

[0109] The term "split intein" refers to a split intein composed of an N-terminal protein splicing region or N-terminal fragment (In, N’ fragment of intein) and a C-terminal protein splicing region or C-terminal fragment (Ic, C fragment of intein), and the gene expressing the precursor protein is split in two open reading frames, and the cleavage site is within the intein sequence.

[0110] The "N-terminal precursor protein" refers to a fusion protein formed by translation of a fusion gene formed by the gene of the N-terminal extein (En) and the N-terminal fragment (In) of the split intein.

[0111] "C-terminal precursor protein" refers to a fusion protein produced after translation of a fusion gene formed by the expression genes of the C-terminal fragment (Ic) of a split intein and the C-terminal exon (Ec).

[0112] The N-terminal fragment (In) or C-terminal fragment (Ic) of a single split intein does not have protein splicing function. After protein translation, In in the N-terminal precursor protein and Ic of the C-terminal precursor protein bind non-covalently through mutual recognition to form a functional intein, which can catalyze the protein trans-splicing reaction, thereby connecting two separate protein exons with a peptide bond (the N-terminal protein exon or N-terminal exon is called En, and the C-terminal protein exon or C-terminal exon is called Ec) (Ozawa. T. Nat Biotechbol. 21(2003)28793).

[0113] Protein trans-splicing refers to a protein splicing reaction mediated by a split intein. During trans-splicing, first, the N-terminal fragment (In) and C-terminal fragment (Ic) of the split intein recognize each other and bind non-covalently. Once bound, its structure folds correctly, and the split intein at this time has a reconstructed active center. Then, the protein splicing reaction is completed according to the typical protein splicing pathway, thereby connecting the exons on both sides.

[0114] In refers to the N-terminal part of a single split intein, and is also called the N-terminal fragment of the split intein or the N-terminal protein splicing region in this article.

[0115] Ic refers to the C-terminal part of a single split intein, and is also called the C-terminal fragment of the split intein or the C-terminal protein splicing region in this article.

[0116] Flanking sequence a is an amino acid sequence that flanks the N-terminal of In and the C-terminal of En, connecting In and En. Here, as Figure 5 shown, the first amino acid adjacent to the N-terminal of In is defined as position -1, the second amino acid residue further towards the N-terminal is position -2, the third amino acid residue is position -3, and so on until En. Generally speaking, the core sequence of flanking sequence a is positions -1 and -2, which is directly related to the splicing efficiency.

[0117] Flanking sequence b is an amino acid sequence that flanks the C-terminal of Ic and the N-terminal of Ec, connecting Ic and Ec. Here, as Figure 5As shown, the first amino acid residue adjacent to the C-terminus of Ic is defined as the +1 position, the second amino acid residue towards the C-terminus is the +2 position, the third amino acid residue is the +3 position, and so on until Ec. Generally speaking, the core sequence of the flanking sequence b is the +1 position and the +2 position, which is directly related to the splicing efficiency.

[0118] Trans-splicing mediated by split inteins, for example as Figure 5 shown, In is separated from the flanking sequence a, Ic is separated from the flanking sequence b, and the flanking sequence a and the flanking sequence b are connected. Thus, En and Ec connected to the flanking sequences are connected, so that the -1 amino acid residue of the flanking sequence a is directly peptide-bonded to the +1 amino acid residue of the flanking sequence b, and the -1 amino acid is at the N-terminus of the +1 amino acid.

[0119] The 20 common amino acids used in the present invention for screening flanking sequences (hereinafter referred to as 20 amino acids) refer to: glycine (G), alanine (A), valine (V), leucine (L), methionine (M), isoleucine (I), serine (S), threonine (T), proline (P), asparagine (N), glutamine (Q), phenylalanine (F), tyrosine (Y), tryptophan (W), lysine (K), arginine (R), histidine (H), aspartic acid (D), glutamic acid (E) and cysteine (C).

[0120] As used herein, "antibody" or "antigen-binding polypeptide" refers to a polypeptide or polypeptide complex that specifically recognizes and binds to an antigen or immunogenic epitope.

[0121] The antibody can be a complete antibody or any antigen-binding fragment or its single chain. Thus, the term "antibody" includes any protein or peptide containing a specific molecule that contains at least a portion of an immunoglobulin molecule that has the biological activity of binding to an antigen or immunogenic epitope. Examples of such cases include, but are not limited to, the complementarity-determining regions (CDRs) of the heavy or light chain or their ligand-binding portions, the variable regions of the heavy or light chain, the constant regions of the heavy or light chain, the framework (FR) regions or any part thereof, or at least a portion of the binding protein.

[0122] As used herein, the term "antibody fragment" or "antigen-binding fragment" is a part of an antibody. The term "antibody fragment" includes aptamers, aptamer enantiomers (spiegelmers) and diabodies, and also includes any synthetic or genetically engineered protein that, like an antibody, can bind to a specific antigen or immunogenic epitope to form a complex.

[0123] "Single-chain variable fragment" or "scFv" refers to a fusion protein of the variable regions of the heavy chain (VH) and light chain (VL) of an immunoglobulin.

[0124] The term "antibody" includes a broad class of polypeptides that can be biochemically recognized. Those skilled in the art will understand that heavy chains are divided into γ, μ, α, δ, ε and have several subclasses (e.g., γ1-4). The nature of this chain determines the "class" of the antibody, such as IgG, IgM, IgA, IgD or IgE. Immunoglobulin subclasses (isotypes) such as IgG1, IgG2, IgG3, IgG4, IgG5, etc. are well characterized and are functionally specific. Those skilled in the art can easily identify each modified form of these classes and isotypes by referring to the present application, and thus, these forms are within the scope of the present application.

[0125] All immunoglobulin classes are clearly within the scope of the present application, and the following discussion will generally be directed to the IgG class of immunoglobulin molecules.

[0126] Regarding IgG, a standard immunoglobulin molecule comprises two identical light chain polypeptides (molecular weight approximately 23,000 daltons) and two identical heavy chain polypeptides (molecular weight approximately 53,000 - 70,000 daltons) that are joined together in a "Y" shape by disulfide bonds.

[0127] The antibodies, antigen-binding polypeptides, their variants or derivatives of the present application include, but are not limited to: polyclonal antibodies, monoclonal antibodies, multispecific antibodies, human antibodies, humanized antibodies, primatized antibodies, or chimeric antibodies, single-chain antibodies, epitope-binding fragments, e.g., Fab, Fab′ and F(ab′)2, Fd, Fvs, single-chain Fvs (scFv), single-chain antibodies, disulfide-linked Fvs (sdFv), fragments containing VL domains or VH domains, fragments produced by a Fab expression library, and anti-idiotypic (Anti-Id) antibodies. The immunoglobulin molecules or antibody molecules of the present application can be of any type (e.g., IgG, IgE, IgM, IgD, IgA and IgY), any class (e.g., IgG1, IgG2, IgG3, IgG4, IgA1 and IgA2) or subclass of immunoglobulin molecules.

[0128] In some instances, e.g., certain immunoglobulins derived from camelids or based on camel immunoglobulins, the complete immunoglobulin molecule can consist only of heavy chains and no light chains. See, e.g., Hamers-Casterman et al., Nature. 363:446-448 (1993).

[0129] Both the light chain and the heavy chain are divided into structural regions and functionally homologous regions. The terms "constant" and "variable" are for functional use. Here, it should be recognized that the light chain variable domain (VL) and the heavy chain variable domain (VH) together determine antigen recognition and specificity. Generally, the number of constant region domains increases with the position away from the antigen-binding site or the amino-terminal end of the antibody. The N-terminal portion is the variable region, while the C-terminal portion is the constant region; the CH3 and CL domains actually contain the carboxyl termini of the heavy chain and the light chain, respectively.

[0130] The antigen-binding site refers to: for any given heavy or light chain variable region, those skilled in the art can readily identify the amino acids that include the CDRs and framework regions, respectively, as they have been clearly defined (see, "Sequences of Proteins of Immunological Interest," Kabat, E., et al., U.S. Department of Health and Human Services, (1983); Chothia and Lesk, J. Mol. Biol., 196: 901-917 (1987), which is hereby incorporated by reference in its entirety into this article).

[0131] When, within the art, a term has two or more definitions and is used and / or acceptable, the definition of the term used herein is intended to include all meanings, unless stated to the contrary expressly.

[0132] The term "complementary determining region" ("CDR") describes the non-contiguous antigen-binding sites present in the variable regions of the heavy and light chain polypeptides. Such specific regions are described by Kabat et al. in U.S. Department of Health and Human Services, "Sequences of Proteins of Immunological Interest" (1983) and by Chothia et al. in J Mol Biol 196: 901-917 (1987), which is incorporated by reference in its entirety into this article. Given the amino acid sequence of the variable region of the antibody, those skilled in the art can generally determine which residues contain a particular CDR.

[0133] As used herein, "Kabat numbering" refers to the numbering system described by Kabat et al., the content of which is recorded in U.S. Department of Health and Human Services, "Sequence of Proteins of Immunological Interest" (1983).

[0134] As used herein, the term "heavy chain constant region" includes the amino acid sequence from an immunoglobulin heavy chain. A polypeptide comprising a heavy chain constant region comprises at least one of the following: a CH1 domain, a hinge (e.g., upper hinge region, middle hinge region, and / or lower hinge region) domain, a CH2 domain, a CH3 domain, or a variant or fragment thereof. For example, an antigen-binding polypeptide used in the present application may comprise a polypeptide chain having a CH1 domain; a polypeptide having a CH1 domain, at least a portion of the hinge domain, and a CH2 domain; a polypeptide chain having a CH1 domain and a CH3 domain; a polypeptide chain having a CH1 domain, at least a portion of the hinge domain, and a CH3 domain, or a polypeptide chain having a CH1 domain, at least a portion of the hinge structure, a CH2 domain, and a CH3 domain. In another embodiment, the polypeptide of the present application comprises a polypeptide chain having a CH3 domain. Additionally, an antibody used in the present application may lack at least a portion of the CH2 domain (e.g., all or a portion of the CH2 domain). As described above, those of ordinary skill in the art will understand that the heavy chain constant regions may be modified such that they differ in amino acid sequence from naturally occurring immunoglobulin molecules.

[0135] The heavy chain constant region of an antibody disclosed herein may be from different immunoglobulin molecules. For example, the heavy chain constant region of a polypeptide may comprise a CH1 domain from an IgG1 molecule and a hinge region from an IgG3 molecule. In another example, the heavy chain constant region may comprise a hinge region that is partially from an IgG1 molecule and partially from an IgG3 molecule. In another example, the heavy chain portion may comprise a chimeric hinge that is part from an IgG1 molecule and part from an IgG4 molecule.

[0136] The term "light chain constant region" includes the amino acid sequence from an antibody light chain. Preferably, the light chain constant region comprises at least one of a constant kappa domain and a constant lambda domain.

[0137] The term "VH domain" includes the amino-terminal variable domain of an immunoglobulin heavy chain, while the term "CH1 domain" includes the first (most often amino-terminal) constant region of an immunoglobulin heavy chain. The CH1 domain is adjacent to the VH domain and is the amino-terminal of the hinge region of an immunoglobulin heavy chain molecule.

[0138] The term "CH2 domain" includes a portion of the heavy chain molecule, which portion ranges, for example, from about residue 244 to residue 360 of an antibody, using a conventional numbering scheme (residues 244 to 360, Kabat numbering system; and residues 231 - 340, EU numbering system; see Kabat et al., U.S. Department of Health and Human Services, "Sequences of Proteins of Immunological Interest" (1983)). The CH2 domain is unique in that it pairs loosely with another domain. Instead, two N-linked branched sugar chains insert between the two CH2 domains of the intact native IgG molecule. It is documented that the CH3 domain extends from the CH2 domain to the C-terminus of the IgG molecule and contains approximately 108 residues.

[0139] As used herein, the terms "specifically bind" or "specific for" generally mean that when an antibody binds to an epitope, the binding through the antigen-binding domain is more favorable compared to binding to a random, unrelated epitope. The term "specificity" is used herein to determine the affinity of a particular antibody for a specific epitope.

[0140] As used herein, the term "treatment" ("treat" or "treatment") refers to therapeutic treatment and prophylactic or preventive measures, wherein a subject is treated to prevent or slow down (alleviate) an adverse physiological change or disorder, such as the development of cancer. Beneficial or desired clinical outcomes include, but are not limited to, alleviation of symptoms, reduction of the extent of the disorder, stabilization (e.g., preventing it from worsening) of the state of the disorder, delay or slowing of disorder progression, improvement or palliation of the state of the disorder, and remission (whether partial or complete), whether or not detectable. "Treatment" can also refer to prolonging survival as compared to the expected survival if not receiving treatment.

[0141] Any of the above antibodies or polypeptides may also include additional polypeptides, for example, an encoded polypeptide as described herein, a signal peptide at the N-terminus of the antibody for directing secretion, or other heterologous polypeptides as described herein.

[0142] In other embodiments, the polypeptides of the present application may contain conservative amino acid substitutions.

[0143] "Conservative amino acid substitution" refers to the replacement of an amino acid residue with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art and include basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), β-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Accordingly, non-essential amino acid residues of an immunoglobulin polypeptide are preferably replaced with other amino acid residues from the same side chain family. In another embodiment, a string of amino acids may be replaced with a structurally similar string of amino acids that differs in sequence and / or in the composition of the side chain family.

[0144] Transient transfection: Transient transfection is one of the ways to introduce DNA into eukaryotic cells. In transient transfection, recombinant DNA is introduced into a highly infectious cell line to obtain transient but high-level expression of the target gene. The transfected DNA does not have to integrate into the host chromosome, and the transfected cells can be harvested in a shorter time than stable transfection, and the target product in the expression supernatant can be detected. BRIEF DESCRIPTION OF THE DRAWINGS

[0145] Figure 1 is a schematic diagram of split intein-mediated homologous polypeptide fragment splicing (A) and a schematic diagram of the primary protein structure of each component (B).

[0146] Figure 2 is a schematic diagram of split intein-mediated heterologous polypeptide fragment splicing (A) and a schematic diagram of the primary protein structure of each component (B).

[0147] Figure 3 is a schematic diagram of split intein-mediated in vitro antibody splicing (A) and a schematic diagram of the primary protein structure of each component (B), and the splicing product is a bispecific antibody. (C) is an exemplary schematic diagram of the amino acid sequence near the splicing site of split intein-mediated antibody splicing, and "X" indicates that the amino acid at this position is any amino acid or deletion.

[0148] Figure 4 is a schematic diagram of the construction of the expression plasmid of component A of the bispecific antibody (A) and a schematic diagram of the construction of the expression plasmid of component B (B).

[0149] Figure 5 Schematic diagram of the flanking sequence numbering.

[0150] Figure 6 is a Western blot detection of the expression supernatant of 293E co-transfected with expression plasmids of amino acid combinations with different positions at -1 of flanking sequence a and +1 of flanking sequence b of intein NpuDnaE.

[0151] Figure 7 shows the Western blot detection of the expression supernatant of 293E cells co-transfected with expression plasmids containing different amino acid combinations at the -1 position (G, V, or A) of the flanking sequence a of intein NpuDnaE, at the +1 position (S) of the flanking sequence b, and different amino acids at the +2 position.

[0152] Figure 8 shows the results of reducing SDS-PAGE and Coomassie Brilliant Blue staining of the expression supernatant of 293E cells co-transfected with an expression plasmid containing different amino acids at the -2 position and G at the -1 position of the flanking sequence a of intein NpuDnaE, and S at the +1 position and E at the +2 position of the flanking sequence b, after affinity purification with protein A.

[0153] Figure 9 Figure 8 also shows the Western blot detection results of the expression supernatant of 293E cells co-transfected with an expression plasmid containing different amino acids at the -1 position, G at the -2 position of the flanking sequence a of intein NpuDnaE, and S at the +1 position and E at the +2 position of the flanking sequence b.

[0154] Figure 10 Figure 8 further shows the Western blot detection results of the expression supernatant of 293E cells co-transfected with an expression plasmid containing G at the -1 position of the flanking sequence a of intein NpuDnaE, and different amino acids at the +3 position, as well as S at the +1 position and E at the +2 position of the flanking sequence b.

[0155] Figure 11 Figure 16 shows the results of reducing SDS-PAGE and Coomassie Brilliant Blue staining of the expression splicing products A1, A10, and A61 of 293E cells co-transfected with expression plasmids containing the corresponding components A and B of intein NpuDnaE, after affinity purification with protein A. Among them, A1 and A10 are positive controls, and the corresponding flanking sequence pairs are A1 - MGG and SVY, A10 - GS and CFN, and the flanking sequence pair of A61 is GK and SEI.

[0156] Figure 12 shows the results of non-reducing SDS-PAGE and Coomassie Brilliant Blue staining of the purified products of component A and component B' with different inteins expressed in 293E cells respectively; (A) Detection of component A, namely Fab4; (B) Detection of component B', namely HAb4; E1, E2, and E3 are products harvested under different elution conditions.

[0157] Figure 13Non-reducing SDS-PAGE and Coomassie Brilliant Blue staining detection of the splicing products of component A and component B' of different inteins. The intein is NpuDnaE, the flanking sequence a is SG, the flanking sequence b is SEI. "Splicing 1" means the concentrations of component A and component B' are 10 uM and 1 uM respectively, and the reaction system contains 2 mM DTT. "Splicing 2" means the concentrations of component A and component B' are 5 uM and 1 uM respectively, and the reaction system contains 2 mM DTT. "No splicing 1" means the concentrations of component A and component B' are 10 uM and 1 uM respectively, and the reaction system does not contain DTT. "No splicing 2" means the concentrations of component A and component B' are 5 uM and 1 uM respectively, and the reaction system does not contain DTT. The control bands are component A as Fab4 (non-reducing), component B' as HAb4 (non-reducing), and monoclonal antibody. Both "Splicing 1" and "Splicing 2" are incubated overnight at 37 °C, and the other groups are stored at 4 °C.

[0158] Figure 14 . Double antigen sandwich ELISA detection results of the splicing products of the intein NpuDnaE with flanking sequence a being SG and flanking sequence b being SEI. Among them, the coating antigen is CD38, and the detection antigen is PD-L1 labeled with horseradish peroxidase (HRP). Detailed implementation mode

[0159] The present invention relates to a preparation method for obtaining bispecific antibodies, which includes: splitting the DNA sequence corresponding to the target antibody, constructing a mammalian cell expression vector through total gene synthesis, purifying the vector, and transiently or stably transfecting mammalian cells such as HEK293 or CHO with the purified vector. The fermentation broth is collected separately, and component A and component B are purified by methods such as protein A, protein L, nickel column, Strep-Tactin affinity chromatography, anti-Flag antibody affinity chromatography, anti-HA antibody affinity chromatography, or cross-linked starch affinity chromatography. The purified component A and component B are subjected to in vitro trans-splicing, and the splicing products are subjected to affinity chromatography corresponding to the tag protein such as nickel column to obtain high-purity bispecific antibodies. The process flow is as Figure 3A shown.

[0160] The antibodies described herein can be from any animal source, including birds and mammals. Preferably, the antibodies are human, murine, donkey, rabbit, goat, guinea pig, camel, llama, horse, or chicken antibodies. In another embodiment, the variable regions can be from chondrichthyes (e.g., from sharks).

[0161] In some embodiments, the antibodies can bind to: therapeutic agents, prodrugs, peptides, proteins, enzymes, viruses, lipids, biological response modifiers, pharmaceuticals, or PEG.

[0162] The antibody can be linked to or fused with a therapeutic agent, which may include a detectable label such as a radiolabel, an immunomodulator, a hormone, an enzyme, an oligonucleotide, a photoactive therapeutic agent or diagnostic agent, a cytotoxic agent, which may be: a drug or toxin, an ultrasound enhancer, a non-radiolabel, combinations thereof and other such components known in the art.

[0163] The antibody is detectably labeled by conjugating it to a chemiluminescent compound. Then, the presence of the chemiluminescent-labeled antigen-binding polypeptide is determined by detecting the luminescence generated during the chemical reaction. Examples of particularly useful chemiluminescent-labeled compounds are luminol, isoluminol, thermatic acridinium esters, imidazoles, acridinium salts and oxalates.

[0164] The antibody can also be detectably labeled using a fluorescent-emitting metal such as 152Eu, or other lanthanide labels. These metals can be linked to the antibody using a metal chelating group such as diethylenetriaminepentaacetic acid (DTPN) or ethylenediaminetetraacetic acid (EDTA).

[0165] The binding specificity of the antigen-binding polypeptide of the present application can be measured by in vitro experiments such as: immunoprecipitation, radioimmunoassay (RIA) or enzyme-linked immunosorbent assay (ELISA).

[0166] Cell lines for producing recombinant polypeptides can be selected and cultured using techniques well known to those skilled in the art.

[0167] To introduce mutations in the nucleotide sequence encoding the antibody of the present application, standard techniques well known to those skilled in the art can be used, including but not limited to: site-directed mutagenesis to generate amino acid substitutions and PCR-mediated mutagenesis. Preferably, the variant (including derivatives), relative to the reference variable heavy chain region, CDR-H1, CDR-H2, CDR-H3, light chain variable region, CDR-L1, CDR-L2 or CDR-L3, encodes fewer than 50 amino acid substitutions, fewer than 40 amino acid substitutions, fewer than 30 amino acid substitutions, fewer than 25 amino acid substitutions, fewer than 20 amino acid substitutions, fewer than 15 amino acid substitutions, fewer than 10 amino acid substitutions, fewer than 5 amino acid substitutions, fewer than 4 amino acid substitutions, fewer than 3 amino acid substitutions, or fewer than 2 amino acid substitutions. Alternatively, mutations can be randomly introduced along all or part of the coding sequence, for example, by saturation mutagenesis the resulting mutants can be screened for biological activity to identify mutants that retain activity.

[0168] The tag proteins used in the present invention can be Fc, oligohistidine (His-tag), Strep-tag, Flag, HA or maltose-binding protein (MBP), etc.

[0169] The transfection used in the present invention can be transient transfection or stable transfection.

[0170] Mammalian cells such as HEK293 or CHO are used in the present invention, but are not limited thereto.

[0171] Liquids containing expression products from mammalian cells, such as fermentation broth and culture supernatant, can be purified by methods such as protein A, protein G, nickel column, Strep-Tactin affinity chromatography, anti-Flag antibody affinity chromatography, anti-HA antibody affinity chromatography, or cross-linked starch affinity chromatography.

[0172] The product obtained by splicing can be subjected to affinity chromatography corresponding to the tag protein to remove unspliced components.

[0173] The gene fragment used in the present invention for constructing the vector can be constructed by total gene synthesis, but is not limited thereto.

[0174] The vectors used in the present invention are pcDNA3.1 or pCHO1.0, but are not limited thereto.

[0175] The restriction endonucleases used in the present invention can include, for example, NotI, NruI, or BamHI-HF, etc., but are not limited thereto.

[0176] BLAST is an alignment program using default parameters. Specifically, the programs are BLASTN and BLASTP. Details of these programs can be obtained at the following Internet addresses: http: / / www.ncbi.nlm.nih.gov / blast / Blast.cgi .

[0177] In a specific embodiment of the present invention, as shown in Figures 1, 2, and 3, a component A expression plasmid (pPa-FSa-In-Tag) and a component B expression plasmid (pTag-Ic-FSb-Pb), or a component A' expression plasmid (pRa-FSa-In-Tag) and a component B' expression plasmid (pTag-Ic-FSb-Rb) can be constructed.

[0178] In another specific embodiment of the present invention, as Figure 4A shown in A and B, Pa-HIn and Pa-L can be constructed into the same plasmid, namely the component A expression plasmid (pBi-Pa-FSa-In-Tag), by molecular cloning methods such as restriction digestion and ligation; or pB'-L, pB'-H, and pB'-FcIc can be constructed into the same plasmid, namely the component B' expression plasmid (pBi-Tag-Ic-FSb-Rb).

[0179] In another specific embodiment of the present invention, the component B expression plasmid can include three types of expression plasmids: pB-L, pB-H, and pB-FcIc.

[0180] In the present invention, Pa is also used to represent the N-terminal protein exon or N-terminal exopeptide of protein P, and is also denoted as Enp; Pb is also used to represent the C-terminal protein exon or C-terminal exopeptide of protein P, and is also denoted as Ecp. Ra is also used to represent the N-terminal protein exon or N-terminal exopeptide of protein R, and is also denoted as En R ; Rb is also used to represent the C-terminal protein exon or C-terminal exopeptide of protein R, and is also denoted as Ec R .

[0181]

[0182]

[0183]

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200] Table 31 shows the amino acid sequence of partial component A containing the NpuDnaE intein.

[0201]

[0202]

[0203] Note: The variable region of the heavy chain of the antibody in Component A is denoted as VHa; the domain sequences such as VHa, CH1, flanking sequence a, and tag protein in the table can be replaced with the protein sequences of other corresponding domains mentioned in this specification.

[0204] Table 32 Amino acid sequences of partial Component B containing the NpuDnaE intein

[0205]

[0206]

[0207] Note: The domain sequences such as Pb, flanking sequence b, and tag protein in the table can be replaced with the protein sequences of other corresponding domains mentioned in this specification.

[0208] Table 33 Component B' containing the NpuDnaE intein

[0209]

[0210]

[0211]

[0212] Note: The domain sequences such as VHa, CH1, flanking sequence a, and tag protein in the table can be replaced with the protein sequences of other corresponding domains mentioned in this specification.

[0213] Examples

[0214] Test methods

[0215] 1. Preparation of recombinant polypeptides

[0216] The DNA sequences in the examples of the present invention were all obtained by reverse translation according to the amino acid sequences and synthesized by Wuhan KingCare.

[0217] The preparation of the recombinant polypeptides involved in the examples was all carried out by the following method: The DNA sequence was ligated with the vector pcDNA3.1 digested with the restriction enzyme EcoRI at 37°C for 30 minutes under the action of a recombinase, and then transformed into Trans10 competent cells by heat shock. After verification by sequencing (Wuhan KingCare), it was transiently transfected into 293E cells (purchased from Thermo Fisher). After expression, purification was carried out.

[0218] 2. The plasmid DNAs involved in the co-transfection in the examples are specifically as follows:

[0219] 1) To express Component A and Component B shown in Figure 1, plasmid pPa-FSa-In-Tag and pTag-Ic-FSb-Pb need to be transfected or co-transfected into 293E cells for expression.

[0220] 2) To express Component A and Component B’ shown in Figure 2, plasmid pPa-FSa-In-Tag and pTag-Ic-FSb-Rb need to be transfected or co-transfected into 293E cells for expression.

[0221] 3) To express Component A shown in Figure 3, plasmid Pa-HIn and Pa-L need to be co-transfected into 293E cells for expression, or plasmid pBi-Pa-FSa-In-Tag can be transfected alone for expression; to express Component B’ shown in Figure 3, plasmid pB’-L, pB’-H and pB’-FcIc need to be co-transfected into 293E cells for expression, or plasmid pBi-Tag-Ic-FSb-Rb can be transfected alone for expression.

[0222] Generally, if two plasmids are co-transfected for expression, the molar ratio of the two plasmids can be 1:1 or any other ratio. If three plasmids are co-transfected for expression, the molar ratio of the three plasmids can be 1:1:1 or any other ratio.

[0223] 3. Purification of Polypeptides with Tag Proteins

[0224] (1) When the tag protein is Fc, affinity chromatography is used, with MabSelect SuRe (GE, catalog number 17-5438-01), 18 ml column.

[0225] (2) When the tag protein is His-tag, affinity chromatography is used, with Ni-NTA (Jiangsu Qianchun, catalog number: A41002-06).

[0226] (3) When the tag protein is Strep-tag, Flag, HA or MBP, etc., the corresponding fillers and buffers for Strep-Tactin affinity chromatography, anti-Flag antibody affinity chromatography, anti-HA antibody affinity chromatography, or cross-linked starch affinity chromatography can be selected respectively.

[0227] (4) Ion exchange chromatography: When Component A (A’) or Component B (B’) does not carry a tag protein, ion exchange chromatography can be used to separate the splicing products according to the difference in isoelectric points. The chromatography filler used can be a cation exchange chromatography filler or an anion exchange chromatography filler, such as Hitrap SP-HP (GE).

[0228] (5) Hydrophobic chromatography. When component A (A’) or component B (B’) is a protein without a tag, the splicing products can be separated by hydrophobic chromatography according to the difference in hydrophobicity. Chromatographic packing materials such as Capto phenyl ImpRes packing material (GE) can be used.

[0229] (6) Molecular sieve. When component A (A’) or component B (B’) is a protein without a tag, the splicing products can be separated by molecular sieve chromatography according to the difference in molecular weight. Chromatographic packing materials such as HiLoad Superdex 200pg (GE) can be used.

[0230] Example 1 Screening of flanking sequence pairs of intein NpuDnaE

[0231] ● Construction of expression plasmids pA-Hln, pA-L, and plasmid (pTag-Ic-FSb-Pb)

[0232] Specifically, in this example, an amino acid selected from one of the 19 amino acids (A, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y) defined in the present invention is denoted as the X amino acid. H refers to a part of the heavy chain of the antibody, and L refers to a part of the light chain of the antibody. For the flanking sequence pairs of the original NpuDnaE, each amino acid is mutated by degenerate primer design to one of the 19 amino acids described in the present invention.

[0233] For the A-HIn expression plasmid (pA-HIn) with mutated flanking sequence pairs, for the sake of simplicity, the plasmid with the X amino acid at the -1 position of flanking sequence a is denoted as pA-HIn(X), and the plasmid with the -2nd and -1st positions independently being the same or different X amino acids is denoted as pA-HIn(XX).

[0234] The sequences involving the X amino acid in pA-HIn(XX) are all obtained by degenerate primer design.

[0235] The expression plasmid of A-Fab1 contains two polypeptides A-HIn and A-L. Among them, the plasmid for polypeptide A-HIn is denoted as pA-HIn(1), and the plasmid for polypeptide A-L is denoted as pA-L; and so on.

[0236] For the expression plasmid of B-FcIc with a mutated flanking sequence pair, similarly, it is denoted as the expression plasmid pTag-Ic-FSb-(B-FcIc). Similarly, the plasmid in which the +1st, +2nd, and +3rd amino acids of the flanking sequence b of the expression plasmid pTag-Ic-FSb-(B-FcIc) are independently the same or different X amino acids is denoted as pTag-Ic-FSb(XXX)-(B-FcIc), and the sequences involving X amino acids are all obtained through degenerate primer design.

[0237] Using the steps and conditions in "Preparation of Recombinant Polypeptides", such as Figure 4A , 4B shown, use the pcDNA3.1 plasmid vector to construct the expression plasmids of the corresponding components respectively according to the compositions shown in Table 1-33.

[0238] ● Screening of the +1st amino acid

[0239] Group the expression plasmid pTag-Ic-FSb-(B-FcIc). Those with the same amino acid residue at the +1st position of the flanking sequence b are grouped into one group, and 19 groups are obtained. At this time, for each group, the amino acid X at the +1st position of its flanking sequence b is determined, and the +2nd and +3rd positions are X, that is, any one of the 19 amino acids.

[0240] Transfect 293E cells with the above 19 groups of pTag-Ic-FSb(XXX)-(B-FcIc) and the corresponding expression plasmids pA-HIn(XX) and pA-L respectively.

[0241] The transfection molar ratio is pTag-Ic-FSb(XXX)-(B-FcIc)∶pA-HIn(XX)∶pA-L = 3∶1∶1.

[0242] At the same time, set up a positive control group separately (the flanking sequence a is AEY, and the flanking sequence b is CFN). The plasmid of the positive control group, its co-transfected plasmid, and the molar ratio are pTag-Ic-FSb-(B-FcIc1)∶pA-HIn(1)∶pA-L = 3∶1∶1.

[0243] Take the supernatant after culturing the transfected cells for 5 days. Detect the protein in the supernatant by Western blot method. The molecular weight marker protein is PageRuler Prestained Protein Ladder, purchased from Thermo, catalog number 26616. The results show that at the position of the splicing product molecular weight of 50 kD, the corresponding band at this position is the complete heavy chain. An obvious band appears in the positive control group (intein + natural flanking sequence pair), and obvious bands appear only when the +1st residue of the flanking sequence b in other groups is serine (S).

[0244] The results show that for NpuDnaE, when the amino acid at the +1 position is S, it has a relatively high cleavage efficiency.

[0245] ● Screening of the amino acid at the -1 position and the amino acid at the +2 position

[0246] The expression plasmid pTag-Ic-FSb(SXX)-(B-FcIc) of the above serine (S) group was further grouped. Those with the same amino acid residue at the +2 position of the flanking sequence b were grouped into one group, and 19 subgroups were obtained. At this time, for each subgroup, the amino acid residue at the +1 position of its flanking sequence b is S, and the amino acid X at the +2 position is determined.

[0247] The expression plasmid pA-HIn(XX) was grouped. Those with the same amino acid residue at the -1 position of the flanking sequence a were grouped into one group, and 19 groups were obtained. At this time, for each group, the amino acid residue at the +1 position of its flanking sequence a is determined.

[0248] The expression plasmids pTag-Ic-FSb(SXX)-(B-FcIc) of the above 19 subgroups, the above 19 expression plasmids pA-HIn(XX), and the plasmid pA-L were co-transfected into 293E cells respectively.

[0249] Primary screening

[0250] In order to reduce the number of experiments in the cross-pairing test, the above 19 expression plasmids pA-HIn(XX) were divided into groups A1 - A6 according to Table 34 below, and the above 19 expression plasmids pTag-Ic-FSb(SXX)-(B-FcIc) were divided into groups B1 - B6 according to Table 35 below.

[0251] Table 34 Grouping of the expression plasmid pA-HIn(X)

[0252]

[0253] Table 35 Grouping of the expression plasmid pTag-Ic-FSb(SXX)-(B-FcIc)

[0254]

[0255] Transfection was carried out in the same way as the pairing in Table 36. The transfection conditions were: the molar ratio of plasmids was pTag-Ic-FSb(SXX)-(B-FcIc)∶pA-HIn(XX)∶pA-L = 3∶1∶1. And positive controls were set up in the same way as above, and 36 groups of transfections from 1-1 to 6-6 were obtained respectively.

[0256] Table 36 Grouping of co-transfection

[0257] Number B1 B2 B3 B4 B5 B6 A1 1-1 1-2 1-3 1-4 1-5 1-6 A2 2-1 2-2 2-3 2-4 2-5 2-6 A3 3-1 3-2 3-3 3-4 3-5 3-6 A4 4-1 4-2 4-3 4-4 4-5 4-6 A5 5-1 5-2 5-3 5-4 5-5 5-6 A6 6-1 6-2 6-3 6-4 6-5 6-6

[0258] The supernatant was collected 5 days after culturing the transfected cells. The protein in the supernatant was detected by Western blot (SDS-PAGE with the reducing agent β-mercaptoethanol, and the detection antibody was HRP-labeled goat anti-human IGG antibody, purchased from Sigma). The results are shown in Figures 6A - 6F . According to the results, at the position of the molecular weight of 50 kD of the intact heavy chain protein of the cleavage product, obvious bands appeared in the positive control group. Among the 36 transfection groups, transfection groups 1-1, 1-4, 1-5, and 1-6 had obvious bands, especially the band in group 1-6 was the most significant, indicating that efficient splicing occurred in this group. The corresponding -1 amino acid residues in this group were: A, V, G, and the corresponding +2 amino acid residues were: R, K, E, D.

[0259] The results show that for NpuDnaE, when the -1 amino acid residue is: A, V, G, and the corresponding +2 amino acid residues are: R, K, E, D, it has a relatively high splicing efficiency.

[0260] All plasmids in group 1-6 were selected for rescreening.

[0261] Rescreening

[0262] All plasmids with the corresponding -1 amino acid residues of: A, V, G and the corresponding +2 amino acid residues of: R, K, E, D in groups 1-6 were paired according to Table 37 and co-transfected.

[0263] Table 37 Grouping of co-transfection for rescreening

[0264]

[0265] Transfection was carried out in the same way as the pairing in Table 37. The transfection conditions were: the molar ratio of plasmids was pTag-Ic-FSb (SRX or SKX or SEX or SDX)-(B-FcIc)∶pA-HIn (XA or XV or XG)∶pA-L = 3∶1∶1. And a positive control was set up in the same way as above.

[0266] The supernatant was collected 5 days after culturing the transfected cells. The protein in the supernatant was detected by Western blot (SDS-PAGE with the reducing agent). The results are shown in Figures 7A - 7C . According to the results, the G-SE group had the most significant 50 kD band. It indicates that efficient splicing occurred in this group.

[0267] The results show that for NpuDnaE, when the +1 amino acid is S, the +2 amino acid is E, and the -1 amino acid is G, it has a relatively high splicing efficiency.

[0268] Therefore, all the plasmids in the pTag-Ic-FSb(SEX)-(B-FcIc):pA-HIn(XG) group were selected for the screening of the -2nd and +3rd amino acids as follows.

[0269] ● Screening of the -2nd amino acid

[0270] The group represented by pA-HIn(XG) was divided into 19 categories according to the different -2nd amino acid X, and these 19 categories were co-transfected with all the plasmids in the group represented by the above pTag-Ic-FSb(SEX)-(B-FcIc) and the plasmid pA-L for primary screening.

[0271] Table 38 Grouping for primary screening of the -2nd amino acid

[0272] Number The - 2nd position of flanking sequence a AG - SE A DG - SE D EG - SE E FG - SE F GG - SE G HG - SE H IG - SE I KG - SE K LG - SE L MG - SE M NG - SE N PG - SE P QG - SE Q RG - SE R SG - SE S TG - SE T VG - SE V WG - SE W YG - SE Y

[0273] Transfection was carried out as described above according to the pairing in Table 38, and 19 groups of transfection, namely AG-SE to YG-SE, were obtained respectively. The transfection conditions were: the molar ratio of plasmids was pTag-Ic-FSb(SEX)-(B-FcIc)∶pA-HIn(XG)∶pA-L = 3∶1∶1. And a positive control was set up in the same way as above.

[0274] After culturing the transfected cells for 5 days, the supernatant was taken. The protein in the supernatant was detected by Western blot method (SDS-PAGE plus reducing agent). The results are shown in Figures 8A - 8C . According to the results, splicing occurred in all 19 groups, and those with relatively high splicing efficiency were DG-SE, FG-SE, LG-SE, NG-SE, GG-SE, SG-SE and WG-SE. Through comprehensive analysis, the groups with the highest efficiency were GG-SE and SG-SE.

[0275] The results show that for NpuDnaE, when the -1st amino acid is G and the -2nd amino acid is selected from: D, F, G, L, N, G, S, W, the splicing efficiency of such flanking sequences is relatively high. In particular, when the -1st amino acid is G and the -2nd amino acid is G or S, the splicing efficiency of the flanking sequence pairs is the highest.

[0276] ● Other schemes for the -1st and -2nd amino acids

[0277] Eight expression plasmids pA-HIn(GA or GE or GK or GQ or GR or GW or GT or GP) were co-transfected with the expression plasmid pTag-Ic-FSb(SEX)-(B-FcIc) and the plasmid pA-L respectively.

[0278] Table 39 Grouping for other screening of the -1st amino acid

[0279]

[0280]

[0281] Transfection was carried out in the same manner as above according to the pairing in Table 39. The transfection conditions were as follows: the molar ratio of plasmids was pTag-Ic-FSb(SEX)-(B-FcIc)∶pA-HIn(GA or GE or GK or GQ or GR or GW or GT or GP)∶pA-L = 3∶1∶1. A positive control was set up in the same manner as above.

[0282] The supernatant was collected 5 days after culturing the transfected cells. The proteins in the supernatant were detected by Western blot (SDS-PAGE with a reducing agent). The results are shown in Figure 9 . According to the results, significant splicing occurred in the transfection groups GA-SE, GK-SE, GS-SE, GQ-SE, GR-SE, GW-SE, and GT-SE.

[0283] The results show that for NpuDnaE, when the amino acid at the -2 position is G and the amino acid at the -1 position is selected from A, K, S, Q, R, W, and T, the splicing efficiency of such flanking sequences is relatively high.

[0284] ● Amino acid at the +3 position

[0285] The group represented by the expression plasmid pTag-Ic-FSb(SEX)-(B-FcIc) was divided into 19 categories according to the different X at the +3 position, and co-transfected with all the plasmids of the group represented by pA-HIn(GX) and the plasmid pA-L for primary screening.

[0286] Table 40 Primary screening grouping of amino acids at the +3 position

[0287] Number The +3rd position of flanking sequence b G - SEA A G - SED D G - SEE E G - SEF F G - SEG G G - SEH H G - SEI I G - SEK K G - SEL L G - SEM M G - SEN N G - SEP P G - SEQ Q G - SER R G - SES S G - SET T G - SEV V G - SEW W G - SEY Y

[0288] Transfection was carried out in the same manner as above according to the pairing in Table 40, and 19 groups of transfection, namely G-SEA to G-SEY, were obtained respectively. The transfection conditions were as follows: the molar ratio of plasmids was pTag-Ic-FSb(SEA~SEY)-(B-FcIc)∶pA-HIn(GX)∶pA-L = 3∶1∶1. A positive control was set up in the same manner as above.

[0289] The supernatant was collected 5 days after culturing the transfected cells. The proteins in the supernatant were detected by Western blot (SDS-PAGE with a reducing agent). The results are shown in Figure 10According to the results, the groups where splicing occurred were: G-SEA, G-SED, G-SEE, G-SEF, G-SEH, G-SEI, G-SEL, G-SEM, G-SES, G-SET, G-SEV, G-SEW, G-SEY. Among them, the groups with significant splicing product bands were: G-SEH, G-SEI, G-SES, G-SET.

[0290] The results show that for NpuDnaE, when the +1 amino acid is S, the +2 amino acid is E, and the +3 amino acid is selected from: A, D, E, F, H, I, L, M, S, T, V, W, Y, it has a relatively high splicing efficiency. In particular, when the +3 amino acid is selected from: H, I, S, T, it has a very high splicing efficiency.

[0291] In summary, the novel flanking sequence pairs of the split intein NpuDnaE that can undergo effective splicing in the present invention are shown in Table 41.

[0292] Table 41 Novel flanking sequence pairs of intein NpuDnaE

[0293]

[0294] Example 2 Splicing comparison between the optimal flanking sequence pair and the known flanking sequence pair

[0295] ● Construction of expression plasmids A-Hln, pA-L, plasmid (pTag-Ic-FSb-Pb)

[0296] Using the conditions in "Preparation of Recombinant Polypeptides", as Figure 4A 、 4B shown, the component expression plasmids of intein NpuDnaE were constructed respectively using the pcDNA3.1 plasmid vector according to the compositions shown in Table 31 and Table 32. The pA-L plasmid used the same pA-L plasmid as in Example 1. For intein NpuDnaE, one of the selected optimal flanking sequence pairs, GK and SET, was used to construct the plasmid pA-HIn(2) corresponding to A-Fab2, and the plasmid pTag-Ic-FSb-(B-FcIc6) corresponding to B-FcIc6.

[0297] The transfection conditions were: the molar ratio of plasmids was pTag-Ic-FSb-(B-FcIc6)∶pA-HIn(2)∶pA-L = 3∶1∶1, and the expression product was A61. Positive controls A1 and A10 were set up. The plasmids corresponding to A1 were pA-HIn(4), pTag-Ic-FSb-(B-FcIc2) and pA-L, and the plasmids corresponding to A10 were pA-HIn(3), pTag-Ic-FSb-(B-FcIc1) and pA-L. The transfection ratio of each plasmid was the same as above.

[0298] The supernatant was taken after culturing the transfected cells for 5 days. Protein A affinity chromatography was performed on the protein in the supernatant. After Protein A affinity chromatography, Coomassie Brilliant Blue staining was performed by SDS-PAGE (adding reducing agent) to detect the protein in the supernatant.

[0299] Figure 11 , The results showed that the appearance of a reduced band near 50 kD indicated that significant splicing of A61 occurred. The positive control A10 could also be spliced, but A1 did not undergo splicing. It was shown that when the flanking sequences a and b were MGG and SVY, efficient splicing could not occur. For the intein NpuDnaE, the corresponding flanking sequence pairs with excellent splicing efficiency were: when flanking sequence a was GK, flanking sequence b was SET.

[0300] The inteins and flanking sequences corresponding to groups A1, A10, and A61 are shown in Table 42.

[0301] Table 42 Different inteins and corresponding effective flanking sequence pairs

[0302]

[0303] Example 3 Intein-mediated in vitro splicing of polypeptide fragments from different protein sources

[0304] ● Vector construction and polypeptide expression

[0305] Using the same conditions as in Example 1, in this example, pcDNA3.1 was used to construct component expression plasmids via the intein NpuDnaE according to the compositions shown in Tables 31 and 33, respectively.

[0306] The expression plasmids for component A in this example were pA-L and pA-HIn(x), where x was different numbers.

[0307] The expression plasmids for component B' in this example were three types: B'-L expression plasmid (pB'-L), B'-H expression plasmid (pB'-H), and B'-FcIc expression plasmid (pB'-FcIcx), where x was different numbers. Among them, pB'-L and B'-H expression plasmids were common among each component B'.

[0308] For the intein NpuDnaE, plasmids pB'-FcIc(1) to B'-FcIc(7) corresponding to B'-HAb1 to B'-HAb7 were constructed.

[0309] Expression and purification of component A:

[0310] Co-transfect each plasmid pA-HIn(x) and plasmid pA-L into CHO cells and culture them at 37°C. The molar ratio of the plasmids is pA-HIn∶pA-L = 1∶1. Harvest the cell supernatant 10 days after transfection. Purify the supernatant using a nickel column chromatography (Jiangsu Qianchun, product number: A41002-06) to obtain the polypeptide fragment of purified component A.

[0311] Expression and purification of component B’:

[0312] Co-transfect plasmid pB’-L, plasmid pB’-H and each plasmid pB’-FcIc into 293E cells and culture them at 37°C. The molar ratio of the plasmids is pB’-L∶pB’-H∶pB’-FcIc = 1∶1∶3. Harvest the cell supernatant 10 days after transfection. Purify the supernatant using a nickel column chromatography to obtain the polypeptide fragment of purified component B’.

[0313] As shown in Table 43, the obtained polypeptide fragments of component A and component B’ are respectively named Fab4 and HAb4.

[0314] Table 43 Polypeptide fragments of obtained component A and component B’

[0315]

[0316] Perform non-reducing SDS-PAGE and Coomassie Brilliant Blue staining on the obtained purified polypeptide fragments of component A and component B’. The results are shown in Figures 12A - 12B .

[0317] E1, E2, and E3 represent elution components with different imidazole concentrations from low to high during the nickel column chromatography process. As Figure 12A can be seen, Fab4 has a relatively high expression level. Moreover, in the Fab4 group, polypeptides with relatively high purity can be obtained by purifying the polypeptides using a nickel column chromatography. As Figure 12B can be seen, HAb4 has a relatively high expression level.

[0318] ● In vitro splicing

[0319] Dialyze the obtained purified polypeptide fragments Fab4 and HAb4 of component A and component B’ respectively into the buffer using a 3kD dialysis bag (purchased from Sigma) at 4°C. The protein concentration of the components is 1 - 10 μmol / L. The buffer contains: 10 - 50 mM Tris / HCl (pH 7.0 - 8.0), 100 - 500 mM NaCl, 0 - 0.5 mM EDTA. Then mix component A (Fab4) and component B’ (HAb4) respectively in a molar ratio of 1∶10 - 10∶1 according to the corresponding serial numbers, add DTT to 0.5 - 5 mM, and incubate overnight at 37°C.

[0320] The obtained spliced product polypeptide was subjected to SDS-PAGE and Coomassie Brilliant Blue staining, and the results are shown in Figure 13 .

[0321] In Figure 13 , "Splicing 1" means that the concentrations of component A and component B' are 10 uM and 1 uM respectively, and the reaction system contains 2 mM DTT; "Splicing 2" means that the concentrations of component A and component B' are 5 uM and 1 uM respectively, and the reaction system contains 2 mM DTT; "No splicing 1" means that the concentrations of the component and component B' are 10 uM and 1 uM respectively, and the reaction system does not contain DTT; "No splicing 2" means that the concentrations of component A and component B' are 5 uM and 1 uM respectively, and the reaction system does not contain DTT; the control bands are component A as Fab4 (non-reduced), component B' as HAb4 (non-reduced), and the monoclonal antibody. Both "Splicing 1" and "Splicing 2" were incubated overnight at 37 °C, and the other groups were stored at 4 °C.

[0322] It can be seen from Figure 13 that the split intein NpuDnaE with the novel flanking sequence pair of the present invention can undergo efficient and effective splicing in vitro, thereby obtaining an in vitro spliced recombinant polypeptide of polypeptide fragments from different proteins, and obtaining the spliced products "Splicing 1" and "Splicing 2". Compared with the monoclonal antibody, the band size of this spliced product is the same, which is 150 kD, proving that the theoretical molecular weight of this product is the same as that of the natural IgG monoclonal antibody.

[0323] ● Detection of biological activity of the spliced product

[0324] The biological activity of the recombinant polypeptide "Splicing 2" was detected by double antigen sandwich ELISA.

[0325] 1) Antigen preparation: For proteins PD-L1 and CD38, they were constructed in a way that only the extracellular domains were selected, and expression plasmids with His tags were constructed. The vector used was pcDNA3.1.

[0326] After construction, 293E cells were used for transient transfection, and expression and purification including two steps of nickel column purification and molecular sieve purification were carried out. After purification, antigen proteins with a purity of not less than 95% detected by SDS-PAGE were obtained.

[0327] The PD-L1 protein was labeled with horseradish peroxidase (HRP).

[0328] 2) Coating with the first antigen: The concentration of the CD38 protein was adjusted to 2 μg / ml, and the enzyme-linked immunosorbent assay (ELISA) plate was coated with 100 μl / well of the liquid containing the CD38 protein and incubated overnight at 4 °C; the supernatant was discarded, and 250 μl of blocking solution (PBS containing 3% BSA) was added to each well;

[0329] 3) Antibody addition: Operate at room temperature according to the experimental design. Dilute the antibody in gradients. The diluent is PBS containing 1% BSA. For example, if the initial concentration of the antibody dilution is 20 μg / ml, perform a 2-fold dilution for 5 concentration gradients. Add the diluted antibody to the ELISA plate wells at 200 μl per well, incubate statically at room temperature for 2 h, and then discard the supernatant;

[0330] 4) Washing: Wash 3 times with 200 μl / well PBST (PBS containing 0.1% Tween 20);

[0331] 5) Incubation with the second antigen: Add the diluted second antigen (HRP-labeled PD-L1 protein). The second antigen is used after being diluted at 1:1000.

[0332] The diluent is PBS containing 1% BSA, and the volume is 100 μl per well. Incubate at room temperature for 1 h;

[0333] 6) Washing: Wash 5 times with 200 μl / well PBST;

[0334] 7) Color development: Add 100 μl / well of TMB color development solution (Prepare color development solutions A and B, purchased from Wuhan Boster Biological Technology Co., Ltd.; mix A and B at 1:1 and use immediately after preparation), and develop color at 37 °C for 5 min.

[0335] 8) Add 100 μl / well of 2 M HCl termination solution. After adding the termination solution, read the absorbance at 450 nm using an ELISA reader within 30 min.

[0336] The ELISA test results of Fab4, HAb4 polypeptide fragments, the unspliced mixture of the two, and the polypeptide fragment Fab4 + HAb4 after in vitro splicing of the two by intein are shown in Figure 14 .

[0337] It can be seen from Figure 14 that Fab4 + HAb4 (splicing 2) has the activity of binding two antigens, CD38 and PD-L1. However, the unspliced mixture in vitro, as well as the individual components A (Fab4) and B (HAb4), do not have the activity of binding two antigens simultaneously.

[0338] The results can prove that by using the intein of the present invention and the novel flanking sequence pair contained therein for splicing, the obtained Fab4 + HAb4 splicing product has good bispecific antibody activity.

[0339] Based on the splicing principle of inteins, according to the molecular weight of the splicing products obtained in the present invention and the results of double antigen sandwich ELISA, it can be speculated that a bispecific antibody with an effective and native-like IgG structure is obtained by the present invention. The test results confirm that the structure of the bispecific antibody is a heterodimer IgG structure composed of two different heavy chains and two different light chains, rather than a mixture of a homodimer IgG structure composed of two identical heavy chains and two identical light chains.

[0340] Industrial applicability

[0341] The present invention provides a method for preparing recombinant polypeptides, especially bispecific antibodies, using split inteins with novel flanking sequence pairs. The split inteins with novel flanking sequence pairs according to the present invention can be widely used in the preparation of recombinant polypeptides in the fields of medicine and bioengineering, especially in the field of antibodies, particularly in the preparation of bispecific antibodies. The bispecific antibodies prepared using the split inteins with novel flanking sequence pairs of the present invention do not have non-natural domains, and their structures are extremely similar to those of natural antibodies (IgA, IgD, IgE, IgG or IgM) and have an Fc domain. The structure of the bispecific antibody is intact and has good stability, and can retain or remove CDC (complement-dependent cytotoxicity) or ADCC (antibody-dependent cell-mediated cytotoxicity) or ADCP (antibody-dependent cell phagocytosis) or FcRn (Fc receptor) binding activity according to different IgG subtypes.

[0342] The bispecific antibodies prepared by the method of the present invention have a long in vivo half-life and low immunogenicity; no linker peptides of any form are introduced, the stability of the antibody molecule is improved, and the immune response in vivo is reduced. The bispecific antibodies prepared by the method of the present invention have the same glycosylation modification as wild-type IgG, resulting in better biological functions, being more stable, and having a long in vivo half-life; by using the in vitro splicing method with inteins, the problems of heavy chain mispairing and light chain mispairing that are extremely likely to occur in traditional methods can be completely avoided.

[0343] The method for preparing bispecific antibodies of the present invention can be used to produce humanized bispecific antibodies and fully human sequence bispecific antibodies. The sequences of such antibodies prepared by the method of the present invention are closer to human antibodies, and the occurrence of immune responses can be effectively reduced. The method for preparing bispecific antibodies of the present invention can construct any bispecific antibody without being restricted by antibody subtypes (IgG, IgA, IgM, IgD, IgE, and light chain K and λ types).

Claims

1. Flanking sequence pairs for split inteins, wherein, the flanking sequence pairs include: flanking sequence a and flanking sequence b; the flanking sequence a is located at the N-terminus of the N-terminal protein splicing region (In) of the split intein, and is between the N-terminal extein (En) and In; the flanking sequence b is located at the C-terminus of the C-terminal protein splicing region (Ic) of the split intein, and is between Ic and the C-terminal extein (Ec); the split intein is NpuDnaE, wherein: the flanking sequence a is GG, SG, GA, GK, GQ, GR, GW, GT, GS, DG, FG, LG, NG, WG, and the flanking sequence b is SE, or the flanking sequence a is DG, FG, GG, LG, NG, SG, WG, GA, GK, GQ, GR, GW, GT or GS, and the flanking sequence b is SEX, where X is any amino acid selected from: A, V, L, M, I, S, T, F, Y, W, H, D, E.

2. Flanking sequence pairs for split inteins, wherein, the flanking sequence pairs include: flanking sequence a and flanking sequence b; the flanking sequence a is located at the N-terminus of the N-terminal protein splicing region (In) of the split intein, and is between the N-terminal extein (En) and In; the flanking sequence b is located at the C-terminus of the C-terminal protein splicing region (Ic) of the split intein, and is between Ic and the C-terminal extein (Ec); the split intein is NpuDnaE, wherein the flanking sequence a is GG or SG, and the flanking sequence b is SET or SEI or SES or SEH; or the flanking sequence a is GA, GK, GQ, GR, GW, GT, GS, and the flanking sequence b is SET or SEI or SES or SEH; or the flanking sequence a is DG, FG, LG, NG, WG and the flanking sequence b is SET or SEI or SES or SEH.

3. The flanking sequence pair for split intein according to claim 1 or 2, wherein, The split intein and the flanking sequence pair are used together for trans-splicing, wherein, the NpuDnaE is composed of In with the sequence of SEQ ID NO:31 and Ic with the sequence of SEQ ID NO:

32.

4. A method for preparing a recombinant polypeptide, which comprises: (1) Providing component A and component B, wherein component A includes flanking sequence a, N-terminal extein En and In, the N-terminus of the flanking sequence a is connected to the C-terminus of the N-terminal extein En, and the C-terminus of the flanking sequence a is connected to In; component B includes flanking sequence b, C-terminal extein Ec and Ic, the C-terminus of the flanking sequence b is connected to the N-terminus of the C-terminal extein Ec, and the N-terminus of the flanking sequence b is connected to Ic; wherein, the flanking sequence a and flanking sequence b are as described in claim 1 or 2, and the coding sequences of the N-terminal extein En and the C-terminal extein Ec are from the same protein or different proteins; and (2) Performing in vitro trans-splicing on the component A and the component B to obtain a recombinant polypeptide, wherein the recombinant polypeptide is a bispecific antibody.

5. The method according to claim 4, wherein in step (1), it includes making cells containing nucleic acid sequences encoding component A and component B express said component A and component B.

6. The method according to claim 4, wherein the N-terminal exon peptide En and the C-terminal exon peptide Ec can be different domains of an antibody.

7. The method according to any one of claims 4-6, wherein, The recombinant polypeptide is obtained by trans-splicing of component A and component B; In component A, the N-terminus of the flank sequence a is connected to the C-terminus of En, and the C-terminus of the flank sequence a is connected to In; In component B, the C-terminus of the flank sequence b is connected to the N-terminus of Ec, and the N-terminus of the flank sequence b is connected to Ic; Wherein, the coding sequences of En and Ec are respectively from the N-terminal part and the C-terminal part of the same protein.

8. The method according to claim 7, wherein a tag protein is connected to the C-terminus of In, or a tag protein is connected to the N-terminus of Ic.

9. The method according to claim 8, wherein the tag protein is selected from SEQ ID NO: 24, 25, 26, 27, 28, 29 or 30.

10. The method according to claim 4, wherein, The recombinant polypeptide is obtained by trans-splicing of component A and component B; In component A, the N-terminus of the flank sequence a is connected to the C-terminus of En, and the C-terminus of the flank sequence a is connected to In; In component B, the C-terminus of the flank sequence b is connected to the N-terminus of Ec, and the N-terminus of the flank sequence b is connected to Ic; Wherein, the coding sequences of En and Ec are from different proteins.

11. The method according to claim 4, wherein, The antibody is of the natural immunoglobulin IgG, IgM, IgA, IgD, or IgE class, or immunoglobulin subclass: IgG1, IgG2, IgG3, IgG4, IgG5, or different classes of light chains: κ, λ; or a single-domain antibody; or The antibody is a full-length antibody or a functional fragment of an antibody.

12. The method according to claim 11, wherein, The functional fragment of the antibody is selected from: the variable region of the heavy chain of the antibody VH, the variable region of the light chain of the antibody VL, the fragment of the constant region of the heavy chain of the antibody Fc, the constant region 1 of the heavy chain of the antibody CH1, the constant region 2 of the heavy chain of the antibody CH2, the constant region 3 of the heavy chain of the antibody CH3, the constant region of the light chain of the antibody CL, or the variable region of a single-domain antibody VHH, one or more of them.

13. The method according to claim 4, wherein, One or more of the same protein, or the different proteins are specific for antigen or epitope A. The antigen A includes: tumor cell surface antigen, immune cell surface antigen, cytokine, cytokine receptor, transcription factor, membrane protein, actin, virus, bacterium, endotoxin, FIXa, FX, CD3, SLAMF7, CD38, BCMA, CD20, CD16, CEA, PD-L1, PD-1, CTLA-4, TIGIT, LAG-3, VEGF, B7-H3, Claudin18.2, TGF-β, Her2, IL-10, Siglec-15, Ras, C-myc, and the epitope A is the immunogenic epitope of the antigen A.

14. The method according to claim 13, wherein, One or more of the same protein, or the different proteins are specific for an antigen or epitope B different from antigen or epitope A. The antigen B includes: tumor cell surface antigen, immune cell surface antigen, cytokine, cytokine receptor, transcription factor, membrane protein, actin, virus, bacterium, endotoxin, FIXa, FX, CD3, SLAMF7, CD38, BCMA, CD20, CD16, CEA, PD-L1, PD-1, CTLA-4, TIGIT, LAG-3, VEGF, B7-H3, Claudin18.2, TGF-β, Her2, IL-10, Siglec-15, Ras, C-myc, and the epitope B is the immunogenic epitope of the antigen B.

15. The method according to claim 14, wherein the recombinant protein is a bispecific antibody that can bind the antigen or epitope A and B simultaneously.

16. The method according to claim 15, wherein the multispecific antibody is a humanized bispecific antibody or a bispecific antibody with a fully human sequence.

Citation Information

Patent Citations

  • Multivalent antibodies and uses therefor

    WO2001077342A1

  • Split inteins and uses thereof

    CN104053779A