Novel covalent protein and use thereof
A yeast display-based high-throughput system optimizes covalent protein development, achieving rapid and specific covalent bonding for enhanced therapeutic efficacy, addressing the limitations of existing monoclonal antibodies and small-molecule drugs.
Patent Information
- Application Number
- PCT/CN2025/094328
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-11
- Filing Date
- 2025-05-12
- Publication Date
- 2025-11-20
AI Technical Summary
Current monoclonal antibody drugs face challenges such as inadequate tissue penetration, undesirable Fc effector function, diminished physical and chemical stability, and high manufacturing costs, while covalent small-molecule drugs have swift covalent inhibition but short in vivo half-life, limiting their potential as standalone drugs. Additionally, existing methods for developing covalent proteins are low throughput and hinder the optimization of crosslinking rates.
A yeast display-based high-throughput selection system is developed to engineer covalent proteins with supercharged inhibition rates by integrating chemical protein modification, using diverse chemical warheads, enabling rapid and specific covalent bonding with target proteins.
The system successfully creates covalent proteins with enhanced covalent inhibition rates, surpassing those of clinically employed small molecules, and demonstrates deep tissue penetration and permanent target inhibition, paving the way for covalent miniproteins as effective standalone drugs.
Smart Images

Figure PCTCN2025094328-FTAPPB-I100001 
Figure PCTCN2025094328-FTAPPB-I100002 
Figure PCTCN2025094328-FTAPPB-I100003
Abstract
Description
NOVEL COVALENT PROTEIN AND USE THEREOF
[0001] CROSS REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of PCT Application PCT / CN2024 / 092697, filed May 11, 2024. The entire content of the foregoing application is incorporated herein by reference.FIELD OF THE INVENTION
[0003] The present disclosure relates to a novel covalent protein, its pharmaceutical composition and use thereof in diagnosing, treating, or preventing diseases such as cancer and autoimmune disorders, which has the advantageous properties including high stability and potency, cost-effective production, enhanced tissue penetration, and the capacity for rapid and permanent target inhibition. The present disclosure also relates to a screening method for developing covalent protein. The present disclosure further details a high-throughput yeast display-based screening method for the development of such covalent proteins. The present disclosure provides novel compounds as crosslinker for the covalent protein.BACKGROUND OF THE INVENTION
[0004] Monoclonal antibody drugs have become the cornerstone of modern therapeutics. Nonetheless, antibodies grapple with various limitations, including suboptimal extravasation resulting in inadequate tissue and tumor penetration, undesirable Fc effector function, diminished physical and chemical stability (in comparison to small molecules) , and elevated manufacturing costs1-4. In response to these challenges, there is a growing emphasis on developing natural or artificially engineered miniproteins as innovative protein drug modalities2, 3, 5. The compact size of miniproteins (ranging from 5 kDa to 15 kDa) positions them to penetrate more effectively into tissues and solid tumors2, 3, 5. Furthermore, their enhanced physical and chemical stability allows for precise and on-demand functional manipulation and streamlined manufacturing processes2, 3, 5. Despite these advantages, miniproteins typically exhibit a brief in vivo half-life (less than 30 minutes) due to their diminutive size, thereby significantly limiting their potential as standalone drugs5, 6. While approaches like PEGylation and Fc domain fusion offer means to extend the half-life of miniproteins6, 7, such modifications come at the expense of their intrinsic size advantage. Consequently, there is a pressing need for novel strategies to fully exploit the therapeutic potential inherent in miniproteins, thereby promoting miniproteins as standalone drug entities.
[0005] In recent years, there have been significant advancements for covalent small-molecule drugs, with numerous approvals for clinical use in treating various ailments, especially for cancers8, 9. These drugs exhibit the unique capability to irreversibly bind target proteins, resulting in heightened inhibitory activity and prolonged effectiveness8-11. Despite their short in vivo half-life, their swift covalent inhibition rate and irreversible mechanism endow them with a sustained period of efficacy8-11. As a case in point, aspirin, with a half-life of 20 minutes, remarkably maintains its effective action for 10 days owing to its covalent inhibition of the target protein12. Notably, the covalent inhibition rate of these small molecules is often exceptionally rapid; for instance, some clinically approved kinase covalent inhibitors boast a covalent target inhibition rate Kinact of 10-3 s-1, with a target inhibition t1 / 2 of ~10 minutes9-11. Inspired by the success of covalent small-molecule drugs, researchers are now exploring the potential of covalent proteins as therapeutic agents13-19. Theoretically, if covalent proteins can swiftly reach the target (< 30 minutes) and form a rapid, specific covalent bond with the target protein (< 30 minutes) , leading to permanent target engagement, these covalent protein drugs could achieve a prolonged duration of action. This might eliminate the necessity for an extended half-life, with the benefit of reducing potential side effects caused by the long circulation lifetime, positioning them as standalone drugs. However, chemical warheads chosen for covalent proteins often have significantly weaker reactivity towards canonical amino acids to ensure stability and target engagement specificity13-17. Current research on covalent protein drugs predominantly utilizes genetic code expansion technology to introduce weak warheads13-17. Reported covalent proteins often exhibit a lengthy reaction time (> 4 hours) , far exceeding their circulation half-life, to achieve partial inhibition (approximately 50-70%covalent inhibition efficiency) of target proteins13-17, 19. This prolonged duration hampers their potential for drug development. Although the covalent crosslinking rate can be optimized by screening different warhead insertion positions using genetic code expansion, this approach entails a low throughput screening with only marginal rate enhancement, hindering the progression of this promising field13-15, 17. The weak reactivity of these chosen warheads also raises questions about the feasibility of developing supercharged (fast crosslinking rate) covalent miniproteins with kinetics comparable to or exceeding those of covalent small molecule drugs.
[0006] The rate of the crosslinking reaction between the covalent protein and the target protein is contingent upon the intrinsic reactivity of the warheads, a factor greatly influenced by their chemical environment and their relative geometry to the target protein (FIG. 1a) 20, 21 Consequently, the choice of chemical warhead, their attachment to the protein, and the local protein sequence proximal to the chemical warhead collectively impact the crosslinking rate. The optimal spatial arrangement of the warhead and the targeted amino acid can significantly enhance the crosslinking reaction rate by reducing entropy, akin to enzyme-substrate recognition (FIG. 1a) 20, 21. However, such optimal spatial prearrangement and chemical environment are difficult to predict or design. Hence, there is a pressing need for a comprehensive high-throughput selection method to effectively screen all potential combinations of influencing factors, facilitating the development of supercharged covalent miniproteins, and by using such a high-throughput selection method, covalent mini-proteins with excellent therapeutic effects can be effectively screened.SUMMARY OF THE INVENTION
[0007] To streamline the development of covalent proteins in a high throughput fashion, the inventors established a yeast display-based selection system. This system facilitates the development of covalent protein drugs boasting supercharged inhibition rates and heightened efficacy by integrating yeast display technology with chemical protein modification, incorporating diverse chemical warheads (FIG. 1a-b) . Leveraging this high-throughput selection system, the inventors have successfully engineered a PD-L1 targeting covalent nanobody with exceptional covalent inhibition rate, surpassing that of clinically employed covalent small molecules. Moreover, the inventors also successfully developed a covalent miniprotein targeting the SARS-CoV-2 RBD, exhibiting even faster covalent inhibition kinetics. The novel high-throughput selection system of the preset disclosure has broad prospects for significantly contributing to the advancement of covalent miniproteins as standalone drugs.
[0008] In one aspect, the present disclosure provides a covalent protein having a structure as shown in Formula (I) :
[0009] wherein
[0010] Ab represents a protein moiety, and R1 is covalently bonded to Ab, preferably through side chain of Cys residue;
[0011] R1 has a structure of -R11-, or -R11-C (O) -R12-R13-, where
[0012] R11 is optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN,
[0013] R12 is selected from the group consisting of NR12a, O, S, and heterocyclylene, R12a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0014] R13 is absent or optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN;
[0015] RA is selected from the group consisting of optionally substituted alkanediyl, and optionally substituted arenediyl;
[0016] Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, alkyl, haloalkyl, and aryl,
[0017] n is an integer selected from 0 and 1;
[0018] Ry is selected from the group consisting of S and P;
[0019] Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, alkyl, haloalkyl, and aryl;
[0020] The bond between Ry and Rz is a single bond or a double bond.
[0021] The present disclosure also provides a covalent protein having a structure as shown in Formula (I-B1) :
[0022] wherein Ab represents a protein moiety, and R1 is covalently bonded to Ab;
[0023] R1 is a linker moiety having a structure of -R11-, or -R11-C (O) -R12-R13-, where
[0024] R11 is optionally substituted alkylene, which is optionally substituted with one or more halo, -OH, or -CN,
[0025] R12 is selected from the group consisting of NR12a, O, S, heterocyclylene, and heteroarylene, R12a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0026] R13 is absent or optionally substituted alkylene, which is optionally substituted with one or more halo, -OH, or -CN;
[0027] Rx is selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions, m is an integer selected from 1 to 4;
[0028] n is an integer selected from 0 and 1.
[0029] In some embodiments, the protein moiety Ab comprises at least one naturally occurring or engineered cysteine (Cys) residue in its amino acid sequence, preferably the cysteine residue is spatially located at or near binding interface region between the protein moiety Ab and a target protein.
[0030] In some embodiments, the protein moiety Ab includes but is not limited to the group consisting of an antibody or antigen-binding fragment thereof; a non-antibody scaffold protein; a cytokine, a growth factor, a hormone, or a variant and functional fragment thereof; a receptor protein or ligand-binding domain thereof; an enzymes or a modulator thereof; a peptide with specific binding activity, and the like.
[0031] In some embodiments, the protein moiety Ab is a PD-L1 targeting nanobody, or a SAR-CoV-2 RBD targeting miniprotein.
[0032] In some embodiments, the protein moiety Ab is the PD-L1 targeting nanobody of the present disclosure as provided hereinafter.
[0033] In some embodiments, the protein moiety Ab is the SAR-CoV-2 RBD targeting miniprotein of the present disclosure as provided hereinafter.
[0034] The present disclosure also provides a pharmaceutical composition comprising a therapeutically effective amount of the covalent protein of the present disclosure, and a pharmaceutically acceptable carrier, excipient, or diluent.
[0035] The present disclosure also provides use of the covalent protein of the present disclosure or the pharmaceutical composition of the present disclosure in the manufacture of a medicament for diagnosing, treating or preventing a disease, disorder, or condition in a subject in need thereof.
[0036] The present disclosure provides the covalent protein of the present disclosure or the pharmaceutical composition of the present disclosure for use in diagnosing, treating or preventing a disease, disorder, or condition in a subject in need thereof.
[0037] The present disclosure also provides a method for diagnosing, treating or preventing a disease, disorder, or condition in a subject in need thereof, comprising administering to the subject an effective amount of the covalent protein of the present disclosure or the pharmaceutical composition of the present disclosure.
[0038] The present disclosure also provides a kit comprising:
[0039] (a) the covalent protein of the present disclosure or the pharmaceutical composition of the present disclosure; and
[0040] (b) instructions for use, said instructions for use directing the diagnosis, treatment or prevention of a disease, disorder, or condition in a subject in need thereof.
[0041] In some embodiments, the disease, disorder, or condition is a disease associated with dysregulated PD-L1 expression or activity including cancer and autoimmune disease, or a disease caused by a coronavirus infection including COVID-19.
[0042] The present disclosure also provides a conjugate obtained by covalently linking the covalent protein of the present disclosure and a target protein via a crosslinker in the covalent protein.
[0043] In another aspect, the present disclosure provides a compound having a structure as shown in Formula (II) :
[0044] wherein
[0045] R2 has a structure of R21-, or R21-C (O) -R22-R23-, where
[0046] R21 is selected from optionally substituted haloalkyl and optionally substituted alkenyl, where haloalkyl or alkenyl can be optionally substituted with one or more halo, -OH, or -CN,
[0047] R22 is selected from the group consisting of NR22a, O, S, and heterocyclylene, R22a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0048] R23 is absent or optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN;
[0049] RA is selected from the group consisting of optionally substituted alkanediyl, and optionally substituted arenediyl;
[0050] Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, alkyl, haloalkyl, and aryl,
[0051] n is an integer selected from 0 and 1;
[0052] Ry is selected from the group consisting of S and P;
[0053] Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, alkyl, haloalkyl, and aryl;
[0054] The bond between Ry and Rz is a single bond or a double bond.
[0055] The present disclosure also provides a compound having a structure as shown in Formula (II) as a crosslinker for covalent protein:
[0056] wherein
[0057] R2 has a structure of R21-, or R21-C (O) -R22-R23-, where
[0058] R21 is optionally substituted haloalkyl, or optionally substituted alkenyl, which haloalkyl, alkenyl is optionally substituted with one or more halo, -OH, or -CN,
[0059] R22 is selected from the group consisting of NR22a, O, S, and heterocyclylene, R22a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0060] R23 is absent or optionally substituted alkylene, which is optionally substituted with one or more halo, -OH, or -CN;
[0061] Rx is selected from the group consisting of -H, -halo, -CN, and -NO2 or other chemical groups with similar functions, m is an integer selected from 1 to 4;
[0062] n is an integer selected from 0 and 1.
[0063] The present disclosure also provides use of the compound having a structure as shown in Formula (II) as a crosslinker for covalent protein.
[0064] The present disclosure also provides a method for preparing a covalent protein of the present disclosure, the method comprising:
[0065] (a) providing a protein comprising at least one cysteine (Cys) residue;
[0066] (b) providing a compound having a structure as shown in Formula (II) ; and
[0067] (c) reacting the compound with the side chain of the at least one cysteine (Cys) residue of the protein to form a covalent bond.
[0068] The detailed description and description of the protein (also referred to as binding protein, i.e., a protein containing at least one cysteine (Cys) residue) and the crosslinker compound in the present disclosure are all applicable to the preparation method.
[0069] In another aspect, the present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0070] In some embodiments, the PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0071] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having deletion, substitution, insertion or addition of one or more amino acid residues as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0072] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations at the position selected from the group consisting of positions 100, 102, 104, 107, 108, 113, 116, 50, 59, 103, 110, 111, 112, 5, 72, 73, 75, 77, 87, 88, 93, and 123 as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0073] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 4 to 38.
[0074] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 4 to 38.
[0075] The present disclosure provides an isolated nucleic acid molecule encoding any of the PD-L1 targeting nanobodies of the present disclosure.
[0076] The present disclosure provides an expression vector comprising the nucleic acid molecule encoding any of the PD-L1 targeting nanobodies of the present disclosure.
[0077] The present disclosure provides a host cell comprising the expression vector of the present disclosure or the nucleic acid molecule encoding any of the PD-L1 targeting nanobodies of the present disclosure.
[0078] In another aspect, the present disclosure provides a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0079] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0080] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having deletion, substitution, insertion or addition of one or more amino acid residues as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0081] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations selected from the group consisting of D11C, K26C, K26G, K26Q, K27R F30C, and Y40C as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0082] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations selected from the group consisting of K26G, K26Q, K27R, and F30C as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0083] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 56 to 61.
[0084] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in any of SEQ ID NOs. 56 to 61.
[0085] The present disclosure provides an isolated nucleic acid molecule encoding any of the SAR-CoV-2 RBD targeting miniproteins of the present disclosure.
[0086] The present disclosure provides an expression vector comprising the nucleic acid molecule encoding any of the SAR-CoV-2 RBD targeting miniproteins of the present disclosure.
[0087] The present disclosure provides a host cell comprising the expression vector of the present disclosure or the nucleic acid molecule encoding any of the SAR-CoV-2 RBD targeting miniproteins of the present disclosure.
[0088] In another aspect, the present disclosure provides a high throughput covalent protein selection method based on yeast display, comprising
[0089] 1) generating a protein library displayed on yeast cells,
[0090] 2) binding a crosslinker to the displayed proteins;
[0091] 3) incubating with target protein; and
[0092] 4) performing FACS sorting.
[0093] In some embodiments, the crosslinker is a compound having a structure as shown in Formula (II) of the present disclosure.
[0094] In some embodiments, in step 2) , before binding the crosslinker to the displayer proteins, a reducing agent is added to treat the yeast cells. In some embodiments, the reducing agent is, for example, DTT, TCEP, or any other appropriate reducing agent in the art. In some embodiments, the reducing agent is added at a concentration range from 0.1 mM to 2 mM at a temperature of 4℃ to 25℃
[0095] In some embodiments, in step 2) , the crosslinker is added, for example, at a concentration of about 0.2 mM to the yeast cells to incubate for, for example, about 2 hours at a temperature of, for example, about 37℃ for attachment on the protein displayed.
[0096] In some embodiments, in step 3) , the yeast cells are incubated with the target protein, for the purpose of covalently binding the target protein to the displayed proteins through the crosslinker. In some embodiments, after binding of the target protein, an acid wash, for example, with a pH of about 2-3 is performed to remove noncovalent binder in order to selectively screen the covalent binder.
[0097] In some embodiments, after the acid wash, FACS sorting is performed to screen candidate covalent protein with high affinity to the target protein.
[0098] In some embodiments, steps 2) to 4) are repeatedly conducted to further screen candidate covalent proteins with high affinity and faster covalent crosslinking rate to the target protein.
[0099] In some embodiments, in each repeat cycle of steps 2) to 4) , the yeast cells are incubated with successively lower concentrations of the target protein, for example, 10 nM for 2nd round, 5 nM for 3rd round, and 3 nM for the 4th round, for the purpose of enriching for binders with higher affinities. In some embodiments, in each repeat cycle of steps 2) to 4) , successively shorter incubation times, for example, 20 min for the 2nd round, 10 min for the 3rd round, and 3 min for the 4th round, for the purpose of enriching for binders with faster reaction rates.
[0100] In some embodiments, the protein library consists of PD-L1 targeting proteins. In some embodiments, the target protein is PD-L1.
[0101] In some embodiments, the protein library is consisting of SARS-CoV-2 RBD targeting proteins. In some embodiments, the target protein is SARS-CoV-2 RBD.
[0102] The present disclosure employs chemical protein conjugation to establish simple and general procedures for producing covalent protein binders. These covalent protein molecules are highly stable, have a low production cost, have deep tissue penetration properties, and can permanently inhibit targets within a short time window. Any latent reactive group could be installed on any protein using the method of the present disclosure while avoiding the genetic code expansion technique. This will significantly broaden the repertoire of covalent protein therapeutic protein scaffolds and small molecule scaffolds. This will also ensure compatibility with high throughput screening (phage display or yeast display) if such screening is necessary to improve drug properties.BRIEF DESCRIPTION OF THE DRAWINGS
[0103] In order to describe more clearly the objects, technical solutions and beneficial effects of the present disclosure, the following accompanying drawings are provided.
[0104] FIG. 1 illustrates establishment of a high throughput covalent protein selection system based on yeast display and chemoselective protein modification. a. Supercharged covalent proteins could be developed through comprehensive screening of diversified crosslinkers and covalent protein sequences. The residues (yellow color) proximal to the crosslinker could greatly influence the covalent crosslinking reaction rate and need to be randomized for selection. b. The general procedure for yeast display-based supercharged covalent protein selection.
[0105] FIG. 2 illustrates development of PD-L1 targeting covalent proteins through diversified crosslinkers. a. Crystal structure of PD-L1 / KN035 complex (PDB 5JDS) in cartoon representation. KN035 is in light blue and PD-L1 is in pale green. Different residues (magenta) on KN035 were chosen for Cys mutation for crosslinker installation to target proximal His or Lys (dark blue) on PD-L1. b. Chemical structures of small molecule crosslinkers. c. Crosslinking analysis of initially selected covalent KN035_6-H (2 μM) with PD-L1 (1 μM) in vitro. Purified proteins were incubated in PBS buffer at 37℃ for indicated time points.
[0106] FIG. 3 illustrates supercharged PD-L1 targeting covalent nanobody selection. a. The construct of covalent protein displayed on yeast. b. Establishing the washing conditions to remove noncovalent binders in order to select covalent binders exclusively. c. Proximal residues are randomized to generate a yeast library for supercharged covalent protein selection. d. supercharged covalent protein selection process and the conditions used for each round of selection.
[0107] FIG. 4 illustrates Characterizations of the selected supercharged covalent PD-L1 targeting nanobody IB101 in vitro. a-b. Covalent crosslinking kinetics study between IB101 and PD-L1. SDS-PAGE analysis of IB101 (10 μM) crosslinking with PD-L1 (1 μM) at indicated time points (a) . Western blotting analysis of IB101 (1 μM) crosslinking with PD-L1 (0.1 μM) at indicated time points (b) . Western blot was visualized using an anti-Histag antibody. c. KN035, KN035_6-H, H9_111A and IB101 binding affinity with PD-L1 measurement. d. The in vitro T cell activation activity of KN035, KN035_108C, KN035_6-H, KN035 (FSY) , H9_111A, IB101, KN035-Fc, and atezolizumab was measured with PD-1 / PD-L1 blockade bioassay. Each sample was tested in triplicates, and the data are presented as the mean ± SEM. Each experiment was repeated at least twice. e. KN035_6-H and IB-101 covalent crosslinking with MC38k / hPD-L1, H460, U87 cell line surface PD-L1 at different concentrations and different incubation times. Data was visualized using anti-PD-L1 western blot analysis. f. IB101 covalent crosslinking with MC38k / hPD-L1, H460, U87 cell line surface PD-L1 at different concentrations upon 15 minutes coincubation. Data was visualized using anti-PD-L1 western blot analysis.
[0108] FIG. 5 illustrates characterizations of IB101 in vivo efficacy and specificity. a. Experiment scheme of the tumor-suppression study. MC38k / hPD-L1 cells were injected (s. c. ) into female hPD-1 / hPD-L1 transgenic mice. Drugs were administrated from day 5, at which point the tumors were around 100 mm3. b. IB101 covalently crosslinked with PD-L1 on tumors engrafted in mice in vivo. c. Evaluation of the tumor suppression efficacy of KN035-Fc and IB101 administered S. C. at 2 mg / kg every 4 days or 2 days. d. Survival rate comparison of different groups. e. Body weight monitoring of mice in different groups during the study, no significant difference was observed. f. Wester-blot analysis to validate IB101 specificity. IB101 only specifically crosslinked with PD-L1 using different cell lines; no obvious off-target crosslinking was observed. An anti-flag antibody was used to detect the flag tag appended at the C terminus of IB101 (upper panel) . An anti-hPD-L1 antibody was used to confirm the crosslinking of IB101 with PD-L1.
[0109] FIG. 6 illustrates supercharged covalent LCB-3 development and the characterizations. a. D11C, K26C, and F30C mutants of LCB3 were modified with diversified crosslinkers to generate covalent LCB3. Different covalent LCB3 (50 μM) were incubated with SARS-CoV-2 RBD (5 μM) at 37℃ for 5 hours to identify the most efficient covalent LCB3. LCB3-F30C showed the highest cross-linking efficiency with the linker 3-H. b. Studies of cross-linking efficiency between F30C_1-CN and SARS-CoV-2 RBD under different time and stoichiometry. c. Covalent crosslinking kinetics study of the optimized supercharged covalent F30C-Mu2_5-NO2. d. SARS-CoV-2 pseudovirus inhibition assay was employed to assess the inhibitory activities of LCB3 muteins and covalent F30C-Mu2_5-NO2. e. SARS-CoV-2 pseudovirus competitive inhibition assay with 2 μg / mL RBD protein. d-e Data were presented as the mean ± SD.
[0110] FIG. 7 illustrates Cys point mutation sites and crosslinkers screening. a. Point mutation sites chosen are spatially proximal to the potential targeting sites H69, K62, and K75 of PD-L1. KN035 and PD-L1 complex are shown in cartoon, KN035 is shown in light blue, and PD-L1 is shown in pale green. Sites mutated to Cys are highlighted in magenta, and corresponding target amino acids are highlighted in blue. b. Covalent binding of KN035-SM to PD-L1 was verified by SDS-PAGE. 6 μM KN035-SM were incubated with 2 μM refolded PD-L1 in PBS buffer at 37℃ for 6 hours before SDS-PAGE and Coomassie blue analysis. c. KN035_6-H covalently bound to PD-L1 on U87 human cancer cell surface. Indicated concentrations of KN035 or KN035_6-H were incubated with U87 for 12 hours, and then the cells were collected and analyzed with western blotting. d. The tandem mass spectrum of KN035_L108C_6-H / PD-L1 complex indicated that Cys108 reacted with His69 of PD-L1. e. LC-MS analysis of purified KN035 (FSY) . f. Time-course study of the crosslinking reaction between KN035 (FSY) and PD-L1, as verified by SDS-PAGE. KN035 (FSY) was incubated with PD-L1 (1 μM) under different molar ratios at 37℃ for the indicated time duration. Note. SM is the abbreviation of different small molecule crosslinkers.
[0111] FIG. 8 illustrates modified crosslinkers screening. a. Covalent binding kinetics of KN035-SM to PD-L1. KN035_L108C and 6-CN crosslinker combination performed the best. 10 μM KN035 proteins coupled with diverse crosslinker small molecules (KN035-SM) were incubated with 2 μM PD-L1 in PBS buffer at 37℃ for the indicated times. b. LC-MS analysis of 6-CN conjugation on KN035_L108C. KN035_L108C protein modified with crosslinker 6-CN had a self-crosslinking side reaction, and Y59F and K50A double mutation abolished the self-crosslinking side reaction. c. Modified crosslinker 6-CN sped up the crosslinking rate.
[0112] FIG. 9 illustrates yeast screening condition optimization. a. Flow cytometry histogram plots showing the effect of different concentrations of DTT on the display efficiency of KN035 on the yeast surface. b. Flow cytometry dot plot indicating the increased covalent binding between PD-L1 and KN035 after treatment of 0.5 mM of DTT for 10 minutes.
[0113] FIG. 10 illustrates site saturation mutagenesis. a. Structure of PD-L1 (pale green) and KN035 (light blue) complex (PDB: 5JDS) is shown as a cartoon presentation. Residues selected for site-saturated mutagenesis are highlighted in magenta. The sequence corresponding to the KN035 CDR regions is shown. The mutated residues are shown in the same color as depicted in the structural model. b. Flow cytometry plots showing binding capability between KN035 and PD-L1 after site-saturated mutagenesis was done at each position as indicated compared to the WT KN035. c. Flow cytometry plots showing binding between indicated KN035 muteins and PD-L1 compared to the WT KN035. d. Protein crosslinking efficiency was confirmed using SDS-PAGE. Indicated KN035 muteins were incubated with 1 μM PD-L1 in PBS buffer at 37℃ for the indicated times. e. Flow cytometry plots showing binding between KN035 double muteins and PD-L1 compared to the WT KN035. f. Protein crosslinking efficiency was confirmed using SDS-PAGE. Indicated double KN035 muteins were incubated with 1 μM PD-L1 in PBS buffer at 37℃ for the indicated times.
[0114] FIG. 11 illustrates highly efficient covalent binders evaluation and optimization. a. Flow cytometry plots showing non-covalent (left panel) and covalent binding (right panel) between selected KN035 muteins and PD-L1 compared to KN035 E102F by yeast surface display. b. Analysis of binding between KN035 muteins (CvPr) and PD-L1 in vitro. Each KN035 mutein (2 μM) was incubated with PD-L1 (1 μM) at 37℃ for the indicated time duration. c. PD-1 / PD-L1 blockade bioassay for evaluating screened covalent KN035 muteins. Three technical repeats were conducted. d. LC-MS spectrum shows CvH9 protein had a self-crosslinking side reaction. e. PD-1 / PD-L1 blockade bioassay for evaluating different CvH9 muteins. f. LC-MS spectrum of IB101 shows Y111A single mutation abolished the self-crosslinking side reaction.
[0115] FIG. 12 illustrates covalent crosslinking kinetic constant measurement. a. IB101 concentration at different time points was measured using gel band intensity, and 1 / [IB101] was plotted against time. Linear regression of the data yielded the apparent first-order rate constant of 0.184 min-1. b. Crosslinking product at different time points was measured using western band intensity, and 1 / [IB101] was plotted against time. Linear regression of the data yielded the apparent first-order rate constant of 0.229 min-1.
[0116] FIG. 13 illustrates pharmacokinetics measurement of IB101 in mice. a, b. IB101 was administered to mice through intravenous (I. V. ) injection (a) or subcutaneous (S. C. ) injection (b) . Blood samples were collected at different time points, and IB101 concentration in each blood sample was then measured using sandwich ELISA.
[0117] FIG. 14 illustrates in vivo distribution of Cy7 labeled KN035 / IB101 / KN035-Fc study in tumor-bearing mice. a-c. LC-MS spectrum of the conjugation reaction for preparing Cy7 labeled KN035 (a) , KN035-Fc (b) , and IB101 (c) . LC-MS spectrum of the reaction mixture confirmed that all the substrates were converted to their corresponding products. d, e. Fluorescence images (d) and quantitative analysis (e) of Cy7 labeled KN035 / IB101 / KN035-Fc proteins enrichment in MC38k / hPD-L1 tumor and their residence time.
[0118] FIG. 15 illustrates that IB101 stability and activity can be maintained using different storage methods. IB101 protein was stored at 4℃ (a) or 25℃ (b) in PBS buffer for up to 1 month. IB101 was lyophilized and stored at 25℃ for 1 month (c) , the lyophilized powder was then dissolved for analysis. The stability was analyzed with LC-MS. d. Studies of cross-linking efficiency between PD-L1 and IB101 stored under different conditions. 2 μM IB101 and 1 μM PD-L1 were incubated at 37℃ in PBS buffer for different times. The crosslinking activity was measured using SDS-PAGE.
[0119] FIG. 16 illustrates development of supercharged covalent LCB3 and the characterizations. a. The structure of LCB3 (gray) and SARS-CoV-2 RBD (cyan) complex (PDB: 7JZM) is shown. Residues selected for cysteine mutation are highlighted in orange, and targeting amino acids of RBD are labeled yellow. b. LC-MS analysis showed F30C_1-CN had a self-crosslinking side reaction at 37℃ upon 2 hours incubation. c. Mutations 1 and 2 of LCB3-F30C were prepared to reduce the self-crosslinking side reaction. d. Two muteins were incorporated with different crosslinkers to generate covalent LCB3 (20 μM) , which were incubated with SARS-CoV-2 wild-type RBD (2 μM) at 37℃ for 0.5 hours to identify the most efficient covalent LCB3. LCB3-F30C-Mu2 showed the highest cross-linking efficiency with the linker 5-NO2. e. Studies of cross-linking efficiency between F30C-Mu2_5-NO2 and SARS-CoV-2 RBD under different times and stoichiometry. f. Kinetics of F30C-Mu2_5-NO2 crosslinking with SARS-CoV-2 RBD. Linear regression of the data yielded the apparent first-order rate constant of 0.86 min-1.
[0120] FIG. 17 illustrates peptide Ac-YGGFLKVDVSHLS (SEQ ID NO. 89) reaction with crosslinker 6-CN analysis using LC_MS. 0.5 mM peptide was incubated with 1 mM crosslinker 6-CN at 37℃ for 16 hours in PBS buffer with 10%DMF. The incubating mixture was analyzed using LC-MS, and no 6-CN modified peptide was observed. Unlabeled small peaks in HPLC trace are nonpeptide-related contaminants.
[0121] FIG. 18 illustrates characterizations of IB101 in vivo efficacy using B16F10 tumor model. a. Experimental design of the in vivo B16F10 tumor-suppression study. b. Evaluation of the tumor suppression efficacy of Atezolizumab and IB101 administered (I. V. ) at 5 mg / kg or 2 mg / kg every 4 days or 2 days. Mice were sacrificed upon tumor volume reaching 1500 mm3. c. Body weight monitoring of mice treated with PBS, Atezolizumab, or IB101, no significant difference was observed.
[0122] FIG. 19 illustrates characterization of IB101_R and IB101_Q. a. Studies of cross-linking efficiency between PD-L1 and IB101_R or IB101_Q. 3 μM IB101 (_Q, _R, or wildtype) and 1 μM PD-L1 were incubated at 37℃ in PBS buffer for different times. The crosslinking activity was measured using SDS-PAGE. b. The in vitro T cell activation activity of IB101_R, IB101_Q, and IB101 was measured using PD-1 / PD-L1 blockade bioassay. c. Measurement of melting temperature of KN035, IB101, IB101_R, and IB101_Q.
[0123] FIG. 20 illustrates IB101_R stability and activity can be maintained after long-time storage. a-c. IB101_R no tag protein was freshly prepared (a) or stored at 37℃ in PBS (pH 5.0) buffer at 4 mg / mL for 2 months (b) or 4.5 months (c) . The stability was analyzed with LC-MS. d. Studies of cross-linking efficiency between PD-L1 and IB101_R stored for 2 months. 6 μM IB101 and 2 μM PD-L1 were incubated at 37℃ in PBS, pH 7.4 buffer for different times. The crosslinking activity was measured using SDS-PAGE.DETAILED DESCRIPTION OF THE INVENTION
[0124] Preferred embodiments of the invention will be described in detail below in conjunction with the accompanying drawings. Experimental methods for which specific conditions are not indicated in the embodiments usually follow conventional conditions, or follow the conditions recommended by the manufacturer. It will be understood that the specific embodiments described herein are intended only to explain the invention and not to limit it.
[0125] Unless otherwise indicated, the present disclosure will be implemented using the conventional techniques of molecular biology (including recombinant techniques) , microbiology, cell biology, biochemistry, and immunology in the art.
[0126] Definition
[0127] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs.
[0128] The articles “a” , “an” , and “the” are used herein to refer to one or to more than one (i.e. to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0129] The use of alternatives (e.g., "or" ) should be understood to mean one, two, or any combination thereof of the alternatives.
[0130] The term "and / or" is to be understood as referring to one or both alternatives.
[0131] Definitions of specific functional groups and chemical terms are described in more detail below. For purpose of this disclosure, the chemical elements are identified in accordance with the Periodic Table of the Elements, CAS version, Handbook of Chemistry and Physics, 75th Edition, inside cover, and specific functional groups are generally defined as described therein. Additionally, general principles of organic chemistry, as well as specific functional moieties and reactivity, are described in Organic Chemistry, Thomas Sorrell, University Science Books, Sausalito, 1999; Smith and March, March’s Advanced Organic Chemistry, 5th Edition, John Wiley & Sons, Inc., New York, 2001; Larock, Comprehensive Organic Transformations, VCH Publishers, Inc., New York, 1989; Carruthers, Some Modem Methods of Organic Synthesis, 3rd Edition, Cambridge University Press, Cambridge, 1987.
[0132] All ranges cited herein are inclusive, unless expressly stated to the contrary.
[0133] When a range of values is listed, it is intended to encompass each value and sub-range within the range. For example, “C1-6” is intended to encompass, C1, C2, C3, C4, C5, C6, C1-6, C1-5, C1-4, C1-3, C1-2, C2-6, C2-5, C2-4, C2-3, C3-6, C3-5, C3-4, C4-6, C4-5, and C5-6. For example, a heteroaromatic ring described as containing from “1 to 4 heteroatoms” means that the ring can contain 1, 2, 3 or 4 heteroatoms. It is also to be understood that any range cited herein includes within its scope all of the sub-ranges within that range. Thus, for example, a heterocyclic ring described as containing from “1 to 4 heteroatoms” is intended to include as aspects thereof, heterocyclic rings containing 2 to 4 heteroatoms, 3 or 4 heteroatoms, 1 to 3 heteroatoms, 2 or 3 heteroatoms, 1 or 2 heteroatoms, 1 heteroatom, 2 heteroatoms, 3 heteroatoms, or 4 heteroatoms.
[0134] When any variable occurs more than one time in any constituent or in Formula (I) or in any other formula depicting and describing the compounds of the present disclosure, its definition at each occurrence is independent of its definition at every other occurrence. Also, combinations of substituents and / or variables are permissible only if such combinations result in stable compounds.
[0135] As used herein, the term “alkyl” refers to a linear or branched chain saturated hydrocarbon group. The term “Ci-j alkyl” refers to an alkyl having i to j carbon atoms. Alkyl groups may contain 1 to 10 carbon atoms, unless otherwise stated. In certain embodiments, alkyl groups contain 1 to 6 carbon atoms (C1-6) , such as, 1 to 5 carbon atoms (C1-5) , 1 to 4 carbon atoms (C1-4) , 1 to 3 carbon atoms (C1-3) , or 1 to 2 carbon atoms (C1-2) . Non-limiting examples of alkyl groups include methyl, ethyl, n-and iso-propyl, n-, sec-, iso-, and tert-butyl, neopentyl, and the like. Alkyl groups may be optionally substituted (i.e., unsubstituted or substituted) , as valency permits, with one, two, three, or, in the case of alkyl groups of two carbons or more, four or more substituents independently selected from the group consisting of: amino; alkoxy; aryl; aryloxy; azido; cycloalkyl; cycloalkyloxy; cycloalkenyl; cycloalkynyl; halogen; heterocyclyl; (heterocyclyl) oxy; heteroaryl; hydroxy; nitro; thiol; silyl; cyano; alkylmercapto; alkylsulfonyl; alkylsulfinyl; alkylsulfenyl; =O; =S; -C (O) R or -SO2R, in which R is amino; and =NR’, in which R’ is H, alkyl, aryl, or heterocyclyl. Each of the substituents may itself be unsubstituted or, as valency permits, substituted with unsubstituted substituent (s) defined herein for each respective group. In certain embodiments, alkyl groups may be optionally substituted with one or more substitutes selected from halogen, C1-4 alkyloxy, C1-4 haloalkyloxy, and C1-4 haloalkylmercapto.
[0136] As used herein, the terms “alkylene” and “alkanediyl” are used interchangeably and refer to a divalent substituent that is a monovalent alkyl having one hydrogen atom replaced with a valency. Alkylene / alkanediyl groups may be unsubstituted or substituted. An optionally substituted alkylene / alkanediyl is an alkylene / alkanediyl that is optionally substituted as described herein for alkyl.
[0137] As used herein, the term “alkenyl” refers to a linear or branched-chain hydrocarbon radical having at least one (such as one, two, or three) carbon-carbon double bond, which may be optionally substituted (i.e., unsubstituted or substituted) independently with one or more substituents described herein, and includes radicals having “cis” and “trans” orientations, or alternatively, “E” and “Z” orientations. Alkenyl groups may contain 2 to 10 carbon atoms, unless otherwise stated. In certain embodiments, alkenyl groups may contain 2 to 6 carbon atoms, such as 2 to 5 carbon atoms, 2 to 4 carbon atoms, 2 to 3 carbon atoms. In certain embodiments, alkenyl groups contain 2 carbon atoms. Non-limiting examples of alkenyl groups include ethylenyl (vinyl) , propenyl, butenyl, pentenyl, 1-methyl-2-buten-1-yl, 5-hexenyl, and the like. An optionally substituted alkenyl is an alkenyl that is optionally substituted as described herein for alkyl.
[0138] As used herein, the term “alkenylene” refers to a divalent substituent that is a monovalent alkenyl having one hydrogen atom replaced with a valency. Alkenylene groups may be unsubstituted or substituted. An optionally substituted alkenylene is an alkenylene that is optionally substituted as described herein for alkyl.
[0139] As used herein, the term “alkynyl” refers to a linear or branched hydrocarbon radical having at least one (such as one, two, or three) carbon-carbon triple bond, which may be optionally substituted (i.e., unsubstituted or substituted) independently with one or more substituents described herein. Alkynyl groups may contain 2 to 10 carbon atoms, unless otherwise stated. In certain embodiments, alkynyl groups may contain 2 to 6 carbon atoms, such as 2 to 5 carbon atoms, 2 to 4 carbon atoms, 2 to 3 carbon atoms. In certain embodiments, alkynyl groups contain 2 carbon atoms. Non-limiting examples of alkynyl groups include ethynyl, 1-propynyl, 2-propynyl, and the like. An optionally substituted alkynyl is an alkynyl that is optionally substituted as described herein for alkyl.
[0140] As used herein, the term “alkynylene” refers to a divalent substituent that is a monovalent alkynyl having one hydrogen atom replaced with a valency. Alkynylene groups may be unsubstituted or substituted. An optionally substituted alkynylene is an alkynylene that is optionally substituted as described herein for alkyl.
[0141] As used herein, the term “cycloalkyl” refers to a partially or fully saturated, monocyclic, or polycyclic carbocyclic ring, which may include fused (when fused with an aryl or a heteroaryl ring, the cycloalkyl is bonded through a non-aromatic ring atom) , spiro, or bridged ring systems. In some embodiments, the cycloalkyl is fully saturated. Cycloalkyl groups may contain 3 to 10 ring forming carbon atoms, unless otherwise stated. In certain embodiments, cycloalkyl groups may contain 3 to 8 ring forming carbon atoms, such as 3 to 7 ring forming carbon atoms, 3 to 6 ring forming carbon atoms, 3 to 5 ring forming carbon atoms, 3 to 4 ring forming carbon atoms, 3 ring forming carbon atoms, 4 ring forming carbon atoms, 5 ring forming carbon atoms, 6 ring forming carbon atoms, 7 ring forming carbon atoms, 8 ring forming carbon atoms, and the like. Particularly, cycloalkyl groups may be monocyclic or bicyclic. Alternatively, bicyclic cycloalkyl groups may include fused, spiro, and bridged cycloalkyl structures. Non-limiting examples of cycloalkyl groups include cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, 1-bicyclo [2.2.1. ] heptyl, 2-bicyclo [2.2.1. ] heptyl, 5-bicyclo [2.2.1. ] heptyl, 7-bicyclo [2.2.1. ] heptyl, and decalinyl. The cycloalkyl group may be optionally substituted (i.e., unsubstituted or substituted) with one, two, three, four, or five substituents independently selected from the group consisting of: alkyl; alkenyl; alkynyl; alkoxy; alkylmercapto; alkylsulfinyl; alkylsulfenyl; alkylsulfonyl; amino; aryl; aryloxy; azido; cycloalkyl; cycloalkyloxy; cycloalkenyl; cycloalkynyl; halogen; heteroalkyl; heteroalkenyl; heteroalkynyl; heterocyclyl; (heterocyclyl) oxy; heteroaryl; hydroxy; nitro; thiol; silyl; cyano; =O; =S; -SO2R, in which R is optionally substituted amino; =NR’, in which R’ is H, alkyl, aryl, or heterocyclyl; and -CON (R″) 2, in which each R″is independently H or alkyl, or both R″, together with the atom to which they are attached, combine to form heterocyclyl. Each of the substituents may itself be unsubstituted or substituted with unsubstituted substituent (s) defined herein for each respective group. In certain embodiments, cycloalkyl groups may be optionally substituted with one or more substitutes selected from C1-4 alkyl, halogen, C1-4 alkyloxy, C1-4 haloalkyloxy, and C1-4 haloalkylmercapto.
[0142] As used herein, the terms “cycloalkylene” and “cycloalkanediyl” are used interchangeably and refer to a divalent substituent that is a cycloalkyl having one hydrogen atom replaced with a valency. Cycloalkylene / cycloalkanediyl groups may be unsubstituted or substituted. An optionally substituted cycloalkylene / cycloalkanediyl is a cycloalkylene / cycloalkanediyl that is optionally substituted as described herein for cycloalkyl.
[0143] As used herein, the term “heterocyclyl” refers to a monocyclic, bicyclic, tricyclic, or tetracyclic ring system having fused, bridged, and / or spiro 3-to 10-membered rings, unless otherwise stated, containing one, two, three, or four heteroatoms independently selected from the group consisting of nitrogen, oxygen, and sulfur as ring forming atoms. In certain embodiments, heterocyclyl groups may be 3-, 4-, 5-, 6-, 7-, 8-, 9-, or 10-membered. In certain embodiments, heterocyclyl groups may be 3-to 9-membered, 3-to 8-membered, 3-to 6-membered, 4-to 10-membered, 4-to 8-membered, 4-to 6-membered, or 5-to 8-membered. In certain embodiments, heterocyclyl groups may contain one, two, or three heteroatoms. In certain embodiments, heterocyclyl may be a monocyclic, bicyclic, tricyclic, or tetracyclic ring system having fused or bridged 5-, 6-, 7-, or 8-membered rings, containing one, two, three, or four heteroatoms independently selected from the group consisting of nitrogen, oxygen, and sulfur. Heterocyclyl can be aromatic or non-aromatic. In certain embodiments, heterocyclyl is non-aromatic. In certain embodiments, non-aromatic 5-membered heterocyclyl has zero or one double bonds, non-aromatic 6-and 7-membered heterocyclyl groups have zero to two double bonds, and non-aromatic 8-membered heterocyclyl groups have zero to two double bonds and / or zero or one carbon-carbon triple bond. In certain embodiments, heterocyclyl is a saturated ring. In certain embodiments, heterocyclyl groups may include up to 9 carbon atoms. Non-aromatic heterocyclyl groups include pyrrolinyl, pyrrolidinyl, pyrazolinyl, pyrazolidinyl, imidazolinyl, imidazolidinyl, piperidinyl, homopiperidinyl, piperazinyl, pyridazinyl, oxazolidinyl, isoxazolidiniyl, morpholinyl, thiomorpholinyl, thiazolidinyl, isothiazolidinyl, thiazolidinyl, tetrahydrofuranyl, dihydrofuranyl, tetrahydrothienyl, dihydrothienyl, dihydroindolyl, tetrahydroquinolyl, tetrahydroisoquinolyl, pyranyl, dihydropyranyl, dithiazolyl, and the like. If the heterocyclic ring system has at least one aromatic resonance structure or at least one aromatic tautomer, such structure is an aromatic heterocyclyl (i.e., heteroaryl) . Non-limiting examples of heteroaryl groups include benzimidazolyl, benzofuryl, benzothiazolyl, benzothienyl, benzoxazolyl, furyl, imidazolyl, indolyl, isoindazolyl, isoquinolinyl, isothiazolyl, isothiazolyl, isoxazolyl, oxadiazolyl, oxazolyl, purinyl, pyrrolyl, pyridinyl, pyrazinyl, pyrimidinyl, qunazolinyl, quinolinyl, thiadiazolyl (e.g., 1, 3, 4-thiadiazole) , thiazolyl, thienyl, triazolyl, tetrazolyl, and the like. The term “heterocyclyl” also includes a heterocyclic compound having a bridged multicyclic structure in which one or more carbons and / or heteroatoms bridges two non-adjacent members of a monocyclic ring, e.g., quinuclidine, tropanes, or diaza-bicyclo [2.2.2] octane. The term “heterocyclyl” includes bicyclic, tricyclic, and tetracyclic groups in which any of the above heterocyclic rings is fused to one, two, or three carbocyclic rings, e.g., an aryl ring, a cyclohexane ring, a cyclohexene ring, a cyclopentane ring, a cyclopentene ring, or another monocyclic heterocyclic ring. Examples of fused heterocyclyl groups include 1, 2, 3, 5, 8, 8a-hexahydroindolizine; 2, 3-dihydrobenzofuran; 2, 3-dihydroindole; and 2, 3-dihydrobenzothiophene. The heterocyclyl group may be unsubstituted or substituted with one, two, three, four or five substituents independently selected from the group consisting of: alkyl; alkenyl; alkynyl; alkoxy; alkylsulfinyl; alkylsulfenyl; alkylsulfonyl; amino; aryl; aryloxy; azido; cycloalkyl; cycloalkoxy; cycloalkenyl; cycloalkynyl; halogen; heteroalkyl; heterocyclyl; (heterocyclyl) oxy; heteroaryl; hydroxy; nitro; thiol; silyl; cyano; -C (O) R or -SO2R, where R is amino or alkyl; =O; =S; =NR’, where R’ is H, alkyl, aryl, or heterocyclyl. Each of the substituents may itself be unsubstituted or substituted with unsubstituted substituent (s) defined herein for each respective group. In certain embodiments, heterocyclyl groups may be optionally substituted with one or more substitutes selected from 4-to 10-membered heterocyclyl, 6-to 10-membered aryl, and 5-to 10-membered heteroaryl.
[0144] As used herein, the term “heterocyclylene” refers to a divalent substituent that is an heterocyclyl having one hydrogen atom replaced with a valency. Heterocyclylene groups may be unsubstituted or substituted. An optionally substituted heterocyclylene is an heterocyclylene that is optionally substituted as described herein for heterocyclyl.
[0145] As used herein, the term “aryl” refers to a mono-, bicyclic, or multicyclic carbocyclic ring system having at least one aromatic rings. Aryl groups may be 6-to 10-membered, unless otherwise stated. In certain embodiments, aryl groups may contain 6 ring forming carbon atoms. All ring forming atoms within a carbocyclic aryl group are carbon atoms. Non-limiting examples of aryl groups include phenyl, naphthyl, 1, 2-dihydronaphthyl, 1, 2, 3, 4-tetrahydronaphthyl, fluorenyl, indanyl, indenyl, and the like. In certain embodiments, aryl is phenyl or naphthyl. In certain embodiments, aryl is phenyl. In the context of the present specification, the terms “aryl” and “aromatic ring” may be used interchangeably. Aryl groups may be unsubstituted or substituted. An optionally substituted aryl group may be an aryl optionally substituted with one, two, three, four, or five substituents independently selected from the group consisting of: alkyl; alkenyl; alkynyl; alkoxy; alkylsulfinyl; alkylsulfenyl; alkylsulfonyl; amino; aryl; aryloxy; azido; cycloalkyl; cycloalkoxy; cycloalkenyl; cycloalkynyl; halogen; heteroalkyl; heteroalkenyl; heteroalkynyl; heterocyclyl; (heterocyclyl) oxy; heteroaryl; hydroxy; nitro; thiol; silyl; - (CH2) n-C (O) OR’; -C (O) R; and -SO2R, in which R is amino or alkyl, R’ is H or alkyl, and n is 0 or 1. Each of the substituents may itself be unsubstituted or substituted with unsubstituted substituent (s) defined herein for each respective group. In certain embodiments, aryl groups may be optionally substituted with one or more substitutes selected from 4-to 10-membered heterocyclyl, C6-10 aryl, and 5-to 10-membered heteroaryl.
[0146] As used herein, the term “arylene” refers to a divalent substituent that is an aryl having one hydrogen atom replaced with a valency. Arylene groups may be unsubstituted or substituted. An optionally substituted arylene is an arylene that is optionally substituted as described herein for aryl.
[0147] As used herein, the term “heteroaryl” refers to a monocyclic ring system, or a fused or bridged bicyclic ring system, in which the ring system contains one, two, three, or four heteroatoms independently selected from the group consisting of nitrogen, oxygen, and sulfur; and at least one of the rings is an aromatic ring. Heteroaryl groups may be 5-to 10-membered, unless otherwise stated. In certain embodiments, heteroaryl groups may be a 5-to 6-membered heteroaryl ring having 1 to 3 heteroatoms independently selected from nitrogen, oxygen, and sulfur; or an 8-to 10-membered bicyclic heteroaryl ring having 1 to 4 heteroatoms independently selected from nitrogen, oxygen, and sulfur. In certain embodiments, heteroaryl groups may contain one, two, or three heteroatoms. In certain embodiments, heteroaryl groups may contain one or two heteroatoms. Non-limiting examples of heteroaryl groups include benzimidazolyl, benzofuryl, benzothiazolyl, benzothienyl, benzoxazolyl, furyl, imidazolyl, indolyl, isoindazolyl, isoquinolinyl, isothiazolyl, isothiazolyl, isoxazolyl, oxadiazolyl, oxazolyl, purinyl, pyrrolyl, pyridinyl, pyrazinyl, pyrimidinyl, qunazolinyl, quinolinyl, thiadiazolyl, thiazolyl, thienyl, triazolyl, tetrazolyl, dihydroindolyl, tetrahydroquinolyl, tetrahydroisoquinolyl, and the like. Heteroaryl groups include at least one ring having at least one heteroatom as described above and at least one aromatic ring. For example, a ring having at least one heteroatom may be fused to one, two, or three carbocyclic rings, e.g., an aryl ring, a cyclohexane ring, a cyclohexene ring, a cyclopentane ring, a cyclopentene ring, or another monocyclic heterocyclic ring. Non-limiting examples of fused heteroaryl groups include 1, 2, 3, 5, 8, 8a-hexahydroindolizine, 2, 3-dihydrobenzofuran, 2, 3-dihydroindole, 2, 3-dihydrobenzothiophene, and the like. In the context of the present disclosure, the terms “heteroaryl” and “heteroaromatic ring” may be used interchangeably. Heteroaryl groups may be unsubstituted or substituted. An optionally substituted heteroaryl group may be a heteroaryl optionally substituted with one, two, three, four, or five substituents independently selected from the group consisting of: alkyl; alkenyl; alkynyl; alkoxy; alkylsulfinyl; alkylsulfenyl; alkylsulfonyl; amino; aryl; aryloxy; azido; cycloalkyl; cycloalkoxy; cycloalkenyl; cycloalkynyl; halogen; heteroalkyl; heteroalkenyl; heteroalkynyl; heterocyclyl; (heterocyclyl) oxy; heteroaryl; hydroxy; nitro; thiol; silyl; - (CH2) n-C (O) OR’; -C (O) R; and -SO2R, in which R is amino or alkyl, R’ is H or alkyl, and n is 0 or 1. Each of the substituents may itself be unsubstituted or substituted with unsubstituted substituent (s) defined herein for each respective group. In certain embodiments, heteroaryl groups may be optionally substituted with one or more substitutes selected from 4-to 10-membered heterocyclyl, C6-10 aryl, and 5-to 10-membered heteroaryl.
[0148] As used herein, the term “heteroarylene” refers to a divalent substituent that is a heteroaryl having one hydrogen atom replaced with a valency. Heteroarylene groups may be unsubstituted or substituted. An optionally substituted heteroarylene is a heteroarylene that is optionally substituted as described herein for heteroaryl.
[0149] As used herein, the term “heteroatom” refers to nitrogen, oxygen, or sulfur, and may include any oxidized form of nitrogen or sulfur, and any quaternized form of a basic nitrogen.
[0150] As used herein, the term “oxo” refers to a divalent oxygen atom and the structure of oxo may be shown as =O.
[0151] As used herein, the term “halogen” (or “halo” ) refers to fluoride, chloride, bromide, and iodide. In certain embodiments, non-limiting examples of halogen include fluoride, chloride, and bromide. In certain embodiments, halogen is chloride or bromide. In certain embodiments, halogen is fluoride.
[0152] As used herein, the term “haloalkyl” refers to an alkyl group as described herein in which one or more of hydrogen atoms have been replaced with one or more halogen atoms independently selected from the group consisting of fluoride, chloride, bromide, and iodide. When a haloalkyl contains more than one halogen atom, the halogen atoms can be the same or be different from each other. Non-limiting examples of haloalkyl groups include -CH2F, -CHF2, -CF3, -CF2Cl, -CH2CF3, -CF2CF3, and the like. In certain embodiments, haloalkyl groups may be perhaloalkyl groups, such as perfluoroalkyl.
[0153] As used herein, the terms “haloalkylene” and “haloalkanediyl” a divalent substituent that is a haloalkyl having one hydrogen atom replaced with a valency. Non-limiting examples of haloalkylene / haloalkanediyl groups include -CHF-, -CF2-, -CH (CF3) -, -CH2CHF-, -CHFCHF-, and the like. In certain embodiments, haloalkylene / haloalkanediyl groups may be perhalo haloalkylene / perhaloalkanediyl groups, such as perfluoro haloalkylene / perhaloalkanediyl.
[0154] As used herein, the term “substituted” , when refers to a chemical group, means that the chemical group has one or more hydrogen atoms that is / are removed and replaced by substituents. The term “substituent” as used herein has the ordinary meaning known in the art and refers to a chemical moiety that is covalently attached to, or if appropriate, fused to, a parent group. It is to be understood that substitution at a given atom is limited by valency. It is understood that the substituent can be further substituted.
[0155] As used herein, the term “optionally substituted” means that the chemical group may have no substituents (i.e., unsubstituted) or may have one or more substituents (i.e., substituted) . It is to be understood that substitution at a given atom is limited by valency.
[0156] The compounds provided herein are described with reference to both generic formulas and specific compounds. In addition, the compounds of the present disclosure may exist in a number of different forms or derivatives, all within the scope of the disclosure. These include, for example, pharmaceutically acceptable salts, tautomers, stereoisomers, racemic mixtures, regioisomers, prodrugs, and active metabolites, and the like. In certain embodiments, the compounds of the disclosure may contain bonds with hindered rotation such that two separate rotomers, or atropisomer, may be separated and may have advantageous biological activity. It is intended that all of the possible atropisomes are included with the scope of this disclosure.
[0157] As used herein, the term "approximately" or "about" means a quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length range of ±15%, ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, or ±1%with respect to a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length.
[0158] Unless otherwise specified, any concentration range, percentage range, ratio range, or integer range shall be understood to include any integer value within said range and, where appropriate, fractions thereof (such as tenths and hundredths of an integer) . When immediately preceded by a numeric or numeric value, the term "approximately" refers to plus or minus 10 percent of the numeric or numeric range.
[0159] Throughout the specification, unless the context otherwise requires, the term "comprise / comprises / comprising" should be understood to mean the inclusion of a specified step, or element, or group of steps or elements, but does not exclude any other step, or element, or group of steps or elements. In particular embodiments, the terms "includes" , "has" , "contains" and "comprises" are used synonymously.
[0160] As used herein, the term "antibody" refers to any form of immunoglobulin molecule that exhibits the desired biological or binding activity. Accordingly, it is used in the broadest sense and specifically covers, but is not limited to, monoclonal antibodies (including full-length monoclonal antibodies) , polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies) , humanized, fully human antibodies, chimeric antibodies and camelidized single domain antibodies.
[0161] As used herein, the term "antibody" includes not only an intact polyclonal or monoclonal antibody, but also, unless otherwise specified, any antigen-binding portion thereof that competes with the intact antibody for specific binding, a fusion protein comprising an antigen-binding portion, and any other modified conformation of an immunoglobulin molecule comprising an antigen recognition site. Antigen-binding portions include, for example, Fab, Fab', F (ab') 2, Fd, Fv, structural domain antibodies (dAb, e.g. shark and camel antibodies) , fragments containing complementary determining regions (CDR) , single-chain variable fragment antibodies (scFv) , maximal antibodies, microbodies, intracellular antibodies, double antibodies, triple antibodies, quadruple antibodies, v-NAR and double scFv, and immune peptides containing at least a portion of an immunoglobulin (sufficient to confer peptide-specific antigen binding) . Antibodies include any class of antibody, such as IgG, IgA or IgM (or subclasses thereof) , and the antibody need not be of any particular class.
[0162] As used herein, the term "nanobody" refers to a single structural domain antibody, which is fragment consisting of a single variable antibody structural domain. Like intact antibodies, they are able to bind selectively to specific antigens. With a molecular weight of only 12-15 kDa, single domain antibodies are much smaller than common antibodies (150-160 kDa) .
[0163] As used herein, the term "anti-PD-L1 antibody" or "anti-PD-1 antibody" refers to an antibody that blocks the binding of PD-L1 expressed on cancer cells to PD-1. In any therapeutic method, drug, and use of the present disclosure in which a human subject is being treated, the anti-PD-L1 antibody specifically binds human PD-L1 and blocks the binding of human PD-L1 to human PD-1, and the anti-PD-1 antibody specifically binds human PD-1 and blocks the binding of human PD-1 to human PD-L1. The antibody may be a monoclonal antibody, a human antibody, a humanized antibody, or a chimeric antibody, and may include a human constant region.
[0164] As used herein, the term "subject" means an animal, such as a mammal, including but not limited to a human, rodent, ape, feline, canine, equine, bovine, porcine, sheep, goat, mammalian laboratory animal, mammalian farm animal, mammalian sport animal, and mammalian pet. Subjects may be male or female and may be of any age, including infants, juveniles, young adults, adults, and elderly subjects. In some embodiments, a subject is an individual in need of treatment for a disease or condition. In some embodiments, the subject receiving the treatment may be a patient who has a condition associated with the treatment or is at risk of developing the condition. In particular embodiments, the subject is a human, such as a human patient. The term is often used interchangeably with "patient" , "test subject" , "treatment subject" , and the like.
[0165] As used herein, the term "pharmaceutically acceptable" indicates that the substance or composition must be chemically and / or toxicologically compatible with the other ingredients comprising the formulation and / or the mammal treated therewith. A "pharmaceutically acceptable carrier" includes any and all physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, absorption delay agents, and the like. Examples of pharmaceutically acceptable carriers include one or more of water, saline, phosphate buffered saline, dextrose, glycerol, ethanol, and the like., and combinations thereof.
[0166] In the prior art, the genetic code expansion technique was used to genetically encode the unnatural amino acid fluorosulfonate-L-tyrosine (FSY) into the epitope domain of human programmed cell death protein-1 (PD-1) , thereby further forming a covalent binding to PD-L1, which established the viability of in vivo application of the covalent protein. Furthermore, whereas most protein drugs must be modified to extend their half-life, the irreversible binding affinity of covalent drugs avoids this requirement because covalent binding decouples drug efficacy from pharmacokinetics. Furthermore, due to its unique mechanism, this could reduce off-target effects. Because covalent linking requires both drug-target binding and covalent warhead-natural residue pairing, the system gains enhanced specificity and target selectivity. As a result, it should provide a highly useful, general platform for converting a wide range of proteins into covalent binders. It is hoped that this technology will hasten the development of therapeutic applications affecting protein therapeutics. However, due to the difficulty of genetic code expansion technology and the lack of compatible high throughput selection methods, wide practical applications remain a challenge.
[0167] Covalent miniproteins are garnering attention as a promising therapeutic avenue. However, current iterations often require a prolonged duration to crosslink with target proteins, far exceeding their circulation half-life. This compromises their target engagement efficiency, thereby limiting their translational viability. Herein, the inventors present a high-throughput covalent protein selection system, integrating yeast display with chemoselective protein modification. This system enables the rapid screening and identification of covalent miniproteins with exceptionally swift (supercharged) kinetics for target crosslinking. Utilizing this system, the inventors developed a serial of PD-L1 targeting covalent nanobody, especially the PD-L1 targeting covalent nanobody, named as IB101 which has excellent crosslinking kinetics (Kinact) of 0.2 min-1 and a half-life (t1 / 2) of ~ 3 minutes. The inventors also engineered a series of SARS-CoV-2 RBD targeting proteins, especially the F30C-Mu2_5-NO2 protein has an excellent Kinact of 0.86 min-1 and a half-life (t1 / 2) of 0.8 minutes. The PD-L1 targeting covalent nanobodies of the present disclosure, especially IB101, achieve complete crosslinking with purified and cell surface PD-L1 within 15 minutes at low nanomolar concentrations. The PD-L1 targeting covalent nanobodies of the present disclosure also demonstrate superior tumor suppression activity compared to the clinically approved KN035-Fc antibody, despite their shorter in vivo half-life.
[0168] All publications and patents mentioned herein are hereby incorporated by reference in their entirety as if each individual publication or patent were specifically and separately indicated as not incorporated by reference. In the event of conflict, this application (including any definitions herein) shall control. However, any references, articles, publications, patents, patent publications and patent applications cited herein are not and shall not be deemed to be admissions or suggestions of any kind which constitute valid prior art or form part of the common knowledge of any country in the world.
[0169] The present disclosure is further described in the following examples, which are not limiting the scope of the present disclosure as described in the claims.
[0170] SEQUENCE LISTING
[0171] COVALENT PROTEINS
[0172] As used herein, the term "covalent protein" refers to a protein molecule comprising at least one covalent bond formed between an atom of said protein molecule, typically an atom within an amino acid residue (either side chain or backbone) , and an atom of a distinct chemical moiety. Said distinct chemical moiety may be endogenous or exogenous to the biological system from which the protein originates. Non-limiting examples of such distinct chemical moieties include: post-translational modification groups (such as phosphate groups, glycosyl groups, ubiquitin molecules, acetyl groups, methyl groups, lipid groups, or atoms participating in disulfide bonds) ; therapeutic agents or drug molecules (such as covalent inhibitors or inactivators) ; probe or label molecules; diagnostic agents; other biomolecules (such as other proteins, peptides, or nucleic acids, forming cross-links or conjugates) ; or synthetic linker molecules. The formation of said covalent bond may result from enzymatic activity, non-enzymatic biological processes, or chemical reaction with an external agent.
[0173] The present disclosure provides a covalent protein having a structure as shown in Formula (I) :
[0174] wherein
[0175] Ab represents a protein moiety, and R1 is covalently bonded to Ab;
[0176] R1 has a structure of -R11-, or -R11-C (O) -R12-R13-, where
[0177] R11 is optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN,
[0178] R12 is selected from the group consisting of NR12a, O, S, and heterocyclylene, R12a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0179] R13 is absent or optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN;
[0180] RA is selected from the group consisting of optionally substituted alkanediyl, and optionally substituted arenediyl;
[0181] Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, alkyl, haloalkyl, and aryl,
[0182] n is an integer selected from 0 and 1;
[0183] Ry is selected from the group consisting of S and P;
[0184] Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, alkyl, haloalkyl, and aryl;
[0185] The bond between Ry and Rz is a single bond or a double bond.
[0186] In some embodiments, R1 is covalently bonded to Ab through side chain of Cys residue.
[0187] In some embodiments, R1 has a structure of -R11-, or -R11-C (O) -R12-R13-.
[0188] In some embodiments, R11 is selected from C1-C8, preferably C1-C6, more preferably C1-C4 alkanediyl optionally substituted with one or more halo, -OH, or -CN. In some preferred embodiments, R11 is selected from the group consisting of -CH2-, -CH2CH2-, -CH2CH2CH2-, and -CH2CH2CH2CH2-, more preferably -CH2-and -CH2CH2-, and even more preferably -CH2-. It should be noted that R11 is the group connected to the protein moiety (Ab) .
[0189] In some embodiments, R12 is selected from the group consisting of NR12a, O, S, and 4-10 membered, preferably 4-8 membered, more preferably 5-6 membered heterocyclylene, for example, 5-or 6-membered heterocyclylene is selected from where R12a is selected from the group consisting of hydrogen, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 heteroalkyl, optionally substituted C3-C8, preferably C3-C6 cycloalkyl, optionally substituted C3-C8, preferably C3-C6 heterocyclyl, optionally substituted C6-C14, preferably C6-C10, more preferably C6 aryl, and optionally substituted 5-14 membered, preferably 5-10 membered, more preferably 5-6 membered heteroaryl.
[0190] In some embodiments, R13 is absent or optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN.
[0191] In some preferred embodiments, R1 is selected from C1-C4 alkanediyl and C1-C4 alkanediyl-C (O) -NH-. In some preferred embodiments, R1 is selected from the group consisting of -CH2-, -CH2CH2-, -CH2CH2CH2-, -CH2CH2CH2CH2-, -CH2-C (O) -NH-, -CH2CH2-C (O) -NH-, -CH2CH2CH2-C (O) -NH-, -CH2CH2CH2CH2-C (O) -NH-, more preferably -CH2-, -CH2CH2-, -CH2-C (O) -NH-, and -CH2CH2-C (O) -NH-, even more preferably, -CH2-and -CH2-C (O) -NH-. It should be noted that the left end of the above group (i.e., the terminal alkanediyl, -CH2-) is the connection position with the protein moiety (Ab) .
[0192] In some embodiments, RA is selected from the group consisting of optionally substituted C1-C8, preferably C1-C6 alkanediyl, and optionally substituted C6-C14, preferably C6-C10, more preferably C6 arenediyl.
[0193] In some preferred embodiments, RA is C1-C8, preferably C1-C6 alkanediyl optionally substituted with one or more halo, -OH, or -CN.
[0194] In some preferred embodiments, RA is C6-C14, preferably C6-C10, more preferably C6 arenediyl optionally substituted with one or more, for example one, two, three, or four Rx, where Rx is selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions. In some embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, and N (Rx1) (Rx2) , where each of Rx1 and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl. In some embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1 and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl. In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0195] In some embodiments, Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, n is 0, then Rw is absent. In some embodiments, n is 1.
[0196] In some embodiments, Ry is S. In some embodiments, Ry is P.
[0197] In some embodiments, Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. The bond between Ry and Rz is a single bond or a double bond. One of ordinary skill in the art will recognize that the bond order between Ry and Rz, specifically whether said bond is a single bond or a double bond, is determined by the chemical identity of the substituents Ry and Rz, as previously defined. The formation of such a bond shall conform to the established principles of valence bond theory, thereby ensuring appropriate chemical valency for the atoms involved.
[0198] In some embodiments, Rz is selected from the group consisting of =O, and =N (Rz2) , the bond between Ry and Rz is a double bond, where Rz2 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, Rz2 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0199] In some embodiments, Rz is selected from the group consisting of -O (Rz1) , and -N (Rz3) (Rz4) , the bond between Ry and Rz is a single bond, where each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0200] In some preferred embodiments, Ry is S, and Rz is selected from the group consisting of =O, and =N (Rz2) , the bond between Ry and Rz is a double bond, where Rz2 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, Rz2 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0201] In some embodiments, Ry is P, and Rz is selected from the group consisting of -O (Rz1) , and -N (Rz3) (Rz4) , the bond between Ry and Rz is a single bond, where each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0202] In some embodiments, RA is C1-C8, preferably C1-C6 alkanediyl optionally substituted with one or more halo, -OH, or -CN. Therefore, the present disclosure provides a covalent protein having a structure as shown in Formula (I-A) :
[0203] wherein,
[0204] p is an integer selected from 1 to 8, preferably 1 to 6, - (CH2) p-optionally substituted with one or more halo, -OH, or -CN;
[0205] each of Ab, R1, Rw, Ry, Rz and n is independently defined as disclosed herein, for example, as defined in the Formula (I) .
[0206] It should be understood that the term "- (CH2) p-" as used herein refers to an alkylene group comprising p carbon atoms, which may be configured as a straight chain or a branched chain structure, provided that the total number of carbon atoms in the chain, whether straight or branched, equals p.
[0207] In some preferred embodiments, R1 has a structure of -R11-C (O) -R12-R13-.
[0208] In some preferred embodiments, R11 is C1-C4 alkanediyl, for example, R11 is selected from the group consisting of -CH2-, -CH2CH2-, -CH2CH2CH2-, and -CH2CH2CH2CH2-, more preferably -CH2-and -CH2CH2-, and even more preferably -CH2-.
[0209] In some preferred embodiments, R12 is NR12a, where R12a is selected from the group consisting of hydrogen, C1-C4 alkyl, for example, R12 is NH.
[0210] In some preferred embodiments, R13 is absent.
[0211] In some preferred embodiments, R1 has a structure of -CH2-C (O) -NH-.
[0212] In some preferred embodiments, p is an integer selected from 1 to 8, preferably 1 to 6, specifically, p is 1, 2, 3, 4, 5, or 6.
[0213] In some preferred embodiments, Rw is absent, or is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, Rw1 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0214] In some preferred embodiments, Ry is S or P.
[0215] In some preferred embodiments, Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0216] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-A1) - (I-A3) :
[0217] wherein,
[0218] p is an integer selected from 1 to 8, preferably 1 to 6, specifically, p is 1, 2, 3, 4, 5, or 6;
[0219] Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl; preferably, Rw1 is selected from the group consisting of H, methyl, ethyl, and phenyl;
[0220] each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl; preferably, each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0221] In some embodiments, RA is optionally substituted C6-C14, preferably C6-C10, more preferably C6 arenediyl. In some embodiments, RA is optionally substituted phenylene. Therefore, the present disclosure provides a covalent protein having a structure as shown in Formula (I-B) :
[0222] wherein
[0223] m is an integer selected from 1 to 4, each of Ab, R1, Rw, Rx, Ry, Rz, and n is independently defined as disclosed herein, for example, as defined in the Formula (I) .
[0224] In some preferred embodiments, R1 has a structure of -R11-C (O) -R12-R13-.
[0225] In some preferred embodiments, R11 is C1-C4 alkanediyl, for example, R11 is selected from the group consisting of -CH2-, -CH2CH2-, -CH2CH2CH2-, and -CH2CH2CH2CH2-, more preferably -CH2-and -CH2CH2-, and even more preferably -CH2-.
[0226] In some preferred embodiments, R12 is NR12a, where R12a is selected from the group consisting of hydrogen, C1-C4 alkyl, for example, R12 is NH.
[0227] In some preferred embodiments, R13 is absent.
[0228] In some preferred embodiments, R1 is selected from the group consisting of -CH2-, and -CH2-C (O) -NH-.
[0229] In some preferred embodiments, m is an integer selected from 1 to 4, specifically, m is 1, 2, 3, or 4, preferably m is 1. Each Rx is independently selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions. In some preferred embodiments, each Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl. In some preferred embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl. In some preferred embodiments, each Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0230] In some preferred embodiments, Rw is absent, or is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, Rw1 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0231] In some preferred embodiments, Ry is S or P.
[0232] In some preferred embodiments, Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0233] The bond between Ry and Rz is a single bond or a double bond.
[0234] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-B-a) - (I-B-c) :
[0235] wherein
[0236] each of Ab, R1, Rw, Rx, Ry, Rz, m and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0237] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0238] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-B1) - (I-B4) :
[0239] wherein,
[0240] each of Ab, R1, Rx, Rw1, Rz1, Rz2, Rz3, Rz4, m and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0241] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-B1-a) - (I-B1-c) :
[0242] wherein,
[0243] each of Ab, R1, Rx, m and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0244] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0245] In some preferred embodiments, the covalent protein having a structure as shown in Formula (I-B1) is selected from the group consisting of
[0246] wherein,
[0247] each of Ab, Rx, and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0248] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0249] In some preferred embodiments, n is 0 or 1. In some preferred embodiments, n is 0. In some preferred embodiments, n is 1.
[0250] In some preferred embodiments, the covalent protein having a structure as shown in Formula (I-B1) is selected from the group consisting of
[0251] wherein,
[0252] each of Ab and Rx is defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0253] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0254] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-B2-a) - (I-B2-c) :
[0255] wherein,
[0256] each of Ab, R1, Rx, Rz2, m and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0257] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0258] In some preferred embodiments, the covalent protein having a structure as shown in Formula (I-B2) is selected from the group consisting of
[0259] wherein,
[0260] each of Ab, Rx, Rz2, and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0261] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0262] In some preferred embodiments, n is 0 or 1. In some preferred embodiments, n is 0. In some preferred embodiments, n is 1.
[0263] In some preferred embodiments, Rz2 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0264] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-B3-a) - (I-B3-c) :
[0265] wherein,
[0266] each of Ab, R1, Rx, Rz3, Rz4, and m is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0267] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0268] In some preferred embodiments, the covalent protein having a structure as shown in Formula (I-B3) is selected from the group consisting of
[0269] wherein,
[0270] each of Rx, Rz3, and Rz4 is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0271] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0272] In some preferred embodiments, each of Rz3 and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl. In some preferred embodiments, Rz3 and Rz4 are methyl.
[0273] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-B4-a) - (I-B4-c) :
[0274] wherein,
[0275] each of Ab, R1, Rx, Rw1, Rz1, and m is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0276] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0277] In some preferred embodiments, the covalent protein having a structure as shown in Formula (I-B4) is selected from the group consisting of
[0278] wherein,
[0279] each of Ab, Rx, Rw1, and Rz1 is independently defined as disclosed herein, for example, as defined in the Formula (I-B) .
[0280] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0281] In some preferred embodiments, each of Rw1 and Rz1 is independently selected from the group consisting of H, methyl, ethyl, and phenyl. In some preferred embodiments, Rw1 is methyl.
[0282] In one aspect, the present disclosure provides a covalent protein having a structure as shown in Formula (I-B1) :
[0283] wherein Ab represents a protein moiety, and R1 is covalently bonded to Ab;
[0284] R1 is a linker moiety having a structure of -R11-, or -R11-C (O) -R12-R13-, where
[0285] R11 is optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN,
[0286] R12 is selected from the group consisting of NR12a, O, S, heterocyclylene, and heteroarylene, R12a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0287] R13 is absent or optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN;
[0288] Rx is selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions, m is an integer selected from 1 to 4;
[0289] n is an integer selected from 0 and 1.
[0290] In some embodiments, R1 is covalently bonded to Ab through side chain of Cys residue.
[0291] In some embodiments, R1 has a structure of -R11-, or -R11-C (O) -R12-R13-.
[0292] In some embodiments, R11 is selected from C1-C8, preferably C1-C6, more preferably C1-C4 alkanediyl optionally substituted with one or more halo, -OH, or -CN. In some preferred embodiments, R11 is selected from the group consisting of -CH2-, -CH2CH2-, -CH2CH2CH2-, and -CH2CH2CH2CH2-, more preferably -CH2-and -CH2CH2-, and even more preferably -CH2-. It should be noted that R11 is the group connected to the protein moiety (Ab) .
[0293] In some embodiments, R12 is selected from the group consisting of NR12a, O, S, and 4-10 membered, preferably 4-8 membered, more preferably 5-6 membered heterocyclylene, for example, 5-or 6-membered heterocyclylene is selected from where R12a is selected from the group consisting of hydrogen, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 heteroalkyl, optionally substituted C3-C8, preferably C3-C6 cycloalkyl, optionally substituted C3-C8, preferably C3-C6 heterocyclyl, optionally substituted C6-C14, preferably C6-C10, more preferably C6 aryl, and optionally substituted 5-14 membered, preferably 5-10 membered, more preferably 5-6 membered heteroaryl.
[0294] In some embodiments, R13 is absent or optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN.
[0295] In some preferred embodiments, R1 is selected from the group consisting of -CH2-, -CH2CH2-, -CH2CH2CH2-, -CH2CH2CH2CH2-, -CH2-C (O) -NH-, -CH2CH2-C (O) -NH-, -CH2CH2CH2-C (O) -NH-, -CH2CH2CH2CH2-C (O) -NH-, more preferably -CH2-, -CH2CH2-, -CH2-C (O) -NH-, and -CH2CH2-C (O) -NH-, even more preferably, -CH2-and -CH2-C (O) -NH-. It should be noted that the left end of the above group (i.e., the terminal alkanediyl, -CH2-) is the connection position with the protein moiety (Ab) .
[0296] In some embodiments, m is an integer selected from 1 to 4, specifically, m is 1, 2, 3, or 4, preferably m is 1. Each Rx is independently selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions. In some embodiments, each Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl. In some embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl. In some preferred embodiments, each Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0297] In some preferred embodiments, the covalent protein has a structure as shown in Formula (I-B1-a) - (I-B1-c) :
[0298] wherein,
[0299] each of Ab, R1, Rx, m and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B1) .
[0300] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0301] In some embodiments, the covalent protein has a structure as shown in Formula (I-B1-a11) - (I-B1-c1) , or Formula (I-B1-a2) - (I-B1-c2) :
[0302] wherein,
[0303] each of Ab, Rx, m and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B1) . In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0304] In some embodiments, the covalent protein having a structure as shown in Formula (I-B1) is selected from the group consisting of
[0305] wherein,
[0306] each of Ab, Rx, and n is independently defined as disclosed herein, for example, as defined in the Formula (I-B1) .
[0307] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0308] In some preferred embodiments, n is 0 or 1. In some preferred embodiments, n is 0. In some preferred embodiments, n is 1.
[0309] In some embodiments, the covalent protein having a structure as shown in Formula (I-B1) is selected from the group consisting of
[0310] wherein,
[0311] Ab and Rx are defined as disclosed herein, for example, as defined in the Formula (I-B1) .
[0312] In some embodiments, Rx is selected from the group consisting of H, F, CN, and NO2, or similar functional groups. In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0313] In some embodiments, the covalent protein having a structure as shown in Formula (I-B1) is selected from the group consisting of
[0314] wherein, Ab is as provided herein.
[0315] In the present disclosure, the protein moiety (Ab) in the covalent protein having the structure shown in Formula (I) (the protein moiety (Ab) is also referred to as "binding protein" or "therapeutic protein" in the present disclosure) refers to a polypeptide or protein that can specifically or non-specifically interact with a predetermined target protein and contains at least one naturally occurring or engineered cysteine (Cys) residue in its amino acid sequence. The side chain thiol group (-SH group) of said cysteine residue is capable of forming a covalent bond with R1 group of the crosslinker moiety in Formula (I) . Advantageously, the cysteine residue is spatially located at or near the binding interface region between the protein moiety (Ab) and the target protein. Consequently, after the protein moiety (Ab) is conjugated with R1 group of the crosslinker moiety, it can effectively direct the crosslinker to one or more reactive groups on the target protein, thereby promoting the formation of a stable covalent bond between them. The conjugation of the protein moiety (Ab) with the crosslinker, and the subsequent covalent binding to the target protein, is intended to enhance the durability and specificity of binding, alter pharmacokinetic properties, or achieve other beneficial biological effects. Meanwhile, the introduction of the crosslinker preferably does not completely eliminate or significantly impair the initial non-covalent binding ability of the protein moiety (Ab) to the target protein.
[0316] As used herein, the term "target protein" refers to any naturally occurring, recombinant, or synthetic protein or polypeptide that is recognized and bound, at least initially through non-covalent interactions, by the protein moiety (Ab) component of a covalent conjugate (protein-crosslinker) . The target protein possesses one or more accessible reactive functional groups, such as, but not limited to, the side chains of amino acid residues (e.g., lysine, cysteine, histidine, tyrosine, serine, threonine) or an N-terminal α-amino group. These reactive functional groups on the target protein are capable of forming a covalent bond with the crosslinker. The target protein can encompass a wide range of biomolecules, including, but not limited to, cell surface proteins (such as receptors or immune checkpoint molecules like PD-L1) , soluble proteins (such as cytokines, chemokines, or growth factors) , viral proteins (such as SARS-CoV-2 RBD) , bacterial proteins, enzymes, enzyme substrates, enzyme inhibitors, antibodies, or any other proteinaceous molecule whose activity, localization, or properties are intended to be modulated or characterized by the formation of said covalent linkage.
[0317] In some embodiments, the protein moiety (Ab) of the present disclosure includes but is not limited to the following: antibodies or their antigen-binding fragments; non-antibody scaffold proteins; cytokines, growth factors, hormones, or their variants and functional fragments; receptor proteins or their ligand-binding domains; ligand proteins or receptor-binding domain thereof; enzymes or their modulators; and peptides with specific binding activity.
[0318] In some embodiments, when the protein moiety (Ab) is an antibody or antigen-binding fragment thereof, it may include, but is not limited to, full-length antibodies (e.g., IgG, IgM, IgA, IgD, IgE antibodies) , antibody fragments (e.g., Fab, Fab', F (ab') 2, Fv, single-chain antibody (scFv) , diabody, triabody, minibody) , nanobodies (also known as VHH) , domain antibodies (e.g., dAb) , or modified antibodies (e.g., chimeric antibodies, humanized antibodies, fully human antibodies, bispecific antibodies, multispecific antibodies) .
[0319] The antibody or its antigen-binding fragment is capable of specifically binding to one or more target proteins, such as cell surface molecules (e.g., immune checkpoint molecules PD-L1, CTLA-4, LAG-3, and the like) , soluble proteins (e.g., cytokines, chemokines) , viral antigens (e.g., the receptor-binding domain (RBD) of the SARS-CoV-2 virus spike protein) , bacterial antigens, and the like.
[0320] The cysteine (Cys) residue can be naturally occurring in the antibody sequence, provided its position and reactivity meet the requirements. However, in preferred embodiments, the Cys residue is introduced at specific positions in the antibody or its fragment by genetic engineering means (such as site-directed mutagenesis) . The site for introducing the Cys residue is preferably located in the variable region (Fv) of the antibody, for example, at a surface-exposed position within or near a complementarity-determining region (CDR; e.g., CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2, CDR-H3) , or at a surface-exposed position in a framework region (FR) that is adjacent to the antigen-binding site. In some cases, a Cys residue may also be introduced at a specific surface-exposed position in the constant region (Fc) , provided that this position allows the crosslinker to effectively approach the target protein after the antibody binds to the target protein.
[0321] The amino acid residues selected for mutation are typically those whose side chains are non-essential for maintaining antibody structure and binding activity, such as serine (Ser) , alanine (Ala) , threonine (Thr) , valine (Val) , leucine (Leu) , isoleucine (Ile) , and the like. For example, in a protein moiety (Ab) targeting programmed death-ligand 1 (PD-L1) , it can be an anti-PD-L1 scFv, nanobody, or Fab fragment, wherein a Cys residue has been introduced at one or more selected positions near the PD-L1 binding interface (e.g., determined by crystal structure analysis or molecular modeling) . For example, in a protein moiety (Ab) targeting the receptor-binding domain (RBD) of Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) , it can be a neutralizing antibody against SARS-CoV-2 RBD (such as IgG, scFv, or VHH) , wherein a Cys residue has been introduced at one or more selected positions near the RBD interaction interface.
[0322] In some embodiments, the protein moiety (Ab) can be miniproteins. These miniproteins generally refer to a class of proteins with small molecular weight (for example, about 20 to about 100 amino acid residues, corresponding to about 3 kDa to about 15 kDa) and well-defined and stable three-dimensional structures. They can be designed based on naturally occurring small protein domains (such as knottins) , obtained by truncating and optimizing existing peptides or proteins, or completely designed de novo. Miniproteins typically stabilize their structure through multiple disulfide bonds or a compact hydrophobic core, giving them high thermal stability and resistance to protease degradation. They can be efficiently engineered to specifically bind a variety of target proteins, such as viral proteins, cell surface receptors, or enzymes, thereby serving as an attractive scaffold. In some specific embodiments, the protein moiety (Ab) can be a miniprotein targeting the receptor-binding domain (RBD) of Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) . The miniprotein is designed or screened to specifically bind SARS-CoV-2 RBD, and either naturally possesses a Cys residue near its interaction interface with RBD, or has a Cys residue engineered at one or more selected positions near its interaction interface with RBD for subsequent conjugation with R1 group of the crosslinker. The selection of the introduction site aims to ensure that the side chain thiol group of the Cys residue has good solvent accessibility and reactivity, and is spatially close to the target protein, without significantly disrupting the overall folding and stability of the miniprotein.
[0323] In some embodiments, when the protein moiety (Ab) is a non-antibody scaffold protein, it can be a protein derived from non-immunoglobulin sources that has been engineered to achieve high affinity and high specificity binding to a specific target. Such proteins typically possess a stable structural core and regions that can be engineered for diverse binding surfaces. Examples of non-antibody scaffold proteins include, but are not limited to: Ankyrin Repeat Proteins (DARPins) , proteins based on fibronectin type III domains (e.g., Adnectins (based on FN3) , Monobodies, Centyrins) , lipocalin-based proteins (Anticalins / Lipocalins) , Affibody molecules (based on the Z domain of Staphylococcal protein A) , Kunitz domain inhibitors, Avimers (multimers based on A domains) , Knottins (often a type of miniprotein themselves) , transport protein-based scaffolds, Cystatins, Stefins, and the like. The Cys residue is typically introduced by site-directed mutagenesis into the variable loop regions responsible for interaction with the target protein, the binding surface, or framework regions adjacent to the binding surface of these scaffold proteins. The selection of the introduction site aims to ensure that the side chain thiol group of the Cys residue has good solvent accessibility and reactivity, and is spatially close to the target protein.
[0324] In some embodiments, when the protein moiety (Ab) is a cytokine, growth factor, hormone, or its variant and functional fragment, it can be a naturally occurring or genetically engineered signaling molecule of this type. Examples include, but are not limited to: interleukins (ILs, such as IL-2, IL-12, IL-15, IL-18 and their functional variants) , interferons (IFNs, such as IFN-α, IFN-β, IFN-γ) , tumor necrosis factor family proteins (TNFs, such as TNF-α) , colony-stimulating factors (CSFs, such as G-CSF, GM-CSF) , chemokines (such as CCL2, CXCL8) , insulin, epidermal growth factor (EGF) , vascular endothelial growth factor (VEGF) , fibroblast growth factors (FGFs) , transforming growth factor-β (TGF-β) , and the like, as well as their biologically active fragments or variants modified to alter stability, receptor binding properties, or immunogenicity. These fragments may, in some cases, be small in size and possess certain characteristics of miniproteins (such as a compact structure) . The Cys residue can be located near its binding interface with the corresponding receptor (which can be considered the target protein) , or at a surface-exposed site that is not directly involved in receptor binding but whose modification maintains binding ability and facilitates crosslinker positioning.
[0325] In some embodiments, when the protein moiety (Ab) is a receptor protein or its ligand-binding domain, it can be a cell surface receptor or its soluble fragment (e.g., ectodomain) capable of binding a specific ligand (which can be considered the target protein) . Examples include, but are not limited to: soluble PD-1, CTLA-4, TNF receptors (such as TNFRSF1A, TNFRSF1B) , the soluble extracellular domain of ACE2 (angiotensin-converting enzyme 2, as a receptor for SARS-CoV-2) or its fragments and variants, and the like. These fragments may also, in some cases, be small in size and possess a stable structure. The Cys residue can be located near its binding interface with the corresponding ligand.
[0326] In some embodiments, when the protein moiety (Ab) is an enzyme or its regulator, the protein moiety (Ab) can be an enzyme, and its target protein is the substrate, product, inhibitor, or allosteric regulator of that enzyme. The Cys residue can be located near the active site or near an allosteric site. Alternatively, the protein moiety (Ab) can be an enzyme inhibitor (such as serine protease inhibitors, serpins, or small molecule protease inhibitor peptides / miniproteins) or an activator, and its target protein is the enzyme.
[0327] In some embodiments, when the protein moiety (Ab) is a peptide with specific binding activity, it can be a linear or cyclic peptide, typically comprising about 5 to about 100 amino acid residues, preferably about 10 to about 50 amino acid residues. These peptides can be obtained through various screening techniques (such as phage display, yeast display, mRNA display, peptide array screening) or rational design, and are capable of specifically binding to a particular target protein. When such peptides are larger in size (e.g., exceeding about 30-40 amino acids) and form stable, compact three-dimensional structures through means such as disulfide bonds, they may be structurally and functionally similar to the aforementioned miniproteins and can serve as a structured binding module. The Cys residue can be naturally occurring in the peptide sequence or synthetically introduced, and its position is crucial for covalent binding to the target protein mediated by the crosslinker.
[0328] The protein moiety (Ab) in the covalent protein of the present disclosure preferably possesses the following properties and structural features:
[0329] Target binding ability: The protein moiety (Ab) should be capable of recognizing and binding its predetermined target protein with a certain affinity (preferably in a non-covalent manner) . The initial non-covalent binding affinity (measured as the dissociation constant, Kd value) can be within a wide range, for example, from millimolar (mM) to picomolar (pM) levels. The subsequent formation of a covalent bond will significantly enhance the apparent affinity or achieve irreversible binding.
[0330] Specificity: Preferably, the protein moiety (Ab) exhibits a high degree of binding specificity for its target protein to minimize cross-reactivity with non-target molecules and potential off-target effects.
[0331] Cysteine (Cys) residue: The protein moiety (Ab) should comprise at least one Cys residue available for reaction with R2 group of the crosslinker (as defined in the CROSSLINKER COMPOUNDS section of the present disclosure) . In some embodiments, multiple such Cys residues may be present, or it may be engineered to ensure that only one or a few specific Cys residues possess the highest reactivity or accessibility, to achieve site-selective conjugation.
[0332] Preferably, the Cys residue is advantageously located at or spatially adjacent to the binding interface between the protein moiety (Ab) and the target protein. The selection of the position should facilitate effective contact of the crosslinker with, and covalent bond formation to, complementary reactive groups on the target protein (such as the side chains of Lys, Cys, His, Tyr, Ser, Thr, or the N-terminal amino group on the target protein) .
[0333] The Cys residue can be naturally occurring in the wild-type sequence of the protein moiety (Ab) , but more commonly, it is introduced at a preselected site via protein engineering techniques (primarily site-directed mutagenesis) . The original amino acid residue replaced by Cys is typically one that has minimal or no negative impact on maintaining the three-dimensional structure, stability, and non-covalent binding capability of the protein to the target protein, such as, but not limited to, serine (Ser) , alanine (Ala) , valine (Val) , leucine (Leu) , isoleucine (Ile) , threonine (Thr) , glycine (Gly) , and the like. The selection of the mutation site should be based on structural information, sequence conservation analysis, and / or computational modeling.
[0334] The side chain thiol group of the Cys residue should possess sufficient solvent accessibility and nucleophilic reactivity to react efficiently and selectively with R2 group of the crosslinker. Its microenvironment (e.g., local pH, charge, steric hindrance) should not significantly inhibit its reactivity. In some cases, it may be necessary to remove or protect other endogenous Cys residues at undesired reactive sites through engineering to improve the selectivity of conjugation and the homogeneity of the product.
[0335] Structural integrity and stability: The protein moiety (Ab) should possess sufficient physical (e.g., thermal stability, resistance to aggregation) and chemical (e.g., pH stability, resistance to oxidation) stability to withstand the requirements of production, purification, conjugation reaction with the crosslinker, storage, and the final application environment (e.g., physiological conditions) . The introduction of Cys residues and conjugation with the crosslinker should not lead to unacceptable unfolding, aggregation, or loss of function of the protein.
[0336] Producibility: The protein moiety (Ab) should be capable of being cost-effectively prepared and purified through conventional biotechnological methods (such as recombinant expression in prokaryotic expression systems like E. coli, or eukaryotic expression systems like yeast, insect cells, mammalian cells such as CHO, HEK293, and the like) or chemical synthesis methods (especially for smaller peptides, miniproteins, and protein fragments) .
[0337] Immunogenicity: For protein moieties (Ab) intended for therapeutic applications, low immunogenicity is preferred. This can be achieved by selecting human sequences, humanization, deimmunization engineering, or in conjunction with other techniques to reduce immunogenicity (such as PEGylation, though care must be taken that PEGylation does not interfere with the Cys reaction and binding function) .
[0338] Molecular size: The molecular size of the protein moiety (Ab) can vary widely, from a few kilodaltons (kDa) (e.g., small peptides, miniproteins, VHH, miniaturized forms of some scaffold proteins) to several hundred kilodaltons (kDa) (e.g., full-length IgG or multimeric protein complexes) .
[0339] Other modifications: The said protein moiety (Ab) can be glycosylated, PEGylated, or undergo other chemical or enzymatic modifications, provided that these modifications do not interfere with its binding to the target protein, do not interfere with the conjugation of the key Cys residue with the crosslinker, and do not interfere with the function of the final covalent protein.
[0340] The protein moiety (Ab) of the present disclosure is a carefully selected or designed biomolecule, the core feature of which lies in its ability to achieve specific covalent linkage with R1 group of a crosslinker through a specific cysteine residue located at or near its binding interface with the target protein. This ultimately promotes the formation of a stable covalent bond between the crosslinker and the target protein. Such a design provides a flexible and powerful platform for developing macromolecular biologic drugs, diagnostic reagents, or research tools with enhanced or novel characteristics, such as irreversible binding, long-lasting effects, or improved targeting.
[0341] In some preferred embodiments, the protein moiety Ab is an antibody or antigen-binding fragment thereof.
[0342] In some preferred embodiments, the protein moiety Ab is an antibody or antigen-binding fragment thereof, targeting an immune checkpoint.
[0343] In some preferred embodiments, the protein moiety Ab is a nanobody targeting an immune checkpoint.
[0344] In some preferred embodiments, the protein moiety Ab is a PD-L1 targeting nanobody.
[0345] In some preferred embodiments, the protein moiety Ab is the PD-L1 targeting nanobody of the present disclosure as provided hereinafter.
[0346] In some preferred embodiments, the protein moiety Ab is the PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NOs. 3, 4 to 16, and 17 to 38.
[0347] In some preferred embodiments, the protein moiety Ab is the PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 3, 4 to 16, and 17 to 38
[0348] In some preferred embodiments, the protein moiety Ab is a protein targeting a viral surface protein, or antigen.
[0349] In some preferred embodiments, the protein moiety Ab is a miniprotein targeting a viral surface protein, or antigen.
[0350] In some preferred embodiments, the protein moiety Ab is a miniprotein targeting a Coronavirus surface protein, or antigen.
[0351] In some preferred embodiments, the protein moiety Ab is a miniprotein targeting a SARS-CoV-2 surface protein, or antigen.
[0352] In some preferred embodiments, the protein moiety Ab is a miniprotein targeting the Receptor Binding Domain (RBD) of SARS-CoV-2.
[0353] In some preferred embodiments, the protein moiety Ab is the SAR-CoV-2 RBD targeting miniprotein of the present disclosure as provided hereinafter.
[0354] In some preferred embodiments, the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NOs. 55, and 56 to 61.
[0355] In some preferred embodiments, the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61.
[0356] The present disclosure provides a covalent protein selected from the group consisting of
[0357] where,
[0358] Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C4 haloalkyl, C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C4 alkyl, and C1-C4 haloalkyl; and
[0359] the protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 3, 4 to 16, and 17 to 38; or a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61.
[0360] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2. In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, and NO2.
[0361] In some preferred embodiments, the covalent protein selected from the group consisting of
[0362] Rx is H, F, CN, or NO2;
[0363] Rx is H, F, CN, or NO2;
[0364] Rx is H;
[0365] Rx is H, F, or NO2;
[0366] Rx is H, CN, or NO2;
[0367] Rx is H, F, or NO2; and
[0368] Rx is H,
[0369] where,
[0370] the protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 3, 4 to 16, and 17 to 38; or a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61.
[0371] In some preferred embodiments, the covalent protein selected from the group consisting of
[0372] where
[0373] Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2, preferably Rx is selected from the group consisting of H, F, CN, and NO2; more preferably Rx is selected from the group consisting of H, CN, and NO2; even more preferably Rx is H or CN;
[0374] the protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 8, 11 to 16; and SEQ ID NOs. 17 to 38, preferably any of SEQ ID NO. 8 and 11, and SEQ ID NOs. 17, and 37 to 38.
[0375] In some preferred embodiments, the covalent protein selected from the group consisting of
[0376] where
[0377] Rx is H or CN;
[0378] the protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 8, 11 to 16; and SEQ ID NOs. 17 to 38, preferably any of SEQ ID NO. 8 and 11, and SEQ ID NOs. 17, and 37 to 38.
[0379] In some preferred embodiments, the covalent protein selected from the group consisting of
[0380] where
[0381] Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2, preferably Rx is selected from the group consisting of H, F, CN, and NO2; more preferably Rx is selected from the group consisting of H, CN, and NO2; even more preferably Rx is H, CN or CN; and
[0382] the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61, preferably the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in any of SEQ ID NOs. 58 and 61;
[0383] In some preferred embodiments, the covalent protein is selected from the group consisting of
[0384] Rx is CN, and the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in SEQ ID NO. 58;
[0385] Rx is H, and the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in SEQ ID NO. 58; and
[0386] Rx is NO2, and the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in SEQ ID NO. 61.
[0387] The present disclosure provides a covalent protein having a structure as shown in Formula (I) for use in covalently linking to a target protein.
[0388] The present disclosure provides use of a covalent protein having a structure as shown in Formula (I) in covalently linking to a target protein.
[0389] The present disclosure also provides use of a covalent protein having a structure as shown in Formula (I) for enhancing the interaction between a binding protein and a target protein.
[0390] In some embodiments, the enhanced interaction is achieved by covalently linking the covalent protein to the target protein with the compound having a structure as shown in Formula (II) .
[0391] CROSSLINKER COMPOUNDS
[0392] As used herein, the term "crosslinker compound" refers to a chemical moiety comprising at least a first reactive group and a second reactive or interactive group, wherein the first reactive group is capable of forming a covalent bond with a protein molecule to yield a "covalent protein" as defined herein. Said covalent bond formation typically occurs via reaction between the first reactive group and an atom within an amino acid side chain of the protein molecule, for example, and without limitation, the sulfur atom of a cysteine side chain. Subsequent to the formation of the "covalent protein" , the crosslinker compound component thereof facilitates an interaction between the "covalent protein" and a second, distinct biomolecule, wherein such interaction may encompass, but is not limited to, antibody-antigen recognition or other specific biomolecular binding events. Furthermore, the crosslinker compound, via its second reactive or interactive group or a functionality derived therefrom after linkage to the protein, is configured to form an additional linkage, notably including a covalent linkage, with said second, distinct biomolecule. The formation of this additional linkage mediated by the crosslinker compound serves to enhance the stability, strength, and / or duration of the association between the "covalent protein" and the second, distinct biomolecule.
[0393] The present disclosure provides a compound having a structure as shown in Formula (II) :
[0394] wherein
[0395] R2 has a structure of R21-, or R21-C (O) -R22-R23-, where
[0396] R21 is selected from optionally substituted haloalkyl and optionally substituted alkenyl, where haloalkyl or alkenyl can be optionally substituted with one or more halo, -OH, or -CN,
[0397] R22 is selected from the group consisting of NR22a, O, S, and heterocyclylene, R22a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0398] R23 is absent or optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN;
[0399] RA is selected from the group consisting of optionally substituted alkanediyl, and optionally substituted arenediyl;
[0400] Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, alkyl, haloalkyl, and aryl,
[0401] n is an integer selected from 0 and 1;
[0402] Ry is selected from the group consisting of S and P;
[0403] Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, alkyl, haloalkyl, and aryl;
[0404] The bond between Ry and Rz is a single bond or a double bond.
[0405] In some embodiments, the compound having a structure as shown in Formula (II) is used as a crosslinker for covalent protein.
[0406] Therefore, the present disclosure provides a compound having a structure as shown in Formula (II) for use as a crosslinker for covalent protein.
[0407] The present disclosure also provides use of a compound having a structure as shown in Formula (II) as a crosslinker for covalent protein.
[0408] In some embodiments, R2 has a structure of R21-, or R21-C (O) -R22-R23-.
[0409] In some embodiments, R21 is selected from optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl and optionally substituted C2-C8, preferably C2-C6, more preferably C2-C4 alkenyl, where said haloalkyl or alkenyl can be optionally substituted with one or more halo, -OH, or -CN. In some preferred embodiments, R21 is selected from the group consisting of -CH2Cl, -CH2Br, -CH2I, -CH2=CH2.
[0410] In some embodiments, R22 is selected from the group consisting of NR22a, O, S, and 4-10 membered, preferably 4-8 membered, more preferably 5-6 membered heterocyclylene, for example, 5-or 6-membered heterocyclylene is selected from where R22a is selected from the group consisting of hydrogen, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 heteroalkyl, optionally substituted C3-C8, preferably C3-C6 cycloalkyl, optionally substituted C3-C8, preferably C3-C6 heterocyclyl, optionally substituted C6-C14, preferably C6-C10, more preferably C6 aryl, and optionally substituted 5-14 membered, preferably 5-10 membered, more preferably 5-6 membered heteroaryl.
[0411] In some embodiments, R23 is absent or optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN.
[0412] In some preferred embodiments, R2 is selected from C1-C4 haloalkyl, C2-C4 alkenyl, C1-C4 haloalkyl-C (O) -NH-, and C2-C4 alkenyl-C (O) -NH-. In some preferred embodiments, R2 is selected from the group consisting of -CH2Cl, -CH2Br, -CH2I, -CH2=CH2, -NH-C (O) -CH2-Cl, -NH-C (O) -CH2-Br, -NH-C (O) -CH2-I, and -NH-C (O) -CH2=CH2.
[0413] In some embodiments, RA is selected from the group consisting of optionally substituted C1-C8, preferably C1-C6 alkanediyl, and optionally substituted C6-C14, preferably C6-C10, more preferably C6 arenediyl.
[0414] In some preferred embodiments, RA is C1-C8, preferably C1-C6 alkanediyl optionally substituted with one or more halo, -OH, or -CN.
[0415] In some preferred embodiments, RA is C6-C14, preferably C6-C10, more preferably C6 arenediyl optionally substituted with one or more, for example one, two, three, or four Rx, where Rx is selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions. In some embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, and N (Rx1) (Rx2) , where each of Rx1 and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl. In some embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1 and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl. In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0416] In some embodiments, Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, n is 0, then Rw is absent. In some embodiments, n is 1.
[0417] In some embodiments, Ry is S. In some embodiments, Ry is P.
[0418] In some embodiments, Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. The bond between Ry and Rz is a single bond or a double bond. One of ordinary skill in the art will recognize that the bond order between Ry and Rz, specifically whether said bond is a single bond or a double bond, is determined by the chemical identity of the substituents Ry and Rz, as previously defined. The formation of such a bond shall conform to the established principles of valence bond theory, thereby ensuring appropriate chemical valency for the atoms involved.
[0419] In some embodiments, Rz is selected from the group consisting of =O, and =N (Rz2) , the bond between Ry and Rz is a double bond, where Rz2 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, Rz2 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0420] In some embodiments, Rz is selected from the group consisting of -O (Rz1) , and -N (Rz3) (Rz4) , the bond between Ry and Rz is a single bond, where each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0421] In some preferred embodiments, Ry is S, and Rz is selected from the group consisting of =O, and =N (Rz2) , the bond between Ry and Rz is a double bond, where Rz2 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, Rz2 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0422] In some embodiments, Ry is P, and Rz is selected from the group consisting of -O (Rz1) , and -N (Rz3) (Rz4) , the bond between Ry and Rz is a single bond, where each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl. In some embodiments, each of Rz1, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0423] In some embodiments, RA is C1-C8, preferably C1-C6 alkanediyl optionally substituted with one or more halo, -OH, or -CN. Therefore, the present disclosure provides a compound having a structure as shown in Formula (II-A) :
[0424] wherein,
[0425] p is an integer selected from 1 to 8, preferably 1 to 6, - (CH2) p-optionally substituted with one or more halo, -OH, or -CN;
[0426] each of R2, Rw, Ry, Rz and n is independently defined as disclosed herein, for example, as defined in the Formula (II) .
[0427] It should be understood that the term "- (CH2) p-" as used herein refers to an alkylene group comprising p carbon atoms, which may be configured as a straight chain or a branched chain structure, provided that the total number of carbon atoms in the chain, whether straight or branched, equals p.
[0428] In some preferred embodiments, R2 has a structure of R21-C (O) -R22-R23-.
[0429] In some preferred embodiments, R21 is C1-C4 haloalkyl, for example C1-C4 bromoalkyl, such as bromomethyl, bromoethyl, bromopropyl, and bromobutyl.
[0430] In some preferred embodiments, R22 is NR22a, where R22a is selected from the group consisting of hydrogen, C1-C4 alkyl, for example, R22 is NH.
[0431] In some preferred embodiments, R23 is absent.
[0432] In some preferred embodiments, R2 has a structure of Br-CH2-C (O) -NH-.
[0433] In some preferred embodiments, p is an integer selected from 1 to 8, preferably 1 to 6, specifically, p is 1, 2, 3, 4, 5, or 6.
[0434] In some preferred embodiments, Rw is absent, or is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, Rw1 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0435] In some preferred embodiments, Ry is S or P.
[0436] In some preferred embodiments, Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0437] In some preferred embodiments, the compound has a structure as shown in Formula (II-A1) - (II-A3) :
[0438] wherein,
[0439] p is an integer selected from 1 to 8, preferably 1 to 6, specifically, p is 1, 2, 3, 4, 5, or 6;
[0440] Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl; preferably, Rw1 is selected from the group consisting of H, methyl, ethyl, and phenyl;
[0441] each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl; preferably, each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0442] In some embodiments, RA is optionally substituted C6-C14, preferably C6-C10, more preferably C6 arenediyl. In some embodiments, RA is optionally substituted phenylene. Therefore, the present disclosure provides a compound having a structure as shown in Formula (II-B) :
[0443] wherein
[0444] each of R2, Rw, Rx, Ry, Rz, m and n is independently defined as disclosed herein, for example, as defined in the Formula (II) .
[0445] In some preferred embodiments, R2 has a structure of R21-, or R21-C (O) -R22-R23-.
[0446] In some preferred embodiments, R21 is C1-C4 haloalkyl, for example C1-C4 bromoalkyl, such as bromomethyl, bromoethyl, bromopropyl, and bromobutyl, or C2-C4 alkenyl, for example, ethenyl, propenyl, and butenyl.
[0447] In some preferred embodiments, R22 is NR22a, where R22a is selected from the group consisting of hydrogen, C1-C4 alkyl, for example, R22 is NH.
[0448] In some preferred embodiments, R23 is absent.
[0449] In some preferred embodiments, R2 is selected from the group consisting of Br-CH2-, Br-CH2-C (O) -NH-, CH2=CH-C (O) -NH-,
[0450] In some preferred embodiments, m is an integer selected from 1 to 4, specifically, m is 1, 2, 3, or 4, preferably m is 1. Each Rx is independently selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions. In some preferred embodiments, each Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl. In some preferred embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl. In some preferred embodiments, each Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0451] In some preferred embodiments, Rw is absent, or is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, Rw1 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0452] In some preferred embodiments, Ry is S or P.
[0453] In some preferred embodiments, Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, and C6-C14, preferably C6-C10, more preferably C6 aryl, for example, each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl.
[0454] The bond between Ry and Rz is a single bond or a double bond.
[0455] In some preferred embodiments, the compound has a structure as shown in Formula (II-B-a) - (II-B-c) :
[0456] wherein
[0457] each of R2, Rw, Rx, Ry, Rz, m and n is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0458] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0459] In some preferred embodiments, the compound has a structure as shown in Formula (II-B1) - (II-B4) :
[0460] wherein,
[0461] each of R2, Rx, Rw1, Rz1, Rz2, Rz3, Rz4, m and n is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0462] In some preferred embodiments, the compound has a structure as shown in Formula (II-B1-a) - (II-B1-c) :
[0463] wherein,
[0464] each of R2, Rx, m and n is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0465] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0466] In some preferred embodiments, the compound having a structure as shown in Formula (II-B1) is selected from the group consisting of
[0467] wherein,
[0468] Rx is defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0469] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0470] In some preferred embodiments, the compound has a structure as shown in Formula (II-B2-a) - (II-B2-c) :
[0471] wherein,
[0472] each of R2, Rx, Rz2, m and n is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0473] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0474] In some preferred embodiments, the compound having a structure as shown in Formula (II-B2) is selected from the group consisting of
[0475] wherein,
[0476] each of Rx, Rz2, and n is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0477] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0478] In some preferred embodiments, n is 0 or 1. In some preferred embodiments, n is 0. In some preferred embodiments, n is 1.
[0479] In some preferred embodiments, Rz2 is selected from the group consisting of H, methyl, ethyl, and phenyl.
[0480] In some preferred embodiments, the compound has a structure as shown in Formula (II-B3-a) - (II-B3-c) :
[0481] wherein,
[0482] each of R2, Rx, Rz3, Rz4, and m is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0483] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0484] In some preferred embodiments, the compound having a structure as shown in Formula (II-B3) is selected from the group consisting of
[0485] wherein,
[0486] each of Rx, Rz3, and Rz4 is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0487] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0488] In some preferred embodiments, each of Rz3 and Rz4 is independently selected from the group consisting of H, methyl, ethyl, and phenyl. In some preferred embodiments, Rz3 and Rz4 are methyl.
[0489] In some preferred embodiments, the compound has a structure as shown in Formula (II-B4-a) - (II-B4-c) :
[0490] wherein,
[0491] each of R2, Rx, Rw1, Rz1, and m is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0492] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0493] In some preferred embodiments, the compound having a structure as shown in Formula (II-B4) is selected from the group consisting of
[0494] wherein,
[0495] each of Rx, Rw1, and Rz1 is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0496] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0497] In some preferred embodiments, each of Rw1 and Rz1 is independently selected from the group consisting of H, methyl, ethyl, and phenyl. In some preferred embodiments, Rw1 is methyl.
[0498] In another aspect, the present disclosure provides a compound having a structure as shown in Formula (II-B1) as a crosslinker for covalent protein:
[0499] wherein
[0500] R2 has a structure of R21-, or R21-C (O) -R22-R23-, where
[0501] R21 is optionally substituted haloalkyl, or optionally substituted alkenyl, which haloalkyl, alkenyl is optionally substituted with one or more halo, -OH, or -CN,
[0502] R22 is selected from the group consisting of NR22a, O, S, and heterocyclylene, R22a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;
[0503] R23 is absent or optionally substituted alkylene, which is optionally substituted with one or more halo, -OH, or -CN;
[0504] Rx is selected from the group consisting of -H, -halo, -CN, and -NO2 or other chemical groups with similar functions, m is an integer selected from 1 to 4;
[0505] n is an integer selected from 0 and 1.
[0506] In some embodiments, the compound having a structure as shown in Formula (II-B1) is used as a crosslinker for covalent protein.
[0507] Therefore, the present disclosure provides a compound having a structure as shown in Formula (II-B1) for use as a crosslinker for covalent protein.
[0508] The present disclosure also provides use of a compound having a structure as shown in Formula (II-B1) as a crosslinker for covalent protein.
[0509] In some embodiments, R2 has a structure of R21-, or R21-C (O) -R22-R23-.
[0510] In some embodiments, R21 is selected from optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl and optionally substituted C2-C8, preferably C2-C6, more preferably C2-C4 alkenyl, where said haloalkyl or alkenyl can be optionally substituted with one or more halo, -OH, or -CN. In some preferred embodiments, R21 is selected from the group consisting of -CH2Cl, -CH2Br, -CH2I, -CH2=CH2.
[0511] In some embodiments, R22 is selected from the group consisting of NR22a, O, S, and 4-10 membered, preferably 4-8 membered, more preferably 5-6 membered heterocyclylene, for example, 5-or 6-membered heterocyclylene is selected from where R22a is selected from the group consisting of hydrogen, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 heteroalkyl, optionally substituted C3-C8, preferably C3-C6 cycloalkyl, optionally substituted C3-C8, preferably C3-C6 heterocyclyl, optionally substituted C6-C14, preferably C6-C10, more preferably C6 aryl, and optionally substituted 5-14 membered, preferably 5-10 membered, more preferably 5-6 membered heteroaryl.
[0512] In some embodiments, R23 is absent or optionally substituted C1-C8, preferably C1-C6, more preferably C1-C4 alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN.
[0513] In some preferred embodiments, R2 is selected from the group consisting of -CH2Cl, -CH2Br, -CH2I, -CH2=CH2, -NH-C (O) -CH2-Cl, -NH-C (O) -CH2-Br, -NH-C (O) -CH2-I, or -NH-C (O) -CH2=CH2.
[0514] In some embodiments, m is an integer selected from 1 to 4, specifically, m is 1, 2, 3, or 4, preferably m is 1. Each Rx is independently selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions. In some embodiments, each Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl. In some embodiments, Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl. In some preferred embodiments, each Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0515] In some embodiments, the compound has a structure as shown in Formula (II-B1-a) - (II-B1-c) :
[0516] wherein,
[0517] each of R2, Rx, m and n is independently defined as disclosed herein, for example, as defined in the Formula (II-B1) .
[0518] In some preferred embodiments, m is 1, Rx is located at the ortho, meta or para position relative to R2.
[0519] In some embodiments, the compound having a structure as shown in Formula (II-B1) is selected from the group consisting of
[0520] wherein,
[0521] each of Rx and n is independently defined as disclosed herein, for example, as defined in the Formula (II-B) .
[0522] In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.
[0523] In some preferred embodiments, n is 0 or 1. In some preferred embodiments, n is 0. In some preferred embodiments, n is 1.
[0524] In some embodiments, the compound having a structure as shown in Formula (II-B1) is selected from the group consisting of
[0525] wherein,
[0526] Rx is defined as disclosed herein. In some embodiments, Rx is selected from the group consisting of H, F, CN, and NO2, or similar functional groups. In some preferred embodiments, Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2. In some preferred embodiments, , Rx is selected from the group consisting of H, F, CN, and NO2.
[0527] In some embodiments, the compound having a structure as shown in Formula (II-B1) is selected from the group consisting of
[0528] USE OF CROSSLINKER COMPOUNDS
[0529] The present disclosure provides a compound having a structure as shown in Formula (II) for use as a crosslinker for covalent protein.
[0530] The present disclosure also provides use of a compound having a structure as shown in Formula (II) as a crosslinker for covalent protein.
[0531] The present disclosure also provides use of a compound having a structure as shown in Formula (II) for covalently linking a binding protein to a target protein.
[0532] The present disclosure also provides use of a compound having a structure as shown in Formula (II) for enhancing the interaction between a binding protein and a target protein.
[0533] In some embodiments, the enhanced interaction is achieved by covalently linking the binding protein and the target protein with the compound having a structure as shown in Formula (II) .
[0534] In some embodiments, the compound of the present disclosure is used as crosslinker and are covalently linked with a protein (also referred to as a binding protein) via at least one cysteine (Cys) residue of the protein to form a covalent protein, for example, a covalent protein of Formula (I) of the present disclosure. In some embodiments, a crosslinker compound of the present disclosure covalently links to the binding protein to form a covalent protein, for example, a covalent protein of formula (I) of the present disclosure; this covalent linkage is achieved, for example, by the reaction of a haloalkyl group of the crosslinker compound with the side chain (i.e., thiol group, -SH) of a cysteine (Cys) residue of the binding protein, or by a Michael addition reaction or a similar nucleophilic addition reaction between an alkenyl group of the crosslinker compound and the thiol group of a cysteine (Cys) residue of the binding protein.
[0535] The covalent protein, for example, a covalent protein of Formula (I) of the present disclosure, comprises a crosslinker moiety, wherein said crosslinker moiety comprises at its terminus a sulfur-containing sulfonyl fluoride group or a phosphorus-containing phosphonyl fluoride group. In some embodiments, the covalent protein covalently binds to a target protein via the terminal sulfonyl fluoride group or phosphonyl fluoride group of said crosslinker moiety to form a complex. The sulfonyl fluoride group or the phosphonyl fluoride group is a highly reactive electrophilic group capable of reacting with nucleophilic amino acid residues on the surface or within a binding pocket of the target protein. In some specific embodiments, said nucleophilic amino acid residues include, but are not limited to, the ∈-amino group of a lysine (Lys) residue, the phenolic hydroxyl group of a tyrosine (Tyr) residue, the hydroxyl group of a serine (Ser) residue, the hydroxyl group of a threonine (Thr) residue, or a nitrogen atom of the imidazole ring of a histidine (His) residue.
[0536] In some embodiments, the compound of the present disclosure is used as a crosslinker to covalently link a protein containing at least one cysteine (Cys) residue (also referred to as a binding protein) with a target protein molecule to form a stable complex. The crosslinker comprises at least two reactive moieties: a first moiety that reacts with a cysteine residue of the binding protein via the R2 group (as defined herein) , and a second moiety, which is a terminal sulfonyl fluoride group or phosphonyl fluoride group, for reacting with the target protein. In this manner, a stable and specifically linked binding protein-linker-target protein complex can be constructed. Such utility is of significant value in drug development (e.g., antibody-drug conjugates, protein degraders) , diagnostic reagents, and fundamental biological research.
[0537] In some preferred embodiments, the covalent binding between the covalent protein and the target protein is highly specific. This specificity can arise from the inherent affinity of the binding protein for the target protein, or from the design of the linker structure (including the RA, Rw, and n moieties as defined herein) , which allows the sulfonyl fluoride group or phosphonyl fluoride group to be precisely directed to the vicinity of a specific nucleophilic residue on the target protein, thereby effecting site-specific covalent conjugation. In certain instances, non-covalent pre-association (e.g., through hydrophobic interactions, hydrogen bonds, or electrostatic interactions) can promote and localize the subsequent covalent reaction.
[0538] Through the aforementioned two-step covalent conjugation strategy, the binding protein and the target protein can be effectively and tightly bound together, forming a stable ternary complex for further functional studies, drug delivery, or diagnostic applications.
[0539] PREPARATION METHOD FOR COVALENT PROTEIN
[0540] The present disclosure also provides a method for preparing a covalent protein of the present disclosure, the method comprising:
[0541] (a) providing a protein comprising at least one cysteine (Cys) residue;
[0542] (b) providing a compound having a structure as shown in Formula (II) ; and
[0543] (c) reacting the compound with the side chain of the at least one cysteine (Cys) residue of the protein to form a covalent bond.
[0544] The detailed description and description of the protein (also referred to as binding protein, i.e., a protein containing at least one cysteine (Cys) residue) and the crosslinker compound in the present disclosure are all applicable to the preparation method.
[0545] HIGH THROUGHPUT COVALENT PROTEIN SELECTION METHOD
[0546] The present disclosure provides a high throughput covalent protein selection method based on yeast display, comprising
[0547] 1) generating a protein library displayed on yeast cells,
[0548] 2) binding a crosslinker to the displayed proteins;
[0549] 3) incubating with target protein; and
[0550] 4) performing FACS sorting.
[0551] In some embodiments, the crosslinker is a compound having a structure as shown in Formula (II) of the present disclosure.
[0552] In some embodiments, in step 2) , before binding the crosslinker to the displayer proteins, a reducing agent is added to treat the yeast cells. In some such embodiments, the reducing agent is selected from the group consisting of dithiothreitol (DTT) , tris (2-carboxyethyl) phosphine (TCEP) , and other suitable reducing agents known in the art. In some embodiments, the reducing agent is added at a concentration range from 0.1 mM to 2 mM at a temperature of 4℃ to 25℃
[0553] In some embodiments, in step 2) , the crosslinker is added, for example, at a concentration of about 0.2 mM to the yeast cells to incubate for, for example, about 2 hours at a temperature of, for example, about 37℃ for attachment on the protein displayed.
[0554] In some embodiments, in step 3) , the yeast cells are incubated with the target protein, for the purpose of covalently binding the target protein to the displayed proteins through the crosslinker. In some embodiments, after binding of the target protein, an acid wash, for example, with a pH of about 2-3 is performed to remove noncovalent binder in order to selectively screen the covalent binder.
[0555] In some embodiments, after the acid wash, FACS sorting is performed to screen candidate covalent protein with high affinity to the target protein.
[0556] In some embodiments, steps 2) to 4) are repeatedly conducted to further screen candidate covalent proteins with high affinity and faster covalent crosslinking rate to the target protein.
[0557] In some embodiments, in each repeat cycle of steps 2) to 4) , the yeast cells are incubated with successively lower concentrations of the target protein, for example, 10 nM for 2nd round, 5 nM for 3rd round, and 3 nM for the 4th round, for the purpose of enriching for binders with higher affinities. In some embodiments, in each repeat cycle of steps 2) to 4) , successively shorter incubation times, for example, 20 min for the 2nd round, 10 min for the 3rd round, and 3 min for the 4th round, for the purpose of enriching for binders with faster reaction rates.
[0558] In some embodiments, the protein library consists of PD-L1 targeting proteins. In some embodiments, the target protein is PD-L1.
[0559] In some embodiments, the protein library consists of SARS-CoV-2 RBD targeting proteins. In some embodiments, the target protein is SARS-CoV-2 RBD.
[0560] PD-L1 TARGETING NANOBODY
[0561] As used herein, the term "protein" refers to a polymer comprising amino acid residues linked together predominantly by peptide bonds. The term is used in its broadest sense and is intended to encompass any chain of two or more amino acid residues, regardless of its length, three-dimensional structure (or lack thereof) , or specific biological function. Accordingly, the term "protein" explicitly includes, without limitation, molecules commonly referred to as peptides, polypeptides, and proteins. The amino acid residues may be naturally occurring (e.g., the 20 common L-amino acids) or non-naturally occurring (e.g., synthetic amino acids, D-amino acids, amino acid analogs) . Furthermore, the term "protein" encompasses molecules that have undergone post-translational modifications (such as, but not limited to, glycosylation, phosphorylation, acetylation, methylation, ubiquitination, lipidation, or disulfide bond formation) , as well as molecules that may be fragments, variants (including mutants) , homologs, or chemically synthesized derivatives of naturally occurring sequences, provided they retain the fundamental polymeric structure of amino acid residues linked by peptide bonds. The protein may be of natural, recombinant, or synthetic origin.
[0562] As used herein, the term "nanobody" refers to an antigen-binding polypeptide comprising, or consisting essentially of, a single variable domain derived from, or structurally analogous to, the variable domain of the heavy chain of a heavy-chain-only antibody (HCAb) . Such HCAbs are known to occur naturally in members of the Camelidae family. Accordingly, a nanobody typically corresponds to the VHH domain (Variable domain of the Heavy chain of a Heavy-chain antibody) . The term "nanobody" is used in its broadest sense to encompass such single variable domains regardless of their specific antigen target, origin (e.g., derived from natural Camelidae HCAbs obtained via immunization, or selected from immune, naive, semi-synthetic, or synthetic libraries) , or method of production (e.g., recombinant expression in prokaryotic or eukaryotic hosts, or chemical synthesis) . Furthermore, the term explicitly includes, without limitation, naturally occurring VHH sequences, genetically engineered variants thereof (such as humanized, camelized, stabilized, or affinity-matured nanobodies) , chimeric nanobodies, functional fragments of a VHH domain that retain antigen-binding capability, and modified nanobodies (e.g., conjugated to other moieties such as labels, therapeutic agents, or half-life extension domains) , provided they retain the characteristic single-domain antigen-binding structure derived from a VHH scaffold.
[0563] As used herein, the term "PD-L1 targeting nanobody" refers to a nanobody, as defined herein, that specifically binds to Programmed Death-Ligand 1 (PD-L1) . A nanobody, as previously defined, comprises, or consists essentially of, a single variable domain (VHH) derived from the heavy chain of a heavy-chain-only antibody or is structurally analogous thereto, encompassing natural, recombinant, synthetic, and modified forms thereof retaining antigen-binding capability. The characteristic of "PD-L1 targeting" signifies that said nanobody recognizes and binds to an epitope present on a PD-L1 protein. Unless specified otherwise, this includes binding to PD-L1 from any relevant species (e.g., human PD-L1, murine PD-L1) , as well as naturally occurring isoforms, allelic variants, or processed forms (such as the extracellular domain) thereof. Specific binding generally implies a measurable binding affinity for PD-L1 that is significantly higher than its affinity for unrelated proteins under comparable conditions. The term "PD-L1 targeting nanobody" therefore encompasses any nanobody exhibiting such specific binding to PD-L1, irrespective of the particular epitope bound on PD-L1 or the functional consequence, if any, of such binding.
[0564] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0565] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having deletion, substitution, insertion or addition of one or more amino acid residues as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0566] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations at the position selected from the group consisting of positions 100, 102, 104, 107, 108, 113, and 116 as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0567] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, each of the amino acid residues at positions 100, 102, 104, 107, 108, 113, and 116 is independently mutate to C as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0568] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations selected from the group consisting of S100C, E102C, P104C, T107C, L108C, G113C, and Q116C.
[0569] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 4 to 10.
[0570] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations at the position selected from the group consisting of position 50, 59, 102, 103, 107, 108, 110, 111, and 112 as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0571] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations at the position selected from the group consisting of position 50, 59, 103, 107, 108, 110, 111, and 112 as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0572] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations at the position selected from the group consisting of position 102, 108, and 111 as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0573] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 50 is mutated to G, T, or A as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0574] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 59 is mutated to F or S as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0575] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 102 is mutated to F or W as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0576] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 103 is mutated to N or D as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0577] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 107 is mutated to N, T, S, or H as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0578] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 108 is mutated to C as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0579] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 110 is mutated to R, H, P, D, or N as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0580] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 111 is mutated to Y, H, S, A, F, or N as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0581] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 112 is mutated to G, S, A, D, or N as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0582] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 113 is mutated to G, A, or T as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0583] The present disclosure provides a PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 11 to 16.
[0584] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0585] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having deletion, substitution, insertion or addition of one or more amino acid residues as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0586] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations at the position selected from the group consisting of position 50, 59, 102, 103, 107, 110, 111, and 112 as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0587] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 50 is mutated to G, T, or A as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0588] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 59 is mutated to F or S as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0589] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 102 is mutated to F or W as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0590] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 103 is mutated to N or D as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0591] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 107 is mutated to N, T, S, or H as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0592] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 110 is mutated to R, H, P, D, or N as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0593] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 111 is mutated to Y, H, S, A, F, or N as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0594] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 112 is mutated to G, S, A, D, or N as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0595] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, the amino acid residue at position 113 is mutated to G, A, or T as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0596] The present disclosure provides a PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 17, and 18 to 36.
[0597] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 3,
[0598] wherein the amino acid residue (s) at positions 100, 102, 104, 107, 108, 113, 116, 50, 59, 103, 110, 111, 112, 5, 72, 73, 75, 77, 87, 88, 93, and 123 are as follows,
[0599] position 100: the amino acid residue at position 100 is selected from the group consisting of Ser (S) , and Cys (C) ;
[0600] position 102: the amino acid residue at position 102 is selected from the group consisting of Glu (E) , Cys (C) , Phe (F) , and Trp (W) ;
[0601] position 104: the amino acid residue at position 104 is selected from the group consisting of Pro (P) , and Cys (C) ;
[0602] position 107: the amino acid residue at position 107 is selected from the group consisting of Thr (T) , Cys (C) , Asn (N) , Ser (S) , and His (H) ;
[0603] position 108: the amino acid residue at position 108 is selected from the group consisting of Leu (L) , and Cys (C) ;
[0604] position 113: the amino acid residue at position 113 is selected from the group consisting of Gly (G) , Cys (C) , Ala (A) , and Thr (T) ;
[0605] position 116: the amino acid residue at position 116 is selected from the group consisting of Gln (Q) , and Cys (C) ;
[0606] position 50: the amino acid residue at position 50 is selected from the group consisting of Lys (K) , Gly (G) , Thr (T) , and Ala (A) ;
[0607] position 59: the amino acid residue at position 59 is selected from the group consisting of Tyr (Y) , Phe (F) , and Ser (S) ;
[0608] position 103: the amino acid residue at position 103 is selected from the group consisting of Asp (D) , and Asn (N) ;
[0609] position 110: the amino acid residue at position 110 is selected from the group consisting of Thr (T) , Arg (R) , His (H) , Pro (P) , Asp (D) , and Asn (N) ;
[0610] position 111: the amino acid residue at position 111 is selected from the group consisting of Ser (S) , Tyr (Y) , His (H) , Ala (A) , Phe (F) , and Asn (N) ;
[0611] position 112: the amino acid residue at position 112 is selected from the group consisting of Ser (S) , Gly (G) , Ala (A) , Asp (D) , and Asn (N) ;
[0612] position 5: the amino acid residue at position 5 is selected from the group consisting of Gln (Q) , and Val (V) ;
[0613] position 72: the amino acid residue at position 72 is selected from the group consisting of Gln (Q) , and Arg (R) ;
[0614] position 73: the amino acid residue at position 73 is selected from the group consisting of Asn (N) , and Asp (D) ;
[0615] position 75: the amino acid residue at position 75 is selected from the group consisting of Ala (A) , and Ser (S) ;
[0616] position 77: the amino acid residue at position 77 is selected from the group consisting of Ser (S) , and Asn (N) ;
[0617] position 87: the amino acid residue at position 87 is selected from the group consisting of Lys (K) , and Arg (R) ;
[0618] position 88: the amino acid residue at position 88 is selected from the group consisting of Pro (P) , and Ala (A) ;
[0619] position 93: the amino acid residue at position 93 is selected from the group consisting of Met (M) , and Val (V) ; and
[0620] position 123: the amino acid residue at position 123 is selected from the group consisting of Gln (Q) , and Leu (L) .
[0621] In some embodiments, the amino acid residue of the PD-L1 targeting nanobody is in unmodified form, protected form or modified form.
[0622] In some embodiments, the remaining amino acid residues are the same as the wild type sequence shown in SEQ ID NO. 3.
[0623] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%and / or less than 100%identity identity as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0624] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least one mutation as defined herein as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0625] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0626] each of the variable amino acids is as defined herein.
[0627] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 4 to 10, 11 to 16, 17 to 38.
[0628] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 4 to 10, 11 to 16, 17 to 38.
[0629] The PD-L1 targeting nanobodies disclosed herein encompass not only the naturally occurring, unmodified forms of the amino acid sequences set forth in the present disclosure, but also chemically or biologically modified derivatives thereof. Such modifications include, but are not limited to, the introduction of substituents and / or the addition of protecting groups to the side chains of amino acid residues. Exemplary modifications include, without limitation, phosphorylation, glycosylation, acetylation, methylation, amidation, sulfonation, and pegylation, as well as other post-translational modifications and artificial synthetic modifications.
[0630] It will be understood by those skilled in the art that, provided the core functional domain of the protein remains substantially unaltered, variations in the type, number, and point of attachment of substituents and protecting groups do not depart from the scope of the present disclosure. Accordingly, equivalent variants of the aforementioned side chain modifications, wherein the amino acid sequence corresponds to the amino acid sequences disclosed herein, are considered reasonable extensions of the scope of the claims and fall within the definition of the PD-L1 targeting nanobodies of the present disclosure.
[0631] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 3,
[0632] wherein the amino acid residue (s) at positions 100, 102, 104, 107, 108, 111, 113, and 116 are as follows,
[0633] position 50: the amino acid residue at position 50 is selected from the group consisting of Lys (K) , and Ala (A) ;
[0634] position 59: the amino acid residue at position 59 is selected from the group consisting of Tyr (Y) , and Phe (F) ;
[0635] position 100: the amino acid residue at position 100 is selected from the group consisting of Ser (S) , and Cys (C) ;
[0636] position 102: the amino acid residue at position 102 is selected from the group consisting of Glu (E) , Cys (C) , Phe (F) , and Trp (W) ;
[0637] position 104: the amino acid residue at position 104 is selected from the group consisting of Pro (P) , and Cys (C) ;
[0638] position 107: the amino acid residue at position 107 is selected from the group consisting of Thr (T) , and Cys (C) ;
[0639] position 108: the amino acid residue at position 108 is selected from the group consisting of Leu (L) , and Cys (C) ;
[0640] position 111: the amino acid residue at position 111 is selected from the group consisting of Ser (S) , and His (H) ;
[0641] position 113: the amino acid residue at position 113 is selected from the group consisting of Gly (G) , and Cys (C) ; and
[0642] position 116: the amino acid residue at position 116 is selected from the group consisting of Gln (Q) , and Cys (C) .
[0643] In some embodiments, the amino acid residue of the PD-L1 targeting nanobody is in unmodified form, protected form or modified form.
[0644] In some embodiments, the remaining amino acid residues are the same as the wild type sequence shown in SEQ ID NO. 3.
[0645] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least one mutation as defined herein as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0646] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0647] each of the variable amino acids is as defined herein.
[0648] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 4 to 10, and 11 to 16.
[0649] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 4 to 10, and 11 to 16.
[0650] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations, for example, 1, 2, 3, 4, 5, or 6 mutations at the position selected from the group consisting of positions 100, 102, 104, 107, 108, 111, 113, and 116 as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0651] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, each of the amino acid residues at positions 100, 102, 104, 107, 108, 113, and 116 is independently mutated to C as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0652] In some embodiments, the amino acid residue of the PD-L1 targeting nanobody is in unmodified form, protected form or modified form.
[0653] In some preferred embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, only one of the amino acid residues at positions 100, 102, 104, 107, 108, 113, and 116 is mutated to C as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0654] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations selected from the group consisting of S100C, E102C, P104C, T107C, L108C, G113C, and Q116C.
[0655] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one mutation selected from the group consisting of S100C, E102C, P104C, T107C, L108C, G113C, and Q116C.
[0656] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 4 to 10.
[0657] In some embodiments, in the amino acid sequence of the PD-L1 targeting nanobody, as compared to the amino acid sequence set forth in SEQ ID NO. 3, the amino acid residue (s) at positions 50, 59, 102, 108, and 111 are as follows,
[0658] position 50: the amino acid residue at position 50 is selected from the group consisting of Lys (K) , and Ala (A) ;
[0659] position 59: the amino acid residue at position 59 is selected from the group consisting of Tyr (Y) , and Phe (F) ;
[0660] position 102: the amino acid residue at position 102 is selected from the group consisting of Glu (E) , Phe (F) , and Trp (W) ;
[0661] position 108: the amino acid residue at position 108 is Cys (C) ; and
[0662] position 111: the amino acid residue at position 111 is selected from the group consisting of Ser (S) , and His (H) .
[0663] In some embodiments, the amino acid residue of the PD-L1 targeting nanobody is in unmodified form, protected form or modified form.
[0664] In some embodiments, the remaining amino acid residues are the same as the wild type sequence shown in SEQ ID NO. 3.
[0665] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0666] each of the variable amino acids is as defined herein.
[0667] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations, for example, 1, 2, or 3 mutations at the position selected from the group consisting of positions 50, 59, 102, 108, and 111 as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0668] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having one or more mutations, for example, 1, 2, or 3 mutations selected from the group consisting of Y50A, F59F, E102F, E102W, L108C, and S111H as compared to the amino acid sequence set forth in SEQ ID NO. 3.
[0669] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 8, and 11 to 16.
[0670] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 8, and 11 to 16.
[0671] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 17,
[0672] wherein the amino acid residue (s) at positions 50, 59, 103, 107, 110, 111, 112, 113, 5, 72, 73, 75, 77, 87, 88, 93, and 123 are as follows,
[0673] position 50: the amino acid residue at position 50 is selected from the group consisting of Gly (G) , Thr (T) , and Ala (A) ;
[0674] position 59: the amino acid residue at position 59 is selected from the group consisting of Phe (F) , and Ser (S) ;
[0675] position 103: the amino acid residue at position 103 is selected from the group consisting of Asp (D) , and Asn (N) ;
[0676] position 107: the amino acid residue at position 107 is selected from the group consisting of Thr (T) , Asn (N) , Ser (S) , and His (H) ;
[0677] position 110: the amino acid residue at position 110 is selected from the group consisting of Arg (R) , His (H) , Pro (P) , Asp (D) , and Asn (N) ;
[0678] position 111: the amino acid residue at position 111 is selected from the group consisting of Ser (S) , Tyr (Y) , His (H) , Ala (A) , Phe (F) , and Asn (N) ;
[0679] position 112: the amino acid residue at position 112 is selected from the group consisting of Ser (S) , Gly (G) , Ala (A) , Asp (D) , and Asn (N) ;
[0680] position 113: the amino acid residue at position 113 is selected from the group consisting of Gly (G) , Ala (A) , and Thr (T) ;
[0681] position 5: the amino acid residue at position 5 is selected from the group consisting of Gln (Q) , and Val (V) ;
[0682] position 72: the amino acid residue at position 72 is selected from the group consisting of Gln (Q) , and Arg (R) ;
[0683] position 73: the amino acid residue at position 73 is selected from the group consisting of Asn (N) , and Asp (D) ;
[0684] position 75: the amino acid residue at position 75 is selected from the group consisting of Ala (A) , and Ser (S) ;
[0685] position 77: the amino acid residue at position 77 is selected from the group consisting of Ser (S) , and Asn (N) ;
[0686] position 87: the amino acid residue at position 87 is selected from the group consisting of Lys (K) , and Arg (R) ;
[0687] position 88: the amino acid residue at position 88 is selected from the group consisting of Pro (P) , and Ala (A) ;
[0688] position 93: the amino acid residue at position 93 is selected from the group consisting of Met (M) , and Val (V) ; and
[0689] position 123: the amino acid residue at position 123 is selected from the group consisting of Gln (Q) , and Leu (L) .
[0690] In some embodiments, the amino acid residue of the PD-L1 targeting nanobody is in unmodified form, protected form or modified form.
[0691] In some embodiments, the remaining amino acid residues are the same as the amino acid sequence shown in SEQ ID NO. 17.
[0692] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0693] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least one mutation as defined herein as compared to the amino acid sequence set forth in SEQ ID NO. 17.
[0694] In some embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0695] each of the variable amino acids is as defined herein.
[0696] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 17, and 18 to 38.
[0697] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 17, and 18 to 38.
[0698] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 17,
[0699] wherein the amino acid residue (s) at positions 50, 59, 103, 107, 110, 111, 112, and 113are as follows,
[0700] position 50: the amino acid residue at position 50 is selected from the group consisting of Gly (G) , Thr (T) , and Ala (A) ;
[0701] position 59: the amino acid residue at position 59 is selected from the group consisting of Phe (F) , and Ser (S) ;
[0702] position 103: the amino acid residue at position 103 is selected from the group consisting of Asp (D) , and Asn (N) ;
[0703] position 107: the amino acid residue at position 107 is selected from the group consisting of Thr (T) , Asn (N) , Ser (S) , and His (H) ;
[0704] position 110: the amino acid residue at position 110 is selected from the group consisting of Arg (R) , His (H) , Pro (P) , Asp (D) , and Asn (N) ;
[0705] position 111: the amino acid residue at position 111 is selected from the group consisting of Ser (S) , Tyr (Y) , His (H) , Ala (A) , Phe (F) , and Asn (N) ;
[0706] position 112: the amino acid residue at position 112 is selected from the group consisting of Ser (S) , Gly (G) , Ala (A) , Asp (D) , and Asn (N) ; and
[0707] position 113: the amino acid residue at position 113 is selected from the group consisting of Gly (G) , Ala (A) , and Thr (T) .
[0708] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0709] each of the variable amino acids is as defined herein.
[0710] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 17, and 18 to 36.
[0711] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 17, and 18 to 36.
[0712] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence having at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 17,
[0713] wherein the amino acid residue (s) at positions 5, 72, 73, 75, 77, 87, 88, 93, and 123 are as follows,
[0714] position 5: the amino acid residue at position 5 is selected from the group consisting of Gln (Q) , and Val (V) ;
[0715] position 72: the amino acid residue at position 72 is selected from the group consisting of Gln (Q) , and Arg (R) ;
[0716] position 73: the amino acid residue at position 73 is selected from the group consisting of Asn (N) , and Asp (D) ;
[0717] position 75: the amino acid residue at position 75 is selected from the group consisting of Ala (A) , and Ser (S) ;
[0718] position 77: the amino acid residue at position 77 is selected from the group consisting of Ser (S) , and Asn (N) ;
[0719] position 87: the amino acid residue at position 87 is selected from the group consisting of Lys (K) , and Arg (R) ;
[0720] position 88: the amino acid residue at position 88 is selected from the group consisting of Pro (P) , and Ala (A) ;
[0721] position 93: the amino acid residue at position 93 is selected from the group consisting of Met (M) , and Val (V) ; and
[0722] position 123: the amino acid residue at position 123 is selected from the group consisting of Gln (Q) , and Leu (L) .
[0723] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0724] each of the variable amino acids is as defined herein.
[0725] The present disclosure provides a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 17, and 37 to 38.
[0726] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 17, and 37 to 38.
[0727] The present disclosure provides an isolated nucleic acid molecule encoding any of the PD-L1 targeting nanobodies of the present disclosure.
[0728] The present disclosure provides an expression vector comprising the nucleic acid molecule encoding any of the PD-L1 targeting nanobodies of the present disclosure.
[0729] The present disclosure provides a host cell comprising the expression vector of the present disclosure or the nucleic acid molecule encoding any of the PD-L1 targeting nanobodies of the present disclosure.
[0730] SAR-COV-2 RBD TARGETING MINIPROTEIN
[0731] As used herein, the term "miniprotein" refers to a polypeptide characterized by a relatively small size, typically comprising less than about 100 amino acid residues (e.g., less than about 80, less than about 70, less than about 60, or less than about 50 amino acid residues) , and further characterized by its ability to adopt a defined and stable three-dimensional conformation. This stable conformation is often conferred by structural constraints such as, but not limited to, one or more disulfide bonds, a tightly packed hydrophobic core, specific secondary structure motifs (e.g., alpha-helices, beta-sheets) , metal coordination, backbone cyclization, or combinations thereof. Miniproteins can be derived from naturally occurring polypeptide sequences (e.g., fragments or domains thereof) , computationally designed de novo, engineered from existing protein or peptide scaffolds, or identified through library screening methods. The amino acid residues comprising the miniprotein may include naturally occurring L-amino acids, D-amino acids, non-natural amino acids, or amino acid analogs, linked predominantly by peptide bonds. The term "miniprotein" is intended to encompass such molecules regardless of their specific biological function or target-binding properties, unless explicitly stated otherwise, and distinguishes them from generally unstructured peptides of similar length by virtue of their defined tertiary structure and conformational stability.
[0732] As used herein, the term "SARS-CoV-2 RBD targeting miniprotein" refers to a miniprotein, as defined herein, that specifically binds to the Receptor Binding Domain (RBD) of the Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) spike (S) protein. A miniprotein, as previously defined, is a polypeptide typically comprising less than about 100 amino acid residues and characterized by its ability to adopt a defined and stable three-dimensional conformation, encompassing molecules of natural, designed, or engineered origin. The characteristic of "SARS-CoV-2 RBD targeting" signifies that said miniprotein recognizes and binds to an epitope present on the SARS-CoV-2 RBD. Unless specified otherwise, this binding includes interaction with the RBD from various SARS-CoV-2 strains or variants of concern or interest (including, but not limited to, reference strains such as Hu-1, or variants such as Alpha, Beta, Gamma, Delta, Omicron and its sublineages, or subsequently emerging variants) , as well as recombinant or fragment forms of the RBD, provided the relevant epitope is present. Specific binding generally implies a measurable binding affinity that is significantly higher than its affinity for unrelated proteins (such as human serum albumin or other viral proteins) under comparable assay conditions. The term therefore encompasses any miniprotein exhibiting such specific binding to the SARS-CoV-2 RBD, irrespective of the particular epitope bound within the RBD, the specific origin or design of the miniprotein, or the functional consequences of such binding (e.g., neutralization of viral entry, inhibition of ACE2 receptor binding) , unless explicitly stated otherwise.
[0733] The present disclosure provides a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0734] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having deletion, substitution, insertion or addition of one or more amino acid residues as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0735] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations selected from the group consisting of D11C, K26C, K26G, K26Q, K27R, F30C, and Y40C as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0736] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations selected from the group consisting of K26G, K26Q, K27R, and F30C as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0737] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in any of SEQ ID Nos. 55, and 56 to 61.
[0738] The present disclosure provides a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 55,
[0739] wherein the amino acid residue (s) at positions 11, 26, 27, 30, and 40 are as follows,
[0740] position 11: the amino acid residue at position 11 is selected from the group consisting of Asp (D) , and Cys (C) ;
[0741] position 26: the amino acid residue at position 26 is selected from the group consisting of Lys (K) , Cys (C) , Gly (G) , and Gln (Q) ;
[0742] position 27: the amino acid residue at position 27 is selected from the group consisting of Lys (K) , and Arg (R) ;
[0743] position 30: the amino acid residue at position 30 is selected from the group consisting of Phe (F) , and Cys (C) ; and
[0744] position 40: the amino acid residue at position 40 is selected from the group consisting of Tyr (Y) , and Cys (C) .
[0745] In some embodiments, the amino acid residue of the SAR-CoV-2 RBD targeting miniprotein is in unmodified form, protected form or modified form.
[0746] In some embodiments, the remaining amino acid residues are the same as the amino acid sequence shown in SEQ ID NO. 55.
[0747] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0748] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having at least one mutation as defined herein as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0749] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0750] each of the variable amino acids is as defined herein.
[0751] The present disclosure provides a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity and / or less than 100%identity as compared to the amino acid sequence set forth in any of SEQ ID NOs. 55, and 56 to 61.
[0752] In some preferred embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61.
[0753] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations, for example, 1, 2, 3, or 4 mutations at the position selected from the group consisting of positions 11, 26, 27, 30, and 40 as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0754] In some embodiments, in the amino acid sequence of the SAR-CoV-2 RBD targeting miniprotein, each of the amino acid residues at positions 11, 26, 30, and 40 is independently mutated to C as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0755] In some preferred embodiments, in the amino acid sequence of the SAR-CoV-2 RBD targeting miniprotein, only one of the amino acid residues at positions 11, 26, 30, and 40 is mutated to C as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0756] In some preferred embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations selected from the group consisting of D11C, K26C, F30C, and Y40C.
[0757] In some preferred embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one mutation selected from the group consisting of D11C, K26C, F30C, and Y40C.
[0758] In some preferred embodiments, the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 56 to 59.
[0759] In some embodiments, in the amino acid sequence of the SAR-CoV-2 RBD targeting miniprotein, as compared to the amino acid sequence set forth in SEQ ID NO. 55, the amino acid residue (s) at positions 26, 27, and 30 are as follows,
[0760] position 26: the amino acid residue at position 26 is selected from the group consisting of Lys (K) , Gly (G) , and Gln (Q) ;
[0761] position 27: the amino acid residue at position 27 is selected from the group consisting of Lys (K) , and Arg (R) ;
[0762] position 30: the amino acid residue at position 30 is selected from the group consisting of Cys (C) .
[0763] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as follows, (the subscripts indicate the positions of amino acid residues which are variable)
[0764] each of the variable amino acids is as defined herein.
[0765] In some embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations, for example, 1, 2, or 3 mutations at the position selected from the group consisting of positions 26, 27, and 30 as compared to the amino acid sequence set forth in SEQ ID NO. 55.
[0766] In some preferred embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations, for example, 1, 2, or 3 mutations selected from the group consisting of K26G, K26Q, K27R, and F30C.
[0767] In some preferred embodiments, the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in any of SEQ ID NOs. 58, 60, and 61.
[0768] The SAR-CoV-2 RBD targeting miniproteins disclosed herein encompass not only the naturally occurring, unmodified forms of the amino acid sequences set forth in the present disclosure, but also chemically or biologically modified derivatives thereof. Such modifications include, but are not limited to, the introduction of substituents and / or the addition of protecting groups to the side chains of amino acid residues. Exemplary modifications include, without limitation, phosphorylation, glycosylation, acetylation, methylation, amidation, sulfonation, and pegylation, as well as other post-translational modifications and artificial synthetic modifications.
[0769] It will be understood by those skilled in the art that, provided the core functional domain of the protein remains substantially unaltered, variations in the type, number, and point of attachment of substituents and protecting groups do not depart from the scope of the present disclosure. Accordingly, equivalent variants of the aforementioned side chain modifications, wherein the amino acid sequence corresponds to the amino acid sequences disclosed herein, are considered reasonable extensions of the scope of the claims and fall within the definition of the SAR-CoV-2 RBD targeting miniproteins of the present disclosure.
[0770] The present disclosure provides an isolated nucleic acid molecule encoding any of the SAR-CoV-2 RBD targeting miniproteins of the present disclosure.
[0771] The present disclosure provides an expression vector comprising the nucleic acid molecule encoding any of the SAR-CoV-2 RBD targeting miniproteins of the present disclosure.
[0772] The present disclosure provides a host cell comprising the expression vector of the present disclosure or the nucleic acid molecule encoding any of the SAR-CoV-2 RBD targeting miniproteins of the present disclosure.
[0773] USAGE AND ADMINISTRATION
[0774] The covalent protein of the present disclosure can be used as medicaments.
[0775] The therapeutic effect of the covalent protein of the present disclosure primarily originates from the biological effect (s) produced upon interaction of the protein moiety (protein moiety Ab) contained therein with its specific target (s) .
[0776] In some embodiments, the protein moiety Ab in the covalent protein of the present disclosure is a nanobody capable of specifically binding to Programmed Death-Ligand 1 (PD-L1) (aPD-L1 targeting nanobody) , the covalent protein of the present disclosure can be used for treating or preventing a disease or condition associated with abnormal expression or dysregulated activity of PD-L1.
[0777] The disease or condition associated with abnormal PD-L1 expression or dysregulated activity may be specifically selected from the following categories:
[0778] (a) Cancer: In such diseases, the expression of PD-L1 is typically upregulated in the tumor microenvironment, aiding tumor cells in evading immune surveillance. The covalent protein of the present disclosure (comprising a PD-L1 targeting nanobody) can, for example, by blocking the interaction of PD-L1 with its receptor (such as PD-1) , restore or enhance anti-tumor immune responses. In some embodiments, said cancer includes, but is not limited to: lung cancer (e.g., non-small cell lung cancer, small cell lung cancer) , melanoma, renal cell carcinoma, bladder cancer (e.g., urothelial carcinoma) , head and neck squamous cell carcinoma, Hodgkin's lymphoma, non-Hodgkin's lymphoma, gastric cancer (e.g., gastric adenocarcinoma) , esophageal cancer (e.g., esophageal squamous cell carcinoma, esophageal adenocarcinoma) , hepatocellular carcinoma, cervical cancer, ovarian cancer, breast cancer (e.g., triple-negative breast cancer) , colorectal cancer, pancreatic cancer, prostate cancer, glioblastoma, or any combination or metastatic forms of the aforementioned cancers.
[0779] (b) Autoimmune diseases: In such diseases, the PD-L1 signaling pathway may be crucial for maintaining immune homeostasis and suppressing autoreactive immune cells. If said PD-L1 targeting nanobody possesses agonistic activity and is capable of enhancing PD-L1-mediated immunosuppressive signals (e.g., by stabilizing PD-L1 expression, promoting its binding to its receptor (s) , or mimicking downstream signals) , then the covalent protein of the present disclosure can be used to treat such diseases. In some embodiments, said autoimmune disease includes, but is not limited to: rheumatoid arthritis, systemic lupus erythematosus, multiple sclerosis, type 1 diabetes, inflammatory bowel disease (e.g., Crohn's disease or ulcerative colitis) , psoriasis (e.g., plaque psoriasis) , autoimmune hepatitis, myasthenia gravis, syndrome, autoimmune thyroiditis (e.g., Hashimoto's thyroiditis) , vasculitis, or graft-versus-host disease.
[0780] In some embodiments, the protein moiety Ab in the covalent protein of the present disclosure is a miniprotein capable of specifically binding to the receptor-binding domain (RBD) of Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) (aSARS-CoV-2 RBD-targeting miniprotein) , the covalent protein of the present disclosure can be used for treating or preventing a disease or condition caused by coronavirus infection (particularly SARS-CoV-2 and its variants) , such as Coronavirus Disease 2019 (COVID-19) .
[0781] Said protein moiety (SARS-CoV-2 RBD-targeting miniprotein) can, for example, neutralize the virus, prevent the binding of the virus to host cell surface receptors (such as ACE2) , thereby inhibiting viral entry into host cells and subsequent replication.
[0782] The treatment or prevention of COVID-19 may include one or more of the following aspects:
[0783] Prophylactic application: for pre-exposure or post-exposure prophylaxis, reducing the risk of an individual contracting SARS-CoV-2 or preventing disease progression after infection.
[0784] Therapeutic application:
[0785] Alleviating one or more symptoms of COVID-19 (e.g., fever, cough, dyspnea, fatigue, loss of taste or smell, etc. ) ;
[0786] Reducing the severity of COVID-19 disease (e.g., progression from mild / moderate to severe / critical illness) ;
[0787] Shortening the duration of illness or time to viral clearance;
[0788] Preventing or mitigating COVID-19-related complications, said complications including, but not limited to: pneumonia, Acute Respiratory Distress Syndrome (ARDS) , Cytokine Release Syndrome (CRS, also known as cytokine storm) , multi-organ dysfunction or failure (e.g., renal failure, heart failure) , thromboembolic events (e.g., deep vein thrombosis, pulmonary embolism) , and long COVID (also known as post-COVID conditions) .
[0789] The covalent protein of the present disclosure is administered by any suitable route in the form of a pharmaceutical composition adapted to such a route, and in a dose effective for the treatment intended. The covalent protein of the present disclosure may be administered orally, rectally, vaginally, parenterally, or topically.
[0790] As used herein, the terms “administration” and “administer” refer to absorbing, ingesting, injecting, inhaling, implanting, or otherwise introducing the covalent protein of the present disclosure, or a pharmaceutical composition thereof. The terms “treatment” and “treat” refer to reversing, alleviating, delaying the onset of, or inhibiting the progress of a “pathological condition” (e.g., a disease, disorder, or condition, or one or more signs or symptoms thereof) described herein. In some embodiments, treatment may be administered after one or more signs or symptoms of a disease or condition have developed or have been observed. In other embodiments, treatment may be administered in the absence of signs or symptoms of the disease or condition. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors) . Treatment may also be continued after symptoms have resolved, for example, to delay or prevent recurrence. As used herein, the terms “disease” , “disorder” , “condition” , and “pathological condition” are used interchangeably.
[0791] Dosage levels for administration can be determined by those skilled in the art by routine experimentation. The dosage regimen for the covalent protein of the present disclosure and / or a composition comprising said protein is based on a variety of factors, including the type, age, weight, sex, and medical condition of the patient; the severity of the condition; the route of administration; and the activity of the particular compound employed. Thus, the dosage regimen may vary widely. For example, dosage levels for the covalent protein of the present disclosure may be from about 0.001 to about 100 mg / kg (i.e., mg per kilogram of body weight) per day. In some embodiments, the total daily dose of a covalent protein of the present disclosure, administered in single or divided doses, may be from about 0.001 to about 10 mg / kg. It is not uncommon that the administration of the covalent protein of the present disclosure may be repeated a plurality of times in a day.
[0792] In some embodiments, the covalent protein of the present disclosure may be administered in combination with one or more of additional therapeutical agents. In some embodiments, non-limiting examples of the additional therapeutical agents may include an anti-tumor agent. The additional therapeutical agent (s) can be administered before, after, or at the same time that the covalent protein of the present disclosure is administered.
[0793] As used herein, the term “anti-tumor agent” refers to any agent which is administered to a subject suffered from a cancer for the purposes of treating the cancer. Conventional surgery or radiotherapy or medicinal therapy may be used in combination with the covalent protein of the present disclosure for cancer treatment.
[0794] PHARMACEUTICAL COMPOSITIONS
[0795] In some aspect, the present disclosure is directed to a pharmaceutical composition comprising the covalent protein as provided herein, and at least one pharmaceutically acceptable diluent, carrier or excipient.
[0796] As used herein, the term “pharmaceutically acceptable diluent, carrier or excipient” refers to a diluent, carrier or excipient which is useful for preparing a pharmaceutical composition that is generally safe, non-toxic, and neither biologically nor otherwise undesirable, and includes diluent, carrier or excipient that is acceptable for veterinary use as well as human pharmaceutical use. A pharmaceutically acceptable diluent, carrier or excipient as used herein includes both one and more than one such diluent, carrier or excipient. The particular diluent, carrier or excipient used will depend upon the means and purpose for which the covalent protein of the present disclosure is being applied. Suitable diluents, carriers and excipients are well known to those skilled in the art and are described in detail in, e.g., Ansel, Howard C, et al., Ansel’s Pharmaceutical Dosage Forms and Drug Delivery Systems. Philadelphia: Lippincott, Williams & Wilkins, 2004; Gennaro, Alfonso R., et al., Remington: The Science and Practice of Pharmacy. Philadelphia: Lippincott, Williams & Wilkins, 2000; and Rowe, Raymond C. Handbook of Pharmaceutical Excipients. Chicago, Pharmaceutical Press, 2005. One or more of buffers, stabilizing agents, surfactants, wetting agents, lubricating agents, emulsifiers, suspending agents, preservatives, antioxidants, opaquing agents, glidants, processing aids, colorants, sweeteners, perfuming agents, flavoring agents, and other known additives may also be included to provide an elegant presentation of the drug (i.e., the covalent protein or pharmaceutical composition as provided herein) or aid in the manufacturing of the pharmaceutical product (i.e., medicament) .
[0797] The compositions of the present disclosure may be formulated in a variety of forms. These include, for example, liquid, semi-solid, and solid dosage forms, such as liquid solutions (e.g., injectable and infusible solutions) , dispersions, or suspensions, tablets, pills, powders, liposomes, suppositories, etc. The form depends on the intended mode of administration and therapeutic application.
[0798] Pharmaceutical compositions of the present disclosure may be prepared by any of the well-known techniques of pharmacy, such as effective formulation and administration procedures. The above considerations in regard to effective formulations and administration procedures are well known in the art, and are described in standard textbooks. Formulation of pharmaceutical products is discussed in, e.g., Hoover, John E., Remington’s Pharmaceutical Sciences, Mack Publishing Co., Easton, Pennsylvania, 1975; Liberman, et al., Eds., Pharmaceutical Dosage Forms, Marcel Decker, New York, N.Y., 1980; and Kibbe, et al., Eds., Handbook of Pharmaceutical Excipients, 3rd Edition, American Pharmaceutical Association, Washington, 1999.
[0799] In some embodiments, the pharmaceutical compositions comprise the covalent protein as provided herein, in combination with one or more of additional therapeutical agents such as an anti-tumor agent, and at least one pharmaceutically acceptable diluent, carrier or excipient.
[0800] In a further aspect, the present disclosure relates to a kit for treating a disease, disorder or condition, which is selected from a cancer, an autoimmune disease, and other immune-related diseases, which comprises a covalent protein as provided herein, or a pharmaceutical composition comprising the covalent protein as provided herein, a container, and optionally a package insert or label indicating treatment. In some embodiments, the kit may further contain one or more additional therapeutical agents such as an anti-tumor agent.
[0801] METHODS OF TREATMENT
[0802] In a further aspect, the present disclosure is directed to a method of treating a disease, disorder, or condition, for example, a cancer, an autoimmune disease, other immune-related diseases, and coronavirus infection in a subject in need thereof, which comprises administering to the subject a therapeutically effective amount of the covalent protein as provided herein.
[0803] As used herein, the term “subject in need thereof” is a subject having a disease, disorder, or condition, for example, a cancer, an autoimmune disease, or other immune-related diseases, or a subject having an increased risk of developing a disease, disorder, or condition, for example, a cancer, an autoimmune disease, other immune-related diseases, or coronavirus infection relative to the population at large. In some embodiments, the subject is a warm-blooded animal. In some embodiments, the warm-blooded animal is a mammal. In some embodiments, the warm-blooded animal is a human.
[0804] The method of treating a disease, disorder, or condition, for example, a cancer, an autoimmune disease, other immune-related diseases, or coronavirus infection as described herein may be used as a monotherapy. As used herein, the term “monotherapy” refers to the administration of a single active or therapeutic compound to a subject in need thereof. In some embodiments, monotherapy will involve administration of a therapeutically effective amount of one of the covalent protein of the present disclosure to a subject in need of such treatment.
[0805] Depending upon the particular disease or condition to be treated, the method of treating a disease, disorder, or condition, for example, a cancer, an autoimmune disease, other immune-related diseases, or coronavirus infection described herein may involve, in addition to administration of the covalent protein, combination therapy of one or more additional therapeutic agent (s) , for example, an anti-tumor agent. As used herein, the term “combination therapy” refers to the administration of a combination of multiple active therapeutic agents. In some embodiments, the covalent protein of the present disclosure may be administered simultaneously, separately or sequentially to treatment with the one or more additional therapeutic agent (s) . For example, the additional therapeutic agent (s) may be administered separately from the covalent protein of the present disclosure, as part of a multiple dosage regimen. Alternatively, the additional therapeutic agent (s) may be part of a single dosage form, mixed with the covalent protein of the present disclosure in a single composition.
[0806] In a further aspect, the present disclosure is directed to use of the covalent protein as provided herein in the manufacture of a medicament for treating a disease, disorder, or condition, for example, a cancer, an autoimmune disease, other immune-related diseases, or coronavirus infection in a subject in need thereof.
[0807] All publications and patents mentioned herein are hereby incorporated by reference in their entirety as if each individual publication or patent were specifically and separately indicated as not incorporated by reference. In the event of conflict, this application (including any definitions herein) shall control. However, any references, articles, publications, patents, patent publications and patent applications cited herein are not and shall not be deemed to be admissions or suggestions of any kind which constitute valid prior art or form part of the common knowledge of any country in the world.
[0808] The present disclosure is further described in the following examples, which are not limiting the scope of the present disclosure as described in the claims.
[0809] EXAMPLES
[0810] GENERAL INFORMATION
[0811] a. Chemicals
[0812] All chemicals were obtained from commercial supplier at the highest commercial quality and used without further purification unless otherwise stated. Tetrahydrofuran (THF) , dicloromethane (DCM) , triethylamine (TEA) and petroleum ether (PE) were purchased from J&K Scientific. N-Bromosuccinimide (NBS) , 1, 8-Diazabicyclo [5.4.0] undec-7-ene (DBU) , Benzoyl peroxide (BPO) , and [4- (acetylamino) phenyl] imidodisulfuryl difluoride (AISF) were purchased from Bidepharm. Analytical thin-layer chromatography (TLC) was carried out on SiLiDa silica gel GF254 plates, using UV at 254 nm or staining with phosphomolybdic acid (PMA) for visualization. Column chromatography was performed with normal phase silica gel (300 -400 mesh) .
[0813] b. NMR
[0814] NMR spectra were recorded on Bruker AVANCE 500 or AVANCE 600 spectrometer in CDCl3 or CD3OD or d6-acetone using tetramethylsilane (TMS) as internal standard unless otherwise stated. Data are presented in the following space: chemical shift, multiplicity, coupling constant in hertz (Hz) , and signal area integration in natural numbers.
[0815] c. Primary cell lines and animals
[0816] The human cancer cells U87 and H460 were purchased from the American Type Culture Collection (ATCC) . U87 was maintained in MEM supplemented with 10 %fetal bovine serum (FBS) (Sigma Cat. F8318) . Jurkat cells were maintained in RPMI 1640 medium (Gibco) supplemented with 10 %FBS. HEK293T cells and human colorectal adenocarcinoma cells Caco-2 were cultured with Dulbecco’s modified Eagle’s medium (DMEM) supplemented with 10 %FBS. All cell lines were cultured in a humidified incubator at 37℃ and 5 %CO2. The culture medium was supplemented with 100 mg / mL streptomycin and 100 U / mL penicillin.
[0817] The care and experimental use of animals followed the guidelines set by the Institutional Animal Care and Use Committee (IACUC) of Westlake University (Hangzhou, China) . The 6-8-week old female C57BL / 6 mice were purchased from the Laboratory Animal Resources Center at Westlake University. Female B6-hPD1 / hPDL1 mice (Strain NO. T004022) , which were genetically modified to express human PD-1 and PD-L1 immune checkpoint proteins, were purchased from GemPharmatech (Nanjing, China) .
[0818] d. Generation of stable cell lines
[0819] To generate stable cell lines, cell membrane-anchored chimeric CD3scFv were constructed by fusing the single chain fragment variable domain (scFv) of OKT3 with human γ1 to the C-terminal domain of mouse CD80 (accession number: NP_033985.3) . The chimeric sequence was synthesized and cloned into a lentivirus vector containing the PuroR gene. The NFAT reporter response element was synthesized and cloned into the pGL4.32 [luc2P NFkB-RE Hygro] vector (Promega; cat. E8491) to generate pGL [luc2P NFAT-RE Hygro] . Human PD-1 (accession number: NP_005009.2) and PD-L1 (accession number: NP_001254635.1) sequences were amplified from the human cDNA library and inserted into the pLVX-CMV600 vector (Addgene, Cat. 110723) , which contained the NeoR gene with the PGK promoter for Neomycin selection. Mouse PD-L1 gRNA #1 (GCCTGCTGTCACTTGCTACG) (SEQ ID NO. 70) and gRNA #2 (TAGAACCCACTGAAAAGATT) (SEQ ID NO. 71) were synthesized and cloned into the CRISPR V2 vector.
[0820] Jurkat cells (clone E6-1) were transfected with pGL [luc2P NFAT-RE Hygro] using BTX Electroporators and subjected to hygromycin (300ug / ml; Thermo Fisher, Cat. 10687010) selection for 10 days. The Jurkat / NFAT-Luc stable clone 72 was selected by single cell plating and confirmed by TCR activation assay. The Jurkat / NFAT-Luc clone 72 was then infected with human PD-1 lentivirus and selected by Geneticin (500 μg / ml; Thermo Fisher, Cat. 10131035) for 2 weeks to generate Jurkat / NFAT-Luc / PD-1. HEK293T cells were infected with human PD-L1 lentivirus and selected by Geneticin for 1 week. Subsequently, the cells were infected with chimeric CD3scFv lentivirus and selected with Puromycin (2ug / ml) for 1 week to generate HEK293T / hPD-L1 / aAPC cells.
[0821] The PD-L1 gene of MC38-K cells was knocked out using mouse PD-L1 CRISPR-Cas9 gRNA lentivirus, followed by puromycin selection for 1 week. The cell pool was then infected with human PD-L1 lentivirus and selected with Geneticin for 1 week. MC38 / mPD-L1KO / hPD-L1 clone 29 were generated through single-cell plating.
[0822] d. LC-MS analysis
[0823] LC-MS analyses were performed using an Agilent 1260-6230 single TOF LC / MS. The column used was Agilent ZORBAX 5 μm 300SB-CN.
[0824] EXAMPLE 1. SYNTHESIS OF CROSSLINKERS
[0825] Preparation of crosslinkers
[0826] Example 1-1 Synthesis of Compound 1 and Its Derivatives
[0827] The known compound 1 and its derivatives were prepared according to the published strategy with some modifications.
[0828] General Procedure A
[0829] STEP 1: This procedure was adapted from De Borggraeve et al. In short, Chamber A of a dried small two-chamber reactor was filled with 1, 1'-sulfonyldiimidazole (SDI, 1.5 eq. ) and KF (4.0 eq. ) . Next, chamber B was charged with the appropriate (hetero) aryl alcohol (1.0 eq. ) , TEA (2.0 eq. ) and DCM. Finally, TFA (1 mL per 1 mmol substrate) was added by injection through the septum in chamber A.
[0830] After stirred at room temperature for 15 hours, one of the caps was carefully removed to release the residual pressure. Next, the content of chamber B was transferred to a round-bottomed flask. The solvent was removed on vacuum and the residue was purified by silica gel column chromatography to afford aryl fluorosulfates.
[0831] STEP 2: To a -20℃ solution of aryl fluorosulfates (1.0 eq. ) in a mixture of MeOH / THF (1: 7) was added NaBH4 (2.0 eq. ) . The reaction mixture was then warmed to 0℃ and stirred for 40 minutes before quenching with 1 M HCl. The result mixture was diluted with H2O and extracted with EtOAc three times. The organic layers were washed with saturated NaCl and dried over anhydrous Na2SO4. The organic phase was concentrated, and the residue was chromatographed on silica gel.
[0832] STEP 3: The above obtained product (1.0 eq. ) was dissolved in anhydrous ether, and 1M PBr3 (0.40 eq. ) dissolved in DCM was added and stirred at 0℃ for 2 hours. After the reaction was completed, ether and saturated NaHCO3 were added to perform extraction. The organic layer was dried over anhydrous Na2SO4, filtered and concentrated under reduced pressure. The residue was purified by silica gel column chromatography to give compound 1 or 1-F.
[0833] General Procedure B
[0834] STEP 1: For the synthesis of 1-CN and 1-NO2, the first step is the same as that of compound 1.
[0835] STEP 2: To a solution of the above obtained product (1.1 eq. ) and NBS (1.0 eq. ) in CCl4 was added BPO (0.1 eq. ) . Then the reaction was stirred at 80℃ for 12 hours. The solvent was evaporated, and the residue was purified by silica gel column chromatography to obtain the title compound.
[0836] 1H NMR (500 MHz, CDCl3) δ 7.50 (d, J = 8.5 Hz, 2H) , 7.32 (d, J = 8.5 Hz, 2H) , 4.48 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 37.89.13C NMR (125 MHz, CDCl3) δ 149.61, 138.67, 131.08, 121.34, 31.34. These data are in agreement with literature data.
[0837] 1H NMR (500 MHz, CDCl3) δ 7.41 -7.38 (m, 1H) , 7.35 -7.33 (m, 1H) , 7.27 -7.25 (m, 1H) , 4.44 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 39.47 (d, J = 10.0 Hz) , -126.47 (d, J = 10.0 Hz) . 13C NMR (125 MHz, CDCl3) δ 153.56 (d, J = 1020 Hz, C-F) , 140.75 (d, J = 25 Hz, C-F) , 136.95 (d, J = 50 Hz, C-F) , 125.78 (d, J = 15 Hz, C-F) , 123.75, 118.75 (d, J = 70 Hz, C-F) , 30.80 (d, J = 10 Hz, C-F) .
[0838] 1H NMR (500 MHz, CDCl3) δ 7.82 -7.80 (m, 1H) , 7.78 -7.76 (m, 1H) , 7.55 -7.53 (m, 1H) , 4.48 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 41.02.13C NMR (125 MHz, CDCl3) δ 13C NMR (126 MHz, C6D6) δ 149.47, 139.99, 135.55, 134.91, 122.99, 112.94, 108.01, 29.65.
[0839] 1H NMR (500 MHz, CDCl3) δ 8.21 (d, J = 2.5 Hz, 1H) , 7.80 (dd, J = 8.5, 2.0 Hz, 1H) , 7.56 (dd, J = 8.5, 1.0 Hz, 1H) , 4.52 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 42.87.13C NMR (125 MHz, CDCl3) δ141.13, 140.30, 135.52, 127.22, 124.34, 29.44
[0840] Example 1-2 Synthesis of Compound 2 and Its Derivatives
[0841] The known compound 2 and its derivatives were prepared according to the following method.
[0842] STEP 1: For the synthesis of 2 and its derivatives, the first step is the same as that of compound 1.
[0843] STEP 2: To a 0℃ solution of the above obtained product aryl fluorosulfate (1.0 eq. ) in DCM under an Ar atmosphere was added TEA (1.2 eq. ) and bromoacetyl bromide (1.1 eq. ) . The reaction mixture was stirred for 2 hours, then warmed to room temperature and quenched with H2O. The layers were separated and the aqueous phase was extracted with DCM. The combined organic layers were washed with 5%HCl, H2O, saturated NaHCO3, brine in turn and dried over anhydrous Na2SO4. After removing the solvent under reduced pressure, the crude product was purified by silica gel column chromatography to obtain the title compound.
[0844] 1H NMR (500 MHz, d6-acetone) δ 9.85 (s, 1H) , 7.87 (d, J = 9.0 Hz, 2H) , 7.51 (d, J = 9.0 Hz, 2H) , 4.06 (s, 2H) . 19F NMR (470 MHz, d6-acetone) δ 35.95.13C NMR (125 MHz, d6-acetone) δ 165.04, 145.74, 139.38, 121.60, 120.99, 29.22.
[0845] 1H NMR (500 MHz, CDCl3) δ 8.27 (s, 1H) , 7.81 -7.79 (m, 1H) , 7.41 -7.37 (m, 1H) , 7.27 -7.25 (m, 1H) , 4.03 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 39.05 (d, J = 10.0 Hz) , -124.64 (d, J = 10.0 Hz) . 13C NMR (125 MHz, CDCl3) δ 163.72, 154.66 (d, J = 1010 Hz, C-F) , 138.34 (d, J = 35 Hz, C-F) , 133.46 (d, J = 55 Hz, C-F) , 123.61, 115.66 (d, J = 15 Hz, C-F) , 109.42 (d, J = 90 Hz, C-F) , 29.06.
[0846] 1H NMR (500 MHz, CDCl3) δ 8.44 (s, 1H) , 8.17 -8.15 (d, J = 2.5 Hz, 1H) , 7.88 (dd, J = 9.0, 2.5 Hz, 1H) , 7.54 (d, J = 9.5 Hz, 1H) , 4.05 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 40.44.13C NMR (125 MHz, CDCl3) δ 164.17, 145.91, 137.92, 125.43, 124.69, 123.24, 112.89, 108.00, 28.83.
[0847] 1H NMR (500 MHz, CDCl3) δ 8.47 (s, 1H) , 8.44 -8.43 (m, 1H) , 8.12 (dd, J = 9.0, 2.5 Hz, 1H) , 7.55 (d, J = 9.0 Hz, 1H) , 4.07 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 42.40.13C NMR (125 MHz, CDCl3) δ164.17, 141.38, 138.06, 137.66, 125.45, 124.63, 117.45, 28.83.
[0848] Example 1-3 Synthesis of Compound 3
[0849] The compound 3 was obtained from commercial supplier and used without further purification.
[0850] Example 1-4 Synthesis of Compound 4
[0851] The compound 4 was prepared according to the following method.
[0852] STEP 1: A solution of 4-aminophenol (2.0 g, 18.3 mmol, 1.0 eq. ) in DCM (16 mL) and saturated NaHCO3 in water (16 mL) was stirred for 10 minutes at room temperature, then acryloyl chloride (1.6 mL, 20.2 mmol, 1.1 eq. ) was added dropwise, and the reaction stirred for an additional 6 hours at room temperature. The resulting solid was collected by filtration, washed with water and dried under vacuum to afford 3.0 g of 4-acrylamido-phenol.
[0853] STEP 2: To a one-dram vial containing the 4-acrylamido-phenol (100 mg, 0.6 mmol, 1.0 eq. ) and
[0854] AISF (231 mg, 0.72 mmol, 1.2 eq. ) was added THF (4 mL) followed by DBU (200 μL, 1.34 mmol, 2.2 eq. ) over a period of 30 seconds. The reaction mixture was stirred at room temperature for 10 minutes and then diluted with EtOAc and washed with 0.5 M HCl and brine. The combined organic
[0855] layers were dried with anhydrous Na2SO4 and concentrated under reduced pressure. The crude residue was purified by silica gel column chromatograph (EtOAc: PE = 30%) to afford 4 (90 mg, yield 60%) .
[0856] 1H NMR (500 MHz, CD3OD) δ 7.82 (dt, J = 9.0, 3.5 Hz, 2H) , 7.42 -7.40 (m, 2H) , 6.46 -6.37 (m, 2H) , 5.80 (dd, J = 9.0, 3.0 Hz, 1H) . 19F NMR (470 MHz, CD3OD) δ 34.97.13C NMR (125 MHz, CD3OD) δ 166.35, 147.44, 140.62, 132.29, 128.66, 122.78, 122.72. These data are in agreement with literature data.
[0857] Example 1-5 Synthesis of Compound 5 and Its Derivatives
[0858] The compound 5 and its derivatives were prepared according to the following method.
[0859] The detailed experimental procedures are the same as the General Procedure B.
[0860] 1H NMR (500 MHz, CDCl3) δ 7.46 -7.45 (m, 2H) , 7.38 (s, 1H) , 7.29 -7.27 (m, 1H) , 4.48 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 37.98.13C NMR (125 MHz, CDCl3) δ 150.19, 140.97, 131.00, 129.41, 121.68, 120.96, 31.43.
[0861] 1H NMR (500 MHz, CDCl3) δ 7.44 -7.42 (m, 1H) , 7.32 -7.29 (m, 1H) , 7.20 -7.17 (m, 1H) , 4.48 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 37.48, -115.34.13C NMR (150 MHz, CDCl3) δ 159.55 (d, J =1000 Hz, C-F) , 145.48, 127.75 (d, J = 70 Hz, C-F) , 123.78 (d, J = 16 Hz, C-F) , 122.93 (d, J = 36 Hz, C-F) , 117.70 (d, J = 96 Hz, C-F) , 23.75 (d, J = 16 Hz, C-F) .
[0862] 1H NMR (500 MHz, CDCl3) δ 8.20 -8.19 (m, 1H) , 7.63 (d, J = 2.5 Hz, 1H) , 7.49 (dd, J = 9.0, 2.5 Hz, 1H) , 4.83 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 39.77.13C NMR (125 MHz, CDCl3) δ 151.86, 147.17, 136.27, 128.11, 125.00, 121.91, 27.39.
[0863] Example 1-6 Synthesis of Compound 6 and Its Derivatives
[0864] The compound 6 and its derivatives were prepared according to the following method.
[0865] The detailed experimental procedures are the same as the General Procedure B.
[0866] 1H NMR (500 MHz, CDCl3) δ 7.57 -7.55 (m, 1H) , 7.46 -7.39 (m, 3H) , 4.53 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 40.52.13C NMR (125 MHz, CDCl3) δ 148.29, 132.47, 130.82, 130.49, 129.24, 121.39, 25.51.
[0867] 1H NMR (500 MHz, CDCl3) δ 7.90 (d, J = 1.5 Hz, 1H) , 7.76 (dd, J = 8.5, 2.0 Hz, 1H) , 7.56 (d, J =8.5 Hz, 1H) , 4.51 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 42.19.13C NMR (125 MHz, CDCl3) δ 150.25, 136.06, 134.23, 132.28, 122.34, 116.57, 113.55, 23.56.
[0868] 1H NMR (500 MHz, CDCl3) δ 8.47 (d, J = 2.5 Hz, 1H) , 8..32 (dd, J = 9.0, 2.5 Hz, 1H) , 7.62 (d, J =9.0 Hz, 1H) , 4.57 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 42.33.13C NMR (125 MHz, CDCl3) δ 151.29, 147.13, 132.41, 127.53, 125.70, 122.30, 23.76.
[0869] Example 1-7 Synthesis of Compound 7 and Its Derivatives
[0870] The compound 7 and its derivatives were prepared according to the following method.
[0871] The detailed experimental procedures are the same as the synthesis of compound 2 and its derivatives.
[0872] 1H NMR (500 MHz, CDCl3) δ 8.29 (s, 1H) , 7.80 (s, 1H) , 7.49 -7.45 (m, 2H) , 7.16 -7.15 (m, 1H) , 4.04 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 38.16.13C NMR (125 MHz, CDCl3) δ 163.91, 150.38, 139.04, 130.92, 119.78, 117.36, 112.84, 29.34.
[0873] 1H NMR (500 MHz, CDCl3) δ 8.52 (s, 1H) , 8.48 -8.47 (m, 1H) , 7.25 -7.22 (m, 1H) , 7.12 -7.11 (m, 1H) , 4.06 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 37.82, -129.55.13C NMR (125 MHz, CDCl3) δ 163.72, 151.28 (d, J = 980 Hz, C-F) , 145.66, 127.30 (d, J = 45 Hz, C-F) , 117.23 (d, J = 35 Hz, C-F) , 116.18 (d, J = 85 Hz, C-F) , 114.19, 28.97.
[0874] 1H NMR (500 MHz, CDCl3) δ 11.40 (s, 1H) , 8.97 (s, 1H) , 8.44 (d, J = 10.0 Hz, 1H) , 7.28 (s, 1H) , 4.12 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 40.39.13C NMR (125 MHz, CDCl3) δ 165.33, 153.63, 136.11, 135.39, 128.44, 116.24, 114.33, 29.07.
[0875] Example 1-8 Synthesis of Compound 8
[0876] The compound 8 was prepared according to the following method.
[0877] The detailed experimental procedures are the same as the synthesis of compound 2 and its derivatives.
[0878] 1H NMR (500 MHz, CDCl3) δ 8.51 (s, 1H) , 8.28 -8.26 (m, 1H) , 7.45 -7.42 (m, 2H) , 7.28 -7.27 (m, 1H) , 4.07 (s, 2H) . 19F NMR (470 MHz, CDCl3) δ 39.51.13C NMR (125 MHz, CDCl3) δ 163.91, 140.42, 129.64, 129.46, 126.34, 123.74, 121.47, 29.39.
[0879] EXAMPLE 2. CONSTRUCTION, EXPRESSION, AND PURIFICATION OF PROTEINS
[0880] The genes encoding KN035 and human hPD-L1 (amino acids 19-239) were synthesized (Genewiz, Suzhou) and cloned into pcDNA3.4 vector with C-terminal His tag. Gene encoding human PD-L1 (amino acid 19-134) was synthesized (Genewiz, Suzhou) and cloned into pET21b vector with C-terminal His tag. Single Cysteine (Cys) mutations were introduced by subcloning using pcDNA3.4-KN035 (WT. ) as a template. Genes encoding KN035 variants with C-terminal His tag, selected from yeast screening, were subcloned into pcDNA3.4 vector separately. Gene encoding H9_111A (IB101) with C-terminal 3 × flag tag was subcloned into pcDNA3.4 vector. LCB3 mini-proteins were constructed and expressed as reported. The gene encoding the protein sequence was synthesized and cloned into a modified pET-29b (+) E. coli plasmid expression vector, with N-terminal 8× His-tag followed by a TEV cleavage site (Genewiz, Suzhou) . Single Cys mutations were introduced at positions D11, K26, F30, and Y40.
[0881] KN035, hPD-L1 (amino acids 19-239) , KN035 variants, and KN035-Fc were expressed and purified from Expi293F cells using transient transfection. Expi293F cells were cultured in SMM 293-TII medium (Sino Biological, Lot. RZ14NO1601) at 37℃ under 5%CO2 in a shaker (140 rpm) . Transient transfections were conducted when cell density reached approximately 1.5x106 / mL using polyethyleneimine (Cat. 24765-1) . 1.5 mg of plasmid was premixed with 3 mg of PEI (1: 2) in 50 mL of fresh medium for 30 minutes before being added into a 1-L cell culture. Medium containing secreted proteins was harvested approximately 84 to 96 hours post-transfection.
[0882] For purification of KN035 and hPD-L1 (amino acids 19-239) : His-tagged proteins were purified by Ni-NTA chromatography. Harvested medium free of cells were loaded onto NiNTA beads (Genscript) and extensively washed with washing buffer (30 mM imidazole, PBS, pH 7.4) . The proteins were then eluted with elution buffer (250 mM imidazole, PBS, pH 7.4) . Eluted proteins were concentrated and then subjected to size-exclusion chromatography (SEC) (Superdex 75 Increase 10 / 300 GL or, Superdex 200 Increase 10 / 300 GL, GE Healthcare) in PBS. The peak fractions were aliquoted and snap frozen for future use.
[0883] For purification of KN035-Fc: the harvested medium free of cells was loaded onto ProteinA beads (GenScript, Cat. L00210-50) and washed with PBS. The proteins were then eluted with 0.1 M glycine, pH 3.0, and then neutralized with 0.2 M NaHCO3, pH 8.4 immediately. Eluted Proteins were concentrated and then subjected to SEC (Superdex 200 Increase 10 / 300 GL, GE Healthcare) in PBS. The peak fractions were aliquoted and snap frozen for future use.
[0884] For purification of H9_111A (IB101) -3 × flag: the harvested medium free of cells was loaded onto Flag affinity resin (GenScript, Cat. L00210-50) and washed with PBS. The proteins were then eluted with 0.2 mg / ml Flag peptides in PBS buffer. Eluted Proteins were concentrated and then subjected to SEC (Superdex 200 Increase 10 / 300 GL, GE Healthcare) in PBS. The peak fractions were aliquoted and snap frozen for future use.
[0885] Human PD-L1 (amino acid 19-134) was expressed in E. coli BL21 (DE3) cells. Following induction of expression with 1mM IPTG at 37℃ for 12 hours, the bacterial culture was centrifuged at 12,000 rpm for 2 min, and the supernatant was removed. The bacterial pellet was resuspended in PBS with protease inhibitor and then lysed by sonification in an ice-water bath. The lysate was then centrifuged at 12,000 g for 30 minutes at 4℃ to recover inclusion bodies and then resuspended in PBS, 10 mM EDTA. The resuspended solution was centrifuged at 12,000 g for 20 minutes at 4℃ and then resuspended the pellet in PBS, 6 M guanidine HCl (GuHCl) before being stirred at 4℃ overnight. The dissolved fraction was clarified by centrifugation at 12,000 g for 30 min at 4℃. The supernatant was loaded onto NiNTA beads (Genscript) and extensively washed with washing buffer (10 mM imidazole, PBS, pH 7.4, 6 M GuHCl) . The proteins were then eluted with elution buffer (250 mM imidazole, PBS, pH 7.4, 6 M GuHCl) . The eluted protein was diluted to below 0.1 mg / mL with PBS, 6 M GuHCl . The diluted protein solution was loaded into a dialysis bag (MWCO = 3 kDa) and the proteins were dialyzed and refolded by decreasing the GuHCl concentration by a gradient (3 M, 1 M, 0.5 M and 0 M) in a buffer containing PBS, pH 7.4. At each GuHCl concentration, dialysis was performed with stirring at 4℃ for 8 hours. The refolded protein solution in 0 M GuHCl buffer was dialyzed 3 times against PBS, concentrated and then subjected to SEC (Superdex 75 Increase 10 / 300 GL, GE Healthcare) in PBS.
[0886] The LCB3 mini protein was expressed in E. coli BL21 (DE3) cells. Following induction of expression with IPTG at 25℃ for 14-16 hours, the bacterial culture was centrifuged at 12,000 rpm for 2 min at 4℃, and the supernatant was removed. The bacterial pellet was resuspended in 50 ml of PBS (pH 7.4) and sonicated for 40 minutes in an ice-water bath. The lysate was then centrifuged at 12,000 g for 30 minutes, and the resulting supernatant was transferred to a new centrifuge tube.
[0887] The supernatant was bound to Ni-NTA resin and incubated for 1-2 hours. The resin was washed with wash buffer (PBS, pH7.4, 60 mM imidazole) for 10 column volumes. The protein was then eluted with elution buffer (PBS, 300 mM imidazole) . The eluted protein solution was concentrated by ultrafiltration centrifugation (Millipore, Cat. UFC900324, MWCO = 3 kDa) followed by further purification using SEC with Superdex 75 Increase 10 / 300 GL equilibrated with PBS.
[0888] All protein samples were characterized with SDS-PAGE and LC-MS, confirming a purity level of over 95%. Protein concentrations were determined using the BCA kit (Beyotime, Cat. P0012) .
[0889] EXAMPLE 3. SCREENING OF CYS MUTANTS FOR KN035 AND CROSSLINKERS
[0890] Chen et. al previously reported a covalent KN035 molecule, “GlueBody” , engineered using genetic code expansion technology14. They showed that GlueBody at 5 μM could achieve ~60%covalent crosslinking with PD-L1 after 5 hours of coincubation, while 1 hour under the same conditions can only crosslink ~10%PD-L114. KN035 was selected as a starting point for exploring new supercharged covalent protein.
[0891] We examined the complex structure of KN035 / PD-L1 and chose S100C, E102C, P104C, T107C, L108C, G113C, Q116C point mutation sites for chemical warheads installation24 (FIG. 2a, FIG. 7a) (SEQ ID NOs. 3 to 10) . These point mutation sites were chosen because they are proximal to the potential chemical warheads targeting sites H69, K62, K75, and K124 of PD-L1 (FIG. 7a) . The fluorosulfonyl or fluorosulfate-based crosslinkers synthesized in Example 1 (FIG. 2b, X = H) were individually attached to each KN035 Cys mutant to assess their covalent crosslinking kinetics with PD-L1. We found that the KN035_L108C and 6-H crosslinker combination (KN035_6-H) performed the best, affording 75%crosslinking with PD-L1 after 6 hours at 2 μM concentration (FIG. 2c, FIG. 7b) . This is comparable to the covalent crosslinking rate of the reported GlueBody (FIG. 7e-f) 14. MS / MS analysis revealed that the crosslinking occurred at His69 of PD-L1 (FIG. 7d) . Since our approach allows convenient, diverse small molecule installation. We synthesized more crosslinkers by attaching different electron-withdrawing groups at the warheads' ortho or para positions, fine-tuning the warheads' reactivity (FIG. 2b, X = -F, -CN, -NO2) . These substitutions could also interact locally with surrounding protein residues to influence the reacting group geometry. We then did the same screening for all KN035 Cys mutants with the modified crosslinkers (FIG. 8a) . We were glad to find that KN035_L108C and 6-CN combination could afford near complete crosslinking with PD-L1 within 4 hours (FIG. 8a) . During the attachment of 6-CN to KN035_L108C, we found the modified KN035_L108C protein had self-crosslinking side reactions with the proximal Tyr59 or Lys50 (FIG. 8b) , Y59F and K50A double mutation (CvKN035) successfully abolished the self-crosslinking side reaction and improved the crosslinking kinetics (FIG. 8b-c) . The convenient crosslinker installation method enabled unlimited possibilities for the choice of warheads. In comparison, introducing a new crosslinker using genetic code expansion technology may require a new tRNA / tRNA synthetase pair to be developed, and this might be a challenging project on its own.
[0892] Although the rate of the covalent CvKN035 improved, it still needs hours for the complete covalent blockade of PD-L1, which is significantly longer than the typical half-life of these miniproteins. Thus, further enhancing the covalent crosslinking rate is required. We then turned to yeast display to systematically screen the protein residues proximal to the chemical warheads, as the local sequence should affect the chemical warhead's chemical environment and conformation relative to the target residue on the PD-L1, thus affecting the crosslinking rate.
[0893] Several important factors have to be considered before establishing a high-throughput yeast display library for covalent KN035 selections. Firstly, our system depends on the chemical installation of crosslinkers on Cys residue. However, the yeast surface displayed Cys could be shielded by forming mixed disulfides with thiol molecules. Thus, it might be necessary to reduce these potential mixed disulfides to expose maximal Cys residues for warhead attachment. However, the most commonly used yeast display technology depends on the interchain disulfide bonds between Aga1p-Aga2p to display the protein (FIG. 1b) 22, 23. Adding reducing agents could potentially remove Aga2p and the displayed protein from the yeast surface, damaging the display system. Secondly, yeast selection has traditionally been used to screen for non-covalent binders; here, we specifically sought to select covalent binders with enhanced crosslinking kinetics. Therefore, it is key to distinguish covalent binders from non-covalent ones.
[0894] To address these potential problems and determine the viable procedure for covalent binder selection, we displayed the KN035_L108C on saccharomyces cerevisiae EBY100 yeast surface using the Aga1p-Aga2p system to explore the selection conditions (FIG. 3a) . Different concentrations of DTT were first added to the KN035_L108C displayed yeast, and 500 uM 10 minutes treatment was found to be tolerated by yeast without losing the displayed proteins (FIG. 9a-b) . Different concentrations of 6-CN were then added, followed by PD-L1. Different washing conditions were then performed to remove the non-covalent binders, including pH 3.0 acid wash. Yeast lacking 6-CN modification was used as a non-covalent binding control. We were glad to find that pH 3.0 acid wash was able to remove non-covalent binders in the control group while not affecting covalent binders (FIG. 3b) . 500 uM 10 minutes DTT treatment could also significantly enhance the covalent target protein binding on yeast surface (FIG. 9b) .
[0895] To explore the proximal residues’ influence on the crosslinking reaction, we generated site saturation mutagenesis libraries of KN035 near the binding interface between KN035 and PD-L1 (FIG. 10a) . The binding affinity of the mutants towards PD-L1 was assessed using flow cytometry.
[0896] Generation of KN035 L108C site-saturation mutagenesis clones
[0897] A set of primers containing NNK at each designed site were synthesized (Genewiz; Table 1) . Each site-saturation mutagenesis (SSM) plasmid was transformed into EBY100 chemically according to the manual from Frozen-EZ yeast transformation II kit (Zymo Research; USA) . All colonies were scrapped from SD agar plate and inoculated into SD-CAA liquid medium and induced in SG-CAA medium.
[0898] Our results revealed that mutations at positions R32, D99, S100, F101, V109, and G113 significantly reduced the binding affinity with PD-L1 (FIG. 10b) . Conversely, mutations at positions E102, S111, and A114 maintained comparable binding affinity with the wild-type KN035 (FIG. 10b) . Subsequently, we combined the libraries containing the favorable mutations and performed covalent protein selection using established screening conditions. After three rounds of screening, mutations S111H, E102F, and E102W were observed to be slightly enriched (FIG. 10c) . To assess the crosslinking efficiency, individual clones displaying the selected mutations were incubated with 5 nM PD-L1 for 5 minutes (FIG. 10c) . Our findings indicated that, compared to the wild-type KN035, all three mutations exhibited only marginal improvement in crosslinking efficiency. To further enhance the crosslinking rate, we evaluated double mutations, specifically S111H, E102F, and S111H, E102W. Although these double mutations resulted in a slightly accelerated crosslinking rate compared to single mutations (FIG. 10e-f) , the improvement achieved was still unsatisfactory. In conclusion, site-saturation mutagenesis highlights the importance of specific residues in the binding interface between KN035 and PD-L1 for maintaining high binding affinity. Significantly enhancing the crosslinking rate with PD-L1 must be the result of a combinatory effort of residues spatially proximal to the crosslinking site.
[0899] We proceeded to construct a randomized library, excluding elements crucial for preserving the binding affinity while ensuring a comprehensive size of 3×107 (FIG. 3c) . Subsequently, the library underwent modification with 6-CN and was co-incubated with PD-L1.
[0900] Generation of KN035 L108C library
[0901] The DNA library of KN035_L108C was constructed by two-step overlap extension PCR. A set of primers containing degenerative codons was synthesized (Genewiz; Table 1) . 1.5 μL of each primer at 10 μM was subsequently used to prepare 50 μl PCR reactions using KOD polymerase (KOD OneTM PCR Master Mix, Toyobo) . The nanobody DNA library pool was successively amplified for yeast transformations with PYAL-PLN. B-F and PYAL-PLN. B-R (Supplementary Table 1) . For yeast transformation, 100 mL of EBY100 yeast were grown to OD600 of 1.5 and treated with 100 mM Lithium acetate and 10 mM DTT, followed by washing with ice-cold water. Then, electrocompetent cells were transformed with 24 μg KN035_L108C insert DNA and 6 μg of pYAL plasmid, digested with BamHI-HF and HindIII-HF (New England BioLabs; USA) , using an ECM 830 Electroporator (BTX-Harvard Apparatus; USA) . Dilutions of transformed yeast were then plated on SD-CAA medium as single colonies to obtain an estimate of library diversity.
[0902] Selection of covalent nanobody binders for hPD-L1
[0903] Yeast cells corresponding to 10 times the size of the library were initially induced, washed, and resuspended in PBSA (PBS pH 7.4, 0.1% (w / v) BSA) . Subsequently, 50 nM of biotinylated hPD-L1 (ACRObiosystems, Cat. PD1-H82E5) was incubated with the yeast cells at 37℃ for 30 minutes. Anti-myc-PE antibody (Cell Signaling Technology; Cat. 3739S) and streptavidin-APC (Biolegend; Cat. 405207) were then added and incubated at room temperature for 30 minutes. The cells were washed at least three times with PBSA before being analyzed using flow cytometry with a CytoFLEX SRT (Backman Coulter; USA) to screen for binders. For all subsequent rounds of selections, 1.5 × 107 induced yeast cells were used. Before introducing the small molecule cross-linker 6-CN, the cells were treated with 500 nM DTT for 10 minutes. During the following three rounds of selections, the yeast cells were incubated with successively lower concentrations of hPD-L1 (10 nM, 5 nM, and 3 nM) to enrich for binders with higher affinities. Similarly, shorter incubation times (20 minutes, 10 minutes, and 3 minutes) were used to enrich for binders with faster reaction rates. After FACS selection, yeast cells were plated as single colonies and plasmids were extracted using the Zymoprep Yeast plasmid miniprep II kit (Zymo Research; Cat. D2004) . These plasmids were then submitted for Sanger sequencing. Additionally, clonal populations were grown individually for crosslinking and binding analysis.
[0904] Table 1. Primers for constructing Yeast library
[0905] Generation of covalent proteins (CvPr)
[0906] The protein with Cys mutation expression and cell harvesting procedure were described as above. Harvested medium free of cells was loaded onto Ni-NTA beads and extensively washed with washing buffer (30 mM imidazole, PBS, pH 7.4) . The proteins were then eluted with elution buffer (250 mM imidazole, PBS, pH 7.4) . Eluted Proteins were incubated with 2 mM DTT for 10 min at room temperature. Concentrated proteins were then subjected to SEC in PBS pH 7.4. The peak fractions obtained from SEC were incubated with crosslinkers at 1: 3 molar ratio for 1h at room temperature. The protein concentration was 20-100 μM. Crosslinkers were dissolved in DMF at a concentration of 30 mM. The final DMF percentage is 0.2%-1%. After confirming the complete coupling of crosslinkers via LC-MS analysis, the incubating mixtures were subjected to SEC in PBS. The peak fractions were aliquoted and snap frozen for future use.
[0907] Crosslinking of covalent proteins with its target in vitro
[0908] Purified KN035 (WT. ) and various covalent KN035 variants proteins were incubated with hPD-L1 (19-239) or refolding PD-L1 (19-134) in PBS buffer at 37℃ with the indicated molar ratio for different time intervals. The concentration of hPD-L1 was 1-2 μM. The purified covalent LCB3 mini-proteins were separately incubated with SARS-CoV-2 RBD at the molar ratio of 10: 1 in PBS buffer at 37℃ for established times. The concentration of RBD was 2-5 μM. After incubation, 4× LDS loading buffer and 1 mM DTT were added into the tubes and heated at 95℃ for 10 minutes. These samples were subsequently separated by 4-20 %SDS-PAGE gel and stained with Coomassie brilliant blue. The kinetics experiments were also conducted with a similar approach, where the time has been set shorter, and the molar ratio of mini binder / RBD has been set at 10: 1. The Gels were analyzed with Image J, and data were plotted using GraphPad Prism V9.5 software.
[0909] Since the initial library exhibited only 8%PD-L1 binders (FIG. 3d) , we initially conducted a binder enrichment experiment without the addition of 6-CN (FIG. 3d) . Starting from the second round of selection, DTT reduction (0.5mM, room temperature, 10 min) was performed to expose the single Cys, 6-CN (2 hours, 37℃, 0.2 mM) modification was then introduced to specifically select covalent binders (FIG. 1b, FIG. 3d) . The PD-L1 concentration and covalent crosslinking reaction time were gradually decreased from round 2 to round 4 to enrich covalent binders with superior affinity and crosslinking kinetics (FIG. 1b, FIG. 3d) . In the fourth round of selection, a reduced PD-L1 concentration of 3 nM and a 3-minute covalent crosslinking reaction were employed (FIG. 1b, FIG. 3d) . KN035 (S111H, E102F) , derived from site saturation mutagenesis, served as a control in the third and fourth round selections to identify superior covalent binders (FIG. 1b, FIG. 3d) . An acid wash (100mM glycine pH 2-3, 500mM NaCl, 0.5%tween-20) was also performed after PD-L1 binding to remove noncovalent binders in order to exclusively select covalent binders from round 2 to round 4.
[0910] The enriched library sequences underwent next-generation sequencing, followed by the selection and expression of 20 individual clones for validation. Among these clones, several exhibited significantly improved crosslinking reaction kinetics, achieving complete crosslinking with PD-L1 within just one hour, and a subset of these clones demonstrated remarkably accelerated kinetics, achieving complete covalent crosslinking within 15 minutes (FIG. 11a-b) . Consequently, we directed our attention to these highly efficient covalent binders and evaluated their PD-1 / PD-L1 blockade activity using cell assays (FIG. 11c) .
[0911] Clone H9 was found to exhibit the highest reactivity in cell-based assays (FIG. 11c) . However, we still observed undesired self-crosslinking side reactions in clone H9, potentially compromising its activity (FIG. 11d) . Structural analysis revealed Tyr111 residue as the potential site for self-crosslinking reactions. To mitigate this issue, we prepared several single-point mutants (Y111A, Y111S, Y111H, and Y111F) and found that H9-Y111A completely eliminated the self-crosslinking side reaction without compromising the PD-L1 blockade activity (FIG. 11e-f) . Consequently, we named it IB101 and prioritized it for further studies.
[0912] EXAMPLE 4. DETERMINATION OF BINDING AFFINITY OF KN035 WITH PD-L1
[0913] We measured the covalent crosslinking reaction kinetics between IB101 (10 μM) and PD-L1 (1 μM) in PBS and found the pseudo-first-order reaction rate is 0.18 min-1, with a t1 / 2 of 3.8 minutes (FIG. 4, FIG. 12) . Remarkably, decreasing the protein concentration 10 times didn’t influence the crosslinking reaction kinetics (FIG. 4, FIG. 12) , further indicating the pseudo-first-order reaction mechanism. This reaction rate is > 100 times faster than the reported GlueBody rate and is also faster than most clinically approved covalent small-molecule drugs targeting the hyper-reactive Cys. Significantly, the crosslinking reaction t1 / 2 of IB101 with PD-L1 is significantly shorter than the in vivo half-life of most miniproteins5, 6, suggesting its potential to overcome the short half-life problem inherited in miniproteins. We conducted a comparative analysis of the binding affinities of KN035, KN035_6-H, H9_111A, and IB101 (FIG. 4c) . KN035 demonstrated a high binding affinity with PD-L1 with a KD of 0.30 nM. The initial screened covalent nanobody, KN035_6-H, exhibited slightly lower affinity toward PD-L1 (KD of 0.57 nM) , primarily attributed to a slower on rate (Kon) . The optimized non-covalent nanobody, H9_111A, displayed enhanced affinity (KD of 0.13 nM) compared to KN035. Notably, the optimized covalent nanobody, IB101, demonstrated an exceptionally high affinity, with a KD value below 0.001 nM. This remarkable affinity primarily stems from an extraordinarily slow off rate (Koff) (FIG. 4c) .
[0914] We comprehensively compared PD-L1 blockade activities among KN035, IB101, engineered intermediate proteins, and two clinically approved antibodies, Envafolimab (KN035-Fc) , and Atezolizumab (FIG. 4d-e) . KN035 nanobody exhibited a PD-1 / PD-L1 blockade with an EC50 value of 59.6 ± 3.2 nM. The L108C mutation resulted in a decrease in EC50 to 462.6 ± 57.7 nM. However, the introduction of 6-H on KN035_L108C (KN035_6-H) substantially restored PD-L1 blockade activity, yielding an EC50 of 95.6 ± 5.4 nM. Notably, this activity surpassed that of the previously reported GlueBody (EC50 of 179.4 ± 9.5 nM) in our experiments. The selected non-covalent H9_111A clone displayed slightly weaker activity (EC50 of 73.9 ± 7.2 nM) compared to KN035. In contrast, IB101 demonstrated an impressive EC50 value of 5.9 ± 0.3 nM, making it ten times more potent than KN035. Remarkably, the activity of IB101 was comparable to that of the two clinically approved antibodies, KN035-Fc (EC50 of 4.7 ± 0.4 nM) and Atezolizumab (EC50 of 3.8 ± 0.9 nM) (FIG. 4e) . It is noteworthy that KN035-Fc and Atezolizumab are dimeric forms, each with two PD-L1 binding sites on a single protein. The comparable activity of IB101 with KN035-Fc and Atezolizumab underscores the remarkable potency of this covalent protein.
[0915] EXAMPLE 5. EVALUATION OF COVALENT CROSSLINKING OF IB101 WITH CELL SURFACE PD-L1
[0916] Crosslinking of IB101 with hPD-L1 on cells
[0917] MC38 / hPD-L1, H460, and U87 cells were seeded in 12-well plates at 5 × 105 per well for 12 hours. KN035 (WT. ) or IB101 was added to the culture media at the indicated concentration in a final volume of 1 mL. Following incubation at 37℃ for an established time, cells were dissociated with 0.25%trypsin-EDTA (GIBCO, Cat. 25200-056) , collected, and lysed by adding 50 μL RIPA (Beyotime, Cat. P0013C) with 1x protease inhibitor cocktail (Cell Signaling Technology, Cat. 5872) , followed by ultrasonication. Protein concentration was quantified using the BCA protein assay kit (Beyotime, Cat. P0010) . The samples were then heated at 100℃ for 10 minutes after adding 4× LDS loading buffer and DTT. Samples were analyzed by western blotting using the anti-human PD-L1 (Abcam, Cat. ab213524; 1: 1000) monoclonal antibody as the primary antibody and HRP-conjugated anti-rabbit IgG (Beyotime, Cat. A0208, 1: 3000) as the secondary antibody. β-actin was used as the internal control. The protein bands were visualized via chemiluminescence using Amersham imager 680.
[0918] we evaluated the covalent crosslinking of IB101 with cell surface PD-L1 (FIG. 4f-g) . Initial experiments with KN035 _6-H (1 μM or 3 μM) required 12 hours to achieve nearly quantitative crosslinking with U87 cell surface PD-L1 (FIG. 7c) . This rate is comparable to the reported covalent Gluebody14, which shares similar in vitro crosslinking reaction kinetics with the purified PD-L1 protein. Subsequent experiments involved incubating IB101 or KN035_6-H at lower concentrations (200 nM, 20 nM) with cancer cells for 4 hours or 45 minutes to assess crosslinking efficiency with cell surface PD-L1 (FIG. 4f) . Notably, IB101 demonstrated complete crosslinking under all conditions, while KN035_6-H (200 nM) achieved only approximately 40%crosslinking of cell surface PD-L1 after 4 hours (FIG. 4f) . Further exploration involved reducing IB101 concentrations to 10 nM, 5 nM, 1 nM (FIG. 4f) . Remarkably, even at 1 nM, IB101 achieved near quantitative crosslinking with cell surface PD-L1 within 45 minutes of coincubation, and > 70%crosslinking at 1 nM with 15 minutes coincubation (FIG. 4e-g) . This exceptional crosslinking efficiency of IB101 was consistently observed in PD-L1-positive MC38k / hPD-L1, human cancer cell lines U87 and H460 (FIG. 4f-g) . Such potent and robust crosslinking efficiency suggests that despite its short half-life, IB101 holds the potential to achieve complete PD-L1 blockade in vivo.
[0919] EXAMPLE 6. EVALUATION OF CROSSLINKING SPECIFICITY OF IB101 WITH HPD-L1 ON CELLS
[0920] Crosslinking specificity of IB101 with hPD-L1 on cells
[0921] MC38 / hPD-L1, H460, and U87 cells were seeded in 12-well plates at 5 × 105 per well for 12 hours. IB101-3 × Flag was added to the culture media at the indicated concentration in a final volume of 1 mL. Following incubation at 37℃ for 6 hours, cells were collected and lysed as described as the above section. Samples were analyzed by western blotting using the anti-Flag (Sigma, Cat. F1804, 1: 1000) antibody as the primary antibody and HRP-conjugated anti-mouse IgG (Beyotime, Cat. A0216, 1: 3000) as the secondary antibody. The samples were analyzed by western blotting using the anti-human PD-L1 as described above. β-actin was used as the internal control. The protein bands were visualized via chemiluminescence using Amersham imager 680.
[0922] We evaluated the specificity and stability of IB101 (FIG. 5f, FIG. 15) . Although IB101 only needs 15 minutes to crosslink with cell surface PD-L1 at 5 nM, we co-incubated IB101 at 50 nM or even 500 nM with different cell lines (MC38k / hPD-L1, H460, U87) for 4 hours (FIG. 5f) . Subsequent anti-flag western blot analysis showed that IB101 only crosslink with PD-L1, and nonspecific crosslinking with other cell surface proteins were not observed, indicating high specificity of IB101 (FIG. 5f) . To assess the stability of IB101, we incubated IB101 in solution at 4℃ or 25℃, or as lyophilized powder at 25℃for one month. Mass-spectrometry and PD-L1 crosslinking analysis indicated that IB101 was completely stable under these conditions, paving the way for further translational studies (FIG. 15) .
[0923] EXAMPLE 7. EVALUATION OF PHARMACOKINETICS OF IB101 FOLLOWING INTRAVENOUS AND SUBCUTANEOUS ADMINISTRATION
[0924] In vivo plasma concentration determination using sandwich enzyme-linked immunosorbent assay (ELISA)
[0925] Female C57BL / 6 mice (7 weeks old) from the Laboratory Animal Resources Center of Westlake University were used for the protein plasma half-life determination experiment. The mice were randomly divided into control and experimental groups. KN035 (WT. ) or IB101 proteins, each bearing C-terminal His tag, at an equimolar concentration of 100 μL, corresponding to 0.2 mg / mL. To assess the plasma half-life of KN035 (WT. ) and IB101 proteins, blood samples were collected by tail bleeding before injection and at various indicated time points after injection. The collected blood samples (in a heparin-coated tube) were centrifuged at 1500 g for 10 minutes to separate the blood plasma, and a protease inhibitor cocktail (Abcam, ab271306) and 5 mM EDTA were added immediately to the separated plasma, which was then snap frozen for further analysis.
[0926] Plasma concentrations of KN035 (WT. ) and IB101 were determined by a sandwich ELISA. In-house-prepared hPD-L1 (19-134) proteins with a C-terminal flag tag was used to coat a 96-well ELISA plate (Biofil, Cat. 190731080) . Following washing with Tris-buffered saline containg Tween 20 (TBST pH 7.4, 0.05%Tween 20) , each well was blocked with 5%nonfat dried milk (Beyotime, P0216) in TBST for 2 hours at room temperature. Subsequently, diluted proteins or blood plasma samples were added to each well and incubated for 2 hours at room temperature.
[0927] Following incubation, each well was treated with 100 μL anti-His HRP antibody (Proteintech, Cat. HRP-66005, 1: 3000 dilution) for 2 hours at room temperature. Between steps, the plate was washed with TBST. Finally, 100 μL TMB (Solarbio Life Sciences. Cat. PR1200) was added and incubated for 15 minutes, followed by the addition of 100 μL 2 M HCl to terminate the reaction. Optical density readings were obtained at 450 nm using a microplate reader (Varioskan LUX) , and data analysis was performed using GraphPad Prism V8.0. The Standard curve was generated using PBS diluted protein samples.
[0928] We evaluated the pharmacokinetics of IB101 following intravenous and subcutaneous administration. IB101 exhibited a brief half-life (t1 / 2) of only 8.5 minutes when administered intravenously and 32.3 minutes when administered subcutaneously (FIG. 13) . Despite this, the plasma concentration of IB101 remained above 5 nM (75 ng / mL) for at least 4 hours (FIG. 13) , allowing efficient crosslinking with cell surface PD-L1 during this timeframe. Pharmacokinetic analysis of the KN035 nanobody revealed comparable results to IB101.
[0929] EXAMPLE 8. ASSESSMENT OF CROSSLINKING EFFICIENCY OF IB101 WITH TUMOR CELL SURFACE PD-L1 IN VIVO
[0930] Crosslinking of IB101 with hPD-L1 in the tumor
[0931] MC38 / hPD-L1 cells (2x106) were resuspended with 100 μL PBS and injected subcutaneously into the flank of 6-week-old female C57BL / 6J mice (Laboratory Animal Resources Center of Westlake University) . Following a 7-day-period, tumors size reached an approximate size of ~160 mm3. Subsequently, 10 μg or 40 μg KN035 (WT. ) or IB101 was administered via injection into the peritumoral area.
[0932] After 4 hours post-administration, mice were euthanized, and tumors were harvested. Tumors were then lysed with 200 μL RIPA (Beyotime, Cat#P0013C) with 1x protease inhibitor cocktail (Cell Signaling Technology, Cat#5872) . Tissue homogenization was achieved using a homogenizer (LUKACEXUYIQI, Cat#LUKYM-I) , followed by further lysis via ultrasonication. Subsequently Western blotting analysis was conducted using the same procedures described as the above section.
[0933] Remarkably, IB101 completely crosslinked PD-L1 on the tumor cell surface at both doses (FIG. 5a-b) . This underscores the ability of IB101 to engage with tumor cell PD-L1 receptors in vivo efficiently.
[0934] Tumor xenograft model study
[0935] A total of 50,000 MC38 / hPD-L1 cells were injected into the flank of the 6-8-week-old B6-hPD1 / hPDL1 mice to induce solid tumor growth. Upon reaching a tumor of approximately 100 mm3, 2 mg / kg KN035-Fc, or IB101 were administrated subcutaneously into the xenograft every 4 days for 3 doses. Tumor growth was monitored by measuring two dimensions using calipers, and tumor volume was calculated using the formula: tumor volume = length × width2 / 2. On day 45, the mice were euthanized.
[0936] the MC38k / hPD-L1 xenograft tumor was engrafted into hPD1 / hPD-L1 transgenic mice to assess the in vivo efficacy of IB101 (FIG. 5a, c-e) . A head-to-head efficacy comparison study was conducted using the clinically approved KN035-Fc (envofolimab) . KN035-Fc, the first subcutaneously dosed drug targeting PD-L1, was administered at a dose of 2 mg / kg every 4 days, three times when the tumor size reached approximately 100 mm3 (FIG. 5a) . Despite IB101 having a significantly shorter plasma half-life (30 minutes) compared to KN035-Fc (2 weeks) 25, IB101 was administered following the same schedule. Additionally, another group received IB101 with a more frequent dosing regime (2 mg / kg every 2 days, five times) . The results revealed that KN035-Fc generally suppressed tumor growth, leading to the complete elimination of tumors in 60%of mice (FIG. 5c-d) . In comparison, IB101 treatment demonstrated consistent tumor regression, with complete tumor elimination achieved after three drug administrations (FIG. 5c-d) . Notably, tumors were eliminated even faster when IB101 was administered more frequently (FIG. 5c-d) . Body weight monitoring indicated that IB101 was well tolerated by mice, with no observable toxicity (FIG. 5e) . The superior tumor suppression activity exhibited by IB101, despite its short in vivo half-life compared to KN035-Fc, underscores the potent efficacy and clinical potential of supercharged covalent mini-proteins.
[0937] Protein labeled with Cyanine 7 (Cy7)
[0938] Recombinant SrtA was expressed and purified as described in the literature. The -LPETGS-sequence was fused to KN035, KN035-Fc, and IB101 C-terminal. Conjugation reactions between proteins and home-made substrate peptides GGGK (N3) were conducted by incubating proteins (1 equivalent (eq. ) ) , substrate peptides (50 eq. ) , and Sortase (0.05 eq. ) in reaction buffer (300 mM Tris-HCl, pH 7.4, 150 mM NaCl, 5 mM CaCl2) for 2 hours at room temperature. The reacting products were purified by removing unreacted proteins and Sortase using Ni-NTA beads. The purity and molecular weight of the product protein were confirmed by LC-MS.
[0939] Subsequently, Sulfo-Cy7-DBCO (Xi'an Qiyue Biology, CAS#Q-0275774) in ddH2O (5 eq. ) was added to N3 modified KN035, KN035-Fc, and IB101 in PBS. The mixture was incubated at room temperature for 2 hours, and the progress of the reaction was monitored using LC-MS. Upon completion of the reaction, mixtures were concentrated and subjected to SEC (Superdex 75 Increase 10 / 300 GL, Superdex 75 Increase 10 / 300 GL, GE Healthcare) in PBS, pH 7.4 separately. The peak fractions were aliquoted and snap frozen for future use.
[0940] We labeled IB101, KN035, and KN035-Fc with Cy7 to assess their in vivo distribution and tumor residence time in mice bearing tumors (FIG. 14a-e) . Following subcutaneous administration of IB101-Cy7, rapid accumulation at the tumor site occurred within 10 minutes (FIG. 14d-e) . In contrast, KN035-Fc took considerably longer to accumulate at the tumor site (FIG. 14d-e) . This swift tumor enrichment dynamic suggests that IB101 can extravasate into tissues and enter the bloodstream, penetrating the tumor more rapidly due to its compact size. The fluorescence signal of IB101-Cy7 at the tumor site peaked around 4 hours post-administration, gradually fading over time but still detectable even after 4 days (FIG. 14d-e) . In comparison, KN035-Fc-Cy7 exhibited a gradual increase in fluorescence signal over several days, surpassing the intensity of IB101-Cy7 in later stages (FIG. 14d-e) . This may be attributed to the prolonged lifespan of KN035-Fc-Cy7 and the continuous accumulation at the tumor site. KN035-Cy7 displayed similar tumor enrichment dynamics to IB101-Cy7, but its fluorescence signal intensity was notably lower overall, indicating less enrichment. These findings highlight the superior and sustained tumor residence of IB101-Cy7, affirming its rapid tissue penetration and potential advantages in therapeutic applications.
[0941] EXAMPLE 9. EVALUATION OF HUMANIZED IB101 (IB101_R AND IB101_Q)
[0942] Construction, expression, and purification of proteins
[0943] The genes encoding His-TEV-IB101_R (SEQ ID NO. 44) and His-TEV-IB101_Q (SEQ ID NO. 45) were subcloned into pcDNA3.4 vector. The genes encoding Atezolizumab heavy chain and light chain were synthesized (Genewiz, Suzhou) and cloned into pcDNA3.4 vector.
[0944] His-TEV-IB101_R and His-TEV-IB101_Q were expressed and purified from Expi293F cells using transient transfection. Transient transfections were conducted when cell density reached around 2 ×106 / mL. A mixture of 0.2 mg of plasmid and 0.4 mg of polyethyleneimine (Cat. 24765-1) in 20 mL of fresh medium was incubated for 30 minutes and then added to a 100-mL cell culture. The medium containing secreted proteins was harvested 96 hours post-transfection. Proteins were captured from the cell supernatant via Ni-NTA chelating resin. Eluted proteins were incubated with 2 mM DTT for 10 min at room temperature. Concentrated proteins were then subjected to SEC (Superdex 200 Increase 10 / 300 GL, Cytiva) in PBS pH 7.4.
[0945] For purification of Atezolizumab: the harvested medium free of cells was loaded onto ProteinA beads (GenScript, Cat. L00210-50) and washed with PBS. The proteins were then eluted with 0.1 M glycine, pH 3.0, and then neutralized with 0.2 M NaHCO3, pH 8.4 immediately. Eluted Proteins were concentrated and then subjected to SEC (Superdex 200 Increase 10 / 300 GL, Cytiva) in PBS. The peak fractions were aliquoted and snap frozen for future use.
[0946] Coupling of 6-CN
[0947] The peak fractions obtained from SEC were incubated with crosslinker 6-CN at a 1: 3 molar ratio for 1 h at room temperature. The protein concentration was 20-100 μM. 6-CN was dissolved in DMF at a concentration of 30 mM. The final DMF percentage is 0.2%-1%. After confirming the complete coupling of 6-CN via LC-MS analysis, the incubating mixtures were subjected to SEC in 50 mM Tris-HCl, pH 7.4, 1 mM DTT, and 100 mM NaCl. The peak fractions were collected for future use.
[0948] Purification of IB101_R no tag and IB101_Q no tag
[0949] Purified His-TEV-IB101_R and His-TEV-IB101_Q were incubated separately with homemade TEV enzyme at 4℃ overnight. Following cleavage, the mixture containing the cleaved proteins (IB101_R no tag (SEQ ID NO. 46) and IB101_Q no tag (SEQ ID NO. 47) ) and the His-tagged TEV enzyme was incubated with Ni-NTA resin. This step allowed the small peptide and TEV enzyme to bind to the resin, while IB101_R no tag and IB101_Q no tag remained in the flow-through. The flow-through was then subjected to SEC in PBS (pH 7.4) . Peak fractions were collected, aliquoted, and snap-frozen for future use.
[0950] It will be understood by those skilled in the art that in the sequences designated "IB101_R no tag" and "IB101_Q no tag" , the initial Glycine (G) is the N-terminal residue remaining after the His-tag has been removed by enzymatic cleavage. The sequences for "IB101_R" and "IB101_Q" are set forth as SEQ ID NO. 37 and SEQ ID NO. 38, respectively.
[0951] Crosslinking of IB101_R no tag and IB101_Q no tag with PD-L1 in vitro
[0952] Purified IB101_R and IB101_Q were incubated with PD-L1 (19-239) in PBS buffer at 37℃ with the indicated molar ratio for different time intervals. The concentration of PD-L1 was 1 μM.
[0953] T cell activity restored by PD-L1 blockade in vitro
[0954] PD-1 / NFAT-luciferase / Jurkat cells were cultured in RPMI 1640 medium supplemented with 10 %FBS, 1 %Penn-Strep, 500 μg / mL neomycin and 250 μg / mL hygromycin B. hPD-L1 / aAPC / HEK cells were cultured in DMEM medium with 10 %FBS, 1 %Penn-Strep, 1 μg / mL puromycin and 500 μg / ml neomycin. For the in vitro efficacy assay, hPD-L1 aAPC / HEK cells were seeded at a density of 50,000 cells per well into a white 96-well microplate (Cellvis, Cat. 062096 ) in 100 μL of growth medium. After 12 hours, half of the medium was removed from the aAPC / HEK cells, and the cells were incubated with 50 μL fresh growth medium supplemented with either PBS or other protein solutions for 15 min. After treatment, half of the medium was removed, followed by adding 100,000 PD-1 NFAT-luciferase / Jurkat cells in 50 μL medium (RPMI 1640, 10 %FBS, 1 %Pen-Strep) . Following 6 hours of co-culture, cells were lysed, and luciferase assay was performed using Luciferase Reporter Gene Assay Kit (Yeasen, Cat#11401ES76) . Luminescence was measured using a microplate reader (Varioskan LUX) .
[0955] B16F10 tumor xenograft model study
[0956] A total of 50,000 B16F10 / hPD-L1 cells were injected into the flank of the 8-week-old B6-hPD1 / hPDL1 female mice to induce solid tumor growth. Upon reaching a tumor of approximately 50 mm3, Atezolizumab and IB101 were administered intravenously. Tumor growth was monitored by measuring two dimensions using calipers, and tumor volume was calculated using the formula: tumor volume = length × width2 / 2. On day 45, the mice were euthanized.
[0957] The B16F10-derived tumor model is characterized by an immune-suppressive microenvironment and only shows limited response to PD-1 / PD-L1 antibody treatment. To investigate whether IB101 could exhibit therapeutic activity against B16F10-derived tumor growth, we conducted an efficacy study comparing IB101 with the FDA-approved Atezolizumab (FIG. 18a) . The results demonstrated that Atezolizumab showed no significant anti-tumor efficacy compared to the PBS group, whereas IB101 treatment consistently suppressed tumor growth in all treated mice (FIG. 18b) . Body weight monitoring indicated that IB101 was well tolerated by mice (FIG. 18c) .
[0958] Evaluation of humanized IB101 (IB101_R and IB101_Q)
[0959] Purified IB101_R no tag and IB101_Q no tag were incubated with PD-L1 (19-239) in PBS buffer at 37℃ with the indicated molar ratio for different time intervals. The concentration of PD-L1 was 1 μM. IB101 (wide type, WT. ) was used as a control.
[0960] We evaluated the crosslinking efficiency of humanized IB101. The crosslinking efficiency of humanized IB101 showed no significant difference from that of IB101 (FIG. 19a) . Then their PD-1 / PD-L1 blockade activity was evaluated using cell assay. IB101_R exhibited an EC50 value of 8.24 ± 0.46 nM, and IB101_Q exhibited an EC50 value of 5.81 ± 0.37 nM, which is similar to the EC50 value of IB101 (7.77 ± 0.61 nM) (FIG. 19b) . Subsequently, the melting temperature (TM) was measured. KN035 exhibited the highest TM (66.8℃) . The TM of IB101 was slightly lower than KN035 (62.1℃) and IB101_R exhibited a much lower TM (60.8℃) , while IB101_Q showed a higher TM than IB101 (64.1℃) (FIG. 19c) .
[0961] Stability of IB101_R
[0962] To evaluate the stability of humanized IB101, we incubated IB101_R in a PBS solution at pH 5.0 at 37℃ for one month or 4.5 months. Mass spectrometry and PD-L1 crosslinking analysis confirmed that IB101_R remained stable after long-term storage (FIG. 20) .
[0963] EXAMPLE 10. DESIGN AND EVALUATION OF COVALENT LCB3 VARIANT (F30C-MU2_5-NO2)
[0964] Pseudotyped SARS-CoV-2 virus inhibition
[0965] The packing of pseudotyped SARS-CoV-2 virus was conducted following the described methods as reported. Pseudotyped SARS-CoV-2 inhibition assays were performed to assess the efficacy of covalent LCB3 mini-proteins. In detail, Caco-2 cells were seeded in 96-well cell culture plates at a density of 1×104 per well and incubated for 24 hours. Covalent LCB3 and wild-type LCB3 were diluted with FBS-free DMEM and mixed with pseudotyped viruses (1: 1, v / v) . The mixture was then incubated at 37℃ for 2 hours and added to the Caco-2 cells. After 12 hours of infection, the culture medium was replaced with fresh DMEM containing 10%FBS, and the cells were further incubated for 36 hours. Subsequently, the cells were lysed with Cell Lysis Buffer (Promega, Madison, WI, USA) , and luciferase activity was detected using a Luciferase Assay System (Promega, Madison, WI, USA) .
[0966] For competitive inhibition assay, pseudotyped SARS-CoV-2 virus was initially co-incubated with either covalent LCB3 or wild-type LCB3 at 37℃ for 2 hours. Then, 2 μg / mL RBD protein was added to the resulting mixture and incubated for an additional hour. The mixture was subsequently introduced to the Caco-2 cell line, and the subsequent steps of assay were carried out as described above. All experimental data were analyzed using GraphPad Prism V9.5 software.
[0967] LCB3 (SEQ ID NO. 55) is a de novo-designed miniprotein targeting the receptor-binding domain (RBD) to inhibit SARS-CoV-226. It has previously undergone engineering into a covalent miniprotein, known as Gluebinder, using genetic code expansion technology, resulting in enhanced activity15. However, the previously engineered covalent LCB3, Gluebinder, exhibited slow crosslinking kinetics, requiring 5 hours to achieve 85%crosslinking15. Our specific aim was to employ our established system to engineer a supercharged covalent LCB3 with extremely fast crosslinking kinetics.
[0968] We analyzed the complex structure of LCB3 / RBD and identified D11, K26, F30, and Y40 as potential single Cys mutation sites for integrating small molecule crosslinkers (FIG. 16a) 26. The amino acid sequences of LCB3_D11C, LCB3_K26C, LCB3_F30C and LCB3_Y40C are as set forth in SEQ ID NOs. 56 to 59, respectively. Initial screening with 8 crosslinkers revealed that the F30C mutant, in combination with crosslinker 3-H, achieved complete crosslinking with RBD within 5 hours (better than Gluebinder) (FIG. 6a) . Subsequently, a second-round screening with additional crosslinkers demonstrated even faster crosslinking kinetics. Notably, F30C modification with 1-CN achieved near-quantitative crosslinking with RBD within just 30 minutes of coincubation (FIG. 6b) . However, it was observed that 2 equivalents of these covalent LCB3 proteins failed to attain similar crosslinking efficiency as 5 or more equivalents. This discrepancy was attributed to significant self-crosslinking side reactions within the covalent protein, compromising their ability to effectively crosslink with the RBD protein (FIG. 16b) . To address this issue, mutations were introduced into covalent LCB3-F30C protein, successfully minimizing the self-crosslinking side reactions (FIG. 16c) 26. The amino acid sequences of LCB3_F30C-Mu1 and LCB3_F30C-Mu2 are as set forth in SEQ ID NOs. 60 and 61, respectively. Upon conducting the crosslinker screening again, it was found that F30C-Mu2_5-NO2 protein exhibited highest crosslinking rate, achieving near-quantitative crosslinking with RBD within just 5 minutes (FIG. 6c, FIG. 16c-f) , the crosslinking kinetics study indicates a t1 / 2 of 0.8 minutes, with Kinact of 0.86 min-1 (FIG. 6c, FIG. 16f) . This remarkable acceleration, approximately 100 times faster than the previously reported covalent Gluebinder15, was attributed to the successful mitigation of self-crosslinking side reactions while retaining the crosslinking efficienty. Given that we had already obtained a supercharged covalent LCB3 protein through single Cys mutants and a diverse crosslinker scan, we opted not to use high-throughput yeast display selection for further optimizing the covalent LCB3.
[0969] LCB3 stands out as a meticulously designed and fully optimized receptor-binding domain (RBD) binder for SARS-CoV-2, and any point mutation on LCB3 is likely to compromise its interaction with RBD according to its single-site saturation mutagenesis scan, subsequently diminishing its inhibitory activity26. F30C-Mu2 (SEQ ID NO. 61) mutation was identified as particularly impactful, leading to a nearly 1000-fold decrease in inhibitory activity (FIG. 6d) . However, through the strategic installation of small molecule crosslinkers, the covalent LCB3 variant (F30C-Mu2_5-NO2) successfully restored the inhibitory activity, which is three times higher than LCB3 (FIG. 6d) .
[0970] To further characterize the irreversible inhibition nature of covalent LCB3 protein, a competitive inhibition assay was performed. Pseudoviruses were initially co-incubated with either covalent LCB3 or wild-type LCB3, followed by the addition of the RBD protein for another 1-hour incubation to compete with the covalent LCB3 or wild-type LCB3. The resulting mixture was subsequently introduced to the ACE2-expressing cells. Notably, the inhibitory activity of wild-type LCB3 was dramatically reduced to 1.75 nM. In contrast, covalent LCB3 retained most of their activity, with an IC50 value of 55 pM (FIG. 6e) . This emphasizes the irreversible nature and advantageous characteristics of covalent LCB3 in target inhibition.
[0971] Technical Advances Achieved by the Present Disclosure
[0972] In recent years, miniproteins have garnered increasing attention for therapeutic applications due to their small size and stable structure, which offer numerous advantages over traditional monoclonal antibodies. However, their intrinsic small size also limits their in vivo circulation lifetime, diminishing their potential as standalone drugs. Inspired by the success of covalent small-molecule drugs, covalent proteins are now being explored as new therapeutic modalities. Nonetheless, currently reported covalent proteins typically require significantly longer time (several hours) than their in vivo circulation half-life to crosslink with target proteins, hindering efficient covalent engagement with target proteins. Moreover, while high-throughput selection methods have been instrumental in the success of many protein therapeutics, a robust and efficient covalent protein selection system for developing supercharged covalent proteins remains elusive. To address this gap, we have combined chemoselective protein modification and yeast display selection, establishing a robust, high-throughput platform for developing supercharged covalent proteins possessing an extremely accelerated crosslinking rate. Utilizing this selection system, we successfully developed a supercharged PD-L1 targeting nanobody IB101 with a target crosslinking rate of 0.2 minutes-1 and a half-life (t1 / 2) of 3 minutes, demonstrated both with purified PD-L1 and cell surface PD-L1. IB101 exhibited superior tumor suppression activity compared to the clinically approved KN035-Fc antibody, despite its significantly shorter in vivo half-life. To showcase the broad applicability of this selection system, we also developed a supercharged covalent miniprotein targeting the RBD, with an extraordinary crosslinking rate of 0.86 minutes-1 (Kinact of 1.4 × 10-2 s-1) and a t1 / 2 of 0.8 minutes.
[0973] Covalent small-molecule drugs typically use mildly reactive warheads to target cysteine residues situated near enzyme-substrate binding pockets due to cysteine's heightened reactivity compared to other amino acids. However, achieving target specificity and increased inhibitory activity often demands substantial medicinal chemistry efforts. Beyond the realm of small molecules, there have been some reports employing high-throughput selection methods, including phage display or mRNA display, to screen covalent peptides. However, these covalent peptides primarily target different enzymes owing to the accessibility of enzyme-substrate binding pockets, and their in vivo efficacy remains elusive27-29. For broader modulations of protein-protein interaction interfaces, larger proteins with well-defined three-dimensional structures remain the optimal choice. This underscores the significant therapeutic potential of covalent proteins. However, doubts persist regarding their potential due to the lack of high-throughput selection methods, the sluggish rate of target crosslinking, and their brief in vivo half-life. Our established high throughput selection system and the developed supercharged covalent proteins precisely addressed these challenges, poised to unleash the full therapeutic potential of covalent proteins.
[0974] Unlike the warheads utilized in covalent small molecule drugs, those chosen for covalent proteins generally exhibit significantly weaker reactivity. For instance, the fluorosulfate-based crosslinker demonstrated no observable reactivity with peptides containing His, Lys, or Tyr residues after incubation at 37℃ with a concentration of 1 mM for 16 hours (FIG. 17) . This intrinsic weak reactivity is essential because the warhead must remain stable in the presence of proteins and all natural amino acids, even during long-term storage. Consequently, the subdued reactivity of these warheads is a primary reason for the slow reported crosslinking rates of current covalent proteins. However, chemical reactivity is not solely determined by intrinsic activity; it can also be profoundly influenced by the surrounding chemical environment and the spatial arrangement (conformation) of the two reacting groups. Our high-throughput selection system excels in efficiently exploring all these influencing factors to identify optimal combinations for generating supercharged covalent proteins using these inherently weak chemical warheads.
[0975] The warheads on covalent proteins were designed to target natural amino acids, which increases the likelihood of self-crosslinking side reactions with nearby residues. These unintended reactions can significantly reduce the activity of the developed covalent protein and necessitate careful engineering to mitigate. In our experience, such self-crosslinking side reactions are frequently encountered. To address this issue, we employ structure-based analysis and mass spectrometry (MS / MS) analysis to identify the residues responsible for self-crosslinking and subsequently mutate them into nonreactive residues. Importantly, our high-throughput covalent protein selection system proves invaluable in eliminating these self-crosslinking clones, as it exclusively selects covalent proteins within our system, effectively excluding those prone to self-crosslinking side reactions.
[0976] Our yeast display selection system relies on the precise Cys chemical modification of the displayed protein, which initially posed challenges due to the presence of Aga1p-Aga2p interchain disulfides. However, we successfully identified conditions that allow maximal modification of the displayed protein's Cys residues without disrupting the critical Aga1p-Aga2p interchain disulfides. Moreover, while chemical modification could potentially occur on other yeast surface proteins, the covalent crosslinking relies on specific protein-protein binding. Therefore, modifications beyond the displayed protein's Cys residues do not crosslink with the target protein, thus ensuring the integrity of the selection system. In the future, alternative yeast display systems utilizing single-chain anchoring proteins beyond Aga1p-Aga2p may offer improved compatibility if interchain disulfides prove problematic for covalent protein selection30.
[0977] With the establishment of this high-throughput covalent protein selection system for supercharged covalent proteins, we believe that a major technical hurdle in the development of covalent protein therapeutics has been overcome. This advancement is poised to significantly propel the discovery of covalent protein therapeutics.
[0978] In view of the many possible embodiments to which the principles of our invention may be applied, it should be recognized that illustrated embodiments are only examples of the invention and should not be considered a limitation on the scope of the invention. Rather, the scope of the invention is defined by the following claims. We therefore claim as our invention all that comes within the scope and spirit of these claims.
[0979] References
[0980] 1. Maute, R.L. et al. Engineering high-affinity PD-1 variants for optimized immunotherapy and immuno-PET imaging. Proc Natl Acad Sci U S A 112, E6506-6514 (2015) .
[0981] 2. Yang, E.Y. & Shah, K. Nanobodies: Next Generation of Cancer Diagnostics and Therapeutics. Front Oncol 10 (2020) .
[0982] 3. Beck, A., Wurch, T., Bailly, C. & Corvaia, N. Strategies and challenges for the next generation of therapeutic antibodies. Nat Rev Immunol 10, 345-352 (2010) .
[0983] 4. Cruz, E. & Kayser, V. Monoclonal antibody therapy of solid tumors: clinical limitations and novel strategies to enhance treatment efficacy. Biologics 13, 33-51 (2019) .
[0984] 5. Crook, Z.R., Nairn, N.W. & Olson, J.M. Miniproteins as a Powerful Modality in Drug Development. Trends in Biochemical Sciences 45, 332-346 (2020) .
[0985] 6. Kontermann, R.E. Strategies for extended serum half-life of protein therapeutics. Curr Opin Biotechnol 22, 868-876 (2011) .
[0986] 7. Ebrahimi, S.B. & Samanta, D. Engineering protein-based therapeutics through structural and chemical design. Nat Commun 14, 2411 (2023) .
[0987] 8. Boike, L., Henning, N.J. & Nomura, D.K. Advances in covalent drug discovery. Nat Rev Drug Discov 21, 881-898 (2022) .
[0988] 9. Abdeldayem, A., Raouf, Y.S., Constantinescu, S.N., Moriggl, R. & Gunning, P.T. Advances in covalent kinase inhibitors. Chem Soc Rev 49, 2617-2687 (2020) .
[0989] 10. Mons, E., Roet, S., Kim, R.Q. & Mulder, M.P.C. A Comprehensive Guide for Assessing Covalent Inhibition in Enzymatic Assays Illustrated with Kinetic Simulations. Curr Protoc 2, e419 (2022) .
[0990] 11. Bauer, R.A. Covalent inhibitors in drug discovery: from accidental discoveries to avoided liabilities and designed therapies. Drug Discov Today 20, 1061-1073 (2015) .
[0991] 12. Awtry, E.H. & Loscalzo, J. Aspirin. Circulation 101, 1206-1218 (2000) .
[0992] 13. Li, Q. et al. Developing Covalent Protein Drugs via Proximity-Enabled Reactive Therapeutics. Cell 182, 85-97 e16 (2020) .
[0993] 14. Zhang, H. et al. Covalently Engineered Nanobody Chimeras for Targeted Membrane Protein Degradation. J Am Chem Soc 143, 16377-16382 (2021) .
[0994] 15. Han, Y. et al. Covalently Engineered Protein Minibinders with Enhanced Neutralization Efficacy against Escaping SARS-CoV-2 Variants. J Am Chem Soc (2022) .
[0995] 16. Wang, N. & Wang, L. Genetically encoding latent bioreactive amino acids and the development of covalent protein drugs. Curr Opin Chem Biol 66, 102106 (2022) .
[0996] 17. Yu, B. et al. Accelerating PERx reaction enables covalent nanobodies for potent neutralization of SARS-CoV-2 and variants. Chem 8, 2766-2783 (2022) .
[0997] 18. Cheng, L., Wang, Y., Guo, Y., Zhang, S.S. & Xiao, H. Advancing protein therapeutics through proximity-induced chemistry. Cell Chem Biol 31, 428-445 (2024) .
[0998] 19. Yu, B., Cao, L., Li, S., Klauser, P.C. & Wang, L. The proximity-enabled sulfur fluoride exchange reaction in the protein context. Chem Sci 14, 7913-7921 (2023) .
[0999] 20. Aqvist, J., Kazemi, M., Isaksen, G.V. & Brandsdal, B.O. Entropy and Enzyme Catalysis. Acc Chem Res 50, 199-207 (2017) .
[1000] 21. Kazemi, M., Himo, F. & Aqvist, J. Enzyme catalysis by entropy without Circe effect. Proc Natl Acad Sci U S A 113, 2406-2411 (2016) .
[1001] 22. Boder, E.T. & Wittrup, K.D. Yeast surface display for screening combinatorial polypeptide libraries. Nat Biotechnol 15, 553-557 (1997) .
[1002] 23. Cherf, G.M. & Cochran, J.R. Applications of Yeast Surface Display for Protein Engineering. Methods Mol Biol 1319, 155-175 (2015) .
[1003] 24. Zhang, F. et al. Structural basis of a novel PD-L1 nanobody for immune checkpoint blockade. Cell Discov 3, 17004 (2017) .
[1004] 25. Papadopoulos, K.P. et al. First-in-Human Phase I Study of Envafolimab, a Novel Subcutaneous Single-Domain Anti-PD-L1 Antibody, in Patients with Advanced Solid Tumors. Oncologist 26, e1514-e1525 (2021) .
[1005] 26. Cao, L. et al. De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science 370, 426-431 (2020) .
Claims
1.A covalent protein having a structure as shown in Formula (I) : whereinAb represents a protein moiety, and R1 is covalently bonded to Ab, preferably through side chain of Cys residue;R1 has a structure of -R11-, or -R11-C (O) -R12-R13-, whereR11 is optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN,R12 is selected from the group consisting of NR12a, O, S, and heterocyclylene, R12a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;R13 is absent or optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN;RA is selected from the group consisting of optionally substituted alkanediyl, and optionally substituted arenediyl;Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, alkyl, haloalkyl, and aryl,n is an integer selected from 0 and 1;Ry is selected from the group consisting of S and P;Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, alkyl, haloalkyl, and aryl;The bondbetween Ry and Rz is a single bond or a double bond.2.The covalent protein according to claim 1, wherein R1 is selected from C1-C4 alkanediyl and C1-C4 alkanediyl-C (O) -NH-, preferably R1 is selected from the group consisting of -CH2-, -CH2CH2-, -CH2CH2CH2-, -CH2CH2CH2CH2-, -CH2-C (O) -NH-, -CH2CH2-C (O) -NH-, -CH2CH2CH2-C (O) -NH-, -CH2CH2CH2CH2-C (O) -NH-, more preferably -CH2-, -CH2CH2-, -CH2-C (O) -NH-, and -CH2CH2-C (O) -NH-, even more preferably -CH2-and -CH2-C (O) -NH-.3.The covalent protein according to claim 1 or 2, wherein RA is selected from the group consisting of optionally substituted C1-C8, preferably C1-C6 alkanediyl, and optionally substituted C6-C14, preferably C6-C10, more preferably C6 arenediyl,preferably, RA is C1-C8, preferably C1-C6 alkanediyl optionally substituted with one or more halo, -OH, or -CN, orRA is C6-C14, preferably C6-C10, more preferably C6 arenediyl optionally substituted with one or more Rx, where Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, N (Rx1) (Rx2) , where each of Rx1 and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl, preferably Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.4.The covalent protein according to any one of claims 1 to 3, wherein the covalent protein has a structure as shown in Formula (I-A) : wherein,p is an integer selected from 1 to 8, preferably 1 to 6, - (CH2) p-optionally substituted with one or more halo, -OH, or -CN.5.The covalent protein according to any one of claims 1 to 4, wherein the covalent protein has a structure as shown in any of Formula (I-A1) - (I-A3) : wherein,p is an integer selected from 1 to 8, preferably 1 to 6.6.The covalent protein according to any one of claims 1 to 3, wherein the covalent protein has a structure as shown in Formula (I-B) : wherein m is an integer selected from 1 to 4.7.The covalent protein according to any one of claims 1 to 3, and 6, wherein the covalent protein has a structure as shown in any of Formula (I-B1) - (I-B4) : wherein m is an integer selected from 1 to 4.8.The covalent protein according to any one of claims 1 to 3, and 6 to 7, wherein the covalent protein has a structure as shown in any of Formula (I-B1-a) - (I-B1-c) , Formula (I-B2-a) - (I-B2-c) , Formula (I-B3-a) - (I-B3-c) , and Formula (I-B4-a) - (I-B4-c) : wherein m is an integer selected from 1 to 4.9.The covalent protein according to any one of claims 1 to 3, and 6 to 8, wherein the covalent protein is selected from the group consisting of 10.A covalent protein having a structure as shown in Formula (I-B1) : wherein Ab represents a protein moiety, and R1 is covalently bonded to Ab;R1 is a linker moiety having a structure of -R11-, or -R11-C (O) -R12-R13-, whereR11 is optionally substituted alkylene, which is optionally substituted with one or more halo, -OH, or -CN,R12 is selected from the group consisting of NR12a, O, S, and heterocyclylene, R12a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;R13 is absent or optionally substituted alkylene, which is optionally substituted with one or more halo, -OH, or -CN;Rx is selected from the group consisting of -H, -halo, -CN, -NO2 and other chemical groups with similar functions, m is an integer selected from 1 to 4;n is an integer selected from 0 and 1.11.The covalent protein according to claim 10, wherein R1 is covalently bonded to Ab through side chain of Cys residue.12.The covalent protein according to claim 10 or 11, wherein R12 is selected from the group consisting of NR12a, O, S, and 5-or 6-membered heterocyclylene, for example, 5-or 6-membered heterocyclylene is selected from R12a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl.13.The covalent protein according to any one of claims 10 to 12, wherein R1 is -CH2-, or -CH2-C (O) -NH-.14.The covalent protein according to any one of claims 10 to 13, wherein each Rx is independently selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl; preferably each Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.15.The covalent protein according to any one of claims 10 to 14, wherein the covalent protein has a structure as shown in Formula (I-B1-a11) - (I-B1-c1) , or Formula (I-B1-a2) - (I-B1-c2) : wherein, Ab, Rx, m and n are as provided in any one of claims 10 to 13.16.The covalent protein according to any one of claims 10 to 15, wherein the covalent protein having a structure as shown in Formula (I-B1) is selected from the group consisting of wherein, Ab, Rx, and n are as provided in any one of claims 10 to 14.17.The covalent protein according to any one of claims 10 to 16, wherein the covalent protein having a structure as shown in Formula (I-B1) is selected from the group consisting of wherein, Ab, and Rx are as provided in any one of claims 10 to 16.18.The covalent protein according to any one of claims 1 to 17, wherein the protein moiety Ab comprises at least one naturally occurring or engineered cysteine (Cys) residue in its amino acid sequence, preferably the cysteine residue is spatially located at or near binding interface region between the protein moiety Ab and a target protein.19.The covalent protein according to any one of claims 1 to 18, wherein the protein moiety Ab includes but is not limited to the group consisting of an antibody or antigen-binding fragment thereof; a non-antibody scaffold protein; a cytokine, a growth factor, a hormone, or a variant and functional fragment thereof; a receptor protein or ligand-binding domain thereof; a ligand protein or receptor-binding domain thereof; an enzymes or a modulator thereof; a peptide with specific binding activity, and the like.20.The covalent protein according to any one of claims 1 to 19, wherein the protein moiety Ab is an antibody or antigen-binding fragment thereof, including, but not limited to IgG, IgM, IgA, IgD, IgE, Fab, Fab', F (ab') 2, Fv, single-chain antibody (scFv) , diabody, triabody, minibody, nanobody, domain antibody, or modified antibody including chimeric antibody, humanized antibody, fully human antibody, bispecific antibody, multispecific antibody.21.The covalent protein according to any one of claims 1 to 20, wherein the protein moiety Ab is an antibody targeting immune checkpoint, or a protein targeting a viral surface protein or antigen.22.The covalent protein according to any one of claims 1 to 21, wherein the protein moiety Ab is a PD-L1 targeting nanobody, or a SAR-CoV-2 RBD targeting miniprotein.23.The covalent protein according to any one of claims 1 to 22, wherein the protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NOs. 3, 4 to 16, and 17 to 38; orthe protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NOs. 55, and 56 to 61.24.The covalent protein according to any one of claims 1 to 23, wherein the protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 3, 4 to 16, and 17 to 38; orthe protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61.25.The covalent protein according to any one of claims 1 to 24, wherein the covalent protein is selected from the group consisting of where,Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C4 haloalkyl, C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C4 alkyl, and C1-C4 haloalkyl; preferably Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2; more preferably, Rx is selected from the group consisting of H, F, CN, and NO2; andprotein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 3, 4 to 16, and 17 to 38; or a SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61.26.The covalent protein according to any one of claims 1 to 25, wherein the covalent protein is selected from the group consisting of whereRx is selected from the group consisting of H, CN, and NO2, preferably Rx is H, and CN, andthe protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 8, 11 to 16; and SEQ ID NOs. 17 to 38, preferably any of SEQ ID NO. 8 and 11, and SEQ ID NOs. 17, and 37 to 38,preferably, the covalent protein is selected from the group consisting ofRx is H or CN, the protein moiety Ab is a PD-L1 targeting nanobody comprising an amino acid sequence as set forth in any of SEQ ID NOs. 8, 11, 17, 37, and 38.27.The covalent protein according to any one of claims 1 to 25, wherein the covalent protein is selected from the group consisting of whereRx is selected from the group consisting of H, F, CN, and NO2, andthe protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in any of SEQ ID NOs. 55, and 56 to 61;preferably, the covalent protein is selected from the group consisting ofRx is CN, and the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in SEQ ID NO. 58;Rx is H, and the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in SEQ ID NO. 58; andRx is NO2, and the protein moiety Ab is a SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence as set forth in SEQ ID NO. 61.28.A pharmaceutical composition comprising a therapeutically effective amount of the covalent protein according to any of claims 1 to 27, and a pharmaceutically acceptable carrier, excipient, or diluent.29.The pharmaceutical composition according to claim 28, wherein said composition is suitable for parenteral administration, oral administration, administration by inhalation, topical administration, or transdermal administration.30.Use of the covalent protein according to any one of claims 1 to 27 or the pharmaceutical composition according to claim 28 or 29 in the manufacture of a medicament for diagnosing, treating or preventing a disease, disorder, or condition in a subject in need thereof.31.The use according to claim 30, wherein the disease, disorder, or condition is a disease associated with dysregulated PD-L1 expression or activity including cancer and autoimmune disease, or a disease caused by a coronavirus infection including COVID-19.32.The covalent protein according to any one of claims 1 to 27 or the pharmaceutical composition according to claim 28 or 29 for use in diagnosing, treating or preventing a disease, disorder, or condition in a subject in need thereof.33.The covalent protein or pharmaceutical composition for use according to claim 32, wherein the disease, disorder, or condition is a disease associated with dysregulated PD-L1 expression or activity including cancer and autoimmune disease, or a disease caused by a coronavirus infection including COVID-19.34.A method for diagnosing, treating or preventing a disease, disorder, or condition in a subject in need thereof, comprising administering to the subject an effective amount of the covalent protein according to any one of claims 1 to 27 or the pharmaceutical composition according to claim 28 or 29.35.The method according to claim 34, wherein the disease, disorder, or condition is a disease associated with dysregulated PD-L1 expression or activity including cancer and autoimmune disease, or a disease caused by a coronavirus infection including COVID-19.36.A kit comprising:(a) the covalent protein according to any one of claims 1 to 27 or the pharmaceutical composition according to claim 28 or 29; and(b) instructions for use, said instructions for use directing the diagnosis, treatment or prevention of a disease, disorder, or condition in a subject in need thereof.37.The kit according to claim 36, wherein the disease, disorder, or condition is a disease associated with dysregulated PD-L1 expression or activity including cancer and autoimmune disease, or a disease caused by a coronavirus infection including COVID-19.38.A compound having a structure as shown in Formula (II) : whereinR2 has a structure of R21-, or R21-C (O) -R22-R23-, whereR21 is selected from optionally substituted haloalkyl and optionally substituted alkenyl, where haloalkyl or alkenyl can be optionally substituted with one or more halo, -OH, or -CN,R22 is selected from the group consisting of NR22a, O, S, and heterocyclylene, R22a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;R23 is absent or optionally substituted alkanediyl, which is optionally substituted with one or more halo, -OH, or -CN;RA is selected from the group consisting of optionally substituted alkanediyl, and optionally substituted arenediyl;Rw is selected from the group consisting of O and N (Rw1) , where Rw1 is selected from the group consisting of H, alkyl, haloalkyl, and aryl,n is an integer selected from 0 and 1;Ry is selected from the group consisting of S and P;Rz is selected from the group consisting of =O, -O (Rz1) , =N (Rz2) , and -N (Rz3) (Rz4) , where each of Rz1, Rz2, Rz3, and Rz4 is independently selected from the group consisting of H, alkyl, haloalkyl, and aryl;The bondbetween Ry and Rz is a single bond or a double bond.39.The compound according to claim 38, R2 is selected from C1-C4 haloalkyl, C2-C4 alkenyl, C1-C4 haloalkyl-C (O) -NH-, and C2-C4 alkenyl-C (O) -NH-, preferably R2 is selected from the group consisting of -CH2Cl, -CH2Br, -CH2I, -CH2=CH2, -NH-C (O) -CH2-Cl, -NH-C (O) -CH2-Br, -NH-C (O) -CH2-I, and -NH-C (O) -CH2=CH2, more preferably R2 is selected from the group consisting of -CH2Br, -NH-C (O) -CH2-Br, and -NH-C (O) -CH2=CH2.40.The compound according to claim 38 or 39, wherein RA is selected from the group consisting of optionally substituted C1-C8, preferably C1-C6 alkanediyl, and optionally substituted C6-C14, preferably C6-C10, more preferably C6 arenediyl,preferably, RA is C1-C8, preferably C1-C6 alkanediyl optionally substituted with one or more halo, -OH, or -CN, orRA is C6-C14, preferably C6-C10, more preferably C6 arenediyl optionally substituted with one or more Rx, where Rx is selected from the group consisting of H, -halo, -CN, -NO2, haloalkyl, alkoxy, N (Rx1) (Rx2) , where each of Rx1 and Rx2 is independently selected from the group consisting of H, alkyl, and haloalkyl, preferably Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.41.The compound according to any one of claims 38 to 40, wherein the compound has a structure as shown in Formula (II-A) : wherein,p is an integer selected from 1 to 8, preferably 1 to 6, - (CH2) p-optionally substituted with one or more halo, -OH, or -CN.42.The compound according to any one of claims 38 to 41, wherein the compound has a structure as shown in any of Formula (II-A1) - (II-A3) : wherein,p is an integer selected from 1 to 8, preferably 1 to 6.43.The compound according to any one of claims 38 to 40, wherein the compound has a structure as shown in Formula (II-B) : wherein m is an integer selected from 1 to 4.44.The compound according to any one of claims 38 to 40, and 43, wherein the compound has a structure as shown in any of Formula (II-B1) - (II-B4) : wherein m is an integer selected from 1 to 4.45.The compound according to any one of claims 38 to 40, and 43 to 44, wherein the compound has a structure as shown in any of Formula (II-B1-a) - (II-B1-c) , Formula (II-B2-a) - (II-B2-c) , Formula (II-B3-a) - (II-B3-c) , and Formula (II-B4-a) - (II-B4-c) : wherein m is an integer selected from 1 to 4.46.The compound according to any one of claims 38 to 40, and 43 to 45, wherein the compound is selected from the group consisting of 47.A compound having a structure as shown in Formula (II-B1) as a linker for covalent protein: whereinR2 has a structure of R21-, or R21-C (O) -R22-R23-, whereR21 is optionally substituted haloalkyl, or optionally substituted alkenyl, which haloalkyl, alkenyl are optionally substituted with one or more halo, -OH, or -CN,R22 is selected from the group consisting of NR22a, O, S, and heterocyclylene, where R22a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl;R23 is absent or optionally substituted alkylene, which is optionally substituted with one or more halo, -OH, or -CN;Rx is selected from the group consisting of -H, -halo, -CN, and -NO2 or other chemical groups with similar functions, m is an integer selected from 1 to 4;n is an integer selected from 0 and 1.48.The compound according to claim 47, wherein the compound has a structure as shown in Formula (II-B1-a) - (II-B1-c) : 49.The compound according to claim 47 or 48, wherein R22 is selected from the group consisting of NR22a, O, S, and 5-or 6-membered heterocyclylene, for example, 5-or 6-membered heterocyclylene is selected from where R22a is selected from the group consisting of hydrogen, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl.50.The compound according to any one of claims 47 to 49, wherein R2 is -CH2Cl, -CH2Br, -CH2I, -NH-C (O) -CH2-Cl, -NH-C (O) -CH2-Br, -NH-C (O) -CH2-I, or -NH-C (O) -CH2=CH2.51.The compound according to any one of claims 47 to 50, wherein each Rx is independently selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl; preferably each Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2.52.The compound according to any one of claims 47 to 51, wherein the compound having a structure as shown in Formula (II-B1) is selected from the group consisting of where Rx is selected from the group consisting of H, -halo, -CN, -NO2, C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl, C1-C8, preferably C1-C6, more preferably C1-C4 alkoxy, and N (Rx1) (Rx2) , where each of Rx1, and Rx2 is independently selected from the group consisting of H, C1-C8, preferably C1-C6, more preferably C1-C4 alkyl, and C1-C8, preferably C1-C6, more preferably C1-C4 haloalkyl; preferably Rx is selected from the group consisting of H, F, CN, NO2, CF3, OCH3, NHCH3, and N (CH3) 2; more preferably, Rx is selected from the group consisting of H, F, CN, and NO2.53.The compound according to any one of claims 47 to 52, wherein the compound having a structure as shown in Formula (II-B1) is selected from the group consisting of 54.The compound according to any one of claims 47 to 53, wherein the compound having a structure as shown in Formula (II-B1) is selected from the group consisting of 55.A PD-L1 targeting nanobody comprising an amino acid sequence having at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NO. 3.56.The PD-L1 targeting nanobody according to claim 58, wherein the PD-L1 targeting nanobody comprises an amino acid sequence having deletion, substitution, insertion or addition of one or more amino acid residues as compared to the amino acid sequence set forth in SEQ ID NO. 3.57.The PD-L1 targeting nanobody according to claim 58 or 59, wherein the amino acid residue (s) at positions 100, 102, 104, 107, 108, 113, 116, 50, 59, 103, 110, 111, 112, 5, 72, 73, 75, 77, 87, 88, 93, and 123 are as follows,position 100: the amino acid residue at position 100 is selected from the group consisting of Ser (S) , and Cys (C) ;position 102: the amino acid residue at position 102 is selected from the group consisting of Glu (E) , Cys (C) , Phe (F) , and Trp (W) ;position 104: the amino acid residue at position 104 is selected from the group consisting of Pro (P) , and Cys (C) ;position 107: the amino acid residue at position 107 is selected from the group consisting of Thr (T) , Cys (C) , Asn (N) , Ser (S) , and His (H) ;position 108: the amino acid residue at position 108 is selected from the group consisting of Leu (L) , and Cys (C) ;position 113: the amino acid residue at position 113 is selected from the group consisting of Gly (G) , Cys (C) , Ala (A) , and Thr (T) ;position 116: the amino acid residue at position 116 is selected from the group consisting of Gln (Q) , and Cys (C) ;position 50: the amino acid residue at position 50 is selected from the group consisting of Lys (K) , Gly (G) , Thr (T) , and Ala (A) ;position 59: the amino acid residue at position 59 is selected from the group consisting of Tyr (Y) , Phe (F) , and Ser (S) ;position 103: the amino acid residue at position 103 is selected from the group consisting of Asp (D) , and Asn (N) ;position 110: the amino acid residue at position 110 is selected from the group consisting of Thr (T) , Arg (R) , His (H) , Pro (P) , Asp (D) , and Asn (N) ;position 111: the amino acid residue at position 111 is selected from the group consisting of Ser (S) , Tyr (Y) , His (H) , Ala (A) , Phe (F) , and Asn (N) ;position 112: the amino acid residue at position 112 is selected from the group consisting of Ser (S) , Gly (G) , Ala (A) , Asp (D) , and Asn (N) ;position 5: the amino acid residue at position 5 is selected from the group consisting of Gln (Q) , and Val (V) ;position 72: the amino acid residue at position 72 is selected from the group consisting of Gln (Q) , and Arg (R) ;position 73: the amino acid residue at position 73 is selected from the group consisting of Asn (N) , and Asp (D) ;position 75: the amino acid residue at position 75 is selected from the group consisting of Ala (A) , and Ser (S) ;position 77: the amino acid residue at position 77 is selected from the group consisting of Ser (S) , and Asn (N) ;position 87: the amino acid residue at position 87 is selected from the group consisting of Lys (K) , and Arg (R) ;position 88: the amino acid residue at position 88 is selected from the group consisting of Pro (P) , and Ala (A) ;position 93: the amino acid residue at position 93 is selected from the group consisting of Met (M) , and Val (V) ; andposition 123: the amino acid residue at position 123 is selected from the group consisting of Gln (Q) , and Leu (L) .58.A PD-L1 targeting nanobody comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NO. 17,wherein the amino acid residue (s) at positions 50, 59, 103, 107, 110, 111, 112, 113, 5, 72, 73, 75, 77, 87, 88, 93, and 123 are as follows,position 50: the amino acid residue at position 50 is selected from the group consisting of Gly (G) , Thr (T) , and Ala (A) ;position 59: the amino acid residue at position 59 is selected from the group consisting of Phe (F) , and Ser (S) ;position 103: the amino acid residue at position 103 is selected from the group consisting of Asp (D) , and Asn (N) ;position 107: the amino acid residue at position 107 is selected from the group consisting of Thr (T) , Asn (N) , Ser (S) , and His (H) ;position 110: the amino acid residue at position 110 is selected from the group consisting of Arg (R) , His (H) , Pro (P) , Asp (D) , and Asn (N) ;position 111: the amino acid residue at position 111 is selected from the group consisting of Ser (S) , Tyr (Y) , His (H) , Ala (A) , Phe (F) , and Asn (N) ;position 112: the amino acid residue at position 112 is selected from the group consisting of Ser (S) , Gly (G) , Ala (A) , Asp (D) , and Asn (N) ;position 113: the amino acid residue at position 113 is selected from the group consisting of Gly (G) , Ala (A) , and Thr (T) ;position 5: the amino acid residue at position 5 is selected from the group consisting of Gln (Q) , and Val (V) ;position 72: the amino acid residue at position 72 is selected from the group consisting of Gln (Q) , and Arg (R) ;position 73: the amino acid residue at position 73 is selected from the group consisting of Asn (N) , and Asp (D) ;position 75: the amino acid residue at position 75 is selected from the group consisting of Ala (A) , and Ser (S) ;position 77: the amino acid residue at position 77 is selected from the group consisting of Ser (S) , and Asn (N) ;position 87: the amino acid residue at position 87 is selected from the group consisting of Lys (K) , and Arg (R) ;position 88: the amino acid residue at position 88 is selected from the group consisting of Pro (P) , and Ala (A) ;position 93: the amino acid residue at position 93 is selected from the group consisting of Met (M) , and Val (V) ; andposition 123: the amino acid residue at position 123 is selected from the group consisting of Gln (Q) , and Leu (L) .59.The PD-L1 targeting nanobody according to any one of claims 58 to 61, wherein the PD-L1 targeting nanobody comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in any of SEQ ID NO. 4 to 38, preferably the PD-L1 targeting nanobody comprises an amino acid sequence as set forth in any of SEQ ID NOs. 4-38.60.An isolated nucleic acid molecule encoding the PD-L1 targeting nanobody according to any of claims 55 to 59.61.An expression vector comprising the nucleic acid molecule according to claim 60.62.A host cell comprising the nucleic acid molecule according to claim 60 or the expression vector according to claim 61.63.A SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NO. 55.64.The SAR-CoV-2 RBD targeting miniprotein according to claim 63, wherein the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having deletion, substitution, insertion or addition of one or more amino acid residues as compared to the amino acid sequence set forth in SEQ ID NO. 55.65.The SAR-CoV-2 RBD targeting miniprotein according to claim 63 or 64, wherein the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence having one or more mutations selected from the group consisting of D11C, K26C, K26G, K26Q, K27R, F30C, and Y40C as compared to the amino acid sequence set forth in SEQ ID NO. 55.66.A SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in SEQ ID NO. 55,wherein the amino acid residue (s) at positions 11, 26, 27, 30, and 40 are as follows,position 11: the amino acid residue at position 11 is selected from the group consisting of Asp (D) , and Cys (C) ;position 26: the amino acid residue at position 26 is selected from the group consisting of Lys (K) , Cys (C) , Gly (G) , and Gln (Q) ;position 27: the amino acid residue at position 27 is selected from the group consisting of Lys (K) , and Arg (R) ;position 30: the amino acid residue at position 30 is selected from the group consisting of Phe (F) , and Cys (C) ; andposition 40: the amino acid residue at position 40 is selected from the group consisting of Tyr (Y) , and Cys (C) .67.The SAR-CoV-2 RBD targeting miniprotein according to any one of claims 63 to 66, wherein the SAR-CoV-2 RBD targeting miniprotein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%identity as compared to the amino acid sequence set forth in any of SEQ ID NOs. 56-61, preferably the SAR-CoV-2 RBD targeting miniprotein comprises an amino acid sequence as set forth in any of SEQ ID NOs. 56-61.68.An isolated nucleic acid molecule encoding the any SAR-CoV-2 RBD targeting miniprotein according to any one of claims 63 to 67.69.An expression vector comprising the nucleic acid molecule according to claim 68.70.A host cell comprising the nucleic acid molecule according to claim 71 or the expression vector according to claim 69.71.A high throughput covalent protein selection method based on yeast display, comprising1) generating a protein library displayed on yeast cells,2) binding a crosslinker to the displayed proteins;3) incubating with target protein; and4) performing FACS sorting.72.The high throughput covalent protein selection method based on yeast display according to claim 68, wherein the crosslinker is a compound according to any one of claims 38 to 54.73.The high throughput covalent protein selection method based on yeast display according to claim 71 or 72, wherein in step 2) , before binding the crosslinker to the displayer proteins, a reducing agent is added to treat the yeast cells.74.The high throughput covalent protein selection method based on yeast display according to any one of claims 71 to 73, wherein in step 3) , after binding of the target protein, an acid wash or other wash method is performed to remove noncovalent binder in order to selectively screen the covalent binder.75.The high throughput covalent protein selection method based on yeast display according to any one of claims 71 to 74, wherein steps 2) to 4) is repeatedly conducted to further screen candidate covalent protein with high affinity and faster covalent crosslinking rate to the target protein.
Citation Information
Patent Citations
Covalent protein drugs developed via proximity-enabled reactive therapeutics (PERX)
CN113179631A
Pro-drug form (p2PDOX) of the highly potent 2-pyrrolinodoxorubicin conjugated to antibodies for targeted therapy of cancer
WO2014124227A1
Compounds comprising cleavable linker and uses thereof
WO2019008441A1
Compounds comprising cleavable linker and uses thereof
WO2020141459A1
Antibody-drug conjugates comprising Anti-b7-h3 antibodies
WO2021260438A1