Design System and Methodology for bioPROTAC Molecules
By using an automated bioPROTAC molecule design system, combined with a protein design module and a ubiquitination probability assessment module, the problems of long cycle and high cost of PROTAC molecule screening in plants in existing technologies have been solved, and bioPROTAC molecule design for efficient screening of targeted degradation proteins has been realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING SHOUJIE DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-03-24
- Publication Date
- 2026-06-02
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, specifically to a design system and method for bioPROTAC molecules. Background Technology
[0002] Targeted protein degradation (TPD) is a revolutionary strategy in drug development. Unlike traditional drugs that inhibit the activity of pathogenic proteins to achieve anti-disease effects, it aims to utilize the cell's own protein clearance system to specifically remove pathogenic proteins. The ubiquitin-proteasome system (UPS) is the most important protein degradation pathway in cells, and the protein hydrolysis-targeted chimera (PROTAC) technology is currently the most mature and fastest-progressing UPS-based TPD strategy.
[0003] Current methods for discovering PROTAC molecules typically involve large-scale experimental screening, which is time-consuming and costly. To facilitate computer-aided PROTAC screening, researchers start with existing molecules in nature and use docking techniques and molecular dynamics simulations to screen candidate molecules capable of binding to target proteins and E3. However, for different target proteins, these methods require re-screening for suitable E3 and PROTAC or molecular gels, resulting in significant computational overhead. Furthermore, compared to animals, plants have greater difficulty absorbing PROTAC molecules in compound form. Therefore, the rational design of bioPROTAC molecules is crucial for applying target protein degradation technology in plants. Summary of the Invention
[0004] This invention presents, for the first time, an automated design system and method for bioPROTAC molecules. By integrating tools originally used for tasks such as structure prediction, this invention applies them to bioPROTAC design, creating a design system and method for E2-based bioPROTAC molecules. Based on this design system and method, experimentally validated bioPROTAC molecules capable of targeting and degrading target proteins have been successfully designed.
[0005] According to one aspect of the present invention, a design system for bioPROTAC molecules is provided, the design system comprising: The binding protein design module is used to screen binding proteins that bind to target proteins; and A ubiquitination probability assessment module is used to screen bioPROTAC molecules that can specifically degrade target proteins; wherein, the ubiquitination probability assessment module includes forming a candidate bioPROTAC molecule with the binding protein and the E2 protein through a linker, and then predicting the structure of the first complex formed by the candidate bioPROTAC molecule, the ubiquitin protein and the target protein.
[0006] In some embodiments, the prediction of the first complex structure formed by the candidate bioPROTAC molecule, ubiquitin protein, and target protein includes: using different random seeds, calculating the distance and angle between the nitrogen atom of the side chain of the lysine of the target protein and the sulfur atom of the cysteine of the ubiquitin linked to the E2 protein in the first complex structure, and statistically analyzing the percentage of complexes with a distance less than 20 nm, preferably less than 18.4 nm, and an angle cosine value greater than 0.75.
[0007] In some embodiments, a first protein structure prediction algorithm is used to sample the complex structure, and then the distance and angle are calculated. In some embodiments, the first protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3, BioEmu, or ESMdiff, preferably Chai-1.
[0008] In some embodiments, the binding protein design module includes: inputting the full length of the target protein or the domains and / or interaction hotspot residues of the target protein into a binding protein design algorithm to obtain candidate binding proteins; and calculating the iptm value of the second complex structure formed by the candidate binding protein and the target protein, and using the candidate binding protein corresponding to the second complex structure with an iptm greater than 0.7 as the binding protein that binds to the target protein.
[0009] In some embodiments, the binding protein design algorithm includes at least one of BindCraft, RFdiffusion, RFdiffusion3, BoltzGen, or PXDesign, preferably BindCraft.
[0010] In some implementations, the target protein’s domains and / or interaction hotspot residues are input into BindCraft to obtain candidate binding proteins.
[0011] In some embodiments, a second protein structure prediction algorithm is used to calculate the iptm value of the second complex structure. In some embodiments, the second protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3, or AlphaFold-Multimer, preferably Chai-1.
[0012] In some embodiments, the design system further includes a target protein three-dimensional structure prediction module for obtaining the three-dimensional structure of the target protein.
[0013] In some embodiments, the design system further includes a target protein surface interaction site and domain prediction module for obtaining target protein domains and target protein hotspot residues for designing binding proteins.
[0014] In some embodiments, the domains of the target protein are predicted using a protein domain prediction algorithm based on the three-dimensional structure of the target protein. In some embodiments, the protein domain prediction algorithm includes at least one of Chainsaw, DPAM, DomainMapper, UniDoc, Merizo, Chainsaw, or TED, preferably Chainsaw.
[0015] In some embodiments, the surface interaction sites of the target protein are predicted using a protein surface interaction site prediction algorithm based on the three-dimensional structure of the target protein. In some embodiments, the protein surface interaction site prediction algorithm includes at least one of MVGNN-PPIS, SurfDock, MaSif-site, or GPsite, preferably MVGNN-PPIS.
[0016] In some embodiments, the three-dimensional structure of the target protein is predicted using a third protein structure prediction algorithm based on the amino acid sequence of the target protein. In some embodiments, the third protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, or AlphaFold3, preferably Chai-1.
[0017] In some embodiments, the E2 protein is selected from UBE2A, UBE2B, UBE2C, UBE2D1, UBE2D2, UBE2D3, UBE2D4, UBE2E1, UBE2E2, UBE2E3, UBE2F, UBE2G1, UBE2G2, UBE2H, UBE2I, UBE2J1, UBE2J2, UBE2K, UBE2L3, UBE2L6, UBE2M, and UBE2 N, UBE2NL, UBE2O, UBE2Q1, UBE2Q2, UBE2QL, UBE2R1, UBE2R2, UBE2S, UBE2T, UBE2U, UBE2V1, UBE2V2, UBE2W, UBE2Z, UVELD, BIRC6, FTS, TSG101, UFC1, or conserved variants or mutants obtained by adding, deleting, substituting, or modifying one or more amino acids compared to them. Preferably, the E2 protein is selected from Ube2A, Ube2B, Ube2D1, UBE2D2, UBE2D3, Ube2D4, or mutants thereof. More preferably, the E2 protein has an amino acid sequence shown in any one of SEQ ID NO: 1-12, or an amino acid sequence having at least 80% sequence identity with any one of SEQ ID NO: 1-12, or an amino acid sequence shown in any one of SEQ ID NO: 1-12 with one or more amino acid additions, deletions, substitutions or modifications.
[0018] In some embodiments, the linker is a polypeptide linker.
[0019] In some embodiments, the connector has a general formula (G) n S) m The amino acid sequence shown, or the sequence with general formula (G n S) m Compared to amino acid sequences with 1, 2, or 3 inserted, substituted, or deleted amino acids, n and m are integers from 1 to 10. Preferably, n is an integer from 1 to 4, and m is an integer from 1 to 3.
[0020] In some embodiments, the linker has an amino acid sequence formed by a polypeptide unit with the sequence GSAGSAAGSGEF (SEQ ID NO: 14) of 1-7 amino acids, or an amino acid sequence having 1, 2, or 3 inserted, substituted, or deleted amino acids compared to that sequence. According to another aspect of the invention, a method for designing a bioPROTAC molecule is provided, the method comprising the following steps: S1: Screening for binding proteins that bind to the target protein; and S2: The binding protein is linked with the E2 protein to form a candidate bioPROTAC, the structure of the first complex formed by the candidate bioPROTAC molecule, ubiquitin protein and target protein is predicted, and bioPROTAC molecules that can specifically degrade the target protein are screened.
[0021] In some embodiments, step S2, the step of predicting the structure of the first complex formed by the candidate bioPROTAC molecule, ubiquitin protein, and target protein, includes: using different random seeds, calculating the distance and angle between the nitrogen atom of the side chain of the lysine of the target protein and the sulfur atom of the cysteine of the ubiquitin linked to the E2 protein in the structure of the first complex, and statistically analyzing the percentage of complexes with a distance less than 20 nm, preferably less than 18.4 nm, and an angle cosine value greater than 0.75.
[0022] In some embodiments, a first protein structure prediction algorithm is used to sample the complex structure, and then the distance and angle are calculated. In some embodiments, the first protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3, BioEmu, or ESMdiff, preferably Chai-1.
[0023] In some embodiments, in step S1: the full length of the target protein or the domains and / or interaction hotspot residues of the target protein are input into a binding protein design algorithm to obtain candidate binding proteins; and the iptm value of the second complex structure formed by the candidate binding protein and the target protein is calculated, and the candidate binding protein corresponding to the second complex structure with an iptm greater than 0.7 is selected as the binding protein that binds to the target protein.
[0024] In some embodiments, the binding protein design algorithm includes at least one of BindCraft, RFdiffusion, RFdiffusion3, BoltzGen, or PXDesign, preferably BindCraft.
[0025] In some implementations, the target protein’s domains and / or interaction hotspot residues are input into BindCraft to obtain candidate binding proteins.
[0026] In some embodiments, a second protein structure prediction algorithm is used to calculate the iptm value of the second complex structure. In some embodiments, the second protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3, or AlphaFold-Multimer, preferably Chai-1.
[0027] In some embodiments, prior to step S1, the design method further includes: obtaining target protein domains and target protein hotspot residues for designing binding proteins based on the three-dimensional structure of the target protein.
[0028] In some embodiments, the domains of the target protein are predicted using a protein domain prediction algorithm based on the three-dimensional structure of the target protein. In some embodiments, the protein domain prediction algorithm includes at least one of Chainsaw, DPAM, DomainMapper, UniDoc, Merizo, Chainsaw, or TED, preferably Chainsaw.
[0029] In some embodiments, the surface interaction sites of the target protein are predicted using a protein surface interaction site prediction algorithm based on the three-dimensional structure of the target protein. In some embodiments, the protein surface interaction site prediction algorithm includes at least one of MVGNN-PPIS, SurfDock, MaSif-site, or GPsite, preferably MVGNN-PPIS.
[0030] In some embodiments, the three-dimensional structure of the target protein is predicted using a protein three-dimensional structure prediction algorithm.
[0031] In some embodiments, the three-dimensional structure of the target protein is predicted using a third protein structure prediction algorithm based on the amino acid sequence of the target protein. In some embodiments, the third protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, or AlphaFold3, preferably Chai-1.
[0032] In some embodiments, the E2 protein is selected from UBE2A, UBE2B, UBE2C, UBE2D1, UBE2D2, UBE2D3, UBE2D4, UBE2E1, UBE2E2, UBE2E3, UBE2F, UBE2G1, UBE2G2, UBE2H, UBE2I, UBE2J1, UBE2J2, UBE2K, UBE2L3, UBE2L6, UBE2M, and UBE2 N, UBE2NL, UBE2O, UBE2Q1, UBE2Q2, UBE2QL, UBE2R1, UBE2R2, UBE2S, UBE2T, UBE2U, UBE2V1, UBE2V2, UBE2W, UBE2Z, UVELD, BIRC6, FTS, TSG101, UFC1, or conserved variants or mutants obtained by adding, deleting, substituting, or modifying one or more amino acids compared to them. Preferably, the E2 protein is selected from Ube2A, Ube2B, Ube2D1, UBE2D2, UBE2D3, Ube2D4, or mutants thereof. More preferably, the E2 protein has an amino acid sequence shown in any one of SEQ ID NO: 1-12, or an amino acid sequence having at least 80% sequence identity with any one of SEQ ID NO: 1-12, or an amino acid sequence shown in any one of SEQ ID NO: 1-12 with one or more amino acid additions, deletions, substitutions or modifications.
[0033] In some embodiments, the linker is a polypeptide linker.
[0034] In some embodiments, the connector has a general formula (G) n S) m The amino acid sequence shown, or the sequence with general formula (G n S) m Compared to amino acid sequences with 1, 2, or 3 inserted, substituted, or deleted amino acids, n and m are integers from 1 to 10. Preferably, n is an integer from 1 to 4, and m is an integer from 1 to 3.
[0035] In some embodiments, the linker has an amino acid sequence formed by a polypeptide unit of 1-7 amino acid sequence GSAGSAAGSGEF (SEQ ID NO: 14), or an amino acid sequence with 1, 2 or 3 inserted, substituted or deleted amino acids compared to it.
[0036] In some embodiments, the design method further includes: validation experiments on the degradation of target proteins by the designed bioPROTAC molecule. In some embodiments, the validation experiments include in vivo validation experiments and / or in vitro validation experiments.
[0037] According to another aspect of the present invention, a bioPROTAC molecule is provided, which is obtained by the design system or design method described in the present invention.
[0038] According to another aspect of the invention, the use of the bioPROTAC molecule described herein in the ubiquitination and degradation of target proteins, regulation of target proteins, study of target protein function, or preparation of medicaments for the prevention or treatment of diseases or conditions mediated by abnormal levels of target proteins is provided. Attached Figure Description
[0039] Figure 1 The flowchart for designing bioPROTAC is shown. A shows the flowchart for designing the binder's structural domains and hotspot residues, B shows the flowchart for designing the binder sequence, and C shows the flowchart for designing the bioPROTAC.
[0040] Figure 2 The three-dimensional structure of maize ZmGID1 predicted by Chai-1 is shown.
[0041] Figure 3 A schematic diagram is shown of residue L49, which is predicted by MVGNN-PPIS to have the highest interaction score in maize ZmGID1.
[0042] Figure 4 The complex structure of ZmGID1 (yellow) and bioPROTAC (binder-1 (dark blue)-linker (purple)-E2 variant (green)) and Ub (red) is shown. Orange represents the active Cys residue on E2 and light blue represents the Lys residue on ZmGID1.
[0043] Figure 5 It shows Figure 4 A magnified view of the structure of the complex.
[0044] Figure 6 The results of the validation of binder interaction and target protein degradation are shown. A shows the fluorescence imaging results of tobacco leaves showing binder interaction with ZmGID1, and B shows the Western blot analysis results of bioPROTAC-mediated target protein degradation. Detailed Implementation
[0045] Numerous specific details are set forth in the following description to provide a full understanding of the invention. However, the invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0046] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0047] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0048] In this invention, the term "multiple" can refer to two or more, and "at least one" can refer to one, two or more.
[0049] The term "and / or" in this invention is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this invention generally indicates that the preceding and following related objects have an "or" relationship.
[0050] The term "targeted protein degradation (TPD)" refers to a technique that utilizes the cell's own protein degradation system to specifically degrade target proteins in order to regulate protein levels or achieve therapeutic purposes.
[0051] The term "Ubiquitin-Proteasome System (UPS)" refers to the primary pathway for protein degradation in eukaryotic cells, responsible for degrading over 80% of intracellular proteins. Essentially, the system is a complex intracellular cascade reaction, primarily composed of ubiquitin molecules (Ub), ubiquitin activator enzyme (E1), ubiquitin conjugate enzyme (E2), ubiquitin ligase (E3), and the 26S proteasome. Its working principle is as follows: In the presence of ATP, the Ub molecule is activated by E1 and transferred to E2. E2 binds to E3, which can target and recognize substrates. Under the action of E3, the Ub molecule is transferred to the substrate. The Ub-labeled substrate can then be recognized and hydrolyzed by the intracellular 26S proteasome.
[0052] PROTAC technology cleverly "hijacks" the cell's natural UPS system. A typical PROTAC molecule is a bifunctional small molecule consisting of a flexible linker that connects an E3 ligand and a target protein ligand. In vivo, it simultaneously binds the target protein and a specific E3, bringing them close together in physical space to form a transient ternary complex, thereby achieving ubiquitination and degradation of the target protein.
[0053] Currently, most PROTACs are used for disease treatment in the form of small chemical molecules. To supplement the small molecule drug development pathway, various TPD (Transplant-to-Drug) schemes based on biologics have been proposed, among which bioPROTACS is the most mature. These fusion proteins are composed of the "degradation domain" and "target recruitment domain" of an E3 ligase. E3-bioPROTACs can be delivered by physical methods or (more commonly) in the form of gene-encoded constructs, which can capture protein targets and bind efficiently to E2 enzymes, thereby accelerating target degradation. This form mimics the natural function of E3 ubiquitin ligases as substrate recruitment factors. While developing these early E3 fusion proteins, researchers have also proposed that E2 enzymes can be used as PROTAC platforms.
[0054] The term "ubiquitin (Ub)" is a small protein that participates in protein degradation and other cellular regulatory processes by tagging target proteins through ubiquitination modification.
[0055] The term "ubiquitin-activating enzyme (E1)" refers to an enzyme that is responsible for activating Ub molecules and forming high-energy linkages to initiate the ubiquitination process.
[0056] The term "ubiquitin-conjugating enzyme (E2)" refers to an enzyme that transfers activated Ub from E1 and, with the assistance of E3 (under natural conditions), covalently links Ub to substrate proteins.
[0057] The term "ubiquitin ligase (E3)" refers to an enzyme that recognizes a substrate protein and catalyzes the transfer of Ub carried by E2 to the substrate, thereby achieving substrate-specific ubiquitination.
[0058] The term "biological proteolysis-targeting chimera (bioPROTAC)" refers to a biotechnological tool that uses a gene-encoded protein fusion to fuse the functional domain of E3 or E2 with a target protein ligand for the degradation of the target protein.
[0059] The term "molecule" has the meaning of any entity having a first domain as defined herein. In a preferred embodiment, the molecule is a polypeptide.
[0060] The term "linker" refers to a (peptide) linker of natural and / or synthetic origin, composed of linear amino acids. Domains in the molecules of this invention can be linked by linkers, wherein each linker is fused to and / or otherwise linked (e.g., via peptide bonds) with at least two polypeptides or domains. Linkers should have a length suitable for linking two or more monomeric domains in this manner, ensuring that the different domains they are linked to fold correctly and are properly presented to perform their biological functions. In various embodiments, linkers have a flexible conformation. Suitable flexible linkers include, for example, those containing glycine, glutamine, and / or serine residues. In some embodiments, the amino acid residues in the linker can be arranged in small repeating units of up to five amino acids.
[0061] The term "interaction hotspot" refers to amino acid residues in a protein that may interact with other proteins.
[0062] In some embodiments, the bioPROTAC molecule design system of the present invention includes the following modules: (1) Target protein three-dimensional structure prediction module The amino acid sequence of the target protein was obtained from a plant genome database. If the three-dimensional structure of the target protein was known (e.g., determined experimentally or existing in a database), the known structure was used directly for subsequent analysis. If the three-dimensional structure of the target protein was unknown, its three-dimensional structure was predicted using the Chai-1 platform. Figure 1 A). Chai-1 is a deep learning-based protein structure prediction tool that can accurately predict the three-dimensional structure of proteins.
[0063] (2) Protein surface interaction sites and domain prediction module MVGNN-PPIS (Multi-View Graph Neural Network for Protein-Protein Interaction Sites prediction) is a novel deep learning model for predicting protein-protein interaction sites (PPIS). Chainsaw is a supervised learning method based on fully convolutional neural networks for domain segmentation of protein structures. The MVGNN-PPIS algorithm is used to predict interaction sites on the surface of target proteins and calculate the interaction score for each site. Simultaneously, the Chainsaw algorithm is used to predict the domains of the target proteins, determining their locations and extents to allow for the selection of domains corresponding to sites with high interaction scores (see [link to relevant documentation]). Figure 1 A).
[0064] (3) Binding protein (binder) sequence design module BindCraft is a powerful automated pipeline for de novo design of protein binding agents. The extracted domains and the amino acid sequences of their high-interaction-score sites are input into the BindCraft model to generate binder sequences capable of binding to target proteins (see [link to BindCraft]). Figure 1 B). The Interface Prediction Template Model (IPTM) score focuses on evaluating the accuracy of interface regions in protein complex structure prediction. It assigns a score between 0 and 1 by comparing the spatial arrangement and interaction patterns of interface residues in the predicted protein complex structure with those in the actual complex structure. A higher score indicates a more accurate prediction of the interface structure. The computational design platform for plant target protein degradation systems uses Chai-1 to predict the complex structure of the binder and the full-length sequence of the target protein. Binders with IPTM values > 0.7 are selected as high-confidence binders for downstream analysis.
[0065] (4) Ubiquitination probability assessment module The ubiquitination probability assessment module includes target protein degradation protein sequence construction and ubiquitination conformation analysis. E2 protein can bind to ubiquitin and transfer it to the target protein. Ubiquitin transfer requires the target protein, E2 protein, and ubiquitin to form a suitable conformation. Specifically, the cysteine thiol group of the E2 protein linking to the ubiquitin needs to be sufficiently close to the side-chain nitrogen atom of the lysine residue of the target protein, and the side-chain nitrogen atom of the lysine residue of the target protein needs to be as close as possible to the cysteine sulfur atom of the E2 protein. For binders with an iptm > 0.7, they are combined with linker sequences of different lengths and a specific E2 protein sequence to form candidate bioPROTACs. Chai-1 is used to predict the conformation of the complex formed by the bioPROTACs, ubiquitin, and target protein. In some embodiments, the linker sequence used is “GSAGSAAGSGEF” repeated 1 to 7 times, for a total of 7 different linker lengths. To collect conformational distribution data, different random seeds were used to predict multiple complex structures for each bioPROTAC. The distance and angle between the nitrogen atom of the lysine side chain of the target protein and the sulfur atom of the cysteine residue of the ubiquitin-linked E2 protein in these complex structures were calculated. The percentage of complexes with a distance less than 18.4 nm and an angle cosine value greater than 0.75 was statistically analyzed. Based on this data, ideal bioPROTAC molecules were screened. Figure 1 C).
[0066] In some embodiments, the bioPROTAC molecule design method of the present invention includes the following steps: (1) such as Figure 1 As shown in Figure A, the platform first uses Chai-1 to predict the three-dimensional structure of the target protein by inputting the sequence of the target protein to be degraded. This three-dimensional structure is then input into Chainsaw, which breaks it down into multiple domains. Simultaneously, this three-dimensional structure is input into MVGNN-PPIS to predict the probability of each residue of the target protein interacting with other proteins, with values ranging from 0 to 1. The interaction probability of each predicted residue, and the average of the interaction probabilities with its adjacent residues, are considered as the interaction score for that residue. (2) For example Figure 1 As shown in Figure B: Input the residues with the highest interaction score and their corresponding domains into BindCraft, and output candidate binder sequences. Use Chai-1 to predict the complex structure of the binder and the full-length target protein sequence, and calculate the IPTM value of the complex structure. Select binder sequences with IPTM values greater than 0.7. (3) The obtained binder sequence is combined with linker sequences of different lengths and a specific E2 protein sequence to form candidate bioPROTACs. Chai-1 is used to predict the conformation of the complex formed by the bioPROTACs, ubiquitin protein, and target protein. To collect conformational distribution, different random seeds are used to predict multiple complex structures for each bioPROTAC. The distance and angle between the nitrogen atom of the lysine side chain of the target protein and the sulfur atom of the cysteine of the ubiquitin linker in these complex structures are calculated. The percentage of complexes with a distance less than 18.4 nm and an angle cosine value greater than 0.75 is counted. Based on this data, ideal bioPROTAC molecules are screened.
[0067] The sequence information involved in this article is shown in Table 1 below.
[0068] Table 1 The following embodiments and accompanying drawings are provided to aid in understanding the present invention. However, it should be understood that these embodiments and drawings are for illustrative purposes only and do not constitute any limitation. The actual scope of protection of the present invention is set forth in the claims. It should be understood that any modifications and changes can be made without departing from the spirit of the present invention.
[0069] Example 1: BioPROTAC Design of Maize ZmGID1 The maize ZmGID1 (uniprot id: A0A1D6M449) protein is the receptor for gibberellin (GA). When GA binds to ZmGID1, it activates the GA signaling pathway, promoting maize plant height. Targeted degradation of ZmGID1 is beneficial for negatively regulating maize plant height, resulting in lodging-resistant maize varieties.
[0070] This embodiment designs a bioPROTAC molecule targeting maize ZmGID1, and the specific steps are as follows.
[0071] (1) such as Figure 1 As shown in Figure A, input the amino acid sequence of the target protein to be degraded, maize ZmGID1 (uniprot id: A0A1D6M449). First, use Chai-1 to predict the three-dimensional structure of the protein, as shown in Figure A. Figure 2 As shown in the figure. The three-dimensional structure was input into Chainsaw and split into multiple domains. Chainsaw split ZmGID1 into three domains: amino acid residues 9 to 49, amino acid residues 63 to 225, and amino acid residues 234 to 349. Simultaneously, this three-dimensional structure was input into MVGNN-PPIS to predict the probability of each residue of the target protein interacting with other proteins, with values ranging from 0 to 1. The predicted interaction probability of each residue, and the average of the interaction probabilities with adjacent residues, were considered as the interaction score for that residue, as shown in Table 2 below. The residue with the highest interaction score was L49, as shown in Table 2. Figure 3 As shown.
[0072] Table 2 (2) For example Figure 1 As shown in B: Due to the short total length of ZmGID1 (351), no domain splitting is performed. The full-length predicted structure of ZmGID1 and the interaction hotspot L49 with the highest interaction score determined in step (1) are input into BindCraft, and candidate binder sequences are output. Exemplary binders are shown in Table 3 below. The complex structure of the binder and the full-length target protein sequence is predicted using Chai-1, and the IPTM value of the complex structure (the average of 5 IPTM values) is calculated. Binder sequences with IPTM values greater than 0.7 are selected. The binder sequences and their corresponding IPTM values are shown in Table 3 below.
[0073] Table 3 (3) such as Figure 1As shown in C: The binder with iptm > 0.7 determined in step (2) is connected to a linker with "GSAGSAAGSGEF (SEQ ID NO: 14)" as the repeating unit at its C-terminus. Then, E2 protein (UBE2D1-V4) is connected to the C-terminus of the linker to form a bioPROTAC and the ubiquitination probability of the bioPROTAC is evaluated to form candidate bioPROTACs. The conformation of the complex formed by bioPROTACs, ubiquitin protein and target protein is predicted using Chai-1. To collect the conformation distribution, different random seeds are used, and multiple complex structures are predicted for each bioPROTAC. After sampling the complex structures using Chai-1, the distance and angle between the nitrogen atom of the lysine side chain of the target protein and the sulfur atom of the cysteine of the ubiquitin linker of the E2 protein are calculated. The percentage of complexes with a distance less than 18.4 nm and an angle cosine value greater than 0.75 is counted. Based on this data, ideal bioPROTAC molecules are screened. Figure 4 and Figure 5 As shown, in the conformation of this complex structure, the distance between the nitrogen atom of the side chain of ZmGID1 lysine and the sulfur atom of the cysteine of the E2 protein linking ubiquitin is less than 18.4 nm (specifically 15.232 nm), and the cosine of the included angle is greater than 0.75 (specifically 0.994), so it is considered a conformation suitable for ubiquitination.
[0074] The calculated complex structures of each bioPROTAC are shown in Table 4 below. The results indicate that when the linker length is 36 (i.e., with 3 repeat units), at least one of the 100 complex structures is suitable for ubiquitination. Finally, the bioPROTAC sequence of the binder-3×repeat unit-E2 protein was obtained.
[0075] Table 4
[0076] Example 2: Combination of protein validation and target protein degradation assay 1. Use the LUC system to detect the physical interaction between the binding protein and ZmGID1. The fusion construct cLUC-ZmGID1 and nLUC-binding protein were transiently co-expressed in tobacco (Nicotiana benthamiana) leaves. Negative controls included combinations of cLUC + nLUC-binding protein, cLUC-ZmGID1 + nLUC, and the empty vector. Forty-eight hours after infiltration, D-luciferin substrate was injected, and luminescence signals were recorded using a chemiluminescence imaging system.
[0077] The specific experimental steps are as follows: The ZmGID1 binder-1 sequence and ZmGID1 protein sequence were ligated into the pCambia1300-5xFLAG-nLUC and pCambia1300-cLUC-3xHA expression vectors, respectively, and fused with nLUC and cLUC for expression. After successfully constructing the vectors, they were transformed into Agrobacterium GV3101 and subjected to a transient tobacco expression experiment according to the method in Example 1. The results were detected 48 hours after injection. D-fluorescein sodium potassium salt (working concentration 150ug / mL) was prepared and injected again into the Agrobacterium liquid injection area. After incubation at room temperature in the dark for 10 min, imaging was performed using a Tanon-4600 chemiluminescence imaging system.
[0078] The transient expression assay for tobacco was performed as follows: Single colonies of Agrobacterium, identified by PCR, were picked and incubated overnight at 28°C and 200 rpm in 1 ml LB liquid medium (Kana + Rif, working concentration of Kana 50 mg / L, working concentration of Rif 40 mg / L). 200 μl of the culture was then added to 10 ml LB liquid medium (Kana + Rif) and incubated at 28°C and 200 rpm for 12-16 h until the bacterial concentration reached OD500. 600 =1.2-1.5. Collect bacteria at 3000 rpm for 10 min, discard the supernatant. Resuspend the bacteria in 10 mM MgCl2 to the specified OD value, add MES (final concentration 10 mM) and AS (final concentration 0.2 mM), and incubate in the dark for 2 h. Inject into tobacco leaves (two 12 mm² injection sites per plant). 2 The mixture consists of two circles (approximately 0.5 ml in total volume). Each group is injected with three tobacco plants, each with three leaves, and cultured according to experimental requirements.
[0079] Figure 6 The results of A showed that a strong luminescent signal was detected in the cLUC-ZmGID1 + nLUC-binding protein combination compared with all controls, indicating that the binding protein specifically interacts with ZmGID1.
[0080] 2. Western blot analysis of bioPROTAC-mediated target protein degradation The FLAG-labeled bioPROTAC and ZmGID1-YFP fusion protein were co-expressed in tobacco leaves, with the amount of bioPROTAC plasmid gradually increasing (0.1, 0.5, 1.0 μg). Total protein was extracted 72 hours after infiltration. BioPROTAC expression was validated using an anti-FLAG antibody, and ZmGID1-YFP protein levels were measured. Ponceau S staining was used as a loading control.
[0081] The specific experimental steps are as follows.
[0082] (1) Sequence acquisition and vector construction: The gene sequence of the target protein ZmGID1 was ligated into the plant binary expression vector pEarleyGate104, placed between the 35S promoter and the OCS terminator, and fused with YFP for expression. A 3xFLAG protein tag was added to the N-terminus of the sequence for subsequent detection of protein expression levels. The bioPROTAC sequence was optimized according to the Nicotiana benthamian codon and synthesized by Suzhou Junji Gene Technology Co., Ltd. The obtained sequence was cloned into the plant binary expression vector pCambia1300-3xFLAG-35S-X. All plasmids were constructed and amplified using Escherichia coli DH5α strain (DL1001, Shanghai Weidi Biotechnology Co., Ltd.). After the expected construction product was verified to be correct by Sanger sequencing, it was transformed into Agrobacterium GV3101 strain (AC1001, Shanghai Weidi Biotechnology Co., Ltd.) for subsequent experiments.
[0083] (b) Transient expression assay in tobacco: Single colonies of Agrobacterium, identified by PCR, were picked and cultured overnight at 28°C and 200 rpm in 1 ml LB liquid medium (Kana+Rif). 200 μl of the culture was added to 10 ml LB liquid medium (Kana+Rif) and cultured at 28°C and 200 rpm for 12-16 h until the bacterial concentration reached OD600 = 1.2-1.5. The colonies were collected at 3000 rpm for 10 min, and the supernatant was discarded. The colonies were resuspended in 10 mM MgCl2 to the specified OD value, and MES (final concentration 10 mM) and AS (final concentration 0.2 mM) were added. The colonies were then incubated in the dark for 2 h. The colonies were then injected into tobacco leaves (two 12 mm² anodes were injected into each treatment group). 2 Each group was injected with three tobacco plants (each with three leaves) in a circular shape (approximately 0.5 ml in total volume), and cultured according to experimental requirements. PDS1+MgCl2 and PDS1+binder were used as control groups.
[0084] (c) Western blot: Samples were taken 72 hours after injection. Two discs were taken from each sample using a 10mm punch, ground, and then 2xSDS loading buffer was added. The mixture was incubated on ice for 10 minutes, then boiled at 100°C for 10 minutes. After cooling, the mixture was centrifuged at 12,000 rpm for 3 minutes, and the supernatant was used for sample loading. Electrophoresis was performed using a 10% SDS-PAGE gel. The protein was then transferred to a 0.45μm PVDF membrane, stained with Ponceau S (P8330, Beijing Solarbio Science & Technology Co., Ltd.), and the residual dye was washed away with TBST. The membrane was blocked with 5% skim milk powder for 1 hour, washed with TBST, and then anti-FLAG-Tag mAb (AE005, Abiotech Co., Ltd.) diluted 1:20,000 in 3% BSA solution was added. The membrane was incubated overnight at 4°C. After primary antibody treatment, the membrane was washed and incubated with a 1:5000 dilution of HRP-conjugated Goat anti-Mouse IgG (H+L) (AS003, Abiotech Biotechnology Co., Ltd.) secondary antibody at room temperature for 1 hour. Unbound antibodies were then washed away, and signal detection was performed using a Clarity Western ELISA kit (1705060; Bio-Rad), with imaging using a Tanon-4600 chemiluminescence imaging system. The 45S large subunit band, stained with Ponceau S, was used as an internal control to assess the degradation of the target protein.
[0085] Figure 6 The results from B showed that the abundance of ZmGID1-YFP decreased in a dose-dependent manner as bioPROTAC expression increased, confirming the effective degradation of the target protein.
[0086] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope defined in claim 1.
Claims
1. A bioPROTAC molecule design system, comprising: A protein-binding design module is used to screen binding proteins that bind to target proteins; and The ubiquitination probability assessment module is used to screen bioPROTAC molecules that can specifically degrade target proteins; The ubiquitination probability assessment module includes: forming a candidate bioPROTAC molecule with the binding protein and the E2 protein through a linker, and then predicting the structure of the first complex formed by the candidate bioPROTAC, the ubiquitin protein, and the target protein.
2. The design system according to claim 1, characterized in that, The structure of the first complex formed by the predicted candidate bioPROTAC molecule, ubiquitin protein, and target protein includes: Using different random seeds, calculate the distance and angle between the nitrogen atom of the side chain of the lysine of the target protein and the sulfur atom of the cysteine of the ubiquitin linked to the E2 protein in the first complex structure, and count the percentage of complexes with a distance less than 20 nm, preferably less than 18.4 nm, and an angle cosine value greater than 0.
75. Preferably, the complex structure is sampled using a first protein structure prediction algorithm, and then the distance and angle are calculated; more preferably, the first protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3, BioEmu or ESMdiff, preferably Chai-1.
3. The design system according to claim 1 or 2, characterized in that, The binding protein design module includes: inputting the full length of the target protein or the domains and / or interaction hotspot residues of the target protein into the binding protein design algorithm to obtain candidate binding proteins; and calculating the iptm value of the second complex structure formed by the candidate binding protein and the target protein, and using the candidate binding protein corresponding to the second complex structure with an iptm greater than 0.7 as the binding protein that binds to the target protein. Preferably, the protein binding design algorithm includes at least one of BindCraft, RFdiffusion, RFdiffusion3, BoltzGen, or PXDesign, with BindCraft being the most preferred. Preferably, the iptm value of the second complex structure is calculated using a second protein structure prediction algorithm; more preferably, the second protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3 or AlphaFold-Multimer, preferably Chai-1.
4. The design system according to any one of claims 1-3, characterized in that, The design system also includes a target protein three-dimensional structure prediction module for obtaining the three-dimensional structure of the target protein; and / or The design system also includes a target protein surface interaction site and domain prediction module, used to obtain target protein domains and target protein hotspot residues for designing binding proteins. Preferably, the target protein's domains are predicted using a protein domain prediction algorithm based on the target protein's three-dimensional structure; more preferably, the protein domain prediction algorithm includes at least one of Chainsaw, DPAM, DomainMapper, UniDoc, Merizo, Chainsaw, or TED, preferably Chainsaw. Preferably, based on the three-dimensional structure of the target protein, a protein surface interaction site prediction algorithm is used to predict the surface interaction sites of the target protein; more preferably, the protein surface interaction site prediction algorithm includes at least one of MVGNN-PPIS, SurfDock, MaSif-site or GPsite, preferably MVGNN-PPIS. Preferably, the three-dimensional structure of the target protein is predicted using a third protein structure prediction algorithm based on the amino acid sequence of the target protein; more preferably, the third protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold or AlphaFold3, preferably Chai-1.
5. A method for designing bioPROTAC molecules, comprising the following steps: S1: Screening for binding proteins that bind to the target protein; and S2: The binding protein is linked with the E2 protein to form a candidate bioPROTAC molecule, the structure of the first complex formed by the candidate bioPROTAC molecule, ubiquitin protein and target protein is predicted, and bioPROTAC molecules that can specifically degrade the target protein are screened.
6. The design method according to claim 5, characterized in that, In step S2, the step of predicting the structure of the first complex formed by the candidate bioPROTAC molecule, ubiquitin protein, and target protein includes: Using different random seeds, calculate the distance and angle between the nitrogen atom of the side chain of the lysine of the target protein and the sulfur atom of the cysteine of the ubiquitin linked to the E2 protein in the first complex structure, and count the percentage of complexes with a distance less than 20 nm, preferably less than 18.4 nm, and an angle cosine value greater than 0.
75. Preferably, the complex structure is sampled using a first protein structure prediction algorithm, and then the distance and angle are calculated; more preferably, the first protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3, BioEmu or ESMdiff, preferably Chai-1.
7. The design method according to claim 5 or 6, characterized in that, In step S1: the full length of the target protein or the domains and / or interaction hotspot residues of the target protein are input into the binding protein design algorithm to obtain candidate binding proteins; and the iptm value of the second complex structure formed by the candidate binding protein and the target protein is calculated, and the candidate binding protein corresponding to the second complex structure with an iptm greater than 0.7 is taken as the binding protein that binds to the target protein. Preferably, the protein binding design algorithm includes at least one of BindCraft, RFdiffusion, RFdiffusion3, BoltzGen, or PXDesign, with BindCraft being the most preferred. Preferably, the iptm value of the second complex structure is calculated using a second protein structure prediction algorithm; more preferably, the second protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold, AlphaFold3 or AlphaFold-Multimer, preferably Chai-1.
8. The design method according to any one of claims 5-7, characterized in that, Prior to step S1, the design method further includes: obtaining target protein domains and target protein hotspot residues for designing binding proteins based on the three-dimensional structure of the target protein; Preferably, the target protein's domains are predicted using a protein domain prediction algorithm based on the target protein's three-dimensional structure; more preferably, the protein domain prediction algorithm includes at least one of Chainsaw, DPAM, DomainMapper, UniDoc, Merizo, Chainsaw, or TED, preferably Chainsaw. Preferably, based on the three-dimensional structure of the target protein, a protein surface interaction site prediction algorithm is used to predict the surface interaction sites of the target protein; more preferably, the protein surface interaction site prediction algorithm includes at least one of MVGNN-PPIS, SurfDock, MaSif-site or GPsite, preferably MVGNN-PPIS. Preferably, the three-dimensional structure of the target protein is predicted by a protein three-dimensional structure prediction algorithm; More preferably, the three-dimensional structure of the target protein is predicted using a third protein structure prediction algorithm based on the amino acid sequence of the target protein; more preferably, the third protein structure prediction algorithm includes at least one of Chai-1, Boltz-1, Boltz-2, Protenix, SeedFold or AlphaFold3, preferably Chai-1; More preferably, the design method further includes: a verification experiment on the degradation of target proteins by the designed bioPROTAC molecule.
9. A bioPROTAC molecule obtained by the design system of claims 1-4 or the design method of any one of claims 5-8.
10. The use of the bioPROTAC molecule of claim 9 in the ubiquitination and degradation of target proteins, regulation of target proteins, research on target protein function, or preparation of medicaments for the prevention or treatment of diseases or conditions mediated by abnormal levels of target proteins.