Transmembrane modulator protein design method based on hinting strategy and generative model
By combining virtual probe identification with hotspots and structural cueing strategies, the reliance on manually specified binding sites in existing technologies has been eliminated. This enables automated, function-guided design of transmembrane proteins, generating stable and functionally specific regulatory proteins and expanding the range of design targets.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-09
AI Technical Summary
Existing AI protein design methods rely heavily on manually specified binding sites, making it difficult to automatically identify unknown allosteric sites in transmembrane proteins such as GPCRs. Furthermore, they lack the ability to directionally regulate membrane environment adaptation and biological function, resulting in poor design performance.
A transmembrane regulatory protein design method based on cueing strategies and generative models is adopted. By using virtual probes to identify hotspot regions and combining site-guided insertion, site blocking preoccupancy, and conformation induction strategies, closed-loop iterative optimization is achieved to generate function-guided transmembrane regulatory protein sequences.
It enables automated, function-directed design of transmembrane proteins, expands the range of designable targets, improves the stability and functional specificity of designed products in the membrane environment, and generates a variety of functional regulatory proteins such as agonists, inhibitors, and biased regulators.
Smart Images

Figure CN122177201A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics, specifically relating to a method for designing transmembrane regulatory proteins based on cue strategies and generative models. Background Technology
[0002] Proteins are the primary carriers of life activities, and their biological functions are determined by one-dimensional sequences composed of amino acids. These sequences fold to form specific three-dimensional spatial structures (also known as conformations). Membrane proteins, represented by G protein-coupled receptors (GPCRs), are key proteins for transmembrane signal transduction in cells. Located on the cell membrane, the conformational change capability of their transmembrane domains determines the specificity of signal transduction. Protein design, as the reverse process of protein folding, aims to achieve desired biological functions by creating amino acid sequences with specific structures. Traditional design processes follow a two-stage "structure-sequence" model: first, the three-dimensional structural framework corresponding to the target function is deduced; then, a sequence stabilizing this framework is designed based on the principle of energy minimization; and finally, its function is verified. Conventional design tasks include binder design and motif scaffolding.
[0003] In recent years, deep learning technology has greatly promoted the development of protein design. Key technical components include: 1) Protein structure prediction: Models represented by AlphaFold2, utilizing co-evolutionary information and attention mechanisms in multiple sequence alignment (MSA), can predict the three-dimensional structure from end to end of the amino acid sequence with near experimental accuracy, and are often used as a verification tool for designed sequences; 2) Protein sequence design (refolding): Graph neural network models represented by ProteinMPNN can quickly generate high-confidence amino acid sequences based on a given three-dimensional backbone; 3) Protein backbone generation: Early methods (such as Rosetta) relied on physical energy functions and fragment assembly, resulting in huge computational overhead. Generative AI methods (such as RFdiffusion) learn the spatial distribution of protein structures and use diffusion models to generate novel backbones with specific geometric features from noise.
[0004] Building upon the aforementioned technical components, the mainstream and effective approach for designing binding proteins targeting specific targets is the hallucination-based binding design process, such as the known BindCraft process. The core idea of this approach is to utilize a high-precision structure prediction network as an "evaluator," or to generate binding proteins under specific constraints. Its typical technical path is as follows: First, the user inputs the target protein structure and manually specifies its surface binding hotspots. Then, the system iteratively optimizes (the hallucination process) by initializing random sequences or noisy structures: the generated portion is compared with the target input structure prediction network, and a composite loss function including binding confidence, structural confidence, and geometric constraints is defined. Gradient descent or Markov Monte Carlo sampling (MCMC) is used to continuously optimize the binder sequence or structure, ensuring geometrical complementarity interactions with the preset hotspot residues. Finally, the generated backbone is redesigned, and AlphaFold2 is used for self-consistency checking to screen for high-binding-energy design candidates.
[0005] While existing techniques based on illusion or diffusion models have achieved some success in the design of soluble protein binders, they still have the following major technical limitations when applied to the design of functional regulators for complex membrane proteins such as GPCRs: First, existing methods rely heavily on prior human knowledge and lack the ability to automatically discover binding sites. They require users to precisely specify the "binding site" and cannot automatically explore and identify unknown allosteric sites on the surface of GPCRs that can effectively regulate receptor function during the design process, which greatly limits the design scope for targets that are difficult to drug.
[0006] Secondly, existing technologies focus only on physical binding and lack targeted regulation of biological functions. Their design logic is based on geometric complementarity and energy minimization, with the goal of generating molecules that can adhere to the target. They lack the perception and guidance of receptor conformational states (such as activated and inactive states), making it difficult to accurately design functional molecules that can specifically open or close downstream signaling pathways.
[0007] Finally, existing technologies lack controllability over membrane environments and specific sequence characteristics. Their training data is mostly based on water-soluble proteins, making it difficult to introduce membrane environment-adaptive constraints during the generation stage. This leads to problems such as folding failure or non-specific adsorption in the designed molecules.
[0008] Overcoming the dependence of existing AI protein design methods on manually specified binding sites and realizing the automatic and precise design of novel protein regulators that can directionally regulate the biological functions of transmembrane proteins (such as GPCRs) in the complex membrane environment is an urgent problem to be solved in this field. Summary of the Invention
[0009] To address the aforementioned issues, this invention proposes a transmembrane regulatory protein design method based on cueing strategies and generative models. This method overcomes the dependence on manually specified binding sites, enables automated and function-guided design of transmembrane regulatory proteins, effectively expands the range of designable targets, and significantly improves the stability and functional specificity of the designed products in the membrane environment.
[0010] The technical solution adopted in this invention is as follows: This invention proposes a method for designing transmembrane regulatory proteins based on a cue strategy and a generative model, comprising the following two stages executed sequentially: Target survey phase: High-throughput automatic identification of binding hotspot regions on the surface of target membrane proteins based on virtual probes; Closed-loop design phase: Based on the aforementioned hotspot regions, an initial backbone set is first generated using a generative diffusion model. Then, a structure cueing strategy guides the protein sequence design model and structure prediction model for iterative optimization to generate transmembrane regulatory protein sequences.
[0011] Furthermore, the target survey phase includes: A generative diffusion model was used to generate multiple virtual probe backbones targeting the membrane protein. Based on protein sequence design models and structure prediction models, a set of conformations of probe-target membrane protein complexes is obtained; The binding hotspot region is determined by analyzing the spatial distribution density of probes in the complex conformation set.
[0012] Furthermore, the determination of binding hotspot regions by analyzing the spatial distribution density of the probes includes: The C-alpha atom coordinates of the probes in the complex conformation set are extracted to form conformation point cloud data; A three-dimensional voxel grid covering a predetermined hotspot region of the target membrane protein is constructed, wherein the size of the three-dimensional voxel grid is 2.0 Å; Calculate the distribution density value of the conformation point cloud data within the three-dimensional voxel grid; A three-dimensional heat map is generated based on the distribution density value, and the high-density areas in the heat map are identified as the combined hotspot areas.
[0013] Furthermore, the closed-loop design phase includes: Based on the control objectives, a structural prompt containing physical constraints is constructed based on the aforementioned hotspot areas; Based on the initial backbone set generated by the structural hints and the generative diffusion model, the protein sequence design model and the structure prediction model are used for closed-loop iterative optimization to screen out transmembrane regulatory protein sequences that meet the preset structural accuracy index.
[0014] Furthermore, structural cues include: Site-guided insertion strategy: The designed sequence is inserted into a functionally unrelated position in the target membrane protein sequence that is adjacent to the binding hotspot region using a flexible linker, resulting in a single-stranded input sequence; Site blocking pre-occupancy strategy: Introduce a known binding protein into the single-stranded input sequence to occupy the non-target binding region of the target membrane protein, thus forming steric hindrance; Conformation induction strategy: A downstream effector protein sequence of the target membrane protein is introduced into the single-stranded input sequence to form a complex input sequence representing a specific functional state, thereby inducing the target membrane protein to present a specific functional conformation.
[0015] Furthermore, the closed-loop iterative optimization includes the following steps: a) Determine the control target and the target target region selected from the combined hotspot region, and use a generative diffusion model to generate an initial skeleton set containing a batch of initial skeletons for the target target region, the initial skeletons being used as skeletons to be optimized; b) Use the protein sequence design model to generate a batch of design sequences for each backbone to be optimized; c) Insert the designed sequence into the target membrane protein sequence using a site-directed insertion strategy to obtain a batch of single-stranded sequences as input sequences; if the regulatory target is the functional regulation of G protein-coupled receptors, a conformation induction strategy is also required to introduce downstream effector protein sequences of the target membrane protein into the single-stranded sequence as input sequences. Each input sequence corresponds to an initial skeleton; d) Utilize the structural prediction model to predict the full atomic structure and its confidence level of the input sequence, extract the predicted backbone part and its corresponding local confidence level from it, and batch predict the backbone for the same design sequence; e) Based on the local confidence index and global self-consistency index of the predicted skeleton, select N best predicted skeletons from the results of step d). f) Select the N best predicted skeletons selected in this iteration as the skeletons to be optimized in the next iteration, and repeat steps b) to e) until all predicted skeletons meet the preset structure accuracy index. g) Output the final transmembrane regulatory protein sequence that meets the criteria.
[0016] Furthermore, the local confidence index is the predicted local distance difference test pLDDT, the global self-consistency index is the global self-consistency root mean square deviation scRMSD between the predicted skeleton and the corresponding initial skeleton; the preset structural accuracy index is that the global self-consistency root mean square deviation scRMSD is less than 1.5 Å and pLDDT is greater than 90.
[0017] Furthermore, the generative diffusion model is the RFDiffusion model, the protein sequence design model is the ProteinMPNN model, and the structure prediction model is the AlphaFold2 model.
[0018] Further, in e), N best predicted skeletons are selected from the results of step d), specifically: those with a global self-consistency root mean square deviation (scRMSD) < 10 Å from the initial skeleton and ranked in the top N of the local confidence index; In g), the design sequences are sorted according to the average index of the batch predicted backbone from the same design sequence, and several best design sequences are output as the generated transmembrane regulatory protein sequences.
[0019] Furthermore, after one round of refolding iteration testing, if the local confidence pLDDT ranking among the top N candidate predicted skeletons shows low global self-consistency with the initial skeleton, such as the proportion of global scRMSD < 20 Å being less than 50% in the first round, it can be determined that there is a risk of off-target convergence (in subsequent rounds, screening by global scRMSD < 10 Å may not yield effective samples for ranking), a site blocking pre-occupancy strategy can be introduced.
[0020] The beneficial effects of this invention are: This invention provides an automated, function-oriented method for designing transmembrane regulatory proteins, completely overcoming the reliance on prior human knowledge in existing technologies. Through high-throughput target surveys based on virtual probes, this method can automatically and efficiently explore and identify unknown allosteric binding hotspots on membrane protein surfaces, eliminating the need for precise pre-specification of binding sites and thus greatly expanding the range of designable targets. Secondly, by constructing a closed-loop iterative design framework based on structural cues, it integrates multiple targeted strategies: using a site-guided insertion strategy to achieve precise spatial localization of anchored proteins; using a site-blocking pre-occupancy strategy to effectively address off-target issues during the design process, guiding the model to converge towards non-dominant binding regions; and using a conformation-inducing strategy to actively regulate receptor conformational states (such as activation or inhibition). These strategies work synergistically to achieve a leap from simple physical binding design to precise biological function design, successfully generating a variety of functional regulatory proteins, including agonists, inhibitors, and biased regulators. Furthermore, this method introduces transmembrane environment adaptability constraints in sequence design and ensures high structural conservation and high confidence of the generated protein through closed-loop iterative optimization, significantly improving the stability and functional specificity of the designed product in complex membrane environments. Validation by examples demonstrates that the designed protein binds tightly to the target and exhibits significant functional regulation, fully proving the accuracy and effectiveness of the method of this invention. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of a high-throughput target survey process based on virtual probes; Figure 2 This is a schematic diagram of a closed-loop iterative design process based on structural cues; Figure 3 This is a schematic diagram of the framework for a transmembrane regulatory protein design method based on cueing strategies and generative models; Figure 4 This is a diagram showing the results of the D1R-anchored exoskeleton protein design based on site-guided insertion in Example 1. In this diagram, a represents the iterative process, including a schematic diagram of the generated protein binding conformation, the SeqLogo diagram of the final round, and the pLDDT mean and variance of the iterative optimization process. In this diagram, b represents the structure of the final round of preferred sequence resolved using cryo-electron microscopy and a structural comparison with the designed structure. Figure 5 This is a diagram showing the results of the site-blocking biased regulator design in Example 2. In this diagram, a is the iterative process, including a schematic diagram of the generated protein binding conformation, the SeqLogo diagram of the final round, and the pLDDT mean and variance of the iterative optimization process. b is the structure of the final round of preferred sequence resolved by cryo-electron microscopy and the structure comparison with the designed structure. Figure 6This is a diagram showing the results of conformation-induced agonist design in Example 3. In this diagram, a represents the iterative process, including a schematic diagram of the generated protein binding conformation, the SeqLogo diagram of the final round, and the pLDDT mean and variance of the iterative optimization process. b shows the structure of the final round of preferred sequences resolved using cryo-electron microscopy and a structural comparison with the designed structure. c shows the results of the corresponding inactivation mutant functional rescue experiment. Detailed Implementation
[0022] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.
[0023] The accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0024] The flowchart shown in the attached diagram is merely an illustrative example and does not necessarily include all steps. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0025] This invention proposes a method for designing transmembrane regulatory proteins based on a cue strategy and a generative model. It mainly includes two stages, with the first stage comprising steps S1 to S2 below, achieving high-throughput target survey based on virtual probes, such as... Figure 1 As shown; the second stage includes the following steps S3 to S4, realizing closed-loop iterative design based on structural cues, such as... Figure 2 As shown.
[0026] like Figure 3 As shown, the specific implementation process of the present invention is as follows: S1, Probe Conformation Generation In this step, a generative diffusion model is used to generate multiple virtual probe backbones for the transmembrane domains of the target membrane protein, and a conformation set of the probe-target membrane protein complex is generated through a protein sequence design model and a structure prediction model.
[0027] In a specific implementation of the present invention, one optional implementation method is as follows: (1.1) Predetermine the hotspot region of the target membrane protein (e.g., GPCR receptor), that is, specify any possible target region covering the protein to be detected, such as the approximate groove of the transmembrane domain.
[0028] (1.2) Input the protein to be detected and any possible target region into the diffusion model. The diffusion model can be RFDiffusion. The diffusion model generates a large number of virtual probe backbones for the target region. In this embodiment, 300 virtual probe backbones are generated for one target region. The amino acid sequence length of the backbone is 30-150, which can be achieved by setting the parameters of the RFDiffusion model.
[0029] (1.3) The generated probe backbone is input into the protein sequence design model for sequence design. One probe backbone can generate a batch of design sequences. The generated design sequences (i.e., the sequences to be designed) and the target membrane protein sequences to be detected (referred to as receptor sequences) are directly input into the structure prediction model to predict the three-dimensional structure of the probe-target protein complex, obtaining a large number of binding conformations. In this embodiment, the protein sequence design model can be ProteinMPNN, which can be set to remove specific amino acids (such as removing alanine) to enhance transmembrane hydrophobic interactions, obtaining a total of 1500-10000 complex three-dimensional structures. The structure prediction model can be AlphaFold2, which can predict a batch of complex three-dimensional structures through the structure prediction model of one design sequence and the target membrane protein sequence to be detected.
[0030] S2, Determining the final target area based on probe distribution density analysis. In this step, the atomic coordinates of the probe in the conformation set of the probe-target membrane protein complex are extracted, a three-dimensional voxel grid is established, and the binding hot spots on the surface of the target membrane protein are identified by spatial density calculation. In a specific implementation of the present invention, one optional implementation method is as follows: (2.1) Extract the C-alpha atoms at the probe positions from the massive binding conformations generated in step (1.3) as conformation point cloud data.
[0031] (2.2) Establish a three-dimensional spatial mesh covering all coordinates in the structure file, divide the space into voxels of a specific size, and calculate the distribution density value of the point cloud in each voxel mesh.
[0032] (2.3) Assign the density attribute of the grid to the calculated point cloud and map the density value to the B-factor (temperature factor) field of the conformation's PDB structure file. Use molecular visualization software to generate a three-dimensional thermogram based on the B-factor value. In this embodiment, PyMOL software is used.
[0033] (2.4) Based on the high-density areas of the heatmap, determine the final target areas suitable for binding on the membrane protein surface. In the PyMOL software used in this embodiment, the high-density areas are displayed in red. Based on the red areas, TM1 / 2 / 4, TM3 / 4 / 5, and TM5 / 6 / 7 can be determined as the final target areas.
[0034] S3, Structural hints for construction This step is based on the final target region. According to the control target and the initial skeleton containing the preset target region, the initial input containing physical constraints is constructed.
[0035] In specific implementations of this invention, the following three strategies are included: Strategy A (Site-Guided Insertion): Designed for anchored proteins. This strategy uses flexible linkers (such as GS-Linker) to directly insert the designed sequence into a functionally unrelated location (such as the N-terminus) in the receptor sequence adjacent to the target region, serving as a single-stranded sequence input to the structure prediction model, thus limiting the search space.
[0036] Strategy B (Site Blocking Pre-occupancy): This strategy targets difficult-to-bind sites, such as off-target events that occur in closed-loop iterative design, causing the designed sequence to bind to non-target regions of the receptor sequence. This strategy introduces known binding proteins (such as the anchoring proteins obtained in Example 1) into the single-stranded sequence input to the structure prediction model, allowing them to occupy dominant binding sites in non-target regions (such as TM1 / 2 / 4), creating steric hindrance and forcing the model to search for non-dominant regions (such as TM3 / 4 / 5).
[0037] Strategy C (Conformation Induction): Design targeting the functional regulation (agonist / inhibitory) proteins of G protein-coupled receptors. This strategy introduces downstream effector protein sequences (such as the Gαs subunit of G proteins) into the single-stranded sequence input to the structure prediction model. That is, the single-stranded sequence and the downstream effector protein sequence are used as inputs to the structure prediction model, which operates in complex mode to induce the receptor to present a specific functional conformation (such as the activated state of TM6 outward shift), thereby screening for proteins that can stabilize this conformation.
[0038] S4, refolding iterative optimization In this step, the optimal backbone is screened through closed-loop iterative optimization of protein sequence design model and structure prediction model, and combined with global self-consistent RMSD screening to generate a high-precision design sequence, which is the final transmembrane regulatory protein sequence.
[0039] In a specific implementation of the present invention, one optional implementation method is as follows: (4.1) Based on the regulatory objective, select the target target region from the final target region, and input the target membrane protein and the selected target target region into the diffusion model. The diffusion model can be RFDiffusion. The diffusion model generates a batch of initial backbones for the target target region, and the initial backbones are used as the backbones to be optimized. In this embodiment, 300 initial backbones are set to be generated for the target target region, which can be achieved by setting the parameters of the RFDiffusion model.
[0040] (4.2) Input the backbone to be optimized into the protein sequence design model. Each backbone generates a batch of sequences to be designed. Insert the sequences to be designed into the receptor sequence using strategy A to obtain a batch of single-stranded sequences as input sequences. You can set the removal of specific amino acids (such as removing alanine) in the protein sequence design model to enhance transmembrane hydrophobic interactions.
[0041] It should be noted that if the regulatory target is the functional regulation of G protein-coupled receptors, strategy C is also required, which involves determining the downstream effector protein sequence of the receptor in addition to the single-stranded sequence, and using them together as the input sequence. Here, each input sequence corresponds to an initial skeleton.
[0042] (4.3) Input the batch input sequences into the structure prediction model to predict the full atomic structure and its confidence of the input sequence. Five full atomic structures are obtained for each sequence. The predicted skeleton part (referred to as the predicted skeleton) is extracted from the full atomic structure of the input sequence. The local confidence of the predicted skeleton (such as pLDDT) and the global self-consistent RMSD between the predicted skeleton and the corresponding initial skeleton are combined to select N best predicted skeletons. In this embodiment, pLDDT is given by the structure prediction model. The local confidence is selected from the designed sequence portion and the corresponding position of the linker sequence portion introduced by strategy A. The average pLDDT is calculated amino acid by amino acid.
[0043] The skeletons were selected according to the following rules: First, predicted skeletons with global scRMSD > 10 Å were removed. The remaining predicted skeletons were then sorted by local confidence level pLDDT, and N best predicted skeletons were selected.
[0044] In one specific embodiment of the present invention, if after several iterations, for example after one round of refolding iteration test, it is observed that the global self-consistency of the local confidence pLDDT ranking among the top N candidate predicted skeletons is low (the proportion of global scRMSD < 20 Å is less than 50%), it can be determined that there is a risk of off-target convergence (in subsequent rounds, screening by global scRMSD < 10 Å may not yield effective samples for ranking), strategy B can be introduced.
[0045] (4.4) The best predicted backbone is used as the new backbone to be optimized. Return to step (4.2) and repeat the backbone screening process until all predicted backbones meet the requirements (e.g., global scRMSD < 1.5 Å and pLDDT > 90). Since each design sequence generated by the protein sequence design model will predict a batch of full-atom structures (i.e., corresponding batch predicted backbones) in the structure prediction model, in the final screening process, the design sequences are sorted according to the average index of the batch predicted backbones from the same candidate design sequence, and the M best design sequences are output as the generated transmembrane regulatory protein sequences. These sequences have high structural conservation and high confidence and can bind to the target region of the receptor sequence.
[0046] To verify the effectiveness of the present invention, the following experiment was designed.
[0047] Example 1: Design of D1R-anchored exoskeleton proteins based on site-guided insertion This embodiment demonstrates how to design an anchored exoskeleton protein that targets the dopamine D1 receptor (D1R)™1 / 2 / 4 region.
[0048] Probe conformation generation: The approximate groove on the receptor surface was selected as the hot spot region, and 300 virtual probe backbones were generated using RFDiffusion. After protein sequence design model and structure prediction model, 1500 binding conformations were obtained.
[0049] The final target region was determined based on probe distribution density analysis: all binding conformations were superimposed, the conformation space was divided into a 2.0 Å voxel grid, the density was calculated, and mapped to the B-factor. Using PyMOL, significant high-density red regions were observed in three areas, indicating that the binding modes predicted by the prediction model were sampled at high frequency in these regions. It can be inferred that there are more easily generated binding modes (binding dominance sites) in these areas, and TM1 / 2 / 4 were identified as the target regions designed in this embodiment.
[0050] Structure suggestion construction: Using the "site-guided insertion" strategy A, the sequence to be designed is inserted into the N end of D1R and connected with a flexible GS-Linker of length 7.
[0051] Iterative optimization: ProteinMPNN is configured to remove alanine, and AlphaFold2 is configured to disable the template. For example... Figure 4 As shown, Figure 4 In step a, after 6 rounds of iteration, the final round converged to obtain a sequence with scRMSD < 1.5Å and pLDDT > 90. At the same time, the pLDDT variance decreased in each round, which can be judged as convergence. The converged sequence is highly conserved, and it can be seen that the transmembrane region is mainly composed of hydrophobic amino acids (mainly red and yellow) while the non-transmembrane region has a small number of polar amino acids (mainly blue and green). Figure 4 In section b, the results were verified: cryo-electron microscopy analysis showed that the designed protein bound tightly to the target site, and the calculated and resolved RMSD of the experimental structure was 1.0 Å, confirming the accuracy of the design.
[0052] Example 2: Design of bias modulators based on site blocking This embodiment addresses the problem that AI tends to combine dominant sites, causing the iteration to fail to converge to the target region.
[0053] Problem description: When attempting to design TM3 / 4 / 5 binding proteins, it was found that the model has a probability of generating proteins that bind to TM1 / 2 / 4, causing it to fail to converge to the expected sites.
[0054] Structural hints were used to construct the structure using "site-guided insertion" strategy A and "site-blocking pre-occupancy" strategy B. In the input data, the GEM_anchor obtained in Example 1 was fixed at positions TM1 / 2 / 4 to form spatial steric hindrance.
[0055] Iterative optimization: This forces the model to search in the TM3 / 4 / 5 region while optimizing the confidence level to obtain a highly conservative sequence sampling batch.
[0056] Result verification: such as Figure 5 As shown, GEM_BAM was successfully obtained. Figure 5 In step a, pLDDT increases in each round according to the mean curve of the amino acid sites. After 5 rounds of iteration, the final round converges to obtain a sequence with scRMSD < 1.5Å and pLDDT > 90. At the same time, the variance of pLDDT decreases in each round, which can be judged as convergence. The converged sequence has high conservation, and it can be seen that the transmembrane region is mainly composed of hydrophobic amino acids (mainly red and yellow) while the non-transmembrane region has a small number of polar amino acids (mainly blue and green). Figure 5 In b, iterative cryo-electron microscopy analysis showed that the target site could be bound even after removing the occupant sequence. The calculated and analyzed RMSD of the experimental structure was 2.7 Å, confirming the accuracy of the design.
[0057] Example 3: Conformation-Induced Agonist Design This example demonstrates how to design ago-PAMs that can activate receptors.
[0058] Structural hints were used to construct the structure using a site-guided insertion strategy (Strategy A) and a conformation-inducible strategy (Strategy C). The input contained the complex sequence of D1R with the downstream G protein Gαs subunit to induce the receptor to present an activated conformation with TM6 outward shift.
[0059] Iterative optimization: Screening for "pincer"-like proteins that can stabilize the activation conformation.
[0060] Result verification: such as Figure 6 As shown, GEM_ago-PAM is obtained. Figure 6 In step a, pLDDT increases in each round according to the mean curve of the amino acid sites. After 5 rounds of iteration, the final round converges to obtain a sequence with scRMSD < 1.5Å and pLDDT > 90. At the same time, the variance of pLDDT decreases in each round, which can be judged as convergence. The converged sequence has high conservation, and it can be seen that the transmembrane region is mainly composed of hydrophobic amino acids (mainly red and yellow) while the non-transmembrane region has a small number of polar amino acids (mainly blue and green). Figure 6 In b, iterative cryo-electron microscopy analysis showed that the protein could stabilize the receptor activation state in the absence of endogenous ligands. The calculated and analyzed RMSD of the experimental structure was 0.9 Å, confirming the accuracy of the design. Cellular function experiments verified the successful restoration of the activity of loss-of-function mutants such as T371K.
[0061] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for designing transmembrane regulatory proteins based on cue strategies and generative models, characterized in that, This includes the following two phases, executed sequentially: Target survey phase: High-throughput automatic identification of binding hotspot regions on the surface of target membrane proteins based on virtual probes; Closed-loop design phase: Based on the aforementioned hotspot regions, an initial backbone set is first generated using a generative diffusion model. Then, a structure cueing strategy guides the protein sequence design model and structure prediction model for iterative optimization to generate transmembrane regulatory protein sequences.
2. The transmembrane regulatory protein design method based on cueing strategy and generative model according to claim 1, characterized in that, The target survey phase includes: A generative diffusion model was used to generate multiple virtual probe backbones targeting the membrane protein. Based on protein sequence design models and structure prediction models, a set of conformations of probe-target membrane protein complexes is obtained; The binding hotspot region is determined by analyzing the spatial distribution density of probes in the complex conformation set.
3. The transmembrane regulatory protein design method based on cueing strategies and generative models according to claim 2, characterized in that, The process of determining binding hotspot regions by analyzing the spatial distribution density of probes includes: The C-alpha atom coordinates of the probes in the complex conformation set are extracted to form conformation point cloud data; A three-dimensional voxel grid covering a predetermined hotspot region of the target membrane protein is constructed, wherein the size of the three-dimensional voxel grid is 2.0 Å; Calculate the distribution density value of the conformation point cloud data within the three-dimensional voxel grid; A three-dimensional heat map is generated based on the distribution density value, and the high-density areas in the heat map are identified as the combined hotspot areas.
4. The transmembrane regulatory protein design method based on cueing strategy and generative model according to claim 1, characterized in that, The closed-loop design phase includes: Based on the control objectives, a structural prompt containing physical constraints is constructed based on the aforementioned combined hotspot areas; Based on the initial backbone set generated by the structural hints and the generative diffusion model, the protein sequence design model and the structure prediction model are used for closed-loop iterative optimization to screen out transmembrane regulatory protein sequences that meet the preset structural accuracy index.
5. The transmembrane regulatory protein design method based on cueing strategies and generative models according to claim 4, characterized in that, Structural cues include: Site-guided insertion strategy: The designed sequence is inserted into a functionally unrelated position in the target membrane protein sequence that is adjacent to the binding hotspot region using a flexible linker, resulting in a single-stranded input sequence; Site blocking pre-occupancy strategy: Introduce a known binding protein into the single-stranded input sequence to occupy the non-target binding region of the target membrane protein, thus forming steric hindrance; Conformation induction strategy: A downstream effector protein sequence of the target membrane protein is introduced into the single-stranded input sequence to form a complex input sequence representing a specific functional state, thereby inducing the target membrane protein to present a specific functional conformation.
6. The transmembrane regulatory protein design method based on cueing strategies and generative models according to claim 5, characterized in that, The closed-loop iterative optimization includes the following steps: a) Determine the control target and the target target region selected from the combined hotspot region, and use a generative diffusion model to generate an initial skeleton set containing a batch of initial skeletons for the target target region, the initial skeletons being used as skeletons to be optimized; b) Use the protein sequence design model to generate a batch of design sequences for each backbone to be optimized; c) Insert the designed sequence into the target membrane protein sequence using a site-directed insertion strategy to obtain a batch of single-stranded sequences as input sequences; if the regulatory target is the functional regulation of G protein-coupled receptors, a conformation induction strategy is also required to introduce downstream effector protein sequences of the target membrane protein into the single-stranded sequence as input sequences. Each input sequence corresponds to an initial skeleton; d) Utilize the structural prediction model to predict the full atomic structure and its confidence level of the input sequence, extract the predicted backbone part and its corresponding local confidence level from it, and batch predict the backbone for the same design sequence; e) Based on the local confidence index and global self-consistency index of the predicted skeleton, select N best predicted skeletons from the results of step d). f) Select the N best predicted skeletons selected in this iteration as the skeletons to be optimized in the next iteration, and repeat steps b) to e) until all predicted skeletons meet the preset structure accuracy index. g) Output the final transmembrane regulatory protein sequence that meets the criteria.
7. The transmembrane regulatory protein design method based on cueing strategies and generative models according to claim 6, characterized in that, The local confidence index is the predicted local distance difference test pLDDT, and the global self-consistency index is the global self-consistency root mean square deviation scRMSD between the predicted skeleton and the corresponding initial skeleton; the preset structural accuracy index is that the global self-consistency root mean square deviation scRMSD is less than 1.5 Å and pLDDT is greater than 90.
8. The transmembrane regulatory protein design method based on cueing strategies and generative models according to claim 6, characterized in that, The generative diffusion model is the RFDiffusion model, the protein sequence design model is the ProteinMPNN model, and the structure prediction model is the AlphaFold2 model.
9. The transmembrane regulatory protein design method based on cueing strategies and generative models according to claim 6, characterized in that, In step e), N best predicted skeletons are selected from the results of step d), specifically those that are: calculated global self-consistency root mean square deviation (scRMSD) from the initial skeleton < 10 Å, and ranked in the top N of the local confidence index. In g), the design sequences are sorted according to the average index of the batch predicted backbone from the same design sequence, and several best design sequences are output as the generated transmembrane regulatory protein sequences.
10. The transmembrane regulatory protein design method based on cueing strategies and generative models according to claim 6, characterized in that, After one round of refolding iteration test, if less than 50% of the candidate predicted skeletons ranked in the top N with local confidence have a global self-consistency of less than 20 Å with the initial skeleton, a site blocking pre-occupancy strategy is introduced.