Model-based polypeptide design method, model training method, apparatus and device
The model-based polypeptide design method accurately integrates unnatural amino acids by leveraging multimodal features and rationality detection, addressing structural compatibility issues and enhancing drug development efficacy.
Patent Information
- Application Number
- JP2025176797
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-06-03
- Filing Date
- 2025-10-20
- Publication Date
- 2026-02-03
AI Technical Summary
Conventional polypeptide design methods struggle to accurately incorporate unnatural amino acids due to their complex structural compatibility, leading to issues such as poor membrane permeability and metabolic stability, limiting their effectiveness in drug development.
A model-based approach that utilizes a pre-trained polypeptide design model to predict and design unnatural amino acids at specific positions in polypeptides by integrating multimodal features like main chain orientation, side chain type, and rigid atom group distribution, combined with rationality detection to ensure structural accuracy.
Enhances the accuracy and efficiency of designing polypeptides with unnatural amino acids, improving their affinity, stability, and druggable properties, making them suitable for clinical research and development.
Smart Images

Figure 2026016511000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of computer technology, particularly to the field of artificial intelligence such as biocomputing, and in particular to a model-based polypeptide design method, model training method, apparatus, and device. [Background technology]
[0002] Peptide design with unnatural amino acids is a cutting-edge technology that overcomes the structural limitations of natural amino acids by integrating artificially synthesized or chemically modified unnatural amino acids into polypeptide chains, thereby enhancing their pharmacological properties and functional diversity.
[0003] As bioactive molecules consisting of 3 to 30 amino acids, polypeptides play important roles in life processes such as signal transduction and immune regulation by specifically recognizing target protein receptors. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides model-based methods for polypeptide design, model training methods, apparatus, and devices. [Means for solving the problem]
[0005] According to one aspect of the present disclosure, there is provided a model-based method for designing a polypeptide, the method including: acquiring a pocket of a target protein and a target polypeptide in which preset positions of a to-be-designed unnatural amino acid are labeled, the pocket of the target protein binds to the target polypeptide via the preset positions; acquiring features of the pocket of the target protein and multimodal features of each known amino acid in the target polypeptide, the multimodal features of each known amino acid including a backbone orientation feature, a backbone rotation feature, a side chain type feature, and a rigid atom group distribution feature; and designing the unnatural amino acid at the preset position in the target polypeptide using a pre-trained polypeptide design model based on the features of the pocket of the target protein and the multimodal features of each known amino acid in the target polypeptide.
[0006] According to another aspect of the present disclosure, there is provided a method for training a polypeptide design model, the method including: obtaining a pocket of a training target protein and a training polypeptide, the training polypeptide binding to the training target protein via an unnatural amino acid at a preset position; obtaining features of the pocket of the training target protein and first multimodal features of amino acids at each position in the training polypeptide, the first multimodal features of the amino acids at each position including a backbone orientation feature, a backbone rotation feature, a side chain type feature, and a rigid atom group distribution feature; and training the polypeptide design model based on the features of the pocket of the training target protein and the first multimodal features of the amino acids at each position in the training polypeptide.
[0007] According to yet another aspect of the present disclosure, there is provided a model-based polypeptide design apparatus, the model-based polypeptide design apparatus comprising: an information acquisition module that acquires a pocket of a target protein and a target polypeptide, wherein the target polypeptide is labeled with preset positions of an unnatural amino acid to be designed, and the pocket of the target protein binds to the target polypeptide via the preset positions; a feature acquisition module that acquires features of the pocket of the target protein and multimodal features of each known amino acid in the target polypeptide, wherein the multimodal features of each known amino acid include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features; and a design module that designs the unnatural amino acid at the preset position in the target polypeptide using a pre-trained polypeptide design model, based on the features of the pocket of the target protein and the multimodal features of each known amino acid in the target polypeptide.
[0008] According to yet another aspect of the present disclosure, there is provided an apparatus for training a polypeptide design model, the apparatus for training a polypeptide design model comprising: an information acquisition module for acquiring a pocket of a training target protein and a training polypeptide, wherein the training polypeptide binds to the training target protein via an unnatural amino acid at a preset position; a feature acquisition module for acquiring features of the pocket of the training target protein and first multimodal features of amino acids at each position in the training polypeptide, wherein the first multimodal features of the amino acids at each position include a main chain orientation feature, a main chain rotation feature, a side chain type feature, and a rigid atom group distribution feature; and a training module for training the polypeptide design model based on the features of the pocket of the training target protein and the first multimodal features of the amino acids at each position in the training polypeptide.
[0009] According to yet another aspect of the present disclosure, there is provided an electronic device comprising at least one processor and a memory communicatively connected to the at least one processor, wherein commands executable by the at least one processor are stored in the memory, and when the commands are executed by the at least one processor, the at least one processor can perform any of the above-described aspects and possible methods.
[0010] According to yet another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform any of the above-described aspects and enabling methods.
[0011] According to yet another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements any of the above-described aspects and possible implementations.
[0012] The techniques of the present disclosure can effectively improve the accuracy of unnatural amino acids at preset positions in designed target polypeptides.
[0013] It should be understood that the contents described in this section are not intended to identify key or essential features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following specification. [Brief explanation of the drawings]
[0014] The accompanying drawings are provided for a better understanding of the present embodiments and are not intended to limit the present disclosure. [Figure 1] 1 is a schematic diagram according to a first embodiment of the present disclosure; [Figure 2] FIG. 10 is a schematic diagram according to a second embodiment of the present disclosure. [Figure 3] 1 is a schematic diagram of the principle of the polypeptide design model provided by the present disclosure. [Figure 4] FIG. 10 is a schematic diagram illustrating a third embodiment of the present disclosure. [Figure 5] FIG. 10 is a schematic diagram illustrating a fourth embodiment of the present disclosure. [Figure 6] FIG. 10 is a schematic diagram illustrating a fifth embodiment of the present disclosure. [Figure 7] FIG. 10 is a schematic diagram illustrating a sixth embodiment of the present disclosure. [Figure 8] FIG. 10 is a schematic diagram illustrating a seventh embodiment of the present disclosure. [Figure 9] FIG. 10 is a schematic diagram illustrating an eighth embodiment of the present disclosure. [Figure 10] FIG. 13 is a schematic diagram illustrating a ninth embodiment of the present disclosure. [Figure 11] FIG. 1 is a block diagram of an electronic device for implementing the method of an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, exemplary embodiments of the present application will be described based on the drawings. For ease of understanding, various details of the embodiments of the present application are included and should be considered as merely examples. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of brevity, the following description will omit descriptions of well-known functions and structures.
[0016] It is clear that the described embodiments are only a part of the embodiments of the present disclosure, but not all of the embodiments, and all other embodiments that can be obtained by a person skilled in the art based on the embodiments of the present disclosure without performing creative work fall within the scope of protection of the present disclosure.
[0017] It should be noted that terminal devices according to embodiments of the present disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers, etc. Display devices include, but are not limited to, devices with display capabilities such as personal computers and televisions.
[0018] Furthermore, the term "and / or" in this specification simply describes a relationship between related objects and means that three relationships can exist. For example, A and / or B can mean three situations: A exists alone, A and B exist simultaneously, and B exists alone. Also, the character " / " in this specification generally means that the related objects before and after it are in an "or" relationship.
[0019] In practical applications, polypeptide drugs offer unique advantages in the hit-to-lead stage of drug development. However, natural polypeptides generally suffer from bottlenecks in drug efficacy, such as poor membrane permeability and metabolic stability. The introduction of unnatural amino acids offers a systematic solution to these problems. The use of cyclic topologies enhances α-helical rigidity, hydrophobic side chain modifications extend plasma half-life, and functional groups enhance target binding specificity, enabling precise control of molecular structure and biological function.
[0020] Conventional techniques have been unable to accurately and effectively design polypeptides because the chemical diversity of unnatural amino acids, such as chiral centers and non-standard functional groups, makes their structural compatibility more complex than that of natural amino acids. Conventional polypeptide design methods can only capture the contextual molecular characteristics of each amino acid in the amino acid sequence of a polypeptide.
[0021] FIG. 1 is a schematic diagram of the first embodiment of the present disclosure. As shown in FIG. 1, this embodiment provides a model-based method for designing a polypeptide, which may specifically include the following steps:
[0022] In S101, a pocket of a target protein and a target polypeptide are obtained, and a preset position of an unnatural amino acid to be designed is labeled in the target polypeptide.
[0023] Specifically, the target polypeptide is a polypeptide designed in this embodiment, and the target protein is a target protein whose binding ability to the target polypeptide is examined in the scenario of this embodiment. The pocket of the target protein binds to the target polypeptide via a preset position.
[0024] Specifically, the pocket of the obtained target protein may specifically include the structure of the pocket portion of the target protein, and accordingly, the obtained target polypeptide further includes amino acid structures at positions other than the preset positions on the polypeptide chain.
[0025] In S102, the pocket features of the target protein and the multimodal features of each known amino acid in the target polypeptide are obtained, and the multimodal features of each known amino acid include main chain direction features, main chain rotation features, side chain type features, and rigid atom group distribution features.
[0026] In particular, the pocket characteristics of the target protein may be digitized characteristics of the structure of the pocket of the target protein using an appropriate molecular description language, in particular in the form of a vector.
[0027] In this embodiment, the multimodality characteristics of each known amino acid in the obtained target polypeptide are exemplified by structural characteristics of four modalities: main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics. In actual applications, the multimodality characteristics of each known amino acid may include other secondary structural characteristics or other tertiary structural characteristics of the target polypeptide, which are not described here. For example, secondary structural characteristics of the target polypeptide may include characteristics related to different fragment folding types in the polypeptide sequence, such as alpha helix and beta fold. Tertiary structural characteristics of the target polypeptide include characteristics related to the three-dimensional coordinates of atoms in the amino acid in the polypeptide sequence.
[0028] Here, the backbone orientation feature, also referred to as the backbone 0 orientation feature, may be in the form of a vector characterizing the spatial position of a residue relative to the global coordinate direction of the target polypeptide in Cartesian coordinates with the C-A (i.e., the C at the α position on the current amino acid) atom as the reference point. The backbone rotation feature, also referred to as the backbone rotation matrix, can be specifically characterized by the rotation parameters of the local coordinate system constructed by the N-C-A-C atoms of the current amino acid relative to the global coordinate system of the target polypeptide, describing the spatial orientation of the residue. The side chain type feature characterizes the side chain type of the current amino acid and is a discrete classification feature, including natural and unnatural amino acids. The rigid atom group distribution feature, specifically, divides the current amino acid side chain into several rigid atom groups. Within each rigid atom group, atoms related to the motion of the current amino acid side chain are clustered. Meanwhile, in a real scenario, atoms related to the motion are considered to be atoms located on the same plane in the structure of the current amino acid.
[0029] The multimodal features of each known amino acid in a target polypeptide obtained in this embodiment can accurately and efficiently characterize the detailed structural information of each amino acid in the target polypeptide, providing effective support for the design of unnatural amino acids at preset positions in the target polypeptide.
[0030] In S103, unnatural amino acids are designed at preset positions in the target polypeptide using a pre-trained polypeptide design model based on the pocket characteristics of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide.
[0031] Specifically, in application, the pocket characteristics of a target protein and the multimodal characteristics of each known amino acid in the target polypeptide can be input into a pre-trained polypeptide design model, which can accurately and efficiently predict unnatural amino acids at preset positions based on the input information, thereby enabling accurate and efficient design of the target polypeptide.
[0032] The model-based polypeptide design method of this embodiment uses a pre-trained polypeptide design model based on the pocket features of a target protein and the multimodal features of each known amino acid in the target polypeptide to accurately and efficiently design unnatural amino acids at preset positions in the target polypeptide, thereby achieving the design of the target polypeptide. Because the multimodal features of each known amino acid in the target polypeptide contain detailed tertiary structure information for each amino acid in the target polypeptide, the accuracy of the unnatural amino acids at preset positions in the designed target polypeptide can be effectively improved.
[0033] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure. The model-based polypeptide design method of this embodiment will be described in more detail in addition to the technical solution of the embodiment shown in Figure 1. As shown in Figure 2, the model-based polypeptide design method of this embodiment may specifically include the following steps:
[0034] In S201, a pocket of a target protein and a target polypeptide are obtained, and the target polypeptide is labeled with a preset position of an unnatural amino acid to be designed.
[0035] In this embodiment, the amino acids at positions other than the preset positions in the target polypeptide may be natural amino acids or approved unnatural amino acids. For example, currently, there are 20 common natural amino acids, and the target may include 160 approved unnatural amino acids. Of course, with the development of medical science and technology, the number of approved unnatural amino acids will continue to increase.
[0036] In S202, the pocket features of the target protein and the multimodal features of each known amino acid in the target polypeptide are obtained, and the multimodal features of each known amino acid include main chain direction features, main chain rotation features, side chain type features, and rigid atom group distribution features.
[0037] When obtaining the rigid atom group distribution feature of this embodiment, the current amino acid side chain is divided into several rigid atom groups based on the dependency of each atom in the current amino acid side chain on the dihedral angle. Specifically, when dividing the side chain, the interlocking behavior of the side chain atoms during dihedral angle rotation is analyzed, and atoms involved in the motion are clustered into the same rigid unit, thereby establishing a quantitatively characterizable rigid atom group distribution. Here, atoms involved in the motion are considered to be atoms located on a single plane. For example, in a specific implementation, each atom on the side chain is sequentially analyzed using a traversal method in order of proximity to the main chain, and atoms located on a single plane, i.e., atoms involved in the motion, are divided into the same rigid unit, thereby achieving division of the rigid atom group corresponding to the current amino acid. By employing this method, the corresponding rigid atom group distribution feature can be obtained accurately and efficiently even for unnatural amino acids with complex side chains.
[0038] The side chain type feature in this embodiment is a discretized feature. As described in the above embodiment, the side chain type may be any of the 20 naturally occurring amino acids or any of the 160 recognized non-natural amino acids. The side chain type feature may be characterized using the identifier of the corresponding type of amino acid, and the identifier of each amino acid is used to uniquely characterize the corresponding amino acid.
[0039] In S203, a polypeptide design model is used to predict the multimodal characteristics of unnatural amino acids at preset positions in the target polypeptide based on the pocket characteristics of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide.
[0040] For example, when this step is specifically realized, the multimodal features of the unnatural amino acids at preset positions in the target polypeptide can be predicted using a polypeptide design model based on the pocket features of the target protein and the multimodal features of each known amino acid in the target polypeptide, with reference to the initial features of the unnatural amino acids at preset positions previously obtained, thereby enabling accurate and efficient prediction of the multimodal features of the unnatural amino acids at preset positions.
[0041] Specifically, before this step is concretely realized, obtaining initial main chain orientation features of unnatural amino acids at preset positions in a target polypeptide based on a Gaussian distribution such as JPEG2026016511000002.jpg716; obtaining initial main-chain rotation characteristics of unnatural amino acids at preset positions in a target polypeptide based on a uniform distribution such as JPEG2026016511000003.jpg719; obtaining an initial side chain type characteristic of an unnatural amino acid at a preset position in a target polypeptide based on a Gaussian distribution such as JPEG2026016511000004.jpg719; The method may further include obtaining an initial distribution profile of rigid atom groups of unnatural amino acids at preset positions in a target polypeptide based on a uniform distribution such as JPEG2026016511000005.jpg720.
[0042] The polypeptide design model of this embodiment may be a conditional flow matching design model that uses initial features of unnatural amino acids at preset positions in a target polypeptide as noise input. The polypeptide design model reduces and recovers the input noise based on the input pocket features of the target protein and multimodal features of known amino acids at positions other than the preset positions in the target polypeptide, and obtains multimodal features of the unnatural amino acids at the preset positions in the target polypeptide by decoding. For example, the multimodal features may include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features of the unnatural amino acids at the preset positions.
[0043] The design principle of the polypeptide design model of this embodiment is considered to be to use a conditional flow matching method to realize prediction of the target distribution from the initial distribution of four types of modalities of unnatural amino acids at preset positions. Here, the initial distribution of main chain direction features, i.e., the main chain direction initial features, is a Gaussian distribution. The initial distribution of the main chain rotation feature is uniformly distributed. The initial distribution of the side chain type features is Gaussian distribution. The initial distribution of the rigid atom group distribution features is uniform. The proposed method matches JPEG2026016511000009.jpg1027. By adopting the above method, the initial features of the unnatural amino acids at the preset positions are determined to be highly reasonable and accurate, providing an effective basis for accurately predicting the multimodal features of the unnatural amino acids at the preset positions in the subsequent noise reduction decoding process.
[0044] FIG. 3 is a schematic diagram illustrating the principle of the polypeptide design model provided by the present disclosure. As shown in FIG. 3, the polypeptide design model may include a pre-processing module, a node encoder, an edge encoder, a fusion processing module, and a noise reduction processing module. The node encoder and edge encoder can extract features between a single amino acid and adjacent amino acids, respectively. For example, the pre-processing module may process input information to obtain the information required for the node encoder and edge encoder, and then input the information to the node encoder and edge encoder, respectively. For each amino acid, the node encoder may generate a residue-level node embedding by fusing the amino acid type feature, local atomic coordinate feature, and main-chain dihedral angle feature. Here, the pre-processing module obtains the local atomic coordinates of the amino acid from a pre-collected amino acid information base based on the type of the current amino acid, and then further obtains the local atomic coordinate feature of the amino acid. This amino acid information base may include information on all natural amino acids and all authenticated unnatural amino acids, and may include, for example, local coordinate information for each atom of the current amino acid in Cartesian coordinates with the CA atom as the reference point for each amino acid. For each amino acid, the preprocessing module may further obtain main-chain dihedral angle information for the current amino acid based on the local coordinate information for the current amino acid. Here, the information on amino acids at preset positions may be subjected to corresponding processing based on the input initial characteristics of the unnatural amino acids at the preset positions.
[0045] Specifically, in the Node Encoder, amino acid types are encoded in the embedding layer, local atomic coordinates are transformed into the local coordinate system, and then expanded into high-dimensional features for each amino acid type, and main-chain dihedral angles are encoded as angles. After the three types of features are joined, they are fused using a multilayer perceptron (MLP). Information leakage is controlled through multiple masks, such as sequence mask, structure mask, and residue mask, and finally a feature vector for each residue is output, effectively integrating protein sequence and structure information.
[0046] Edge encoding focuses on modeling the relationships between residues. Pair embeddings represent type interactions for amino acid pairs, while relative position embeddings encode sequence spacing. After Gaussian transformation of interatomic distances, distance embeddings are generated using MLPs and combined with learnable distance coefficients to achieve dynamic weighting. At the same time, angle encoding captures the spatial orientation relationships of paired residues. All features are embedded using the output edges of the fusion network, fully characterizing the structural proximity, sequence correlation, and spatial interactions between residues. Edges between adjacent amino acids in a polypeptide can be pre-determined or pre-determined based on a detection policy in a preprocessing module. For example, an edge is considered to exist if the distance between the closest atom pairs of two amino acids is less than 4 angstroms. Similarly, edges between a pocket in a target protein and an unnatural amino acid at a preset position in the target polypeptide can be detected in the preprocessing module. There may be one, two, or multiple edges between the pocket of the intended target protein and the unnatural amino acid at the preset position.
[0047] As shown in Figure 3, the fusion processing module fuses all feature information encoded by the Node Encoder and Edge Encoder, and the noise reduction processing module performs noise reduction decoding based on the fused information to recover and output multimodal features of unnatural amino acids at preset positions.
[0048] According to the operating principle of the above polypeptide design model, it can be seen that this polypeptide design model is a conditional flow matching model implemented based on a graph neural network, in which the initial features of the unnatural amino acid at the preset position of the target polypeptide are input into the polypeptide design model as noise, and the noise is decoded and restored based on the other input features to obtain the multimodal features of the unnatural amino acid at the preset position.
[0049] In S204, the structure of the unnatural amino acid at the preset position in the target polypeptide is determined based on the multimodal characteristics of the unnatural amino acid.
[0050] Specifically, based on the multimodal characteristics of the unnatural amino acid at the preset position predicted by the polypeptide design model, in combination with information such as the bond length and bond angle of the chemical molecule, the coordinates of each atom in the unnatural amino acid at the preset position can be determined, and the structure of the unnatural amino acid can be obtained.
[0051] In S205, rationality detection is performed based on the structure of the pocket of the target protein, the structure of each known amino acid in the target polypeptide, and the structure of the unnatural amino acid at the preset position.
[0052] In S206, in response to the detected irrationality, the unnatural amino acid designed into the preset position is discarded.
[0053] The rationality detection in this embodiment can be considered as atomic collision detection. For example, by knowing the pocket structure of the target protein, the structures of each known amino acid in the designed target polypeptide, and the structure of the unnatural amino acid at the preset position, the coordinates of each atom in the pocket structure of the target protein and the coordinates of each atom in the structure of the amino acid at each position in the target polypeptide can be determined. The presence or absence of atomic collisions can then be further detected. Specifically, this means detecting the presence or absence of collisions between each atom in the unnatural amino acid at the designed preset position and each atom in the pocket structure and each atom in the amino acid at another position in the target polypeptide. If a collision exists, the unnatural amino acid designed at the preset position in the target polypeptide is considered to be irrational and can be discarded. If it is detected as rational, the unnatural amino acid designed at the preset position is retained, and further target polypeptide design is realized.
[0054] The model-based polypeptide design method of this embodiment uses a polypeptide design model to predict the multimodal features of unnatural amino acids at preset positions in the target polypeptide based on the pocket features of the target protein and the multimodal features of each known amino acid in the target polypeptide, and then determines the structure of the unnatural amino acid at the preset position in the target polypeptide. Because the multimodal features contain rich internal structural information of amino acids, the technical solution of this embodiment can accurately and efficiently design unnatural amino acids in polypeptides, and can also accurately and efficiently design polypeptides.
[0055] Furthermore, according to the technical solution of this embodiment, rationality detection can be performed based on the pocket structure of the target protein, the structure of each known amino acid in the target polypeptide, and the structure of unnatural amino acids at preset positions, making it possible to easily eliminate irrationally designed polypeptides and ensure the accuracy and effectiveness of polypeptides in subsequent research and development and testing stages.
[0056] Through testing and verification, the polypeptide designed using the technical solution of this example can effectively improve the affinity, and at the same time, improve the druggable properties, such as stability and half-life, making it suitable for the next stage of clinical research and development.
[0057] The technical solution of this embodiment, which is a polypeptide design containing engineered unnatural amino acids, can be widely applied in fields such as G protein-coupled receptor (GPCR) targeting drug development, anti-infection, and antibacterial treatment.
[0058] 4 is a schematic diagram according to the third embodiment of the present disclosure. As shown in FIG. 4, the method for training a polypeptide design model according to this embodiment may specifically include the following steps:
[0059] In S401, a pocket of a training target protein and a training polypeptide that binds to the training target protein via an unnatural amino acid at a preset position are obtained.
[0060] It should be noted that the training polypeptide may also include multiple amino acids at other positions.
[0061] In S402, a pocket feature of the training target protein and a first multimodal feature of an amino acid at each position in the training polypeptide are obtained, and the first multimodal feature of an amino acid at each position includes a main chain direction feature, a main chain rotation feature, a side chain type feature, and a rigid atom group distribution feature.
[0062] The specific implementation method of this step can be referred to the related description of the embodiment shown in FIG. 1 or FIG. 2 above, and will not be described in detail here.
[0063] In S403, a polypeptide design model is trained based on the pocket features of the training target protein and the first multimodal features of the amino acids at each position in the training polypeptide.
[0064] The method for training a polypeptide design model of this embodiment trains the polypeptide design model based on the pocket features of the training target protein and the first multimodal features of the amino acids at each position in the training polypeptide. Because the multimodal features contain rich internal structural information of the amino acids, the technical solution of this embodiment can effectively improve the accuracy of the trained polypeptide design model.
[0065] Figure 5 is a schematic diagram of a fourth embodiment of the present disclosure. The training method for a polypeptide design model of this embodiment is in addition to the technical solution of the embodiment shown in Figure 4 above, and further describes the technical solution of the present disclosure in more detail. As shown in Figure 5, the training method for a polypeptide design model of this embodiment may specifically include the following steps:
[0066] In S501, a pocket of a training target protein and a training polypeptide that binds to the training target protein via an unnatural amino acid at a preset position are obtained.
[0067] In S502, pocket features of the training target protein and first multimodal features of amino acids at each position of the training polypeptide are obtained, and the first multimodal features of amino acids at each position include main chain direction features, main chain rotation features, side chain type features, and rigid atom group distribution features.
[0068] In concrete implementation, step S102 in the embodiment shown in FIG. 1 or step S202 in the embodiment shown in FIG. 2 can be referred to, and therefore, detailed description will not be given here.
[0069] In S503, a polypeptide design model is used to predict a second multimodal feature of an unnatural amino acid at a preset position in the training polypeptide based on the pocket feature of the training target protein and the first multimodal feature of an amino acid at each position other than the preset position in the training polypeptide.
[0070] The specific embodiment is the same as the embodiment of step S203 in the embodiment shown in Figure 2. For details, please refer to the relevant description of the above embodiment, and therefore, detailed description will not be given here.
[0071] In S504, a first loss function is constructed based on the second multimodal features and the first multimodal features of the unnatural amino acids at the preset positions.
[0072] For example, when this step is specifically implemented, it may include the following steps:
[0073] (a1) Based on the second multimodal feature and the first multimodal feature of the unnatural amino acid at the preset position, a main chain orientation feature difference, a main chain rotation feature difference, and a side chain type feature difference are obtained, respectively.
[0074] Here, the main chain orientation feature difference may be the difference between the main chain orientation feature of the unnatural amino acid at that preset position in the second multimodal feature and the main chain orientation feature of the unnatural amino acid at that preset position in the first multimodal feature, and may be expressed, for example, as a vector difference.
[0075] Accordingly, the main chain selection feature difference is similar to the main chain direction feature difference value taking scheme.
[0076] The side chain type feature difference can be considered to be the probability that the type of amino acid at that preset position in the first multimodal feature is an unnatural amino acid, since the first multimodal feature is a multimodal feature corresponding to the training polypeptide. The second multimodal feature is a multimodal feature of an unnatural amino acid at that preset position predicted against a mask of unnatural amino acids at the preset position. In this case, the probability that the type of amino acid at that preset position in the second multimodal feature is an unnatural amino acid can be recorded as p, where p is greater than 0 and less than 1. In this case, the side chain type feature difference can be equal to 1-p.
[0077] (b1) A first loss function is constructed based on the main chain direction characteristic difference, the main chain rotation characteristic difference, and the side chain type characteristic difference.
[0078] Specifically, the first loss function may be a weighted sum of the main chain orientation characteristic difference, the main chain rotation characteristic difference, and the side chain type characteristic difference. The weights corresponding to each characteristic can be set based on experience and practical needs, but are not limited thereto.
[0079] In S505, parameters of the polypeptide design model are adjusted with the aim of convergence of the first loss function.
[0080] In this embodiment, a single training data set including a training target protein pocket and a training polypeptide is used to perform supervised training of a polypeptide design model. In practical applications, a polypeptide design model can be trained using multiple similar training data sets in the above-described training method until the number of training rounds reaches a preset threshold, or until the first loss function consistently converges after multiple consecutive training rounds and the training is completed. In this case, a polypeptide design model can be obtained.
[0081] The polypeptide design model training method of this embodiment uses an authenticated training polypeptide as a positive sample and trains the polypeptide design model through supervised training, thereby effectively improving the accuracy of the trained polypeptide design model.
[0082] Figure 6 is a schematic diagram of the fifth embodiment of the present disclosure. In addition to the technical solution of the embodiment shown in Figure 4 above, the training method for a polypeptide design model of this embodiment further describes the technical solution of the present disclosure in more detail. As shown in Figure 6, the training method for a polypeptide design model of this embodiment may specifically include the following steps:
[0083] In S601, a pocket of a training target protein and a training polypeptide that binds to the training target protein via an unnatural amino acid at a preset position are obtained.
[0084] In S602, a replacement amino acid at a preset position in the training polypeptide is obtained.
[0085] For example, an unnatural amino acid at a preset position in a training polypeptide is reverse-modified to obtain a replacement amino acid, or a target amino acid that is least similar to the unnatural amino acid at the preset position is obtained as a replacement amino acid from a pre-constructed amino acid similarity matrix.
[0086] For example, by performing reverse modification on the side chain conformation of a positive sample of a verified polypeptide sequence, such as mirror-inverting the direction of the charged group, a negative sample with specific differences in physicochemical properties can be generated, effectively resolving the lack of experimental data for unnatural amino acids.
[0087] In this embodiment, the number of rows and columns in the pre-constructed amino acid similarity matrix may be equal to the sum of the number of natural amino acids and the number of authenticated non-natural amino acids. Each row and each column corresponds to one amino acid. The numerical value at each position in a matrix such as (i, j) is equal to the similarity between the two amino acids in the row and column corresponding to that position. The similarity between two amino acids is equal to the weighted sum of the structural similarity and the pharmacokinetic similarity of the two amino acids. Here, the structural similarity and the pharmacokinetic similarity can be calculated using known testing methods or a pre-trained neural network model.
[0088] The replacement amino acid selected in this embodiment is used to generate negative samples, so it is first positioned in the row corresponding to the unnatural amino acid at the preset position in the amino acid similarity matrix, and then the target amino acid corresponding to the column with the least similarity is obtained as the replacement amino acid.
[0089] In this embodiment, by doing as described above, side chain substitution can be performed while the main chain structure is kept stable, and therefore accurate substituted amino acids can be generated.
[0090] In S603, negative samples of the training polypeptide are generated based on the substituted amino acids.
[0091] Specifically, a non-natural amino acid at a preset position in a training polypeptide is replaced with a replacement amino acid to create a negative sample of the training polypeptide. Specifically, the replacement amino acid may be a natural amino acid or a non-natural amino acid. The replacement amino acid is the amino acid that is reversely modified or least similar to the target polypeptide, so the negative sample is an unreasonable polypeptide.
[0092] In S604, a pocket feature of the training target protein, a first multimodal feature of the amino acid at each position in the training polypeptide, and a third multimodal feature of the amino acid at each position in the negative sample are obtained, and the multimodal feature of the amino acid at each position includes a main chain direction feature, a main chain rotation feature, a side chain type feature, and a rigid atom group distribution feature.
[0093] In concrete implementation, reference can be made to step S102 in the embodiment shown in FIG. 1 or step S202 in the embodiment shown in FIG. 2, so a detailed description will not be given here.
[0094] In S605, a polypeptide design model is trained by comparative learning based on the pocket features of the training target protein, the first multimodal features of the amino acids at each position in the training polypeptide, and the third multimodal features of the amino acids at each position in the negative sample.
[0095] This step S605, when specifically implemented, may include, for example, the following steps:
[0096] (a2) Using the polypeptide design model, predict a first predicted multimodal feature of an unnatural amino acid at a preset position in the training polypeptide based on the pocket feature of the training target protein and the first multimodal feature of an amino acid at each position other than the preset position in the training polypeptide.
[0097] (b2) Predicting second predicted multimodal features of replacement amino acids at preset positions in the negative sample using the polypeptide design model based on the pocket features of the training target protein and the third multimodal features of amino acids at each position other than the preset positions in the negative sample.
[0098] (c2) adjusting parameters of the polypeptide design model such that a first predicted probability that the type of the preset position in the first predicted multimodal feature is an unnatural amino acid is greater than a second predicted probability that the type of the preset position in the second predicted multimodal feature is a substituted amino acid.
[0099] In this embodiment, the training polypeptide is a verified polypeptide, which is used as a positive sample in this embodiment. However, in actual application scenarios, the number of positive samples is small, so the training effect may not be ideal when using the supervised training method of the embodiment shown in Figure 5 above. In this embodiment, to improve the training effect, more negative samples can be constructed based on the positive samples, and more positive and negative sample pairs can be generated. This allows the polypeptide design model to be trained through comparative learning using positive and negative sample pairs.
[0100] The polypeptide design model training method of this embodiment trains the polypeptide design model through comparative learning, which can generate a large number of positive and negative sample pairs and efficiently train the polypeptide design model, thereby effectively improving the accuracy of the trained polypeptide design model.
[0101] 7 is a schematic diagram of a sixth embodiment of the present disclosure. This embodiment provides a model-based polypeptide design device 700, including: an information acquisition module 701 that acquires a pocket of a target protein and a target polypeptide that labels a preset position of an unnatural amino acid to be designed, the pocket of the target protein binding to the target polypeptide via the preset position; a feature acquisition module 702 that acquires features of the pocket of the target protein and multimodal features of each known amino acid in the target polypeptide, the multimodal features of each known amino acid including a main chain orientation feature, a main chain rotation feature, a side chain type feature, and a rigid atom group distribution feature; and a design module 703 that designs an unnatural amino acid for the preset position in the target polypeptide using a pre-trained polypeptide design model based on the features of the pocket of the target protein and the multimodal features of each known amino acid in the target polypeptide.
[0102] The model-based polypeptide design apparatus 700 of this embodiment employs the above-described modules to achieve model-based polypeptide design, and the implementation principles and technical effects are the same as those of the above-described related method embodiments. For details, please refer to the description of the above-described related method embodiments, and therefore, a detailed description will not be given here.
[0103] 8 is a schematic diagram according to a seventh embodiment of the present disclosure. The model-based polypeptide design device 800 of this embodiment will be described in more detail below, in addition to the technical solution of the embodiment shown in FIG. 7 described above. As shown in FIG. 8, the model-based polypeptide design device 800 of this embodiment includes an information acquisition module 801, a feature acquisition module 802, and a design module 803, which are modules with the same names and functions as those shown in FIG. 7.
[0104] The design module 803 includes a prediction unit 8031 that predicts the multimodal characteristics of the unnatural amino acid at the preset position in the target polypeptide using the polypeptide design model based on the pocket characteristics of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, and a determination unit 8032 that determines the structure of the unnatural amino acid at the preset position in the target polypeptide based on the multimodal characteristics of the unnatural amino acid.
[0105] Further, optionally, in one embodiment of the present disclosure, the prediction unit 8031 predicts the characteristics of the unnatural amino acid at a preset position in the target polypeptide using the polypeptide design model, based on the pocket characteristics of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, and with reference to the initial characteristics of the unnatural amino acid at the preset position previously obtained.
[0106] Further, optionally, in one embodiment of the present disclosure, feature acquisition module 802 acquires initial main chain orientation characteristics for the unnatural amino acids at the preset positions in the target polypeptide based on a Gaussian distribution, acquires initial main chain rotation characteristics for the unnatural amino acids at the preset positions in the target polypeptide based on a uniform distribution, acquires initial side chain type characteristics for the unnatural amino acids at the preset positions in the target polypeptide based on a Gaussian distribution, and acquires initial distribution characteristics of rigid atom groups for the unnatural amino acids at the preset positions in the target polypeptide based on a uniform distribution.
[0107] Furthermore, optionally, as shown in FIG. 8 , in one embodiment of the present disclosure, the device 800 for designing unnatural amino acids in polypeptides further includes a detection module 804 for detecting rationality based on the structure of the pocket of the target protein, the structure of each known amino acid in the target polypeptide, and the structure of the unnatural amino acid at the preset position, and a processing module 805 for discarding the unnatural amino acid designed for the preset position in response to detection of irrationality.
[0108] The model-based polypeptide design apparatus 800 of this embodiment employs the above-described modules to achieve model-based polypeptide design, and the implementation principles and technical effects are the same as those of the above-described related method embodiments. For details, please refer to the description of the above-described related method embodiments, and therefore, a detailed description will not be given here.
[0109] 9 is a schematic diagram of an eighth embodiment of the present disclosure. This embodiment provides an apparatus 900 for training a polypeptide design model, the apparatus including: an information acquisition module 901 for acquiring a pocket of a training target protein and a training polypeptide that binds to the training target protein via an unnatural amino acid at a preset position; a feature acquisition module 902 for acquiring features of the pocket of the training target protein and first multimodal features of amino acids at each position in the training polypeptide, the first multimodal features of the amino acids at each position including a main chain orientation feature, a main chain rotation feature, a side chain type feature, and a rigid atom group distribution feature; and a training module 903 for training the polypeptide design model based on the features of the pocket of the training target protein and the first multimodal features of the amino acids at each position in the training polypeptide.
[0110] The implementation principle and technical effect of the polypeptide design model training device 900 of this embodiment, which employs the above-mentioned modules to train a polypeptide design model, are the same as those of the above-mentioned related method embodiments, and details can be found in the above-mentioned related method embodiments, so a detailed description will not be given here.
[0111] 10 is a schematic diagram of a ninth embodiment of the present disclosure. In addition to the technical solution of the embodiment shown in FIG. 9 described above, the training device 1000 for a polypeptide design model of this embodiment further describes the technical solution of the present disclosure in more detail. As shown in FIG. 10, the training device 1000 for a polypeptide design model of this embodiment includes an information acquisition module 1001, a feature acquisition module 1002, and a training module 1003, which have the same names and functions as those shown in FIG. 9.
[0112] As shown in FIG. 10 , in this embodiment, the training module 1003 includes a prediction unit 10031 that predicts second multimodal features of unnatural amino acids at the preset positions in the training polypeptide using the polypeptide design model based on pocket features of the training target protein and first multimodal features of amino acids at each position other than the preset positions in the training polypeptide; a construction unit 10032 that constructs a first loss function based on the second multimodal features of the unnatural amino acids at the preset positions and the first multimodal features; and an adjustment unit 10033 that adjusts parameters of the polypeptide design model to aim for convergence of the first loss function.
[0113] Further, optionally, in one embodiment of the present disclosure, the construction unit 10032 obtains a main chain orientation feature difference, a main chain rotation feature difference, and a side chain type feature difference based on the second multimodal feature and the first multimodal feature of the unnatural amino acid at the preset position, respectively, and constructs the first loss function based on the main chain orientation feature difference, the main chain rotation feature difference, and the side chain type feature difference.
[0114] Further, optionally, as shown in FIG. 10 , in one embodiment of the present disclosure, the training apparatus 1000 for a polypeptide design model further includes an information acquisition module 1001 for acquiring replacement amino acids at the preset positions in the training polypeptide, and a generation module 1004 for generating negative samples of the training polypeptide based on the replacement amino acids.
[0115] Further, optionally, in one embodiment of the present disclosure, the training module 1003 trains the polypeptide design model by comparative learning based on pocket features of the training target protein, first multimodal features of amino acids at each position in the training polypeptide, and third multimodal features of amino acids at each position in the negative sample.
[0116] Further, optionally, in one embodiment of the present disclosure, the information acquisition module 1001 reverse-modifies the unnatural amino acid at the preset position in the training polypeptide to obtain the replacement amino acid, or obtains the target amino acid that has the least similarity to the unnatural amino acid at the preset position from a pre-constructed amino acid similarity matrix as the replacement amino acid.
[0117] Further, optionally, in one embodiment of the present disclosure, the training module 1001 uses the polypeptide design model to predict a first predicted multimodal feature of an unnatural amino acid at the preset position in the training polypeptide based on a pocket feature of the training target protein and a first multimodal feature of an amino acid at each position other than the preset position in the training polypeptide; and uses the polypeptide design model to predict a second predicted multimodal feature of a replacement amino acid at the preset position in the negative sample based on a pocket feature of the training target protein and a third multimodal feature of an amino acid at each position other than the preset position in the negative sample; and adjusts parameters of the polypeptide design model so that a first predicted probability that the type of the preset position in the first predicted multimodal feature is the unnatural amino acid is greater than a second predicted probability that the type of the preset position in the second predicted multimodal feature is the replacement amino acid.
[0118] The implementation principle and technical effect of the polypeptide design model training apparatus 1000 of this embodiment, which employs the above-mentioned modules to train a polypeptide design model, are the same as those of the above-mentioned related method embodiments, and details thereof can be found in the above-mentioned related method embodiments, so a detailed description thereof will not be given here.
[0119] In the technical solution disclosed herein, the acquisition, storage, application, etc. of relevant users' personal information shall comply with the provisions of relevant laws and regulations and shall not violate public order and morals.
[0120] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0121] 11 is a schematic block diagram of an electronic device 1100 that may be used to implement embodiments of the present disclosure. The electronic device may represent various forms of digital computers, such as laptops, desktop computers, workbenches, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as PDAs, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and are not intended to limit the implementation of the present disclosure as described and / or claimed herein.
[0122] 11, device 1100 includes a computing means 1101 that can perform various appropriate operations and processes in accordance with a computer program stored in a read-only memory (ROM) 1102 or loaded from a storage means 1108 into a random access memory (RAM) 1103. The RAM 1103 may store various programs and data necessary for the operation of electronic device 1100. The computing means 1101, ROM 1102, and RAM 1103 are connected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0123] Several components of device 1100 are connected to I / O interface 1105, including input means 1106, e.g., a keyboard, a mouse, etc., output means 1107, e.g., various types of displays, speakers, etc., storage means 1108, e.g., a magnetic disk, an optical disk, etc., and communication means 1109, e.g., a network card, a modem, a wireless communication transceiver, etc. The communication means 1109 enables electronic device 1100 to exchange information / data with other devices via, e.g., a computer network of the Internet and / or various telecommunication networks.
[0124] The computing means 1101 may be various general-purpose and / or special-purpose processing components having processing capabilities and calculation capabilities. Some examples of the computing means 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing means 1101 executes various methods and processes described above, such as the methods described herein. For example, in some embodiments, the methods described herein may be implemented as a computer software program physically embodied in a machine-readable medium, such as the storage means 1108. In some embodiments, some or all of the computer program can be loaded and / or installed into the electronic device 1100 via the ROM 1102 and / or the communication means 1109. When the computer program is loaded into the RAM 1103 and executed by the computing means 1101, it can perform one or more steps of the methods described herein. Alternatively, in other embodiments, the computing means 1101 may be configured in any other suitable manner (eg, via firmware) to perform the methods described herein.
[0125] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), field programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being embodied in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor may be a special-purpose or general-purpose programmable processor that can receive data and commands from a storage system, at least one input device, and at least one output device, and forward data and commands to the storage system, the at least one input device, and the at least one output device.
[0126] Program code for implementing the methods of the present disclosure can be written using any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus such that, when executed by the processor or controller, the program code performs the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, as a stand-alone package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0127] In the context of this disclosure, a machine-readable medium is a tangible medium that can contain or store a program used by or in connection with a command execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media include one or more line-based electrical connections, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0128] To provide for user interaction, the systems and techniques described herein may be implemented on a computer that includes a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) for providing input by the user to the computer. Other types of devices may also be used to provide for user interaction. For example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user may be received in any form (including sound input, speech input, or tactile input).
[0129] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network (LAN), a wide area network (WAN), and an internetwork.
[0130] The computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server arises by virtue of computer programs running on the corresponding computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server in a blockchain combination.
[0131] It should be understood that steps can be rearranged, added, or deleted using the various types of flows shown above. For example, the steps described in this application can be performed in a parallel order, a sequential order, or can be performed in a different order, and are not limited thereto, as long as the desired results of the technical solution disclosed in this application can be achieved.
[0132] The above specific embodiments do not constitute limitations on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, partial combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A model-based method for designing polypeptides, comprising: obtaining a pocket of a target protein and a target polypeptide in which a preset position of a to-be-designed unnatural amino acid is labeled, and the pocket of the target protein binds to the target polypeptide via the preset position; Obtaining pocket features of the target protein and multimodal features of each known amino acid in the target polypeptide, wherein the multimodal features of each known amino acid include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features; designing unnatural amino acids at the preset positions in the target polypeptide using a pre-trained polypeptide design model based on pocket characteristics of the target protein and multimodal characteristics of each known amino acid in the target polypeptide; A method for designing a polypeptide, comprising:
2. 2. The method of claim 1, designing unnatural amino acids at the preset positions in the target polypeptide using a pre-trained polypeptide design model based on pocket characteristics of the target protein and multimodal characteristics of each known amino acid in the target polypeptide, predicting multimodal characteristics of unnatural amino acids at the preset positions in the target polypeptide using the polypeptide design model based on characteristics of the pocket of the target protein and multimodal characteristics of each known amino acid in the target polypeptide; determining a structure of the unnatural amino acid at the preset position in the target polypeptide based on the multimodal characteristics of the unnatural amino acid; A method for designing a polypeptide, comprising:
3. 3. The method of claim 2, predicting a multimodal characteristic of an unnatural amino acid at the preset position in the target polypeptide using the polypeptide design model based on a pocket characteristic of the target protein and a multimodal characteristic of each known amino acid in the target polypeptide, comprising: A method for designing a polypeptide, comprising: predicting the characteristics of an unnatural amino acid at a preset position in the target polypeptide using the polypeptide design model, based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, and referring to the initial characteristics of the unnatural amino acid at the preset position previously obtained.
4. 4. The method of claim 3, and before predicting the characteristics of the unnatural amino acid at the preset position in the target polypeptide using the polypeptide design model, with reference to the initial characteristics of the unnatural amino acid at the preset position previously obtained based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, further comprising: obtaining initial main chain orientation characteristics of the unnatural amino acid at the preset position in the target polypeptide based on a Gaussian distribution; obtaining initial main-chain rotation characteristics of unnatural amino acids at the preset positions in the target polypeptide based on a uniform distribution; obtaining an initial side chain type property for the unnatural amino acid at the preset position in the target polypeptide based on a Gaussian distribution; obtaining an initial distribution profile of rigid atom groups of unnatural amino acids at the preset positions in the target polypeptide based on a uniform distribution; A method for designing a polypeptide, comprising:
5. 3. The method of claim 2, designing unnatural amino acids at the preset positions in the target polypeptide using a pre-trained polypeptide design model based on the pocket characteristics of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, further comprising: performing a rationality detection based on the structure of the pocket of the target protein, the structure of each known amino acid in the target polypeptide, and the structure of the unnatural amino acid at the preset position; In response to detecting an irrationality, discarding the unnatural amino acid designed for the preset position. A method for designing a polypeptide, comprising:
6. 1. A method for training a polypeptide design model, comprising: obtaining a pocket of a training target protein and a training polypeptide, wherein the training polypeptide binds to the training target protein via an unnatural amino acid at a preset position; obtaining pocket features of the training target protein and first multimodal features of amino acids at each position in the training polypeptide, the first multimodal features of the amino acids at each position including a main chain orientation feature, a main chain rotation feature, a side chain type feature, and a rigid atom group distribution feature; training the polypeptide design model based on pocket features of the training target protein and first multimodal features of amino acids at each position in the training polypeptide; A method for training a polypeptide design model, comprising:
7. 7. The method of claim 6, training the polypeptide design model based on pocket features of the training target protein and first multimodal features of amino acids at each position in the training polypeptide, predicting, using the polypeptide design model, a second multimodal signature of an unnatural amino acid at the preset position in the training polypeptide based on the pocket signature of the training target protein and a first multimodal signature of an amino acid at each position other than the preset position in the training polypeptide; constructing a first loss function based on a second multimodal feature of the unnatural amino acid at the preset position and the first multimodal feature; adjusting parameters of the polypeptide design model with the goal of convergence of the first loss function; A method for training a polypeptide design model, comprising:
8. 8. The method of claim 7, constructing a first loss function based on a second multimodal feature of the unnatural amino acid at the preset position and the first multimodal feature, obtaining a main chain orientation feature difference, a main chain rotation feature difference, and a side chain type feature difference based on the second multimodal feature of the unnatural amino acid at the preset position and the first multimodal feature, respectively; constructing the first loss function based on the main chain direction characteristic difference, the main chain rotation characteristic difference, and the side chain type characteristic difference; A method for training a polypeptide design model, comprising:
9. 7. The method of claim 6, before training the polypeptide design model based on the pocket features of the training target protein and the first multimodal features of the amino acids at each position in the training polypeptide, further comprising: obtaining a replacement amino acid at the preset position in the training polypeptide; generating a negative sample of the training polypeptide based on the substituted amino acids; A method for training a polypeptide design model, comprising:
10. 10. The method of claim 9, training the polypeptide design model based on pocket features of the training target protein and first multimodal features of amino acids at each position in the training polypeptide, training the polypeptide design model by comparative learning based on the pocket features of the training target protein, the first multimodal features of the amino acids at each position in the training polypeptide, and the third multimodal features of the amino acids at each position in the negative sample; A method for training a polypeptide design model, comprising:
11. 10. The method of claim 9, Obtaining a replacement amino acid at the preset position in the training polypeptide comprises: reverse-modifying the unnatural amino acid at the preset position in the training polypeptide to obtain the replacement amino acid; or obtaining, as the replacement amino acid, a target amino acid that has the smallest similarity to the unnatural amino acid at the preset position from a pre-constructed amino acid similarity matrix; A method for training a polypeptide design model, comprising:
12. 11. The method of claim 10, training the polypeptide design model by comparative learning based on pocket features of the training target protein, first multimodal features of amino acids at each position in the training polypeptide, and third multimodal features of amino acids at each position in the negative sample; predicting a first predicted multimodal signature for an unnatural amino acid at the preset position in the training polypeptide using the polypeptide design model based on the pocket signature of the training target protein and a first multimodal signature for an amino acid at each position other than the preset position in the training polypeptide; predicting second predicted multimodal features of the replacement amino acids at the preset positions in the negative samples using the polypeptide design model based on pocket features of the training target proteins and third multimodal features of amino acids at each position other than the preset positions in the negative samples; adjusting parameters of the polypeptide design model such that a first predicted probability that the type of the preset position in the first predicted multimodal feature is the unnatural amino acid is greater than a second predicted probability that the type of the preset position in the second predicted multimodal feature is the substituted amino acid. A method for training a polypeptide design model, comprising:
13. A model-based polypeptide design system, comprising: an information acquisition module that acquires a pocket of a target protein and a target polypeptide, wherein the target polypeptide is labeled with a preset position of an unnatural amino acid to be designed, and the pocket of the target protein binds to the target polypeptide via the preset position; a feature acquisition module that acquires pocket features of the target protein and multimodal features of each known amino acid in the target polypeptide, wherein the multimodal features of each known amino acid include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features; a design module that designs unnatural amino acids at the preset positions in the target polypeptide using a pre-trained polypeptide design model based on pocket characteristics of the target protein and multimodal characteristics of each known amino acid in the target polypeptide; A polypeptide design device comprising:
14. 1. A training device for a polypeptide design model, comprising: an information acquisition module that acquires a pocket of a training target protein and a training polypeptide, the training polypeptide binding to the training target protein via an unnatural amino acid at a preset position; a feature acquisition module that acquires pocket features of the training target protein and first multimodal features of amino acids at each position in the training polypeptide, the first multimodal features of the amino acids at each position including a main chain orientation feature, a main chain rotation feature, a side chain type feature, and a rigid atom group distribution feature; a training module that trains the polypeptide design model based on pocket features of the training target protein and first multimodal features of amino acids at each position in the training polypeptide; A training device for a polypeptide design model, comprising:
15. at least one processor; a memory communicatively coupled to the at least one processor; An electronic device, wherein the memory stores commands executable by the at least one processor, and when the commands are executed by the at least one processor, the at least one processor can perform the method of any one of claims 1 to 5 or claims 6 to 12.
16. A non-transitory computer readable storage medium having stored thereon computer instructions, the computer instructions causing a computer to perform the method of any one of claims 1 to 5 or claims 6 to 12.
17. A computer program product which, when executed by a processor, implements the method according to any one of claims 1 to 5 or claims 6 to 12.