Model-based polypeptide design method, model training method, device and equipment
By employing a model-based peptide design approach, non-natural amino acids can be designed using the multimodal characteristics of target proteins and peptides. This approach addresses the issue of insufficient accuracy in peptide design, improves the drug performance of peptides, and is applicable to the development of G protein-coupled receptor targeted drugs and anti-infective therapy.
Patent Information
- Application Number
- CN202510733056.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Existing peptide design schemes cannot accurately and effectively capture the multimodal characteristics of non-natural amino acids, resulting in insufficient accuracy in peptide design and affecting the drug-likeness of drug development.
A model-based peptide design approach is adopted. By obtaining the pocket of the target protein and the multimodal features of the target peptide, a pre-trained peptide design model is used to design non-natural amino acids in reserved positions in the target peptide. The accuracy of the design is ensured by combining rationality detection.
It improves the design accuracy of non-natural amino acids in peptides, enhances peptide affinity and medicinal properties such as stability and half-life, and is suitable for advancing to the clinical development stage.
Smart Images

Figure CN120877938B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to the field of artificial intelligence technology such as biological computing, and particularly to a model-based peptide design method, model training method, apparatus, and device. Background Technology
[0002] Peptide design with unnatural amino acids is a cutting-edge technology that integrates artificially synthesized or chemically modified unnatural amino acids into the peptide chain, overcoming the structural limitations of natural amino acids to enhance their pharmacological properties and functional diversity.
[0003] As bioactive molecules composed of 3-30 amino acids, peptides play a crucial role in life processes such as signal transduction and immune regulation by specifically recognizing target protein receptors. Summary of the Invention
[0004] This disclosure provides a model-based peptide design method, model training method, apparatus, and device.
[0005] According to one aspect of this disclosure, a model-based peptide design method is provided, comprising:
[0006] Obtain a pocket of a target protein and a target peptide, wherein the target peptide is marked with a reserved position for a non-natural amino acid to be designed; the pocket of the target protein binds to the target peptide through the reserved position;
[0007] The characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide are obtained; the multimodal characteristics of each known amino acid include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics;
[0008] Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, a pre-trained peptide design model is used to design non-natural amino acids at the reserved positions in the target peptide.
[0009] According to another aspect of this disclosure, a method for training a peptide design model is provided, comprising:
[0010] Obtain a pocket of the training target protein and a training peptide, wherein the training peptide binds to the training target protein via a non-natural amino acid at a predetermined position;
[0011] The characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at each position in the training peptide are obtained; the first multimodal characteristics of the amino acids at each position include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics;
[0012] The peptide design model is trained based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide.
[0013] According to another aspect of this disclosure, a model-based peptide design apparatus is provided, comprising:
[0014] An information acquisition module is used to acquire the pocket of a target protein and the target peptide, wherein the target peptide is marked with a reserved position for a non-natural amino acid to be designed; the pocket of the target protein binds to the target peptide through the reserved position;
[0015] The feature acquisition module is used to acquire the features of the pocket of the target protein and the multimodal features of each known amino acid in the target polypeptide; the multimodal features of each known amino acid include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features;
[0016] The design module is used to design non-natural amino acids at the reserved positions in the target peptide based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, using a pre-trained peptide design model.
[0017] According to another aspect of this disclosure, a training apparatus for a peptide design model is provided, comprising:
[0018] An information acquisition module is used to acquire the pocket of the training target protein and the training peptide, wherein the training peptide binds to the training target protein through a non-natural amino acid at a preset position;
[0019] The feature acquisition module is used to acquire the features of the pocket of the training target protein and the first multimodal features of amino acids at each position in the training peptide; the first multimodal features of the amino acids at each position include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features;
[0020] The training module is used to train the peptide design model based on the features of the pocket of the training target protein and the first multimodal features of the amino acids at each position in the training peptide.
[0021] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0022] At least one processor; and
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.
[0025] According to yet another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.
[0026] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.
[0027] According to the technology disclosed herein, the accuracy of non-natural amino acids at reserved positions in the designed target peptide can be effectively improved.
[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0029] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0030] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0031] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0032] Figure 3 This disclosure provides a schematic diagram illustrating the principle of a peptide design model.
[0033] Figure 4 This is a schematic diagram according to the third embodiment of the present disclosure;
[0034] Figure 5 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0035] Figure 6 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0036] Figure 7 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0037] Figure 8 This is a schematic diagram according to the seventh embodiment of the present disclosure;
[0038] Figure 9 This is a schematic diagram according to the eighth embodiment of the present disclosure;
[0039] Figure 10 This is a schematic diagram according to the ninth embodiment of the present disclosure;
[0040] Figure 11 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0041] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0042] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0043] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.
[0044] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0045] In practical applications, although peptide drugs exhibit unique advantages in the hit-to-lead stage of drug development, natural peptides generally suffer from drug-forming bottlenecks such as poor membrane permeability and low metabolic stability. The introduction of non-natural amino acids provides a systematic solution to this problem: enhancing α-helix rigidity through cyclic topology, prolonging plasma half-life through hydrophobic side chain modification, and improving target binding specificity through functionalized group grafting, thereby achieving precise regulation from molecular structure to biological function.
[0046] In existing technologies, the chemical diversity of non-natural amino acids, such as chiral centers and non-standard functional groups, makes their structural compatibility more complex than that of natural amino acids. Existing peptide design schemes can only capture the contextual molecular features of each amino acid in the peptide's amino acid sequence, and cannot perform accurate and effective peptide design.
[0047] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1 As shown, this embodiment provides a model-based peptide design method, which may specifically include the following steps:
[0048] S101. Obtain the pocket of the target protein and the target peptide, and mark the reserved positions of the non-natural amino acids to be designed in the target peptide.
[0049] Specifically, the target polypeptide is the polypeptide to be designed in this embodiment, and the target protein is the target protein whose binding performance with the target polypeptide is to be studied in this embodiment. The target protein pocket binds to the target polypeptide through a reserved site;
[0050] Specifically, the obtained pocket of the target protein may include the structure of the pocket portion of the target protein. Correspondingly, the obtained target polypeptide may also include the amino acid structures at various positions on the polypeptide chain other than the reserved positions.
[0051] S102. Obtain the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide; the multimodal characteristics of each known amino acid include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics;
[0052] Specifically, the features of the pocket of the target protein can be digital features of the structure of the pocket of the target protein using a suitable molecular description language, specifically in the form of vectors.
[0053] In this embodiment, the multimodal features of each known amino acid in the target polypeptide are exemplified by four structural features: main chain orientation, main chain rotation, side chain type, and rigid atom group distribution. In practical applications, the multimodal features of each known amino acid may also include features of other secondary or tertiary structures of the target polypeptide, which will not be elaborated further here. For example, the secondary structure features of the target polypeptide may include the folding type of different segments in the polypeptide sequence, alpha helices, beta folds, and other related features. The tertiary structure features of the target polypeptide include features related to the three-dimensional coordinates of atoms in the amino acids of the polypeptide sequence.
[0054] Among these features, the main chain orientation feature, also known as the main chain 0-direction feature, can be represented as a vector, using Cartesian coordinates with the CA (C atom at the α position of the current amino acid) as the reference point, and its orientation relative to the global coordinates of the target polypeptide, characterizing the spatial positioning of residues. The main chain rotation feature, also known as the main chain rotation matrix, can be characterized by the rotation parameters of the local coordinate system constructed by the N-CA-C atoms of the current amino acid relative to the global coordinate system of the target polypeptide, describing the spatial orientation of residues. The side chain type feature characterizes the side chain type of the current amino acid, a discrete classification feature including natural and non-natural amino acids. The rigid atom group distribution feature divides the current amino acid side chain into several rigid atom groups; within each rigid atom group, the atoms with motion associations on the current amino acid side chain are clustered together. In practical scenarios, motion association atoms can be considered as atoms located on the same plane in the structure of the current amino acid.
[0055] The multimodal features of each known amino acid in the target peptide obtained in this embodiment can accurately and effectively characterize the detailed structural information of each amino acid in the target peptide, providing effective support for designing non-natural amino acids with reserved positions in the target peptide.
[0056] S103. Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, a pre-trained peptide design model is used to design non-natural amino acids in the reserved positions of the target peptide.
[0057] Specifically, in use, the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide can be input into a pre-trained peptide design model. Based on the input information, the peptide design model can accurately and effectively predict the non-natural amino acids at the reserved positions, thereby achieving accurate and effective design of the target peptide.
[0058] The model-based peptide design method in this embodiment can accurately and effectively design non-natural amino acids at reserved positions in the target peptide by leveraging the characteristics of the target protein's pockets and the multimodal features of each known amino acid in the target peptide, using a pre-trained peptide design model. Because the multimodal features of each known amino acid in the target peptide are used, carrying detailed tertiary structural information of each amino acid, the accuracy of designing non-natural amino acids at reserved positions in the target peptide can be effectively improved.
[0059] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure; the model-based peptide design method of this embodiment, in the above... Figure 1 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 2As shown, the model-based peptide design method of this embodiment may specifically include the following steps:
[0060] S201. Obtain the pocket of the target protein and the target peptide, and mark the reserved positions of the non-natural amino acids to be designed in the target peptide.
[0061] In this embodiment, the amino acids at positions other than the reserved positions of the target peptide can be natural amino acids or certified non-natural amino acids. For example, currently, there are 20 commonly used natural amino acids, and there are 160 certified non-natural amino acids. Of course, with the development of medical technology, the number of certified non-natural amino acids will continue to increase.
[0062] S202. Obtain the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide; the multimodal characteristics of each known amino acid include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics;
[0063] In this embodiment, the rigid atom group distribution characteristics are acquired based on the dependence of each atom on the dihedral angle of the current amino acid side chain, dividing the current amino acid side chain into several rigid atom groups. Specifically, by analyzing the linkage of side chain atoms during dihedral angle rotation, atoms with motion association are clustered into the same rigid unit, thereby establishing a quantifiable rigid atom group distribution. Here, atoms with motion association can also be considered as atoms located on a single plane. For example, in a concrete implementation, atoms on the side chain can be analyzed sequentially according to their distance from the main chain from closest to furthest, and atoms located on a single plane, i.e., atoms with motion association, can be grouped into the same rigid unit, thus achieving the division of the rigid atom group corresponding to the current amino acid. Even for non-natural amino acids with complex side chains, this method can accurately and efficiently acquire the corresponding rigid atom group distribution characteristics.
[0064] The side chain type feature in this embodiment is a discrete feature. As described in the previous embodiment, the side chain type can be any one of the 20 natural amino acids or any one of the 160 certified non-natural amino acids. The side chain type feature can be characterized by the identifier of the corresponding amino acid type, with each amino acid identifier uniquely representing the corresponding amino acid.
[0065] S203. Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, a peptide design model is used to predict the multimodal characteristics of non-natural amino acids in the reserved positions of the target peptide.
[0066] For example, in the specific implementation of this step, based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, and referring to the initial characteristics of the non-natural amino acids at the reserved positions obtained in advance, a peptide design model can be used to predict the multimodal characteristics of the non-natural amino acids at the reserved positions in the target peptide, which can achieve accurate and effective prediction of the multimodal characteristics of the non-natural amino acids at the reserved positions.
[0067] Specifically, before this step is implemented, the following steps may also be included:
[0068] Based on Gaussian distribution, such as Obtain the initial characteristics of the main chain orientation of non-natural amino acids at reserved positions in the target polypeptide;
[0069] Based on a uniform distribution such as U(SO(3)), the initial characteristics of the main chain rotation of non-natural amino acids at reserved positions in the target polypeptide are obtained;
[0070] Based on Gaussian distribution, such as Initial characteristics of the side chain type of non-natural amino acids at reserved positions in the target peptide;
[0071] Based on a uniform distribution such as U([0,2π) 5 This method obtains the initial distribution characteristics of the rigid atom set of non-natural amino acids at reserved positions in the target polypeptide.
[0072] The peptide design model in this embodiment can be a conditional flow matching design model, in which the initial features of the non-natural amino acids at the reserved positions in the target peptide can be used as noise input. The peptide design model performs noise reduction and recovery on the input noise based on the features of the pocket of the input target protein and the multimodal features of the known amino acids at each position other than the reserved positions in the target peptide. The multimodal features of the non-natural amino acids at the reserved positions in the target peptide are decoded. For example, it can include the main chain direction features, main chain rotation features, side chain type features, and rigid atom group distribution features of the non-natural amino acids at the reserved positions.
[0073] The design principle of the peptide design model in this embodiment can be considered as using conditional flow matching to predict the four modalities of non-natural amino acids at reserved positions from their initial distribution to the target distribution. Specifically, the initial distribution of the main chain direction features, i.e., the initial main chain direction features, conforms to a Gaussian distribution. The initial distribution of the main chain rotation feature, i.e. the initial main chain rotation feature, conforms to a uniform distribution U(SO(3)); the initial distribution of the side chain type feature, i.e. the initial side chain type feature, conforms to a Gaussian distribution. The initial distribution of the rigid atom set conforms to a uniform distribution U([0, 2π).5 By adopting the above method, the initial features of non-natural amino acids at the reserved positions are set very reasonably and accurately, providing an effective basis for the subsequent noise reduction and decoding process to accurately predict the multimodal features of non-natural amino acids at the reserved positions.
[0074] Figure 3 This is a schematic diagram illustrating the principle of a peptide design model provided in this disclosure. For example... Figure 3 As shown, the peptide design model can include a preprocessing module, a node encoder, an edge encoder, a fusion processing module, and a noise reduction module. The node encoder and edge encoder can extract features of individual amino acids and features between adjacent amino acids, respectively. For example, the preprocessing module can process the input information to obtain the information required by the node encoder and edge encoder, and input them to the node encoder or edge encoder respectively. For each amino acid, the node encoder can generate residue-level node embeddings by fusing amino acid type features, local atomic coordinate features, and main chain dihedral features. The local atomic coordinate features can be obtained by the preprocessing module based on the type of the current amino acid, from a pre-collected amino acid database. This amino acid database can include information on all natural amino acids and all certified non-natural amino acids, such as the local coordinate information of each atom in the current amino acid in Cartesian coordinates with CA atoms as the reference point. For each amino acid, the preprocessing module can also obtain the main chain dihedral information of the current amino acid based on its local coordinate information. Specifically, the information on amino acids at reserved positions can be processed based on the initial characteristics of the non-natural amino acids at those reserved positions.
[0075] Specifically, in the Node Encoder, amino acid types are encoded through the Embedding layer, local atomic coordinates are expanded into high-dimensional features according to amino acid types after local coordinate system transformation, and dihedral angles of the main chain are encoded using angles. These three types of features are concatenated and then fused using a Multilayer Perceptron (MLP). Multiple masks, such as sequence masks, structural masks, and residue masks, are used to control information leakage, ultimately outputting a feature vector for each residue. This effectively handles the integration of protein sequence and structural information.
[0076] Edge Encoding focuses on modeling relationships between residues. Amino acid pairs represent type interactions through Pair Embedding, while Relative Position Embedding encodes sequence spacing. Atom distances are transformed using a Gaussian kernel and then processed through an MLP to generate distance features (Distance Embedding), which are dynamically weighted using learnable distance coefficients. Simultaneously, the dihedral angles of paired residues are captured through angle encoding to capture spatial orientation associations. All features are fused into edge embeddings, comprehensively characterizing the structural adjacency, sequence correlation, and spatial interactions between residues. The edges between adjacent amino acids in the peptide can be pre-determined or pre-determined by the preprocessing module based on a detection strategy. For example, if the nearest atom pair of two amino acids is less than 4 angstroms apart, it is considered to have an edge. Similarly, the edges between the pocket of the target protein and the non-natural amino acids at the reserved positions of the target peptide can be detected in the preprocessing module. There can be one, two, or more edges between the pocket of the target protein and the non-natural amino acids at the reserved positions.
[0077] like Figure 3 As shown, the fusion processing module is used to fuse all feature information encoded by the Node Encoder and Edge Encoder, and the noise reduction processing module performs noise reduction decoding based on the fused information to recover and output the multimodal features of non-natural amino acids at preset positions.
[0078] Based on the working principle of the above peptide design model, it can be seen that the peptide design model is a conditional flow matching model based on graph neural network. The initial features of the non-natural amino acids at the reserved position of the target peptide are used as noise input into the peptide design model. The peptide design model decodes and recovers the noise based on other input features to obtain the multimodal features of the non-natural amino acids at the reserved position.
[0079] S204. Based on the multimodal characteristics of non-natural amino acids, determine the structure of non-natural amino acids at reserved positions in the target polypeptide;
[0080] Specifically, based on the multimodal characteristics of unnatural amino acids at reserved positions predicted by peptide design models, and combined with information such as bond lengths and bond angles of chemical molecules, the coordinates of each atom in the unnatural amino acid at the reserved position can be determined, thereby obtaining the structure of the unnatural amino acid.
[0081] S205. Based on the structure of the pocket of the target protein, the structure of each known amino acid in the target polypeptide, and the structure of non-natural amino acids at the reserved positions, a rationality test is performed.
[0082] S206. In response to the detection of inappropriateness, discard non-natural amino acids designed for reserved positions.
[0083] The rationality check in this embodiment can be considered as atomic collision detection. For example, knowing the structure of the pocket of the target protein, the structure of each known amino acid in the designed target peptide, and the structure of the non-natural amino acid at the reserved position, the coordinates of each atom in the pocket structure of the target protein and the coordinates of each atom in the structure of each amino acid at each position in the target peptide can be obtained. Then, the existence of atomic collisions can be further detected; specifically, this refers to detecting whether each atom in the non-natural amino acid at the designed reserved position collides with each atom in the pocket structure and each atom in other positions of the target peptide. If collisions occur, the non-natural amino acid designed for the reserved position in the target peptide is considered unreasonable and can be discarded. If the detection is reasonable, the non-natural amino acid designed for the reserved position is retained, thereby realizing the design of the target peptide.
[0084] The model-based peptide design method of this embodiment uses a peptide design model to predict the multimodal characteristics of unnatural amino acids at reserved positions in the target peptide based on the features of the pockets of the target protein and the multimodal characteristics of each known amino acid in the target peptide. This allows for the determination of the structure of the unnatural amino acids at the reserved positions in the target peptide. Since the multimodal characteristics carry rich internal structural information of the amino acids, the technical solution of this embodiment can accurately and effectively design unnatural amino acids in peptides, thereby accurately and effectively designing peptides.
[0085] Furthermore, the technical solution of this embodiment can also perform rationality testing based on the structure of the pocket of the target protein, the structure of each known amino acid in the target peptide, and the structure of non-natural amino acids at the reserved positions, so as to remove unreasonably designed peptides and ensure the accuracy and effectiveness of peptides in subsequent research and development and experimental stages.
[0086] Experimental verification shows that the peptides designed using the technical solution of this embodiment can have their affinity effectively improved; at the same time, their medicinal properties, such as stability and half-life, are also improved, making them more suitable for advancing to the next stage of clinical research and development.
[0087] The technical solution of this embodiment, with its designed peptides containing non-natural amino acids, can be widely applied in the development of G protein-coupled receptor (GPCR) targeted drugs, as well as in anti-infection and antibacterial therapies.
[0088] Figure 4 This is a schematic diagram based on the third embodiment of the present disclosure; as shown Figure 4As shown, the training method for the peptide design model in this embodiment may specifically include the following steps:
[0089] S401. Obtain the pocket of the training target protein and the training peptide. The training peptide binds to the training target protein through non-natural amino acids at a preset position.
[0090] It should be noted that the training peptide also includes amino acids at several other positions.
[0091] S402. Obtain the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at each position in the training peptide; the first multimodal characteristics of amino acids at each position include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics.
[0092] The specific implementation method for this step can be found above. Figure 1 or Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.
[0093] S403. Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide, the peptide design model is trained.
[0094] The peptide design model training method of this embodiment trains the peptide design model based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at each position in the training peptide. Since the multimodal characteristics carry rich internal structural information of amino acids, the technical solution of this embodiment can effectively improve the accuracy of the trained peptide design model.
[0095] Figure 5 This is a schematic diagram according to the fourth embodiment of this disclosure; the training method of the peptide design model in this embodiment, as described above... Figure 4 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 5 As shown, the training method for the peptide design model in this embodiment may specifically include the following steps:
[0096] S501. Obtain the pocket of the training target protein and the training peptide. The training peptide binds to the training target protein through non-natural amino acids at a preset position.
[0097] S502. Obtain the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at each position in the training peptide; the first multimodal characteristics of amino acids at each position include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics.
[0098] For specific implementation, please refer to the above. Figure 1 Step S102 of the illustrated embodiment or Figure 2 Step S202 in the illustrated embodiment will not be described again here.
[0099] S503. Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at positions other than the preset position in the training peptide, a peptide design model is used to predict the second multimodal characteristics of non-natural amino acids at the preset position in the training peptide.
[0100] The specific implementation method is the same as described above. Figure 2 The implementation of step S203 in the illustrated embodiment is the same. For details, please refer to the relevant records of the above embodiments, which will not be repeated here.
[0101] S504. Construct a first loss function based on the second multimodal features and the first multimodal features of non-natural amino acids at preset positions;
[0102] For example, this step can be implemented by including the following steps:
[0103] (a1) Based on the second multimodal features and the first multimodal features of non-natural amino acids at preset positions, the main chain direction feature difference, main chain rotation feature difference, and side chain type feature difference are obtained respectively;
[0104] The main chain orientation feature difference can be taken as the difference between the main chain orientation feature of the non-natural amino acid at the preset position in the second multimodal feature and the main chain orientation feature of the non-natural amino acid at the preset position in the first multimodal feature, and can be represented in the form of vector difference.
[0105] The corresponding main chain selection feature difference is determined in a similar way to the main chain direction feature difference.
[0106] Regarding the side chain type feature difference, the first multimodal feature is the multimodal feature corresponding to the training peptide. Therefore, the probability that the amino acid type at the preset position in the first multimodal feature is a non-natural amino acid can be considered equal to 1. The second multimodal feature is the mask for the non-natural amino acid at the preset position, and the predicted multimodal feature of the non-natural amino acid at the preset position. At this time, the probability that the amino acid type at the preset position in the second multimodal feature is a non-natural amino acid can be denoted as p, where p is greater than 0 and less than 1. At this time, the side chain type feature difference can be equal to 1-p.
[0107] (b1) Construct the first loss function based on the main chain direction feature difference, the main chain rotation feature difference, and the side chain type feature difference.
[0108] Specifically, the weighted sum of the main chain direction feature difference, the main chain rotation feature difference, and the side chain type feature difference can be used as the first loss function. The weights corresponding to each feature can be set based on experience or requirements, and are not limited here.
[0109] S505. Adjust the parameters of the peptide design model with the goal of achieving convergence of the first loss function.
[0110] In this embodiment, a supervised training of a peptide design model is performed using a single training dataset that includes a pocket of training target proteins and a training peptide. In practical applications, multiple similar training datasets can be used, and the peptide design model can be trained using the above training method until the number of training iterations reaches a preset threshold, or the first loss function consistently converges in multiple rounds of training, at which point the training terminates, and the peptide design model is obtained.
[0111] The peptide design model training method in this embodiment uses certified training peptides as positive samples and trains the peptide design model through supervised training, which can effectively improve the accuracy of the trained peptide design model.
[0112] Figure 6 This is a schematic diagram according to the fifth embodiment of this disclosure; the training method of the peptide design model in this embodiment, in the above... Figure 4 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 6 As shown, the training method for the peptide design model in this embodiment may specifically include the following steps:
[0113] S601. Obtain the pocket of the training target protein and the training peptide. The training peptide binds to the training target protein through a non-natural amino acid at a preset position.
[0114] S602. Obtain the replacement amino acids at preset positions in the training peptide;
[0115] For example, a non-natural amino acid at a preset position in a training peptide can be reverse-modified to obtain a replacement amino acid; or a target amino acid with the lowest similarity to a non-natural amino acid at a preset position can be obtained from a pre-constructed amino acid similarity matrix and used as a replacement amino acid.
[0116] For example, by reversing the conformation of the positive sample side chain of a validated polypeptide sequence, such as by mirroring and flipping the orientation of charged groups, negative samples with specific differences in physicochemical properties can be generated, effectively solving the problem of scarce experimental data for non-natural amino acids.
[0117] In this embodiment, the number of rows and columns in the pre-constructed amino acid similarity matrix can be equal to the sum of the number of natural amino acids and the number of certified non-natural amino acids; each row and each column corresponds to one amino acid. The value of each position in the matrix, such as (i,j), is equal to the similarity between the two amino acids in the row and column corresponding to that position; the similarity between the two amino acids is equal to the weighted sum of the structural similarity and the pharmacophore distribution similarity between the two amino acids; wherein, the structural similarity and the pharmacophore distribution similarity can be calculated using known experimental methods or using a pre-trained neural network model.
[0118] Since the replacement amino acid selected in this embodiment is used to generate negative samples, the row corresponding to the non-natural amino acid at the preset position is first located in the amino acid similarity matrix, and then the target amino acid corresponding to the column with the lowest similarity is obtained as the replacement amino acid.
[0119] In this embodiment, the side chain can be replaced while maintaining the main chain conformation stability, thereby generating accurate replacement amino acids.
[0120] S603. Based on the replacement of amino acids, negative samples of training peptides are generated.
[0121] Specifically, this involves replacing non-natural amino acids at predetermined positions in the training peptide with alternative amino acids, creating a negative sample of the training peptide. The alternative amino acids can be either natural or non-natural. Because the alternative amino acids are obtained through reverse modification or by using the amino acid with the lowest similarity, the negative sample is an unreasonable peptide relative to the target peptide.
[0122] S604. Obtain the characteristics of the pocket of the training target protein, the first multimodal characteristics of amino acids at each position in the training peptide, and the third multimodal characteristics of amino acids at each position in the negative sample; the multimodal characteristics of amino acids at each position include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics;
[0123] For specific implementation, please refer to the above. Figure 1 Step S102 of the illustrated embodiment or Figure 2 Step S202 in the illustrated embodiment will not be described again here.
[0124] S605. Based on the features of the pocket of the training target protein, the first multimodal features of amino acids at each position in the training peptide, and the third multimodal features of amino acids at each position in the negative sample, the peptide design model is trained through comparative learning.
[0125] For example, the specific implementation of step S605 may include the following steps:
[0126] (a2) Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at other positions in the training peptide besides the preset position, a peptide design model is used to predict the first predicted multimodal characteristics of non-natural amino acids at the preset position in the training peptide.
[0127] (b2) Based on the features of the pocket of the training target protein and the third multimodal features of amino acids at other positions in the negative sample besides the preset position, a peptide design model is used to predict the second predicted multimodal features of the replacement amino acids at the preset position in the negative sample.
[0128] (c2) The parameters of the peptide design model are adjusted with the first prediction probability that the type of the preset position in the first prediction multimodal feature is a non-natural amino acid and the second prediction probability that the type of the preset position in the second prediction multimodal feature is a replaced amino acid as the target.
[0129] In this embodiment, the training peptide is a validated peptide, which is used as a positive sample. However, in real-world applications, due to the limited number of positive samples, the above-mentioned peptide is used... Figure 5 The supervised training method in the illustrated embodiment may not produce ideal training results. In this embodiment, to improve training results, more negative samples can be constructed based on positive samples, thereby generating more positive-negative sample pairs; thus, positive-negative sample pairs can be used to train the peptide design model through comparative learning.
[0130] The peptide design model training method in this embodiment trains the peptide design model through comparative learning. Because of this training method, a large number of positive and negative sample pairs can be constructed, thereby effectively training the peptide design model and improving the accuracy of the trained peptide design model.
[0131] Figure 7 This is a schematic diagram according to the sixth embodiment of the present disclosure; this embodiment provides a model-based peptide design apparatus 700, including:
[0132] Information acquisition module 701 is used to acquire the pocket of the target protein and the target peptide, wherein the target peptide is marked with a reserved position for a non-natural amino acid to be designed; the pocket of the target protein binds to the target peptide through the reserved position;
[0133] The feature acquisition module 702 is used to acquire the features of the pocket of the target protein and the multimodal features of each known amino acid in the target polypeptide; the multimodal features of each known amino acid include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features;
[0134] Design module 703 is used to design non-natural amino acids at the reserved positions in the target polypeptide based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, using a pre-trained polypeptide design model.
[0135] The model-based peptide design device 700 of this embodiment achieves the same implementation principle and technical effect of model-based peptide design by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0136] Figure 8 This is a schematic diagram according to the seventh embodiment of the present disclosure; the model-based peptide design device 800 of this embodiment, in the above... Figure 7 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 8 As shown, the model-based peptide design device 800 of this embodiment includes... Figure 7 The modules with the same name and function shown are: information acquisition module 801, feature acquisition module 802, and design module 803.
[0137] Design module 803 includes:
[0138] The prediction unit 8031 is used to predict the multimodal characteristics of the non-natural amino acids at the reserved positions in the target polypeptide based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, using the polypeptide design model.
[0139] The determining unit 8032 is used to determine the structure of the non-natural amino acid at the reserved position in the target polypeptide based on the multimodal characteristics of the non-natural amino acid.
[0140] Further optionally, in one embodiment of this disclosure, the prediction unit 8031 is configured to:
[0141] Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, and referring to the initial characteristics of the non-natural amino acids at the reserved positions obtained in advance, the peptide design model is used to predict the characteristics of the non-natural amino acids at the reserved positions in the target peptide.
[0142] Further optionally, in one embodiment of this disclosure, the feature acquisition module 802 is configured to:
[0143] Based on Gaussian distribution, the initial characteristics of the main chain orientation of non-natural amino acids at the reserved positions in the target polypeptide are obtained;
[0144] Based on uniform distribution, the initial main chain rotation characteristics of the non-natural amino acids at the reserved positions in the target polypeptide are obtained;
[0145] Based on Gaussian distribution, the initial characteristics of the side chain type of the non-natural amino acid at the reserved position in the target polypeptide are obtained;
[0146] Based on uniform distribution, the initial distribution characteristics of the rigid atom set of non-natural amino acids at the reserved positions in the target polypeptide are obtained.
[0147] Further optional, such as Figure 8 As shown, in one embodiment of this disclosure, the design device 800 for non-natural amino acids in the polypeptide further includes:
[0148] The detection module 804 is used to perform a rationality test based on the structure of the pocket of the target protein, the structure of each known amino acid in the target polypeptide, and the structure of the non-natural amino acid at the reserved position;
[0149] Processing module 805 is configured to discard the non-natural amino acid designed for the reserved position in response to the detection of an inappropriateness.
[0150] The model-based peptide design device 800 of this embodiment achieves the same implementation principle and technical effect of model-based peptide design by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0151] Figure 9 This is a schematic diagram according to the eighth embodiment of the present disclosure; this embodiment provides a training device 900 for a peptide design model, including:
[0152] Information acquisition module 901 is used to acquire the pocket of training target protein and training peptide, wherein the training peptide binds to the training target protein through a non-natural amino acid at a preset position;
[0153] The feature acquisition module 902 is used to acquire the features of the pocket of the training target protein and the first multimodal features of amino acids at each position in the training peptide; the first multimodal features of the amino acids at each position include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features;
[0154] Training module 903 is used to train the peptide design model based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide.
[0155] The peptide design model training device 900 of this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0156] Figure 10 This is a schematic diagram according to the ninth embodiment of this disclosure; the training device 1000 for the peptide design model in this embodiment, in the above... Figure 9 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 10 As shown, the training device 1000 for the peptide design model in this embodiment includes... Figure 9 The modules with the same name and function shown are: information acquisition module 1001, feature acquisition module 1002, and training module 1003.
[0157] like Figure 10 As shown, in this embodiment, the training module 1003 includes:
[0158] The prediction unit 10031 is used to predict the second multimodal features of non-natural amino acids at the preset position in the training peptide based on the features of the pocket of the training target protein and the first multimodal features of amino acids at positions other than the preset position in the training peptide, using the peptide design model.
[0159] Construction unit 10032 is used to construct a first loss function based on the second multimodal feature of the non-natural amino acid at the preset position and the first multimodal feature;
[0160] The adjustment unit 10033 is used to adjust the parameters of the peptide design model with the goal of achieving convergence of the first loss function.
[0161] Further optionally, in one embodiment of this disclosure, the construction unit 10032 is used for:
[0162] Based on the second multimodal features and the first multimodal features of the non-natural amino acids at the preset positions, the main chain direction feature difference, the main chain rotation feature difference, and the side chain type feature difference are obtained respectively.
[0163] The first loss function is constructed based on the main chain direction feature difference, the main chain rotation feature difference, and the side chain type feature difference.
[0164] Further optional, such as Figure 10 As shown, in one embodiment of this disclosure, the training device 1000 for the peptide design model further includes a generation module 1004:
[0165] Information acquisition module 1001 is used to acquire the replacement amino acid at the preset position in the training peptide;
[0166] The generation module 1004 is used to generate negative samples of the training peptide based on the replaced amino acids.
[0167] Further optionally, in one embodiment of this disclosure, the training module 1003 is used for:
[0168] Based on the features of the pocket of the training target protein, the first multimodal features of amino acids at each position in the training peptide, and the third multimodal features of amino acids at each position in the negative sample, the peptide design model is trained through comparative learning.
[0169] Further optionally, in one embodiment of this disclosure, the information acquisition module 1001 is used for:
[0170] The non-natural amino acid at the preset position in the training peptide is reverse-modified to obtain the replaced amino acid; or
[0171] From a pre-constructed amino acid similarity matrix, the target amino acid with the lowest similarity to the non-natural amino acid at the preset position is selected as the replacement amino acid.
[0172] Further optionally, in one embodiment of this disclosure, the training module 1001 is used for:
[0173] Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at positions other than the preset position in the training peptide, the peptide design model is used to predict the first predicted multimodal characteristics of non-natural amino acids at the preset position in the training peptide.
[0174] Based on the features of the pocket of the training target protein and the third multimodal features of amino acids at positions other than the preset position in the negative sample, the peptide design model is used to predict the second predicted multimodal features of the replacement amino acid at the preset position in the negative sample.
[0175] The parameters of the peptide design model are adjusted with the goal of having a first prediction probability that the type of the preset position in the first predicted multimodal feature is the non-natural amino acid and a second prediction probability that the type of the preset position in the second predicted multimodal feature is the substituted amino acid.
[0176] The peptide design model training device 1000 of this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0177] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0178] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0179] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0180] like Figure 11 As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded from storage unit 1108 into random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.
[0181] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0182] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the methods described above in this disclosure. For example, in some embodiments, the methods described above in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 11011. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the methods described above in this disclosure can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the methods described above in this disclosure by any other suitable means (e.g., by means of firmware).
[0183] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0184] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0185] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0186] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0187] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0188] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0189] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0190] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A model-based peptide design method, comprising: Obtain the pocket of the target protein and the target peptide, wherein the target peptide is marked with reserved positions for non-natural amino acids to be designed; The pocket of the target protein binds to the target polypeptide through the reserved site; The characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide are obtained; The multimodal characteristics of each of the known amino acids include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics; Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, a pre-trained peptide design model is used to design non-natural amino acids at the reserved positions in the target peptide.
2. The method according to claim 1, wherein, Based on the characteristics of the pocket in the target protein and the multimodal characteristics of each known amino acid in the target peptide, a pre-trained peptide design model is used to design non-natural amino acids at the reserved positions in the target peptide, including: Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target polypeptide, the polypeptide design model is used to predict the multimodal characteristics of the non-natural amino acids at the reserved positions in the target polypeptide. Based on the multimodal characteristics of the non-natural amino acids, the structure of the non-natural amino acids at the reserved positions in the target polypeptide is determined.
3. The method according to claim 2, wherein, Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, the peptide design model is used to predict the multimodal characteristics of the non-natural amino acids at the reserved positions in the target peptide, including: Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, and referring to the initial characteristics of the non-natural amino acids at the reserved positions obtained in advance, the peptide design model is used to predict the characteristics of the non-natural amino acids at the reserved positions in the target peptide.
4. The method according to claim 3, wherein, Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, and referring to the pre-obtained information, before predicting the characteristics of the non-natural amino acids at the reserved positions in the target peptide using the peptide design model, the method further includes: Based on Gaussian distribution, the initial characteristics of the main chain orientation of non-natural amino acids at the reserved positions in the target polypeptide are obtained; Based on uniform distribution, the initial main chain rotation characteristics of the non-natural amino acids at the reserved positions in the target polypeptide are obtained; Based on Gaussian distribution, the initial characteristics of the side chain type of the non-natural amino acid at the reserved position in the target polypeptide are obtained; Based on uniform distribution, the initial distribution characteristics of the rigid atom set of non-natural amino acids at the reserved positions in the target polypeptide are obtained.
5. The method according to any one of claims 2-4, wherein, Based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, after designing the non-natural amino acid at the reserved position in the target peptide using a pre-trained peptide design model, the method further includes: Based on the structure of the pocket of the target protein, the structure of each known amino acid in the target polypeptide, and the structure of the non-natural amino acid at the reserved position, a rationality test is performed; In response to the detection of inappropriateness, the non-natural amino acid designed for the reserved position is discarded.
6. A training method for a peptide design model, comprising: Obtain a pocket of the training target protein and a training peptide, wherein the training peptide binds to the training target protein via a non-natural amino acid at a predetermined position; The characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide are obtained; The first multimodal characteristics of the amino acids at each position include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics; The peptide design model is trained based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide.
7. The method according to claim 6, wherein, Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide, the peptide design model is trained, including: Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at positions other than the preset position in the training peptide, the peptide design model is used to predict the second multimodal characteristics of non-natural amino acids at the preset position in the training peptide. A first loss function is constructed based on the second multimodal features of the non-natural amino acids at the preset positions and the first multimodal features; The parameters of the peptide design model are adjusted with the goal of achieving convergence of the first loss function.
8. The method according to claim 7, wherein, Based on the second multimodal features of the non-natural amino acids at the preset positions and the first multimodal features, a first loss function is constructed, including: Based on the second multimodal features and the first multimodal features of the non-natural amino acids at the preset positions, the main chain direction feature difference, the main chain rotation feature difference, and the side chain type feature difference are obtained respectively. The first loss function is constructed based on the main chain direction feature difference, the main chain rotation feature difference, and the side chain type feature difference.
9. The method according to claim 6, wherein, Before training the peptide design model based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide, the method further includes: Obtain the replacement amino acid at the preset position in the training peptide; Based on the replaced amino acids, negative samples of the training peptide are generated.
10. The method according to claim 9, wherein, Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of the amino acids at each position in the training peptide, the peptide design model is trained, including: Based on the features of the pocket of the training target protein, the first multimodal features of amino acids at each position in the training peptide, and the third multimodal features of amino acids at each position in the negative sample, the peptide design model is trained through comparative learning.
11. The method according to claim 9, wherein, Obtaining the replacement amino acid at the preset position in the training peptide includes: The non-natural amino acid at the preset position in the training peptide is reverse-modified to obtain the replaced amino acid; or From a pre-constructed amino acid similarity matrix, the target amino acid with the lowest similarity to the non-natural amino acid at the preset position is selected as the replacement amino acid.
12. The method according to claim 10, wherein, Based on the features of the pocket of the training target protein, the first multimodal features of amino acids at each position in the training peptide, and the third multimodal features of amino acids at each position in the negative samples, the peptide design model is trained through contrastive learning, including: Based on the characteristics of the pocket of the training target protein and the first multimodal characteristics of amino acids at positions other than the preset position in the training peptide, the peptide design model is used to predict the first predicted multimodal characteristics of non-natural amino acids at the preset position in the training peptide. Based on the features of the pocket of the training target protein and the third multimodal features of amino acids at positions other than the preset position in the negative sample, the peptide design model is used to predict the second predicted multimodal features of the replacement amino acid at the preset position in the negative sample. The parameters of the peptide design model are adjusted with the goal of having a first prediction probability that the type of the preset position in the first predicted multimodal feature is the non-natural amino acid and a second prediction probability that the type of the preset position in the second predicted multimodal feature is the substituted amino acid.
13. A model-based peptide design device, comprising: The information acquisition module is used to acquire the pocket of the target protein and the target peptide, wherein the target peptide is marked with reserved positions for non-natural amino acids to be designed. The pocket of the target protein binds to the target polypeptide through the reserved site; The feature acquisition module is used to acquire the features of the pocket of the target protein and the multimodal features of each known amino acid in the target polypeptide; the multimodal features of each known amino acid include main chain orientation features, main chain rotation features, side chain type features, and rigid atom group distribution features; The design module is used to design non-natural amino acids at the reserved positions in the target peptide based on the characteristics of the pocket of the target protein and the multimodal characteristics of each known amino acid in the target peptide, using a pre-trained peptide design model.
14. A training device for a peptide design model, comprising: An information acquisition module is used to acquire the pocket of the training target protein and the training peptide, wherein the training peptide binds to the training target protein through a non-natural amino acid at a preset position; The feature acquisition module is used to acquire the features of the pocket of the training target protein and the first multimodal features of amino acids at each position in the training peptide; The first multimodal characteristics of the amino acids at each position include main chain orientation characteristics, main chain rotation characteristics, side chain type characteristics, and rigid atom group distribution characteristics; The training module is used to train the peptide design model based on the features of the pocket of the training target protein and the first multimodal features of the amino acids at each position in the training peptide.
15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-5 or 6-12.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5 or 6-12.
17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5 or 6-12.
Citation Information
Patent Citations
Model training method and device
CN116130024A
Computational method for designing enzymes for incorporation of non natural amino acids into proteins
US20040053390A1