Candidate polypeptide drug generation method and system based on artificial intelligence
Through the artificial intelligence-based polypeptide drug generation method, new candidate polypeptide molecules are generated using diffusion model and inverse folding technology, solving the problem of low prediction efficiency of polypeptide drug toxicity, and achieving efficient and accurate identification and generation of candidate drug molecules.
Patent Information
- Application Number
- CN202510274649.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-10
AI Technical Summary
In the prior art, the prediction efficiency of polypeptide drug toxicity is low, resulting in a hysteresis of the efficiency of polypeptide drug development.
Using the artificial intelligence-based candidate polypeptide drug generation method, we collect sequence data, structural data and related property tags of polypeptide drug molecules and their target proteins, use the diffusion model to generate the polypeptide backbone, and generate new candidate polypeptide molecules through reverse folding.
It realizes the rapid identification and generation of potential candidate drug molecules accurately, efficiently and at low cost, effectively solving the problem that polypeptides are difficult to model in free state, and combines the physical characteristics of the protein to adjust the diffusion process to make it conform to the constraints of the actual chemical environment.
Smart Images

Figure CN120183538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of drug discovery, and in particular, to a method and system for generating candidate polypeptide drugs based on artificial intelligence. Background Art
[0002] Drug discovery is the process of discovering new candidate drugs. Polypeptides are composed of 20 amino acids through dehydration condensation, and different combinations can form a huge chemical space. Identifying different combinations through wet experiments will consume a huge amount of time, financial resources and material resources. Past polypeptide drug design often relied on identifying natural functional polypeptides or obtaining polypeptide fragments from functional proteins based on expert experience. This discovery and identification method has slowed down the efficiency of polypeptide drug development. With the development of artificial intelligence, AI technology has provided new tools and methods for drug discovery. AI can quickly identify potential drug targets and candidate drug molecules from a large amount of biological data through deep learning and machine learning algorithms, further accelerating the efficiency of drug development. Summary of the Invention
[0003] In order to solve the technical problems existing in the toxicity prediction of polypeptide drugs in the prior art, the present invention provides a method and system for generating candidate polypeptide drugs based on artificial intelligence.
[0004] The present invention is realized through the following technical solutions:
[0005] A method for generating candidate polypeptide drugs based on artificial intelligence, comprising:
[0006] S1: Collect sequence data, structure data and related property labels of polypeptide drug molecules and their corresponding target proteins, where the related properties include functions and chemical and physical properties;
[0007] S2: Utilize the interaction site information between the polypeptide and its target, fix the structure of the target protein, and generate the backbone of the polypeptide using a diffusion model;
[0008] S3: Reverse-fold the generated polypeptide backbone to generate new candidate polypeptide molecules and analyze the prediction results.
[0009] Further, the data collected in step S1 includes sequence data and structure data of existing drug molecules and target proteins, sequence data and bioactivity data of experimentally identified functional polypeptides.
[0010] Further, the process of constructing the diffusion model in step S2 includes:
[0011] S21. Data preprocessing: Obtain relevant protein interaction data, and extract hydrophobicity, surface charge, affinity, activity and spatial geometry information by processing the interaction site information;
[0012] S22. Model the target protein using a graph convolutional neural network;
[0013] S23. Construct a diffusion model: Fix the target protein and the binding protein in geometric space and gradually add noise to it;
[0014] S24. Multi-task learning. Introduce a multi-task learning framework during the training of the diffusion model to predict binding affinity, geometric structure, and activity simultaneously.
[0015] Furthermore, in step S22, through a heterogeneous graph convolutional neural network, using Cα atoms as nodes and chemical bonds between adjacent atoms as edges for modeling, and through an information propagation mechanism, transmit the edge features composed of different chemical bonds to the nodes, thereby capturing the key features of the protein structure and outputting a feature vector.
[0016] Furthermore, the diffusion formula of the diffusion model in step S23 is as follows:
[0017]
[0018]
[0019] Among them, is the polypeptide backbone structure at the t-th diffusion step, is the time step control parameter, is the polypeptide backbone structure at the (t - 1)-th diffusion step, is the noise updated according to neighboring nodes, is the neural network, is the feature of node i, is the updated feature of the neighboring nodes of node i.
[0020] Furthermore, the inverse folding in step S3 includes:
[0021] S31. Data preprocessing: Obtain the sequence and structure information of relevant polypeptides, introduce polypeptide dynamic structure simulation data, and use molecular dynamics simulation to extract the structural change characteristics of polypeptides in different environments;
[0022] S32. Obtain the feature vector of the polypeptide sequence. Use a word embedding model to map each amino acid to a feature vector, capture the semantic relationship between amino acids, and map it to a polypeptide sequence feature vector with a fixed dimension;
[0023] S33: Obtain the structural feature vector. Input the dynamic structural features of the polypeptide into a heterogeneous graph convolutional neural network. The heterogeneous graph convolutional neural network inputs the atomic types, positions, and chemical bonds of the polypeptide backbone and adopts information edge transfer;
[0024] S34. Input the feature vector and structural feature vector of the polypeptide sequence into the neural network of the Transformer architecture to predict the three-dimensional coordinates of C, H, O, and N of the polypeptide and learn the mapping relationship from structure to sequence;
[0025] S35. Use the Monte Carlo algorithm to decode the latent vectors generated from the polypeptide backbone to generate new polypeptide sequences;
[0026] S36. Use the automatic learning strategy to select sequences with higher information entropy for experimental iteration, calculate the feature representation and predicted output distribution for each sequence, and calculate the uncertainty.
[0027] Further, the update formula in step S33 is:
[0028]
[0029] where, is the feature of node i in the l-th layer, is the feature of node i in the (l + 1)-th layer, is the adjacency matrix, representing the connection between nodes i and j, is the trainable parameter in the l-th layer, is the Sigmod activation, where i and j represent node i and node j respectively, is the set of adjacent nodes of node i.
[0030] Further, the description of learning the mapping relationship from structure to sequence is as follows:
[0031] Loss =
[0032] where, represents the th amino acid, N represents the length of the polypeptide chain, and stur represents the structure of the polypeptide.
[0033] The present invention also provides an artificial intelligence-based candidate polypeptide drug generation system, based on the above-mentioned artificial intelligence-based polypeptide drug toxicity prediction method, which includes:
[0034] A data acquisition module, which is used to collect sequence data, structural data, and related property labels of polypeptide drug molecules and their corresponding target proteins;
[0035] A polypeptide backbone generation module, which is used to utilize the interaction site information between the polypeptide and its target, fix the structure of the target protein, and generate the backbone of the polypeptide using the diffusion model;
[0036] A polypeptide backbone inverse folding module, which is used to perform inverse folding on the generated polypeptide backbone to generate new candidate polypeptide molecules and analyze the prediction results.
[0037] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, on which program instructions for a method for generating candidate polypeptide drugs based on artificial intelligence are stored. The program instructions for the method for generating candidate polypeptide drugs based on artificial intelligence can be executed by one or more processors to implement the steps of the method for generating candidate polypeptide drugs based on artificial intelligence as described above.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0039] The present invention realizes accurate, efficient, and low-cost rapid identification and generation of potential candidate drug molecules, effectively solves the structural information of proteins and polypeptides, solves the problem that it is difficult to model polypeptides in the free state, and adjusts the diffusion process in combination with the physical properties of proteins to make it conform to the constraints of the actual chemical environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0041] Figure 1 is a schematic flowchart of a method for generating candidate polypeptide drugs based on artificial intelligence according to an embodiment of the present application;
[0042] Figure 2 is a schematic diagram of Monte Carlo simulation sampling according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0044] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention can also be implemented or applied through other different specific implementation manners, and various modifications or changes can be made to the details in this specification based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The form, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the layout form of its components may also be more complex.
[0046] See Figure 1 , a method for generating candidate polypeptide drugs based on artificial intelligence, comprising:
[0047] S1: Collect sequence data, structural data, and related property tags of polypeptide drug molecules and their corresponding target proteins, where the related properties include functions and chemical and physical properties;
[0048] Specifically, the collected data includes sequence data and structural data of existing drug molecules and target proteins, sequence data of experimentally identified functional polypeptides, and bioactivity data;
[0049] S2: Utilize the interaction site information between polypeptides and their targets, fix the structure of the target protein, and generate the backbone of the polypeptide using a diffusion model to improve the generation effect of the model;
[0050] The construction process of the diffusion model is as follows:
[0051] S21. Data preprocessing: Obtain relevant protein interaction data, and extract hydrophobicity, surface charge, affinity, activity, and spatial geometry information by processing the interaction site information;
[0052] S22. Model the target protein using a graph convolutional neural network: Fully consider the differences in the properties and side chain attributes of different amino acids in the protein structure. Through a heterogeneous graph convolutional neural network, use the Cα atom as a node and the chemical bonds between adjacent atoms as edges for modeling, and through an information propagation mechanism, transfer the edge features composed of different chemical bonds to the nodes to capture the key features of the protein structure and output the polypeptide graph feature vector. This method can effectively solve the structural information of proteins and polypeptides and solve the problem of difficult modeling of polypeptides in the free state;
[0053] S23. Construct a diffusion model: Fix the target protein and the binding protein in geometric space and gradually add noise to it. The edge weight weighted diffusion method is introduced in the addition process, and the physical properties of the binding protein (such as charge distribution and bond energy) are used to adjust the diffusion process to make it conform to the constraints of the actual chemical environment. The formula for the diffusion process is:
[0054]
[0055]
[0056] Among them, is the polypeptide backbone structure at the t-th diffusion step, is the time step control parameter, is the polypeptide backbone structure at the (t - 1)-th diffusion step, is the noise updated according to neighboring nodes, is the neural network, is the feature of node i, is the updated feature of the neighboring nodes of node i.
[0057] Introduce the conditional feature c into the diffusion model, and guide the generation of the polypeptide structure according to the local characteristics of the target protein. During the denoising process, calculate the noise prediction value through the heterogeneous graph convolutional neural network and the conditional feature c:
[0058]
[0059] Among them, is the graph feature of the target protein, is the graph feature of the polypeptide, is the polypeptide backbone structure at the t-th diffusion step, represents the heterogeneous graph convolutional neural network model.
[0060] S24. Multi-task learning: To simultaneously predict binding affinity, geometric structure, and activity, introduce a multi-task learning framework in the training of the diffusion model. The loss function is in the form of multi-objective optimization, and the expression of the loss function is as follows:
[0061]
[0062] Among them, , , are the weights of different tasks, used to balance the influence of multiple objectives and fine-tune for the target later, is the binding site prediction loss, is the loss for binding geometric structure prediction, is the binding energy prediction loss.
[0063] is used to measure the difference between the predicted affinity of the model and the true affinity. The mean square error (MSE) loss function is adopted:
[0064]
[0065] Among them, N represents the number of samples, represents the th molecule, represents the th molecule's predicted affinity, represents the The experimental affinity of a molecule, i.e., the Kd value.
[0066]
[0067] Where M is the number of atoms, and represent the predicted and true atomic coordinates respectively, K is the number of bonds, and are the predicted and true bond lengths respectively.
[0068] It is used to measure the difference between the predicted activity of the model and the true activity. The mean squared error (MSE) loss function is adopted:
[0069]
[0070] Where N represents the number of samples, represents the th molecule, represents the predicted activity of the th molecule, represents the experimental activity (IC50 value) of the th molecule.
[0071] After training is completed, the inverse process of the diffusion model starts from the geometric features of the target protein and gradually denoises the Gaussian noise at the active site of the target protein to generate the binding polypeptide of the target protein. The relevant description is as follows. At time t, the model predicts the denoised state based on the current state and the geometric features of the target protein :
[0072]
[0073] Where is the trained model.
[0074] S3: Reverse fold the generated polypeptide backbone to generate new candidate polypeptide molecules and analyze the prediction results.
[0075] The reverse folding process is as follows:
[0076] S31. Data preprocessing: Obtain the sequence and structure information of the relevant polypeptide, introduce the polypeptide dynamic structure simulation data, and use molecular dynamics simulation to extract the structural change characteristics of the polypeptide in different environments. Assume the dynamic conformation of the polypeptide is , and the time series structure calculated by molecular dynamics simulation is:
[0077]
[0078] Among them, are the coordinates of the i-th Cα atom, N is the amino acid length of the polypeptide, and through dimensionality reduction or clustering, dynamic structure features are extracted as the input for the subsequent model.
[0079] S32. Obtain the feature vector of the polypeptide sequence. Using the word embedding model, map each amino acid to a feature vector to capture the semantic relationship between amino acids and map it to a polypeptide sequence feature vector of a fixed dimension.
[0080] S33: Obtain the structure feature vector. Input the dynamic structure features of the polypeptide into the heterogeneous graph convolutional neural network. The heterogeneous graph convolutional neural network inputs the atomic types, positions, and chemical bonds of the polypeptide backbone and adopts information edge transmission. The specific update formula is:
[0081]
[0082] Among them, is the feature of node i in the l-th layer, is the feature of node i in the (l + 1)-th layer, is the adjacency matrix, representing the connection between nodes i and j, is the trainable parameter in the l-th layer, is the Sigmod activation, where i and j represent nodes i and j respectively, is the set of adjacent nodes of node i.
[0083] S34. Input the feature vector of the polypeptide sequence and the structure feature vector into the neural network of the Transformer architecture to predict the three-dimensional coordinates of C, H, O, and N of the polypeptide and learn the mapping relationship from structure to sequence. The relevant expression is as follows:
[0084] Loss =
[0085] Among them, represents the th amino acid, N represents the length of the polypeptide chain, and stur represents the structure of the polypeptide.
[0086] S35. Use the Monte Carlo algorithm to decode the latent vectors generated from the polypeptide backbone obtained in S2 to generate a new polypeptide sequence.
[0087] S36. Use the automatic learning strategy to select sequences with higher information entropy for experimental iteration, calculate the feature representation and the predicted output distribution ( ) for each sequence, and calculate the uncertainty:
[0088]
[0089] Among them, represents the feature vector of the sequence, represents the predicted th category, is the predicted probability of the polypeptide affinity predicted by the model.
[0090] Calculate the uncertainty of the newly generated polypeptide, select the polypeptides with top rankings for biological experiment verification, and use the bioactivity information obtained from the biological experiments for model parameter tuning.
[0091] Preferably, the data set in step S31 is used for model training, parameter tuning and evaluation. Training is carried out using the training set of polypeptide sequences and structures to learn the mapping relationship between the sequences and structures of bioactive polypeptides.
[0092] After training is completed, the model generates a latent vector representation by inputting the backbone of the bioactive polypeptide. These latent vectors capture the key feature information of the polypeptide sequence. Further, by operating on and sampling the latent vectors, new polypeptide sequences are generated, and these sequences theoretically have bioactive characteristics similar to or better than the input polypeptide.
[0093] Preferably, in step S35, the Monte Carlo sampling method is used to sample the latent vectors of the model to generate new polypeptide sequences. By sampling different latent vectors, the latent space of the polypeptide sequences can be explored, thereby generating polypeptide sequences with different bioactivities.
[0094] In this embodiment, the present invention realizes accurate, efficient and low-cost rapid identification and generation of potential candidate drug molecules, effectively solves the structural information of proteins and polypeptides, solves the problem that it is difficult to model polypeptides in the free state, and adjusts the diffusion process in combination with the physical properties of proteins to make it conform to the constraints of the actual chemical environment.
[0095] The embodiment of the present invention also proposes an artificial intelligence-based candidate polypeptide drug generation system, based on the above-mentioned artificial intelligence-based candidate polypeptide drug generation method, including:
[0096] A data acquisition module, which is used to acquire the sequence data, structure data and related property labels of polypeptide drug molecules and their corresponding target proteins;
[0097] A polypeptide backbone generation module, which is used to utilize the interaction site information between the polypeptide and its target, fix the structure of the target protein, and generate the backbone of the polypeptide using a diffusion model;
[0098] A polypeptide backbone inverse folding module, which is used to perform inverse folding on the generated polypeptide backbone to generate new candidate polypeptide molecules and analyze the prediction results.
[0099] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which program instructions of a method for generating a candidate polypeptide drug based on artificial intelligence are stored. The program instructions of the method for generating a candidate polypeptide drug based on artificial intelligence can be executed by one or more processors to implement the steps of the method for generating a candidate polypeptide drug based on artificial intelligence as described above.
[0100] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for generating candidate peptide drugs based on artificial intelligence, characterized in that: include: S1: Collect sequence data, structural data and related property labels of peptide drug molecules and their corresponding target proteins, wherein the related properties include functional and chemical properties; S2: Using the interaction site information of the peptide and its target, the target protein structure is fixed and the peptide backbone is generated using a diffusion model; S3: Reverse fold the generated polypeptide backbone to generate new candidate polypeptide molecules, and analyze the prediction results.
2. The method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to claim 1, characterized in that: The data collected in step S1 include sequence data and structural data of existing drug molecules and target proteins, and sequence data and biological activity data of experimentally identified functional polypeptides.
3. The method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to claim 1, characterized in that: The diffusion model building process in step S2 includes: S21. Data preprocessing: Obtain relevant protein interaction data and extract hydrophobicity, surface charge, affinity, activity and spatial geometry information by processing the interaction site information; S22, Modeling target proteins using graph convolutional neural networks; S23, build a diffusion model: fix the target protein and the binding protein in the geometric space and gradually add noise to it; S24. Multi-task learning, introduces a multi-task learning framework in the training of diffusion models to simultaneously predict binding affinity, geometry and activity.
4. The method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to claim 3, characterized in that: In step S22, a heterogeneous graph convolutional neural network is used to model the Cα atoms as nodes and the chemical bonds between adjacent atoms as edges. The edge features composed of different chemical bonds are transferred to the nodes through an information propagation mechanism, thereby capturing the key features of the protein structure and outputting a polypeptide graph feature vector.
5. The method for predicting polypeptide drug toxicity based on artificial intelligence according to claim 3, characterized in that: The diffusion formula of the diffusion model in step S23 is as follows: , ,in, is the peptide backbone structure at the diffusion step t, is the time step control parameter, is the peptide backbone structure at the diffusion step t-1, is the noise updated according to the neighboring nodes, is a neural network, is the characteristic of the i-node, It is the update feature of the neighbor nodes of node i.
6. The method for predicting peptide drug toxicity based on artificial intelligence according to claim 1, characterized in that: The reverse folding in step S3 includes: S31. Data preprocessing: Obtain the sequence and structural information of related peptides, introduce the dynamic structure simulation data of peptides, and use molecular dynamics simulation to extract the structural change characteristics of peptides in different environments; S32, obtaining a feature vector of the polypeptide sequence, using a word embedding model to map each amino acid into a feature vector, capturing the semantic relationship between the amino acids, and mapping it into a polypeptide sequence feature vector of fixed dimension; S33: Obtaining a structural feature vector, inputting the dynamic structural features of the peptide into a heterogeneous graph convolutional neural network, which inputs the atomic types, positions, and chemical bonds of the peptide skeleton and uses information edge transfer; S34, input the feature vector of the peptide sequence and the structural feature vector into the neural network of the Transformer architecture, predict the three-dimensional coordinates of the C, H, O, and N of the peptide, and learn the mapping relationship from structure to sequence; S35, decoding the latent vector generated from the polypeptide backbone using a Monte Carlo algorithm to generate a new polypeptide sequence; S36. Use the automatic learning strategy to select sequences with higher information entropy for experimental iterations, calculate the feature representation and predicted output distribution for each sequence, and calculate the uncertainty.
7. The method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to claim 1, characterized in that: The updating formula in step S33 is: ,in, is the feature of node i in layer l, is the feature of node i in the l+1th layer, is the adjacency matrix, which represents the connection between nodes i and j. is the trainable parameter of the lth layer, is Sigmod activation, i, j represent nodes i and j respectively, is the set of adjacent nodes of node i.
8. The method for predicting polypeptide drug toxicity based on artificial intelligence according to claim 6, characterized in that: The mapping relationship from structure to sequence is described as follows: Loss = ,in, Representative amino acids, N represents the length of the polypeptide chain, and stur represents the structure of the polypeptide.
9. An artificial intelligence-based candidate peptide drug generation system, based on the artificial intelligence-based candidate peptide drug generation method according to any one of claims 1 to 8, comprising: A data acquisition module, which is used to collect sequence data, structural data and related property labels of peptide drug molecules and their corresponding target proteins; A peptide skeleton generation module, which is used to utilize the interaction site information between the peptide and its target, fix the target protein structure, and generate the peptide skeleton using a diffusion model; The polypeptide backbone reverse folding module is used to reverse fold the generated polypeptide backbone to generate new candidate polypeptide molecules and analyze the prediction results.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions for a method for generating candidate polypeptide drugs based on artificial intelligence, and the program instructions for the method for generating candidate polypeptide drugs based on artificial intelligence can be executed by one or more processors to implement the steps of the method for generating candidate polypeptide drugs based on artificial intelligence as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Protein engineering system and platform based on large language model
CN116631514A
Protein reverse folding design method and electronic equipment
CN116682487A
Discrete graph probability denoising diffusion model for protein sequence generation
CN116994642A
Construction method and device of hyperbolic discrete diffusion model of three-dimensional RNA structure inverse folding
CN119339781A
Protein design using diffusion models operating on full atom representations
WO2024240774A1
Cited By
Molecular optimization method, system and device based on neural network and storage medium
CN121237260A