Polypeptide drug toxicity prediction method and system based on artificial intelligence
Through artificial intelligence-based methods, a polypeptide toxicity prediction model is constructed, combined with kinetic modeling and word embedding technology, the time-consuming and cost-effectiveness of traditional polypeptide drug toxicity detection technology is solved, and the high accuracy and efficiency of polypeptide toxicity prediction is achieved.
Patent Information
- Application Number
- CN202510274647.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Traditional polypeptide drug toxicity detection techniques are time-consuming and difficult to quickly and effectively identify the toxicity of polypeptides. Especially due to the complexity of the peptide and the uncertainty in the metabolism process, the toxicity prediction task is full of challenges.
Using an artificial intelligence-based method, a Bayesian regression model is used to predict the rate constant of peptide metabolism by collecting sequence data and properties of peptide molecules, and a convolutional neural network is used to construct a toxicity prediction model, combining kinetic modeling to calculate the rate of change of concentration over time, and predict the toxicity of peptides and their hydrolysates.
Through the combination of intelligent word segmentation, word embedding vectors and kinetic modeling, the peptide information is extracted more accurately, achieving high accuracy prediction of peptide toxicity and reducing the time and cost of toxicity assessment.
Smart Images

Figure CN120183537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of drug screening, and in particular to a method and system for predicting the toxicity of polypeptide drugs based on artificial intelligence. Background Art
[0002] Drug screening is a complex and time-consuming process, and the toxicity assessment of drugs is one of the key steps. Traditional compound toxicity detection techniques usually require biochemical tests, cell experiments, and even animal models. These methods not only consume a large amount of time but also are costly. The toxicity of a drug determines whether it can pass experiments and reviews. Rapid and effective identification of the toxicity of polypeptides is a key step in the development of polypeptide drugs. However, due to the complexity and diversity of polypeptides, as well as the uncertainty of the hydrolysis products generated during their metabolism in the body, this prediction task remains challenging. Summary of the Invention
[0003] In order to solve the technical problems existing in the prediction of the toxicity of polypeptide drugs in the prior art, the present invention provides a method and system for predicting the toxicity of polypeptide drugs based on artificial intelligence.
[0004] The present invention is achieved by the following technical solutions:
[0005] A method for predicting the toxicity of polypeptide drugs based on artificial intelligence includes:
[0006] S1: Collect the sequence data and properties of polypeptide molecules; the sequence data and properties of the polypeptide molecules include the sequence of the polypeptide molecules, relevant bioactivity data, physicochemical property parameters, known toxicity information, and possible metabolite data;
[0007] S2: Construct a prediction model for the rate constant of polypeptide metabolism based on the Bayesian regression model;
[0008] S3: Construct a toxicity prediction model to predict the toxicity scores of polypeptide molecules and their hydrolysis products, including extracting fragment sequences from the polypeptide sequence to generate sequence features of polypeptides and polypeptide fragments; inputting the selected feature vectors into the trained toxicity prediction model to predict polypeptide toxicity, constructing a kinetic simulation-based polypeptide toxicity prediction model, and obtaining the toxicity prediction results of polypeptide molecules and their hydrolysis products.
[0009] Further, in step S2, the prediction model for the rate constant of polypeptide metabolism constructs the relationship between the molecular weight, hydrophobicity, molecular fragment descriptors, charge distribution, number of hydrogen bond donors and acceptors and the metabolic rate constant of polypeptides and their metabolites.
[0010] Further, step S2 also includes feature selection and importance evaluation. Using feature importance evaluation, key features in the prediction of the metabolic rate constant are screened out, and the contribution of each feature to the model prediction is quantified.
[0011] Furthermore, step S3 further includes:
[0012] S31: Data preprocessing, including obtaining a drug data set, and by processing the original biochemical test data, extracting molecular weight, hydrophobicity, molecular fragment descriptors, charge distribution, polypeptide sequence, polypeptide toxicity data, metabolite data, and the number of secondary structures of the polypeptide;
[0013] S32. Perform word segmentation on the polypeptide sequence, use a sliding window of length T to obtain polypeptide fragments, count the frequencies of the fragments to form a list of key fragments, and construct the feature space of the sequence;
[0014] S33. Use a word embedding model to represent the amino acid sequence, map each amino acid to a high-dimensional vector, capture the semantic relationships and interactions between amino acids, and obtain the feature vector of the polypeptide sequence;
[0015] S34. Based on kinetic modeling, calculate the rate of change of concentration over time, use a convolutional neural network to construct a polypeptide toxicity prediction model, and establish the relationship between the polypeptide and toxicity.
[0016] Furthermore, when calculating the rate of change of concentration over time based on kinetic modeling, it is assumed that the metabolic pathway of polypeptide P in the body conforms to first-order reaction kinetics, and the rate constant is predicted by the model in S2, which is expressed as:
[0017] i = 1, 2, …, n
[0018] where, [p] and are the concentrations of the polypeptide and the polypeptide metabolite respectively, k and are the rate constants of the polypeptide and the polypeptide metabolite respectively, is the rate of change of the polypeptide concentration [p] over time t, is the rate of change of the metabolite concentration over time.
[0019] At time t, the metabolite concentration is = , the peak time is , substituting back into the formula gives the peak concentration .
[0020] Furthermore, when using a convolutional neural network to construct a polypeptide toxicity prediction model and establish the relationship between the polypeptide and toxicity, it includes: According to the relationship between the metabolite concentration and toxicity, the toxicity contribution can be expressed as:
[0021]
[0022] where, is the peak concentration of the metabolite, is the model weight, is the non - linear activation function, and i is the serial number of the polypeptide metabolite.
[0023] Further, the step S34 further includes using the data set in S31 to divide into a training set, a validation set, and a test set for the training, parameter tuning, and evaluation of the model. Use the labeled polypeptide toxicity training set for training to learn the sequence features and combination relationships of polypeptide toxicity.
[0024] Further, after the training in the step S34 is completed, use the validation set to evaluate the trained model, use the test set to evaluate the performance of the model on unseen data, test its accuracy and generalization ability in predicting polypeptide toxicity, and optimize and improve the model according to the result analysis.
[0025] The present invention also provides an artificial - intelligence - based polypeptide drug toxicity prediction system, based on an artificial - intelligence - based polypeptide drug toxicity prediction method as described above, which includes:
[0026] A data acquisition module, which is used to acquire the sequence data and property labels of polypeptide molecules and identify the sequence combination features in polypeptide drug molecules;
[0027] A polypeptide metabolism rate constant prediction module, which is used to predict the rate constant of polypeptide metabolism;
[0028] A toxicity prediction module, which is used to extract fragment sequences from the polypeptide sequence to generate the sequence features of polypeptides and polypeptide fragments; input the selected feature vectors into the trained toxicity prediction model to predict drug properties, construct a machine - learning model for kinetic simulation, and obtain the toxicity prediction results of polypeptide molecules and their hydrolysis products.
[0029] In addition, to achieve the above - mentioned purpose, the present invention also provides a computer - readable storage medium. Program instructions of an artificial - intelligence - based polypeptide drug toxicity prediction method are stored on the computer - readable storage medium. The program instructions of the artificial - intelligence - based polypeptide drug toxicity prediction method can be executed by one or more processors to implement the steps of an artificial - intelligence - based polypeptide drug toxicity prediction method as described above.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] By combining intelligent word segmentation and context statistical information, and simultaneously combining the word embedding vectors and kinetic modeling of polypeptides, the information of polypeptides is fully extracted, and the advantages of both are combined to more accurately predict the toxicity of polypeptides. At the same time, by predicting the overall toxicity, the corresponding amino acid combination information is extracted from the context word frequency information, making the prediction result of the toxic fragment more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0033] Figure 1 is a schematic flowchart of a method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to an embodiment of the present application;
[0034] Figure 2 is a schematic diagram of dynamic input and fragment prediction according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0036] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0037] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention and are not drawn according to the number, shape, and size of the components in actual implementation. The actual form, number, and ratio of each component during actual implementation can be arbitrarily changed, and the component layout form may also be more complex.
[0038] See Figure 1 , a method for predicting the toxicity of polypeptide drugs based on artificial intelligence, comprising the following steps:
[0039] S1: Collect the sequence data and property labels of polypeptide molecules to ensure the accuracy and integrity of the data;
[0040] The sequence data and properties of the polypeptide molecule include the sequence of the polypeptide molecule, related bioactivity data, physicochemical property parameters, known toxicity information, and possible metabolite data.
[0041] S2: Construct a prediction model for the rate constant of polypeptide metabolism, including:
[0042] S21. Feature extraction: Extract the physicochemical properties of the molecule through the molecular feature generation tool RDKit, including molecular weight, hydrophobicity, molecular fragment descriptors, charge distribution, number of hydrogen bond donors and acceptors;
[0043] S22. Model construction: Use the Bayesian regression model to construct the relationship between molecular weight, hydrophobicity, molecular fragment descriptors, charge distribution, number of hydrogen bond donors and acceptors and the metabolic rate constants of the polypeptide and its metabolites. For the regression task of the rate constant, the optimization objective is to minimize the mean squared error, and its mathematical expression is:
[0044]
[0045] where N is the number of samples, is the true rate constant of the i-th sample, is the predicted value of the i-th sample.
[0046] S23. Feature selection and importance evaluation: Use feature importance evaluation to screen out the key features in the prediction of the metabolic rate constant. Use SHAP analysis to quantify the contribution of each feature to the model prediction, identify the key features and improve the interpretability of the model.
[0047] S24. Calculate the rate constant through Bayesian regression to provide input parameters for kinetic modeling.
[0048] S3: Construct a toxicity prediction model to predict the toxicity scores of polypeptide molecules and their hydrolysis products, including:
[0049] Extract fragment sequences from the polypeptide sequence to generate sequence features of the polypeptide and polypeptide fragments; input the selected feature vectors into the trained toxicity prediction model to predict polypeptide toxicity, construct a polypeptide toxicity prediction model for kinetic simulation, and obtain the toxicity prediction results of polypeptide molecules and their hydrolysis products. Specifically, it includes the following steps:
[0050] S31. Data preprocessing: Obtain the drug data set, and extract molecular weight, hydrophobicity, molecular fragment descriptors, charge distribution, polypeptide sequence, polypeptide toxicity data, metabolite data, and the number of secondary structures of the polypeptide by processing the original biochemical test data;
[0051] S32. Perform word segmentation on the polypeptide sequence, obtain polypeptide fragments using a sliding window of length T, count the frequencies of the fragments to form a list of key fragments, and construct the feature space of the sequence. This process captures the information in the polypeptide sequence by extracting global and local sequence features. Optionally, T is selected as 3.
[0052] It also includes statistically analyzing the sequences with high word frequencies by performing word segmentation on the sequences. The core process is as follows:
[0053] a. Traverse each position in the sequence, sample using a window of length T, and perform sliding window sampling from the N-terminus to the C-terminus of the polypeptide sequence;
[0054] b. Generate candidate field information, store the candidate fields, and perform statistical information;
[0055] c. Repeat the above operations on new sequences, maintain a dictionary of statistical information until there are no new sequences in the training set.
[0056] d. Set a threshold, remove fragment combinations that appear less than the threshold from the dictionary, and count the probability information.
[0057] S33. Use a word embedding model to represent the amino acid sequence, map each amino acid to a high-dimensional vector, capture the semantic relationships and interactions between amino acids, and obtain the feature vector of the polypeptide sequence.
[0058] It also includes: Combining the word segmentation method in S32, when the length of the polypeptide sequence is n and the window length is 3, generate m - 2 subsequences, generate a potential representation of length 3, and use 0 padding in the potential representation of the subsequences to make the sequence length n. Concatenate the potential representations of the segmented subsequences and the potential representation of the polypeptide sequence so that the model can obtain the features of both the polypeptide sequence and the polypeptide fragments when extracting polypeptide sequence information.
[0059] S34. Calculate the rate of change of concentration with time based on kinetic modeling. Assume that the metabolic pathway of polypeptide P in the body conforms to first-order reaction kinetics, and the rate constant is predicted by the model in S2. Its expression is:
[0060] i = 1, 2, …, n
[0061] where, [p] and are the concentrations of the polypeptide and the polypeptide metabolite respectively, k and are the rate constants of the polypeptide and the polypeptide metabolite respectively, is the rate of change of the polypeptide concentration [p] with time t, is the concentration of the metabolite The rate of change with time.
[0062] At time t, the metabolite concentration is = , the peak time is , substituting back into the formula gives the peak concentration .
[0063] Furthermore, in constructing the polypeptide toxicity prediction model using a convolutional neural network, establishing the relationship between polypeptides and toxicity includes: According to the relationship between metabolite concentration and toxicity, the toxicity contribution can be expressed as:
[0064]
[0065] where is the peak concentration of the metabolite, is the model weight, is the non-linear activation function, and i is the serial number of the polypeptide metabolite.
[0066] Using the dataset in S31, divide it into a training set, a validation set, and a test set for model training, parameter tuning, and evaluation. Use the labeled polypeptide toxicity training set for training to learn the sequence features and combined relationships of polypeptide toxicity.
[0067] After training is completed, use the validation set to evaluate the trained model, and use the test set to evaluate the performance of the model on unseen data, test its accuracy and generalization ability in predicting polypeptide toxicity, and optimize and improve the model according to the result analysis.
[0068] In step S34, the toxicity scores of polypeptides and polypeptide hydrolysis products are predicted through the prediction model. Combining the word frequency information obtained from step 2 with the prediction results of the toxicity scores, deduce the combined fragments that cause toxicity.
[0069] In this embodiment, the present invention combines intelligent word segmentation and context statistical information, and at the same time combines the word embedding vectors of polypeptides, so that the information of polypeptides can be fully extracted and the advantages of both can be combined, and the prediction of polypeptide toxicity is more accurately realized. At the same time, by predicting the overall toxicity, the corresponding amino acid combination information is extracted from the context word frequency information, making the prediction results of toxicity fragments more accurate.
[0070] The embodiment of the present invention also proposes an artificial intelligence-based polypeptide drug toxicity prediction system, based on the above-mentioned artificial intelligence-based polypeptide drug toxicity prediction method, including:
[0071] A data acquisition module, which is used to acquire the sequence data and property labels of polypeptide molecules and identify the sequence combination characteristics in polypeptide drug molecules;
[0072] A polypeptide metabolism rate constant prediction module, which is used to predict the rate constant of polypeptide metabolism;
[0073] A toxicity prediction module, which is used to extract fragment sequences from a polypeptide sequence to generate sequence features of the polypeptide and polypeptide fragments; input the selected feature vectors into a trained toxicity prediction model to predict drug properties, construct a machine learning model for kinetic simulation, and obtain the toxicity prediction results of the polypeptide molecule and its hydrolysis products.
[0074] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which program instructions for the method for predicting the toxicity of polypeptide drugs based on artificial intelligence are stored. The program instructions for the method for predicting the toxicity of polypeptide drugs based on artificial intelligence can be executed by one or more processors to implement the steps of the method for predicting the toxicity of polypeptide drugs based on artificial intelligence as described above.
[0075] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for predicting the toxicity of peptide drugs based on artificial intelligence, characterized in that: include: S1: Collecting sequence data and properties of polypeptide molecules; the sequence data and properties of polypeptide molecules include the sequence of polypeptide molecules, relevant biological activity data, physicochemical property parameters, known toxicity information and possible metabolite data; S2: Construct a rate constant prediction model for peptide metabolism based on Bayesian regression model; S3: Construct a toxicity prediction model to predict the toxicity scores of polypeptide molecules and their hydrolysis products, including extracting fragment sequences from polypeptide sequences to generate sequence features of polypeptides and polypeptide fragments; input the selected feature vectors into the trained toxicity prediction model to predict polypeptide toxicity, construct a polypeptide toxicity prediction model for kinetic simulation, and obtain toxicity prediction results for polypeptide molecules and their hydrolysis products.
2. The method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to claim 1, characterized in that: The rate constant prediction model for polypeptide metabolism in step S2 constructs the relationship between molecular weight, hydrophobicity, molecular fragment descriptors, charge distribution, the number of hydrogen bond donors and acceptors, and the metabolic rate constants of the polypeptide and its metabolites.
3. The method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to claim 2, characterized in that: The step S2 also includes feature selection and importance assessment, and the key features in the prediction of metabolic rate constants are screened out by feature importance assessment, and the contribution of each feature to the model prediction is quantified.
4. The method for predicting polypeptide drug toxicity based on artificial intelligence according to claim 1, characterized in that: The step S3 further comprises: S31: Data preprocessing, including obtaining drug data sets, extracting molecular weight, hydrophobicity, molecular fragment descriptors, charge distribution, peptide sequence, peptide toxicity data, metabolite data, and the number of secondary structures of peptides by processing raw biochemical test data; S32, performing a word segmentation operation on the polypeptide sequence, using a sliding window of length T to obtain polypeptide fragments, counting the frequencies of the fragments to form a key fragment list, and constructing a feature space of the sequence; S33. Use the word embedding model to characterize the amino acid sequence, map each amino acid into a high-dimensional vector, capture the semantic relationship and interaction between amino acids, and obtain the feature vector of the polypeptide sequence; S34. Calculate the rate of change of concentration over time based on kinetic modeling, use convolutional neural networks to build a peptide toxicity prediction model, and construct the relationship between peptides and toxicity.
5. The method for predicting polypeptide drug toxicity based on artificial intelligence according to claim 4, characterized in that: The calculation of the change rate of concentration over time based on kinetic modeling assumes that the metabolic pathway of polypeptide P in the body conforms to first-order reaction kinetics, and the rate constant is predicted by the S2 model, which is expressed as: i=1,2,…,n Among them, [p] and are the concentrations of peptide and peptide metabolite, k and are the rate constants of peptide and peptide metabolite, respectively, is the rate at which the peptide concentration [p] changes with time t, is the concentration of metabolites speed of change over time; At time t, the concentration of metabolites is = , the peak time is , peak concentration .
6. The method for predicting polypeptide drug toxicity based on artificial intelligence according to claim 4, characterized in that: The use of convolutional neural networks to construct a peptide toxicity prediction model and construct the relationship between peptides and toxicity includes: according to the relationship between metabolite concentration and toxicity, the toxicity contribution can be expressed as: ,in, is the peak concentration of metabolites, is the model weight, is a nonlinear activation function, and i is the serial number of the polypeptide metabolite.
7. The method for predicting the toxicity of polypeptide drugs based on artificial intelligence according to claim 1, characterized in that: The step S34 also includes using the data set in S31 to divide the training set, validation set and test set for model training, parameter adjustment and evaluation, using the labeled polypeptide toxicity training set for training, and learning the sequence characteristics and combination relationships of polypeptide toxicity.
8. The method for predicting polypeptide drug toxicity based on artificial intelligence according to claim 4, characterized in that: The step S34 also includes, after the training is completed, using the validation set to evaluate the trained model, using the test set to evaluate the performance of the model on unseen data, testing its accuracy and generalization ability in predicting polypeptide toxicity, and tuning and improving the model based on the analysis of the results.
9. An artificial intelligence-based peptide drug toxicity prediction system, based on the artificial intelligence-based peptide drug toxicity prediction method according to any one of claims 1 to 8, comprising: A data acquisition module is used to collect sequence data and property labels of polypeptide molecules and identify sequence combination features in polypeptide drug molecules; A rate constant prediction module for polypeptide metabolism, which is used to predict the rate constant of polypeptide metabolism; The toxicity prediction module is used to extract fragment sequences from polypeptide sequences and generate sequence features of polypeptides and polypeptide fragments; the selected feature vectors are input into the trained toxicity prediction model to predict drug properties, and a machine learning model for kinetic simulation is constructed to obtain toxicity prediction results for polypeptide molecules and their hydrolysis products.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions of a method for predicting the toxicity of a polypeptide drug based on artificial intelligence, and the program instructions of the method for predicting the toxicity of a polypeptide drug based on artificial intelligence can be executed by one or more processors to implement the steps of the method for predicting the toxicity of a polypeptide drug based on artificial intelligence as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method for predicting the toxicity of polypeptides
CN111128295A
Screening of bovine whey protein source antioxidant peptide based on molecular simulation technology
CN115631794A
Metabokinetics and toxicity prediction method based on graph representation multi-task learning
CN116343930A
Polypeptide immunocompetence prediction and generation method fusing polypeptide physicochemical properties, sequence characteristics and word vector embedding
CN119068997A