Intelligent training method, prediction method, device and medium for pharmacokinetic properties of polypeptide drugs
By constructing a prediction model for the pharmacokinetic properties of peptide drugs and combining multiple machine learning and deep learning algorithms, the problems of accuracy in predicting the ADMET properties of peptide drugs and applicability across biological systems were solved, thus achieving efficient and personalized peptide drug development.
Patent Information
- Application Number
- CN202510440181.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Existing technologies make it difficult to accurately predict the absorption, distribution, metabolism, excretion and toxicity properties of peptide drugs, resulting in low R&D efficiency. The lack of systematic modeling across species, organs and cell lines leads to inaccurate prediction results.
A variety of machine learning algorithms and deep learning architectures are used, combined with peptide sequences and molecular descriptors, to construct a prediction model for the pharmacokinetic properties of peptide drugs. These algorithms include random forest, XGBoost, support vector machine, decision tree, gradient boosting tree, and LightGBM algorithms. A peptide half-life prediction model is constructed by combining graph neural networks and the AlphaPeptDeep transfer learning architecture. A dual-path structure of graph convolutional networks and multi-layer perceptrons is used for toxicity prediction.
It achieves efficient and accurate prediction of the ADMET properties of peptide drugs, adapts to different modification types, takes biological differences into consideration, provides personalized prediction results, and improves the efficiency and reliability of peptide drug development.
Smart Images

Figure CN120280000B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of drug analysis technology, and in particular to an intelligent training method, prediction method, device and medium for the pharmacokinetic properties of polypeptide drugs. Background Art
[0002] With the rapid development of peptide drug research and development and the growing demand for personalized treatments, the construction of artificial intelligence-based peptide absorption, distribution, metabolism, excretion, and toxicity (ADMET) property prediction platforms has become a core requirement for novel drug development. Peptide drugs, due to their high specificity and low immunogenicity, have become a hot topic in the treatment of tumors, metabolic diseases, and other fields.
[0003] However, its complex physicochemical properties lead to significant heterogeneity in ADMET properties. Traditional experimental methods require several months of animal experiments and preclinical evaluations, which seriously restricts R&D efficiency. Summary of the Invention
[0004] This application proposes an intelligent training method, prediction method, device and medium for the pharmacokinetic properties of polypeptide drugs, which can solve one of the problems existing in the background technology.
[0005] To achieve the above objectives, this application adopts the following technical solutions:
[0006] In a first aspect, a method for training a peptide drug pharmacokinetic property prediction model is provided, the training method comprising:
[0007] Obtaining polypeptide drug property prediction training data, wherein the training data is divided into at least two data sets, each of which is used to predict a pharmacokinetic property of a polypeptide drug; and
[0008] The prediction model is trained using the training data, where the prediction model includes at least two prediction sub-models.
[0009] Based on the above technical solution, the trained prediction model can be used to predict the pharmacokinetic properties of at least two peptide drugs. Compared with traditional experimental methods, it is efficient, accurate and comprehensive. In addition, this comprehensive and systematic prediction platform can promote the unification of model data standards developed by different research teams, effectively ensuring the actual value of peptide drugs in early screening and optimization.
[0010] In a possible design of the first aspect, the first data set includes: a polypeptide sequence, a SMILES string, and permeability data of the polypeptide in different cell lines; the pharmacokinetic property of the first polypeptide drug is an absorption distribution property, and the absorption distribution property includes: LogD 7.4 , permeability, blood-brain barrier membrane permeability and bioavailability, the first prediction sub-model is used to calculate the molecular descriptors, polypeptide descriptors and fingerprints that characterize the physicochemical properties, topological properties and electronic properties of the polypeptide drug based on the first data set, and predict the pharmacokinetic properties of the first polypeptide drug based on the molecular descriptors, polypeptide descriptors and fingerprints. The first prediction sub-model adopts random forest, XGBoost, support vector machine, decision tree, gradient boosting tree and LightGBM algorithms. The first prediction sub-model is combined with graph neural network when predicting permeability.
[0011] Based on the above technical solution, it can adapt to peptides with different modification types and provide personalized ADMET predictions; the use of molecular descriptors, peptide descriptors and fingerprints for prediction can accurately describe key features such as conformational changes of peptides, effectively improving prediction accuracy; and based on the significant impact of modifications on peptide permeability, protein binding ability, etc., it can ensure that there is no significant deviation between the predicted results and the actual situation.
[0012] In a possible design manner of the first aspect, the first prediction sub-model is further configured to:
[0013] Combining the polypeptide sequence and the SMILES string using a molecular descriptor calculation method to obtain a corresponding molecular descriptor; and
[0014] The optimal feature subset was obtained from the molecular descriptors using variance filtering, correlation filtering and random forest recursive feature elimination.
[0015] In a possible design method of the first aspect, the second data set includes: polypeptide sequence, modification information, half-life data of natural peptide drugs in different species and different organs, half-life data of modified peptides in different species and different organs, and polypeptide retention time data; the second pharmacokinetic property of the polypeptide drug is a metabolic excretion property, and the metabolic excretion property is a half-life; the second prediction sub-model adopts the AlphaPeptDeep transfer learning architecture, and the AlphaPeptDeep transfer learning architecture learns the polypeptide retention time in the pre-training stage, and shares the learned knowledge to adjust the weight in the prediction stage.
[0016] Based on the above technical solution, it can adapt to peptides with different modification types and provide personalized ADMET predictions; it fully considers the biological differences across species, organs, and cell lines, making the predictions adaptable to various scenarios, ensuring the applicability and reliability of ADMET prediction tools in peptide drug development; it can effectively capture the stability characteristics of peptides in different environments, thereby achieving more accurate half-life predictions; based on the important influence of modification on the metabolic stability of peptides, it can ensure that the predicted results will not deviate significantly from the actual situation.
[0017] In a possible design manner of the first aspect, the third data set includes: a polypeptide sequence, a SMILES string, a toxicity positive or negative marker, and a hemolytic activity value, the pharmacokinetic property of the third polypeptide drug is toxicity, and the third prediction sub-model adopts a dual-path structure of a graph convolutional network and a multi-layer perceptron, the graph convolutional network is used to extract molecular graph features, the multi-layer perceptron is used to extract key descriptors, and the third prediction sub-model is also used to fuse the molecular graph features and the key descriptors.
[0018] In a possible design manner of the first aspect, the third prediction sub-model is specifically used to:
[0019] performing a first processing on the third data set to classify the polypeptide drug into toxic and non-toxic categories; and
[0020] For peptide drugs that are judged to be toxic, the toxicity is classified into different levels and the hemolytic toxicity level is quantified.
[0021] In a possible design method of the first aspect, the evaluation indicators of the prediction model include: classification task indicators and regression task indicators, the classification task indicators include: accuracy, area under the ROC curve and F1 value, and the regression task indicators include: determination coefficient, root mean square error and mean absolute error.
[0022] In a second aspect, a method for intelligently predicting the pharmacokinetic properties of a polypeptide drug is provided, wherein the method comprises:
[0023] Obtaining data on peptide drugs to be predicted; and
[0024] The prediction model trained as described above is used to process the data of the polypeptide drug to be predicted to obtain a prediction result.
[0025] In a third aspect, an electronic device is provided, comprising: a processor, and a memory coupled to the processor, the memory being used to store a computer program; the processor being used to execute the computer program stored in the memory, so that the electronic device performs the training method as any possible implementation in the first aspect, or performs the prediction method described in the second aspect.
[0026] In a fourth aspect, a computer-readable storage medium is provided, comprising a computer program or instructions, which, when executed on a computer, enables the computer to execute the training method according to any possible implementation of the first aspect, or to execute the prediction method according to the second aspect.
[0027] In a fifth aspect, a computer program product is provided, comprising: a computer program or instructions, which, when the computer program or instructions are run on a computer, enables the computer to execute the training method as any possible implementation method in the first aspect, or execute the prediction method in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0029] Figure 1 This is a schematic diagram of a method for predicting the in vivo pharmacokinetic properties of a polypeptide drug provided in an embodiment of the present application;
[0030] Figure 2 This is an application process for step-by-step prediction of the toxicity of a novel polypeptide drug provided in the examples of this application;
[0031] Figure 3 This is a schematic diagram of a novel peptide toxicity step-by-step prediction framework MLR-GAT provided in the examples of the present application. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0033] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0035] Before introducing the embodiments of the present application, a brief description of the current technical research of the present application is given:
[0036] The current mainstream ADMET prediction technology focuses on two types of drugs: small molecules and proteins. Its method system is difficult to adapt to the particularity of peptides. First, small molecule ADMET prediction is mostly based on deep learning methods of molecular descriptors or SMILES strings, which usually assume that the molecules are rigid and have difficulty capturing their conformational changes in different environments. The secondary structure and flexible segments of peptides may change significantly in the solvent, pH value or receptor environment, affecting transmembrane permeability and metabolic stability. Protein ADMET prediction mainly relies on sequence information or coarse-grained molecular docking, which is suitable for protein-protein interactions and metabolic enzyme identification. However, the metabolic stability of peptides is affected by enzymatic degradation sites and local conformations, and existing protein methods are difficult to apply directly.
[0037] Therefore, the current peptide ADMET prediction methods have the following deficiencies:
[0038] 1. Lack of Targetedness: Current ADMET prediction methods primarily target small molecules or proteins and fail to fully consider the unique properties of peptides. For example, small molecule methods often rely on molecular descriptors or SMILES sequences, while protein methods focus on sequence information or whole-protein docking simulations. Both methods struggle to accurately describe key features of peptides, such as conformational changes and enzymatic degradation sites, resulting in limited prediction accuracy.
[0039] 2. Lack of modification information: Peptide drugs often contain multiple chemical modifications (such as methylation, phosphorylation, glycosylation, etc.), which have a significant impact on their metabolic stability, transmembrane permeability, and protein binding ability. However, existing ADMET prediction methods are usually based on natural amino acid composition, which makes it difficult to effectively model and evaluate the pharmacokinetic behavior of modified peptides, resulting in a large deviation between the predicted results and the actual situation.
[0040] 3. Lack of consideration of biological differences: Existing ADMET prediction methods are usually trained and validated based on a single species (such as humans), a specific organ (such as the liver), or a specific cell line (such as Caco-2), making it difficult to accurately predict the pharmacokinetic properties of peptides in different biological systems. However, the absorption, distribution, metabolism, and excretion processes of peptide drugs often involve interactions between multiple organs, as well as differences between different species. For example, the metabolic pathways of certain peptides in rodents may be significantly different from those in humans, making it difficult to directly extrapolate animal experimental data to clinical applications. Therefore, the lack of systematic modeling across species, organs, and cell lines limits the applicability and reliability of existing ADMET prediction tools in peptide drug development.
[0041] 4. Lack of a comprehensive prediction platform: Current peptide ADMET prediction research is still in its early stages and lacks a systematic, integrated prediction platform. Existing methods often focus on a specific property and fail to combine multiple prediction models to provide comprehensive ADMET assessments. Furthermore, the inconsistent data standards of models developed by different research teams make the comparison and application of results difficult, limiting the practical value of peptide drugs in early screening and optimization.
[0042] To address the aforementioned technical deficiencies, the present invention provides an intelligent prediction method for the in vivo pharmacokinetics of peptide drugs, which effectively learns peptide sequence features and improves the accuracy and robustness of peptide drug ADMET property prediction. As will be apparent from the following description, the present invention also relates to a method for training an intelligent prediction model for the in vivo pharmacokinetics of peptide drugs.
[0043] The present invention provides an intelligent prediction method for the pharmacokinetics of polypeptide drugs in vivo. Figure 1 Specifically, the following are included:
[0044] S1. Data collection and preprocessing: Collect the sequence and modification information of the predicted peptide and the experimental data related to the corresponding ADMET, including physicochemical properties, bioavailability, blood-brain barrier permeability, permeability, half-life and toxicity, and perform data cleaning and standardization to ensure high quality and consistency of the data;
[0045] S2. Molecular descriptor calculation: Calculate the molecular descriptors of the peptide, including physicochemical properties, topological properties, electronic properties, etc., to provide input features for subsequent machine learning models;
[0046] S3. Machine learning model construction: Use multiple machine learning algorithms for training to predict the physical and chemical properties of peptides, and improve the prediction performance of the model through cross-validation and hyperparameter optimization;
[0047] S4. Peptide half-life prediction: Based on the AlphaPeptDeep deep learning framework, we use transfer learning methods to predict peptide half-life to improve the robustness of model prediction;
[0048] S5. Peptide toxicity prediction: The peptide toxicity step-by-step prediction framework MLR-GAT is used to achieve accurate prediction of peptide toxicity;
[0049] S6. Model performance evaluation: Evaluate the performance of the prediction model for each property separately and retain the best model;
[0050] S7. Construction of a peptide ADEMT property prediction platform: Integrate the above-mentioned machine learning and deep learning models to build an integrated peptide ADMET prediction platform, providing a visual interface where users can input peptide sequences and modification information to obtain ADMET prediction results.
[0051] Specifically, the S1 data collection and preprocessing methodology involved manually collecting and screening data from multiple open-access databases and peer-reviewed literature. All data were categorized by ADMET properties and further subdivided into multiple subcategories based on the specific significance of each endpoint. To ensure data consistency and reliability, a series of rigorous preprocessing steps were performed, including missing and duplicate value handling.
[0052] Specifically, the S2 molecular descriptor calculation method includes two main computational methods for peptide characterization: one for molecular descriptors and fingerprints calculated based on SMILES strings, and the other for peptide feature descriptors calculated based on sequence information. The descriptors are then subjected to variance filtering, correlation filtering, and random forest recursive feature elimination to select the optimal feature subset.
[0053] Specifically, the S3 machine learning model construction method includes: making predictions based on traditional machine learning methods, combining six advanced algorithms (random forest (RF), extreme gradient boosting (XGBoost), support vector machine (SVM), decision tree (DT), gradient boosting tree (GBT) and light gradient boosting machine (LightGBM)) with different descriptors to build and optimize the model.
[0054] Specifically, the S4 peptide half-life prediction method includes: applying the AlphaPeptDeep deep learning architecture and introducing a transfer learning strategy. The core idea of this method is to use the potential physicochemical relationship between peptide retention time and half-life to establish a shared knowledge learning model to improve prediction accuracy. During the implementation process, AlphaPeptDeep extracts the sequence characteristics and physicochemical properties of peptides through a deep neural network, and learns the retention time pattern of peptides in the pre-training stage. Subsequently, in the peptide half-life prediction task, the model shares the learned knowledge and further adjusts the weights to more accurately model the factors affecting half-life. Ultimately, this method can effectively capture the stability characteristics of peptides under different environments, thereby achieving more accurate half-life predictions.
[0055] Specifically, if Figure 2 As shown, the S5 peptide toxicity prediction method includes: proposing a new peptide toxicity step-by-step prediction framework MLR-GAT, which can handle multi-task learning and multi-classification problems at the same time. First, the model will preliminarily classify the input peptides and divide them into two categories: toxic and non-toxic. Subsequently, for peptides judged to be toxic, the toxicity categories are further refined, including hemostatic injury toxicity, cytotoxicity, G protein coupled receptor toxicity, neurotoxicity, cytolytic toxicity, and hemolytic toxicity, in order to more accurately identify their potential hazards. On this basis, for neurotoxins, the model further subdivides their mechanisms of action, including acetylcholine receptor inhibition toxicity, Na + , K + , Ca 2+ Ion channel toxicity, in addition, the system also introduces a regression model to HC 50 The values were predicted to quantify the toxicity level.
[0056] Specifically, the S6 model evaluation method includes: for classification tasks, the evaluation indicators include accuracy (ACC):
[0057]
[0058] Area under the ROC curve (AUC):
[0059]
[0060] F1 value (F1):
[0061]
[0062] Among them, i is the category index, TP is a true positive example, that is, a positive example is predicted and is actually a positive example; FP is a false positive example, a positive example is predicted but is actually a negative example; FN is a false negative example, that is, a negative example is predicted but is actually a positive example; TN is a true negative example, a negative example is predicted but is actually a negative example.
[0063] For regression tasks, the evaluation metrics include the coefficient of determination (R 2 ):
[0064]
[0065] Root Mean Square Error (RMSE):
[0066]
[0067] Mean Absolute Error (MAE):
[0068]
[0069] Where N is the number of samples, y i is the true value, is the predicted value, is the average of the true values.
[0070] Specifically, the S7 peptide ADEMT property prediction platform is constructed using an intuitive user interface, with easy navigation through five functional modules: Home, Services, Documentation, Help, and About. The Home module provides users with an overview of the platform's features and functions, helping them quickly understand its overall framework and objectives. The Services module focuses on the prediction and calculation of key ADMET properties, integrating multiple efficient prediction tools and supporting SMILES string input or molecular structure drawing via a structure editor. The platform also requires users to provide peptide sequence information and offers 35 modification types for half-life prediction. After users submit relevant data, the platform automatically calculates molecular descriptors and performs predictions via a web server. The results are displayed in real time on an interactive page, providing detailed prediction results and corresponding molecular information. The Documentation module details the data collection, model construction, and algorithm implementation processes. The Help module provides users with a concise tutorial to help them quickly master the platform's usage. The About.com module showcases information about the R&D team, introducing the platform's development background and team members.
[0071] In summary, compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0072] 1. This platform is the first comprehensive platform capable of comprehensively predicting peptide ADMET properties. By combining machine learning and deep learning, an efficient and accurate peptide ADMET prediction method and platform has been established.
[0073] 2. Ability to characterize modified peptides, thereby improving prediction accuracy. Specifically designed for peptide ADMET prediction, it can adapt to peptides with different modification types and provide personalized predictions.
[0074] 3. Biological differences are fully considered in the prediction of some properties, making the prediction adaptable to various scenarios.
[0075] The above-mentioned embodiments of the present application are described in detail below through a combination of multiple application examples.
[0076] Application Example 1: Peptide Property Prediction Experiment Based on Machine Learning
[0077] This application example uses a variety of advanced machine learning methods to build a high-performance prediction model to verify the effectiveness and applicability of the methods.
[0078] (1) Data source and preprocessing
[0079] This study used peptide and related property data from public databases and literature. All data were classified according to ADMET properties and further subdivided into multiple subcategories based on the specific significance of each endpoint, including absorption distribution properties such as LogD 7.4 , permeability, blood-brain barrier membrane permeability, bioavailability; metabolic excretion properties include half-life; toxicity includes overall toxicity, which can be divided into G protein-coupled receptor toxicity, cytotoxicity, hemostatic damage toxicity, cytolytic toxicity, neurotoxicity, hemolytic toxicity; among them, neurotoxicity is divided into Na + , K + , Ca 2+ Ion channel toxicity and acetylcholine receptor inhibition toxicity, hemolytic toxicity also includes hemolytic activity index HC 50 . Finally, 17 independent data sets were constructed for ADMET predictive modeling. To ensure the consistency and reliability of the data, a series of rigorous preprocessing steps were performed: 1) The labels and experimental values of all data were extracted, the units were unified, and records with missing values were eliminated. For regression data, if a molecule has multiple data entries, the arithmetic mean is taken. 2) Only peptide sequences containing 2 to 50 amino acids are retained. 3) All peptide sequences and SMILES are checked and verified one by one. In the case of sequence missing, the corresponding peptide sequence is manually extracted according to SMILES; for peptides with missing SMILES, the natural peptide is converted using the RDkit package, and the modified peptide is manually modified according to the modification characteristics. 4) Duplicates are eliminated by calculating the InChIKey value to ensure the uniqueness and accuracy of the data. After the above processing, a total of 36,643 high-quality data were obtained, covering the prediction information of 27 key endpoints.
[0080] (2) Descriptor calculation
[0081] Peptide characterization is primarily achieved through two computational methods: molecular descriptors and fingerprints calculated based on SMILES strings, and peptide feature descriptors calculated based on sequence information. Specifically, MOE software and PyBioMed were used to calculate SMILES-based small molecule descriptors, while RDKit was used to calculate MACCS fingerprints, PubChem fingerprints, and Pharmacophore ErG fingerprints. PyBioMed and modlAMP were used to calculate sequence-based peptide descriptors. During the feature selection phase, descriptors with variances less than 0.01 were first eliminated. Features with inter-descriptor correlations greater than 0.99 were de-redundant, and one highly correlated descriptor was randomly removed. Next, a random forest recursive feature elimination method was applied in conjunction with a five-fold cross-validation strategy to select the optimal subset for subsequent modeling.
[0082] (3) Model construction
[0083] Six machine learning algorithms (Random Forest RF, XGBoost, Support Vector Machine SVM, Decision Tree DT, Gradient Boosting Tree GBT and LightGBM) were used to train with different descriptors, and the hyperparameters were adjusted to optimize the model performance. Specifically, LogD 7.4 , bioavailability, and blood-brain barrier membrane permeability mainly use small molecule and peptide descriptors calculated using PyBioMed and modlAMP as input; in particular, for permeability, we built models for the commonly used in vitro models Caco-2, PAMPA, and RRCK, and calculated different molecular descriptors and molecular fingerprints, including MOE2D descriptors, peptide descriptors, small molecule descriptors, and MACCS fingerprints. In addition, the use of GNN further improved the permeability prediction performance of the model on large data sets. By taking SMILES and sequences as input, GNN integrates molecular graph and descriptor information from two pathways, and extracts key features through the attention mechanism to generate the final prediction results. All data sets were randomly divided into training and test sets at a ratio of 0.75:0.25, and 5-fold cross-validation was performed on the training set to ensure that the model has good generalization ability.
[0084] (4) Experimental results
[0085] Regression models were successfully constructed for LogD7.4 and permeability, and binary classification models were successfully constructed for bioavailability and blood-brain barrier permeability. The evaluation results of the models on the test set are detailed in Tables 1 and 2.
[0086] Table 1: Results of machine learning modeling regression tasks
[0087] Model <![CDATA[Q 2 ]]> <![CDATA[RMSE cv ]]> MAE cv ]]> [R 2 ]]> <![CDATA[RMSE test ]]> <![CDATA[MAE test ]]> <![CDATA[LogD 7.4 ]]> 0.82 0.466 0.324 0.818 0.491 0.318 Caco2-L 0.411 0.747 0.551 0.435 0.765 0.582 Caco2-C 0.582 0.425 0.307 0.527 0.472 0.335 Caco2-A 0.50 0.548 0.388 0.476 0.573 0.408 PAMPA-C 0.646 0.428 0.317 0.657 0.423 0.308 RRCK-C 0.691 0.34 0.269 0.623 0.42 0.32
[0088] Among them, LogD 7.4 The best model is the GBT-based model, which has an R 2 It reached 0.820, which is higher than 0.818 in cross-validation, reflecting its superior predictive performance. Caco2-L represents the permeability model of linear peptides in the Caco-2 cell line (the best model is the SVR model based on the MOE2D descriptor), Caco2-C represents the permeability model of cyclic peptides in the Caco-2 cell line (the best model is the LightGBM model based on the MOE2D descriptor), Caco2-A represents the permeability model of linear and cyclic peptides in the Caco-2 cell line (the best model is the RF model based on the MOE2D descriptor), PAMPA-C represents the permeability model of cyclic peptides in the PAMPA cell line (the best model is the RF model based on the MOE2D descriptor), and RRCK-C represents the permeability model of cyclic peptides in the RRCK cell line (the best model is the RF model based on the MOE2D descriptor). That is, we took into account the phenomenon that different types of peptides have different permeabilities in different cell lines. As can be seen from the table, the performance of the model in the test set is not inferior to the cross-validation results in the training set, which reflects the excellent robustness of the model.
[0089] Table 2: Results of machine learning modeling classification tasks
[0090]
[0091] Among them, the best model for bioavailability is the model based on LightGBM, with an ACC of 0.811, an AUC of 0.901, and an f1 of 0.761 in cross-validation, and an ACC of 0.844, an AUC of 0.90, and an f1 of 0.80 in the test set; the best model for blood-brain barrier membrane permeability is the model based on RF, with an ACC of 0.808, an AUC of 0.894, and an f1 of 0.805 in cross-validation, and an ACC of 0.78, an AUC of 0.889, and an f1 of 0.789 in the test set, indicating that the model performs equally well in the test set, reflecting the model's good robustness and generalization ability.
[0092] Application Example 2: Peptide Half-Life Prediction Based on Transfer Learning
[0093] This application example uses the AlphaPeptDeep transfer learning architecture to improve the prediction ability of peptide half-life.
[0094] (1) Data collection and benchmark dataset construction
[0095] Data were collected from public peptide half-life databases and literature, and statistical analysis was performed on the collected data according to different species, organs, and types. Ultimately, peptide half-life data under five different conditions were obtained, including: half-life of native peptides in human blood (HBN), half-life of modified peptides in human blood (HBM), half-life of native peptides in mouse blood (MBN), half-life of modified peptides in mouse blood (MBM), and half-life of modified peptides in mouse intestine (MIM). These data contain 117, 187, 106, 182, and 378 peptides and their corresponding half-lives, respectively, totaling 970 high-quality data points.
[0096] (2) Model pre-training and fine-tuning
[0097] First, a deep learning model was trained on a large peptide retention time dataset to learn the physicochemical properties of peptides. It was then fine-tuned on a half-life dataset to enhance its predictive power. The dataset was split into training and test sets with a 75%:25% ratio.
[0098] (3) Experimental results
[0099] The coefficient of determination (R 2 ), root mean square error (RMSE), and mean absolute error (MAE) indicators were used for evaluation. The results showed that for the five data sets HBN, HBM, MBN, MBM, and MIM, the model R after transfer learning was significantly better than the best model results of traditional machine learning. 2 The improvements were 9%, 32%, 16.4%, 44%, and 15%, respectively, demonstrating the high applicability of transfer learning in small sample scenarios. The model results after transfer learning are detailed in Table 3.
[0100] Table 3: Half-life prediction results using the AlphaPeptDeep deep learning architecture
[0101] Model <![CDATA[Q 2 ]]> <![CDATA[RMSE cv ]]> <![CDATA[MAE cv ]]> <![CDATA[R 2 ]]> <![CDATA[RMSE test ]]> <![CDATA[MAE test ]]> HBN 0.86 202.34 68.222 0.84 272.2 114.19 HBM 0.92 133.71 56.95 0.9 147.65 81.03 MBN 0.997 12.72 5.56 0.984 23.15 11.49 MBM 0.87 0.60 0.30 0.93 0.49 0.34 MIM 0.93 0.62 0.62 0.94 0.54 0.34
[0102] HBN refers to the half-life of the natural peptide in human blood, HBM refers to the half-life of the modified peptide in human blood, MBN refers to the half-life of the natural peptide in mouse blood, MBM refers to the half-life of the modified peptide in mouse blood, and MIM refers to the half-life of the modified peptide in mouse intestine. In other words, we have taken into account the modification information of the peptide and the differences in its half-life in different species and organs. As can be seen from the table, the R 2 The values are all above 0.8, and the test set results are similar to the training set results, which reflects the powerful performance of transfer learning in few-sample tasks. Among them, the MBN model has the best overall performance.
[0103] Application Example 3: Application of the MLR-GAT Framework
[0104] This application example verifies the effectiveness of the proposed MLR-GAT framework in the peptide toxicity prediction task.
[0105] (1) Data collection and cleaning
[0106] Peptide data known to be toxic were collected from public databases and supplemented with hemolytic peptide sequences from the hemolytic database. Negative data exclude any records with toxic functions from publicly available data sources to ensure that these sequences do not contain any functional features that may overlap with positive peptides to improve the accuracy of the model. The dataset consists of peptides containing 50 or fewer amino acids. By removing duplicates and sequences containing non-standard amino acid residues (such as X, B, Z, etc.), a total of 1,743 positive data and an equal amount of negative data were obtained.
[0107] (2) Model structure
[0108] This example, MLR-GAT, combines a hybrid representation of molecular graphs and molecular descriptors, employing a dual-path architecture consisting of a graph convolutional network (GCN) and a multi-layer perceptron (MLR). The model is designed to process different types of input data through two separate paths, ultimately fusing the features of both for accurate predictions. The input to path one is the SMILES representation of a molecule, a common string representation of molecular structure. The SMILES string is first converted into a molecular graph representation, where nodes represent atoms and edges represent chemical bonds between atoms. Next, using a graph convolutional network (GCN), the model captures the relationships between atoms and their neighbors, extracting essential features of atoms and bonds. GCNs effectively propagate information within graph-structured data, gradually learning the relationships between nodes (atoms) and their neighbors, and ultimately generating high-level features of the molecular graph. These features, through further feature conversion layers, form the final representation of the molecular graph, which reflects the potential behavior and properties of molecules in chemical reactions. Path 2 takes the peptide sequence and SMILES string as input to calculate the small molecule descriptor and peptide descriptor. The small molecule descriptor is a quantitative representation of the molecular structure, usually including multiple physicochemical properties such as molecular weight, polar surface area, and molecular folding, which can help the model understand the basic properties of the molecule. The peptide descriptor is based on the analysis of the peptide sequence, capturing the composition, arrangement order and other biological characteristics of amino acids. Path 2 extracts key descriptors through a multi-layer perceptron (MLP). The output results of the two paths are fused to generate a comprehensive feature representation. This representation combines the structural information of the molecular graph and the quantitative features of the descriptor to provide a more comprehensive molecular representation for the model to perform subsequent prediction tasks. The fused features are input into the fully connected layer for final prediction. The fully connected layer maps the extracted features to the output space of the target task through a deep linear transformation. In order to further optimize the performance of the model in multi-task learning, a shared weighted network structure is adopted. This structure allows the model to share some network parameters while assigning different attention weights to each task. In this way, the model can dynamically adjust its learning focus according to the characteristics of each task, thereby improving the prediction accuracy of each task ( Figure 3 ).
[0109] (3) Optimization strategy
[0110] ① Data partitioning: The dataset is divided into training set, validation set, and test set according to the ratio of 8:1:1, and the model is trained through 5 random splits.
[0111] ②Loss function: For binary classification tasks, use the BCEWithLogitsLoss loss function:
[0112] Loss=-(y n log(σ(x n))+(1-y n )log(1-σ(x n )))
[0113] Among them, x n is the sample input, y n For the corresponding output, σ(x) is the Sigmoid function:
[0114]
[0115] For multi-classification tasks, the CrossEntropyLoss loss function is used:
[0116]
[0117] Where N is the number of samples, Indicates that the i-th sample is in the true category y i The predicted probability of .
[0118] For regression tasks, the MSELoss loss function is used:
[0119]
[0120] Where N is the number of samples, and x n is the sample input, y n is the corresponding output.
[0121] ③ Data Imbalance Handling: Binary classification tasks adjust the weight of positive samples (pos_weight), while multi-classification tasks adjust the weight of each class (weight). To address data imbalance, the binary classification task adjusts the weight of positive samples (pos_weight), while multi-classification tasks adjust the weight of each class (weight). The dataset is divided into training, validation, and test sets in an 8:1:1 ratio, and the model is trained using five random splits.
[0122] (4) Experimental results
[0123] ACC, AUC, and F1-score are used for classification tasks, and R is used for regression tasks. 2 , MAE, and RMSE were used for evaluation. The specific results are shown in Tables 4 and 5.
[0124] Table 4: Regression task prediction results using the MLR-GAT peptide step-by-step prediction architecture
[0125] Model <![CDATA[Q 2 ]]> <![CDATA[RMSE cv ]]> <![CDATA[MAE cv ]]> <![CDATA[R 2 ]]> RMSE test ]]> <![CDATA[MAE test ]]> <![CDATA[HC 50 ]]> 0.703 0.376 0.278 0.474 0.479 0.360
[0126] We can see that the Q2 of the HC50 prediction on the validation set is 0.703, and the R2 on the test set is 0.474, indicating that the model has a strong ability to fit the validation set. The result of the test set is slightly lower than its performance on the validation set, indicating that the model faces certain challenges when facing unknown data.
[0127] Table 5: Prediction results of classification tasks using the MLR-GAT peptide hierarchical prediction architecture
[0128] Model <![CDATA[ACC cv ]]> <![CDATA[AUC cv ]]> <![CDATA[f1 cv ]]> <![CDATA[ACC test ]]> <![CDATA[AUC test ]]> <![CDATA[f1 test ]]> <![CDATA[Toxin binary ]]> 0.894 0.965 0.904 0.794 0.885 0.813 <![CDATA[Toxin type ]]> 0.951 0.995 0.84 0.885 0.949 0.634 <![CDATA[Toxin neuro ]]> 0.923 0.994 0.923 0.736 0.905 0.683
[0129] The results in Table 5 show that the multi-task multi-classification model MLR-GAT constructed in Example 3 has good performance in the toxicity classification prediction task. binary The toxicity binary classification model has an ACC of 0.794, an AUC of 0.885, and an f1 score of 0.813 on the test set; type The six-classification model of toxicity has an ACC of 0.885, an AUC of 0.949, and an f1 score of 0.634 on the test set; neuro The neurotoxicity four-classification model achieved an ACC of 0.736, an AUC of 0.905, and an f1 score of 0.683 on the test set. These results indicate that the multi-task model can accurately capture peptide sequence information and achieve accurate toxicity prediction.
[0130] Application Example 4: Construction of a Peptide Property Prediction System Based on an Intelligent Computing Platform
[0131] This application example provides a scalable intelligent computing platform that integrates multiple algorithms for peptide property prediction and supports model training, reasoning, and visual analysis.
[0132] The underlying architecture of this application example is deployed on a cloud server, supporting efficient programming development and model training. The platform adopts a layered architecture design, which clearly separates business logic from the user interaction interface to simplify system upgrades and maintenance. A dynamic interactive interface is built on the front end to enhance user experience and operational fluency. The back end combines efficient request processing and load balancing mechanisms to ensure stable operation and high concurrency support of the system. In terms of data storage, the platform uses a reliable database system to achieve efficient data management and retrieval. The platform integrates a variety of advanced machine learning frameworks to support model training and reasoning, and combines chemical informatics methods for molecular descriptor calculations. As an open service system, this platform allows users to directly access its core functions, providing convenient and efficient peptide ADMET property prediction capabilities.
[0133] An embodiment of the present application also provides an electronic device, comprising: a processor, and a memory coupled to the processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method described in any one of the above embodiments.
[0134] The electronic device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The electronic device may include, but is not limited to, a processor and a memory.
[0135] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire device using various interfaces and lines.
[0136] The memory may be used to store the computer program, and the processor implements various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.
[0137] The memory may primarily include a program storage area and a data storage area, wherein the program storage area may store an operating system, at least one application required for a function, and the like; and the data storage area may store data created based on the use of the mobile phone, and the like. Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0138] The embodiment of the present application also provides a storage medium, which is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0139] An embodiment of the present application further provides a computer program product, including: a computer program or instructions, which, when executed on a computer, causes the computer to execute any of the above-mentioned possible implementation methods.
[0140] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.
Claims
1. A method for training a prediction model for the pharmacokinetic properties of a polypeptide drug, characterized in that: The training method comprises: Obtaining polypeptide drug property prediction training data, wherein the training data is divided into at least two data sets, each of which is used to predict a pharmacokinetic property of a polypeptide drug; and Using the training data, training the prediction model, wherein the prediction model includes at least two prediction sub-models; The first data set includes: peptide sequence, SMILES string and peptide permeability data in different cell lines. The pharmacokinetic properties of the first peptide drug are absorption and distribution properties, which include: LogD 7.4 , permeability, blood-brain barrier membrane permeability and bioavailability, the first prediction sub-model is used to calculate the molecular descriptors, polypeptide descriptors and fingerprints that characterize the physicochemical properties, topological properties and electronic properties of the polypeptide drug based on the first data set, and predict the pharmacokinetic properties of the first polypeptide drug based on the molecular descriptors, polypeptide descriptors and fingerprints. The first prediction sub-model adopts random forest, XGBoost, support vector machine, decision tree, gradient boosting tree and LightGBM algorithms. The first prediction sub-model is combined with graph neural network when predicting permeability.
2. The training method according to claim 1, wherein: The first prediction sub-model is further specifically used for: Combining the polypeptide sequence and the SMILES string using a molecular descriptor calculation method to obtain a corresponding molecular descriptor; and The optimal feature subset was obtained from the molecular descriptors using variance filtering, correlation filtering and random forest recursive feature elimination.
3. The training method according to claim 1, wherein: The second data set includes: peptide sequence, modification information, half-life data of natural peptides in different species and different organs, half-life data of modified peptides in different species and different organs, and peptide retention time data; the second pharmacokinetic property of the peptide drug is the metabolic excretion property, and the metabolic excretion property is half-life; the second prediction sub-model adopts the AlphaPeptDeep transfer learning architecture, which learns the peptide retention time in the pre-training stage and shares the learned knowledge to adjust the weight in the prediction stage.
4. The training method according to claim 1, wherein: The third data set includes: polypeptide sequence, SMILES string, toxicity positive or negative marker, and hemolytic activity value. The pharmacokinetic property of the third polypeptide drug is toxicity. The third prediction sub-model adopts a dual-path structure of a graph convolutional network and a multi-layer perceptron. The graph convolutional network is used to extract molecular graph features, and the multi-layer perceptron is used to extract key descriptors. The third prediction sub-model is also used to fuse the molecular graph features and the key descriptors.
5. The training method according to claim 4, wherein: The third prediction sub-model is specifically used for: performing a first processing on the third data set to classify the polypeptide drug into toxic and non-toxic categories; and For peptide drugs that are judged to be toxic, the toxicity is classified into different levels and the hemolytic toxicity level is quantified.
6. The training method according to claim 1, wherein: The evaluation indicators of the prediction model include: classification task indicators and regression task indicators. The classification task indicators include: accuracy, area under the ROC curve and F1 value. The regression task indicators include: determination coefficient, root mean square error and mean absolute error.
7. An intelligent prediction method for the pharmacokinetic properties of polypeptide drugs, characterized in that: The prediction method comprises: Obtaining data on peptide drugs to be predicted; and The prediction model trained by the training method according to any one of claims 1 to 6 is used to process the polypeptide drug data to be predicted to obtain a prediction result.
8. An electronic device, characterized in that: The electronic device includes: a processor, and a memory coupled to the processor, The memory is used to store computer programs; and The processor is configured to execute the computer program stored in the memory, so that the electronic device executes the training method according to any one of claims 1 to 6, or executes the prediction method according to claim 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a computer program or instructions, which, when executed on a computer, causes the computer to execute the training method according to any one of claims 1 to 6, or the prediction method according to claim 7.
Citation Information
Patent Citations
ADMET prediction electronic device based on multi-model integration and method thereof
CN115620828A
System, method and program for pharmacokinetic parameter prediction of peptide sequence by mathematical model
US20100121791A1