Intelligent training method, prediction method, equipment and medium for pharmacokinetic properties of polypeptide drug

By constructing a prediction model for pharmacokinetic properties of polypeptide drugs, using machine learning and deep learning algorithms, the problems of low efficiency and insufficient accuracy of prediction of ADMET properties of polypeptide drugs are solved, and personalized and comprehensive prediction effects are achieved, improving the efficiency and accuracy of polypeptide drug development.

CN120280000AActive Publication Date: 2025-07-08CENT SOUTH UNIV

Patent Information

Application Number
CN202510440181.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately predict the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of polypeptide drugs. Traditional experimental methods take a long time and lack a systematic prediction platform, making it difficult to adapt to the unique properties and biological differences of the polypeptide.

Method used

A variety of machine learning and deep learning algorithms are used to construct a prediction model for pharmacokinetic properties of polypeptide drugs, and feature extraction is performed through molecular descriptors and polypeptide descriptors, combining transfer learning and graph convolution networks to build a personalized ADMET prediction platform to support the comprehensive prediction of polypeptide drugs.

Benefits of technology

It realizes efficient and accurate prediction of the properties of ADMET of polypeptide drugs, adapts to different modification types and biological differences, provides personalized prediction results, and improves the efficiency and reliability of polypeptide drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280000A_ABST
    Figure CN120280000A_ABST
Patent Text Reader

Abstract

According to the intelligent training method, the prediction method, the equipment and the medium for the pharmacokinetic properties of the polypeptide drugs, based on the prediction model obtained through training, the pharmacokinetic properties of at least two polypeptide drugs can be predicted, and compared with a traditional experiment method, the intelligent training method has the advantages of being efficient, accurate and comprehensive in effect and high in practicability. And the comprehensive systematic prediction platform can promote the unification of model data standards developed by different research teams, and effectively ensures the actual value of the polypeptide drug in early screening and optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of pharmaceutical analysis, and particularly to an intelligent training method, a prediction method, a device and a medium for the pharmacokinetic properties of polypeptide drugs. Background Art

[0002] With the booming development of polypeptide drug research and development and the increasing demand for personalized treatment, building an artificial intelligence-based platform for predicting the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of polypeptides has become a core requirement for the development of new drugs. Polypeptide drugs have become a hot topic in the fields of treating tumors, metabolic diseases, etc. due to their advantages such as high specificity and low immunogenicity.

[0003] However, their complex physicochemical properties lead to significant heterogeneity in ADMET properties. Traditional experimental methods require months of animal experiments and preclinical evaluations, seriously restricting the R & D efficiency. Summary of the Invention

[0004] The present application provides an intelligent training method, a prediction method, a device and a medium for the pharmacokinetic properties of polypeptide drugs, which can solve one of the problems existing in the background art.

[0005] To achieve the above object, the present application adopts the following technical solutions:

[0006] In a first aspect, a training method for a pharmacokinetic property prediction model of polypeptide drugs is provided. The training method includes:

[0007] Obtaining polypeptide drug property prediction training data, where the training data is divided into at least two groups of data sets, and each data set is used to predict a pharmacokinetic property of a polypeptide drug; and

[0008] Using the training data to train the prediction model, where the prediction model includes at least two prediction sub-models.

[0009] Based on the above technical solution, based on the trained prediction model, the pharmacokinetic properties of at least two polypeptide drugs can be predicted. Compared with traditional experimental methods, it has the effects of high efficiency, accuracy and comprehensiveness. And this comprehensive and systematic prediction platform can promote the unification of the model data standards developed by different research teams, effectively ensuring the practical value of polypeptide drugs in early screening and optimization.

[0010] In a possible design of the first aspect, the first data set includes: polypeptide sequences, SMILES strings, and permeability data of polypeptides in different cell lines. The pharmacokinetic properties of the first polypeptide drug are absorption and distribution properties, and the absorption and distribution properties include: LogD 7.4 , permeability, blood-brain barrier membrane permeability, and bioavailability. The first predictor model is used to calculate molecular descriptors, polypeptide descriptors, and fingerprints characterizing the physicochemical properties, topological properties, and electronic properties of the polypeptide drug based on the first data set, and predict the pharmacokinetic properties of the first polypeptide drug based on the molecular descriptors, polypeptide descriptors, and fingerprints. The first predictor model uses algorithms such as random forest, XGBoost, support vector machine, decision tree, gradient boosting tree, and LightGBM. The first predictor model combines a graph neural network when predicting permeability.

[0011] Based on the above technical solutions, it is possible to adapt to polypeptides of different modification types and provide personalized ADMET predictions. Using molecular descriptors, polypeptide descriptors, and fingerprints for prediction can accurately describe key features such as conformational changes of polypeptides, effectively improving the prediction accuracy. And since modifications have an important impact on polypeptide permeability, protein binding ability, etc., it can ensure that there is no large deviation between the prediction results and the actual situation.

[0012] In a possible design of the first aspect, the first predictor model is specifically further used for:

[0013] Using a molecular descriptor calculation method, combining the polypeptide sequence and the SMILES string to obtain corresponding molecular descriptors; and

[0014] Using variance filtering, correlation filtering, and random forest recursive feature elimination method to obtain a preferred feature subset from the molecular descriptors.

[0015] In a possible design of the first aspect, the second data set includes: polypeptide sequences, modification information, half-life data of natural peptide drugs in different species and different organs, half-life data of modified peptides in different species and different organs, and polypeptide retention time data; the pharmacokinetic properties of the second polypeptide drug are metabolism and excretion properties, and the metabolism and excretion property is the half-life; the second predictor model uses the AlphaPeptDeep transfer learning architecture, and the AlphaPeptDeep transfer learning architecture learns the polypeptide retention time in the pre-training stage and shares the learned knowledge to adjust the weights in the prediction stage.

[0016] Based on the above technical solution, it can adapt to polypeptides of different modification types and provide personalized ADMET prediction; fully consider the biological differences across species, organs, and cell lines, enabling the prediction to adapt to multiple scenarios and ensuring the applicability and reliability of the ADMET prediction tool in polypeptide drug development; effectively capture the stability characteristics of polypeptides in different environments, thereby achieving more accurate half-life prediction; based on the fact that modification has an important impact on polypeptide metabolic stability, etc., therefore, it can ensure that there is no large deviation between the prediction result and the actual situation.

[0017] In a possible design of the first aspect, the third data set includes: polypeptide sequences, SMILES strings, toxicity positive or negative labels, and hemolytic activity values. The pharmacokinetic property of the third polypeptide drug is toxicity. The third prediction sub-model adopts a dual-path structure of a graph convolutional network and a multi-layer perceptron. The graph convolutional network is used to extract molecular graph features, and the multi-layer perceptron is used to extract key descriptors. The third prediction sub-model is also used to fuse the molecular graph features and the key descriptors.

[0018] In a possible design of the first aspect, the third prediction sub-model is specifically used for:

[0019] Perform a first processing on the third data set to classify the toxicity and non-toxicity of the polypeptide drug; and

[0020] For the polypeptide drug determined to be toxic, perform a hierarchical classification of toxicity and quantify the hemolytic toxicity level.

[0021] In a possible design of the first aspect, the evaluation metrics of the prediction model include: classification task metrics and regression task metrics. The classification task metrics include: accuracy, area under the ROC curve, and F1 value. The regression task metrics include: coefficient of determination, root mean square error, and mean absolute error.

[0022] In the second aspect, an intelligent prediction method for the pharmacokinetic property of a polypeptide drug is provided, characterized in that the prediction method includes:

[0023] Obtain the data of the polypeptide drug to be predicted; and

[0024] Use the trained prediction model as described above to process the data of the polypeptide drug to be predicted and obtain a prediction result.

[0025] In a third aspect, an electronic device is provided, which includes: a processor, and a memory coupled to the processor, where the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the training method according to any possible implementation manner in the first aspect, or executes the prediction method in the second aspect.

[0026] In a fourth aspect, a computer-readable storage medium is provided, including a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method according to any possible implementation manner in the first aspect, or execute the prediction method in the second aspect.

[0027] In a fifth aspect, a computer program product is provided, including: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method according to any possible implementation manner in the first aspect, or execute the prediction method in the second aspect. Description of the Drawings

[0028] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or related technologies will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is a schematic diagram for constructing a method for predicting the in vivo pharmacokinetic properties of a polypeptide drug provided by an embodiment of the present application;

[0030] Figure 2 It is an application process for gradually predicting the toxicity of a novel polypeptide drug provided by an embodiment of the present application;

[0031] Figure 3 It is a schematic diagram of a novel polypeptide toxicity gradual prediction framework MLR-GAT provided by an embodiment of the present application. Detailed Embodiments

[0032] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0033] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the sequence in the flowchart. The terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0035] Before introducing the embodiments of this application, a brief description of the current technical research of this application is given first:

[0036] Currently, the mainstream ADMET prediction technologies mainly focus on two types of drugs, small molecules and proteins, and their method systems are difficult to adapt to the particularity of polypeptides. First, small molecule ADMET prediction is mostly based on deep learning methods of molecular descriptors or SMILES strings, usually assuming molecular rigidity and difficult to capture conformational changes in different environments. The secondary structure and flexible segments of polypeptides may change significantly in solvent, pH value or receptor environment, affecting transmembrane permeability and metabolic stability. Protein ADMET prediction mainly relies on sequence information or coarse-grained molecular docking, which is suitable for protein-protein interactions and metabolic enzyme recognition. However, the metabolic stability of polypeptides is affected by enzymatic cleavage sites and local conformations, and existing protein methods are difficult to apply directly.

[0037] Therefore, there are the following deficiencies in the current polypeptide ADMET prediction methods:

[0038] 1. Lack of pertinence: The current ADMET prediction methods mainly target small molecules or proteins and do not fully consider the unique properties of polypeptides. For example, small molecule methods often rely on molecular descriptors or SMILES sequences, while protein methods focus on sequence information or whole protein docking simulations, both of which are difficult to accurately describe key features such as conformational changes and enzymatic cleavage sites of polypeptides, resulting in limited prediction accuracy.

[0039] 2. Lack of modification information: Polypeptide drugs often contain various chemical modifications (such as methylation, phosphorylation, glycosylation, etc.), and these modifications have important effects on their metabolic stability, transmembrane permeability and protein binding ability. However, the existing ADMET prediction methods are usually based on natural amino acid compositions and are difficult to effectively model and evaluate the pharmacokinetic behaviors of modified polypeptides, resulting in a large deviation between the prediction results and the actual situation.

[0040] 3. Lack of consideration for biological differences: Existing ADMET prediction methods are usually trained and validated based on a single species (such as humans), a specific organ (such as the liver), or a specific cell line (such as Caco-2), making it difficult to accurately predict the pharmacokinetic properties of polypeptides in different biological systems. However, the processes of absorption, distribution, metabolism, and excretion of polypeptide drugs often involve interactions between multiple organs and differences between different species. For example, the metabolic pathways of certain polypeptides in rodents may be significantly different from those in humans, making it difficult to directly extrapolate animal experiment data to clinical applications. Therefore, the lack of systematic modeling across species, organs, and cell lines limits the applicability and reliability of existing ADMET prediction tools in the development of polypeptide drugs.

[0041] 4. Lack of a comprehensive prediction platform: Current research on polypeptide ADMET prediction is still in its early stages, lacking a systematic and integrated prediction platform. Existing methods often focus on a specific property and fail to combine multiple prediction models to provide a comprehensive ADMET assessment. In addition, the data standards of models developed by different research teams are not unified, making it difficult to compare and apply the results, which limits the practical value of polypeptide drugs in early screening and optimization.

[0042] To address the above technical deficiencies, the embodiments of the present application provide an intelligent prediction method for the in vivo pharmacokinetics of polypeptide drugs, which can effectively learn the sequence characteristics of polypeptides and improve the accuracy and robustness of the prediction of the ADMET properties of polypeptide drugs. It can be clearly concluded from the following description that the embodiments of the present application also relate to a training method for an intelligent prediction model of the in vivo pharmacokinetics of polypeptide drugs.

[0043] The embodiments of the present application provide an intelligent prediction method for the in vivo pharmacokinetics of polypeptide drugs, as Figure 1 described, specifically including:

[0044] S1. Data collection and preprocessing: Collect the sequence and modification information of the polypeptide to be predicted and the corresponding ADMET-related experimental data, including physicochemical properties, bioavailability, blood-brain barrier permeability, permeability, half-life, and toxicity, etc., and perform data cleaning and standardization processing to ensure the high quality and consistency of the data;

[0045] S2. Molecular descriptor calculation: Calculate the molecular descriptors of the polypeptide, including physicochemical characteristics, topological properties, electronic properties, etc., to provide input features for subsequent machine learning models;

[0046] S3. Machine learning model construction: Use multiple machine learning algorithms for training to predict the physicochemical properties of the polypeptide, and improve the prediction performance of the model through cross-validation and hyperparameter optimization;

[0047] S4, Polypeptide half-life prediction: Based on the AlphaPeptDeep deep learning framework, a transfer learning method is used for polypeptide half-life prediction to improve the robustness of model prediction;

[0048] S5, Polypeptide toxicity prediction: The step-by-step prediction framework MLR-GAT for polypeptide toxicity is adopted to achieve accurate prediction of polypeptide toxicity;

[0049] S6, Model performance evaluation: The performance of the prediction models for each property is evaluated respectively, and the best models are retained;

[0050] S7, Construction of a polypeptide ADEMT property prediction platform: Integrate the above machine learning and deep learning models to build an integrated polypeptide ADMET prediction platform, provide a visual interface, and users can input polypeptide sequences and modification information to obtain ADMET prediction results.

[0051] Specifically, the S1 data collection and preprocessing method includes: Some data manually collected and screened from multiple open-access databases and peer-reviewed literature. All data is classified according to ADMET attributes and further subdivided into multiple subcategories according to the specific meaning of each endpoint. To ensure the consistency and reliability of the data, a series of strict preprocessing steps are performed, including missing value handling, duplicate value handling, etc.

[0052] Specifically, the S2 molecular descriptor calculation method includes: The characterization of polypeptide molecules is mainly achieved through two calculation methods: one is the molecular descriptors and fingerprints calculated based on SMILES strings, and the other is the polypeptide feature descriptors calculated based on sequence information. And variance filtering, correlation filtering, and random forest recursive feature elimination method are used to select the optimal feature subset for the descriptors.

[0053] Specifically, the S3 machine learning model construction method includes: Based on traditional machine learning methods for prediction, six advanced algorithms (Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Machine (SVM), Decision Tree (DT), Gradient Boosting Tree (GBT), and Light Gradient Boosting Machine (LightGBM)) are combined with different descriptors to build and optimize the model.

[0054] Specifically, the S4 polypeptide half-life prediction method includes: applying the AlphaPeptDeep deep learning architecture and introducing a transfer learning strategy. The core idea of this method is to utilize the potential physicochemical association between the polypeptide retention time and the half-life to establish a shared knowledge learning model, thereby improving the prediction accuracy. In the implementation process, AlphaPeptDeep extracts the sequence features and physicochemical properties of the polypeptide through a deep neural network and learns the retention time pattern of the polypeptide during the pre-training stage. Subsequently, in the polypeptide half-life prediction task, the model shares the learned knowledge and further adjusts the weights to more accurately model the influencing factors of the half-life. Finally, this method can effectively capture the stability characteristics of the polypeptide in different environments, thus achieving more accurate half-life prediction.

[0055] Specifically, as Figure 2 shown, the S5 polypeptide toxicity prediction method includes: proposing a new polypeptide toxicity step-by-step prediction framework MLR-GAT, which can handle multi-task learning and multi-classification problems simultaneously. First, the model will initially classify the input polypeptide into two categories: toxic and non-toxic. Subsequently, for the polypeptides determined to be toxic, the toxicity categories are further refined, including hemostatic injury toxicity, cytotoxicity, G protein-coupled receptor toxicity, neurotoxicity, cytolytic toxicity, and hemolytic toxicity, to more accurately identify their potential hazards. On this basis, for neurotoxins, the model further subdivides their mechanisms of action, including acetylcholine receptor inhibition toxicity, Na + , K + , Ca 2+ ion channel toxicity. In addition, the system also introduces a regression model to predict the HC 50 value to quantify the toxicity level.

[0056] Specifically, the S6 model evaluation method includes: for classification tasks, the evaluation metrics include accuracy (ACC):

[0057]

[0058] Area under the ROC curve (AUC):

[0059]

[0060] F1 value (F1):

[0061]

[0062] where i is the class index, TP is the true positive, that is, predicted as positive and actually positive; FP is the false positive, predicted as positive but actually negative; FN is the false negative, that is, predicted as negative but actually positive; TN is the true negative, predicted as negative and actually negative.

[0063] For regression tasks, the evaluation metrics include the coefficient of determination (R 2 ):

[0064]

[0065] Root Mean Square Error (RMSE):

[0066]

[0067] Mean Absolute Error (MAE):

[0068]

[0069] where N is the number of samples, y i is the true value, is the predicted value, is the average value of the true values.

[0070] Specifically, the method for building the S7 polypeptide ADEMT property prediction platform includes: building a platform that provides an intuitive user interface, facilitating access to five major functional modules through a navigation bar: Home, Services, Documentation, Help, About us. Among them, the Home module provides an overview of the platform's functions and features for users, helping users quickly understand the overall framework and objectives of the platform. The Services module focuses on the prediction and calculation of key ADMET properties, integrating multiple efficient prediction tools, supporting SMILES string input or drawing molecular structures through a structure editor. In addition, the platform also requires users to provide the sequence information of the polypeptide and offers 35 modification types for half-life prediction. After users submit relevant data, the platform automatically calculates molecular descriptors and makes predictions through the Web server background, and the results are displayed in real time through an interactive page, providing detailed prediction results and corresponding molecular information. The Documentation module elaborates in detail on the data collection, model construction, and algorithm implementation processes. The Help module provides users with a concise teaching guide to help users quickly master the usage method of the platform. The About us module displays information about the R & D team and introduces the development background of the platform and team members.

[0071] In summary, compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0072] 1. This platform is the first comprehensive platform that can achieve a comprehensive prediction of polypeptide ADMET properties. By combining machine learning and deep learning, an efficient and accurate polypeptide ADMET prediction method and platform have been established.

[0073] 2. It can realize the characterization of modified peptides, thereby improving the prediction accuracy. It is specifically designed for peptide ADMET prediction, can adapt to peptides of different modified types, and provide personalized predictions.

[0074] 3. Biological differences are fully considered in the prediction of some properties, making the prediction adaptable to a variety of scenarios.

[0075] The above-mentioned embodiments of the present application are described in detail below through a combination of multiple application examples.

[0076] Application Example 1: Peptide property prediction experiment based on machine learning

[0077] This application example uses a variety of advanced machine learning methods to build a high-performance prediction model to verify the effectiveness and applicability of the methods.

[0078] (1) Data source and preprocessing

[0079] This study used peptides and their related property data from public databases and literature. All data were classified according to ADMET properties and further subdivided into multiple subcategories based on the specific significance of each endpoint, including absorption distribution properties such as LogD 7.4 , permeability, blood-brain barrier membrane permeability, bioavailability; metabolic excretion properties include half-life; toxicity includes overall toxicity, which can be specifically divided into G protein-coupled receptor toxicity, cytotoxicity, hemostatic injury toxicity, cytolytic toxicity, neurotoxicity, and hemolytic toxicity; among them, neurotoxicity is further divided into Na + , K + , Ca 2+ Ion channel toxicity and acetylcholine receptor inhibition toxicity, hemolytic toxicity also includes hemolytic activity index HC 50 . Finally, 17 independent data sets were constructed for ADMET prediction modeling. To ensure the consistency and reliability of the data, a series of rigorous preprocessing steps were performed: 1) The labels and experimental values ​​of all data were extracted, the units were unified, and records with missing values ​​were removed. For regression data, if a molecule has multiple data entries, the arithmetic mean is taken. 2) Only peptide sequences containing 2 to 50 amino acids are retained. 3) All peptide sequences and SMILES are checked and verified one by one. In the case of sequence missing, the corresponding peptide sequence is manually extracted according to SMILES; for peptides with missing SMILES, natural peptides are converted using the RDkit package, and modified peptides are manually modified according to the modification characteristics. 4) Duplicates are eliminated by calculating the InChIKey value to ensure the uniqueness and accuracy of the data. After the above processing, a total of 36,643 high-quality data were obtained, covering the prediction information of 27 key endpoints.

[0080] (2) Descriptor calculation

[0081] The characterization of polypeptide molecules is mainly achieved through two computational methods: one is molecular descriptors and fingerprints calculated based on SMILES strings, and the other is polypeptide feature descriptors calculated based on sequence information. Specifically, the MOE software and PyBioMed are used for the calculation of small molecule descriptors based on SMILES strings, while RDKit is used to calculate MACCS fingerprints, PubChem fingerprints, and Pharmacophore ErG fingerprints; PyBioMed and modlAMP are used to calculate polypeptide descriptors based on sequence information. In the feature selection stage, descriptors with a variance less than 0.01 are first removed, and redundancy processing is performed on features with a correlation greater than 0.99 between descriptors, randomly removing one of the highly correlated descriptors. Next, the random forest recursive feature elimination method is applied in combination with a 5-fold cross-validation strategy to select the optimal subset for subsequent modeling.

[0082] (3) Model construction

[0083] Six machine learning algorithms (Random Forest RF, XGBoost, Support Vector Machine SVM, Decision Tree DT, Gradient Boosting Tree GBT, and LightGBM) are used to train in combination with different descriptors, and hyperparameters are adjusted to optimize the model performance. Specifically, LogD 7.4 , bioavailability, and blood-brain barrier membrane permeability mainly use small molecule and polypeptide descriptors calculated by PyBioMed and modlAMP as inputs; in particular, for permeability, we constructed models for the commonly used in vitro models Caco-2, PAMPA, and RRCK respectively, and calculated different molecular descriptors and molecular fingerprints, including MOE2D descriptors, peptide descriptors, small molecule descriptors, and MACCS fingerprints. In addition, using GNN further improved the permeability prediction performance of the model on large datasets. By taking SMILES and sequences as inputs, GNN integrates molecular graph and descriptor information from two paths and extracts key features through an attention mechanism to generate the final prediction result. All datasets are randomly divided into training sets and test sets in a ratio of 0.75:0.25, and 5-fold cross-validation is performed on the training sets to ensure that the model has good generalization ability.

[0084] (4) Experimental results

[0085] Regression models were successfully constructed for LogD7.4 and permeability properties; binary classification models were successfully constructed for bioavailability and blood-brain barrier membrane permeability. The evaluation results of the models on the test sets are shown in Tables 1 and 2.

[0086] Table 1: Results of machine learning modeling regression tasks

[0087] Model <![CDATA[Q 2 > <![CDATA[RMSE cv > <![CDATA[MAE cv > <![CDATA[R 2 > <![CDATA[RMSE test > <![CDATA[MAE test > <![CDATA[LogD 7.4 > 0.82 0.466 0.324 0.818 0.491 0.318 Caco2-L 0.411 0.747 0.551 0.435 0.765 0.582 Caco2-C 0.582 0.425 0.307 0.527 0.472 0.335 Caco2-A 0.50 0.548 0.388 0.476 0.573 0.408 PAMPA-C 0.646 0.428 0.317 0.657 0.423 0.308 RRCK-C 0.691 0.34 0.269 0.623 0.42 0.32

[0088] Among them, the best model for LogD 7.4 is the GBT-based model, and its R 2 in the test set reaches 0.820, higher than 0.818 in cross-validation, reflecting its superior predictive performance. Caco2-L represents the permeability model of linear peptides in the Caco-2 cell line (the best model is the SVR model based on MOE2D descriptors), Caco2-C represents the permeability model of cyclic peptides in the Caco-2 cell line (the best model is the LightGBM model based on MOE2D descriptors), Caco2-A represents the permeability model of linear and cyclic peptides together in the Caco-2 cell line (the best model is the RF model based on MOE2D descriptors), PAMPA-C represents the permeability model of cyclic peptides in the PAMPA cell line (the best model is the RF model based on MOE2D descriptors), RRCK-C represents the permeability model of cyclic peptides in the RRCK cell line (the best model is the RF model based on MOE2D descriptors), that is, we considered the phenomenon that the permeability of different types of polypeptides varies in different cell lines. As can be seen from the table, the performance of the model in the test set is not inferior to the results of cross-validation in the training set, reflecting the excellent robustness of the model.

[0089] Table 2: Results of machine learning modeling classification tasks

[0090]

[0091] Among them, the best model for bioavailability is the LightGBM-based model, with an ACC of 0.811, an AUC of 0.901, and an f1 of 0.761 in cross-validation, and an ACC of 0.844, an AUC of 0.90, and an f1 of 0.80 in the test set; the best model for blood-brain barrier membrane permeability is the RF-based model, with an ACC of 0.808, an AUC of 0.894, and an f1 of 0.805 in cross-validation, and an ACC of 0.78, an AUC of 0.889, and an f1 of 0.789 in the test set, indicating that the performance of the model in the test set is also excellent, reflecting the good robustness and generalization ability of the model.

[0092] Application Example 2: Prediction of polypeptide half-life based on transfer learning

[0093] This application example uses the AlphaPeptDeep transfer learning architecture to improve the prediction ability of polypeptide half-life.

[0094] (1) Data collection and construction of benchmark dataset

[0095] Data were collected from the public polypeptide half-life database and the literature, and the collected data were statistically analyzed by different species, organs, and types. Finally, peptide half-life data under five different conditions were obtained, including: the half-life of natural peptides in human blood (HBN), the half-life of modified peptides in human blood (HBM), the half-life of natural peptides in mouse blood (MBN), the half-life of modified peptides in mouse blood (MBM), and the half-life of modified peptides in mouse intestine (MIM). These data contain 117, 187, 106, 182, and 378 peptides and their corresponding half-lives respectively, totaling 970 high-quality data.

[0096] (2) Model Pretraining and Fine-tuning

[0097] First, a deep learning model was trained on a large polypeptide retention time dataset to learn the physicochemical properties of polypeptides. Then, it was fine-tuned on the half-life dataset to enhance the prediction ability. The dataset was divided into a training set and a test set according to the ratio of 75%:25%.

[0098] (3) Experimental Results

[0099] The coefficient of determination (R 2 ), root mean square error (RMSE), and mean absolute error (MAE) metrics were used for evaluation. The results showed that for the five datasets HBN, HBM, MBN, MBM, and MIM, compared with the results of the best traditional machine learning models, the R 2 values of the model after transfer learning increased by 9%, 32%, 16.4%, 44%, and 15% respectively, indicating that the transfer learning method has high applicability in small-sample scenarios. The results of the model after transfer learning are shown in Table 3.

[0100] Table 3: Half-life Prediction Results Using the AlphaPeptDeep Deep Learning Architecture

[0101] Model <![CDATA[Q 2 > <![CDATA[RMSE cv > <![CDATA[MAE cv > <![CDATA[R 2 > <![CDATA[RMSE test > <![CDATA[MAE test > HBN 0.86 202.34 68.222 0.84 272.2 114.19 HBM 0.92 133.71 56.95 0.9 147.65 81.03 MBN 0.997 12.72 5.56 0.984 23.15 11.49 MBM 0.87 0.60 0.30 0.93 0.49 0.34 MIM 0.93 0.62 0.62 0.94 0.54 0.34

[0102] Among them, HBN refers to the half-life of natural polypeptides in human blood, HBM refers to the half-life of modified polypeptides in human blood, MBN refers to the half-life of natural polypeptides in mouse blood, MBM refers to the half-life of modified polypeptides in mouse blood, and MIM refers to the half-life of modified polypeptides in mouse intestine. That is, we considered the modification information of polypeptides and the half-life differences in different species and different organs. It can be seen from the table that the R 2 values of all models are above 0.8, and the test set results are similar to the training set results, reflecting the powerful performance of transfer learning in few-shot tasks. Among them, the overall performance of the MBN model is the best.

[0103] Application Example 3 Application of the MLR-GAT Framework

[0104] This application example verifies the effectiveness of the proposed MLR-GAT framework in the task of polypeptide toxicity prediction.

[0105] (1) Data collection and cleaning

[0106] Collect polypeptide data known to be toxic from public databases and supplement the hemolytic peptide sequences in the hemolysis database. Negative data excludes any records with toxic functions from publicly authorized and accessible data sources to ensure that these sequences do not contain any functional features that may cross-react with positive peptides, thereby improving the accuracy of the model. The dataset consists of peptides with 50 or fewer amino acids. Sequences with duplicates and non-standard amino acid residues (such as X, B, Z, etc.) are removed, resulting in 1743 positive data and an equal amount of negative data.

[0107] (2) Model structure

[0108] In this example, MLR-GAT combines a hybrid representation of molecular graphs and molecular descriptors, adopting a dual-path structure of graph convolutional network (GCN) and multi-layer perceptron (MLR). The design of this model processes different types of input data through two paths respectively, and finally fuses the features of both to make accurate predictions. The input of Path 1 is the SMILES representation of a molecule, which is a common string description of molecular structures. The SMILES string is first converted into a molecular graph representation, where the nodes in the graph represent atoms and the edges represent chemical bonds between atoms. Next, through the graph convolutional network (GCN), the model captures the relationships between atoms and their adjacent atoms, extracting the basic features of atoms and bonds. GCN can effectively propagate information in graph-structured data, gradually learning the mutual relationships between nodes (atoms) and their neighbor nodes, and finally generating the high-order features of the molecular graph. These features pass through a further feature transformation layer to form the final representation of the molecular graph, which can reflect the potential behaviors and properties of molecules in chemical reactions. Path 2 takes the polypeptide sequence and the SMILES string as inputs, calculates small molecule descriptors and polypeptide descriptors. Small molecule descriptors are quantitative characterizations of molecular structures, usually including multiple physical and chemical properties such as molecular weight, polar surface area, and molecular folding degree, which can help the model understand the basic properties of molecules. Polypeptide descriptors are based on the analysis of polypeptide sequences, capturing the composition, arrangement order of amino acids, and other biological features. Path 2 extracts key descriptors through a multi-layer perceptron (MLP). The output results of the two paths are fused to generate a comprehensive feature representation. This representation combines the structural information of the molecular graph and the quantitative features of the descriptors, providing a more comprehensive molecular representation for the model to perform subsequent prediction tasks. The fused features are input into a fully connected layer for final prediction. The fully connected layer maps the extracted features to the output space of the target task through deep linear transformations. To further optimize the performance of the model in multi-task learning, a shared weighted network structure is adopted. This structure allows the model to share some network parameters while assigning different attention weights to each task. In this way, the model can dynamically adjust its learning focus according to the characteristics of each task, thereby improving the prediction accuracy of each task ( Figure 3 )。

[0109] (3) Optimization strategies

[0110] ① Data partitioning: The dataset is partitioned into a training set, a validation set, and a test set in the ratio of 8:1:1, and the model is trained through 5 random splits.

[0111] ② Loss function: For binary classification tasks, the BCEWithLogitsLoss loss function is used:

[0112] Loss = -(y n log(σ(x n))+(1 - y n ) log(1 - σ(x n )))

[0113] where x n is the sample input, y n is the corresponding output, and σ(x) is the Sigmoid function:

[0114]

[0115] For multi - classification tasks, the CrossEntropyLoss loss function is used:

[0116]

[0117] where N is the number of samples, represents the predicted probability of the i - th sample in the true class y i .

[0118] For regression tasks, the MSELoss loss function is used:

[0119]

[0120] where N is the number of samples, where x n is the sample input, y n is the corresponding output.

[0121] ③ Data imbalance handling: For binary - classification tasks, adjust the positive - sample weight (pos_weight), and for multi - classification tasks, adjust the class weights (weight). For the data imbalance problem, binary - classification tasks adjust the positive - sample weight (pos_weight), while multi - classification tasks adjust the weights of each class (weight). The dataset is divided into a training set, a validation set, and a test set in the ratio of 8:1:1, and the model is trained through 5 random splits.

[0122] (4) Experimental results

[0123] For classification tasks, ACC, AUC, F1 - score are used, and for regression tasks, R 2 , MAE, RMSE are used for evaluation. The specific results are shown in Tables 4 and 5.

[0124] Table 4: Prediction results of regression tasks using the MLR - GAT polypeptide progressive prediction architecture

[0125] Model <![CDATA[Q 2 > <![CDATA[RMSE cv > <![CDATA[MAE cv > <![CDATA[R 2 > <![CDATA[RMSE test > <![CDATA[MAE test > <![CDATA[HC 50 > 0.703 0.376 0.278 0.474 0.479 0.360

[0126] It can be seen that the Q2 of the HC50 prediction on the validation set is 0.703, and the R2 on the test set is 0.474, indicating that the model has a strong fitting ability for the validation set. The results of the test set are slightly lower than its performance on the validation set, indicating that the model faces certain challenges when dealing with unknown data.

[0127] Table 5: Prediction Results of Classification Tasks Using the MLR-GAT Polypeptide Hierarchical Prediction Architecture

[0128] Model <![CDATA[ACC cv > <![CDATA[AUC cv > <![CDATA[f1 cv > <![CDATA[ACC test > <![CDATA[AUC test > <![CDATA[f1 test > <![CDATA[Toxin binary > 0.894 0.965 0.904 0.794 0.885 0.813 <![CDATA[Toxin type > 0.951 0.995 0.84 0.885 0.949 0.634 <![CDATA[Toxin neuro > 0.923 0.994 0.923 0.736 0.905 0.683

[0129] The results in Table 5 show that the multi-task multi-classification model MLR-GAT constructed in Example 3 has good performance in the toxicity classification prediction task. Among them, Toxin binary refers to the toxicity binary classification model, whose ACC on the test set is 0.794, AUC is 0.885, and f1 score is 0.813; Toxin type refers to the toxicity six-classification model, whose ACC on the test set is 0.885, AUC is 0.949, and f1 score is 0.634; Toxin neuro refers to the neurotoxicity four-classification model, whose ACC on the test set is 0.736, AUC is 0.905, and f1 score is 0.683. These results all show that the multi-task model can accurately capture polypeptide sequence information and achieve accurate toxicity prediction.

[0130] Construction of a Polypeptide Property Prediction System Based on an Intelligent Computing Platform in Application Example 4

[0131] This application example provides an integrated, multi-algorithm, and scalable intelligent computing platform for polypeptide property prediction, supporting model training, inference, and visual analysis.

[0132] The underlying architecture of this application example is deployed on a cloud server, supporting efficient programming development and model training. The platform adopts a layered architecture design, clearly separating business logic from the user interface to simplify system upgrades and maintenance. The front-end constructs a dynamic interactive interface to enhance the user experience and operation fluency. The back-end combines efficient request processing and load balancing mechanisms to ensure the stable operation of the system and support for high concurrency. In terms of data storage, the platform adopts a reliable database system to achieve efficient data management and retrieval. The platform integrates multiple advanced machine learning frameworks to support model training and inference, and combines cheminformatics methods for molecular descriptor calculation. As an open service system, this platform allows users to directly access its core functions, providing convenient and efficient polypeptide ADMET property prediction capabilities.

[0133] An embodiment of the present application further provides an electronic device, including: a processor, and a memory coupled to the processor, where the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above embodiments.

[0134] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory.

[0135] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire device through various interfaces and lines.

[0136] The memory may be used to store the computer program. The processor realizes various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.

[0137] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0138] The embodiments of the present application also provide a storage medium, which is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0139] The embodiments of the present application also provide a computer program product, including: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is enabled to execute the method of any of the above possible implementation manners.

[0140] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.

Claims

1. A training method for a prediction model of the pharmacokinetic properties of a polypeptide drug, characterized in that, The training method includes: Obtaining polypeptide drug property prediction training data, where the training data is divided into at least two data sets, and each data set is used to predict a pharmacokinetic property of a polypeptide drug; and Using the training data to train the prediction model, where the prediction model includes at least two prediction sub-models.

2. The training method according to claim 1, wherein The first data set includes: a polypeptide sequence, a SMILES string, and permeability data of the polypeptide in different cell lines. The pharmacokinetic properties of the first polypeptide drug are absorption and distribution properties, and the absorption and distribution properties include: LogD 7.4 , permeability, blood-brain barrier membrane permeability, and bioavailability. The first predictive sub-model is used to calculate molecular descriptors, polypeptide descriptors, and fingerprints characterizing the physicochemical properties, topological properties, and electronic properties of the polypeptide drug based on the first data set, and predict the pharmacokinetic properties of the first polypeptide drug according to the molecular descriptors, polypeptide descriptors, and fingerprints. The first predictive sub-model adopts algorithms such as random forest, XGBoost, support vector machine, decision tree, gradient boosting tree, and LightGBM. The first predictive sub-model combines a graph neural network when predicting permeability.

3. The training method according to claim 2, wherein The first prediction sub-model is specifically further used for: Using a molecular descriptor calculation method to combine the polypeptide sequence and the SMILES string to obtain corresponding molecular descriptors; and Using variance filtering, correlation filtering, and random forest recursive feature elimination method to obtain a preferred feature subset from the molecular descriptors.

4. The training method according to claim 1, characterized in that The second data set includes: polypeptide sequence, modification information, half-life data of natural peptides in different species and different organs, half-life data of modified peptides in different species and different organs, and polypeptide retention time data; the second pharmacokinetic property of the polypeptide drug is metabolic excretion property, and the metabolic excretion property is half-life; the second prediction sub-model adopts the AlphaPeptDeep transfer learning architecture, and the AlphaPeptDeep transfer learning architecture learns the polypeptide retention time in the pre-training stage and shares the learned knowledge to adjust the weights in the prediction stage.

5. The training method according to claim 1, wherein The third data set includes: polypeptide sequence, SMILES string, toxicity positive or negative label, and hemolytic activity value. The third pharmacokinetic property of the polypeptide drug is toxicity. The third prediction sub-model adopts a dual-path structure of a graph convolutional network and a multi-layer perceptron. The graph convolutional network is used to extract molecular graph features, and the multi-layer perceptron is used to extract key descriptors. The third prediction sub-model is also used to fuse the molecular graph features and the key descriptors.

6. The training method according to claim 5, wherein The third prediction sub-model is specifically used for: Performing a first process on the third data set to classify the toxicity and non-toxicity of the polypeptide drug; and For the polypeptide drug determined to be toxic, performing a step-by-step classification of toxicity levels and quantifying the hemolytic toxicity level.

7. The training method according to claim 1, characterized in that The evaluation metrics of the prediction model include: classification task metrics and regression task metrics. The classification task metrics include: accuracy, area under the ROC curve, and F1 value. The regression task metrics include: coefficient of determination, root mean square error, and mean absolute error.

8. An intelligent prediction method for the pharmacokinetic properties of a polypeptide drug, characterized in that, The prediction method includes: Obtaining data of the polypeptide drug to be predicted; and Using the prediction model trained as described in any one of claims 1-7 to process the data of the polypeptide drug to be predicted to obtain a prediction result.

9. An electronic device, characterized in that, The electronic device includes: a processor, and a memory coupled to the processor, The memory is used to store a computer program; and The processor is used to execute the computer program stored in the memory so that the electronic device executes the training method described in any one of claims 1-7, or executes the prediction method described in claim 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instructions, which, when running on a computer, cause the computer to execute the training method described in any one of claims 1-7, or execute the prediction method described in claim 8.

Citation Information

Patent Citations

  • Logistic regression-based pharmacokinetic parameter prediction method for drug compound

    CN111833971A

  • Product activity value and ADMET property prediction method and system based on BPMLP-XGBoost

    CN114649065A

  • ADMET prediction electronic device based on multi-model integration and method thereof

    CN115620828A

  • Drug property prediction method and system based on pharmacokinetic model

    CN115691703A

  • Pharmacokinetics and toxicity prediction method based on coarse and fine granularity classification

    CN116206704A

Cited By

  • Neural network-based acetylcholin esterase inhibitor screening method

    CN121148477A

  • A neural network-based method for screening acetylcholinesterase inhibitors

    CN121148477B

  • Multi-modal umami peptide recognition method and system based on molecular map and sequence characteristics

    CN121415874A

  • Multi-modal umami peptide recognition method and system based on molecular graph and sequence characteristics

    CN121415874B