A machine learning-based target-specific virtual screening method and system

By constructing a target-specific scoring function through the integration of multi-conformation target features and multiple machine learning models, the problem that a single model cannot capture the complex interactions between compounds and targets is solved, enabling more efficient and accurate virtual screening and improving the efficiency and accuracy of drug development.

CN120072032BActive Publication Date: 2025-12-19SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510062182.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-12-19
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

In existing technologies, single machine learning models cannot fully capture the complex interactions between compounds and targets when constructing target-specific scoring functions, resulting in low prediction accuracy and efficiency during virtual screening.

Method used

By utilizing the dynamic features of multi-conformation targets, a target-specific scoring function is constructed through multiple machine learning models. Furthermore, the advantages of each model are integrated through a model ensemble method to build an ensemble model with high stability and excellent prediction performance.

Benefits of technology

It significantly improves the adaptability and predictive ability of virtual screening, enhances the efficiency and accuracy of drug screening, reduces experimental resource consumption, and shortens the drug development cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072032B_ABST
    Figure CN120072032B_ABST
Patent Text Reader

Abstract

The application provides a target-specific virtual screening method and system based on machine learning. Active molecules and inactive molecules are docked with multiple conformations of a target, and protein-ligand interaction features of the docked molecules and related features of the ligands are extracted as input features of a machine learning model. A scoring function model is constructed using multiple machine learning models, and based on the advantages of multiple target-specific scoring function models, an integrated target-specific scoring function model is finally obtained by combining a model integration method. The integrated model can more effectively process complex protein-ligand interaction data, has excellent prediction performance and virtual screening capability, provides more reliable prediction results, and improves the virtual screening capability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of bioinformatics analysis, and particularly relates to a target specificity virtual screening method and system based on machine learning. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] Virtual screening (VS) plays an important role in the discovery of lead compounds, which can screen potential active molecules from a large number of compounds with the help of molecular docking technology. Among them, scoring functions (SF) are often used to evaluate the binding affinity of compounds and targets, and accordingly rank the compounds. Therefore, the accuracy of SF directly affects the efficiency and accuracy of virtual screening. The classical scoring function can be described as a linear additive form to represent the relationship between the binding strength and the characteristics used to characterize the protein-ligand interaction. However, in fact, such a linear relationship does not always exist, so it is difficult to fully consider the complex protein-ligand interaction when predicting the binding affinity of ligands and proteins, and the performance of the classical scoring function may not be sufficient to meet the needs when dealing with large-scale compound libraries.

[0004] In order to improve the efficiency and accuracy of virtual screening, machine learning-based scoring function (MLSF) is developed and applied, which does not need to define the function form in advance, but relies on machine learning algorithm to learn the function form from data, so it shows more accurate prediction accuracy than the classical scoring function. According to different application scope, MLSF can be divided into general MLSF and target-specific MLSF. General MLSF is usually trained on a wide range of data sets, trying to cover a variety of different molecules and situations, but it is not optimized for specific targets, so its applicability is limited in structure-based virtual screening. Target-specific MLSF is a method designed specifically for a particular target, which can more accurately capture the interaction pattern between the specific target and the ligand, and is usually superior to general MLSF in prediction performance.

[0005] In the process of constructing target-specific scoring function based on machine learning for virtual screening, a single machine learning model is usually relied on. However, due to the different focuses of different models in feature selection, hypothesis space and optimization target, the scoring function constructed by a single model may not be able to fully capture the complex interaction between compounds and targets, resulting in lower prediction accuracy and efficiency in the virtual screening process. SUMMARY

[0006] To overcome the limitations of the prior art, the present application comprehensively captures the diversity of protein-ligand interactions by introducing the dynamic characteristics of multi-conformation targets, constructs target-specific scoring functions using multiple machine learning models, and fuses the advantages of each model through model integration methods to construct target-specific scoring functions with high stability, excellent prediction performance, and outstanding virtual screening capabilities, which can better reflect the complexity of real biological systems, significantly improve adaptability and prediction ability, while improving drug screening efficiency and accuracy, reducing experimental resource consumption, and shortening the drug development cycle.

[0007] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0008] In a first aspect, the present application provides a target-specific virtual screening method based on machine learning, comprising:

[0009] obtaining training data, the training data comprising: multi-conformation structures of targets and active molecules and inactive molecules of targets;

[0010] molecular docking the active molecules and inactive molecules of the target with the multi-conformation structures of the target; after molecular docking is completed, extracting protein-ligand interaction features and ligand-related features and performing feature engineering;

[0011] training the features processed through feature engineering on multiple machine learning models to obtain multiple trained target-specific scoring function models, integrating the multiple target-specific scoring function models to obtain an integrated model;

[0012] molecular docking the compounds in the compound library to be screened with the multi-conformation structures of the target, extracting corresponding protein-ligand interaction features and ligand-related features, and obtaining a prediction result based on the extracted features through the integrated model, and screening candidate compounds with high activity potential according to the prediction result.

[0013] In a second aspect, the present application provides a target-specific virtual screening system based on machine learning, comprising:

[0014] a training data acquisition module configured to obtain training data, the training data comprising: multi-conformation structures of targets and active molecules and inactive molecules of targets;

[0015] a docking module configured to molecularly dock the active molecules and inactive molecules of the target with the multi-conformation structures of the target;

[0016] a feature extraction module configured to, after molecular docking is completed, extract protein-ligand interaction features and ligand-related features and perform feature engineering;

[0017] a machine learning scoring function model construction module configured to train the features after feature engineering on multiple machine learning models to obtain multiple target-specific scoring function models after training;

[0018] a machine learning scoring function model integration module configured to integrate the multiple target-specific scoring function models to obtain an integrated model;

[0019] a virtual screening module configured to perform molecular docking of compounds in a compound library to be screened to multiple conformational structures of a target, extract corresponding protein-ligand interaction features and ligand-related features, and obtain a prediction result based on the extracted features through the integrated model, and screen out candidate compounds with high activity potential according to the prediction result.

[0020] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, and computer instructions stored in the memory and running on the processor, when the computer instructions are run by the processor, the method of the first aspect is completed.

[0021] In a fourth aspect, the present application provides a computer readable storage medium for storing computer instructions, when the computer instructions are executed by a processor, the method of the first aspect is completed.

[0022] In a fifth aspect, the present application provides a computer program product comprising a computer program, when the computer program is executed by a processor, the method of the first aspect is completed.

[0023] The above one or more technical solutions have the following beneficial effects:

[0024] In the present application, active molecules and inactive molecules are molecularly docked to multiple conformations of a target, protein-ligand interaction features and ligand-related molecular features are extracted as input features of a machine learning model; then, multiple machine learning models are used to construct target-specific scoring function models, and the advantages of different models are fused through model integration method to generate an integrated scoring function model. The integrated model can more efficiently process complex protein-ligand interaction data and provide more stable and reliable prediction results, thereby significantly improving the ability and efficiency of virtual screening.

[0025] The advantages of the additional aspects of the present application will be partially given in the following description, partially will become apparent from the following description, or will be understood by the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0026] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification. The embodiments of these drawings are set to explain the application, and do not constitute an improper limitation to the application.

[0027] Figure 1 A step flow chart for constructing a scoring function targeting hepatitis B virus nucleocapsid protein based on a machine learning model in the first embodiment of the application;

[0028] Figure 2 A step flow chart for feature engineering provided in the first embodiment of the application;

[0029] Figure 3 A structural schematic diagram of a machine learning-based target-specific virtual screening system provided in the second embodiment of the application. DETAILED DESCRIPTION

[0030] The application will be further described below in conjunction with the drawings and embodiments.

[0031] It should be noted that the following description is intended to further explain the specific embodiments of the application, all of which are exemplary descriptions and do not constitute a limitation on the protection scope of the application. Unless explicitly defined otherwise in the context, all technical and scientific terms used in this specification should be interpreted according to the conventional understanding of ordinary skilled persons in the art.

[0032] The various embodiments of the application and the features therein can be freely combined without conflict, thereby providing more technical variants and application schemes. Although the embodiments listed in the application contain specific technical details, it should be understood that these technical details should not limit the protection scope of the application. On the contrary, the protection scope of the application should be determined by the claims, and should cover all equivalent structures and variations within the spirit and scope of the application.

[0033] For example, although in the specific embodiments of the application, four kinds of machine learning models such as support vector machine, random forest, extreme gradient boosting and artificial neural network are used to construct the scoring function, and voting method is used for integration, other types of machine learning models and model integration methods can also be used to achieve similar functions without deviating from the technical scope of the application.

[0034] In addition, adaptive adjustments may be made to the construction of the data set, feature extraction and processing, model training and evaluation steps during implementation, and these adjustments should also be included in the protection scope of the application. Therefore, the application should not be limited to the specific embodiments in the detailed description, but should cover all variations and improvements made according to the basic principles of the application

[0035] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0036] It should be noted that the terms used herein are only intended to describe specific embodiments and are not intended to limit exemplary embodiments according to the present application.

[0037] Embodiment one

[0038] This embodiment is illustrated by taking the Hepatitis B Virus (HBV) Capsid protein (Cp) target as an example. HBV Cp plays a crucial role in various stages of the viral life cycle, especially in the maintenance and replication of viral genome cccDNA. The diversity of Cp function means that Cp inhibitors targeting Cp have the potential to inhibit multiple aspects of the viral replication cycle and inhibit the amplification of cccDNA in cells. Therefore, HBV Cp has become an important target for the development of antiviral drugs. By targeting HBV Cp, the viral life cycle can be effectively interfered with, and viral replication and transmission can be blocked, which has the potential to cure HBV. Research has found that the target-specific scoring function constructed based on machine learning and high-precision molecular docking methods not only has stronger classification accuracy for active molecules and inactive molecules, but also has significantly better enrichment ability for active compounds than classic scoring functions.

[0039] Reference Figure 1 , shows the step flow chart of the target-specific virtual screening method based on machine learning provided by the embodiments of the present application, specifically comprising the following steps:

[0040] Step 101: Construct a data set targeting HBV Cp, the data set including structure data, activity data, and Cp multi-conformation structure data of active molecules and inactive molecules, wherein the activity data includes IC 50 , EC 50 , K d , and K i .

[0041] In the specific implementation of this example, the step of constructing a data set targeting HBV Cp includes:

[0042] The structure data and activity data of active molecules and inactive molecules are preferably from literature, PubChem, ChEMBL, BindingDB, DUD-E, laboratory internal experimental data, etc., but are not limited to these sources; the multi-conformation structure data of the target is preferably from the database of experimentally resolved target structures or theoretically modeled predicted target structures, but is not limited to the target multi-conformation structures from the above-mentioned sources.

[0043] As an embodiment, the structure data and activity data of active molecules and inactive molecules are derived from the three data sources of PubChem, ChEMBL and BindingDB, and the multi-conformation structure data of HBV Cp is derived from the RCSB PDB data source and the Cp structure predicted by AlphaFold2. As shown in Table 1, the data composition of the data set targeting HBV Cp is given.

[0044] Table 1 Data composition of the data set targeting HBV Cp

[0045]

[0046] Table 1 (continued)

[0047]

[0048]

[0049] Based on the structure data and activity data of active molecules and inactive molecules, the compounds are pre-processed and screened: a series of pre-processing is performed on the compound structure to obtain standardized structures, which are saved as SDF files, wherein the pre-processing includes, for example, removing organometallic compounds, protonating and hydrogenating. At the same time, repeated compounds with large differences in inhibitory effect (EC 50 / IC 50 / K d / K i difference of one order of magnitude are removed, while for repeated compounds with a difference within one order of magnitude, their average values are taken to ensure the accuracy of the data.

[0050] Based on the screened compounds, the data is arranged according to the following rules: compounds with EC 50 , IC 50 , K d , K i values below 10 μM are classified as active molecules of Cp, and compounds with EC 50 , IC 50 , K d , K i values higher than 50 μM are classified as inactive molecules. At the same time, for EC 50 , IC 50 , K d, K i Compounds with activity values between 10 and 50 mM are discarded.

[0051] Based on the finished compounds, structural clustering is performed, the best structure with activity value is retained in similar compound structure, the uniqueness of each compound structure is ensured, and after clustering, 212 active compounds and 94 inactive compounds with unique structure are collected. Subsequently, in the DUD-E database, 4257 decoys, i.e. inactive molecules, are generated based on active compounds, and finally 212 active compounds and 4351 inactive compounds are obtained, the ratio of active molecules to inactive molecules is about 1:20, the construction of the compound dataset is completed, the active molecule label is set to 1, and the inactive molecule label is set to 0, which is convenient for subsequent construction of machine learning model.

[0052] In the RCSB PDB database, the structure files of HBV Cp multi-conformation are obtained, and the PDB numbers are 5WRE and 6J10 respectively; based on the amino acid sequence of HBV Cp, the Cp structure is predicted by AlphaFold2, and the subsequent pretreatment of the above protein is carried out, including adding missing atoms and residues, automatically determining the protonation state of the residues according to the physiological pH condition and adding hydrogen atoms, etc. The multi-conformation structure data of HBV Cp is successfully constructed.

[0053] After the construction of the dataset targeting HBV Cp, step 102 is performed.

[0054] Step 102: Molecular docking of the collected active molecules and inactive molecules with the multi-conformation of HBV Cp, wherein the target conformation includes the crystal structure derived from experimental analysis and the structure predicted by theoretical model.

[0055] This embodiment adopts a molecular docking method with conformer-dependent quantum mechanics charges (MDCC) described in patent CN115206441A. This method is based on the Glide SP method in the traditional molecular docking method, i.e. , based on different conformations of small molecule ligands, the charge fitting is performed by quantum chemistry calculation method, and multiple sets of RESP charge parameters are obtained for each drug molecule according to its different conformations, fully considering the influence of conformation change on charge distribution.

[0056] Among them, the specific implementation steps of the MDCC docking method are as follows: first, the quantum chemistry calculation software (Gaussian16) is used to optimize the structure of the collected active and inactive small molecule ligands, then The Conformational Search module in Schrodinger performs conformational search to obtain the top 10 conformations of the small molecule ligand with the lowest energy; then, the top 10 conformations of the small molecule ligand with the lowest energy are subjected to structure optimization and frequency analysis using Gaussian16, the restrained electrostatic potential (RESP) charges of each conformation are calculated using Multiwfn 3.8 software, and the charges are assigned to the 10 conformations of the optimized small molecule ligand, thereby obtaining the charge distribution of the small molecule ligand in different conformations; finally, based on the multi-conformation dependent charges of the small molecule ligand, the docking method of MDCC is used to dock the small molecule ligand with the protein target, and the docking method of MDCC is used to dock the small molecule ligand with the protein target. Molecular docking is performed under the OPLS_2005 force field mode.

[0057] After the molecular docking is completed, step 103 is performed.

[0058] Step 103: Extract and process protein-ligand interaction features and related features of the ligand, specifically including:

[0059] After the molecular docking is completed, three kinds of protein-ligand interaction features in the MDCC docking method are extracted: the first kind is the energy auxiliary term separated from Glide SP, including docking score, GlideScore (g score ), lipophilic term (lipo), hydrogen bond term (h bond ), metal binding term (metal), reward or penalty term (rewards), van der Waals energy (e vdw ), electrostatic energy (e coul ), rotatable bond penalty term (e rotb ), active site polarity interaction term (e site ), model score (e model ), and modified electrostatic-van der Waals interaction energy (energy), the second kind is the energy term based on the decomposition of amino acid residues separated from Glide SP, and the third kind is the entropy effect of protein-ligand interaction.

[0060] The related features of the ligand include the partition coefficient (logP) of the ligand, the number of hydrogen bond acceptors (HBA), the number of hydrogen bond donors (HBD), the number of rotatable bonds (RB), and the total partial charge (FC).

[0061] The specific implementation method for extracting the entropy effect characteristics of protein-ligand interaction is: using the density functional theory B3LYP functional in quantum chemistry method, combining the 6-311G(d, p) basis set, and through the D3 dispersion correction of Grimme (combined with the Becke-Johnson correction) to perform structure optimization and frequency analysis on the small molecule ligand, and finally obtain the total entropy effect of the ligand. Under the assumption that the ligand is completely captured by the protein, the entropy effect of the protein-ligand interaction is approximately equal to the total entropy effect of the ligand in the free state.

[0062] The protein-ligand interaction feature and ligand feature processing method can be selected from one or more of the following: standardization, normalization, logarithmic transformation, discretization, one-hot encoding, label encoding, target encoding, frequency encoding, recursive feature elimination, polynomial feature generation, cross-feature generation, principal component analysis, linear discriminant analysis, singular value decomposition, and other processing methods suitable for specific application scenarios.

[0063] As an implementation, the feature processing method is performed in the following four steps in turn as shown in the following four steps: Figure 2

[0064] Remove features with variance less than a set threshold such as 0.05, removing features that change little in the data set and contribute less to the model;

[0065] Calculate the mutual information between each feature and the target variable activity, and only keep the features with mutual information greater than 0;

[0066] Calculate the Pearson correlation coefficient matrix between the remaining features, identify highly correlated feature pairs with a correlation coefficient greater than a set coefficient threshold such as 0.95, and optionally remove highly correlated features, keeping only one of them;

[0067] Through recursive feature elimination and cross-validation, evaluate the importance of each feature, remove the least important features step by step, and select the optimal 39 features.

[0068] After completing feature extraction and processing, step 104 is performed.

[0069] Step 104: Use four machine learning models to build a scoring function targeting HBV Cp.

[0070] In a specific implementation, the present embodiment uses four common machine learning models: support vector machines, random forests, extreme gradient boosting, and artificial neural networks. These models are mainly used for classification tasks and can efficiently process complex data in high-dimensional space, accurately capture complex interactions between molecules and proteins, and provide highly accurate and well-generalized scoring functions.

[0071] ​It should be noted that the four machine learning models of support vector machine, random forest, extreme gradient boosting and artificial neural network are used to construct the scoring function as an embodiment, but other types of machine learning models can also be used to achieve similar functions without departing from the scope of the present application.

[0072] Next, the working principles and advantages of the four models will be introduced in detail:

[0073] Support vector machine solves the classification problem by finding the best separating hyperplane in the feature space through a nonlinear kernel function to map the data to a high-dimensional feature space. The kernel function is a very important hyperparameter of SVM, mainly including linear kernel, polynomial kernel, sigmoid kernel, radial basis function (RBF) kernel, etc., among which RBF is the most commonly used and best performing method. SVM performs excellently in handling small sample data and high-dimensional features, so it is very suitable for constructing the scoring function targeting Cp.

[0074] Random forest is an ensemble learning algorithm that uses decision trees (DTs) as base learners and combines the ideas of random sampling and "bagging" to integrate the prediction results of multiple DTs to improve overall performance. The introduction of randomness effectively reduces the risk of overfitting of the model and significantly enhances its noise resistance. Random forest constructs multiple decision trees by randomly sampling the data set with replacement, and the difference between these trees makes the model have stronger generalization ability. Compared with a single decision tree, RF is more robust and can automatically evaluate the importance of features, so it is very suitable for processing data targeting Cp to help construct a reliable scoring function.

[0075] Extreme gradient boosting is an ensemble learning algorithm based on gradient boosting, which constructs a series of weak classifiers (usually decision trees), adjusts the weights step by step, and corrects the errors in the previous round of classification to continuously improve the accuracy of the model. The algorithm uses an incremental learning strategy, adjusting the model weights according to the errors of the previous round each time, so that it quickly converges to the optimal result, which is particularly suitable for processing complex and noisy data. XGBoost also has excellent computing efficiency, can efficiently process large-scale data, and can avoid overfitting through feature parallelism, memory optimization and regularization.

[0076] An artificial neural network (ANN) is a mathematical model that applies structures similar to the connections of brain synapses for information processing, consisting of input layer, hidden layer and output layer, and the neurons are connected through weights. The input data is propagated forward in the network, layer by layer, and processed through a nonlinear activation function, and finally a prediction result is generated in the output layer. ANN calculates the output error through the backpropagation algorithm and adjusts the weight of each connection to continuously optimize the model performance. The multi-layer network structure and nonlinear activation function of ANN enable it to effectively capture complex nonlinear relationships in data and extract multi-level feature representations, with strong fitting and generalization capabilities. Therefore, ANN performs particularly well in handling complex tasks such as high-dimensional, nonlinear features.

[0077] In this embodiment, support vector machines, random forests and artificial neural networks are implemented using the SVC class, RandomForestClassifier class and MLPClassifier class in the open source Python library scikit-learn, respectively, while extreme gradient boosting is implemented based on the XGBClassifier class in the open source Python library XGBoost.

[0078] After completing the construction of the scoring function based on the machine learning model targeting Cp, step 105 is performed.

[0079] Step 105: Based on the scoring function model, model training and evaluation are performed.

[0080] The target HBV Cp dataset obtained in step 101 is split into a training set and a test set in a ratio of 4:1 based on the skeleton structure, with the training set used for model training and cross-validation and the test set used for model evaluation.

[0081] In the model training phase, hyperparameter optimization and model generalization capability evaluation are simultaneously achieved through nested cross-validation, and the optimal hyperparameter combination of each machine learning model is finally determined. As shown in Table 2, the optimal hyperparameter combinations of the four MLSF models are listed.

[0082] Table 2: Hyperparameter definitions and optimal configurations of four MLSF models

[0083]

[0084] After completing model training and evaluation, step 106 is performed.

[0085] Step 106: Using appropriate model integration methods such as voting, averaging, stacking, blending, boosting, etc., the trained multiple models are integrated into an integrated MLSF model with excellent prediction performance and virtual screening capability. The specific implementation steps are as follows:

[0086] In the present application, the weighted voting method is used for model fusion, and each model is assigned a weight according to its performance. In the MLSF model, the PR curve reflects the relationship between the true positive rate and the recall rate, and the area under the curve, i.e. AUPR, is considered to better reflect the enrichment ability of the MLSF model for positive samples. By comparing the AUPRs of the four machine learning scoring function models, it is found that the XGBoost model has the best enrichment ability for positive samples, so the XGBoost model is assigned the highest weight, i.e. the first weight such as 0.4, and the remaining three models (such as support vector machine, random forest and artificial neural network) are assigned a second weight such as 0.2; of course, the remaining three models can also be assigned different weights according to the AUPR.

[0087] For each small molecule compound y, the prediction results of all models are the activity probabilities. The activity probabilities predicted by each model are integrated by the weighted voting method, and the specific formula is as follows:

[0088] P final (y)=ω XGBoost P XGBoost (y)+ω SVM P SVM (y)

[0089] +ω RF P RF (y)+ω ANN P ANN (y)

[0090] where P final (y) is the weighted prediction probability of the integrated model for the activity of each small molecule compound, ω XGBoost , ω SVM , ω RF , ω ANN are the weights assigned to XGBoost, SVM, RF, and ANN, respectively, P XGBoost (y), P SVM (y), P RF (y), and P ANN (y) are the prediction probabilities of XGBoost, SVM, RF, and ANN for the activity of each small molecule compound, respectively. The model weights in this formula can be adjusted and customized according to different needs for model performance in specific scenarios.

[0091] The weighted prediction activity probability P final (y) is compared with the customized prediction activity probability threshold 0.5. If P final (y)>0.5, the compound is determined to be an active compound, and P final (y)<0.5, the compound is determined to be an inactive compound.

[0092] Finally, the Ensemble MLSE model targeting Cp was obtained, named as MLSE(Ensemble).

[0093] As shown in Table 3, the prediction performance of the four MLSE models targeting Cp and the MLSE(Ensemble) model in the test set was compared and analyzed. The results showed that the key indicators such as Recall, F1-score, Accuracy and AUROC of the MLSE(Ensemble) model in the test set were better than those of the single four MLSE models, which showed its excellent prediction performance.

[0094] Table 3 Prediction performance of machine learning scoring functions in the test set

[0095]

[0096]

[0097] Subsequently, the virtual screening ability of the single four MLSE models targeting Cp, the MLSE(Ensemble) model and the classic scoring function GlideScore-SP was analyzed and compared, and the virtual screening ability index of the scoring function was shown in Table 4.

[0098] Table 4 Evaluation index of virtual screening ability of scoring function

[0099]

[0100] BEDROC and EF α The calculation formula is as follows:

[0101]

[0102] Wherein, n, N are the number of positive samples and the total number of samples, R α is the proportion of positive samples in the total samples (i.e. n / N), r i is the ranking value of the i-th positive sample in all samples, and a is the weight given to the molecules with high ranking, when a is set to 20.0, 80.5 and 321.9, respectively, 8%, 2% and 0.5% of the molecules with high ranking account for 80% of the total score, and in this experiment, the parameter a is set to 20.0.

[0103]

[0104] Wherein, NTB α represents the number of active molecules in the top (e.g. a is 1%, 5% or 10%) of all molecules, NTB total represents the number of all active molecules.

[0105] As shown in Table 5, the virtual screening ability of the MLSF model and the classic scoring function GlideScore-SP on the test set is compared. The results show that the integrated MLSF model, i.e. MLSF (Ensemble), has the best virtual screening ability, which is significantly better than the single MLSF model and the classic scoring function GlideScore-SP. MLSF (Ensemble) helps to improve the virtual screening efficiency. At the same time, MLSF (Ensemble) performs significantly better in key performance indicators (such as accuracy, recall rate and AUC value), which shows that it not only improves the prediction performance, but also significantly improves the virtual screening efficiency while maintaining the screening accuracy.

[0106] Table 5 Comparison of virtual screening ability of MLSF and classic scoring function GlideScore-SP targeting Cp on test set

[0107]

[0108] Table 5 (continued)

[0109]

[0110]

[0111] In summary, the MLSF (Ensemble) model targeting Cp constructed by the model integration method in this embodiment has significantly better prediction performance and virtual screening ability on the test set than the single MLSF model. At the same time, all MLSF models perform better than the classic scoring function GlideScore-SP. The above results show that MLSF (Ensemble) is an efficient virtual screening method and has application potential for discovering potential Cp inhibitors.

[0112] Example Two

[0113] Referring to Figure 3 , a structure schematic diagram of a target-specific virtual screening system based on machine learning provided by an embodiment of the present application is shown, which comprises:

[0114] A training data acquisition module configured to acquire training data, the training data comprising: multi-conformational structures of a target and active molecule and inactive molecule data of the target;

[0115] A docking module configured to perform molecular docking of the active molecule and the inactive molecule of the target with the multi-conformational structures of the target;

[0116] A feature extraction module configured to extract protein-ligand interaction features and related features of ligands and perform feature engineering after the molecular docking is completed;

[0117] a machine learning scoring function model construction module configured to train the features after feature engineering on the plurality of machine learning models to obtain a plurality of trained target-specific scoring function models;

[0118] a machine learning scoring function model integration module configured to integrate the plurality of target-specific scoring function models to obtain an integrated model;

[0119] a virtual screening module configured to perform molecular docking of a compound in a compound library to be screened to the multi-conformation structure of the target, extract corresponding protein-ligand interaction features and ligand-related features, and obtain a prediction result based on the extracted features through the integrated model, and screen out candidate compounds with high activity potential according to the prediction result.

[0120] In more embodiments, there are also provided:

[0121] An electronic device includes a memory and a processor, and computer instructions stored in the memory and run on the processor, when the computer instructions are run by the processor, the method described in embodiment one is completed. For brevity, this will not be described here.

[0122] It should be understood that in the embodiments, the processor can be a central processing unit CPU, and the processor can also be other general-purpose processors, digital signal processors DSPs, application-specific integrated circuits ASICs, ready-to-program gate arrays FPGA or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0123] The memory can include read-only memory and random access memory, and provide instructions and data to the processor, and a part of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.

[0124] A computer readable storage medium for storing computer instructions, when the computer instructions are executed by a processor, the method described in embodiment one is completed.

[0125] The method in embodiment one can be directly embodied as a hardware processor to execute and complete, or be executed and completed by a combination of hardware and software modules in the processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory, and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0126] A computer program product comprising a computer program which, when executed by a processor, implements the method described in embodiment one.

[0127] The present application also provides at least one computer program product tangibly stored on a non-transitory computer readable storage medium. The computer program product includes computer executable instructions, for example, instructions embodied in program modules, executed by devices at the objects' real or virtual processors to perform the processes / methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types. In various embodiments, the functionality of program modules can be combined or split between program modules, as desired, and machine executable instructions for program modules can be executed within local or distributed devices. In distributed devices, program modules can be located in local and remote memory storage devices.

[0128] Computer program code for carrying out operations of the present application can be written in one or more programming languages. These computer program code can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, which executes via the processor of the computer or other programmable data processing apparatus, transforms the computer or other programmable data processing apparatus into a particular machine implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be embodied on a computer readable medium, which can be a non-transitory computer readable medium. The computer readable medium can include a computer data signal embodied in a carrier wave, a computer readable medium, or other medium, tangibly embodied in a computer data base or program modules, and / or computer program code.

[0129] In the context of the present application, the computer program code or related data can be carried by any suitable carrier, to enable the devices, apparatus or processors to perform the various processes and operations described above. Examples of carriers include signals, computer readable media, and the like. Examples of signals can include electrical, optical, radio, sound or other forms of propagated signals, such as carrier waves, infrared signals, and the like.

[0130] The present application can be adapted to the construction of data sets, feature extraction and processing, model selection, model training and evaluation, model integration steps, and these adaptations should also be included within the scope of the present application. Thus, the present application should not be limited to the specific embodiments in the detailed description, but should cover all variations and modifications made according to the basic principles of the present application.

[0131] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the present embodiment can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software manner depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0132] Although the specific embodiments of the present application are described above in combination with the drawings, it is not a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications or variations made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the scope of protection of the present application.

Claims

1. A machine learning based target specific virtual screening method, characterized in that, The method comprises the following steps: obtaining training data, wherein the training data comprises a plurality of conformations of a target and active molecule and inactive molecule data of the target; molecular docking is performed on the active molecule and the inactive molecule of the target and the plurality of conformations of the target, after the molecular docking is completed, protein-ligand interaction features and ligand-related features are extracted and feature engineering is performed; the features processed by the feature engineering are trained on a plurality of machine learning models to obtain a plurality of trained target-specific scoring function models, the plurality of target-specific scoring function models are integrated to obtain an integrated model; molecular docking is performed on a compound in a compound library to be screened and the plurality of conformations of the target, corresponding protein-ligand interaction features and ligand-related features are extracted, and a prediction result is obtained based on the extracted features by using the integrated model, and a candidate compound with high activity potential is screened according to the prediction result; the extracted protein-ligand interaction features and ligand-related features are subjected to feature engineering, specifically: features with a variance less than a set threshold are removed; mutual information between each feature and a target variable is calculated, and features with a mutual information less than or equal to 0 are removed; Pearson correlation coefficients between the remaining features are calculated, and features with a Pearson correlation coefficient greater than a set coefficient threshold are retained as one of the features; the importance of each feature is evaluated by recursive feature elimination and cross-validation to determine the finally retained features.

2. The machine learning based target specific virtual screening method as claimed in claim 1, wherein, The molecular docking method adopts a Glide SP method or a Glide XP method in Schrödinger.

3. The machine learning based target specific virtual screening method as claimed in claim 1, wherein, The protein-ligand interaction features comprise: energy auxiliary items separated from the Glide SP, including docking scores, GlideScores, lipophilic items, hydrogen bond items, metal binding items, reward or penalty items, van der Waals energy, electrostatic energy, rotatable bond penalty items, active site polarity interaction items, model scores and modified electrostatic-van der Waals interaction energy; energy items separated from the Glide SP and decomposed based on amino acid residues; entropy effects of protein-ligand interactions; the ligand-related features comprise: a partition coefficient of the ligand, a number of hydrogen bond acceptors, a number of hydrogen bond donors, a number of rotatable bonds and total partial charges.

4. The machine learning based target specific virtual screening method as claimed in claim 3, wherein, The entropy effect calculation method of the protein-ligand interaction is as follows: a density functional theory B3LYP functional in quantum chemistry is combined with a 6-311G(d,p) basis set, and a small molecule ligand is subjected to structure optimization and frequency analysis by means of a D3 dispersion correction of Grimme, and finally a total entropy effect of the ligand is obtained, and the entropy effect of the protein-ligand interaction is equal to the total entropy effect of the ligand in a free state under the assumption that the ligand is completely captured by the protein.

5. The machine learning based target specific virtual screening method as claimed in claim 1, wherein, The plurality of target-specific scoring function models are integrated by using a weighted voting method to obtain an integrated model, specifically: according to the size of the area under the PR curve of each target-specific scoring function model, a corresponding weight is given to the prediction result of each target-specific scoring function model. The prediction results of each target-specific scoring function model to which a corresponding weight is assigned are added to obtain a prediction result of the ensemble model.

6. A machine learning based target specific virtual screening system, characterized in that, The method comprises the following steps: a training data acquisition module configured to acquire training data, the training data comprising a multi-conformation structure of a target and active molecule and inactive molecule data of the target; a docking module configured to perform molecular docking of the active molecule and inactive molecule of the target and the multi-conformation structure of the target; a feature extraction module configured to, after the molecular docking is completed, extract protein-ligand interaction features and ligand-related features and perform feature engineering; the extracted protein-ligand interaction features and ligand-related features are subjected to feature engineering, specifically: features with a variance less than a set threshold are removed; mutual information between each feature and a target variable is calculated, and features with a mutual information less than 0 are removed; Pearson correlation coefficients between the remaining features are calculated, and features with a Pearson correlation coefficient greater than a set coefficient threshold are retained as one of the features; importance of each feature is evaluated through recursive feature elimination and cross-validation to determine the final retained features; a machine learning scoring function model construction module configured to train the features subjected to feature engineering on multiple machine learning models to obtain multiple trained target-specific scoring function models; a machine learning scoring function model ensemble module configured to ensemble the multiple target-specific scoring function models to obtain an ensemble model; a virtual screening module configured to perform molecular docking of a compound in a compound library to be screened and the multi-conformation structure of the target, extract corresponding protein-ligand interaction features and ligand-related features, and obtain a prediction result based on the extracted features through the ensemble model, and screen out candidate compounds with high activity potential according to the prediction result.

7. An electronic device, comprising: A computer program product comprising a memory and a processor and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, the method of any one of claims 1-5 is completed.

8. A computer-readable storage medium, characterized in that, A computer program product for storing computer instructions, when the computer instructions are executed by a processor, the method of any one of claims 1-5 is completed.

9. A computer program product, characterised in that, A computer program product for storing computer instructions, when the computer instructions are executed by a processor, the method of any one of claims 1-5 is completed.

Citation Information

Patent Citations

  • Construction method and application of senile breast cancer senescence scoring model

    CN117746983A

  • Methods and systems for machine-learning based molecule generation and scoring

    US20250014688A1