Target specificity virtual screening method and system based on machine learning

By introducing dynamic characteristics of multi-conformational targets and integrated methods of multiple machine learning models in virtual screening, a target-specific scoring function model is constructed, which solves the problem of low accuracy and efficiency of virtual screening prediction in the prior art, and achieves more efficient and accurate drug screening.

CN120072032AActive Publication Date: 2025-05-30SHANDONG UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510062182.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art is difficult to adequately capture the complex interactions between compounds and targets in virtual screening, resulting in low prediction accuracy and efficiency.

Method used

By introducing the dynamic characteristics of multi-conformational targets, a variety of machine learning models are used to construct target-specific scoring functions, and the advantages of each model are fused through model integration methods to generate an integrated scoring function model.

Benefits of technology

It significantly improves the ability and efficiency of virtual screening, can better reflect the complexity of real biological systems, improve adaptability and prediction capabilities, reduce experimental resource consumption, and shorten drug development cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072032A_ABST
    Figure CN120072032A_ABST
Patent Text Reader

Abstract

The invention provides a target specificity virtual screening method and system based on machine learning, and the method comprises the steps: carrying out the molecular docking of active molecules and inactive molecules with a plurality of conformations of a target, and extracting the protein-ligand interaction characteristics of the docking molecules and the related characteristics of ligands as the input characteristics of a machine learning model; a scoring function model is constructed by adopting a plurality of machine learning models, and an integrated target specificity scoring function model is finally obtained based on the advantages of a plurality of target specificity scoring function models in combination with a model integration method. The integrated model can more effectively process complex protein-ligand interaction data, has excellent prediction performance and virtual screening capability, provides a more reliable prediction result, and improves the virtual screening capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to bioinformatics analysis, and particularly relates to a target-specific virtual screening method and system based on machine learning. Background Art

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] Virtual Screening (VS) plays an important role in the discovery of lead compounds. It can screen out potential active molecules from a large number of compounds by means of technologies such as molecular docking. Among them, Scoring Functions (SF) are often used to evaluate the affinity of compounds binding to targets, and the compounds are ranked accordingly. Therefore, the accuracy of SF directly affects the efficiency and accuracy of virtual screening. Classical scoring functions can be described in a linear summation form to represent the relationship between the binding strength and the characteristics used to characterize the protein-ligand interaction. However, in fact, this linear relationship does not always exist. Therefore, it is difficult to fully consider the complex protein-ligand interactions when predicting the binding affinity of ligands and proteins, and when dealing with large-scale compound libraries, the performance of classical scoring functions may not be sufficient to meet the requirements.

[0004] To improve the efficiency and accuracy of virtual screening, Machine Learning-Based Scoring Function (MLSF) has been developed and applied. It does not require a predefined function form, but instead relies on machine learning algorithms to learn the function form from data, thus showing a more accurate prediction accuracy than classical scoring functions. According to different application scopes, MLSF can be divided into general MLSF and target-specific MLSF. General MLSF is usually trained on a wide range of datasets, attempting to cover various different molecules and situations, but it is insufficiently optimized for specific targets, so its applicability in structure-based virtual screening is limited. Target-specific MLSF is a class of methods designed specifically for specific targets, which can more precisely capture the interaction patterns between specific targets and ligands, and usually has better prediction performance than general MLSF.

[0005] In the process of constructing a target-specific scoring function based on machine learning for virtual screening, a single machine learning model is usually relied on. However, due to the different focuses of different models on feature selection, hypothesis space, and optimization objectives, the scoring function constructed by a single model may not be able to fully capture the complex interactions between compounds and targets, resulting in low prediction accuracy and efficiency in the virtual screening process. Summary of the Invention

[0006] To overcome the limitations of the prior art, the present invention comprehensively captures the diversity of protein-ligand interactions by introducing the dynamic characteristics of multi-conformational targets, constructs target-specific scoring functions using multiple machine learning models, and integrates the advantages of each model through model integration methods to construct a target-specific scoring function with high stability, excellent prediction performance, and outstanding virtual screening ability, which can better reflect the complexity of real biological systems, significantly improve adaptability and prediction ability, and at the same time improve the efficiency and accuracy of drug screening, reduce the consumption of experimental resources, and shorten the drug development cycle.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In the first aspect, the present invention provides a machine learning-based target-specific virtual screening method, including:

[0009] Obtaining training data, where the training data includes: the multi-conformational structure of the target and the data of active molecules and inactive molecules of the target;

[0010] Performing molecular docking on the active molecules and inactive molecules of the target with the multi-conformational structure of the target; after the molecular docking is completed, extracting the protein-ligand interaction features and the relevant features of the ligand and performing feature engineering;

[0011] Training the features processed by feature engineering on multiple machine learning models respectively to obtain multiple trained target-specific scoring function models, and integrating the multiple target-specific scoring function models to obtain an integrated model;

[0012] Performing molecular docking on the compounds in the compound library to be screened with the multi-conformational structure of the target, extracting the corresponding protein-ligand interaction features and ligand-related features, and obtaining a prediction result through the integrated model based on the extracted features, and screening out candidate compounds with high activity potential according to the prediction result.

[0013] In the second aspect, the present invention provides a machine learning-based target-specific virtual screening system, including:

[0014] A training data acquisition module, which is configured to: obtain training data, where the training data includes: the multi-conformational structure of the target and the data of active molecules and inactive molecules of the target;

[0015] A docking module, which is configured to: perform molecular docking on the active molecules and inactive molecules of the target with the multi-conformational structure of the target;

[0016] A feature extraction module, which is configured to: after the molecular docking is completed, extract the protein-ligand interaction features and the relevant features of the ligand and perform feature engineering;

[0017] A machine learning scoring function model construction module, which is configured to: train the features processed by feature engineering on multiple machine learning models respectively to obtain multiple trained target-specific scoring function models;

[0018] A machine learning scoring function model integration module, which is configured to: integrate multiple target-specific scoring function models to obtain an integrated model;

[0019] A virtual screening module, which is configured to: perform molecular docking on the compounds in the compound library to be screened with the multiple conformations of the target, extract the corresponding protein-ligand interaction features and ligand-related features, and obtain a prediction result through the integrated model based on the extracted features, and screen out candidate compounds with high activity potential according to the prediction result.

[0020] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0021] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.

[0022] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0023] The above one or more technical solutions have the following beneficial effects:

[0024] In the present invention, active molecules and inactive molecules are subjected to molecular docking with multiple conformations of the target, and protein-ligand interaction features and related molecular features of the ligand are extracted as input features of the machine learning model. Subsequently, multiple machine learning models are used to construct target-specific scoring function models, and the advantages of different models are fused through a model integration method to generate an integrated scoring function model. This integrated model can more efficiently process complex protein-ligand interaction data, provide more stable and reliable prediction results, and thus significantly improve the ability and efficiency of virtual screening.

[0025] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0026] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0027] Figure 1 It is a flowchart of the steps for constructing a scoring function targeting hepatitis B virus capsid protein based on a machine learning model in the first embodiment of the present invention;

[0028] Figure 2 It is a flowchart of the steps for feature engineering provided in the first embodiment of the present invention;

[0029] Figure 3 It is a schematic structural diagram of a machine learning-based target-specific virtual screening system provided in the second embodiment of the present invention. Detailed Embodiments

[0030] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0031] It should be noted that the following description is intended to further explain the specific embodiments of the present invention. All content is illustrative and does not limit the protection scope of the present invention. Unless otherwise clearly specified in the context, all technical and scientific terms used in this specification should be interpreted according to the conventional understanding of those of ordinary skill in the art.

[0032] The various embodiments of the present invention and the features therein can be freely combined without conflict, thereby providing more technical variants and application solutions. Although the embodiments listed in the present invention include specific technical details, it should be understood that these technical details should not limit the protection scope of the present invention. On the contrary, the protection scope of the present invention should be determined by the claims and should cover all equivalent structures and modifications within the spirit and scope of the present invention.

[0033] For example, although in the specific embodiments of the present invention, four machine learning models, such as support vector machine, random forest, extreme gradient boosting, and artificial neural network, are used to construct the scoring function and the voting method is used for integration, other types of machine learning models and model integration methods can also be used to achieve similar functions without departing from the technical scope of the present invention.

[0034] In addition, during the implementation process, adaptive adjustments may be made to the construction of the dataset, feature extraction and processing, model training and evaluation steps, and these adjustments should also be included in the protection scope of the present invention. Therefore, the present invention should not be limited to the specific embodiments in the detailed description, but should cover all variations and improvements made based on the basic principles of the present invention.

[0035] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0037] Example 1

[0038] This example is illustrated by taking the Hepatitis B Virus (HBV) capsid protein (Cp) as a target. HBV Cp plays a crucial role in all stages of the viral life cycle, especially in maintaining and replicating the viral genomic cccDNA. The diversity of Cp functions means that Cp inhibitors targeting Cp have the potential to inhibit multiple links in the viral replication cycle and inhibit the amplification of cccDNA in cells. Therefore, HBV Cp has become an important target for antiviral drug development. By targeting HBV Cp, the viral life cycle can be effectively disrupted, virus replication and transmission can be blocked, and there is the potential to cure HBV. It has been found that the target-specific scoring function constructed based on machine learning and high-precision molecular docking methods not only has stronger classification accuracy for active and inactive molecules, but also significantly outperforms classical scoring functions in terms of the enrichment ability of active compounds.

[0039] Refer to Figure 1 , which shows the step flow chart of the machine learning-based target-specific virtual screening method provided by the embodiments of the present application, specifically including the following steps:

[0040] Step 101: Construct a data set targeting HBV Cp, which includes the structural data, activity data, and multi-conformation structural data of Cp of active and inactive molecules. Among them, the activity data includes IC 50 , EC 50 , K d , K i .

[0041] In the specific implementation of this example, the steps of constructing a data set targeting HBV Cp include:

[0042] The structural data and activity data of active molecules and inactive molecules are preferably selected from data sources such as literature, PubChem, ChEMBL, BindingDB, DUD-E, and in-house laboratory experimental data, but are not limited to these sources; the multi-conformational structural data of the target is preferably selected from the experimentally resolved target structure database or the target structure predicted based on a theoretical model, but is not limited to the multi-conformational structures of the target from the above sources.

[0043] As an implementation, the structural data and activity data of active molecules and inactive molecules are sourced from the three major data sources of PubChem, ChEMBL, and BindingDB, and the multi-conformational structural data of HBV Cp is sourced from the RCSB PDB data source and the Cp structure predicted by AlphaFold2. Table 1 shows the data composition of the dataset targeting HBV Cp.

[0044] Table 1 Data Composition of the Dataset Targeting HBV Cp

[0045]

[0046] Continued Table 1

[0047]

[0048]

[0049] Based on the structural data and activity data of active molecules and inactive molecules, compound preprocessing and screening are carried out: a series of preprocessing is performed on the compound structure to obtain a standardized structure, which is saved as an SDF file, where the preprocessing includes removing organometallic compounds, protonation, and hydrogenation, etc. At the same time, for duplicate compounds with a large difference in inhibitory effect (EC 50 / IC 50 / K d / K i differing by an order of magnitude), they are removed, while for duplicate compounds with a difference within an order of magnitude, their average value is taken to ensure the accuracy of the data.

[0050] Based on the screened compounds, the data is sorted according to the following rules: compounds with EC 50 、IC 50 、K d 、K i values below 10 μM are classified as active molecules of Cp, and compounds with EC 50 、IC 50 、K d 、K i values above 50 μM are classified as inactive molecules. At the same time, for EC 50 、IC 50 、K d, K i Active compounds with uncertain activity values between 10 μM and 50 μM are discarded.

[0051] Based on the sorted compounds, structural clustering is performed. Among similar compound structures, the structure with the best activity value is retained to ensure the uniqueness of each compound structure. After clustering, a total of 212 active compounds and 94 inactive compounds with unique structures are collected. Subsequently, in the DUD-E database, 4257 decoy molecules, that is, inactive molecules, are generated based on the active compounds. Finally, 212 active compounds and 4351 inactive compounds are obtained. The ratio of active molecules to inactive molecules is approximately 1:20, completing the construction of the compound dataset. The active molecule label is set to 1, and the inactive molecule label is set to 0 for facilitating the construction of subsequent machine learning models.

[0052] In the RCSB PDB database, the structural files of multiple conformations of HBV Cp are obtained, and their PDB numbers are 5WRE and 6J10 respectively; based on the amino acid sequence of HBV Cp, the Cp structure is predicted by AlphaFold2. Subsequently, the above proteins are preprocessed, including adding missing atoms and residues, automatically determining the protonation state of residues according to physiological pH conditions, and adding hydrogen atoms, etc., successfully constructing the multi-conformation structure data of HBV Cp.

[0053] After constructing the dataset targeting HBV Cp, step 102 is executed.

[0054] Step 102: Docking the collected active and inactive molecules with multiple conformations of HBV Cp, where the target conformations include the crystal structures resolved by experiments and the structures predicted by theoretical models.

[0055] This embodiment uses a molecular docking method based on conformer-dependent charges (a Molecular Docking Method with Conformer-dependent Quantum Mechanics Charges, MDCC) described in Patent CN115206441A. This method is based on the Glide SP method in the traditional molecular docking method, and based on different conformations of small molecule ligands, charge fitting is performed through quantum chemical calculation methods. Multiple sets of RESP charge parameters are obtained for each drug molecule according to its different conformations, fully considering the influence of conformational changes on charge distribution.

[0056] Among them, the specific implementation steps of the MDCC docking method are as follows: First, use quantum chemical calculation software (Gaussian16) to optimize the structures of the collected active and inactive small molecule ligands, and then use The Conformational Search module in performs conformational search to obtain the top 10 conformations with the lowest energy of the small molecule ligand; subsequently, the top 10 conformations with the lowest energy of the small molecule ligand are subjected to structure optimization and frequency analysis using Gaussian16, and the Restrained Electro Static Potential (RESP) charges of each conformation are calculated using Multiwfn 3.8 software and assigned to the 10 conformations of the optimized small molecule ligand, thereby obtaining the charge distributions of the small molecule ligand in different conformations; finally, based on the multi-conformation-dependent charges of the small molecule ligand, using

[0057] After completing the molecular docking, step 103 is executed.

[0058] Step 103: Extract and process the protein-ligand interaction features and the relevant features of the ligand, specifically including:

[0059] After the molecular docking is completed, three protein-ligand interaction features in the MDCC docking method are extracted: the first is the energy auxiliary term separated from Glide SP, including the docking score, GlideScore (g score ), lipophilic term (lipo), hydrogen bond term (h bond ), metal binding term (metal), reward or penalty term (rewards), van der Waals energy (e vdw ), electrostatic energy (e coul ), rotatable bond penalty term (e rotb ), active site polar interaction term (e site ), model score (e model ), and modified electrostatic - van der Waals interaction energy (energy), the second is the energy term decomposed based on amino acid residues separated from Glide SP, and the third is the entropy effect of the protein-ligand interaction.

[0060] The relevant features of the ligand include the partition coefficient of the ligand (logP), the number of hydrogen bond acceptors (HBA), the number of hydrogen bond donors (HBD), the number of rotatable bonds (RB), and the total partial charge (FC).

[0061] The specific implementation method for extracting the entropy effect characteristics of protein-ligand interactions is as follows: The density functional theory B3LYP functional in quantum chemistry methods is used, combined with the 6-311G(d,p) basis set, and the small molecule ligand is structurally optimized and frequency analyzed through Grimme's D3 dispersion correction (combined with Becke-Johnson correction). Finally, the total entropy effect of the ligand is obtained. Assuming that the ligand is completely captured by the protein, the entropy effect of protein-ligand interaction is approximately equal to the total entropy effect of the ligand in the free state.

[0062] The methods for processing protein-ligand interaction characteristics and ligand characteristics can be selected from one or more of the following: standardization, normalization, logarithmic transformation, discretization, one-hot encoding, label encoding, target encoding, frequency encoding, recursive feature elimination, polynomial feature generation, cross-feature generation, principal component analysis, linear discriminant analysis, singular value decomposition, and other processing methods suitable for specific application scenarios.

[0063] As an implementation method, the feature processing method is carried out in the following four steps in sequence as Figure 2 shown:

[0064] Remove features with variances less than a set threshold such as 0.05, and remove features that vary little in the dataset and contribute little to the model.

[0065] Calculate the mutual information between each feature and the target variable activity, and only retain features with mutual information greater than 0.

[0066] Calculate the Pearson correlation coefficient matrix between the remaining features, identify feature pairs with high correlation, that is, correlation coefficients greater than a set coefficient threshold such as 0.95, and optionally remove highly correlated features, only retaining one of them.

[0067] Evaluate the importance of each feature through recursive feature elimination and cross-validation, gradually remove the least important features, and select the optimal 39 features.

[0068] After completing feature extraction and processing, execute step 104.

[0069] Step 104: Use four machine learning models to construct a scoring function targeting HBV Cp.

[0070] In specific implementation, this embodiment uses four common machine learning models: support vector machine, random forest, extreme gradient boosting, and artificial neural network. These models are mainly used for classification tasks, can efficiently process complex data in high-dimensional spaces, accurately capture the complex interactions between molecules and proteins, and provide highly accurate and well-generalized scoring functions.

[0071] It should be noted that, as an implementation, four machine learning models, namely support vector machine, random forest, extreme gradient boosting, and artificial neural network, are used to construct a scoring function. However, other types of machine learning models can also be used to achieve similar functions without departing from the scope of the present invention.

[0072] Next, the working principles and advantages of these four models will be introduced in detail:

[0073] The support vector machine maps data into a high-dimensional feature space through a non-linear kernel function and solves the linear SVM by finding the optimal separation hyperplane in the feature space to handle classification problems. The kernel function is a very important hyperparameter of the SVM, mainly including linear kernel, polynomial kernel, sigmoid kernel, radial basis function (RBF) kernel, etc. Among them, RBF is the most commonly used and has the best prediction performance. The SVM performs excellently in dealing with small sample data and high-dimensional features, so it is very suitable for constructing a scoring function for targeted Cp.

[0074] Random forest is an ensemble learning algorithm that uses decision trees (DTs) as the base learners and combines the ideas of random sampling and "bagging" to integrate the prediction results of multiple DTs to improve the overall performance. The introduction of randomness effectively reduces the risk of overfitting of the model and significantly enhances its anti-noise ability. Random forest constructs multiple decision trees by randomly sampling the data set with replacement. The differences between these trees enable the model to have stronger generalization ability. Compared with a single decision tree, RF is more robust and can automatically evaluate the importance of features. Therefore, it is very suitable for processing targeted Cp data and helping to construct a reliable scoring function.

[0075] Extreme gradient boosting is an ensemble learning algorithm based on gradient boosting. By constructing a series of weak classifiers (usually decision trees), it gradually adjusts the weights to correct the errors in the previous round of classification, thereby continuously improving the accuracy of the model. This algorithm adopts an incremental learning strategy, adjusting the model weights according to the errors in the previous round each time, so that it can quickly converge to the optimal result, especially suitable for dealing with complex and noisy data. XGBoost also has excellent computational efficiency, can efficiently process large-scale data, and avoids overfitting through feature parallelism, memory optimization, and regularization.

[0076] An artificial neural network is a mathematical model that processes information using a structure similar to the synaptic connections in the brain. It consists of an input layer, hidden layers, and an output layer, with neurons connected by weights. Input data undergoes forward propagation through the network, being passed layer by layer and processed by non-linear activation functions, and finally generating prediction results at the output layer. The ANN calculates the output error through the backpropagation algorithm and adjusts the weights of each connection to continuously optimize the model performance. Its multi-layer network structure and non-linear activation functions enable the ANN to effectively capture complex non-linear relationships in the data and extract multi-level feature representations, having strong fitting ability and generalization ability. Therefore, the ANN performs particularly well in handling complex tasks, such as data with significant high-dimensional and non-linear features.

[0077] In this embodiment, the support vector machine, random forest, and artificial neural network are respectively implemented using the SVC class, RandomForestClassifier class, and MLPClassifier class in the open-source Python library scikit-learn, while the extreme gradient boosting is implemented based on the XGBClassifier class in the open-source Python library XGBoost.

[0078] After constructing the scoring function for the targeted Cp based on the machine learning model, step 105 is executed.

[0079] Step 105: Based on the scoring function model, perform model training and evaluation.

[0080] The targeted HBV Cp dataset obtained in step 101 is split into a training set and a test set according to a ratio of 4:1 and based on the backbone structure. The training set is used for model training and cross-validation, and the test set is used for model evaluation.

[0081] In the model training stage, hyperparameter tuning and model generalization ability evaluation are simultaneously achieved through nested cross-validation, and finally the optimal hyperparameter combination for each machine learning model is determined. As shown in Table 2, the optimal hyperparameter combinations of 4 MLSF models are listed.

[0082] Table 2 Hyperparameter definitions and optimal configurations of four MLSF models

[0083]

[0084] After completing model training and evaluation, step 106 is executed.

[0085] Step 106: Adopt appropriate model integration methods, such as voting method, mean method, stacking method, hybrid method, boosting method, etc., to fuse multiple trained models into an integrated MLSF model with excellent prediction performance and virtual screening ability. The specific implementation steps are as follows:

[0086] In the example of this application, the weighted voting method is used for model fusion, and weights are assigned to each model according to its performance. In the MLSF model, the PR curve reflects the relationship between the true positive rate and the recall rate, and the area under the curve, i.e., AUPR, is considered to better reflect the enrichment ability of the MLSF model for positive samples. By comparing the AUPR of the four machine learning scoring function models, it is found that the XGBoost model has the best enrichment ability for positive samples. Therefore, the XGBoost model is given the highest weight, i.e., the first weight, such as 0.4, and the remaining three models (such as support vector machine, random forest, and artificial neural network) are all given the second weight, such as 0.2; of course, different weights can also be assigned to the remaining three models according to AUPR.

[0087] For each small molecule compound y, the prediction results of all models are their activity probabilities. The activity probabilities predicted by each model are integrated by the weighted voting method, and the specific formula is as follows:

[0088] P final (y) = ω XGBoost P XGBoost (y) + ω SVM P SVM (y)

[0089] + ω RF P RF (y) + ω ANN P ANN (y)

[0090] Among them, P final (y) is the weighted prediction probability of the integrated model for the activity of each small molecule compound, ω XGBoost , ω SVM , ω RF , ω ANN are the weights assigned to XGBoost, SVM, RF, and ANN respectively, and P XGBoost (y), P SVM (y), P RF (y), P ANN (y) are the prediction probabilities of XGBoost, SVM, RF, and ANN for the activity of each small molecule compound respectively. The model weights in this formula can be adjusted and customized according to different requirements for model performance in specific scenarios;

[0091] The weighted predicted activity probability P final (y) is compared with the custom predicted activity probability threshold of 0.5. If P final (y) > 0.5, the compound is determined to be an active compound; if P final (y) < 0.5, the compound is determined to be an inactive compound;

[0092] Finally, a targeted Cp integrated MLSF model was obtained, named MLSF (Ensemble).

[0093] As shown in Table 3, the prediction performance of the four Cp-targeted MLSF models and the MLSF (Ensemble) model in the test set was compared and analyzed. The results show that the key indicators such as Recall, F1-score, Accuracy and AUROC of the MLSF (Ensemble) model on the test set are better than those of the four single MLSF models, demonstrating its excellent prediction performance.

[0094] Table 3 Prediction performance of machine learning scoring functions on the test set

[0095]

[0096]

[0097] Subsequently, the virtual screening capabilities of the single MLSF, MLSF (Ensemble) and the classic scoring function GlideScore-SP for the four targeted Cp were analyzed and compared, as shown in Table 4, which shows the virtual screening capability indicators of the scoring function.

[0098] Table 4 Evaluation indicators of virtual screening ability of scoring function

[0099]

[0100] BEDROC and EF α The calculation formula is as follows:

[0101]

[0102] Among them, n and N are the number of positive samples and the total number of samples respectively, and R α is the proportion of positive samples in the total samples (i.e. n / N), r i is the ranking value of the ith positive sample among all samples, α is the weight given to the top-ranked molecules. When α is set to 20.0, 80.5, and 321.9, it corresponds to 8%, 2%, and 0.5% of the top-ranked molecules accounting for 80% of the total score, respectively. In this experiment, the parameter α is set to 20.0.

[0103]

[0104] Among them, NTB α represents the number of active molecules among all molecules ranked at the top (e.g., α is 1%, 5%, or 10%), NTB total Represents the number of all active molecules.

[0105] As shown in Table 5, the virtual screening ability performance of the MLSF model and the classical scoring function GlideScore-SP on the test set was compared. The results show that the integrated MLSF model, namely MLSF(Ensemble), has the best virtual screening ability, significantly superior to the single MLSF model and the classical scoring function GlideScore-SP. MLSF(Ensemble) helps to improve the virtual screening efficiency. At the same time, MLSF(Ensemble) shows significant advantages in key performance indicators (such as accuracy, recall rate, and AUC value), indicating that it can not only improve the prediction performance but also significantly enhance the virtual screening efficiency while maintaining the screening accuracy.

[0106] Table 5 Comparison of the virtual screening abilities of MLSF targeting Cp and the classical scoring function GlideScore-SP on the test set

[0107]

[0108] Continued Table 5

[0109]

[0110]

[0111] In summary, the MLSF(Ensemble) model targeting Cp constructed by the model integration method in this embodiment has significantly better prediction performance and virtual screening ability on the test set than the single MLSF model. At the same time, the performance of all MLSF models is better than that of the classical scoring function GlideScore-SP. The above results indicate that MLSF(Ensemble) is an efficient virtual screening method and has the application potential to discover potential Cp inhibitors.

[0112] Example Two

[0113] Refer to Figure 3 , which shows a schematic structural diagram of a target-specific virtual screening system based on machine learning provided by an embodiment of the present application, including:

[0114] A training data acquisition module, which is configured to: acquire training data, where the training data includes: multiple conformational structures of a target and active molecule and inactive molecule data of the target;

[0115] A docking module, which is configured to: perform molecular docking on the active molecules and inactive molecules of the target with the multiple conformational structures of the target;

[0116] A feature extraction module, which is configured to: after the molecular docking is completed, extract protein-ligand interaction features and relevant features of the ligand and perform feature engineering;

[0117] A machine learning scoring function model construction module, which is configured to: train the features processed by feature engineering on multiple machine learning models respectively to obtain multiple trained target-specific scoring function models;

[0118] A machine learning scoring function model integration module, which is configured to: integrate multiple target-specific scoring function models to obtain an integrated model;

[0119] A virtual screening module, which is configured to: perform molecular docking on the compounds in the compound library to be screened and the multi-conformational structures of the target, extract the corresponding protein-ligand interaction features and ligand-related features, and obtain a prediction result through the integrated model based on the extracted features, and screen out candidate compounds with high activity potential according to the prediction result.

[0120] In more embodiments, there is also provided:

[0121] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0122] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0123] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0124] A computer-readable storage medium for storing computer instructions, which when executed by a processor, completes the method described in Embodiment 1.

[0125] The method in Embodiment 1 can be directly implemented by a hardware processor to complete, or implemented by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0126] A computer program product includes a computer program which, when executed by a processor, implements the method described in Embodiment 1.

[0127] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the process / method described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed, and the machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.

[0128] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0129] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0130] In the present invention, adaptive adjustments may be made to the construction of the data set, feature extraction and processing, model selection, model training and evaluation, and model integration steps, and these adjustments should also be included within the scope of protection of the present invention. Therefore, the present invention should not be limited to the specific embodiments in the detailed description, but should cover all variations and improvements made based on the basic principles of the present invention.

[0131] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0132] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A target-specific virtual screening method based on machine learning, characterized in that: include: Acquiring training data, wherein the training data includes: multi-conformational structures of the target and data of active molecules and inactive molecules of the target; Molecular docking is performed on the active molecules and inactive molecules of the target with the multi-conformation structure of the target. After the molecular docking is completed, the protein-ligand interaction characteristics and the related characteristics of the ligand are extracted and feature engineering is performed; Using the features processed by feature engineering to train multiple machine learning models respectively, to obtain multiple trained target-specific scoring function models, and integrating the multiple target-specific scoring function models to obtain an integrated model; The compounds in the compound library to be screened are molecularly docked with the multi-conformational structure of the target, the corresponding protein-ligand interaction characteristics and ligand-related characteristics are extracted, and the prediction results are obtained through the integrated model based on the extracted characteristics, and candidate compounds with high activity potential are screened out according to the prediction results.

2. A target-specific virtual screening method based on machine learning as claimed in claim 1, characterized in that: Molecular docking method using The Glide SP method or Glide XP method in .

3. A target-specific virtual screening method based on machine learning as claimed in claim 1, characterized in that: The protein-ligand interaction characteristics include: Energy assistance terms isolated from Glide SP, including docking score, GlideScore, lipophilic term, hydrogen bonding term, metal binding term, bonus or penalty term, van der Waals energy, electrostatic energy, rotatable bond penalty term, active site polar interaction term, model score, and modified electrostatic-van der Waals interaction energy; Energy terms based on amino acid residue decomposition obtained from Glide SP; Entropic effects on protein-ligand interactions; Relevant characteristics of the ligand include: the partition coefficient, the number of hydrogen bond acceptors, the number of hydrogen bond donors, the number of rotatable bonds, and the total partial charge of the ligand.

4. The protein-ligand interaction feature according to claim 3, characterized in that: The entropy effect calculation method of the protein-ligand interaction is: using the density functional theory B3LYP functional in the quantum chemistry method, combined with the 6-311G (d, p) basis set, and performing structural optimization and frequency analysis on the small molecule ligand through Grimme's D3 dispersion correction, and finally obtaining the total entropy effect of the ligand. Under the condition that the ligand is completely captured by the protein, the entropy effect of the protein-ligand interaction is equal to the total entropy effect of the ligand in the free state.

5. A target-specific virtual screening method based on machine learning as claimed in claim 1, characterized in that: Feature engineering is performed on the extracted protein-ligand interaction features and ligand related features, specifically: Remove features whose variance is less than the set threshold; Calculate the mutual information between each feature and the target variable, and remove the features whose mutual information is not greater than 0; Calculate the Pearson correlation coefficient between the remaining features, and retain one of the features in the feature pairs whose Pearson correlation coefficient is greater than the set coefficient threshold; Through recursive feature elimination and cross-validation, the importance of each feature is evaluated and the final retained features are determined.

6. A target-specific virtual screening method based on machine learning as claimed in claim 1, characterized in that: The weighted voting method is used to integrate multiple target-specific scoring function models to obtain an integrated model, which is as follows: According to the area under the PR curve of each target-specific scoring function model, a corresponding weight is assigned to the prediction result of each target-specific scoring function model; The prediction results of each target-specific scoring function model assigned with a corresponding weight are added together to obtain the prediction result of the integrated model.

7. A target-specific virtual screening system based on machine learning, characterized in that: include: A training data acquisition module is configured to: acquire training data, wherein the training data includes: multi-conformational structures of the target and data of active molecules and inactive molecules of the target; A docking module, which is configured to: perform molecular docking of the active molecules and the inactive molecules of the target with the multi-conformation structure of the target; A feature extraction module, which is configured to: extract protein-ligand interaction features and ligand-related features after molecular docking and perform feature engineering; A machine learning scoring function model building module is configured to: train multiple machine learning models with the features processed by feature engineering to obtain multiple trained target-specific scoring function models; A machine learning scoring function model integration module is configured to: integrate multiple target-specific scoring function models to obtain an integrated model; The virtual screening module is configured to: perform molecular docking on the compounds in the compound library to be screened and the multi-conformational structure of the target, extract the corresponding protein-ligand interaction characteristics and ligand-related characteristics, and obtain prediction results through the integrated model based on the extracted characteristics, and screen out candidate compounds with high activity potential according to the prediction results.

8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.

9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for detecting binding specificity between ligand and target and drug screening method

    CN102798708A

  • Drug target virtual screening method based on interactive fingerprints and machine learning

    CN106446607A

  • Method for improving virtual screening capability of docking software based on machine learning algorithm

    CN111402967A

  • Multi-target drug screening method based on ensemble learning and hybrid neural network

    CN113066525A

  • Construction method and application of senile breast cancer senescence scoring model

    CN117746983A