Drug screening and optimizing method and drug screening system for PDL1 inhibitor based on random forest
Through drug screening and optimization methods based on random forest algorithms, a high-precision QSAR model is constructed to predict the activity level of PDL1 inhibitors and perform structural optimization, which solves the problems of low response rate and low R&D efficiency of PDL1 inhibitors in the prior art, and achieves efficient and accurate drug screening and structural optimization.
Patent Information
- Application Number
- CN202510297313.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The response rate of existing PDL1 inhibitors is low and the drug resistance is frequent. Traditional drug screening methods are time-consuming and costly, and it is difficult to effectively correlate the complex nonlinear relationship between the structure and activity of the compound, resulting in inefficient drug development.
A high-precision quantitative structure-activity relationship (QSAR) model was constructed by a drug screening and optimization method based on a random forest algorithm, and the PDL1 inhibitory activity level of candidate compounds was predicted through the mapping relationship between compound structural characteristics and biological activity, and structural optimization was performed.
It significantly improves the accuracy and reliability of the prediction of PDL1 inhibitor activity, reduces the blindness of experimental verification, reduces the initial R&D costs, and provides structural interpretability and wide applicability.
Smart Images

Figure CN120220889A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of drug screening, and specifically relates to a method for screening and optimizing PDL1 inhibitors based on random forest. Background Art
[0002] Programmed death protein 1 (PD1) is an immune checkpoint receptor that regulates T cell responses. Its ligand, programmed death ligand 1 (PD-L1), is usually overexpressed on the surface of tumor cells. When PD1 binds to PD-L1, T cell responses are inhibited, leading to tumor immune resistance. Checkpoint inhibitors that block the formation of the PD1 / PD-L1 complex have attracted great interest in cancer immunotherapy. In recent years, immune checkpoint inhibitors (such as PD1 / PDL1 inhibitors) have made breakthrough progress in the field of tumor immunotherapy. By blocking the binding of PD1 and PD-L1, the anti-tumor immune response of T cells can be activated. However, existing PDL1 inhibitors have problems such as low response rate and frequent drug resistance, and there is an urgent need to develop new inhibitors with high activity and high selectivity. Traditional drug screening methods rely on high-throughput experimental screening and structural modification, but such methods are time-consuming, costly, and difficult to effectively correlate the complex non-linear relationship between compound structure and activity, seriously restricting the efficiency of drug research and development.
[0003] Although computational screening techniques based on quantitative structure-activity relationship (QSAR) can partially alleviate the above problems, traditional QSAR models (such as linear regression, support vector machines, etc.) have limited ability to process high-dimensional chemical features and are difficult to capture the complex interactions between molecular structural features and biological activities, resulting in insufficient prediction accuracy. In addition, existing models have significant shortcomings in feature selection, noise data robustness, and generalization ability, especially for the screening of inhibitors targeting PDL1, and an efficient and reliable prediction system has not been formed.
[0004] Due to its ensemble learning characteristics, anti-overfitting ability, and adaptability to high-dimensional data, the random forest algorithm shows potential in the field of drug discovery. However, systematic screening and optimization research on PDL1 inhibitors is still blank. The existing technology does not fully combine the feature importance evaluation function of the random forest to guide structural optimization, nor does it have customized model parameters and verification strategies for the PDL1 target. Therefore, developing a research method for PDL1 inhibitor based on the random forest algorithm, integrating data-driven screening and directed structural optimization, has important value for improving the discovery efficiency of candidate compounds and shortening the research and development cycle. Summary of the Invention
[0005] In view of the problems of low efficiency, high cost, and insufficient accuracy of the prediction model in the existing PDL1 inhibitor screening technology, the present invention proposes a drug screening and optimization method based on the random forest algorithm, aiming to construct a high-precision quantitative structure-activity relationship (QSAR) model to achieve efficient screening of high-activity PDL1 inhibitors and accelerate the development process of candidate compounds through structure optimization guidance.
[0006] A drug screening and optimization method for PDL1 inhibitors based on random forest of the present invention includes the following steps:
[0007] (1) Obtain a data set containing the structural information of known PDL1 inhibitor compounds and the corresponding bioactivity data, and randomly divide the training set and the test set.
[0008] (2) Use the random forest algorithm to learn the training set data and construct a quantitative structure-activity relationship prediction model. The quantitative structure-activity relationship prediction model predicts the PDL1 inhibitory activity level of candidate compounds through the mapping relationship between compound structural features and bioactivity.
[0009] (3) Input the test set data into the trained random forest model for prediction, and calculate the model performance indicators by comparing the prediction results with the measured data.
[0010] (4) Screen high-activity PDL1 inhibitors from the candidate compound library based on the established quantitative structure-activity relationship prediction model, and optimize the structure of the screening results.
[0011] Further, in step (1), the random division method uses the stratified random sampling method, and the division ratio of the training set to the test set is 3:1 to 4:1.
[0012] Further, in step (2), the structural features include but are not limited to: molecular descriptors, molecular fingerprints, and three-dimensional pharmacophore features, and feature selection is performed through feature importance evaluation.
[0013] Further, the parameter settings of the random forest algorithm are: the number of decision trees is 500 - 1000, the maximum tree depth is 10 - 20 layers, and the minimum number of samples per node is 1.
[0014] Further, in step (2), the inhibition activity level division standard is: compounds with an IC 50 value ≤ 1000 nM are defined as high-activity inhibitors, and compounds with an IC 50 value > 1000 nM are defined as low-activity inhibitors.
[0015] Further, in step (3), the model performance evaluation indicators include at least two combinations of indicators such as accuracy, recall rate, AUC value, and F1 score.
[0016] Furthermore, the structural optimization in step (4) includes: based on the feature importance ranking output by the model, performing directional modification on the molecular skeleton, substituent groups or functional groups of the candidate compounds.
[0017] The method of the present invention is applied to the research and development of anti-tumor drugs, and the research and development of the anti-tumor drugs is for the development of specific inhibitors targeting the PD1 / PDL1 immune checkpoint pathway.
[0018] A drug screening system includes:
[0019] A data preprocessing module for performing the dataset partitioning described in claim 1;
[0020] A model training module for constructing the QSAR model described in claim 1;
[0021] An activity prediction module for performing the biological activity prediction described in claim 1;
[0022] A result output module for generating a visual screening report and structural optimization suggestions.
[0023] The present invention has the following beneficial effects:
[0024] Efficient screening: By processing high-dimensional chemical features through the random forest algorithm, the accuracy and reliability of PDL1 inhibitor activity prediction are significantly improved (AUC value ≥ 0.90) (see Figure 2 ), reducing the blindness of experimental verification;
[0025] Cost optimization: Combining traditional experimental screening with computational prediction to reduce the initial R & D cost;
[0026] Structural interpretability: Using feature importance analysis to clarify key pharmacophore groups, providing quantitative guidance for subsequent structural optimization;
[0027] Wide applicability: The model supports multi-scenario screening of small molecule compound libraries and virtual compound libraries, and can be extended to the research and development of other immune checkpoint inhibitors. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is the tSNE distribution diagram after dimensionality reduction of the compound database of the present invention;
[0029] Figure 2 It is the ROC result curve diagram of the classification result after prediction for the test set of the present invention;
[0030] Figure 3 It is the fusion diagram of the classification result after prediction for the test set of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clearly understood, the spirit of the content disclosed by the present invention will be described in detail below. After any person skilled in the relevant technical field understands the embodiments of the content of the present invention, the techniques taught by the content of the present invention can be changed and modified without departing from the spirit and scope of the content of the present invention.
[0032] The exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.
[0033] Embodiment
[0034] Step 1: Data Preparation and Preprocessing
[0035] (1) Collect a compound dataset containing more than 8,000 known PDL1 inhibitors, covering IC50 values, molecular descriptors (such as logP, molecular weight, topological polar surface area), molecular fingerprints (ECFP6, MACCS), and three-dimensional pharmacophore features;
[0036] (2) Divide the dataset into a training set (70%) and a test set (30%) by stratified random sampling to ensure that the proportions of high-activity (IC50 ≤ 1000 nM) and low-activity (IC50 > 1000 nM) categories are consistent;
[0037] (3) Normalize the molecular features (Z-score normalization) and filter redundant features (correlation coefficient threshold of 0.9).
[0038] Step 2: Random Forest Model Construction and Parameter Configuration
[0039] Use the RandomForestClassifier in the Scikit-learn library to construct a model with the following parameter configurations:
[0040] Gini coefficient (Gini coefficient), used for node splitting
[0041] Current value = None, allowing the tree to grow until all leaves are pure
[0042] Current value = 2, the minimum number of samples for splitting to prevent overfitting
[0043] Current value = 1, the minimum number of samples in a leaf node
[0044] Current value = sqrt, each tree randomly selects sqrt(n_features) features
[0045] Current value = True, enabling sampling with replacement to improve model stability
[0046] Current value = None, assuming data class balance
[0047] Current value = 42. Fix the random seed to ensure reproducibility of the results.
[0048] Description of key parameter optimization:
[0049] n_estimators: After cross-validation testing, 100 decision trees can balance prediction accuracy (AUC ≥ 0.92) and training time (< 10 minutes per thousand samples);
[0050] max_depth = None: Allow the tree to grow completely to capture complex feature relationships, and cooperate with post-pruning strategies to avoid overfitting;
[0051] max_features ='sqrt': Each tree randomly selects √n features (n is the total number of features) to reduce the impact of collinearity between features;
[0052] bootstrap = True: Increase the diversity of base learners through bootstrap sampling to improve the generalization ability of the ensemble model.
[0053] Step 3: Model training and validation
[0054] (1) Input the training set into the model for training, and use 5-fold cross-validation to adjust hyperparameters;
[0055] (2) Use the test set to evaluate the model performance and calculate the following metrics:
[0056] Accuracy ≥ 85%
[0057] AUC value ≥ 0.90
[0058] Recall of high-activity categories ≥ 80% (see Figure 2 )
[0059] (3) Output the feature importance ranking (such as the number of hydrogen bond donors in the pharmacophore, the proportion of hydrophobic groups, etc.), and screen the top 20% of features with importance scores for subsequent optimization.
[0060] Step 4: Inhibitor screening and structure optimization
[0061] (1) Input the virtual compound library (≥ 10,000 molecules) into the trained model to predict the PDL1 inhibitory activity level;
[0062] (2) Screen the top 5% of compounds predicted to be highly active (about 500), and perform targeted modification in combination with the feature importance results:
[0063] (3) Perform secondary prediction on the optimized compounds, and retain the candidate molecules with AUC confidence ≥ 0.95 for in vitro experimental verification.
[0064] Step 5: Experimental Verification and Iterative Optimization
[0065] (1) Perform PDL1 / PD1 binding inhibition experiments (ELISA method) on the top 50 candidate compounds selected by the model;
[0066] (2) Feed the experimental data back to the training set and update the model parameters (such as adjusting class_weight to handle the new data imbalance problem);
[0067] (3) Repeat steps 2 - 4 for iterative optimization until high - activity lead compounds with IC50 ≤ 10 nM are obtained.
Claims
1. A drug screening and optimization method for PDL1 inhibitors based on random forests, characterized in that The following steps are involved: (1) obtaining a data set containing structural information of known PDL1 inhibitor compounds and corresponding biological activity data, and randomly dividing the data into a training set and a test set; (2) using a random forest algorithm to learn the training set data and construct a quantitative structure-activity relationship prediction model, wherein the quantitative structure-activity relationship prediction model predicts the PDL1 inhibitory activity level of the candidate compound through a mapping relationship between the compound structural features and the biological activity; (3) Input the test set data into the trained random forest model for prediction, and calculate the model performance index by comparing the prediction results with the measured data; (4) Based on the established quantitative structure-activity relationship prediction model, the candidate compound library was screened for highly active PDL1 inhibitors, and the screening results were structurally optimized.
2. The drug screening and optimization method for PDL1 inhibitors based on random forests according to claim 1, characterized in that The random division in step (1) adopts a stratified random sampling method, and the ratio of the training set to the test set is 3:1 to 4:
1.
3. The drug screening and optimization method of PDL1 inhibitor based on random forest according to claim 1, characterized in that The structural features in step (2) include: molecular descriptors, molecular fingerprints and three-dimensional pharmacophore features, and feature selection is performed through feature importance evaluation.
4. The drug screening and optimization method for PDL1 inhibitors based on random forests according to claim 1, characterized in that The parameters of the random forest algorithm are set as follows: the number of decision trees is 500-1000, the maximum tree depth is 10-20 layers, and the minimum number of node samples is 1.
5. The drug screening and optimization method for PDL1 inhibitors based on random forests according to claim 1, characterized in that The inhibitory activity classification standard in step (2) is: IC 50 Compounds with values ≤1000 nM were defined as highly active inhibitors, and IC 50 Compounds with values > 1000 nM were defined as low activity inhibitors.
6. The drug screening and optimization method for PDL1 inhibitors based on random forests according to claim 1, characterized in that The model performance evaluation indicators in step (3) include: a combination of at least two indicators among accuracy, recall, AUC value, and F1 score.
7. The drug screening and optimization method for PDL1 inhibitors based on random forests according to claim 1, characterized in that The structural optimization in step (4) includes: based on the feature importance ranking output by the model, targeted modification of the molecular skeleton, substituent group or functional group of the candidate compound.
8. The method as claimed in claim 1 is applied to the development of anti-tumor drugs, wherein the anti-tumor drug development is the development of specific inhibitors for the PD1 / PDL1 immune checkpoint pathway.
9. A drug screening system, characterized in that include: A data preprocessing module, used to perform the data set division as claimed in claim 1; A model training module, used to execute the QSAR model construction according to claim 1; An activity prediction module, used to perform the biological activity prediction according to claim 1; The result output module is used to generate visual screening reports and structural optimization suggestions.