Method for screening FXR regulator with anti-HBV activity based on machine learning
By constructing a machine learning-based screening method, integrating the random forest model and in vitro cell activity verification, the problem of lack of effective screening of FXR modulators in the existing technology is solved, and efficient and accurate screening effect is achieved, providing support for the application of FXR modulators in anti-HBV treatment.
Patent Information
- Application Number
- CN202510115291.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The lack of effective machine learning models in the prior art for screening FXR modulators for anti-HBV activity limits the application potential of FXR modulators in anti-HBV therapy.
A machine learning-based screening method was constructed to screen out potential anti-HBV activity FXR regulators by integrating random forest models and in vitro cell activity verification. The method includes obtaining FXR regulator data from the ChEMBL database, performing preprocessing and feature calculations, establishing a machine learning model for training and verification, and finally screening out compounds with potential anti-HBV activity.
The screening efficiency and accuracy of FXR modulators against HBV activity was improved, and multiple compounds with potential anti-HBV activity were successfully screened, providing strong support for the application of FXR modulators in anti-HBV therapy.
Smart Images

Figure CN120108563A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of drug discovery and relates to a method for screening FXR regulators with anti-HBV activity based on machine learning. Background Art
[0002] FXR is a member of the nuclear receptor superfamily NRs and is highly expressed in the gut-liver axis. Based on its role in regulating bile acid homeostasis, glucose and lipid metabolism, inflammatory response, etc., FXR has become a target of great concern in the development of drugs for liver diseases such as PBC, NAFLD, NASH, and liver fibrosis. In recent years, studies have shown that exogenous FXR modulators have a significant inhibitory effect on HBV replication, and currently there are two synthetic non-steroidal FXR modulators that are in clinical research for their anti-HBV effects. It can be seen that FXR modulators have great application potential in the treatment of liver-related diseases and anti-HBV. However, the current research on FXR modulators for anti-HBV is still limited, and there is no relevant machine learning model for screening FXR modulators with anti-HBV activity. With the continuous development of machine learning technology, machine learning models can be used to extract useful features from a large number of compounds and biological information to screen active compounds. The present invention constructs an efficient machine learning model that can accurately screen out potential FXR modulators with anti-HBV activity, providing strong support for the application research of FXR modulators in anti-HBV. Summary of the invention
[0003] The purpose of the present invention is to propose a method for screening FXR modulators with anti-HBV activity based on machine learning, aiming to assist drug developers in efficiently and accurately discovering potential FXR modulators with anti-HBV activity. The method of the present invention improves the prediction accuracy by integrating the random forest screening model and in vitro cell activity verification, thereby improving the efficiency and accuracy of drug screening.
[0004] To achieve the above object, the technical solution of the present invention is as follows: as a first aspect, a method for screening FXR modulators with anti-HBV activity based on machine learning is provided, comprising the following steps:
[0005] Step S1, obtain FXR modulator active and inactive compounds from the ChEMBL database, and collect FXR modulator binding ability data of these compounds, including: affinity values of ligand binding to FXR, half-activation concentration values EC generated after ligand binding to FXR 50 , half inhibitory concentration value IC 50 and SMILES string. Preferably, the collected data is preprocessed, including data cleaning, data deduplication, unified concentration units and label extraction. The label is set to whether it has FXR regulatory activity.
[0006] Step S2, cluster analysis is performed on database compounds using molecular fingerprints to construct a FXR modulator compound library.
[0007] Preferably, the data in the FXR modulator compound library are divided into a training set and a test set, and the proportion of compounds labeled as having FXR modulating activity and compounds without FXR modulating activity in the training set is balanced.
[0008] Step S3, feature calculation is performed on the preprocessed data, and the following 12 molecular fingerprint descriptors are selected as training features using PaDEL-Descriptor: AtomPairs2D-FingerprintCount, AtomPairs2D-Fingerprinter, Estate-Fingerprinter, Extended-Fingerprinter, Fingerprinter, GraphOnly-Fingerprinter, KlekotaRoth-FingerprintCount, KlekotaRothFingerprinter, MACCS-Fingerprinter, Pubchem-Fingerprinter, Substructure-FingerprintCount, SubstructureFingerprinter. The 12 molecular fingerprint descriptors are used individually and in combination.
[0009] Step S4, establishing a machine learning model for training to obtain a FXR regulatory activity prediction model.
[0010] Specifically, a variety of machine learning methods are used to establish a prediction model for the activity of FXR modulators, such as random forest (RF) and support vector machine (SVM) algorithms, and the comprehensive performance of the two algorithms is compared. A combination of one or more molecular fingerprint descriptors is used as a training feature to establish a prediction model for the activity of FXR modulators and compare the effects. The FXR regulatory activity prediction model obtained by calculating Accuracy, Precision, Recall, F1 value, and AUC evaluation is used to optimize the training features and machine learning methods.
[0011] Step S5, using the preferred training features and machine learning method (the present invention uses random forest) to train a FXR modulator activity prediction model, input the test compound (activity unknown) data, perform activity prediction, and screen multiple test compounds with top FXR modulator activity prediction results as candidate compounds.
[0012] Specifically, the test compound data comes from the FDA_HY-L022 Library compound library.
[0013] Step S6, as a preferred step, in order to screen out potential FXR modulators, further analysis is performed on compounds with higher activity predicted by the RF model. Three crystal structures of FXR and ligand complexes are selected, and the candidate compounds obtained in step S5 are scored using molecular docking technology. By calculating the binding energy scores of the compounds with the FXR receptor, the scoring order is arranged, and a plurality of top-ranked compounds with potential FXR modulating activity are screened out. In the present invention, 252 potential modulators are obtained.
[0014] Step S7, FXR modulator activity screening. In order to screen out potential FXR modulators, the activity of compounds in the 252 drug libraries (FDA_HY-L022 Library) was verified in HEK-293T-FXR-Luc cells. Specifically, a cell line containing FXR response elements and capable of stably expressing luciferase was constructed, and the screened compounds were screened at a concentration of 10 μM. Compounds with a screening regulation factor greater than 130% were considered agonists; compounds with a screening regulation factor less than 50% were considered antagonists or inverse agonists; as a control, the adjustment factor of DMSO added was 100%.
[0015] Step S8, anti-HBV activity phenotypic verification. In order to screen out the anti-HBV activity of potential FXR modulators, the 95 compounds screened in S7 were verified for anti-HBV activity in HepG2.2.15 cells. Specifically, the HepG2.2.15 cell ELISA method was used to determine the effects of FXR agonist compounds on the secretion of HBeAg and HBsAg antigens, and the MTT method was used to test cell viability to exclude the influence of cytotoxicity on the anti-HBV activity test.
[0016] Step S9, finally screening to obtain potential new FXR modulator compounds.
[0017] Furthermore, in step S1, data preprocessing includes:
[0018] Data cleaning: Delete the FXR modulator binding ability data with unavailable activity and the molecules with failed three-dimensional conformation;
[0019] Data deduplication: Duplicates are removed by ID, and unique IDs are retained; the different FXR regulator binding capacity values (EC 50 value), retain the lowest value;
[0020] Unified concentration units: The concentration units of database compounds are not unified. In order to facilitate subsequent screening, the data units are uniformly converted to nM (nmol / L).
[0021] Tag extraction: Using FXR modulator binding capacity values (EC 50 The activity threshold was set to 2000 nM, and molecules with values less than the activity threshold were marked as category "1" and regarded as active regulators of FXR. Otherwise, they were regarded as weak or inactive regulators and marked as category "0", which was the activity label.
[0022] Furthermore, the cluster analysis in step S2 uses the MACCS molecular fingerprint in RDKIT to calculate the Tanimoto coefficient to obtain molecular fingerprint clustering.
[0023] Specifically, the threshold was set to 80%, and the results showed that the database had 192 compound clusters.
[0024] Furthermore, the selection of the classification prediction model in step S4 is performed by calculating the accuracy, precision, recall, F1-score, and AUC, and by comprehensively comparing the scores of each indicator to optimize the training features and machine learning methods. The calculation formulas for each indicator are as follows:
[0025] (1) Accuracy: Accuracy is the most common performance evaluation metric for classification problems. It measures the proportion of samples correctly predicted by the model. However, in the case of an unbalanced class distribution, accuracy can be misleading, so other metrics need to be combined to evaluate model performance.
[0026]
[0027] (2) Precision and Recall: Precision and recall are important indicators for the problem of imbalanced class distribution. Precision measures the proportion of true positive examples among the positive examples predicted by the model, while recall measures the proportion of positive examples correctly predicted by the model to the actual positive examples.
[0028]
[0029] (3) F1-score: F1-score is the harmonic mean of precision and recall, which takes into account the performance of both. It is a comprehensive performance indicator that can provide a more comprehensive evaluation when dealing with imbalanced datasets.
[0030]
[0031] (4) AUC is calculated based on the area of the ROC curve (Receiver Operating Characteristic Curve), also known as the AUC-ROC curve (Area Under the Receiver Operating Characteristic Curve), which is an important tool for evaluating the performance of binary classification models. This curve plots the relationship between the true positive rate (True Positive Rate) and the false positive rate (False Positive Rate) at different thresholds, which can intuitively understand the performance of the model in the classification task. The area under the curve represented by the AUC-ROC curve can reflect the model's ability to distinguish between positive and negative examples, thereby evaluating the quality of the model. The ROC curve is a curve drawn with FPR (False Positive Rate) as the horizontal axis and TPR (True Positive Rate, i.e. Recall) as the vertical axis. The value range of AUC is between 0 and 1, and the larger the value, the better the model performance.
[0032] FPR=FP / (TN+FP)
[0033] TPR=Recall=TP / (TP+FN)
[0034] Among them, TP represents true positive examples, FP represents false positive examples, FN represents false negative examples, and TN (True Negatives) represents true negative examples, that is, the number of samples correctly predicted as negative classes.
[0035] Furthermore, in step S4, after comparing the SVM and RF machine learning algorithms, the optimal model is selected by taking the average AUC of 10 cross-validations performed under different hyperparameter tuning settings.
[0036] Furthermore, in step S4, RF is implemented by calling the RandomForest-Classifier module through the SKLEARN library. The model k-fold cross validation uses 10 cross validations, 500 decision trees, a maximum depth of 30 for each tree, and a minimum number of samples required for node splitting of 2. The random forest classifier with a random seed of 42 is set to correct the model instability caused by overfitting of a single decision tree. The SVM model is constructed using the SVM module of the SKLEARN library. The optimal hyperparameters of the algorithm are determined through grid search and 10 cross validations. In order to conduct a comprehensive and rigorous evaluation of the model performance, it is necessary to calculate and summarize the average values of various classification indicators under the optimal parameter settings.
[0037] Furthermore, in step S5, the RF parameters are set as follows: the total number of samples is 1781, the ratio of the training set to the test set is 8:2, the training set has 1425 compounds, the test set has 356 compounds, and there are two classifications: 'active' and 'inactive', and RF and SVM modeling is performed.
[0038] Furthermore, in step S6, the top 10% of compounds with higher activity prediction probability were selected for molecular docking, and the binding energy score ≤-10 kcal / mol was set as the threshold for screening, and a total of 252 compounds were determined to be compounds with potential FXR regulating activity.
[0039] Furthermore, in step S7, 95 drug library (FDA_HY-L022 Drug Library) compounds with a regulation activity of 130% or more at a concentration of 10 μM were considered agonists; compounds with a regulation multiple of less than 50% were screened and considered antagonists or inverse agonists.
[0040] Furthermore, in step S8, a total of 37 FXR agonists were analyzed for their anti-HBV activity, of which 2 compounds had an HBeAg inhibition rate of >50%; 6 small molecules had an HBsAg inhibition rate of >50%, and MTT was used to detect cell viability, and 36 compounds had a survival rate greater than 70%, indicating that 36 FXR agonists had anti-HBV activity without affecting cell viability.
[0041] As a second aspect, an FXR modulator with anti-HBV activity obtained by machine learning screening is provided, which is any compound shown in the structural formula FL113, FL116, FL889, FL917, FL1387, FL1578, FL1855:
[0042]
[0043] The beneficial effects of the present invention include: the present invention adopts advanced machine learning technology to provide an efficient method for the screening of FXR modulators, which is of great significance in the field of anti-HBV therapeutic drugs. The present invention has developed an innovative screening process: first, the compounds in the database are processed using molecular fingerprints; second, the database is trained using advanced support vector machine models and random forest models to improve the accuracy and stability of the model in predicting the activity of FXR modulators; then, molecular docking is used to assist in the screening of FXR modulators; finally, through in vitro activity verification, it is effectively tested whether the model has learned key structural features from the training data set to ensure the screening ability of the model. The experimental results show that the model shows good performance on the test set, with an average AUC value of 0.93, which proves its predictive accuracy in the screening of FXR modulators. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a flow chart of the method for discovering FXR modulators of anti-HBV activity based on machine learning according to the present invention;
[0045] Figure 2 It is the structure of the HBV inhibitor with FXR regulating activity described in the present invention. DETAILED DESCRIPTION
[0046] The invention pre-processes the collected compound data, including data cleaning, data deduplication and label extraction; constructs a molecular fingerprint database in a compound database by using molecular fingerprint extraction; constructs a classification prediction model of FXR modulator active compounds by using random forest (RF) and support vector machine (SVM) algorithms, determines the optimal parameters of the algorithms by setting 10 cross-validations, selects the optimal model of the two models by multi-dimensional comparison, and finally selects the RF (random forest) model to predict the activity of drug library (FDA_HY-L022 Library) molecules; performs molecular docking analysis on compounds with higher predicted activity probability, and screens them again; uses a HEK-293T-FXR-Luc stable transfection cell line containing an FXR response element (FXRE) and capable of stably expressing luciferase to experimentally detect the FXR regulation activity, and uses GW4064 as a positive control.
[0047] The following will be combined Figure 1 , the technical scheme and beneficial effects of the present invention are described in detail.
[0048] The present invention provides a method for predicting the activity of FXR modulators based on machine learning, which specifically comprises the following steps:
[0049] S1, FXR regulator data were collected from the ChEMBL database, and then the collected data were preprocessed, including data cleaning, data deduplication, and label extraction.
[0050] Data collection: FXR modulator data were collected from the ChEMBL database, including affinity values, EC 50 Value, IC 50 values and SMILES strings, and 1781 FXR regulator-related data were collected;
[0051] Table 1 Statistics of sample numbers of each data set
[0052] Dataset Positive Sample Negative samples total Training set 711 714 1425 Test Set 178 178 356 total 889 892 1781
[0053] Data preprocessing mainly includes data cleaning, data deduplication and label extraction.
[0054] Data cleaning mainly involves deleting active unavailable data and useless molecules such as metals and metal oxides.
[0055] Data deduplication mainly involves removing duplicates after standardizing the SMIlES string, retaining the only SMIlES as the molecular representation; different EC values generated when the same inhibitor molecule acts on different types of FXR mutants 50 value, keep the lowest value;
[0056] The affinity value of the ligand extracted by the tag and the binding to FXR, and the half-activation concentration value EC generated after the ligand binds to FXR 50 , half inhibitory concentration value IC 50 The activity threshold was selected as 2000nM, and the EC 50 Molecules with values less than the activity threshold were labeled as class "1" and considered as active modulators of FXR, otherwise they were considered as weak or inactive inhibitors and labeled as class "0".
[0057] S2, use relevant tools to calculate features and filter the calculated descriptors, and then divide the dataset into training set and test set.
[0058] The feature calculation uses 12 molecular fingerprint descriptors in PaDEL-Descriptor: AtomPairs2D-FingerprintCount, AtomPairs2D-Fingerprinter, Estate-Fingerprinter, Extended-Fingerprinter, Fingerprinter, GraphOnly-Fingerprinter, KlekotaRoth-FingerprintCount, KlekotaRothFingerprinter, MACCS-Fingerprinter, Pubchem-Fingerprinter, Substructure-FingerprintCount, SubstructureFingerprinter.
[0059] The fingerprint descriptor is the binary value corresponding to the feature. For example, each feature of the MACCS fingerprint corresponds to a specific chemical substructure, such as a hydroxyl group, a benzene ring, or a nitrogen atom. If the structure exists, the value of the binary bit corresponding to the feature is 1, otherwise it is 0.
[0060] The classification prediction model is selected by calculating the accuracy, precision, recall, F1-score, and AUC, and by comprehensively comparing the scores of each indicator. The calculation formulas for each indicator are as follows:
[0061] (1) Accuracy: Accuracy is the most common performance evaluation metric for classification problems. It measures the proportion of samples correctly predicted by the model. However, in the case of an unbalanced class distribution, accuracy can be misleading, so other metrics need to be combined to evaluate model performance.
[0062]
[0063] (2) Precision and Recall: Precision and recall are important indicators for the problem of imbalanced class distribution. Precision measures the proportion of true positive examples among the positive examples predicted by the model, while recall measures the proportion of positive examples correctly predicted by the model to the actual positive examples.
[0064]
[0065] (3) F1-score: F1-score is the harmonic mean of precision and recall, which takes into account the performance of both. It is a comprehensive performance indicator that can provide a more comprehensive evaluation when dealing with imbalanced datasets.
[0066]
[0067] (4) AUC is calculated based on the area of the ROC curve (Receiver Operating Characteristic Curve), also known as the AUC-ROC curve (Area Under the Receiver Operating Characteristic Curve), which is an important tool for evaluating the performance of binary classification models. This curve plots the relationship between the true positive rate (True Positive Rate) and the false positive rate (False Positive Rate) at different thresholds, which can intuitively understand the performance of the model in the classification task. The area under the curve represented by the AUC-ROC curve can reflect the model's ability to distinguish between positive and negative examples, thereby evaluating the quality of the model. The ROC curve is a curve drawn with FPR (False Positive Rate) as the horizontal axis and TPR (True Positive Rate, i.e. Recall) as the vertical axis. The value range of AUC is between 0 and 1, and the larger the value, the better the model performance.
[0068] FPR=FP / (TN+FP)
[0069] TPR=Recall=TP / (TP+FN)
[0070] Among them, TP represents true positive examples, FP represents false positive examples, FN represents false negative examples, and TN (True Negatives) represents true negative examples, that is, the number of samples correctly predicted as negative classes.
[0071] Table 2 compares the effects of two machine learning models, random forest (RF) and support vector machine (SVM) models, and also shows the performance comparison of the two models under different feature screening methods, including accuracy, AUC, precision, recall and F1 value (F1-score). Taking all factors into consideration, the random forest model of the present invention shows the best performance. Table 3 compares the results of the calculation of different molecular fingerprint descriptor combinations by random forest, and the results show that the fingerprint indicators of Substructure Fingerprint Count all achieve very good results.
[0072] Table 2 Comparison of calculations of twelve molecular fingerprint descriptors using random forest (RF) and support vector machine (SVM)
[0073]
[0074]
[0075] Table 3 Results of molecular fingerprint descriptor combination calculation RF model
[0076]
[0077] S3, the classification model is applied to the external validation set for further validation and evaluation, the activity of unknown molecules is predicted, and the prediction result of whether each molecule has FXR regulatory activity is output.
[0078] S4, molecular docking was performed on the top 10% of compounds with higher predicted probability of activity by combining with the docking model, and the binding energy score was scored. The threshold of the binding energy score was set as Docking Score ≤ -10 kcal / mol. A total of 252 compounds were identified as compounds with potential FXR regulatory activity.
[0079] S5, FXR regulatory activity screening:
[0080] (1) Construction of HEK-293T-FXR-Luc cell line: First, the high-purity, endotoxin-free lentiviral vector and its auxiliary packaging vector plasmid were extracted and co-transfected into HEK-293T cells using HG transgene reagent. After a series of culture and treatments, including adding Enhancing buffer after transfection, replacing fresh culture medium, collecting and concentrating the cell supernatant rich in lentiviral particles, the virus titer was finally measured and calibrated in HEK-293T cells to ensure that the quality and quantity of the virus met the experimental requirements. Subsequently, HEK-293T cells were infected with the packaged CMV-Luc-PGK lentivirus. After cell inoculation, virus infection, replacement of culture medium, detection of infection efficiency and resistance screening, the HEK-293T-FXR-Luc stable strain was successfully constructed.
[0081] (2) Validation of HEK-293T-FXR-Luc cells: In order to verify the stability and functionality of the constructed HEK-293T-FXR-Luc cell line, we conducted a series of experiments. First, the cell morphology was observed under an inverted microscope to ensure that the cell confluence was suitable for the experiment. Then, through steps such as cell resuspension, inoculation of 96-well plates, and replacement of culture medium, the cells were tested for luciferase activity using the GMOne-Step Luciferase Reporter Gene Assay Kit.
[0082] (3) Determination of the initial DMSO concentration in drug screening: Before drug screening, the toxicity of DMSO to HEK-293T-FXR-Luc cells was first determined. Ten concentration gradients of DMSO from 0‰ to 9‰ were set for verification, and the growth state and luciferase activity of cells under different concentrations of DMSO were observed, thereby determining the appropriate DMSO concentration range.
[0083] (4) Determination of the effective concentration and EC of the positive drug GW4064 50 Determination: By comparing the effects of different concentrations of GW4064 on the luciferase activity of HEK-293T-FXR-Luc cells, we finally determined that 8μM GW4064 was the optimal positive drug concentration. This result provides an important reference for our subsequent drug screening.
[0084] (5) Activity analysis of the FXR regulatory effect of the test molecule at the cellular level: After determining the DMSO concentration and GW4064 concentration, the screening of the actual compound began. Through cell testing, the patented compound was found to have good FXR regulatory activity.
[0085] Table 4 Regulatory activity of compounds on intracellular FXR
[0086]
[0087]
[0088] Table 5. Structures of FXR modulator compounds
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098] S6, FXR modulator in vitro anti-HBV activity test: The sample solution was prepared by dissolving the aforementioned 95 FXR modulator compounds in DMSO, and the solution concentration was 30 μM. The ELISA method was used to determine the effects of 95 FXR modulators on the secretion of HBeAg and HBsAg antigens. The specific experimental operation steps are as follows: First, the HepG2.2.15 cells with appropriate cell status were washed with PBS, trypsin was added to disperse the cells, and culture medium (MEM+10% FBS+380mg / mL G418) was added and pipetted into a single cell suspension; then, the suspension was transferred to an EP tube (15mL), centrifuged to remove the supernatant, and then 2mL of complete culture medium (MEM+10% FBS+380μg / mL G418) was added and pipetted into a single cell suspension, and re-seeded in a 48-well plate, with about 3000 cells per well, and placed in an incubator at 37°C, 5% CO 2 Finally, the culture medium was discarded, and the cells were treated with the prepared sample solution (final concentration 30 μM), lamivudine was used as a positive control, and cells without sample treatment were used as blank controls. The cells were placed in an incubator (37°C, 5% CO 2 ) and the culture fluid was collected after 72 hours of culture. The samples were tested using the HBeAg detection kit (Shanghai Kehua Bioengineering Co., Ltd., National Medical Device No. 20163400144) and the HBsAg detection kit (Shanghai Kehua Bioengineering Co., Ltd., National Medicine No. S10910113). The absorbance value (OD value) was measured using an ELISA reader at a detection wavelength of 450nm and a reference wavelength of 630nm. The inhibition rate of the compound on HBsAg and HBeAg was calculated according to the formula as follows:
[0099]
[0100] S7, FXR modulators in vitro MTT assay to determine compound cell viability
[0101] (1) HepG2.2.15 cells in appropriate cell states were washed with PBS, the cell culture medium was removed, 0.25% trypsin was added to digest the cells, culture medium (MEM+10% FBS+380 mg / mL G418) was added and pipetted into a single cell suspension; then, the suspension was transferred to an EP tube (15 mL), centrifuged to remove the supernatant, and 2 mL of complete culture medium (MEM+10% FBS+380 mg / mL G418) was added and pipetted into a single cell suspension, which was re-seeded in a 96-well plate with approximately 3,000 cells per well;
[0102] (2) After 12 hours of adhesion, the culture medium was removed and 100 μL of culture medium (MEM + 10% EBS + 380 mg / mL G418) containing different concentrations of the test sample (diluted from below 200 μM to 1.56 μM) was added. Blank wells (containing only culture medium) and control wells (solvent control without adding the test sample) were also set up;
[0103] (3) Cells and drugs were maintained at 37°C, 5% CO 2 After 24 hours of co-incubation, 10 μL of MTT solution (10 mg / mL) was added to each well. After 4 hours of co-incubation, 100 μL of MTT solution was added. After incubation at 37°C overnight, the absorbance (A value) of each well was measured at a wavelength of 550 nm.
[0104] (4) Calculate the cell survival rate. Survival rate = (experimental well - blank well) / (control well - blank well) × 100%. The experiment was repeated three times.
[0105] Among them, the anti-HBV activity of 37 FXR agonists, ie, FXR modulators with a regulatory activity greater than 130%, was analyzed (Table 6). The results showed that 6 small molecules had an inhibition rate of >50% on cell secretion of HBsAg, and 2 small molecules had an inhibition rate of >50% on HBeAg. Figure 2 The structures of compounds with regulatory activity greater than 130% in HEK-293T-FXR-Luc cell assays, inhibition rates of HBsAg secretion by cells greater than 50% in HepG 2.2.15 cell assays, and inhibition rates of HBeAg greater than 50% were displayed. MTT assay of cell viability showed that 36 compounds had a viability greater than 70%, indicating that 36 FXR agonists had anti-HBV activity without affecting cell viability.
[0106] Table 6 Anti-HBV activity screening results
[0107]
[0108] Therefore, the present invention proposes an efficient and accurate FXR modulator screening method based on machine learning, which not only successfully constructs a virtual screening model for FXR modulators with an AUC value of up to 0.93, demonstrating excellent data set prediction capabilities, but also verifies the reliability of the model through cell experiments. Furthermore, in the test for anti-HBV activity, this method successfully screened out 6 active molecules that have both significant FXR agonist activity and the ability to inhibit HBsAg and HBeAg secretion, providing new candidate compounds for the research and development of FXR-related drugs, demonstrating the great potential and application value of the present invention in the field of drug discovery.
Claims
1. A method for screening FXR modulators for anti-HBV activity based on machine learning, characterized in that: The following steps are involved: Collect compound data, including affinity values, EC 50 Value, IC 50 Value and SMILES string; set the label to whether it has FXR regulatory activity, cluster the compounds using molecular fingerprints, and build a FXR modulator compound library; Calculate the molecular fingerprint descriptor of the compound as a training feature, establish a machine learning model for training, obtain the FXR regulatory activity prediction model, input the test compound data, and screen multiple test compounds with top FXR regulatory activity prediction results as candidate compounds; The candidate compounds were screened for FXR modulator activity and phenotype verification of anti-HBV activity to screen out potential new FXR modulator compounds.
2. The method for screening FXR modulators for anti-HBV activity based on machine learning according to claim 1, characterized in that: After the compound data is collected, the method further includes: cleaning the data, removing duplicates, unifying concentration units, and extracting labels; The setting label is whether there is FXR regulatory activity, specifically: setting EC 50 The value of 2000 nM is the activity threshold; The clustering of compounds using molecular fingerprints is specifically as follows: using MACCS molecular fingerprints, calculating Tanimoto coefficients to obtain similarity for clustering evaluation, which is used to indicate the molecular types in the database; After constructing the FXR modulator compound library, the method further includes: dividing the library into a training set and a test set, and adjusting the balance of the proportion of compounds labeled as having FXR modulating activity and compounds not having FXR modulating activity in the training set; The test compound data comes from the FDA_HY-L022 Library compound library.
3. The method for screening FXR modulators for anti-HBV activity based on machine learning according to claim 1, characterized in that: The molecular fingerprint descriptor is selected from AtomPairs2D-FingerprintCount, One or a combination of AtomPairs2D-Fingerprinter, Estate-Fingerprinter, Extended-Fingerprinter, Fingerprinter, GraphOnly-Fingerprinter, KlekotaRoth-FingerprintCount, KlekotaRothFingerprinter, MACCS-Fingerprinter, Pubchem-Fingerprinter, Substructure-FingerprintCount, SubstructureFingerprinter.
4. The method for screening FXR modulators for anti-HBV activity based on machine learning according to claim 1, characterized in that: After the FXR regulatory activity prediction model is obtained, the method further includes: calculating the FXR regulatory activity prediction model obtained by Accuracy, Precision, Recall, F1 value, and AUC evaluation, and using it to optimize the training features and machine learning methods.
5. The method for screening FXR modulators for anti-HBV activity based on machine learning according to claim 1, characterized in that: After screening a plurality of test compounds with top predicted results of FXR regulation activity, the method further comprises: The crystal model of FXR and ligand complex was used to simulate the candidate compounds using molecular docking technology, and combined with energy scoring, several top-ranked compounds were screened out.
6. The method for screening FXR modulators for anti-HBV activity based on machine learning according to claim 1, characterized in that: The steps of screening the FXR regulatory activity include: A cell line containing FXR response elements and capable of stably expressing luciferase was constructed, and candidate compounds were screened at a concentration of 10 μM. Compounds with a regulation fold greater than 130% were considered agonists; compounds with a regulation fold less than 50% were considered antagonists or inverse agonists.
7. The method for screening FXR modulators for anti-HBV activity based on machine learning according to claim 6, characterized in that: The steps of anti-HBV activity phenotype verification include: The HepG2.2.15 cell ELISA method was used to determine the effects of FXR agonist compounds on the secretion of HBeAg and HBsAg antigens, and the MTT method was used to test cell viability to exclude the influence of cytotoxicity on the anti-HBV activity test.
8. An FXR modulator with anti-HBV activity obtained based on machine learning screening, characterized in that: The FXR modulator with anti-HBV activity is any compound represented by the structural formula FL113, FL116, FL889, FL917, FL1387, FL1578, FL1855:
Citation Information
Patent Citations
Liver disease-related biomarkers and methods of use thereof
CN108990420A
Genome recombination fingerprint for characterizing hHRD homologous recombination deficiency and identification method thereof
CN110241198A
Virtual screening method of IRAK1 kinase inhibitor and drug lead compound
CN112259175A
Application of anti-inflammatory agent in preparation of medicine for treating or preventing viral hepatitis
CN115337310A
Farnesol X receptor antagonist and virtual screening method and application thereof
CN116130027A