Method for predicting the three-dimensional folding of a g-protein coupled receptor protein and the binding model of a drug molecule

By optimizing the flexible regions and drug structure sites of the GPCR model through three-dimensional sequence alignment and template structure selection, and combining cross-validation and machine learning algorithms, the problem of accurately predicting the three-dimensional structure of GPCRs and the binding mode of drug molecules was solved, achieving higher accuracy prediction results.

CN116386764BActive Publication Date: 2026-02-06ALPHAMOL SCIENCE LTD (SHANGHAI)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211727118.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-02-06
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

Existing technologies lack precision and applicability in predicting the three-dimensional structure of G protein-coupled receptor proteins (GPCRs) and drug molecule binding modes, thus failing to effectively guide drug development.

Method used

We used three-dimensional sequence alignment and template structure to select the initial GPCR model, optimized the flexible region and drug structure site, and combined cross-validation of various types of GPCR-drug molecule artificial intelligence models. The optimal binding mode was selected through cross-validation, and prediction models were constructed using machine learning algorithms such as Scikit-learn, XGBoost, Auto-Sklearn, H2O and PyCaret.

Benefits of technology

It achieves accurate prediction of the three-dimensional structure of GPCRs and precise capture of drug molecule binding modes, improving prediction accuracy and surpassing the performance of AlphaFold2 and RosettaFold, providing a solid foundation for GPCR drug design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386764B_ABST
    Figure CN116386764B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting a G protein-coupled receptor protein three-dimensional folding and a drug molecule binding model. The method comprises the following steps: selecting a plurality of target GPCR initial models by using three-dimensional sequence alignment and template three-dimensional structure selection, and optimizing a flexible area of the GPCR initial model and a drug structure site to obtain a target GPCR model for molecular docking and drug design; and selecting an optimal GPCR and drug molecule binding mode by cross-verification of prediction performances of a plurality of types of GPCR-drug molecule artificial intelligence models. The application improves the prediction accuracy of the G protein-coupled receptor protein three-dimensional folding and the GPCR drug molecule binding mode of a drug target, and can also accurately capture the interaction mode of the related drug molecule and the GPCR target.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of medicine, and particularly relates to a method for predicting a G protein-coupled receptor protein (GPCR) three-dimensional folding and drug molecule binding model. BACKGROUND

[0002] Membrane proteins are a large class of lipid-soluble proteins embedded in the cell membrane, which are relatively water-soluble proteins. Membrane proteins include GPCRs, ion channels, transport proteins, nuclear proteins, ion pumps, etc. Water-soluble proteins include enzymes, kinases, chaperones, etc. According to statistics, membrane proteins account for more than 70% of the entire drug targets on the market, and water-soluble proteins account for only about 20% of the proportion, and other targets are nucleic acid macromolecules. GPCRs are also known as seven-transmembrane helix proteins, which are the most important membrane protein drug targets. Currently, nearly 40% of the drugs on the market are developed for GPCRs.

[0003] Proteins are the basis of life, and the three-dimensional structure of proteins directly determines their physiological functions. Therefore, analyzing and predicting the three-dimensional structure of proteins has become an important and indispensable basic link in the development of new drugs. Although traditional experimental methods (including cryo-electron microscopy, X-ray diffraction, and nuclear magnetic resonance) have made great progress in the field of water-soluble protein structure analysis, the structure analysis of membrane proteins has been very slow. This is mainly due to the fact that the expression and purification technology and conditions of membrane proteins are not yet mature. It often takes years to analyze a new membrane protein structure. With the rapid development of biotechnology and artificial intelligence (AI) technology, the use of computers to predict protein folding and three-dimensional structure has made great breakthroughs. For example, Google's protein folding prediction tool AlphaFold2 has improved the prediction accuracy. In addition, another protein folding tool RosettaFold can also accurately predict the three-dimensional structure of proteins.

[0004] However, with the release of the AlphaFold2 tool and the public disclosure of the predicted target structure models, more and more structural biologists have found that AlphaFold2 is better for smaller water-soluble proteins, but very poor for membrane protein structure prediction. In February 2021, a number of American Academy of Sciences, including Alex T. Brunger, published a review article entitled "The protein-folding problem: Not yet solved" in the international top academic journal "Nature". They openly stated that they did not agree with the AlphaFold2 team's claim that the protein folding problem had been completely solved by AI. Subsequently, Michael J. E. Sternberg et al. published a paper entitled "The AlphaFold Database of Protein Structures: A Biologist's Guide" in the academic journal "Journal of Molecular Biology". The article clearly pointed out that during the evaluation of multiple membrane proteins, the AlphaFold2 prediction had no reliable correlation with the experimental structure. In April 2022, Professor Brian Roth, a well-known structural biologist and pharmacologist in the field of GPCRs, published an article entitled "What's next for AlphaFold and the AI protein-folding revolution" in the "Nature" journal. He pointed out in the paper that "there are dozens of GPCR structures in his laboratory that have been resolved but have not been published. Half of the structures have decent AlphaFold predictions, but more than half of the structures have no value in AlphaFold predictions. For example, some structures are predicted by AlphaFold to be highly reliable, but experimental structures show that the predictions are completely wrong." In July 2022, Professor Eric Xu of the Shanghai Institute of Materia Medica of the Chinese Academy of Sciences published an academic paper entitled "AlphaFold2 versus experimental structures: evaluation on G protein-coupled receptors" in the "Acta Pharmacologica Sinica" journal. He pointed out that "AlphaFold2 can only predict the general skeleton of GPCR target proteins, and the prediction of the key transmembrane helix and drug molecule structure pocket of GPCRs is very different from the experimental structure. AlphaFold2's prediction of GPCR membrane protein structure cannot be applied to drug development guidance related research work."

[0005] In summary, the existing scheme for predicting the three-dimensional structure of GPCRs and the binding mode of drug molecules to GPCRs needs to be improved to further improve the reliability and scope of the prediction. SUMMARY

[0006] The purpose of the present application is to overcome the above-mentioned defects of the prior art and provide a method for predicting the three-dimensional folding of G protein-coupled receptor proteins and the binding model of drug molecules. The method comprises the following steps:

[0007] A plurality of target GPCR initial models are selected using three-dimensional sequence alignment and template three-dimensional structure, and the target GPCR model for molecular docking and drug design is obtained by optimizing the flexible regions of the GPCR initial models and the drug structure sites.

[0008] The prediction performance of various types of GPCR-drug molecule artificial intelligence models is verified by cross-validation, and the best GPCR-drug molecule binding mode is selected.

[0009] The input of the GPCR-drug molecule artificial intelligence model is a three-dimensional interaction fingerprint feature of a ligand in a receptor, and each column represents the number of interaction bonds between the corresponding ligand and the amino acid in the current column. The output is the root mean square deviation (RMSD) value, which is used to measure the distance between the predicted 3D conformation and the real conformation.

[0010] Compared with the prior art, the present application has the advantages of providing a method for accurately predicting the three-dimensional folding of the most important drug target G protein-coupled receptor protein (GPCR) and its drug molecule binding mode. Not only can the three-dimensional structure of GPCR be accurately predicted, but also the binding mode of drug molecules to GPCR can be accurately predicted, which provides a strong guarantee for accurate GPCR drug design.

[0011] Other features and advantages of the present application will become apparent from the following detailed description of exemplary embodiments thereof, which is to be taken in conjunction with the appended drawings. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments of the present application and, together with the description, serve to explain the principles of the application.

[0013] Figure 1 is a flowchart of a method for predicting the three-dimensional folding of G protein-coupled receptor proteins (GPCRs) and the binding model of drug molecules according to an embodiment of the present application;

[0014] Figure 2 is a schematic diagram of the calculation process for predicting and optimizing the three-dimensional folding model of GPCRs according to an embodiment of the present application;

[0015] Figure 3 is a schematic diagram of the GPCR-drug molecule binding AI model building process according to an embodiment of the present application;

[0016] Figure 4 is a schematic diagram of the GPCR amino acid relative position renumbering and extension according to an embodiment of the present application;

[0017] Figure 5 is a schematic diagram of the interaction fingerprint matrix generation according to an embodiment of the present application;

[0018] Figure 6 is a schematic diagram of the prediction results on 5 new GPCR structures according to an embodiment of the present application;

[0019] Figure 7 is a schematic diagram of the comparison of the prediction results with AlphaFold2 and RosettaFold prediction results according to an embodiment of the present application;

[0020] Figure 8 is a schematic diagram of the precise prediction of the complex drug small molecule and GPCR receptor binding model according to an embodiment of the present application;

[0021] Figure 9 is a schematic diagram of the comparison of the three-dimensional folding prediction results on orphan receptor GPR158 with the prior art according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] Various exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of the components and steps, numerical expressions, and numerical values set forth in these embodiments are not limiting to the scope of the present application unless otherwise specifically stated.

[0023] The following description of at least one exemplary embodiment is merely exemplary in nature and is in no way intended to limit the scope of the application, its application, or uses.

[0024] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail herein. However, the techniques, methods, and devices should be considered part of the specification, if appropriate.

[0025] In all of the examples shown and discussed herein, any specific values should be interpreted as merely exemplary, and not as a limitation. Thus, other examples of the exemplary embodiments can have different values.

[0026] It should be noted that like reference numerals and letters refer to like items throughout the several views of the drawings, and that discussion of one item in a drawing does not necessitate further discussion of that item in subsequent drawings.

[0027] Referring to Figure 1 As shown in the provided method for predicting the three-dimensional folding of G protein-coupled receptor protein (GPCR) and the model of drug molecule binding, the method comprises the following steps. Step S110, selecting a plurality of target GPCR initial models by using three-dimensional sequence alignment and template three-dimensional structure, and optimizing the flexible region of the GPCR initial model and the drug structure site to obtain a target GPCR model for molecular docking and drug design; and step S120, selecting the best GPCR-drug molecule binding mode by cross- verifying the prediction performance of various types of GPCR-drug molecule artificial intelligence models. In short, the technical solution of the present application mainly comprises the processes of GPCR folding model prediction and optimization, GPCR-drug molecule AI model construction, and GPCR-drug molecule binding mode prediction. In the following, the details will be described in combination with the drawings.

[0028] Specifically, the method for predicting the three-dimensional folding of G protein-coupled receptor protein (GPCR) and the model of drug molecule binding comprises the following steps.

[0029] Step S1.1, predicting and optimizing the GPCR three-dimensional folding model

[0030] Referring to Figure 2 As shown in the provided method for predicting the three-dimensional folding of G protein-coupled receptor protein (GPCR) and the model of drug molecule binding, the method comprises the following steps. Step S110, selecting a plurality of target GPCR initial models by using three-dimensional sequence alignment and template three-dimensional structure, and optimizing the flexible region of the GPCR initial model and the drug structure site to obtain a target GPCR model for molecular docking and drug design; and step S120, selecting the best GPCR-drug molecule binding mode by cross- verifying the prediction performance of various types of GPCR-drug molecule artificial intelligence models. In short, the technical solution of the present application mainly comprises the processes of GPCR folding model prediction and optimization, GPCR-drug molecule AI model construction, and GPCR-drug molecule binding mode prediction. In the following, the details will be described in combination with the drawings.

[0031] Step S1.1.1, template selection and three-dimensional sequence alignment

[0032] First, find the experimental structure with high homology as the homologous template for computer modeling in the protein structure database website (www.rcsb.org). After determining the primary sequence of the target GPCR amino acid, the secondary structure is predicted.

[0033] In order to determine the helix secondary structure, first find the polypeptide chain atoms of each amino acid, and then calculate the hydrogen bond between the i-th and i+4-th amino acids on the polypeptide chain. The helix structure is supported by the hydrogen bond vortex. In addition, the central carbon atom (carbon alpha) on the backbone of all amino acids is defined as "CA" in the structure file. CA is surrounded by three main functional groups: secondary amine, carboxyl and side chain. According to the understanding, the carboxyl and side chain are covalently bonded to CA through carbon atoms. The carbon combined with the side chain is annotated as "CB" in the structure file. This information allows the carbon of the side chain and the carbon of the carboxyl to be distinguished. Therefore, the atoms on the polypeptide chain are identified by calculating the covalent distance between the carbon of the carboxyl and the nitrogen of the secondary amine and carbon alpha.

[0034] After obtaining the three-dimensional information of hydrogen, nitrogen and oxygen in the polypeptide chain and ensuring that the sequence length of the protein itself is greater than 4 amino acids, the hydrogen bond calculation of the polypeptide chain can be performed. Because the experimental structure lacks information of hydrogen atoms, nitrogen can be used instead of hydrogen when calculating the hydrogen bond. Among all amino acids, proline is a special structure of amino acid. When calculating the helix sequence, proline can be automatically included in the helix sequence. Finally, the helix length must be greater than 2 to be marked as a helix.

[0035] The transmembrane helix region, extracellular region and intracellular region in the primary sequence structure are determined. The hydrophobic and hydrophilic part information of the transmembrane helix region (TM) in the primary sequence is analyzed as a constraint condition for subsequent modeling. The conserved amino acids of each transmembrane helix and the position of the disulfide bond of the flexible loop region of the extracellular ECL2 are determined. The primary sequence of the target GPCR is three-dimensionally aligned with the homologous experimental structure to ensure that all conserved amino acids are correctly aligned.

[0036] Step S1.1.2, generating a structural model

[0037] Using three-dimensional sequence alignment and template three-dimensional structure, for example, 50-100,000 initial models of the target GPCR are generated. The structural stability is evaluated by the CHARMM molecular force field energy scoring function, and the energy minimum model is selected as the initial optimization model.

[0038] Step S1.1.3, optimizing the flexible region and drug structure site

[0039] Because there are many flexible regions in the non-transmembrane region of the GPCR and the binding site of the drug molecule is very close to these regions, it is necessary to optimize the structure of these regions. First, the output initial optimization model in step S1.1.2 is constructed by Gromacs software to build a cell membrane physiological environment. The relevant GPCR will be embedded in the cell membrane. The intracellular and extracellular of the cell membrane are filled with water molecules, and contain 0.15 moles of NaCl salt molecules, simulating the situation of the relevant system in the real physiological state. Then, the position restriction is added to the main chain part of the transmembrane helix region of the GPCR, and long-time scale all-atom molecular dynamics simulation (simulation time length is 500-1000 ns) is performed. The simulation molecular force field can be CHARMM or Amber force field. Finally, the last 50 ns of the simulation process is clustered and analyzed, and the representative structure in the largest cluster is selected for energy minimization optimization and used as the final model of the target GPCR for subsequent molecular docking and drug design.

[0040] Step S1.2, constructing a GPCR-drug molecule AI model

[0041] Combination Figure 3 As shown, constructing a GPCR-drug molecule AI model includes the following steps:

[0042] Step S1.2.1, data collection, cleaning and expansion

[0043] Firstly, all the experimental structures of GPCRs and their drug molecules complex were downloaded from PDB public database, and the corresponding eight-point Uniprot ID was downloaded. Polypeptide molecules with more than 10 amino acids were deleted from the relevant structures. All the structures of GPCRs were aligned to a GPCR structure, and the selected superimposition template PDB number was 5C1M.

[0044] All the molecular structure information and biological activity information (including Kd, Ki, EC 50 , IC 50 values) of the reported GPCR target molecules were downloaded from Pubchem, CHEMBL, PDBBIND, BindingDB and some commercial databases. Molecules with activity greater than 1 μM were deleted from the relevant databases, and finally 148733 compounds were left.

[0045] Step S1.2.2, preparing model data

[0046] Step S1.2.2.1, molecular clustering

[0047] The relevant molecules were docked into their experimental structures by computer cross-molecular docking (multiple molecular docking software was docked at the same time, and the intersection of the results was taken). The original experimental structure molecules can be used as reference ligands for molecular docking. The variance RMSD (root mean square deviation) of the atomic-atomic direct distance between the new molecules and the molecules in the experimental structure was calculated. According to the RMSD value, clustering was performed to obtain 20+1 different classifications. Finally, the structure of the new molecule and the related GPCR complex was obtained and energy minimization was performed using Gromacs.

[0048] Step S1.2.2.2, renumbering GPCR amino acids

[0049] The amino acid numbering of GPCRs usually uses Ballesteros-Weinstein numbering, which redefines the position of amino acids as a sequence in the transmembrane, rather than the entire protein sequence. Starting from the motif, the amino acid closest to the middle of the transmembrane is preferentially selected as the first amino acid, which is annotated at the 50th position in the transmembrane. Then, the numbering is extended to the entire transmembrane sequence, as shown in Figure 4 . The length of the transmembrane needs to be strictly controlled. When the length of the transmembrane sequence is too short, according to the number of amino acids before and after the middle residue, amino acids are added at the starting transmembrane region or the terminal transmembrane region. Conversely, a too long transmembrane sequence should delete some amino acids at the starting or terminal transmembrane region according to similar restrictions.

[0050] Step S1.2.2.3, Calculate GPCR-ligand interaction fingerprints

[0051] The molecular interactions between the receptor and the ligand can be characterized by different parameters, including: the type of interaction, the amino acids involved in the interaction, the atoms of the amino acids and the ligand involved, and the three-dimensional positions of the atoms (x, y, and z). The types of interactions include: hydrophobic contact, hydrogen bond, water bridge, salt bridge, pi stacking interaction, pi-cation interaction, and halogen bond.

[0052] Two interaction fingerprint matrix files are generated after the calculation. These data are roughly divided into three types of information: complex structure information identifier, interaction type count, and amino acid count of interaction. The difference between the two result files is the display of the amino acid type in the third type of information, see Figure 5 , where the interaction fingerprint is shown with Ballesteros-Weinstein numbering, the type of interaction includes the amino acid involved or not involved in the interaction.

[0053] Step S1.2.3, Establish AI model

[0054] Step S1.2.3.1, Select features

[0055] The input X is a 3D interaction fingerprint of a ligand in a receptor, and each column represents the number of interaction bonds between the ligand and the amino acid in the current column, which is a continuous integer; the output y is an RMSD value, representing the distance between the 3D conformation and the real conformation, which is a continuous floating-point value. According to the molecular fingerprint features generated by the 83852 complexes in step S1.2.2.3, the features are relatively sparse, in order to make the machine learning model learn the fingerprint features better, the experiment briefly selects the features and reduces the dimension. Finally, the data dimension is: 83852 rows, 326 columns.

[0056] Step S1.2.3.1.1 Feature selection based on standard deviation

[0057] After normalizing the data, calculate the standard deviation of each column feature. The standard deviation reflects the dispersion degree of each sample feature, and a smaller standard deviation represents that the feature in all samples is closer to the average value and is less likely to have a specific effect on the target value. For example, taking 0.05 as the threshold, remove the column of features with a standard deviation less than 0.05.

[0058] Step S1.2.3.1.2, Feature selection based on correlation

[0059] The correlation degree between each two features can be determined by calculating the size of the correlation coefficient between the features, and the value interval is between [-1, 1]. For example, the Pearson correlation coefficient is used to calculate the correlation coefficient of each column of the molecular fingerprint. If the correlation coefficient is large (for example, greater than 0.9), it is considered that the two columns of fingerprints are strongly correlated, and one column is removed. The calculation method is as follows:

[0060]

[0061] where X, Y are two columns of fingerprints, μ is the mean of the fingerprint column, σ is the standard deviation of the fingerprint column, and E is the expected value.

[0062] Step S1.2.3.1.3. Z-score standardization of the screened samples

[0063] The screened sample fingerprint x is standardized, and the mean of the processed data is 0 and the standard deviation is 1. For example, the standardization method is as follows:

[0064]

[0065] where μ is the mean of the fingerprint column and σ is the standard deviation of the fingerprint column.

[0066] Finally, a total of 322 columns of fingerprints were selected from the 875 columns of fingerprints as features for the machine learning model.

[0067] Step S1.2.3.2. Establishing a regression model

[0068] In one embodiment, the regression model can use Scikit-learn, XGBoost, Auto-Sklean, H2O, Pycaret tools to establish the model.

[0069] Scikit-learn includes a variety of classification, regression and clustering algorithms, including support vector machines, random forests, K-means and other machine learning algorithm models.

[0070] Auto-sklearn includes 12 machine learning models: AdaBoost, ard_regression, decision_tree, extra_trees, gaussian_process, gradient_boosting, k_nearest_neighbors, liblinear_svr, libsvm_svr, mlp (Multi-layer Perceptron), random_forest, sgd (Stochastic Gradient Descent).

[0071] H2O AutoML can be used to automatically train and tune many models within a user-specified time limit. H2O provides many explainability methods for AutoML objects (model groups) as well as individual models. Explanations can be automatically generated and a simple interface is provided to explore and explain AutoML models. To build a predictive model using H2O, the training set accepts the regression module of the H2O toolkit. The model is built using the handout strategy, which divides the data into a training set of 80% (where the training set also uses a 5-fold cross-validation approach), a test set of 20%, and a runtime of 1 hour, 2 hours, 3 hours, 12 hours, and 24 hours.

[0072] PyCaret is a low-code Python machine learning library that brings together a variety of machine learning methods based on the popular R Caret library, and short lines of code can complete data preprocessing, modeling, and finally model preprocessing with minimal human effort. In addition, the ability to compare and tune many models using simple commands can simplify efficiency and productivity while reducing the time it takes to create useful models. The PyCaret team added NVIDIA GPU support in version 2.2, including all the latest and greatest versions in RAPIDS. With GPU acceleration, PyCaret modeling time can be 2 to 200 times faster, depending on the workload. To build a predictive model using PyCaret, the entire dataset is passed to the regression module of PyCaret 2.2, which splits the dataset into training and test sets by default, containing 80% (67081) and 20% (16771) records, respectively. All 19 regression models available in the machine learning library and framework are trained on the training set and ranked according to their R2 scores.

[0073] Step S1.2.3.2.1, split data

[0074] To more accurately establish an RMSD prediction regression model for GPCR complexes, existing data is split to train and validate the model. In the experiment, five-fold cross-validation can be used to split and train the data. Specifically, the data is divided into 5 groups, each group of data is used as a validation set for validation, and the remaining 4 groups are used as a model training set, and 4 models are obtained by cycling 4 times. The average RMSD error is calculated by averaging the errors obtained from the 4 models.

[0075] In execution, it is completed using the KFold command in the scikit-learn toolkit. The shuffle parameter is set to True, which means that the data will be shuffled during each division to ensure the reliability of the results, and the random_state parameter is set to ensure the reproducibility of the experiment.

[0076] Step S1.2.3.2.2, adjusting parameters

[0077] The parameters of the model are adjusted by using Grid Search in Scikit-learn, which is an exhaustive search method for adjusting parameters, that is, in all candidate parameters, through loop traversal, trying every possibility, and obtaining the best parameter set as the final parameter of the model. The candidate parameter range is determined by conventional experience, and the required parameters and range to be adjusted are different for different models. The specific model used and the selected parameter range are described in the following model selection step.

[0078] Step S1.2.3.2.3, selecting a model

[0079] According to the object of the present application, that is, predicting the RMSD value of the GPCR complex from the calculated molecular fingerprint, a regression model is established to fit the relationship between the characteristic value and the target value. A linear regression model is used as a benchmark, and various regression models are tried, including ridge regression, Lasso, support vector machine, decision tree, random forest, K-neighbor, and XGBoost model. The following is a specific introduction of each type of model.

[0080] (a) Linear Regression

[0081] Linear regression is a regression analysis that uses the least squares function to model the relationship between molecular fingerprints and RMSD. It is the simplest and fastest, and is used here as a benchmark for comparison.

[0082] Execution tool: Sklearn.linear_model.LinearRegression

[0083] Parameter adjustment range: none

[0084] (b) Support Vector Machine (SVR)

[0085] Support vector machine is a supervised learning model that can be used for classification and regression problems. Its basic principle is to fit a hyperplane that maximizes the data on the surface. Its characteristics are the use of kernel, sparse solution and VC control margin and the number of support vectors.

[0086] Execution tool: Sklearn.svm.LinearSVR

[0087] (c) Ridge Regression and Lasso Regression

[0088] Ridge and LASSO are both multiple linear regression models, the difference is that LASSO regression model changes the penalty term of the loss function from L2 norm to L1 norm, which can reduce the unimportant regression coefficient to 0, achieve the purpose of eliminating variables, and make the prediction more accurate.

[0089] Execution tool: Sklearn.linear_model.Ridge Sklearn.linear_model.lasso

[0090] (d) Decision Tree

[0091] Decision tree is a basic classification and regression method with tree structure, and regression decision tree mainly refers to CART (classification and regression tree) algorithm, the value of internal node feature is "yes" and "no", which is a binary tree structure.

[0092] Execution tool: sklearn.tree.DecisionTreeRegressor

[0093] (e) K-Nearest Neighbors (KNN)

[0094] K-Nearest Neighbors regression algorithm finds the K nearest neighbor samples of the complex, and assigns the average property of these complexes to the complex to get the corresponding RMSD value of the complex.

[0095] Execution tool: sklearn.neighbors.KNeighborsRegressor

[0096] (f) Random Forest

[0097] Random forest is composed of multiple independent decision trees, and the prediction result is obtained in parallel by randomly selecting complexes and fingerprints. By averaging the results of all trees, the regression prediction result of the whole forest is obtained.

[0098] Execution tool: sklearn.ensemble.RandomForestRegressor

[0099] (g) XGBoost

[0100] XGBoost is also an integrated model based on decision regression tree, which combines multiple related decision trees for prediction. Unlike random forest, in XGBoost, the input samples of the next decision tree are related to the training and prediction of the previous decision tree.

[0101] Implementation tool: xgboost.XGBRegressor (Python API)

[0102] Step S1.2.3.2.4, select evaluation index for measuring model performance

[0103] (a) coefficient of determination (R 2 score)

[0104] The coefficient of determination is a commonly used score in regression models, which can be used as a standard for measuring the prediction ability of the model, representing the percentage of the square of the correlation between the predicted value and the actual value of the target variable. The closer to 1, the better the prediction effect, and the calculation method is:

[0105]

[0106] Wherein, is the predicted RMSD value, y i is the true value, is the true average value.

[0107] (b) mean absolute error (MAE)

[0108] The mean absolute error refers to the average distance between the predicted value of the model and the true sample value. The smaller the value, the better the prediction effect of the model, and the calculation method is:

[0109]

[0110] (c) mean square error (MSE)

[0111] The mean square error refers to the average of the square of the distance between the predicted value of the model and the true sample value. The smaller the value, the better the prediction effect of the model, and the calculation method is:

[0112]

[0113] Step S1.2.3.2.5, analyze test results

[0114] The test results are shown in Table 1. Among all the test models, pycaret, H2O, auto-sklearn, and XGBoost have excellent R 2 , MAE and MSE scores.

[0115] Table 1: Experimental results

[0116] train R 2 ]] test R 2 ]]> MAE MSE auto-sklearn 0.9377 0.8317 1.0488 2.4206 H2O 0.9676 0.8485 1.0230 2.1430 Pycaret 0.9970 0.8543 0.9106 2.0761 xgboost (baseline model) 0.9968 0.8351 0.2675 0.1649 Linear Regressor 0.5060 0.5007 0.5564 0.4992 SVR linear 0.4929 0.4878 0.5501 0.5121 Ridge 0.5060 0.5008 0.5564 0.4992 Lasso 0.4843 0.4812 0.5676 0.5187 Decison Tree 0.7215 0.6594 0.4129 0.3405 RF 0.7716 0.7204 0.3830 0.2796 KNN 0.8603 0.7797 0.2997 0.2203

[0117] Step S1.3, prediction of GPCR-drug molecule binding mode

[0118] Step S1.3.1 prediction process

[0119] Firstly, the reliable three-dimensional folding model of the relevant GPCR target point is obtained by using the algorithm and process in Figure 2 After obtaining the chemical structure of the relevant research drug small molecule, it is docked by the method in Figure 3 100 different conformations are generated. Finally, cross-prediction is performed by the AI models in S1.2, such as pycaret, H2O, auto-sklearn, XGBoost, etc., and the intersection is selected as the final best prediction binding mode.

[0120] Step S1.3.2, prediction result

[0121] In order to further verify whether the present application can correctly predict the three-dimensional structure folding of GPCR and the binding mode of related drug molecules and GPCR, thereby providing theoretical guidance for GPCR drug research and development, five new GPCR target points (including: APJ, GPR139, KOR, NPY1R, NMUR2 receptor) are predicted. The average variance (RMSD, the lower the value, the higher the accuracy) of the atomic-atomic distance is very small compared with the experiment, and the model accuracy is very high. The RMSD of the overall structure main chain is: The RMSD of the main chain in the key transmembrane helix region is: Figure 6 is the prediction result of five new GPCR structures, in which the dark model is the experimental structure, and the light model is the prediction result of the present application. It can be seen from Figure 6 that the relevant prediction model is completely consistent with the experimental structure after superposition.

[0122] Figure 7 is the comparison of the prediction results of the present application with the prediction results of AlphaFold2 and RosettaFold. Download from the official model data website of AlphaFold2 (https: / / alphafold.ebi.ac.uk) and the official website of RosettaFold GPCR model (http: / / www.rosettaGPCR.org), and compare with the experimental structure. By comparing the RMSD of the relevant models, it is found that the prediction result of the present application for the structure of five new GPCR target points is 60% more accurate than AlphaFold2, and 100% more accurate than RosettaFold.

[0123] In addition, the AI model of the present application also accurately predicts the binding model of drug molecules and GPCRs. The prediction of KOR receptors is currently very challenging because the drug small molecules are very large. After verification, the present application can successfully capture all important interactions of the complex drug molecules, and the prediction results are highly consistent with the experimental structure, see Figure 8 , wherein all key interactions in the experimental structure (left) are captured by the prediction of the present application (right).

[0124] To further verify the reliability of the present application, a new Class C GPCR receptor is predicted, see Figure 9 , the three-dimensional folding prediction result of the orphan receptor GPR158 is shown, wherein the dark model represents the experimental structure, the light model represents the predicted structure, the left corresponds to the present application, the middle corresponds to AlphaFold2, and the right corresponds to RosettaFold. Experimental results show that the RMSD of the predicted model of the present application and the experimental structure is only ; while the RMSD predicted by AlphaFold2 and RosettaFold is and , which is very different from the experimental structure, and even the secondary structure is predicted to be wrong.

[0125] In summary, the present application provides a new method for predicting the three-dimensional folding of the most important drug target G protein-coupled receptor protein (GPCR) and the binding mode of GPCR drug molecules. Experimental results show that the prediction accuracy of the present application is more accurate than that of Google AlphaFold2, and it can also accurately capture the interaction mode of the related drug molecules and GPCR targets, thereby providing a solid foundation for structure-based drug design of GPCR targets.

[0126] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions embodied therewith, which instructions are used to program computers to implement the various aspects of the present application.

[0127] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0128] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0129] Computer readable program instructions for carrying out operations of the present application can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.

[0130] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0131] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0132] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0133] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0134] Embodiments of the present application have been described above, and the description is intended to be illustrative, and not restrictive, of the disclosed embodiments. Many modifications and variations of the disclosed embodiments are possible in light of the above teachings. It is therefore to be understood that within the scope of the disclosed embodiments, modifications and variations of the disclosed embodiments can be practiced. It is also to be understood that the specific order or hierarchy of steps in the processes disclosed is an illustration of exemplary processes. Based upon the description and illustrations provided herein, those skilled in the art will understand that changes can be made to the order of steps in the processes and that many of the individual steps can be modified or eliminated. Additionally, the description and illustrations provided herein are not meant to limit the scope of the disclosed embodiments. The scope of the disclosed embodiments is limited only by the claims.

Claims

1. A method for predicting the three-dimensional folding of a G protein-coupled receptor protein and a drug molecule binding model, comprising the following steps: selecting a plurality of target GPCR initial models using three-dimensional sequence alignment and template three-dimensional structure optimization, and obtaining a target GPCR model for molecular docking and drug design by optimizing the flexible regions of the GPCR initial model and the drug structure site; selecting the best GPCR-drug molecule binding mode by cross- validating the prediction performance of a plurality of types of GPCR-drug molecule artificial intelligence models; wherein the input of the GPCR-drug molecule artificial intelligence model is a three-dimensional interaction fingerprint feature of a ligand in a receptor, and each column represents the number of interaction bonds between the corresponding ligand and the amino acid in the current column, and the output is the root mean square deviation (RMSD) value, which is used to measure the distance between the predicted 3D conformation and the real conformation; wherein the target GPCR model for molecular docking and drug design is obtained according to the following steps: selecting an experimental structure as a homologous template for computer modeling according to homology, and determining the primary sequence of the target GPCR amino acid and predicting its secondary structure by performing three-dimensional sequence alignment between the primary sequence of the target GPCR and the homologous experimental structure, and then obtaining a plurality of target GPCR initial models; evaluating the structural stability of the GPCR initial model by a CHARMM molecular force field energy scoring function, and selecting the lowest energy model as the initial optimization model; performing all-atom molecular dynamics simulation on the initial optimization model, and performing cluster analysis on the conformation within a set time range during the simulation to select representative structures as the target GPCR model for molecular docking and drug design, wherein the molecular force field of the simulation is CHARMM or Amber force field; wherein the GPCR-drug molecule artificial intelligence model is constructed according to the following steps: obtaining experimental structures of GPCR and its drug molecule complexes from public databases, as well as molecular structure information and biological activity information of related GPCR targets, and selecting a plurality of sample compounds based on the biological activity information; calculating the GPCR-drug molecule interaction fingerprint for the sample compounds to obtain the molecular interaction information between the receptor and the ligand; based on the molecular interaction information of the receptor and the ligand, establishing the GPCR-drug molecule artificial intelligence model to determine the input and output of the model.

2. The method of claim 1, wherein, The GPCR-drug molecule artificial intelligence model includes one or more of ridge regression model, Lasso model, support vector machine model, decision tree model, random forest model, K nearest neighbor model and XGBoost model.

3. The method of claim 1, wherein, The input features of the GPCR-drug molecule artificial intelligence model are determined according to the following steps: performing preliminary feature screening and dimensionality reduction on the molecular fingerprint features generated by a plurality of binders to obtain dimensionally reduced molecular fingerprint feature data performing data normalization on the dimensionally reduced molecular fingerprint feature data, calculating the standard deviation of each column feature, and removing the column of features with a standard deviation less than a set standard deviation threshold to obtain the first screened feature data; For the feature data after the first screening, the correlation coefficient between the features is calculated, and a column with a correlation coefficient greater than a set correlation coefficient threshold is removed to obtain the second screening feature data, wherein the correlation coefficient is used to represent the correlation degree between each two features, and the value interval is between [-1, 1]; For the second screening feature data, after Z-score standardization processing, it is used as the input feature of the GPCR-drug molecule artificial intelligence model.

4. The method of claim 3, wherein, The correlation coefficient threshold is set to 0.9, and the correlation coefficient is Pearson correlation coefficient, and the calculation method is: wherein, is the mean of the two columns of fingerprints, is the mean of the column of fingerprints, is the standard deviation of the column of fingerprints, is the expected value.

5. The method of claim 1, wherein, During the cross-validation process of the GPCR-drug molecule artificial intelligence model, the following indexes are used to evaluate the prediction performance of the model: Determination coefficient, indicating the percentage of the square of the correlation degree between the predicted value and the actual value of the target variable; Mean absolute error, indicating the average value of the distance between the predicted value and the real sample value of the model; Mean square error, indicating the average value of the square of the distance between the predicted value and the real sample value of the model.

6. The method of claim 1, wherein, For the GPCR-drug molecule interaction fingerprint, the molecular interaction between the receptor and the ligand is characterized by the interaction type, the amino acid involved in the interaction, the atoms of the amino acid and the ligand involved, and the three-dimensional position of the atoms, wherein the interaction type includes: hydrophobic contact, hydrogen bond, water bridge, salt bridge, π stacking interaction, π-cation interaction and halogen bond.

7. A computer readable storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to realize the steps of the method according to any one of claims 1 to 6.

8. A computer device comprising a memory and a processor, having stored on the memory a computer program capable of running on the processor, characterized in that, The processor executes the computer program to realize the steps of the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Protein-ligand interaction fingerprint spectrum-based drug target prediction method

    CN107038348A

  • Screening method of compound of targeted G protein coupled receptor

    CN113270153A