Method for predicting drug discovery target proteins, and system for predicting drug discovery target proteins

The method and system leverage clinical data and machine learning to predict drug target proteins by analyzing drugs with low disease likelihood and binding proteins, addressing the challenge of identifying target proteins for diverse diseases.

JP7792680B2Active Publication Date: 2025-12-26NAT UNIV CORP TOKAI NAT HIGHER EDUCATION & RES SYST
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021170944
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-22
Filing Date
2021-10-19
Publication Date
2025-12-26
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

Identifying drug target proteins is challenging due to the vast number of proteins in the body with unclear functions, and many diseases lack clear therapeutic drugs, necessitating new methods for predicting drug target proteins.

Method used

A method and system that utilize clinical data to calculate the likelihood of disease occurrence, analyze drugs with low likelihood as disease preventive drugs, and predict binding proteins using interaction, gene expression profile, and chemical structure data to identify drug discovery target proteins.

Benefits of technology

Enables the prediction of drug target proteins for various diseases, including those without established treatments, by integrating clinical data and machine learning models to identify proteins that bind to disease-preventive drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007792680000006
    Figure 0007792680000006
  • Figure 0007792680000007
    Figure 0007792680000007
  • Figure 0007792680000008
    Figure 0007792680000008
Patent Text Reader

Abstract

To provide a predictive method and a predictive system for predicting the drug discovery target protein of disease for the therapeutic purpose.SOLUTION: In a method for predicting the drug discovery target protein of disease for the therapeutic purpose, the disease for therapeutic purpose is selected, in which the drug target protein is desired to be predicted, for that disease, an analysis step S11 is performed, in which the clinical data analysis is first performed, the drug having the effective possibility is identified, and a prediction step S21 is performed for this drug, in which the binding protein prediction for the compound is performed. Then, the predicted drug discovery target protein is displayed on a monitor or the like as appropriate S31.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and a system for predicting target proteins for drug discovery. [Background technology]

[0002] Drug targets are biomolecules such as proteins that can lead to the treatment of diseases. Drugs are designed to bind to target biomolecules and control them by inhibiting or activating them. Selecting an inappropriate target reduces the success rate of drug development, so identifying target molecules is an important challenge in drug development.

[0003] Patent Document 1 discloses that a dataset of protein-protein interactions, which includes attributes of the three-dimensional structure of the protein-protein interaction, attributes of existing drugs / compounds that act on each protein that constitutes the protein-protein interaction, and attributes of the biological functions of each protein that constitutes the protein-protein interaction, is used as a positive example and a negative example, and machine learning is performed to construct a mathematical model that predicts protein-protein interactions that have the potential to be drug targets.

[0004] Patent Document 2 provides a method for quickly and appropriately selecting candidate drugs effective against a disease from among existing drugs. Patent Document 2 discloses the following: A drug search device 1 is provided with a drug response DB 13 that stores gene expression level data for multiple known drugs. Expression level data for a sample and a control is input, and variation data representing the difference in expression level between the sample and the control is calculated. A network structure is estimated in which the multiple known drugs and the disease are nodes, using the expression level data for the multiple known drugs read from the drug response DB 13 and the disease variation data as parameters. A network clustering is performed on the multiple known drugs and the disease based on the network structure. A drug having an inverse correlation with the disease is selected from drugs classified into the same cluster as the disease, and data on the selected drug is output. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-165230 [Patent Document 2] Japanese Patent Application Laid-Open No. 2016-099674 Summary of the Invention [Problem to be solved by the invention]

[0006] Identifying drug target proteins, which are proteins that can be targeted by drug discovery, is important in pharmaceutical development. However, the number of proteins in the body is enormous, and some of their functions are unclear. In addition, proteins whose relationships are not immediately clear may be closely related to tissues or organs of the body. Furthermore, there are many intractable diseases of unknown cause for which no therapeutic drugs or treatment methods have been established. In response to these issues, new methods for predicting drug target proteins are needed in addition to the construction of mathematical models in Patent Document 1 and the drug discovery in Patent Document 2. Under such circumstances, an object of the present invention is to predict drug discovery target proteins for diseases to be treated. [Means for solving the problem]

[0007] The present inventors have conducted extensive research to solve the above problems and have found that the following inventions meet the above objectives, thereby completing the present invention.

[0008] <1> A method for predicting a drug discovery target protein for a disease to be treated, comprising: a step of calculating the likelihood of a disease to be treated when a drug is administered from clinical data, and analyzing a drug that has a low likelihood of causing the disease as a disease preventive drug; and predicting a binding protein to the compound of the disease preventive drug based on a prediction score using one or more data selected from the group consisting of interaction data on compounds and proteins, gene expression profile data, and chemical structure data, The method for predicting a drug discovery target protein includes predicting the binding protein as a drug discovery target protein for the disease. <2> The step of predicting predicts the binding protein as one or more prediction scores selected from a group of prediction scores consisting of the following (1) to (3): <1> The prediction method described in (1) Using the chemical structure data and the interaction data, a predicted compound is selected from registered compounds known to bind to a protein, and the predicted compound is a registered compound having a high similarity in chemical structure to the compound of the disease preventive drug. (2) Using the gene expression profile data and the interaction data, a predicted compound is selected from registered compounds known to bind to a protein, and the predicted compound is a registered compound having a gene expression profile highly similar to that of the compound of the disease preventive drug. (3) A prediction score for predicting the binding protein of the disease preventive drug using a machine learning prediction model for compound-protein interactions. <3> In the predicting step, two or more prediction scores are predicted from the group of prediction scores, and an integrated score is calculated to rank the plurality of candidates for the binding protein. <2> The prediction method described in <4> The clinical data is data collected in a drug adverse event reporting system. <1> ~ <3> 1. A prediction method according to any one of the preceding claims. <5> In the analyzing step, two or more disease preventive drugs are extracted from the candidate disease preventive drugs; In the prediction step, binding protein candidates for each of the plurality of disease preventive drugs are extracted, and the binding protein candidates for the plurality of disease preventive drugs are combined to predict a drug discovery target protein. <1> ~ <4> 1. A prediction method according to any one of the preceding claims. <6> A system for predicting a drug discovery target protein for a disease to be treated, comprising: an analysis unit that calculates the likelihood of disease occurring when a drug is administered based on clinical data and analyzes drugs that have a low likelihood of disease occurrence as disease preventive drugs; a prediction unit that predicts a binding protein to the compound of the disease preventive drug based on a prediction score using one or more data selected from the group consisting of interaction data, gene expression profile data, and chemical structure data regarding compounds and proteins, A drug discovery target protein prediction system that predicts the binding protein as a drug discovery target protein for the disease. [Effects of the Invention]

[0009] According to the present invention, it is possible to predict target proteins for drug discovery for diseases to be treated. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a flow diagram of an embodiment of a prediction method of the present invention. [Figure 2] FIG. 10 is a flow diagram of another embodiment of the prediction method of the present invention. [Figure 3] 1 is a schematic diagram of an embodiment of a prediction system of the present invention; [Figure 4] FIG. 10 is a diagram for explaining an example of a prediction process of the present invention. [Figure 5] FIG. 2 is a diagram for explaining a part of a process according to an example of a prediction process of the present invention. [Figure 6] FIG. 2 is a diagram for explaining a part of a process according to an example of a prediction process of the present invention. [Figure 7] FIG. 2 is a diagram for explaining a part of a process according to an example of a prediction process of the present invention. [Figure 8] FIG. 10 is a diagram for explaining a part of a process according to another example of the prediction process of the present invention. [Figure 9] FIG. 10 is a diagram for explaining a part of a process according to another example of the prediction process of the present invention. [Figure 10]FIG. 10 is a diagram for explaining a part of a process according to another example of the prediction process of the present invention. [Figure 11] FIG. 10 is a diagram for explaining another example of the prediction process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0011] The following describes in detail an embodiment of the present invention, but the following description of the constituent elements is one example (typical example) of an embodiment of the present invention, and the present invention is not limited to the following content unless the gist of the present invention is changed. Note that when the expression "to" is used in this specification, it is used as an expression that includes the numerical values ​​before and after it.

[0012] [Prediction method of the present invention] The method for predicting a drug target protein of the present invention is a method for predicting a drug target protein for a disease to be treated, and includes the steps of: calculating, from clinical data, the likelihood of the disease to be treated occurring when a drug is administered; analyzing drugs with a low likelihood of causing the disease as disease preventive drugs; and predicting binding proteins for the compound of the disease preventive drug based on a prediction score using one or more data selected from the group consisting of interaction data, gene expression profile data, and chemical structure data regarding compounds and proteins, and predicting the binding proteins as drug target proteins for the disease. In this application, the method for predicting a drug target protein of the present invention may also be simply referred to as the prediction method of the present invention.

[0013] [Prediction system of the present invention] The drug discovery target protein prediction system of the present invention is a system for predicting drug discovery target proteins for a disease to be treated, and includes an analysis unit that calculates the likelihood of disease onset when a drug is administered from clinical data and analyzes drugs with a low likelihood of disease onset as disease preventive drugs, and a prediction unit that predicts binding proteins for the compound of the disease preventive drug based on a prediction score using one or more data selected from the group consisting of interaction data, gene expression profile data, and chemical structure data regarding compounds and proteins, and predicts the binding proteins as drug discovery target proteins for the disease. In the present application, the drug discovery target protein prediction system of the present invention may also be simply referred to as the prediction system of the present invention.

[0014] According to the prediction method and prediction system of the present invention, it is possible to predict target proteins for drug discovery. Note that in the present application, the prediction method of the present invention can also be performed using the prediction system of the present invention, and the corresponding configurations in the present application can be mutually utilized.

[0015] In studying the prediction of drug discovery target proteins, the inventors considered combining the use of clinical data with technology to predict compound binding proteins. Clinical data is clinical data on existing drugs, etc. Analysis of clinical data can reveal that existing drugs have therapeutic effects on diseases other than those targeted by the drug. Meanwhile, technology to identify compounds similar to a compound is being investigated.

[0016] The present inventors discovered that by combining these pieces of information, it may be possible to predict proteins that bind to a drug that has an effect on a certain disease, based on the structure of the drug compound and omics data, and thus arrived at the present invention. This method predicts the drug target protein itself for a disease from information on the binding protein. This has the advantage that it can identify drug target proteins that are difficult to predict based on conventional knowledge, and can also be used to develop new drugs targeting those drug target proteins.

[0017] [Flowchart for predicting drug target proteins] Fig. 1 is a flow diagram of an embodiment of the prediction method of the present invention. A disease for which a therapeutic target protein is to be predicted is selected, and for that disease, an analysis step S11 is first performed to analyze clinical data and identify potentially effective drugs. For these drugs, a prediction step S21 is performed to predict binding proteins for the compound. The predicted drug target protein is then displayed on a monitor or the like as appropriate in step S31.

[0018] FIG. 2 is a flow chart of another embodiment of the prediction method of the present invention. FIG. 2 shows a more detailed example of the flow chart. First, a disease for which a drug discovery target protein is to be predicted is input. Next, an odds ratio, which indicates the likelihood of the disease occurring when a drug is administered, is calculated using clinical data analysis. Drugs with an odds ratio of <1.0, which indicates a significantly low likelihood of the disease, are selected. Drugs with a low likelihood of the disease occurring are extracted as disease preventive drugs that are expected to have a preventive effect.

[0019] Next, in the binding protein prediction, interaction data, omics data, and chemical structure data are used to make a prediction using a prediction score based on the chemical structure and omics data. By performing this prediction using the prediction score for multiple drugs, binding proteins common to these drugs are selected. As a result, proteins predicted to have a high probability of being binding proteins can be predicted as drug discovery target proteins.

[0020] [Drug target protein prediction system] Figure 3 is a schematic diagram of an embodiment of the prediction system of the present invention. The prediction system 1 inputs a disease for which a drug discovery target protein is to be predicted into an input unit 2. It has a control unit 4 that performs analysis in an analysis unit that determines disease preventive drugs for this prediction, and prediction in a prediction unit that determines binding proteins. This control unit 4 uses data from a data unit 3 that collects data on the disease to be treated and data for analysis and prediction. The prediction system 1 also has a memory 5 that stores programs for analysis and prediction, analysis results, prediction results, etc. It also has a display unit 6 that displays information related to its control and the prediction results.

[0021] The data unit 3 can use data such as clinical data 31, interaction data 32, omics data 33, and chemical structure data 34 obtained, set, or collected from external databases as appropriate. Furthermore, based on the output of the display unit 6, drug discovery target proteins and their prediction scores can be displayed on an external monitor 61 or the like. These are configured by appropriately adopting a supercomputer, monitor, etc.

[0022] [Diseases for treatment] The prediction method of the present invention is a method for predicting drug discovery target proteins for a disease to be treated. The disease to be treated is a disease selected as a target for treatment by the prediction method of the present invention. The disease to be treated can be a disease for which a therapeutic drug exists, a disease for which a therapeutic drug does not exist, a disease for which a therapeutic drug is difficult to obtain, a disease for which a therapeutic drug is difficult to produce, an intractable disease that is difficult to treat, a disease with a small number of patients, etc.

[0023] [Analysis process] The analyzing step (analysis step) is a step of calculating the likelihood of disease occurring when a drug is administered from clinical data, and analyzing drugs (compounds) that have a low likelihood of disease occurrence as disease preventive drugs.

[0024] [Clinical data] Clinical data is data collected on drugs and their adverse drug reactions. Data collected in adverse drug event reporting systems can be used as clinical data. For example, databases such as FAERS and JADER are publicly available and can be used as clinical data. These databases can be used as is, but if they contain noise, such as obvious clerical errors or inconsistencies in expression, they can be processed to reduce or remove the noise before being used as clinical data in the analysis process. Furthermore, for diseases targeted for treatment, rather than simply analyzing them by disease name, data included in the clinical data can be filtered for factors such as age and gender, which affect the efficacy of drugs and the appropriateness of treatment depending on the patient.

[0025] FAERS (FDA Adverse Events Reporting System) is a database made public by the US Food and Drug Administration (FDA). FAERS contains information such as DEMO, DRUG, REAC, OUTC, RPSR, INDI, and THER. DEMO is basic patient information such as gender, age, date of onset of adverse event, and country where the adverse event occurred. DRUG is information such as drug name, administration route, and dosage. REAC is the name of the adverse event. OUTC is the case outcome. RPSR is the adverse event information source. INDI is the indication. THER is information such as the start date of administration, end date of administration, and treatment period. FAERS has over 16,000 registered drugs and approximately 40 million registered adverse events, and is updated regularly.

[0026] "JADER" (Japanese Adverse Drug Event Report database) is a database published by the Pharmaceuticals and Medical Devices Agency (PMDA). JADER has over 3,000 registered drugs and approximately 10 million registered adverse events, and is updated regularly.

[0027] [Odds ratio] In the analysis step, the likelihood of a disease occurring when a drug is administered is calculated from the clinical data. This likelihood of a disease occurring is also called the odds ratio. Clinical data from medication can be classified into the following four types for the new disease to be treated. In this explanation, an existing drug will be referred to as existing drug α, and the disease to be treated will be referred to as disease β. Existing drug α is clinical data for a condition in which there is no existing drug that treats disease β. Category A: Patients taking existing drug α and with disease β Category B: Patients taking existing drug α and not with disease β Category C: Patients with disease β who are not taking existing drug α Category D: Patients who are not taking existing drug α and are not patients with disease β

[0028] Based on this classification, the odds ratio can be expressed by the following formula: Odds ratio = "Class A x Class D" / "Class B x Class C" That is, it shows the ratio of the group of patients who took the drug (category A) and the group of patients who did not take the drug (category D) to the group of patients who did not take the drug but were not patients (category B) and the group of patients who did not take the drug (category C). If this odds ratio is high, it indicates that disease β occurs even when taking existing drug α, and that disease β is unlikely to occur even when not taking the drug, and it is thought that existing drug α is not expected to be a preventive or therapeutic drug. If this odds ratio is low, it indicates that people who take existing drug α are unlikely to develop disease β, and those who do not take it are likely to develop disease β, and it is thought that existing drug α is expected to be a preventive or therapeutic drug. An odds ratio can be determined to be significantly low by a low p-value (e.g., p<0.05) in a statistical test.

[0029] Using this odds ratio, the likelihood of disease occurrence when a drug is administered can be calculated, and an analysis can be performed to identify drugs with low odds ratios as disease preventive drugs. One or more disease preventive drugs can be analyzed to extract them as candidates, and multiple disease preventive drug candidates, such as two or more, three or more, or four or more, may be identified. The upper limit on the number of disease preventive drug candidates does not need to be particularly set, but may be set to 100 or less, 50 or less, 30 or less, etc., taking into consideration the efficiency of data processing and the influence of the prediction score of the drug discovery target protein.

[0030] Regarding odds ratios, please refer to Reference 1, "Horinouchi et al, Renoprotective effects of a factor Xa inhibitor: fusion of basic research and a database analysis, Sci Rep. 2018; 8: 10858."

[0031] [Disease prevention drug] In the analysis process, drugs with a low likelihood of causing disease are extracted from clinical data as disease preventive drugs. Disease preventive drugs are drugs that are thought to potentially prevent the disease they are intended to treat. Among these drugs, those whose active ingredients or main component compounds have been identified are targeted for prediction of binding proteins.

[0032] [Prediction Process] The prediction step (prediction step) is a step of predicting binding proteins for the compound of the disease preventive drug based on a prediction score using one or more data selected from the group consisting of interaction data, gene expression profile data, and chemical structure data regarding the compound and the protein.

[0033] [Binding of compounds to proteins] Compounds bind to proteins, and the compounds affect the function of the bound protein, activating or inhibiting the protein. In this application, a protein to which a compound binds is called a binding protein. The learning data used for machine learning in the prediction step can include interaction data, gene expression profile data, chemical structure data, and other data. These are then subjected to machine learning to predict proteins that bind to a drug designated as a disease preventive drug, focusing on its chemical structure.

[0034] [Interaction Data] The interaction data is data regarding the interaction between a compound and a protein. Examples of databases that contain this data include ChEMBL, MATADOR, Drug Bank, PDSP-Ki, KEGG DRUG, BindingDB, and Therapeutic Target Database. One or more databases from the group consisting of these may be used, and a plurality of databases may also be used. Databases other than those mentioned above may also be used for the data regarding the interaction between a compound and a protein.

[0035] ChEMBL is a database of drug-like biologically active small molecules provided by EBI.

[0036] MATADOR (Manually Annotated Targets and Drugs Online Resource) is a database of protein chemical functions.

[0037] Drug Bank, developed by the University of Alberta, is a database that collects and organizes information on drugs, including FDA-approved drugs and drugs under clinical trials, and their target proteins.

[0038] The Psychoactive Drug Screening Program Ki Database (PDSP-Ki) is a public domain resource that provides information on the interactions between drugs and molecular targets.

[0039] KEGG DRUG is a database that centrally aggregates pharmaceutical information from Japan, the United States, and Europe in terms of chemical structure and ingredients.

[0040] BindingDB is a database of measured binding affinities, focusing primarily on interactions between drug-like molecules and proteins that are potential drug targets.

[0041] The Therapeutic Target Database (TTD) is a database that provides information on known therapeutic protein and nucleic acid targets.

[0042] [Gene expression profile data] Gene expression profile data is related to omics data (omics information). Omics data is comprehensive information about biomolecules, specifically a compilation of various comprehensive molecular information called genomes, transcriptomes, proteomes, metabolomes, interactomes, and cellomes. Gene expression profile data can be used to identify disease preventive drugs and compounds with known binding proteins that share similar gene expression profiles.

[0043] [Chemical structure data] Chemical structure data is data for identifying the characteristics of a chemical structure based on the chemical structure description formula of a compound. If a disease preventive drug compound has a similar chemical structure to a known ligand, it is possible to identify a target protein that has an acceptor corresponding to that ligand. Therefore, chemical structure data can be used to identify compounds registered as known ligands that have a similar chemical structure to a disease preventive drug compound.

[0044] [Prediction Score] The predicting step predicts binding proteins for the disease preventive drug compound based on a prediction score using one or more data selected from the group consisting of interaction data, gene expression profile data, and chemical structure data regarding these compounds and proteins.

[0045] The prediction score is a score that serves as an index of the possibility of a binding protein. A high prediction score can be considered to be a highly likely binding protein, and a low prediction score can be considered to be a low possibility of a binding protein. The prediction score may be a value that is easy to calculate depending on the method for calculating the prediction score, or the predicted results for a large number of compound-protein combinations may be normalized to make them easier to compare. Using such a prediction score, the binding protein is predicted as a drug discovery target protein for the disease.

[0046] The prediction score in the prediction step can be, for example, a group of prediction scores consisting of the following (1) to (3). Binding proteins can be predicted as one or more prediction scores selected from this group of prediction scores. Note that there may be multiple methods for determining prediction scores belonging to each of (1) to (3). Therefore, multiple prediction scores belonging to (1) can be used, or multiple prediction scores belonging to (2), or multiple prediction scores belonging to (3).

[0047] (1) A prediction score is calculated from registered compounds known to bind to proteins using chemical structure and interaction data, taking into account the similarity of their chemical structure to disease preventive drug compounds.

[0048] The predicted score according to (1) can be used to calculate the similarity using a descriptor such as KCF-S.

[0049] For calculating this prediction score, for example, the following literature on similarity search and KCF-S descriptors can be referenced:

[0050] ·Reference 2-1 "Kotera et al., 2013, BMC Syst. Biol. Kotera, M., Tabei, Y., Yamanishi, Y., Moriya, Y., Tokimatsu, T., Kanehisa, M., and Goto, S.,"KCF-S: KEGG Chemical Function and Substructure for improved interpretability and prediction in chemical bioinformatics", BMC Systems Biology, 7(Suppl 6):S2, 2013” ·Reference 2-2 "Sawada, R., Iwata, M., Umezaki, M., Usui, Y., Kobayashi, T., Kubono, T., Hayashi, S., Kadowaki, M., and Yamanishi, Y., "KampoDB, database of predicted targets and functional annotations of natural medicines", Scientific Reports, 8:11216, 2018." References 2-3 include "Tabei, Y., Kishimoto, A., Kotera, M., and Yamanishi, Y., "Succinct Interval Splitting Tree for Scalable Similarity Search of Compound-Protein Pairs with Property Constraints," Proceedings of the 19th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD2013), pp. 176-184, ACM New York, NY, USA, 2013."

[0051] (2) A prediction score is calculated from registered compounds known to bind to proteins using gene expression profile data and interaction data, taking into account the similarity of the gene expression profile of the disease preventive drug compound.

[0052] For the prediction score related to (2), for example, the following literature on gene expression profile data and omics data can be referenced.

[0053] ·Reference 3-1 "Lamb, J. et al., "The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease". Science, 313, 1929-1935, 2006" ·Reference 3-2 "Subramanian, A. et al., "A next generation connectivity Map: L1000 platform and the first 1,000,000 profiles". Cell, 171, 1437-1452, 2017" Reference 3-3 "Iwata, M. et al., "Pathway-based drug repositioning for cancers: computational prediction and experimental validation", Journal of Medicinal Chemistry, 61(21), 9583-9595, 2018" can be referred to.

[0054] Furthermore, when determining the prediction score in (2), data obtained by evaluating the gene expression profile of a disease preventive drug may be used as appropriate, or if known gene expression profile data is available, that data may be used.

[0055] (3) A prediction score for predicting the binding protein of the disease preventive drug using a machine learning prediction model for compound-protein interactions.

[0056] The prediction score according to (3) uses a prediction model that has undergone machine learning to determine whether or not there is an interaction between a compound and a protein, such as binding. Various types of machine learning can be used for this, but more specifically, the following methods can be used:

[0057] (3A) The prediction score of a binding protein prediction model that predicts binding proteins of disease preventive drugs using a machine learning prediction model that separates interacting pairs and interacting pairs of compounds and proteins as training data.

[0058] The prediction score according to (3A) can be calculated, for example, by pairwise kernel regression or support vector machine within the framework of the kernel method. The following documents can be referred to for predictions based on the chemical structure of a compound according to this prediction score or on the gene expression pattern of a compound in human-derived cells. ·Reference 4-1 "Yamanishi, Y.,"Supervised Bipartite Graph Inference", Advances in Neural Information Processing Systems 21 (Koller, D., Schuurmans, D., Bengio, Y. and Bottou, L. eds.), 1841-1848, MIT Press, Cambridge, MA, 2009." ·Reference 4-2 "Bleakley, K. and Yamanishi, Y., "Supervised prediction of drug-target interactions using bipartite local models", Bioinformatics, 25, 2397-2403, 2009." ·Reference 4-3 “Yamanishi et al, Bioinformatics, 2008” ·Reference 4-4 “Yamanishi, Adv Neural Inf Process Syst, 2009” ·Reference 4-5 “Bleakley et al., Bioinformatics, 2009” ·Reference 4-6 “Tabei et al, BMC Systems Biology, 2013” ·Reference 4-7 “Hizukuri et al, BMC Med Genomics, 2015” ·Reference 4-8 “Iwata et al, Sci rep, 2017; Sawada et al, Sci rep, 2018”

[0059] (3B) A prediction score that predicts the binding proteins of disease preventive drugs using machine learning to classify the presence or absence of interactions between compounds and proteins in the space of compound structures.

[0060] The prediction score according to (3B) is based on sparse modeling and can be calculated by, for example, L1 regularized logistic regression. The following literature can be referred to as an example of a method using such logistic regression:

[0061] ·Reference 5-1 "Tabei, Y., Kotera, M., Sawada, R., and Yamanishi, Y.,"Network-based characterization of drug-protein interaction signatures with a space-efficient approach", BMC Systems Biology, 13(Suppl 2):39, 2019.” ·Reference 5-2 “Tabei, Y., Pauwels, E., Stoven, V., Takemoto, K., and Yamanishi, Y., "Identification of chemogenomic features from drug-target interaction networks using interpretable classifiers", Bioinformatics, 28, i487-i494, 2012.

[0062] (3C) A prediction score for predicting the binding proteins of disease preventive drugs using a prediction model that uses machine learning of binding proteins based on chemical structure data and interaction data from registered compounds known to bind to proteins using a graph convolutional neural network based on chemical structure.

[0063] The prediction score for (3C) is based on a graph convolutional neural network, a type of deep learning model, and the following literature can be referenced:

[0064] ·Reference 6-1 "Fukunaga, I., Sawada, R., Shibata, T., Kaitoh, K., Sakai, Y., and Yamanishi, Y.,"Prediction of the Health Effects of Food Peptides and Elucidation of the Mode-of-action Using Multi-task Graph Convolutional Neural Network", Molecular Informatics, 39(1-2):e1900134, 2020.” ·Reference 6-2 “Altae-Tran et al., 2017, ACS Cent. Sci., 3, 4, 283-293”

[0065] (3) For the prediction score, the following literature on binding protein prediction based on chemogenomics methods can be referenced.

[0066] ·Reference 7-1 "Yamanishi, Y., Kotera, M., Moriya, Y., Sawada, R., Kanehisa, M., and Goto, S. "DINIES: drug-target interaction network inference engine based on supervised analysis", Nucleic Acids Research, 42, W39-W45, 2014" ·Reference 7-2 "Yamanishi, Y., Araki, M., Gutteridge, A., Honda, W., and Kanehisa, M., "Prediction of drug-target interaction networks from the integration of chemical and genomic spaces", Bioinformatics, 24, i232-i240, 2008."

[0067] (3) For the prediction score, the following literature on binding protein prediction based on transcriptomics methods can be referenced.

[0068] ·Reference 8-1 "Sawada, R., Iwata, M., Tabei, Y., Yamato, H., and Yamanishi, Y., "Predicting inhibitory and activatory drug targets by chemically and genetically perturbed transcriptome signatures", Scientific Reports, 8:156, 2018." ·Reference 8-2 "Iwata, M., Sawada, R., Iwata, H., Kotera, M., and Yamanishi, Y., "Elucidating the modes of action for bioactive compounds in a cell-specific manner by large-scale chemically-induced transcriptomics", Scientific Reports, 7:40164, 2017." ·Reference 8-3 "Hizukuri, Y., Sawada, R., and Yamanishi, Y., "Predicting target proteins for drug candidate compounds based on drug-induced gene expression data in a chemical structure-independent manner", BMC Medical Genomics, 8:82 (10 pages), 2015."

[0069] (3) For the prediction score, the following literature on binding protein prediction based on phenomics methods can be referenced.

[0070] ·Reference 9-1 Takarabe, M., Kotera, M., Nishimura, Y., Goto, S., and Yamanishi, Y., "Drug target prediction using adverse event report systems: a pharmacogenomic approach", Bioinformatics, 28, i611-i618, 2012. ·Reference 9-2 "Yamanishi, Y., Kotera, M., Kanehisa, M., and Goto, S., "Drug-target interaction prediction from chemical, genomic and pharmacological data in an integrated framework", Bioinformatics, 26, i246-i254, 2010."

[0071] FIG. 4 is a diagram illustrating an example of the prediction process of the present invention. FIG. 4 particularly illustrates the flow of calculating the prediction score (1) and the prediction score (2) of the prediction score group. Based on the compound data of the drug (medicine), similar compounds related to the compound structure data are identified, as shown in the upper part. This particularly relates to the prediction score (1) above. Similar compounds using this compound data can be compared with compound-protein interaction (binding) data, and those with known interactions with proteins can be used. For example, the similarity between a disease preventive drug compound and a similar compound can be used as the prediction score.

[0072] Similarly, the bottom panel of Figure 4 identifies similar compounds based on omics data such as gene expression profiles. This is particularly relevant to the prediction score mentioned in (2) above. Similar compounds using this omics data can be compared with compound-protein interaction (binding) data to identify compounds with known protein interactions. For example, the similarity between a disease preventive drug compound and a similar compound based on omics data can be used as a prediction score.

[0073] Figures 5, 6, and 7 are diagrams illustrating some of the steps involved in an example of the prediction process of the present invention. The prediction score (1) can be calculated using Figures 5 to 7, and can be performed with reference to KCF-S (reference: Kotera et al., 2013, BMC Syst. Biol.). Other descriptors or fingerprints, such as ECFP and DRAGON, may also be used. Feature vectors generated by a graph convolutional neural network may also be used.

[0074] As shown in Figure 5, in order to find compounds with a high degree of similarity to the chemical structure of a disease preventive drug, the presence or absence and number of predetermined partial structures are identified. There can be a large number of predetermined partial structures, such as 500,000, but in Figure 5, for example, a compound is identified as having one aromatic carboxylic acid-like structure on the far left, one ester-like structure second from the left, two carboxylic acid-like structures third from the left, no pyrazolidine-like structure second from the right (0), and three ethoxy-like structures on the far right.

[0075] Next, based on the results of identifying the partial structure of the disease preventive drug, compounds that are similar in terms of the presence or absence of partial structures and the number of such structures are identified. The similarity is determined by the method shown in Figure 6. Here, an example is shown in which the similarity between compound X and compound Y is calculated. The partial structures of compound X and compound Y are quantified and compared to quantify the similarity. The calculated similarity is 1 if they are a perfect match, and 0 if they are completely different. The closer the value is to 1, the higher the similarity. The compound with the highest similarity may be extracted, or multiple compounds may be extracted in descending order of similarity.

[0076] Figure 7 shows an example of quantifying the predicted score of binding proteins by comparing compounds with high chemical structure similarity, as shown in Figures 5 and 6, with interaction data. When a preventive drug is input, similar compounds are extracted in descending order of similarity to the compound of the preventive drug, along with their compound similarity. These similar compounds are then identified as binding proteins in the interaction data. The similarity of each similar compound is then used as the score for the binding protein. For example, a binding protein of a similar compound with a similarity of 0.91 can be assigned a predicted score of 0.91 for the drug discovery target protein prediction (Estimated Target Protein). Similarly, a binding protein of a similar compound with a similarity of 0.80 can be assigned a predicted score of 0.80. A similarity of 0.77 can be assigned a predicted score of 0.77.

[0077] [Machine Learning] The prediction step can be a step of predicting binding proteins of disease preventive drugs using a neural network model for separating interacting pairs from non-interacting pairs, using data such as interaction data, gene expression profile data, and chemical structure data as training data. The prediction step generates a trained model (trained neural network model) relating to the binding between compounds and binding proteins from these data and uses this model. Other machine learning models, such as kernel methods and sparse classifiers, may also be used. By using such a trained model, when the disease preventive drug compounds extracted in the analysis step are input as unknown data, the output binding proteins are identified as drug discovery target proteins.

[0078] In the prediction step, trained models are created so that predictions can be made based on the similarity of chemical structures, the similarity of gene expression profiles, the chemical structure using a graph convolutional neural network, etc. Only one of these trained models may be used, or multiple trained models may be used in combination.

[0079] [Kernel method] 8, 9, and 10 are diagrams illustrating a part of the steps according to another example of the prediction step of the present invention. These particularly relate to the prediction score of (3) above, which uses machine learning. This predicts binding proteins using the kernel method.

[0080] As shown in Figure 8, we attempt to formulate prediction of interactions between compounds and proteins from the perspective of machine learning. To this end, we create a trained model that predicts interactions between compounds and proteins whose interactions are unknown, using known interactions between compounds and proteins as training data.

[0081] As shown in Figure 9, here, pairs of compounds and proteins are created, and whether they are interacting pairs or non-interacting pairs is known is used as training data. Pairwise learning of compounds and proteins is performed, and a feature space is created for predicting whether or not a compound and protein interact.

[0082] The data framework used for machine learning of interactions can be chemogenomics, phenomics, transcriptomics, etc. These are used as information on compound similarity and protein similarity in the training data to predict interactions.

[0083] Chemogenomics combines information from the chemical space (chemical structure of compounds) and the genome space (protein sequence and structure). For example, see the aforementioned references "Yamanishi et al., Bioinformatics, 2008," "Yamanishi, Adv Neural Inf Process Syst, 2009," "Bleakley et al., Bioinformatics, 2009," and "Tabei et al., BMC Systems Biology, 2013."

[0084] The phenomics approach combines information from the pharmacological space, which relates to phenotypes on the human body such as headaches, nausea, elevated mood, changes in blood pressure, and fluctuations in disease markers, with information from the genomic space, which relates to protein sequences and structures.

[0085] Transcriptomics combines information from the transcriptional space regarding compound-responsive gene expression and the genomic space regarding protein gene expression. For example, see references such as "Hizukuri et al., BMC Med Genomics, 2015" and "Iwata et al., Sci rep, 2017; Sawada et al., Sci rep, 2018."

[0086] [Logistic Regression] Logistic regression can be used to classify all compounds in a space of descriptors into those with and without interaction with proteins. Logistic regression can be calculated using the following logistic regression formula based on L1 regularization. In this formula, the symbols are as follows: Xi: compound descriptor i. n: total number of compounds. w: descriptor weight. C: regularization parameter. This allows learning to increase the weight of descriptors that are effective in classification. A learning method based on L2 regularization may also be used. Other sparse classifiers, such as support vector machines with L1 regularization, may also be used.

[0087]

number

[0088] [Neural Networks] Figure 11 is a diagram illustrating another example of the prediction process of the present invention. The reference "Altae-Tran et al., 2017, ACS Cent. Sci." can be used as a reference. Here, graph convolution is performed based on the chemical structure of a drug compound to construct a neural network that determines whether or not it has an effect on the input layer. Graph convolution involves extracting graph topology and atomic features (1. extract graph topology and atom features), applying graph convolution and pooling (2. apply graph convolutions and pools), applying graph gathering (3. apply graph gather), and applying a dense layer (4. apply dense layer). This graph convolution is used to construct neural networks for the input, intermediate, and output layers to determine whether or not an interaction exists. This neural network can be used for the prediction score described above in (3C). Either single-task learning, which builds a predictive model for each protein, or multi-task learning, which builds predictive models for all proteins simultaneously, can be used.

[0089] [Integrated score] The prediction step can involve predicting two or more prediction scores from a group of prediction scores, calculating an integrated score, and ranking the multiple candidate binding proteins. For example, once the prediction scores (1) to (3) are calculated, they can be normalized within each index and summed to obtain the integrated score. For a binding protein corresponding to a certain disease preventive drug, if the prediction scores calculated are 0.9 from (1), 0.8 from (2), 0.95 from (3), 0.85 from (3), and 0.8 from (3), the total integrated score will be 4.3.

[0090] The prediction step preferably involves extracting multiple disease preventive drugs in the analysis step and calculating a prediction score by combining the prediction scores of these disease preventive drugs. This may be referred to as an overall score. For example, an example will be described in which two disease preventive drug candidates and three binding proteins are extracted. For the first disease preventive drug (a) extracted from the odds ratio, prediction scores and integrated scores for the candidate binding proteins, protein A, protein B, and protein C, are calculated. Next, for the second disease preventive drug (b), prediction scores and integrated scores for the candidate binding proteins, protein A, protein B, and protein C, are calculated. By combining the prediction scores for the first disease preventive drug (a) and the second disease preventive drug (b), a comprehensive ranking of protein A, protein B, and protein C can be determined from multiple perspectives, which is expected to improve reliability, etc.

[0091] Furthermore, these integrated scores and overall scores may be the sum or weighted sum of the prediction scores of each disease preventive drug.

[0092] In this way, the present invention makes it possible to predict drug discovery target proteins. Drug discovery target proteins predicted by the present invention can contribute to the efficient discovery of drug discovery target proteins, and can also extract proteins whose in vivo mechanism of action is unknown and which would be overlooked by conventional methods. Furthermore, predicting drug discovery target proteins in this way is expected to facilitate the investigation of therapeutic drugs for diseases and contribute to improving the efficiency of development of therapeutic drugs, etc.

[0093] [Prediction flow] The flow according to one example of the present invention will be described in more detail below.

[0094] We analyzed a certain disease α and investigated drug discovery target proteins for disease α. Disease α can be, for example, an intractable disease for which there are few effective therapeutic drugs.

[0095] [Analysis 1. Analysis process using FAERS] The odds ratios of disease α were analyzed using data from the FAERS (Adverse Drug Event Reporting System) database. The results are shown in Table 1. These drugs had small odds ratios and were selected as candidates for disease prevention drugs.

[0096] [Table 1]

[0097] [Prediction 1. Drug discovery target protein prediction] The binding proteins of the disease preventive drug compounds extracted by the analysis in Analysis 1 were evaluated. Using a database of compound-protein interactions, a convolutional neural network was trained using the data registered in the database as training data to create a trained model. This trained model was used to predict the binding proteins of the compounds that are the active substances of each drug. Other prediction models, such as kernel methods and sparse classifiers, may also be used. The sum of the predicted scores for the binding proteins of each disease preventive drug compound was used as the integrated score. The predicted scores for each compound and the total score, which is their sum, are shown in Table 2.

[0098] For each compound, the target protein is predicted and its prediction score can be calculated. Furthermore, by summing the prediction scores for multiple compounds, proteins with higher overall prediction scores can be identified. These results can be used to develop therapeutic drugs using these proteins as drug discovery targets.

[0099] [Table 2]

[0100] [Analysis example of irritable bowel syndrome] Below is an example of a case study that predicted drug discovery targets for irritable bowel syndrome.

[0101] [Analysis 1-1. Analysis process using FAERS] The odds ratio of irritable bowel syndrome was analyzed using data from the FAERS (Adverse Drug Events Reporting System) database. As a result, the following drugs were identified as candidates for disease prevention drugs, as they had small odds ratios.

[0102] [Table 3]

[0103] [Prediction 1-1. Prediction of drug discovery target proteins] Binding proteins were evaluated for the disease prevention drug compounds extracted by the analysis in Analysis 1-1. Using databases related to compound-protein binding, including ChEMBL, MATADOR, Drug Bank, PDSP-Ki, KEGG DRUG, BindingDB, and Therapeutic Target Database, we used compound-protein binding data registered in these databases as training data to train similarity search, graph convolutional neural networks, and logistic regression models to create trained models. Cross-validation was performed on the training data to find the hyperparameters that maximized prediction accuracy, and each model was trained using the optimized hyperparameters. These trained models were used to predict the binding proteins for the compounds that are the active ingredients of each drug. In addition, the sum of the predicted scores of the binding proteins of each disease preventive drug compound was used as the integrated score.

[0104] When the top 10 proteins with the highest scores were predicted, the predicted scores were, as shown in the table below, TDP1, KCNH2, OPRK1, NR1I2, ORM1, TP53, OPRD1, OPRM1, HTR2A, and HTR3A.

[0105] In fact, OPRK1, OPRM1, and HTR3A corresponded to proteins that are known drug discovery targets for irritable bowel syndrome. In other words, this is an example of a known drug discovery target for irritable bowel syndrome being reproduced using the proposed method. Proteins other than known drug discovery targets are expected to be candidates for new drug discovery targets for irritable bowel syndrome, and these proteins can be used as drug discovery targets for the development of therapeutic drugs.

[0106] [Table 4] [Industrial Applicability]

[0107] The present invention can be used to predict target proteins for drug discovery and is industrially useful. [Explanation of symbols]

[0108] 1. Prediction System 2 Input section 3 Data section 31 Clinical Data 32 Interaction Data 33 Omics Data 34 Chemical structure data 4. Control section 5. Memory 6 Display section 61 External Monitor

Claims

1. A method for predicting a drug discovery target protein for a disease to be treated, comprising: a step of calculating the likelihood of a disease to be treated when a drug is administered using clinical data, and extracting drugs for which the likelihood of the disease to be treated is lower than a threshold value as disease preventive drugs; calculating a prediction score, which is an index showing the possibility that a protein is a binding protein to the compound of the disease preventive drug, using one or more types of data selected from the group consisting of interaction data, gene expression profile data, and chemical structure data regarding compounds and proteins, and predicting a binding protein to the compound of the disease preventive drug based on the magnitude of the prediction score; The method for predicting a drug discovery target protein includes predicting the binding protein as a drug discovery target protein for the disease.

2. The prediction method according to claim 1, wherein the predicting step predicts the binding protein as one or more prediction scores selected from a group of prediction scores consisting of the following (1) to (3): (1) Using the chemical structure data and the interaction data, a registered compound that has a high degree of similarity in chemical structure to the compound of the disease preventive drug is determined as a predicted compound from registered compounds that are known to bind to a protein. A predicted score of the chemical structure (2) Using the gene expression profile data and the interaction data, a registered compound having a gene expression profile highly similar to that of the compound of the disease preventive drug is selected as a predicted compound from registered compounds known to bind to a protein. A gene expression prediction score (3) A prediction model obtained by machine learning using at least one of the gene expression profile data and the chemical structure data, and the interaction data as learning data, and a prediction score for predicting binding proteins of the disease preventive drug using the prediction model for separating interacting pairs having an interaction (binding) between a compound and a protein and non-interacting pairs having no interaction.

3. The prediction method according to claim 2, wherein the predicting step predicts two or more prediction scores from the group of prediction scores, calculates an integrated score, and ranks the plurality of candidate binding proteins.

4. The prediction method according to any one of claims 1 to 3, wherein the clinical data is data collected in a drug adverse event reporting system.

5. In the extracting step, two or more types of disease preventive drugs are extracted from candidate disease preventive drugs; The prediction method according to any one of claims 1 to 4, wherein in the predicting step, binding protein candidates for each of the plurality of disease preventive drugs are extracted, and a drug discovery target protein is predicted by combining these binding protein candidates for the plurality of disease preventive drugs.

6. A system for predicting a drug discovery target protein for a disease to be treated, comprising: an analysis unit that calculates the likelihood of disease when a drug is administered using clinical data and extracts drugs with a disease likelihood lower than a threshold as disease preventive drugs; a prediction unit that calculates a prediction score, which is an index showing the possibility that a protein is a binding protein to the compound of the disease preventive drug, using one or more types of data selected from the group consisting of interaction data, gene expression profile data, and chemical structure data regarding compounds and proteins, and predicts a binding protein to the compound of the disease preventive drug based on the magnitude of the prediction score; A drug discovery target protein prediction system that predicts the binding protein as a drug discovery target protein for the disease.

Citation Information

Patent Citations

  • Method and system for predicting protein-protein interaction as drug target

    JP2010165230A

  • Recording medium storing gene information / medical care information, and management system and method of gene information / medical care information using the recording medium

    JP2016014984A

  • Medicament search device, medicament search method and program

    JP2016099674A

  • Combined Affinity Prediction System and Method

    JP2017520868A

  • Computational systems for biomedical data

    US20080082522A1