Prediction method for drug effect of compound

The method addresses slow drug development by predicting drug efficacy through binding affinity profiles and machine learning, enabling effective prediction for new drug modalities and diseases without existing treatments.

JP2025159843APending Publication Date: 2025-10-22UNIV OKAYAMA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024062646
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-09
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing drug development methods are slow and costly, and computational methods for predicting drug efficacy are limited by the lack of pharmacological action information for new drug modalities.

Method used

A method using deep learning to predict drug efficacy by creating binding affinity profiles between compounds and in vivo proteins, integrating these profiles, and applying pathway analysis, similarity search, and machine learning to predict efficacy.

Benefits of technology

Enables the prediction of pharmacological efficacy for all molecules in the biological world, including diseases without existing drugs, and supports multi-modality drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025159843000001_ABST
    Figure 2025159843000001_ABST
Patent Text Reader

Abstract

To provide a prediction method and device for drug effect of a compound that can predict drug effect on all molecules present in the living world, and further can predict drug effect on a disease for which there is no existing drug available.SOLUTION: There is provided a prediction method for drug effect of a compound, which includes: a process S2 of creating a first coupling affinity profile as a coupling affinity profile of a prediction object compound group and an in vivo protein group; and a process S5 of predicting drug effect of the prediction object compound group based upon the created first coupling affinity profile and a second coupling affinity profile as a coupling affinity profile of a compound group whose drug effect information is present and the in vivo protein group.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a method for predicting the efficacy of a compound, and an apparatus and program for predicting the efficacy of a compound. [Background technology]

[0002] Although many drugs have been developed in recent years, there are still many rare and intractable diseases for which no effective treatment exists, and there is a strong demand for the development of new therapeutic drugs for these diseases. However, new drug development using traditional medical modalities such as small molecule compounds has been sluggish, and the enormous time and cost required to develop a single new drug is a problem. As a solution to this problem, attention is being paid to the development of new pharmaceuticals using new drug discovery modalities (antibody drugs, recombinant proteins, cell therapy, etc.).

[0003] The search method for new drug discovery modalities has mainly been experimental search methods based on inferences from past examples, existing mechanisms, structural similarities, etc. However, there is a limit to the number of pairings that can be experimentally confirmed, and in order to conduct a comprehensive search for drug discovery modalities, it is essential to develop a computational method for predicting drug efficacy.

[0004] As described above, conventional techniques are limited in the number of pairings that can be experimentally verified. Furthermore, a certain amount of pharmacological action information is required to predict drug efficacy using a computer. However, many new drug discovery modalities are still in the development stage, and there is little known information about pharmacological action, making it difficult to apply computational methods. The present inventors have reported new computational methods to address these issues in Non-Patent Documents 1 and 2.

[0005] Non-Patent Document 1 reports that a new computational method was developed to predict potential drug discovery targets and new drug indications for systematic drug repositioning using large-scale compound-protein interactome data, and that a statistical model was constructed to predict new drug indications for a wide range of diseases with various molecular characteristics based on the drug target profile.

[0006] Non-Patent Document 2 reports that a new in silico model has been constructed to predict compound-induced side effects and estimate the underlying mechanisms with high versatility by integrating comprehensive prediction of potential compound-protein interactions (CPIs) with machine learning, and that cross-validation experiments have shown that the proposed CPI-based model has higher or equal performance than conventional compound molecular structure-based models. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] J Chem Inf Model. 2015 Dec 28;55(12):2717-30 [Non-patent document 2] J Toxicol Sci. 2020;45(3):137-149 Summary of the Invention [Problem to be solved by the invention]

[0008] The present disclosure aims to provide a method for predicting the efficacy of a compound. [Means for solving the problem]

[0009] As a result of intensive research conducted by the inventors to achieve the above-mentioned objective, they discovered that it is possible to predict drug efficacy by creating a binding affinity profile between a drug with known efficacy and an in vivo protein, and a binding affinity profile between a compound with unknown efficacy and an in vivo protein, integrating these profiles, and using a deep learning model.

[0010] The present disclosure was completed based on these findings and further investigations, and includes, for example, the subject matter described in the following sections.

[0011] Item 1. A method for predicting the efficacy of a compound, comprising: (1) creating a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is to be predicted and a group of in vivo proteins; and (2) predicting the efficacy of the target compounds based on the first binding affinity profile prepared in step (1) and a binding affinity profile (second binding affinity profile) between the compounds and the in vivo proteins for which efficacy information exists; A method comprising: Item 2. A method for predicting the efficacy of a compound, comprising: (1) creating a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is predicted and a group of in vivo proteins; (1') creating a binding affinity profile (second binding affinity profile) between a group of compounds for which efficacy information exists and a group of in vivo proteins; and (2) predicting the efficacy of the group of target compounds based on the first and second binding affinity profiles created in steps (1) and (1'); A method comprising: Item 3. The method according to Item 1 or 2, wherein in the steps (1) and (1'), a binding affinity profile between the group of compounds and the group of proteins is created by docking simulation. Item 4. The method according to any one of Items 1 to 3, wherein in step (2), the efficacy of the group of target compounds is predicted based on the similarity between the first and second binding affinity profiles by applying the binding affinity profiles to pathway analysis, similarity search, or statistical methods and machine learning. Item 5. The method according to any one of Items 1 to 3, wherein in the step (2), the efficacy of the group of target compounds is predicted based on the similarity between the first and second binding affinity profiles by applying statistical techniques and machine learning to the binding affinity profiles. Item 6. The method according to Item 5, wherein the statistical method and machine learning are deep learning. Item 7. (3) A step of selecting compounds from a group of predicted target compounds that have a high similarity to the binding affinity profile of a group of compounds for which efficacy information exists for a specific disease as candidate compounds having efficacy for the specific disease. Item 7. The method according to any one of Items 1 to 6, further comprising: Item 8. The method according to any one of Items 2 to 7, wherein the group of in vivo proteins used in step (1) is the same as the group of in vivo proteins used in step (1'). Item 9. The method according to any one of Items 1 to 8, wherein the number of endogenous proteins used in steps (1) and (1') is 5,000 or more. Item 10. The method according to any one of Items 1 to 9, wherein a group of compounds with known pharmacological effects is used as the group of compounds for which pharmacological effect information exists. Item 11. The method according to any one of Items 1 to 10, wherein a group of compounds to which pharmacological information is assigned obtained by analyzing clinical data is used as the group of compounds for which pharmacological information is available. Item 12. The method according to any one of Items 1 to 11, wherein the group of compounds for which efficacy is predicted is at least one selected from the group consisting of low molecular weight compounds, medium molecular weight compounds, proteins, and individual active ingredients of herbal medicines. Item 13. A device for predicting the efficacy of a compound, a binding affinity profile creation unit that creates a binding affinity profile between a group of compounds to be predicted for drug efficacy and a group of in vivo proteins; and a prediction unit that predicts the efficacy of a group of compounds to be predicted based on the binding affinity profile created by the binding affinity profile creation unit and the binding affinity profile between a group of compounds for which efficacy information exists and a group of in vivo proteins; An apparatus comprising: Item 14. A device for predicting the efficacy of a compound, a first binding affinity profile creation unit that creates a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is to be predicted and a group of in vivo proteins; a second binding affinity profile creation unit that creates a binding affinity profile (second binding affinity profile) between a group of compounds for which drug efficacy information exists and a group of in vivo proteins; and a prediction unit that predicts the efficacy of the prediction target compound group based on the first binding affinity profile created by the first binding affinity profile creation unit and the second binding affinity profile created by the second binding affinity profile creation unit; An apparatus comprising: Item 15. A program for predicting the efficacy of a compound, (1) creating a binding affinity profile between a group of compounds whose efficacy is predicted and a group of in vivo proteins; and (2) A step of predicting the efficacy of a group of compounds to be predicted based on the binding affinity profile created in step (1) and the binding affinity profile between a group of compounds for which efficacy information exists and a group of in vivo proteins. A program that causes a computer to execute the following. Item 16. A program for predicting the efficacy of a compound, (1) creating a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is predicted and a group of in vivo proteins; (1') creating a binding affinity profile (second binding affinity profile) between a group of compounds for which efficacy information exists and a group of in vivo proteins; and (2) predicting the efficacy of the group of target compounds based on the first and second binding affinity profiles created in steps (1) and (1'); A program that causes a computer to execute the following. [Effects of the Invention]

[0012] The present disclosure provides a novel and effective means for predicting the pharmacological efficacy of compounds. The method of the present disclosure makes it possible to predict the pharmacological efficacy of all molecules present in the biological world, and also makes it possible to predict the pharmacological efficacy of diseases for which no existing drugs exist.

[0013] The drug efficacy prediction technology disclosed herein is a method for predicting the drug efficacy of new modalities, and is expected to become one of the important basic technologies for multi-modality drug discovery and development. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a flowchart showing the processing steps of a method for predicting the efficacy of a compound. [Figure 2] 2 is a block diagram of an apparatus for carrying out the method shown in FIG. 1. [Figure 3] FIG. 1 shows the flow of in silico screening using protein binding affinity profiles. [Figure 4] Figure 1 shows plasma protein-in vivo protein docking scores. [Figure 5] FIG. 1 shows drug-protein binding affinity in vivo. [Figure 6] FIG. 1 shows the results of an analysis of drugs that may have a preventive effect on amyotrophic lateral sclerosis (ALS) using medical big data. [Figure 7] FIG. 1 is a diagram showing a Transformer-based AI prediction model used in the examples. [Figure 8] FIG. 1 shows the results of predicting the efficacy of plasma proteins. [Figure 9] 1 is a table showing the prediction accuracy of the efficacy of existing plasma protein drugs. DETAILED DESCRIPTION OF THE INVENTION

[0015] Each embodiment included in the present disclosure will be described in more detail below. The present disclosure preferably includes, but is not limited to, a method for predicting the efficacy of a compound, and the like. The present disclosure includes all that is disclosed in the present specification and that can be recognized by a person skilled in the art.

[0016] Method for predicting efficacy of compounds The method for predicting the efficacy of a compound of the present disclosure includes: (1) creating a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is to be predicted and a group of in vivo proteins; and (2) predicting the efficacy of the target compounds based on the first binding affinity profile prepared in step (1) and a binding affinity profile (second binding affinity profile) between the compounds and the in vivo proteins for which efficacy information exists; (hereinafter, also referred to as the "prediction method of the present disclosure").

[0017] Another embodiment of the prediction method of the present disclosure comprises: (1) creating a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is predicted and a group of in vivo proteins; (1') creating a binding affinity profile (second binding affinity profile) between a group of compounds for which efficacy information exists and a group of in vivo proteins; and (2) predicting the efficacy of the group of target compounds based on the first and second binding affinity profiles created in steps (1) and (1'); Includes.

[0018] The terms and phrases used in this specification are explained in detail below.

[0019] Compounds in the group of compounds for which efficacy prediction is performed and compounds for which efficacy information exists can be used without particular limitations, and examples include (inorganic and organic) low molecular weight compounds, medium molecular weight compounds (e.g., peptides), and biopolymers (e.g., proteins, nucleic acids, polysaccharides). Furthermore, individual active ingredients of herbal medicines can also be used as such compounds. Examples of low molecular weight compounds include compounds with a molecular weight of approximately 1000 or less, and examples of medium molecular weight compounds include compounds with a molecular weight of approximately 1000 to 3000.

[0020] A medicinal effect refers to the resulting effect (efficacy) when a compound (drug) is applied to a target. The types of diseases that are the target of the medicinal effect in the present disclosure are not particularly limited, and examples include skin diseases, otic diseases, musculoskeletal diseases, ophthalmic diseases, digestive diseases, circulatory diseases, nervous system diseases, mental disorders, respiratory diseases, dental and oral diseases, endocrine and metabolic diseases, blood and hematopoietic diseases, infectious diseases, malignant neoplasms, congenital and genetic diseases, etc.

[0021] The prediction method of the present disclosure is a method for predicting drug efficacy in an organism, and the organism here is, for example, a mammal such as a human, monkey, rat, mouse, rabbit, horse, cow, goat, sheep, dog, or cat, and preferably a human.

[0022] The term "group of compounds for which efficacy is predicted" refers to a group of compounds for which efficacy is predicted in the prediction method of the present disclosure. The group of compounds for which efficacy is predicted can be used without any particular limitation, and examples thereof include low molecular weight compounds, medium molecular weight compounds, proteins, and individual active ingredients of herbal medicines. Compounds for which efficacy information is already known can also be used to predict new efficacy. The number of compounds for which efficacy is predicted is not particularly limited as long as computational resources allow. Approximately 200 million compounds registered in the CAS registry are included, and the number may be, for example, 1 to 5,000. The upper or lower limit of this range can be, for example, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, or 4,000.

[0023] The "group of compounds for which pharmacological efficacy information is available" can be any group of compounds for which available information on pharmacological efficacy is available, and examples of such compounds include groups of compounds with known pharmacological efficacy (e.g., groups of compounds for which pharmacological efficacy information is registered in publicly known databases such as KEGG DRUG, DrugBank, PubChem, and ChEMBL), and groups of compounds to which pharmacological efficacy information (e.g., applicable diseases) has been assigned, obtained by analyzing clinical data (e.g., medical big data) using methods such as those described below. The "group of compounds for which pharmacological efficacy information is available" can be used without any particular limitation, and particularly includes low-molecular-weight compounds, medium-molecular-weight compounds, proteins, and the like. The number of "group of compounds for which pharmacological efficacy information is available" can be, for example, 1 to 7,000, and the upper and lower limits of this range can be, for example, 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, and 6,000.

[0024] "In vivo protein group" refers to a group of proteins present inside a living organism, and examples of such proteins include plasma proteins, membrane proteins, enzymes, hormones, cytokines, ligands, receptors, regulatory proteins, cytoskeleton, and secreted proteins. The living organism in this context is preferably the same as the target organism for which the drug efficacy is predicted, and is, for example, a mammal such as a human, monkey, rat, mouse, rabbit, horse, cow, goat, sheep, dog, or cat, with a human being being preferred. The number of "in vivo protein group" may be, for example, 5,000 or more species, and the upper or lower limit of the number of "in vivo protein group" may be, for example, 6,000, 7,000, 8,000, 9,000, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, 16,000, 17,000, 18,000, 19,000, or 20,000 species. The in vivo protein group includes all proteins present in the living world, 1×10 7 Seeds may also be used.

[0025] In the present disclosure, "binding affinity" refers to the activity of non-covalent interactions between a group of compounds whose efficacy is predicted or a group of compounds whose efficacy information exists and an in vivo protein. Furthermore, "binding affinity profile" refers to information on the level of binding affinity between each compound in a group of compounds whose efficacy is predicted or a group of compounds whose efficacy information exists and an in vivo protein.

[0026] Process (1) In step (1), a binding affinity profile between a group of target compounds whose efficacy is predicted and a group of in vivo proteins is created. In the present disclosure, the binding affinity profile between a group of target compounds whose efficacy is predicted and a group of in vivo proteins is sometimes referred to as a first binding affinity profile. In step (1), when a binding affinity profile between N types of target compounds whose efficacy is predicted and M types of in vivo proteins is to be obtained, a total of N × M binding affinity profiles can be created.

[0027] The first binding affinity profile can be created using various known computational techniques. The first binding affinity profile can also be created, for example, by docking simulation. Docking simulation can be performed using dedicated software such as MEGADOCK, AutoDock Vina, and rDOCK. To perform such docking simulation, data on the three-dimensional structure of a compound such as a protein is generally required. Such three-dimensional structure data can be obtained from known databases, such as the Protein Data Bank (PDB) and AlphaFold DB. Furthermore, docking simulations can be performed for low molecular weight compounds using structural formula files such as MOL files. Such structural formula files such as MOL files can be obtained from known databases. The docking score obtained by the docking simulation calculation can then be used as the binding affinity value. Furthermore, experimentally determined data can also be used as part of the first binding affinity profile, if necessary.

[0028] Process (1') In step (1'), a binding affinity profile between a group of compounds for which efficacy information exists and a group of in vivo proteins is created. In the present disclosure, the binding affinity profile between a group of compounds for which efficacy information exists and a group of in vivo proteins is sometimes referred to as a second binding affinity profile. In step (1'), when a binding affinity profile between A types of compounds for which efficacy information exists and M types of in vivo proteins is to be obtained, a binding affinity profile consisting of a total of A × M types can be created.

[0029] The second binding affinity profile can be created using various known computational techniques. The second binding affinity profile can also be created, for example, by docking simulation. Docking simulation can be performed using dedicated software such as MEGADOCK, AutoDock Vina, or rDOCK. To perform such docking simulation, data on the three-dimensional structure of a compound such as a protein is generally required. Such three-dimensional structure data can be obtained from known databases, such as the Protein Data Bank (PDB) and AlphaFold DB. Furthermore, docking simulations can be performed for low molecular weight compounds using structural formula files such as MOL files. Such structural formula files such as MOL files can be obtained from known databases. The docking score obtained by the docking simulation calculation can then be used as the binding affinity value. Furthermore, experimentally determined data can also be used as part of the second binding affinity profile, if necessary.

[0030] The order of carrying out step (1) and step (1') may be any; step (1') may be carried out after step (1), or step (1') may be carried out after step (1).

[0031] In order to predict the efficacy of the group of compounds to be predicted in step (2), the first and second binding affinity profiles are compared, so it is preferable that the group of in vivo proteins used in step (1) and the group of in vivo proteins used in step (1') are the same.

[0032] In addition to compounds with known efficacy, compounds with assigned efficacy information obtained by analyzing clinical data may also be used as the group of compounds for which efficacy information is available. Using compounds with assigned efficacy information obtained by analyzing clinical data makes it possible to predict efficacy even for diseases for which no existing drugs exist. The method for assigning efficacy information by analyzing clinical data is not particularly limited, and for example, a known method reported by the present inventors in a literature article (Kidney Int. 2021 Apr;99(4):885-899) can be used. In the literature, a side effect database was used to explore potential pharmacological effects for drug repositioning when a certain drug is used, and the validity of this exploratory method was reported to have been confirmed by verification using basic pharmacological techniques and retrospective observational studies. Furthermore, by using a side effect database (e.g., VigiBase) to calculate the odds ratio of the likelihood that a certain disease will not occur when a certain drug is used, drugs with a low odds ratio can be assigned as drugs with therapeutic and preventive effects for a certain disease.

[0033] Process (2) In step (2), the efficacy of the target compounds is predicted based on the first binding affinity profile prepared in step (1) and the binding affinity profile between the compounds for which efficacy information exists and the in vivo proteins. The "binding affinity profile between the compounds for which efficacy information exists and the in vivo proteins" may be one prepared in advance, or the second binding affinity profile prepared in step (1') may be used.

[0034] In step (2), the efficacy of the group of predicted target compounds is predicted based on the similarity between the first and second binding affinity profiles by applying pathway analysis, similarity search, statistical methods, machine learning, and the like to the first and second binding affinity profiles. That is, the efficacy of compounds for which efficacy information exists that have binding affinity profiles highly similar to the binding affinity profile of the target compound whose efficacy is predicted can be predicted as the efficacy of the target compound. Here, the probability of each efficacy can be assigned depending on the degree of similarity with the second binding affinity profile.

[0035] Pathway analysis is a method of extracting pathways involving intracellular proteins with high or low binding affinity from binding affinity profiles, and predicting drug efficacy based on the similarity of the pathways with which each compound group interacts.

[0036] Similarity search is a method for predicting the efficacy of a group of compounds to be predicted by comparing the similarity between a first binding affinity profile and a second binding affinity profile.

[0037] The statistical methods and machine learning are not particularly limited as long as they can predict the efficacy of a group of target compounds based on the similarity between the first and second binding affinity profiles, and examples include deep learning (e.g., Transformer), neural networks, support vector machines, random forests, logistic regression, and principal component analysis.

[0038] In step (2), it is desirable to predict the efficacy of the group of target compounds based on the similarity between the first and second binding affinity profiles by applying statistical methods and machine learning to the binding affinity profile. As the statistical methods and machine learning, deep learning is particularly preferred. When machine learning is used, a trained prediction model of efficacy may be created using a previously created second binding affinity profile, and the efficacy of the group of target compounds may be predicted using the trained prediction model.

[0039] Process (3) The prediction method of the present disclosure may further include the following step (3).

[0040] In step (3), compounds from the group of predicted target compounds that have a high degree of similarity to the binding affinity profile of a group of compounds for which efficacy information exists for a specific disease are selected as candidate compounds that have efficacy for the specific disease.

[0041] As described above, in step (2), the first and second binding affinity profiles are subjected to pathway analysis, similarity search, statistical methods, machine learning, and the like, whereby the likelihood of each compound having a pharmacological effect corresponding to a group of compounds for which pharmacological information exists can be assigned to the group of predicted target compounds according to the degree of similarity between the first and second binding affinity profiles. Therefore, compounds in the group of predicted target compounds that are highly similar to the second binding affinity profile for a specific disease and are predicted to have a high likelihood of pharmacological effect on the disease can be selected as candidate compounds having pharmacological effect on the specific disease. The candidate compounds selected in this case may be one type, or two or more types of compounds. When two or more types of compounds are selected, compounds having a certain level of similarity or higher may be selected as candidate compounds.

[0042] The prediction method of the present disclosure is a novel and effective means for predicting the pharmacological efficacy of compounds. The prediction method of the present disclosure makes it possible to predict the pharmacological efficacy of all molecules present in the biological world, and also makes it possible to predict the pharmacological efficacy of diseases for which no existing drugs exist.

[0043] The prediction method disclosed herein is capable of predicting the efficacy of new modalities, and is expected to become one of the important basic technologies for multi-modality drug development.

[0044] Forecasting process Next, an example of the processing procedure of the prediction method of the present disclosure will be described. Fig. 1 is a flowchart showing the processing procedure of the prediction method of the present disclosure, and Fig. 2 is a block diagram of an apparatus 1 for executing steps S1 to S5 of the method.

[0045] The device 1 can be configured using a general-purpose computer or the like, and includes a data processing unit 2, an auxiliary storage device 3, an input unit 4, a display unit 5, and a communication interface unit (communication I / F unit) 6. The data processing unit 2 is configured as software, while the auxiliary storage device 3, the input unit 4, the display unit 5, and the communication I / F unit 6 are configured as hardware. The device 1 further includes, as hardware components, a processor such as a CPU or GPU, and a main storage device such as a DRAM or SRAM.

[0046] The data processing unit 2 is a functional block realized by a processor executing a program (a compound efficacy prediction program 37) of the present disclosure. The device 1 includes, as its functional blocks, a first binding affinity profile creation unit 21, a second binding affinity profile creation unit 22, and an efficacy prediction unit 23. These units can be realized in software by the processor of the device 1 reading the program of the present disclosure into a main storage device and executing it. The program 37 may be downloaded to the device 1 via a network 7 such as the Internet, or may be installed in the device 1 via a computer-readable non-transitory recording medium such as a CD-R on which the program 37 is recorded.

[0047] In the flowchart shown in FIG. 1, in step S1, the device 1 acquires input of three-dimensional structure data 31 of a group of target compounds whose efficacy is to be predicted and three-dimensional structure data 32 of a group of in vivo proteins via the input unit 4, the communication I / F unit 6, etc., and stores the data in the auxiliary storage device 3.

[0048] In step S2, the first binding affinity profile creation unit 21 of the device 1 creates a first binding affinity profile based on the three-dimensional structure data 31 of the group of target compounds whose efficacy is to be predicted and the three-dimensional structure data 32 of the group of in vivo proteins, which were input in step S1. The calculated first binding affinity profile data 33 is stored in the auxiliary storage device 3. Step (1) in the claims corresponds to step S2.

[0049] In step S3, the device 1 acquires input of three-dimensional structure data 34 of a group of compounds for which drug efficacy information exists via the input unit 4, the communication I / F unit 6, etc., and stores the data in the auxiliary storage device 3.

[0050] In step S4, the second binding affinity profile creation unit 22 of the device 1 creates a second binding affinity profile based on the three-dimensional structure data 34 of the compound group for which drug efficacy information exists and the three-dimensional structure data 32 of the in vivo protein group input in steps S3 and S1. The calculated second binding affinity profile data 35 is stored in the auxiliary storage device 3. Step (1') in the claims corresponds to step S4.

[0051] In step S5, the efficacy prediction unit 23 of the device 1 predicts the efficacy of the prediction target compound group based on the similarity between the first and second binding affinity profiles by applying pathway analysis, similarity search, statistical methods, machine learning, etc. to the first binding affinity profile 33 created in step S2 and the second binding affinity profile 35 created in step S4. The predicted efficacy data 36 is stored in the auxiliary storage device 3. The predicted efficacy is displayed on the display unit 5. The device 1 terminates the process after calculating the efficacy of the compound. Step (2) in the claims corresponds to step S5.

[0052] The auxiliary storage device 3 is a non-volatile storage device that stores an operating system (OS), various control programs, data generated by the programs, etc., and is configured by, for example, a flash memory, an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc. In this embodiment, the auxiliary storage device 3 stores prediction target compound group data 31, in vivo protein group data 32, a first binding affinity profile 33, efficacy information-containing compound group data 34, a first binding affinity profile 35, efficacy data 36, ​​and a compound efficacy prediction program 37.

[0053] The input unit 4 may be configured with, for example, a mouse, a keyboard, etc., and the display unit 5 may be configured with, for example, a liquid crystal display, an organic EL display, etc. The communication I / F unit 6 transmits and receives data to and from external devices via a wired or wireless network. The communication I / F unit 6 may be configured with various wired or wireless connections such as Ethernet (registered trademark), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0054] Other embodiments Although the present disclosure has been described above with reference to specific embodiments, the present disclosure is not limited to the above-described embodiments.

[0055] In the above embodiment, a step can be further added of selecting, based on the efficacy predicted in step S5, compounds from the group of predicted target compounds that have a high similarity to the binding affinity profile of a group of compounds for which efficacy information exists for a specific disease, as candidate compounds having efficacy for the specific disease.

[0056] In the above embodiment, steps S3 and S4 are performed after steps S1 and S2, but steps S1 and S2 may be performed after steps S3 and S4. In that case, in step S3, input of three-dimensional structure data 32 of a group of in vivo proteins is acquired via the input unit 4, communication I / F unit 6, etc., and stored in the auxiliary storage device 3. Alternatively, steps S1 and S3 may be performed first, and then steps S2 and S4 may be performed.

[0057] In the above embodiment, steps S3 and S4 are performed, but steps S3 and S4 may be omitted. In that case, the second binding affinity profile is created in advance, and data 35 of the second binding affinity profile is stored in the auxiliary storage device 3. In this case, a trained prediction model of drug efficacy may be created using the second binding affinity profile created in advance and stored in the auxiliary storage device 3, and the drug efficacy of the group of compounds to be predicted may be predicted in step S5 using this trained prediction model.

[0058] In the above embodiment, the device 1 is realized as an integrated device, but the device 1 does not need to be an integrated device, and the processor, memory, auxiliary storage device 3, etc. may be located in different places and connected to each other via a network. The input unit 4 and the display unit 5 also do not need to be located in one place, and may be located in different places and connected to each other so as to be able to communicate via a network.

[0059] In the above embodiment, the functional blocks 21 to 23 constituting the data processing unit 2 are realized by software, but some or all of these functional blocks 21 to 23 may be realized as hardware. The processing of the functional blocks 21 to 23 constituting the data processing unit 2 does not need to be performed by a single processor, but may be distributed and processed by multiple processors. Some or all of the functions of the data processing unit 2 and the data items in the auxiliary storage device 3 may be cloud-based in an external server device connected via the communication I / F unit 6.

[0060] It should be noted that, in this specification, the term "comprising" includes "consisting essentially of" and "consisting of." Furthermore, the present disclosure encompasses all arbitrary combinations of the constituent elements described in this specification.

[0061] Furthermore, the various characteristics (properties, structures, functions, etc.) described in each embodiment of the present disclosure above may be combined in any way to specify the subject matter encompassed by the present disclosure, i.e., the present disclosure encompasses all subject matter consisting of any combination of the combinable characteristics described herein. [Example]

[0062] The present invention will be described in more detail below with reference to examples, but the present invention is not limited to these examples.

[0063] Example Because typical drugs exert their efficacy by interacting with target proteins in the body, protein interactions are important for predicting drug efficacy. In this study, we performed drug efficacy prediction and plasma protein-disease pairings, taking into account the proteins they interact with in the body, using proteins present in plasma as an example. First, we created (1) binding affinity profiles between plasma proteins and in vivo proteins, and (2) binding affinity profiles between drugs with known efficacy and in vivo proteins. Furthermore, (3) for diseases for which no therapeutic drugs exist, we used medical big data to add diseases for which preventive effects are expected as drug efficacy. Finally, (4) we constructed a drug efficacy prediction deep learning model from the binding affinity profiles and efficacy information of drugs with known efficacy, applied it to the binding affinity profiles of plasma proteins, and predicted novel drug efficacies for plasma proteins (Figure 3). Details are provided below.

[0064] (1) Creation of in vivo protein binding affinity profiles of plasma proteins First, the three-dimensional structures of 119 plasma proteins were downloaded from the AlphaFold DB. Next, the three-dimensional structures of 5,291 in vivo proteins (approximately 3,000 plasma proteins and 3,000 membrane proteins) to be used for docking were also downloaded from the AlphaFold DB. Binding affinities were then calculated using the protein-protein docking software MEGADOCK, and docking scores between the 119 plasma proteins and the 5,291 in vivo proteins were determined (Figure 4).

[0065] (2) Creation of in vivo protein binding affinity profiles of drugs For 1,110 drugs registered in VigiBase for which efficacy information was available (drug efficacy: 446 types, efficacy information was obtained from KEGG DRUG), the binding affinity between these drugs and the 5,291 in vivo proteins extracted in (1) was calculated using AutoDock Vina (Figure 5).

[0066] (3) Allocation of applicable diseases using medical big data Using VigiBase, the WHO adverse drug reaction database (containing information on approximately 25 million adverse drug reactions reported from over 130 WHO member countries), we calculated the odds ratio of the likelihood that a certain disease will not occur when a certain drug is used. As an analysis example, we show drugs that may have a preventive effect on amyotrophic lateral sclerosis (ALS). After analyzing 855 drugs using VigiBase, we detected signals in 10 drugs (Figure 6).

[0067] (4) Drug efficacy prediction using AI machine learning models Using the pharmacological information of these small molecule drugs, we constructed a Transformer-based AI prediction model (Figure 7) and confirmed that it was an AI model capable of predicting the drug's efficacy with high accuracy (Table 1).

[0068] [Table 1]

[0069] The constructed AI prediction model utilizes the Encoder portion of the Transformer. When a protein-binding affinity profile is input, it outputs a score indicating whether the drug is effective. When a binding affinity profile is input to the Encoder, the Encoder's attention mechanism processes it, outputting a corrected binding affinity profile in which the affinity for proteins strongly associated with the drug's efficacy is weighted. Furthermore, the number of proteins, which is the characteristic size of the binding affinity profile, is maintained during processing, resulting in a profile of the same size as the input. The profile output from the Encoder is converted into a score indicating the efficacy of each drug using a linear transformation and activation function. Figure 8 shows the results of applying this AI model to existing blood products (plasma proteins). We then verified the prediction accuracy of the original efficacy (antithrombotic agents (ATC classification B01): 5 drugs, antihemorrhagic agents (ATC classification B02): 16 drugs), and predicted the efficacy of each existing drug in the first and second places (Figure 9). [Explanation of symbols]

[0070] 1 device 2 Data Processing Unit 21 First binding affinity profile creation unit 22 Secondary Binding Affinity Profile Creation Section 23 Drug Efficacy Prediction Department 3 Auxiliary storage 31 Predicted target compound group data 32 In vivo protein group data 33 First Binding Affinity Profile 34 Drug efficacy information existing compound group data 35 Secondary Binding Affinity Profile 36 Efficacy Data 37 Compound efficacy prediction program 4 Input section 5 Display section 6. Communication interface section (communication I / F section) 7 Network

Claims

1. A method for predicting the efficacy of a compound, comprising: (1) creating a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is to be predicted and a group of in vivo proteins; and (2) predicting the efficacy of the target compounds based on the first binding affinity profile prepared in step (1) and a binding affinity profile (second binding affinity profile) between the compounds for which efficacy information exists and the in vivo proteins; A method comprising:

2. A method for predicting the efficacy of a compound, comprising: (1) creating a binding affinity profile (first binding affinity profile) between a group of compounds whose efficacy is to be predicted and a group of in vivo proteins; (1') creating a binding affinity profile (second binding affinity profile) between a group of compounds for which drug efficacy information exists and a group of in vivo proteins; and (2) predicting the efficacy of the group of target compounds based on the first and second binding affinity profiles prepared in steps (1) and (1'); A method comprising:

3. The method according to claim 1 or 2, wherein in the steps (1) and (1'), a binding affinity profile between a group of compounds and a group of proteins is created by docking simulation.

4. The method according to claim 1 or 2, wherein in the step (2), the efficacy of the group of target compounds is predicted based on the similarity between the first and second binding affinity profiles by applying the binding affinity profiles to pathway analysis, similarity search, or statistical methods and machine learning.

5. The method according to claim 1 or 2, wherein in the step (2), the efficacy of the group of target compounds is predicted based on the similarity between the first and second binding affinity profiles by applying statistical techniques and machine learning to the binding affinity profiles.

6. The method of claim 5 , wherein the statistical techniques and machine learning are deep learning.

7. (3) A step of selecting compounds from the group of predicted target compounds that have a high similarity to the binding affinity profile of a group of compounds for which efficacy information exists for a specific disease as candidate compounds having efficacy for the specific disease. The method of claim 1 or 2, further comprising:

8. The method according to claim 2, wherein the group of in vivo proteins used in step (1) is the same as the group of in vivo proteins used in step (1').

9. The method according to claim 1 or 2, wherein the number of endogenous proteins used in steps (1) and (1') is 5,000 or more.

10. The method according to claim 1 or 2, wherein a group of compounds whose pharmacological effects are known is used as the group of compounds for which pharmacological effect information exists.

11. The method according to claim 1 or 2, wherein a group of compounds to which pharmacological information is assigned obtained by analyzing clinical data is used as the group of compounds for which pharmacological information is available.

12. The method according to claim 1 or 2, wherein the group of compounds for which the efficacy of the drug is predicted is at least one selected from the group consisting of low molecular weight compounds, medium molecular weight compounds, proteins, and individual active ingredients of herbal medicines.

13. An apparatus for predicting the efficacy of a compound, comprising: a binding affinity profile creation unit that creates a binding affinity profile between a group of compounds to be predicted for drug efficacy and a group of in vivo proteins; and a prediction unit that predicts the efficacy of a group of compounds to be predicted based on the binding affinity profile created by the binding affinity profile creation unit and the binding affinity profile between a group of compounds for which efficacy information exists and a group of in vivo proteins; An apparatus comprising:

14. A program for predicting the efficacy of a compound, (1) creating a binding affinity profile between a group of compounds whose efficacy is predicted and a group of in vivo proteins; and (2) predicting the efficacy of the target compounds based on the binding affinity profile created in step (1) and the binding affinity profiles between the compounds for which efficacy information exists and the in vivo proteins; A program that causes a computer to execute the following.