Prediction device, prediction method, and prediction program

The integration of chemical, pharmacological, and cellular information using machine learning in the prediction device and method addresses inefficiencies in drug discovery, enabling precise prediction of drug efficacy and side effects, thereby enhancing the drug development process.

JP2025134955APending Publication Date: 2025-09-17TRES ALCHEMIX CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025107559
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Existing drug discovery methods rely heavily on human expertise and are inefficient, with low success rates and high costs, and AI-based methods lack comprehensive integration of chemical, cellular, and clinical data to accurately predict drug efficacy and side effects.

Method used

A prediction device and method that integrates chemical, pharmacological, and cellular-level biological information using machine learning to estimate and predict drug efficacy and side effects, incorporating features like chemical structure, pharmacokinetics, and drug administration data to train a model that can predict in vivo pharmacodynamics.

Benefits of technology

Enables accurate and efficient drug development by predicting drug efficacy and side effects, reducing the need for extensive experimental verification and optimizing the drug discovery process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025134955000001_ABST
    Figure 2025134955000001_ABST
Patent Text Reader

Abstract

To provide a prediction device, a prediction method, and a prediction program that promote integration of chemical information of compounds, action information at a cellular level, and information in an organism or in clinical use, and that predict desired drug development, and also effects and adverse reactions of the drug, in a more accurate and efficient manner.SOLUTION: In a pharmacokinetic prediction system based on AI called ADMET prediction (Adsorption, Distribution, Metabolism, Excretion, Toxicity), a computer of a prediction device 1 executes the functions of: an acquisition section for acquiring chemical substance information and pharmacological information of a drug; an estimation section for estimating estimated information of a drug by conducting machine learning using the chemical substance information and the pharmacological information; and an output section for predicting and outputting both effects and adverse reactions of the drug in an organism by re-training a machine learning model based on the estimated information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a prediction device, a prediction method, and a prediction program, and more particularly to a device, a method, and a program for predicting a drug that will exhibit optimal efficacy and its efficacy, etc., by machine learning based on the properties of the drug. [Background technology]

[0002] In traditional drug discovery methods, developers would predict and generate effective pharmaceutical compounds based primarily on past knowledge and experience, and then evaluate the efficacy and side effects of the pharmaceutical compounds through experimental verification. However, with traditional methods, the success of drug discovery depends on the developer's proficiency and expertise, and often relies on accidental discoveries. As a result, the success rate of new drug development is extremely low, and it takes a huge amount of money and time to discover a single new drug.

[0003] In recent years, artificial intelligence (AI)-based drug discovery methods have attracted attention as a way to reduce the enormous development costs of traditional drug discovery. For example, molecular docking, which predicts the binding affinity between target proteins and drug candidate compounds, has been used to predict the optimal drug compound with the highest binding affinity. However, this method requires accurate identification of the target protein's three-dimensional structure and the drug-binding site on the protein based on that structure. Therefore, this method cannot be applied to proteins whose three-dimensional structure has not been accurately identified, such as membrane proteins including drug transporters. Furthermore, the value obtained from this method is a predicted value of binding affinity with a specific protein. Therefore, additional experiments are required to determine how the binding affinity contributes to biological activity.

[0004] AI-based drug discovery, focusing on the physiological activity of drugs—that is, drug discovery methods that directly predict quantitative structure-activity relationships—has been under consideration for many years. However, existing methods focus only on pharmacological activity against specific proteins, and their use is limited to searching for drugs with structures similar to existing drugs with known efficacy. This is because in vitro data is essential to actually confirm pharmacological activity or biological efficacy, and an astronomical amount of experiments is required to evaluate any compound.

[0005] As examples of drug discovery techniques using artificial intelligence (AI), for example, a method of graphing artificial compound data having molecular structure data and using a multilayer neural network (see Patent Document 1, etc.), and a method of specifying a target biopolymer and acquiring the three-dimensional structure of a compound for which binding is predicted, and then using machine learning to predict the binding between the three-dimensional structure of the biopolymer and the three-dimensional structure of the compound (see Patent Document 2, etc.) have been proposed. However, the data acquired in these cases is one-sided, and there is little connection between the compound information and the mechanism of action in the body. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2020-9203 [Patent Document 2] Japanese Patent Application Publication No. 2019-28879 Summary of the Invention [Problem to be solved by the invention]

[0007] Thereafter, the inventors have made extensive studies on integrating chemical information on compounds, information on their actions at the cellular level, and information on living organisms or clinical situations when predicting the structure of desired drugs.

[0008] The present invention has been made in consideration of the above points, and provides a prediction device, a prediction method, and a prediction program for more accurately and efficiently developing desired drugs and predicting the efficacy and side effects of the drugs by promoting the integration of chemical information of compounds, information on their actions at the cellular level, and biological or clinical information. [Means for solving the problem]

[0009] That is, the prediction device of the embodiment is characterized by including an acquisition unit that acquires chemical substance information and pharmacological information of the drug, an estimation unit that estimates estimated information of the drug by performing machine learning using the chemical substance information and the pharmacological information, and an output unit that predicts and outputs both the pharmacological effects and side effects of the drug on a living body by retraining a machine learning model based on the estimated information.

[0010] The acquisition unit of the prediction device may further acquire biological information of the drug at a cellular level, and the estimation unit may perform machine learning using the chemical substance information, pharmacological information, and biological information at a cellular level. Furthermore, the estimation unit may estimate characteristic information of the drug based on a chemical structure included in the chemical substance information and drug administration information included in the pharmacological information.

[0011] The estimation unit may estimate pharmacokinetic information of a drug based on the chemical structure included in the chemical substance information and the drug administration information included in the pharmacological information. Furthermore, the estimation unit may estimate biological information based on the chemical structure included in the chemical substance information and the drug administration information included in the pharmacological information.

[0012] Furthermore, the prediction device may be provided with a preliminary estimation unit that preliminary estimates the chemical structure of the drug based on the drug's characteristic information. Furthermore, when predicting the estimated information, the estimation unit may retrain the machine learning model using other estimated information. Additionally, the machine learning may be a neural network, and an autoencoder may be used.

[0013] Furthermore, in the prediction device, the chemical substance information of the drug may include chemical structure information and property information, the chemical structure information may be a feature based on the chemical structure of the drug, and the property information may be a feature based on the chemical properties and physical properties of the drug. Also, the biological information at the cell level may be a feature based on information on the behavior of the drug when administered to cultured cells.

[0014] Furthermore, the pharmacological information may include features based on the effects of the drug on the living body and the dynamics of the drug in the body, and drug administration information, and the drug administration information may be features based on the drug administration means and biological information at the time of drug administration.

[0015] Furthermore, when predicting and outputting both the efficacy and side effects of a drug on a living body, the output unit of the prediction device may also predict and output a drug administration plan.

[0016] Furthermore, the acquisition unit of the prediction device may include a crawling unit that acquires at least one of chemical substance information of a drug, cellular-level action information of a drug, or pharmacological information of a drug from a website on the Internet.

[0017] In addition, the prediction method of the embodiment is characterized in that the computer executes the following steps: an acquisition step of acquiring chemical substance information and pharmacological information of the drug; an estimation step of estimating estimated information of the drug by performing machine learning using the chemical substance information and the pharmacological information; and an output step of predicting and outputting both the efficacy and side effects of the drug on a living body by retraining a machine learning model of the machine learning based on the estimated information. Furthermore, the acquisition step in the prediction method may further acquire cellular-level biological information of the drug, and the estimation step may perform machine learning using the chemical substance information, pharmacological information, and cellular-level biological information.

[0018] In addition, the prediction program of the embodiment is characterized in that the computer is implemented with an acquisition function that acquires chemical substance information and pharmacological information of the drug, an estimation function that estimates estimated information of the drug by performing machine learning using the chemical substance information and the pharmacological information, and an output function that predicts and outputs both the pharmacological effect and side effects of the drug on a living body by retraining a machine learning model of the machine learning based on the estimated information. Furthermore, the acquisition function in the prediction program may further acquire cellular-level biological information of the drug, and the estimation function may perform machine learning using the chemical substance information, pharmacological information, and cellular-level biological information. [Effects of the Invention]

[0019] The prediction device of the present invention includes an acquisition unit that acquires chemical substance information and pharmacological information of a drug, an estimation unit that estimates estimated information about the drug by performing machine learning using the chemical substance information and pharmacological information, and an output unit that predicts and outputs both the drug's efficacy and side effects on a living body by retraining a machine learning model based on the estimated information.Therefore, by integrating the chemical information of a compound and its action information, it becomes possible to accurately and efficiently develop a desired drug, and further to predict the drug's efficacy and side effects.

[0020] Similarly, in prediction methods and programs, chemical information on compounds and information on their actions can be integrated to enable accurate and efficient drug discovery of desired drugs, and furthermore, prediction of the efficacy and side effects of those drugs. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a schematic diagram illustrating a configuration of a prediction device and its system according to an embodiment. [Figure 2] FIG. 2 is a block diagram showing the configuration of functional units of a computer. [Figure 3] FIG. 1 is a first schematic diagram illustrating a machine learning process. [Figure 4] FIG. 2 is a second schematic diagram showing the machine learning process. [Figure 5]FIG. 3 is a third schematic diagram showing the machine learning process. [Figure 6] FIG. 4 is a fourth schematic diagram showing the machine learning process. [Figure 7] FIG. 5 is a fifth schematic diagram showing the machine learning process. [Figure 8] 1 is a flowchart illustrating a prediction method according to an embodiment. [Figure 9] 10 is a flowchart illustrating a prediction method according to another embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0022] One problem with drug discovery using computational methods (in silico) is that the design process does not take into account the dynamics of a drug in the body, i.e., pharmacokinetics, or the events that a drug ultimately causes in the body (pharmacodynamics: efficacy and side effects) as the final output.In recent years, AI-based pharmacokinetic prediction systems known as ADMET prediction (Adsorption, Distribution, Metabolism, Excretion, Toxicity) have begun to be developed as a method for predicting adverse drug events.

[0023] However, existing prediction methods are limited in their predictive capabilities, such as the affinity of a compound for a specific drug transporter or for an enzyme that is likely to cause cardiotoxicity when inhibited. This is primarily due to the fact that building an AI system for ADMET prediction requires a huge amount of experimental data, and data is scarce for some items. Furthermore, ADMET prediction and efficacy prediction are independent systems, and there is currently no AI system that integrates both perspectives to effectively and directly predict the final in vivo pharmacodynamics of a drug.

[0024] The main reason why it is difficult to build an AI model that can accurately predict pharmacodynamics, the ultimate goal of drug discovery, is the need to comprehensively and quantitatively grasp factors such as how drugs interact with the in vivo drug transport and drug metabolism systems, how much a drug ultimately reaches and binds to target and non-target cells and proteins, and how this changes cells, and to integrate this with clinical information.In addition, the amount of clinical data that directly records drug efficacy and side effects is limited, and it is currently difficult to train a machine learning model using clinical data alone.

[0025] The prediction device, prediction method, and prediction program of the present embodiment begin by extracting the structure and chemical and physical properties of a drug (compound). Then, stepwise training is performed on cellular-level action information, ultimately leading to a prediction of in vivo efficacy and side effects (pharmacodynamics), the most important information in drug discovery. The prediction device, prediction method, and prediction program then enable the final trained model to quantitatively predict the efficacy and side effects of an unknown drug candidate based on relatively limited information, such as its structure, enabling the selection of promising drug candidates based on the predictions. Additionally, by reading the output of the intermediate layers of the machine learning model, pharmacokinetic information for the drug candidate (compound) can also be simultaneously obtained, which can be used to predict the dosage plan for the selected drug candidate. The candidate drug (compound) can include various types of small- or medium-molecular-weight organic compounds with molecular weights of approximately 100 to 5,000, as well as peptides, nucleic acids, metal-containing compounds, and other compounds. The method can also be extended to the development of macromolecular pharmaceuticals such as monoclonal antibodies.

[0026] In describing the prediction device, prediction method, and prediction program of the embodiment, the prediction device will be described together with the prediction method. The schematic diagram of FIG. 1 shows the configuration of the prediction device 1 of the embodiment. The prediction device 1 is connected to an internet line (not shown). Through the internet line, the prediction device 1 can acquire information published on websites by individuals, research institutions, public and private data centers, medical institutions, etc. Examples of information that can be acquired include various information related to drug discovery, such as the physical properties of compounds, biological experimental data, clinical data, and basic medical and clinical papers. It is also possible to use information and data entered by the user of the prediction device 1.

[0027] In FIG. 1, chemical substance information J1 of a drug, cellular-level biological information J2 of the drug on cells, and pharmacological information J3 of the drug on a living organism are input to a prediction device 1. Then, the prediction device 1 (prediction method, prediction program) uses machine learning trained using each of these pieces of information to output information useful for clinical medicine, such as the drug's efficacy and side effects K1 and even a medication plan K2. The medication plan K2 will be described later. The prediction device 1 includes a computer 10, and is equipped with hardware such as a processing element 11 (such as a GPU), a ROM 12, a RAM 13, and a storage unit 14. The computer 10 of the prediction device 1 is composed of various electronic computers (computational resources), such as a supercomputer, a mainframe, a workstation, or a cloud computing system.

[0028] When each functional unit of the computer 10 of the prediction device 1 in Fig. 1 is realized by software, the computer 10 is realized by executing instructions of a program, which is software that realizes each function. The recording medium that stores this program can be a "non-transitory tangible medium," such as a CD, DVD, semiconductor memory, or programmable logic circuit. In addition, this program may be supplied to the computer 10 of the prediction device 1 via any transmission medium (communication network, broadcast wave, etc.) that can transmit the program.

[0029] The storage unit of the computer 10 of the prediction device 1 is provided with a storage device such as an HDD or SSD. The storage unit 14 may be an external server (not shown). The storage unit 14 stores various data, information, a prediction program, various data required to execute the program, and the like. Furthermore, each functional unit that performs various calculations, operations, and the like is a computing element 11. In addition, input devices (not shown) such as a keyboard and a mouse, a display unit (display device such as a monitor), an output device that outputs data, and the like may also be appropriately connected to the prediction device 1 (computer 10).

[0030] The functional units in the processing element 11 of the computer 10 of the prediction device 1 are shown in the schematic block diagram of Fig. 2. The functional units include a crawling unit 100, an acquisition unit 110, a preliminary estimation unit 120, an estimation unit 130, an output unit 140, etc. Processing and execution in the prediction device 1 are realized in software terms by a prediction program loaded into main memory, etc.

[0031] The acquisition unit 110 acquires chemical substance information J1 of the drug and pharmacological information J3 of the drug. In addition to these, the acquisition unit 110 also acquires cellular-level biological information J2 of the drug (see FIG. 1). The information J1, J2, and J3 are acquired via the crawling unit 100, which will be described later, or by the user of the prediction device 1 inputting the necessary information.

[0032] The drug chemical substance information J1 is a feature based on at least the chemical structure of the drug. The information included in this drug chemical substance information J1 includes chemical structure information and chemical and physical property information (characteristic information). The drug chemical and physical property information (characteristic information) includes, for example, the molecular weight, charge, LogP, pKa, hydrogen bonding, molecular orbital, solubility, and membrane permeation rate of the drug (compound). In addition, the drug chemical substance information J1 includes feature amounts that can be calculated based on the two-dimensional and three-dimensional structures of the compound related to at least the drug's chemical structure. For example, feature amounts based on the two-dimensional structure, such as the number of atoms and a structural fragment list, and feature amounts based on the three-dimensional structure, such as the molecular volume, molecular surface area, bond length, and rotational degrees of freedom, are included. Therefore, the chemical substance information J1 is generated as feature amounts based on the chemical structure information and characteristic information. The chemical structure information is a feature amount based on the chemical structure of the drug. The acquisition unit 110 acquires these various properties.

[0033] Cellular-level biological information J2 of drugs is primarily a quantified indicator of the behavior and effects of compounds on cells. It is generated as a feature based on information obtained when a drug is administered to cultured cells. This information can be obtained in vitro and is broader than information obtained purely through biochemistry. For example, (in the case of intracellular targets) it includes the cellular bioavailability of the drug compound, as well as molecular biological information such as epigenetic changes in DNA, RNA expression, and increases or decreases in transcription factors after drug administration, as well as comprehensive (omics) changes in protein expression and intracellular metabolic products. It also includes macroscopic cell biological events (e.g., cell size after culture, cytokine production, and apoptosis).

[0034] Pharmacological information J3 is information on the in vivo action of a drug, as well as information on the effects of a drug when administered to a living body (drug administration information). Information on the in vivo action of a drug is obtained from so-called clinical data (in vivo data), and specifically includes information on the pharmacokinetics of a drug in the body, as well as drug efficacy and side effects (pharmacodynamics).

[0035] Drug administration information is a value that affects the pharmacokinetics, efficacy, and side effects of a drug when it is administered into the body. Pharmacokinetics in the body, efficacy, and side effects vary depending on the drug's administration method (dosage, administration route, etc.). In addition, although clinical data does not directly describe values, the in vivo environment (e.g., target extracellular tissue information) of a patient with a disease differs from that of a healthy subject, which affects pharmacokinetic information, efficacy, and side effects. Therefore, information that affects pharmacokinetic information, efficacy, and side effects is included in pharmacological information J3 as drug administration information. Specifically, this information includes the drug's administration method (dosage, administration route, etc.) and biological information at the time of drug administration (age, gender, extracellular tissue information, etc.).

[0036] When the acquisition unit 110 acquires data, the acquisition unit 110 includes a crawling unit 100, which enables automated acquisition. The crawling unit 100 acquires at least one of, or two or all of, chemical substance information J1 of a drug, biological information J2 of a drug at the cell level, and pharmacological information J3 of a drug from websites on the Internet. Websites contain information made public by various facilities such as research institutes, national and public / private data centers, and medical institutions. The acquisition unit 110 may also use information and data obtained by a user.

[0037] In addition to chemical and physical knowledge and papers about drugs, knowledge at the cellular and clinical levels, papers, and reference values ​​for administration are constantly increasing. Therefore, in order to constantly obtain the latest information and data and update the machine learning model, the computer 10 of the prediction device 1 obtains public information on websites via the crawling unit 100 and stores it in the storage unit 14. Note that the user of the prediction device 1 may also store the information and data in the storage unit 14.

[0038] The preliminary estimation unit 120 preliminarily estimates the chemical structure of the drug based on the drug characteristic information (see FIG. 3). The drug characteristic information is acquired via the crawling unit 100 or by the user of the prediction device 1 inputting the necessary information.

[0039] The preliminary estimation unit 120 serves to preliminarily estimate the chemical structure of the drug in the subsequent estimation unit 130 based on the characteristic information. Specifically, the characteristic information includes, for example, chemical structure information including the surface area and molecular weight of the drug compound, charge information including electron density, polarizability, acidity, hydrogen bonding, hydrophilicity, hydrophobicity, molecular orbital energy level, membrane permeation rate, solubility, etc., as well as the rate of change of the drug's chemical structure and characteristic information under a specific environment, for example, the rate of change of the molecular structure and charge distribution when interacting with polar groups such as amino acids and phosphates and non-polar groups such as alkyl groups and phenyl groups. In addition, the chemical substance information includes, along with the characteristic information, feature quantities that can be calculated based on the two-dimensional and three-dimensional structures of the compound related to at least the drug's chemical structure, such as the number of atoms and a structural fragment list based on the two-dimensional structure, and chemical structure information such as molecular volume, molecular surface area, bond length, and rotational degrees of freedom based on the three-dimensional structure.

[0040] In the schematic diagram of the machine learning process in Figure 3, when selecting the molecular structure of a drug, feature quantities based on the chemical and physical properties of the drug (drug characteristic information) are incorporated into the input layer in the diagram. Note that the machine learning process in the preliminary estimation unit 120 may be incorporated optionally depending on the process, such as when the selection of the molecular structure of a drug is required.

[0041] Based on the combination of input values, the chemical structure of the drug is output by machine learning. The output format may be, for example, SMILES (Simplified Molecular Input Line Entry System) notation. A neural network or the like is used for machine learning in the preliminary estimation unit 120, the estimation unit 130, and the output unit 140.

[0042] Here, when machine learning is performed in the computing element, machine learning is performed using a neural network or the like. Other machine learning methods include association rule learning, random forest, support vector machine, etc. Data processing using a programming language and natural language processing can be added to the information acquired through the crawling unit 100.

[0043] If the prediction accuracy does not reach the target value after training using the above input items, the autoencoder or other device is retrained using chemical structures, and additional latent features are extracted to compensate for the insufficient predictions. These automatically extracted features are added as input values ​​to the input value items, and the machine learning model is retrained to improve prediction accuracy. In this way, preliminary network training in machine learning is performed.

[0044] The estimation unit 130 estimates estimated information about a drug by performing machine learning using the chemical substance information J1 and pharmacological information J3. Furthermore, the estimation unit 130 estimates estimated information about a drug by performing machine learning using the chemical substance information J1, pharmacological information J3, and cellular-level biological information J2. To improve the accuracy of machine learning in the output unit 140, it is desirable to add pharmacokinetic information, cellular-level biological information, and estimated information in addition to characteristic information, as in this embodiment. Additionally, the estimation unit 130 estimates characteristic information about a drug based on the chemical structure included in the chemical substance information J1 and the drug administration information included in the pharmacological information J3, estimates pharmacokinetic information about a drug based on the chemical structure included in the chemical substance information J1 and the drug administration information included in the pharmacological information J3, and estimates biological information J2 based on the chemical structure included in the chemical substance information J1 and the drug administration information included in the pharmacological information J3. Furthermore, when predicting estimated information, the estimation unit 130 retrains the machine learning model using additional estimated information.

[0045] The schematic diagram of the machine learning process in Figure 4 shows the process of outputting drug characteristic information from limited chemical substance information of the drug. Specifically, the chemical structure of the drug is input, and the feature values ​​of each item of the drug characteristic information are output. In this case, as shown in the schematic diagram in Figure 4, feature values ​​of drug administration information such as drug dose, administration route, and target tissue information may be added to the input values.

[0046] In the machine learning processing stage of Figure 4, the network is trained to calculate the same output value even when the feature values ​​of the drug administration information on the input layer are changed appropriately. The reason for this is that the purpose of this stage is to extract feature values ​​based on the pure chemical and physical properties of the drug, and to have the machine learning model learn input-output relationships that are independent of input values ​​such as drug dosage, administration route, and target tissue information. Note that the network in the schematic diagram of Figure 4 differs from the preliminary network in Figure 3 in that the weights are randomly initialized before training.

[0047] As shown in the schematic diagram of the machine learning model in Figure 4, a perspective that takes into account the administration of drugs to living organisms, such as drug dosage, administration route, and target tissue information, has been added to the previous machine learning that only associated drug property information and chemical structure (chemical substance information). As a result, the machine learning model will be trained taking into account not only the chemical structure but also the aforementioned administration of drugs to living organisms, thereby improving prediction accuracy.

[0048] Next, to improve the prediction accuracy of the output section, training of a machine learning model for estimating the in vivo pharmacokinetics of a drug may be performed using chemical substance information and pharmacological information of the drug as input. In this case, when retraining the machine learning model of Figure 4, the only chemical substance information required for the input data is the chemical structure. Furthermore, as shown in the schematic diagram of Figure 5, the output layer outputs the same items as the output (in silico) of the chemical data processing layer in Figure 4, in addition to items related to the pharmacokinetics of the drug (features based on information on the behavior of the drug when administered to a living body).

[0049] The reason for this is that when a machine learning model trained to output pharmacokinetics undergoes transfer learning as shown in Figure 6 below, the output (characteristic information) of the first layer is also directly used to facilitate output prediction. To accommodate changes in the characteristic information of a drug under a specific environment (e.g., acidic conditions), the output characteristic information differs from the output shown in Figure 4 above by retraining the in silico output items to reflect changes in target tissue information, which is one of the input values. After this retraining, the learnable parameters of all or some of the units in the first layer may be frozen, and these values ​​may not be updated during subsequent training.

[0050] In training to estimate pharmacokinetics, as shown in the schematic diagram in Figure 5, based on input values ​​for "drug chemical substance information (chemical structure)," "drug dosage," and "route of administration," values ​​of pharmacokinetic-related features such as "bioavailability (an index of how much of an administered drug can reach the bloodstream)," "maximum blood concentration," "area under the blood concentration-time curve," "volume of distribution," "renal clearance," and "hepatic metabolic rate" are output as features based on information on the behavior of the drug when administered to a living body. Here, the output values ​​related to pharmacokinetics are trained so that input values ​​not related to pharmacokinetics (target tissue information in Figure 5) are output consistently regardless of the value they take.

[0051] Regarding bioavailability, the system learns how bioavailability varies depending on the route of administration of the compound used as a drug. The "Route of Administration" field is entered as 0 (intravenous), 1 (oral), 2 (rectal), etc. For example, when intravenous is entered, the bioavailability is 100%, but for other routes, for example, 1% for oral and 50% for rectal are output.

[0052] Subsequently, to further improve the prediction accuracy of the output unit 140, training of a machine learning model may be performed on multiple cellular-level biological information (so-called in vitro data) for the drug's effect on cells. Here, the data is broadly divided into three stages: "cellular bioavailability," "intermediate data including molecular biological data," and "macroscopic biological data." These are feature items based on information about the behavior of a drug when administered to cultured cells. The order of this training is based on biological logic; for example, training on RNA expression is performed before protein expression. Note that the order of training may be reversed, and machine learning does not necessarily require all stages.

[0053] The main purpose of training using multiple cell-level biological information (in vitro data) is to reduce the amount of data required for training the next output layer (see the final drug efficacy and side effect prediction layer (in vivo) in Figure 7). For this reason, the training content is adjusted appropriately so that the final drug efficacy and side effect predictions made by the machine learning model can achieve sufficient accuracy.

[0054] Here too, chemical substance information and pharmacological information of the drug are input, and cellular-level biological information of the drug, i.e., information on its behavior when administered to cultured cells, is estimated. The schematic diagram of the machine learning process in Figure 6 shows the process of outputting cellular-level biological information by retraining the machine learning model that estimates the characteristic information and pharmacokinetic information in Figure 5 above.

[0055] Obviously, in the feature quantities (in vitro data) corresponding to the cellular biological information of a drug on cells, the amount of drug added during culture generally does not match the dosage item. Therefore, the dosage corresponding to each data item can be estimated in advance using the following method. For example, after training of the model shown in the schematic diagram in Figure 5 is completed, the chemical structure, administration route, and target tissue information used when acquiring the in vitro data are fixed. Then, the dose value is varied, and changes in the pharmacokinetic information value are examined. Next, a dose is searched for such that the drug concentration in the culture medium used in the in vitro system is the same as the extracellular fluid concentration of the drug in the tissue estimated from the pharmacokinetic prediction model. In this case, the composition information of the culture medium is input into the fixed target tissue information. Because the pharmacokinetic information value varies depending on the administration route, the dosage is estimated independently for each input value (0: intravenous, 1: oral, 2: rectal, etc.).

[0056] Cellular bioavailability, shown in Layer 3a of the schematic diagram in Figure 6, is a value indicating how much of a drug compound penetrates into different cells, and values ​​such as the amount of drug in the cells / the amount of drug in the culture medium are used. Measurements are taken of a wide variety of cells over a specific culture period under various dosages / culture conditions (target tissue information). In addition to cellular bioavailability, the output of the first layer, "characteristic information," and the output of the second layer, "pharmacokinetic information," are also output under the target tissue information.

[0057] In the schematic diagram of Figure 6, layer 3b acquires molecular biology data and comprehensive (omics) data on changes in protein expression, etc. Examples include RNA expression in cells and transcription factor expression. RNA expression information of various cells measured over a specific culture period and under various dosages / culture conditions (target tissue information) is learned. This training is performed after training on cellular bioavailability in layer 3a.

[0058] Macroscopic biological characteristics are used as data for multiple change features (in vitro data) in the third layer c of the schematic diagram in Figure 6. For example, the size of cells after culture, the viability after cell culture, protein expression inside and outside the cells that can be obtained by measurements such as a flow cytometer, and the amount of cytokine production by each cell are trained as output values ​​for the features.

[0059] The output unit 140 receives the estimated information from the estimation unit 130 and performs machine learning to retrain a machine learning model (machine learning model), thereby predicting and outputting the efficacy and side effects of the drug on a living body. Of course, depending on the settings, the output unit 140 may be configured to output the prediction of either the efficacy or side effects.

[0060] The processing in the output unit 140 is based on all of the output values ​​from each layer up to this point, and outputs the drug's efficacy and side effects K1, which is the main objective, and also the drug's dosage plan K2 as information accompanying the efficacy and side effects (see Figure 1). In the schematic diagram of Figure 7, the machine learning model trained in the estimation unit 130 is retrained using clinical data, and the drug's efficacy and side effects are output. The clinical data is obtained from all or part of the data published on a website. The crawling unit 100 mentioned above is used to obtain the data.

[0061] As with previous machine learning processes, the input values ​​are the drug's chemical structure and drug administration information. The schematic diagram in Figure 7 uses "drug dosage," "administration route," and "target tissue information." The drug dosage can be the dose per compound or the total dose over the administration period. "Target tissue information" is input based on the extracellular environment of the tissue, as estimated from the patient's main complaint. Because this data cannot be obtained from published clinical trials, physiological values ​​can be used for convenience. However, changes in the extracellular fluid of the target tissue may occur depending on the disease. For example, in the case of severe inflammation in the tissue, changes in the external environment also vary depending on the severity of the inflammation, so a machine learning model with a certain degree of flexibility is constructed. Other factors such as age and medical history can be considered in the model, but as shown in the figure, a model that does not consider these factors can also be used. Although this can be problematic, as with administration period, ignoring these factors in machine learning models allows for a larger data volume.

[0062] One output is the ratio of "people who have confirmed efficacy (or side effects)" to "the number of clinical trial participants in the clinical trial." The direct meaning of "drug efficacy" is the remission rate of the disease. However, focusing on this value makes it difficult to grasp the overall effect of the drug compound on the body. Therefore, the output value also includes data that indirectly indicate disease remission, obtained from blood tests and other tests typically performed in clinical trials. For example, in the case of rheumatoid arthritis, in addition to the data on "rheumatoid arthritis remission rate," test results for CRP, a commonly tested inflammatory marker, are added as an item called "CRP reduction rate." Additionally, output forms of side effects include numerical indicators that suggest the likelihood of side effects being confirmed per unit number of people, such as "occurrence of headache after administration: 80%," "rash: 50%," and "probability of abnormal values ​​in liver and biliary markers: 20%."

[0063] In this way, even drugs used for "rheumatoid arthritis" but not for "systemic lupus erythematosus" can be consistent in the category of a decrease in CRP, and both can be interpreted as drugs that "suppress inflammation." Regarding side effects, both subjective items such as "headache," as in the example above, and objective items such as blood markers like renal markers can be used. Furthermore, items such as "decreased blood pressure" that are a medicinal effect in the case of hypertension medications and also a side effect in the case of anaphylactic shock are not standardized, and separate items are created for the medicinal effect, "decreased blood pressure," and the side effect, "decreased blood pressure."

[0064] Because efficacy and side effects are so diverse, there is concern that there may be missing values ​​in the data for each drug in the clinical database. To compensate for these missing values, the drug data in the clinical database is input into the trained machine learning model (the machine learning model in the estimation section), and a correlation coefficient between all drugs is calculated based on all or part of the output values. This is then used to create a similarity index between drugs. When a missing value is identified for a drug (compound), it is possible to compensate for the missing value using the effects recorded in the clinical database of another drug (compound) that shows a high correlation with the drug (compound).

[0065] As explained above, the prediction device 1 (and its prediction method and prediction program) can use machine learning to simulate and predict either the efficacy or side effects, or both, of a candidate drug based on its chemical structure. In other words, the prediction device 1 (and its prediction method and prediction program) learns chemical information (chemical substance information about the drug), biological information (biological information about the drug at the cellular level), and clinical information (pharmacological information about the drug in the living body), thereby predicting not only the drug's chemical properties but also its cellular action, efficacy, and side effects in the living body, based solely on the drug's chemical structure. Furthermore, because pharmacokinetic values ​​are calculated during this process, it is also possible to predict the drug's dosing plan. The dosing plan here indicates the route of administration, frequency of administration, and dosage. Even for conditions for which there is currently no cure, the efficacy of the candidate drug can be estimated to a certain extent based on the side effects of the candidate drug, as well as biological information at the cellular level and blood markers. In this way, the chemical structure of a candidate drug with potential efficacy can be estimated without actually synthesizing the candidate drug, and it is even possible to predict the efficacy, side effects, and dosage plan of the candidate drug. This can greatly contribute to simplifying the screening process in new drug development (drug discovery) and selecting the direction of synthesis of candidate drugs.

[0066] An example of drug discovery based on prediction of the efficacy and side effects of an unknown drug using a machine learning model trained using the prediction device 1 is summarized below. First, compounds are randomly selected to create a comprehensive list of candidate drugs. At this time, compounds that are difficult to synthesize based on existing methods for evaluating the difficulty of synthesis will be excluded from List-1. All types of compounds are then input into the trained network of the machine learning model, and the remission rate (efficacy) and side effects of the disease being cured for each compound are obtained. The obtained efficacy and side effects are compared with the standard values, and drugs whose efficacy falls below the standard or whose side effects exceed the standard are excluded, and List-2 is generated with the remaining drugs as candidates. After that, drugs with partially modified chemical structures of each drug in List-2 are mass-produced, drugs that are difficult to synthesize are excluded, and the efficacy and side effects of the remaining drugs are predicted. Drugs that do not meet the standard values ​​are excluded, and List-3 is generated based on the remaining drugs. The compound newly added to List 3 is partially modified and evaluated in the same way. This process of partially modifying and evaluating this compound is repeated in an attempt to iteratively optimize efficacy and side effects, and when no significant improvement in efficacy or side effects is observed for any drug on the list, the repeated process is terminated, and finally the drugs on the list are ranked, with the top-ranked drug selected as the final candidate drug.

[0067] It is considered appropriate to prioritize input of information such as 0 (intravenous) for the drug administration route. However, the value for dose (administration amount) has not been determined. The value is adjusted as appropriate, and the changes in efficacy and side effects are verified to determine a realistic dose range. Only drugs that exceed the standard efficacy and fall below the side effect standard are retained on the list as drug candidates. Regarding the administration route, it is possible to change the administration route conditions depending on the purpose, such as when oral medication is required in addition to intravenous administration, and evaluate efficacy and side effects.

[0068] In drug discovery based on prediction of the efficacy and side effects of unknown drugs using a trained machine learning model using the prediction device 1 of the embodiment, multiple candidate drugs can be organized in order of efficacy and side effects. This makes it easier to determine the priority of which candidate drugs should be selected from multiple candidate drugs and submitted to each phase of testing. Since the efficacy and side effects of candidate drugs can be predicted in advance, candidate drugs can be prioritized, improving the efficiency of planning and managing pharmaceutical trials, etc.

[0069] Hereinafter, the prediction method and prediction program in the computer 10 of the prediction device 1 will be described with reference to the flowcharts in Figs. 8 and 9. The prediction method is executed by the processing element 11 of the computer based on the prediction program. The prediction program causes the computer 10 of Figs. 1 and 2 to execute a crawling function, an acquisition function, a preliminary estimation function, an estimation function, an output function, etc. Note that each function overlaps with the description of the prediction device 1 described above, so details will be omitted.

[0070] As can be seen from the flowchart in FIG. 8, the processing of the arithmetic element 11 of the computer 10 includes various steps, such as a crawling step (S100), an acquisition step (S110), an estimation step (S130), and an output step (S140). Of course, various steps necessary for the operation of the computer 10 itself are naturally included. In the flowchart shown in FIG. 8, the processing of the crawling step (S100) is not essential. As mentioned above, if the user of the prediction device 1 manually inputs information and data, the processing of the crawling step (S100) can be omitted.

[0071] The crawling function acquires at least one of chemical substance information of a drug, cellular-level biological information of a drug on cells, or pharmacological information of a drug from a website on the Internet (S100; crawling step). The acquisition function acquires chemical substance information of a drug and pharmacological information of a drug (S110; acquisition step). The acquisition function can also acquire chemical substance information of a drug, pharmacological information of a drug, and cellular-level biological information of a drug. The estimation function estimates estimated information of a drug by performing machine learning using the chemical substance information and pharmacological information (S130; estimation step). The estimation function can also estimate estimated information of a drug by performing machine learning using the chemical substance information, pharmacological information, and cellular-level biological information. The output function retrains a machine learning model based on the estimated information, thereby predicting and outputting both the efficacy and side effects of the drug on a living body (S140; output step).

[0072] Regarding another embodiment of the prediction method and prediction program in the computer 10 of the prediction device 1, in the flowchart of FIG. 9, the processing of the arithmetic element 11 of the computer 10 includes various steps, such as a crawling step (S100), an acquisition step (S110), a preliminary estimation step (S120), an estimation step (S130), and an output step (S140). Naturally, various steps necessary for the operation of the computer 10 itself are naturally included. Here, a preliminary estimation function is provided. The preliminary estimation function preliminarily estimates the chemical structure of a drug based on characteristic information (S120; preliminary estimation step). Even in the flowchart shown in FIG. 9, the processing of the crawling step (S100) is not essential. As mentioned above, if information and data are input by the user of the prediction device 1, the processing of the crawling step (S100) can be omitted.

[0073] The computer program of the present invention described above may be recorded on a processor-readable recording medium, and the recording medium may be a "non-transitory tangible medium" such as a disk, card, semiconductor memory, programmable logic circuit, etc.

[0074] The computer program can be implemented using, for example, a scripting language such as ActionScript or JavaScript (registered trademark), an object-oriented programming language such as Objective-C or Java (registered trademark), or a markup language such as HTML5. [Explanation of symbols]

[0075] 1 Prediction device 10. Computers 11 Computing elements 12 ROM 13 RAM 14 Storage section 100 Crawling Section 110 Acquisition Department 120 Preliminary Estimation Section 130 Estimation part 140 Output section J1 Drug Chemical Substance Information J2 Cellular level biological information of drugs J3 Pharmacological information of drugs K1 Drug efficacy and side effects K2 Medication Plan

Claims

1. an acquisition unit that acquires chemical substance information of a drug and pharmacological information of the drug; an estimation unit that estimates estimated information about the drug by performing machine learning using the chemical substance information and the pharmacological information; and an output unit that predicts and outputs both the efficacy and side effects of the drug on a living body by retraining the machine learning model based on the estimated information. A prediction device characterized by:

2. The acquisition unit further acquires biological information of the drug at a cellular level, The prediction device according to claim 1 , wherein the estimation unit performs machine learning using the chemical substance information, the pharmacological information, and the cellular level biological information.

3. The prediction device according to claim 1 , wherein the estimation unit estimates the drug characteristic information based on a chemical structure included in the chemical substance information and drug administration information included in the pharmacological information.

4. The prediction device according to claim 1 , wherein the estimation unit estimates the pharmacokinetic information of the drug based on a chemical structure included in the chemical substance information and drug administration information included in the pharmacological information.

5. The prediction device according to claim 2 , wherein the estimation unit estimates the biological information at the cell level based on a chemical structure included in the chemical substance information and drug administration information included in the pharmacological information.

6. The prediction device according to claim 3 , further comprising a preliminary estimation unit that preliminarily estimates the chemical structure of the drug based on the characteristic information of the drug.

7. The prediction device according to claim 1 , wherein the estimation unit retrains the machine learning model using other estimated information when predicting the estimated information.

8. The prediction device according to claim 1 , wherein the machine learning is a neural network and an autoencoder is used.

9. The prediction device described in claim 3, wherein the chemical substance information of the drug includes chemical structure information and the characteristic information, the chemical structure information is a feature based on the chemical structure of the drug, and the characteristic information is a feature based on the chemical properties and physical properties of the drug.

10. The prediction device according to claim 2 , wherein the biological information at the cell level is a feature quantity based on information on behavior of the drug when administered to cultured cells.

11. The prediction device according to claim 3, wherein the pharmacological information includes features based on the effect of the drug on the living body and the dynamics of the drug in the body, and the drug administration information, and the drug administration information is features based on the drug administration means and biological information at the time of drug administration.

12. The prediction device according to claim 1 , wherein the output unit predicts and outputs a medication plan for the drug when predicting and outputting both the efficacy and side effects of the drug on the living body.

13. The prediction device according to claim 2, wherein the acquisition unit includes a crawling unit that acquires at least one of chemical substance information of the drug, cellular level biological information of the drug, or pharmacological information of the drug from a website on the Internet.

14. The computer an acquisition step of acquiring chemical substance information of a drug and pharmacological information of the drug; an estimation step of estimating estimated information about the drug by performing machine learning using the chemical substance information and the pharmacological information; and an output step of predicting and outputting both the efficacy and side effects of the drug on a living body by retraining the machine learning model based on the estimated information. A prediction method characterized by:

15. The obtaining step further obtains biological information of the drug at a cellular level, The prediction method according to claim 14 , wherein the estimation step performs machine learning using the chemical substance information, the pharmacological information, and the cellular level biological information.

16. On the computer, an acquisition function for acquiring chemical substance information of a drug and pharmacological information of the drug; an estimation function of estimating estimated information about the drug by performing machine learning using the chemical substance information and the pharmacological information; and an output function for predicting and outputting both the efficacy and side effects of the drug on a living body by retraining the machine learning model based on the estimated information. A prediction program characterized by:

17. The acquisition function further acquires biological information of the drug at a cellular level, The prediction program according to claim 16 , wherein the estimation function performs machine learning using the chemical substance information, the pharmacological information, and the cellular level biological information.

Citation Information

Patent Citations

  • Connectivity prediction method, apparatus, program, recording medium, and production method of machine learning algorithm

    JP2019028879A

  • Deep layer learning method and apparatus of chemical compound characteristic prediction using artificial chemical compound data, as well as chemical compound characteristic prediction method and apparatus

    JP2020009203A