Prediction device, prediction method, and prediction program

By integrating the chemical information of compounds, cell-level action information and bio-organism information, and using machine learning models to predict the efficacy and side effects of the agent, the problem of relying on experience in traditional drug research and development is solved, and efficient drug research and development and accurate drug efficacy prediction are achieved.

CN120266136APending Publication Date: 2025-07-04SANLIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280102162.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional drug research and development relies on developers' experience, has low success rate and high cost. Existing artificial intelligence-based methods cannot accurately predict the efficacy and side effects of drugs in the body, especially insufficient predictions for membrane proteins and drug transporters with unresolved three-dimensional structures.

Method used

By integrating chemical information of compounds, cell-level action information and biologic information, machine learning models are used to predict the efficacy and side effects of the agent, including obtaining chemical substance information, pharmacological information and cell-level biological information of the agent, and conducting multi-level information training to predict pharmacokinetics and pharmacokinetics.

Benefits of technology

It realizes accurate prediction of the efficacy and side effects of drugs in the body, simplifies the process of new drug development, improves the efficiency and success rate of drug research and development, and is suitable for the development of low-molecular, medium-molecular compounds and polymer pharmaceuticals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266136A_ABST
    Figure CN120266136A_ABST
Patent Text Reader

Abstract

The purpose of the present invention is to integrate chemical substance information of a compound and information obtained when a medicament is administered to cells, promote biological or clinical information integration, accurately and efficiently implement research and development of a desired medicament, and predict the efficacy and side effects of the medicament. To this end, provided is a prediction device comprising: an acquisition unit that acquires chemical substance information and pharmacological information of a drug; an estimation unit that estimates estimation information of the drug by performing machine learning using the acquired chemical substance information and pharmacological information; and an output unit for retraining the machine learning model on the basis of the estimation information, thereby predicting and outputting the efficacy and side effects of the drug on the living body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a prediction device, a prediction method, and a prediction program. Specifically, the present invention relates to a device, a method, and a program for predicting a drug that exhibits the best drug efficacy by performing machine learning based on the properties of a drug and predicting the drug efficacy of the drug. Background Art

[0002] In conventional drug development methods, based mainly on past knowledge and experience, developers predict and generate effective pharmaceutical compounds and evaluate the drug efficacy and side effects of the pharmaceutical compounds through experimental verification. However, in conventional methods, the success or failure of drug development often depends on the proficiency and expertise of the developers and often relies on accidental discoveries. Therefore, the success rate of developing new drugs is extremely low, and developing a new drug requires a large amount of cost and time.

[0003] In order to reduce the huge development cost in conventional drug development, in recent years, drug development methods using artificial intelligence (AI) have received extensive attention. For example, the molecular docking method is used to predict the binding affinity between a target protein and a drug candidate compound, and a method for predicting an optimal pharmaceutical compound based on its highest binding degree is proposed. However, this method requires accurate identification of the three-dimensional structure of the target protein and its drug binding site for prediction. Therefore, this method cannot be applied to membrane proteins such as drug transporters whose three-dimensional structures have not been accurately analyzed. In addition, the numerical values obtained by this method are only prediction values of the binding affinity with a specific protein. Therefore, further experiments are still required to determine the degree of association with biological activity.

[0004] In recent years, a drug development method that focuses on the physiological activity of a drug and uses artificial intelligence (AI), that is, a drug development method for directly predicting the quantitative structure-activity relationship, has been studied for a long time. However, existing methods only focus on the pharmacological activity against a specific protein, and their application scope is limited to finding candidate drugs similar in structure to drugs known to have drug efficacy. This is because in vitro experimental data is indispensable for actually confirming pharmacological activity or biological effects, and to evaluate all compounds, an astronomical number of experiments need to be performed.

[0005] As one of the methods for drug development using artificial intelligence (AI), for example, there has been proposed a method of graphing artificial compound data having molecular structure data and using a multi-layer neural network (see Patent Document 1, etc.); and a method of obtaining the three-dimensional structure of a specified target biopolymer and a compound predicting its binding property, and predicting the binding property between the three-dimensional structure of the biopolymer and the three-dimensional structure of the compound by machine learning (see Patent Document 2, etc.). However, the data obtained by these methods is one-sided, and the correlation between compound information and its mechanism of action in vivo is weak. Prior Art Documents Patent Documents

[0006] Patent Document 1: Japanese Patent Laid-Open No. 2020-09203 Patent Document 2: Japanese Patent Laid-Open No. 2019-28879 Summary of the Invention

Problems to be Solved by the Invention

[0007] After that, the inventors et al. deeply explored a method of integrating chemical information of compounds, action information at the cell level, and relevant information in vivo or in clinical practice during the process of predicting the structure of the required drug.

[0008] The present invention is proposed in view of the above problems, aiming to promote the integration of chemical information of compounds, action information at the cell level, and information in vivo or in clinical practice, so as to more accurately and efficiently achieve drug development of the required drug, and further provide a prediction device, a prediction method, and a prediction program for predicting the efficacy and side effects of the drug.

Means for Solving the Problems

[0009] That is, the prediction device of the present embodiment is characterized by including: an acquisition unit that acquires chemical substance information of a drug and pharmacological information of the drug; a estimation unit that performs machine learning by using the chemical substance information and the pharmacological information, thereby estimating estimation information of the drug; and an output unit that retrains a machine learning model based on the estimation information, thereby predicting and outputting the efficacy and side effects of the drug on a living body.

[0010] In addition, the acquisition unit of the prediction device may further acquire biological information of the drug at the cell level; the estimation unit uses the chemical substance information, the pharmacological information, and the biological information at the cell level for machine learning. Further, the estimation unit may also estimate characteristic information of the drug based on the chemical structure included in the chemical substance information and the drug administration information included in the pharmacological information.

[0011] In addition, the estimation unit estimates the pharmacokinetic information of the drug based on the chemical structure included in the chemical substance information and the drug administration information included in the pharmacological information. Further, the estimation unit estimates the biological information based on the chemical structure included in the chemical substance information and the drug administration information included in the pharmacological information.

[0012] Further, the prediction device further includes a pre-estimation unit for pre-estimating the chemical structure of the drug based on the characteristic information of the drug. In addition, when estimating the prediction estimation information, the estimation unit retrains the machine learning model using other estimation information. In addition, the machine learning is a neural network and an autoencoder is adopted.

[0013] Further, in the prediction device, the chemical substance information of the drug includes chemical structure information and characteristic information. The chemical structure information is a feature quantity based on the chemical structure of the drug, and the characteristic information is a feature quantity based on the chemical and physical properties of the drug. The biological information at the cell level is a feature quantity based on the behavioral information when the drug is administered to cultured cells.

[0014] Further, the pharmacological information includes the characteristic quantities based on the effects of the drug on the living body and its pharmacokinetics in the body, and the drug administration information, which is a feature quantity based on the drug administration method and the living body information at the time of administration.

[0015] Further, when predicting and outputting the efficacy and side effects of the drug on the living body, the output unit of the prediction device also predicts and outputs the drug administration plan.

[0016] Further, the acquisition unit includes a crawling unit for acquiring at least one of the chemical substance information of the drug, the biological information of the drug at the cell level, or the pharmacological information of the drug from a WEB site on the Internet.

[0017] In addition, the prediction method of the present embodiment is characterized in that a computer executes the following steps: an acquisition step of acquiring the chemical substance information of the drug and the pharmacological information of the drug; an estimation step of estimating the estimation information of the drug by performing machine learning using the chemical substance information and the pharmacological information; and an output step of retrainging the machine learning model based on the estimation information to predict and output the efficacy and side effects of the drug on the living body. Further, the acquisition step in the prediction method further includes acquiring the biological information of the drug at the cell level, and the estimation step uses the chemical substance information, the pharmacological information, and the biological information at the cell level for machine learning.

[0018] In addition, the prediction program of the present embodiment is characterized in that the following functions are executed by a computer: an acquisition function for acquiring chemical substance information and pharmacological information of a medicine; a presumption function for performing machine learning by using the chemical substance information and the pharmacological information to presume presumption information of the medicine; and an output function for retraining a machine learning model based on the presumption information, thereby predicting and outputting the efficacy and side effects of the medicine on a living body. Further, the acquisition function in the prediction program further includes acquiring biological information of the medicine at the cellular level, and the presumption function performs machine learning by using the chemical substance information, the pharmacological information, and the biological information at the cellular level.

Effects of the Invention

[0019] The prediction device of the present embodiment includes: an acquisition unit for acquiring chemical substance information and pharmacological information of a medicine; a presumption unit for performing machine learning by using the chemical substance information and the pharmacological information to presume presumption information of the medicine; and an output unit for retraining a machine learning model based on the presumption information, thereby predicting and outputting the efficacy and side effects of the medicine on a living body. Therefore, by integrating the chemical information of the compound with its action information, it is possible to more accurately and efficiently achieve drug development of the required medicine, and further predict the efficacy and side effects of the medicine.

[0020] Similarly, in the prediction method and the prediction program, it is also possible to achieve more accurate and efficient drug development of the required medicine by integrating the chemical information of the compound and its action information, and further predict the efficacy and side effects of the medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram showing an outline of the prediction device and its system configuration of the present embodiment. Figure 2 It is a block diagram showing a configuration of computer functional units. Figure 3 It is a first schematic diagram showing machine learning processing. Figure 4 It is a second schematic diagram showing machine learning processing. Figure 5 It is a third schematic diagram showing machine learning processing. Figure 6 It is a fourth schematic diagram showing machine learning processing. Figure 7 It is a fifth schematic diagram showing machine learning processing. Figure 8 It is a flowchart showing the prediction method of the present embodiment. Figure 9 It is a flowchart for explaining the prediction method of other embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0022] As one of the problems existing in the in - silico drug development method using computers, usually, in the drug design process, the dynamic changes of drugs in the body, that is, the perspective of pharmacokinetics, and the final action results of the drug in the body (i.e., pharmacodynamics and side effects) are not included as the final output goals in the drug development process. In recent years, as a method for predicting drug adverse reactions, an artificial - intelligence - based pharmacokinetic prediction system called the ADMET prediction system (Absorption, Distribution, Metabolism, Excretion, Toxicity) has begun to be developed.

[0023] However, the elements that can be predicted by existing prediction methods are still very limited. For example, it can only predict the affinity of a compound for a specific drug transporter or the affinity for an enzyme that may cause cardiotoxicity when inhibited. This is mainly because constructing an AI system for ADMET prediction relies on a large amount of experimental data, and the data volume of some projects is still insufficient. In addition, the ADMET prediction system and the prediction system related to pharmacodynamics are currently independent of each other, and there is no AI system that can integrate the two and effectively and directly predict the final pharmacodynamic effect of drugs in the body (Pharmacodynamics).

[0024] The main reason why it is difficult to construct an artificial - intelligence model that can accurately predict the pharmacodynamic effect (Pharmacodynamics) in drug development is that it is necessary to integrate and quantitatively master multiple complex factors. For example, how the drug interacts with the drug transport system and the metabolic system in the body, and finally to what extent it reaches and binds to target cells or non - target cells and proteins, and how these processes specifically trigger changes at the cellular level. In addition, the above - mentioned information needs to be combined with clinical - medicine - related data. Moreover, the number of clinical data directly recording pharmacodynamics and side effects is very limited. Therefore, it is still a great challenge to train machine - learning models only with existing clinical data at the present stage.

[0025] Therefore, the prediction device, prediction method, and prediction program provided by this embodiment start from the molecular structure of a drug (compound) and the extraction of its chemical and physical properties, gradually training multi-level information including the mechanism of action at the cellular level, and finally achieving the prediction of the in vivo manifestations of drug efficacy and side effects (i.e., Pharmacodynamics). Such a prediction is precisely the most crucial core content in the drug R & D process. With the aid of this prediction device, prediction method, and prediction program, the finally constructed trained model can quantitatively predict the efficacy and side effects of a drug based on relatively limited information such as the structure of unknown candidate drugs, and accordingly screen out potential candidate drugs. In addition, by reading the output results of the intermediate layer of the machine learning model, relevant insights into the pharmacokinetics of candidate drugs (compounds) can also be obtained simultaneously, which further helps to predict the most suitable dosing regimen for the drug. The candidate drugs (compounds) described herein include not only low-molecular or medium-molecular organic compounds with a molecular weight of approximately 100 to 5000, but also various types such as peptides, nucleic acids, metal-containing compounds, etc. Furthermore, this technology is also applicable to the development of high-molecular medical products such as monoclonal antibodies.

[0026] When explaining the prediction device, prediction method, and prediction program of this embodiment, the structure of the prediction device will be described in conjunction with the processing steps of the prediction method. Figure 1 The schematic diagram shows the composition of the prediction device 1 of this embodiment. This prediction device 1 is connected to an external network through the Internet (not shown). Through the Internet connection, this prediction device 1 can obtain publicly available information from public websites such as individuals, research institutions, national or private data centers, and medical institutions. The information obtained includes, but is not limited to: all kinds of information related to drug R & D, such as the physicochemical properties of compounds, biological experimental data, clinical data, and papers related to basic medicine and clinical practice. In addition, information or data input by the user of the prediction device 1 can also be used as the basis for model operation and prediction.

[0027] In Figure 1Among them, the chemical substance information J1 of the medicament, the biological information J2 of the medicament at the cellular level of the cells, and the pharmacological information J3 of the medicament for the living body are input into the prediction device 1. Then, the prediction device 1 (prediction method, prediction program) can output the efficacy and side effects K1 of the medicament, as well as effective information for clinical medicine, such as the dosing regimen K2, etc., by using machine learning trained based on these information. The dosing regimen K2 will be described in detail in the subsequent section. The prediction device 1 includes a computer 10, and is equipped with arithmetic elements 11 (such as GPUs, etc.), a ROM 12, a RAM 13, a storage unit 14, etc. on the hardware. The computer 10 of the prediction device 1 is composed of various electronic computers (computing resources) such as supercomputers, hosts, workstations, cloud computing systems, etc.

[0028] If implemented by software Figure 1 each functional part of the computer 10 of the prediction device 1 shown, then the computer 10 realizes its functions by executing instructions of software programs that implement each function. The recording medium storing this program can be a "non-temporary tangible computer-readable storage medium", for example, CD, DVD, semiconductor memory, programmable logic circuit, etc. In addition, this program can also be supplied to the computer 10 of the prediction device 1 through any transmission medium (communication network, broadcast wave, etc.) that can transmit this program.

[0029] The storage unit of the computer 10 of the prediction device 1 is equipped with storage devices such as HDDs or SSDs. The storage unit 14 can also be an external server (not shown). The storage unit 14 is used to store various data, information, prediction programs, and various data required for executing this program, etc. In addition, the components that execute various calculations, operations, etc. are the arithmetic elements 11. Also, input devices such as keyboards and mice (not shown), a display unit (such as a display device such as a monitor), output devices for outputting data, etc. can also be appropriately connected to the prediction device 1 (computer 10).

[0030] Each functional part in the arithmetic element 11 of the computer 10 of the prediction device 1, such as Figure 2 shown in the block diagram. Each functional part includes a scraping part 100, an acquisition part 110, a preliminary estimation part 120, an estimation part 130, an output part 140, etc. The processing and execution in the prediction device 1 are realized by relying on software such as prediction programs loaded into the main memory.

[0031] The acquisition part 110 is used to acquire the chemical substance information J1 of the medicament and the pharmacological information J3 of the medicament. In addition, the acquisition part 110 also acquires the biological information J2 of the medicament at the cellular level (see Figure 1 ). The acquisition of each piece of information J1, J2, J3 can be realized by the scraping part 100 described later, or the user of the prediction device 1 manually inputs the required information.

[0032] The chemical substance information J1 of the pharmaceutical agent is a characteristic quantity generated based on the chemical structure at least included in the pharmaceutical agent. The information included in the chemical substance information J1 of the pharmaceutical agent includes chemical structure information, chemical properties, and physical property information (characteristic information). The chemical properties and physical property information (characteristic information) of the pharmaceutical agent include, for example: the molecular weight, charge, LogP, pKa, hydrogen bond binding property, molecular orbital, solubility, membrane penetration rate, etc. of the pharmaceutical agent (compound). In addition, it also includes characteristic quantities that can be calculated based on the two-dimensional structure and three-dimensional structure of the compound associated with at least the chemical structure of the pharmaceutical agent. For example, it also includes characteristic quantities such as the number of atoms and the list of structural fragments based on the two-dimensional structure, and the molecular volume, molecular surface area, bond length, rotational freedom, etc. based on the three-dimensional structure. Therefore, the chemical substance information J1 of the pharmaceutical agent is generated as a characteristic quantity based on its chemical structure information and characteristic information. It should be noted that the chemical structure information is a characteristic quantity generated based on the chemical structure of the pharmaceutical agent. The acquisition unit 110 is used to acquire the above various properties.

[0033] The biological information J2 of the pharmaceutical agent at the cellular level mainly refers to the index obtained by quantifying the behavior and influence of the compound on cells, and is generated as a characteristic quantity of the information obtained when the pharmaceutical agent is administered to cultured cells. That is, this information can be obtained through in vitro experiments, and its scope is wider than the information obtained solely by biochemical methods. For example, (in the case of intracellular targets) it includes molecular biology-related information such as the cellular bioavailability of the pharmaceutical agent compound, DNA changes, RNA expression changes, and increases or decreases in transcription factors caused by epigenetic changes after the administration of the pharmaceutical agent, as well as omics changes in protein expression, intracellular metabolites, etc. In addition, it also includes macroscopic cell biological phenomena such as cell volume after culture, cytokine production, and apoptosis.

[0034] The pharmacological information J3 refers to, in addition to the information on the action of the pharmaceutical agent at the in vivo level, the relevant information (pharmaceutical agent administration information) that may affect its action when the drug is administered to the living body. The information on the action of the pharmaceutical agent at the in vivo level is obtained through so-called clinical data (in vivo data). Specifically, it includes pharmacokinetic information (Pharmacokinetics) and pharmacodynamic effects and side effects (Pharmacodynamics) in the body.

[0035] Pharmaceutical administration information refers to the parameters that affect the pharmacokinetics, efficacy, and side effects of a drug when it is administered into the body. The pharmacokinetics, efficacy, and side effects of a drug in the body can vary depending on the administration method of the drug (including dosage, route of administration, etc.). In addition, although the specific values of these parameters are not directly recorded in clinical data, since there are differences between patients suffering from actual diseases and healthy individuals in their in-vivo environments (such as extracellular tissue information as the target), it also affects pharmacokinetic information, efficacy, and side effects. Therefore, any information that affects pharmacokinetic information, efficacy, and side effects is included in the pharmacological information J3 as pharmaceutical administration information. Specifically, it includes the administration method of the drug (dosage, route of administration, etc.) and the biological information of the body when the drug is administered (such as age, gender, extracellular tissue information, etc.).

[0036] When the acquisition unit 110 acquires information, the acquisition unit 110 is equipped with a crawling unit 100, and with the help of this crawling unit 100, automated acquisition of information can be achieved. The crawling unit 100 can acquire at least any one, two, or all of the chemical substance information J1 of the drug, the biological information J2 of the drug at the cellular level, or the pharmacological information J3 from websites existing on the Internet. The websites are information sources publicly disclosed by various facilities such as research institutions, national and public and private data centers, and medical institutions. In addition, the acquisition unit 110 can also use information or data obtained by the user himself / herself for information acquisition.

[0037] Due to the increasing amount of information such as chemical and physical knowledge of drugs, related papers, and knowledge, papers, and dosing standards at the cellular level and clinical level. Therefore, in order to continuously obtain the latest information and data and update the machine learning model, the computer 10 of the prediction device 1 acquires from the publicly available information on the website through the crawling unit 100 and stores it in the storage unit 14. In addition, the user of the prediction device 1 can also store information and data in the storage unit 14.

[0038] The preliminary estimation unit 120 makes a preliminary presumptive estimation of the chemical structure of the drug based on the characteristic information of the drug (see Figure 3 ). The characteristic information of the drug can be obtained by the crawling unit 100 or by the user of the prediction device 1 inputting the required information.

[0039] The pre-estimation unit 120 is responsible for the role of pre-estimating the chemical structure of the drug based on characteristic information in the subsequent estimation unit 130. Specifically, for example, it includes chemical structure information such as the surface area and molecular weight of the drug compound, charge information such as electron density and polarizability, acidity, water hydrogen bond ability, hydrophilicity, hydrophobicity, molecular orbital energy level, membrane permeability rate, solubility, etc. In addition, it also includes the chemical structure change rate of the drug and the change rate of characteristic information under specific environments, such as the change rate of molecular structure and charge distribution when interacting with polar groups such as amino acids and phosphates and non-polar groups such as alkyl groups and phenyl groups, etc., which can all be used as the above-mentioned characteristic information. In addition, it also includes the characteristic quantities of the two-dimensional and three-dimensional structures of the compound calculated based on the chemical structure at least related to the drug. For example, characteristic quantities such as the number of atoms and the list of structural fragments based on the two-dimensional structure, and chemical structure information such as molecular volume, molecular surface area, bond length, and rotational freedom based on the three-dimensional structure. Such information is also included in the chemical substance information as characteristic information.

[0040] In Figure 3 In the schematic diagram of the machine learning process shown, when selecting the molecular structure of the drug, the characteristic quantities based on the chemical properties and physical properties (characteristic information of the drug) of the drug are embedded in the input layer in the figure. It should be noted that in the machine learning process in the pre-estimation unit 120, the selection process of whether to embed the molecular structure of the drug can be determined as an optional item according to actual needs.

[0041] Based on the combination of input values, the chemical structure of the drug is output through machine learning. For example, a representation method that can be expressed using the SMILES notation (Simplified Molecular Input Line Entry System) can be used as an output form. In the machine learning in the pre-estimation unit 120, the estimation unit 130, and the output unit 140, neural networks, etc. can also be used.

[0042] When performing machine learning in the computing element, machine learning is also carried out through methods such as neural networks. Other machine learning methods such as association rule learning, random forests, and support vector machines can also be adopted. For the information obtained by the crawling unit 100, data processing can also be carried out through programming languages or processed using natural language processing.

[0043] In addition, after training using the above input items, if the prediction accuracy does not reach the target value, a model such as an autoencoder will be retrained based on the chemical structure to extract additional potential feature amounts for compensating the insufficient part of the prediction. The automatically extracted features are added to the input items as input values, and the machine learning model is retrained to improve the prediction accuracy. In this way, the network training for preprocessing in machine learning can be executed.

[0044] The estimation unit 130 estimates the estimation information of the medicine by performing machine learning using the chemical substance information J1 and the pharmacological information J3. In addition, the estimation unit 130 also estimates the estimation information of the medicine by performing machine learning by combining the chemical substance information J1, the pharmacological information J3, and the biological information J2 at the cell level. To improve the accuracy of machine learning in the output unit 140, as shown in this embodiment, in addition to the characteristic information, pharmacokinetic information and biological information at the cell level can be further used as the estimation information. In addition, the estimation unit 130 estimates the characteristic information of the medicine based on the chemical structure included in the chemical substance information J1 and the medicine administration information included in the pharmacological information J3; it can also estimate the pharmacokinetic information of the medicine based on the chemical structure included in the above chemical substance information J1 and the medicine administration information included in the pharmacological information J3; it can also estimate the biological information J2 based on the chemical structure included in the chemical substance information J1 and the medicine administration information included in the pharmacological information J3. In addition, when predicting the estimation information, the estimation unit 130 also uses other estimation information to further train the machine learning model to improve the prediction accuracy.

[0045] Figure 4 The schematic diagram of the machine learning process in shows the processing flow of outputting the characteristic information of the medicine from the limited chemical substance information of the medicine. Specifically, the chemical structure of the input drug is input, and each characteristic amount in the characteristic information of the drug is output. In this process, as shown in the schematic diagram of Figure 4 , the characteristic amounts of the medicine administration information such as the medicine administration amount, the administration route, and the target tissue information can be used as input values for processing.

[0046] In Figure 4 the machine learning processing stage, even if the characteristic amounts of the medicine administration information in the input layer change appropriately, the model is trained to output the same result state. The reason for this is that the purpose of this stage is to extract characteristic amounts from the pure chemical and physical properties of the medicine and enable the machine learning model to learn an input-output relationship that does not depend on input variables such as the medicine administration amount, the administration route, and the target tissue information. In addition, Figure 4 the network in the schematic diagram is different from the preliminary network in Figure 3 , and its weights are randomly initialized before training.

[0047] As Figure 4 shown in the schematic diagram of the machine learning model, in the traditional machine learning model where the characteristic information of a pharmaceutical agent is associated with its chemical structure (chemical substance information), a new perspective has been added that takes into account aspects of in-vivo administration of the pharmaceutical agent, such as the dosage of the pharmaceutical agent, the route of administration, and target tissue information. Through this improvement, subsequent training of the machine learning model is carried out not only based on the chemical structure but also on the basis of considering the above-mentioned in-vivo administration of the pharmaceutical agent, thereby improving its prediction accuracy.

[0048] Next, in order to improve the prediction accuracy of the output unit, the chemical substance information and pharmacological information of the drug can be used as input information, and the dynamic changes of the drug in the living body can be inferred to train the machine learning model. During this process, when further training the Figure 4 machine learning model, the chemical substance information required for the input data is only the chemical structure. In addition, as Figure 5 shown in the schematic diagram, in addition to the terms related to the dynamic changes of the drug (characteristic quantities based on the behavioral information of the drug after in-vivo administration), the output layer also outputs the same terms as those output (in silico) in the Figure 4 chemical data processing layer.

[0049] The reason for this is that when transfer learning (as shown later in Figure 6 ) is carried out to train the machine learning model for outputting the dynamic changes of the drug, the output (characteristic information) of the first layer will be directly utilized to facilitate output prediction. In order to cope with the changes in the characteristic information of the pharmaceutical agent under certain specific environments (such as in an acidic environment), the output characteristic information is different from that in Figure 4 , and it is necessary to further train the changes in the in silico output terms according to the target tissue information, which is one of the input values. After such retraining, the learnable parameters of all or part of the units in the first layer can be frozen to ensure that these values are no longer updated in subsequent training.

[0050] In the training for inferring pharmacokinetics, as Figure 5 shown in the schematic diagram, based on the input values "chemical substance information (chemical structure) of the pharmaceutical agent", "dosage of the pharmaceutical agent", and "route of administration", values of characteristic quantities related to Pharmacokinetics, such as "Bioavailability (the degree to which the administered pharmaceutical agent reaches the blood)", "maximum plasma concentration", "area under the plasma concentration-time curve", "volume of distribution", "renal clearance", and "hepatic metabolic rate", are output as the characteristic quantities corresponding to the behavioral information after the pharmaceutical agent is administered to the living body. Here, for input values that are not directly related to pharmacokinetics (such as Figure 5the target tissue information shown), during training, the output value related to Pharmacokinetics can be kept constantly output regardless of any value taken by this type of input,

[0051] Regarding Bioavailability, through machine learning, the model can learn the changes in Bioavailability of the drug candidate compound under different administration routes. Among them, the "administration route" is encoded as a numerical value as an input item. For example: 0 represents intravenous injection, 1 represents oral administration, 2 represents rectal administration, etc. When the input is intravenous injection, its Bioavailability output is 100%; while for other routes, such as oral administration, the output is 1%, and rectal administration will output 50%.

[0052] Next, in order to further improve the prediction accuracy of the output unit 140, machine learning model training can also be performed on the biological information of the drug at multiple cell levels in the cell (i.e., the so-called in vitro experimental data, in vitro data). During this process, the biological information used can be roughly divided into three stages: "cellular bioavailability", "intermediate data containing molecular biology data", and "biological data at the macroscopic level". These information constitute the characteristic quantity items corresponding to the behavior information shown when the drug is applied to the cultured cells. The order of this training process can be set according to biological logic. For example, the training of RNA expression is prior to the training of the protein expression stage. It should be noted that adjusting the training order is allowed, and in the implementation of machine learning, it is not required to include all the above stages.

[0053] It should be noted that the main purpose of the training based on the biological information of multiple cell levels (i.e., in vitro data) is to reduce the amount of data required during the subsequent training of the output layer (see Figure 7 the final pharmacodynamic and side effect prediction layer (in vivo) shown). Therefore, in order to enable the machine learning model to predict the final pharmacodynamic and side effects with sufficient accuracy, its training content can be appropriately adjusted as needed.

[0054] Here, the chemical substance information and pharmacological information of the drug can also be used as inputs, and then the biological information of the drug at the cell level, that is, the behavior information when the drug is administered to the cultured cells, can be deduced. Figure 6 The schematic diagram of the machine learning process shown represents the processing flow of re-training the machine learning model used to deduce the characteristic information and pharmacokinetic information described above Figure 5 to output the biological information at the cell level.

[0055] As is well known, in the characteristic quantities (i.e., in vitro data) corresponding to the biological information at the cellular level representing the effect of a drug on cells, the amount of drug added during the culture process is not essentially the same as the dosage item in the usual sense. Therefore, the following method can be used to pre-estimate the dosage corresponding to various types of data. For example, after the model training shown in the schematic diagram of Figure 5 is completed, the chemical structure, administration route, and target tissue information used to obtain in vitro data are set as fixed values. Then, by changing the value of dose (dosage), the change in pharmacokinetic information is observed. Subsequently, a dose is explored such that the drug concentration in the culture medium used in the in vitro system is consistent with the extracellular fluid concentration in the tissue estimated by the pharmacokinetic prediction model. During this process, the composition information of the culture medium is input into the fixed input target tissue information. Since the value of pharmacokinetic information varies depending on the administration route, for each type of route in the input values, such as 0: intravenous injection, 1: oral administration, 2: rectal administration, etc., the corresponding dosage can be estimated separately and independently.

[0056] As Figure 6 shown in the third a layer of the schematic diagram, the cellular bioavailability at the cellular level refers to the degree to which a drug compound enters the interior of different types of cells, and its value can be expressed by the ratio of the drug concentration in the cell to the drug concentration in the culture medium, etc. During a specific culture period, for multiple types of cells, measurements are carried out under different drug dosages and culture conditions (i.e., target tissue information). During this measurement process, in addition to the cellular bioavailability at the cellular level, the first-layer output "characteristic information" and the second-layer output "pharmacokinetic information" under the conditions of the target tissue information are also output.

[0057] In Figure 6 the third b layer of the schematic diagram shown, data related to comprehensive (omics) changes such as molecular biology data and protein expression changes are obtained. For example, it includes information such as the expression of RNA in cells and the expression of transcription factors. The above expression information of RNA, etc. is measured for multiple cells under different drug dosages and culture conditions (i.e., target tissue information) during a specific culture period, and the expression information of RNA, etc. in the cells is obtained and used as training data to be input into the model for learning. The above training is executed after the model training of the cellular bioavailability at the cellular level in the third a layer is completed.

[0058] In Figure 6In the third c layer of the schematic diagram shown, biological feature data observed from a macroscopic perspective is used as multiple variable feature quantities (i.e., in vitro data). For example, data that can be used for training as the output value of the feature quantity includes: the cell size after cell culture, the viability after cell culture, the protein expression levels inside and outside the cells obtained by measurement methods such as flow cytometry, the amount of cytokines produced by various types of cells, etc.

[0059] The output unit 140 retrains the machine learning model (i.e., the machine learning model) by inputting the estimation information obtained by the estimation unit 130 and performing machine learning, so as to predict and output the efficacy and side effects of the drug on the living body. Of course, according to the setting, only one of the predicted results of the efficacy or side effect can be output.

[0060] The processing of the output unit 140 is based on all the output values of the previous layers, outputs the efficacy and side effects K1 of the drug as the main target, and further outputs the drug administration plan K2 as the accompanying information related to the efficacy and side effects (see Figure 1 ). In Figure 7 the schematic diagram shown, the machine learning model trained in the estimation unit 130 is retrained by using clinical data, so as to output the efficacy and side effects of the drug. The clinical data can be obtained from all or part of the data publicly available on the website, and the aforementioned scraping unit 100 can be used for acquisition.

[0061] The input values are the same as those in the previous machine learning process. In addition to the chemical structure of the drug, they also include the administration information of the drug. In Figure 7 the schematic diagram, "drug dosage", "administration route", and "target tissue information" are used. The drug dosage can be the dosage of each administration of the drug compound or the total dosage of the entire administration period. For the "target tissue information", the input is the extracellular environment information of the tissue cells that can be inferred from the disease complaint. Since this data cannot be obtained from public clinical trials, it can be conveniently set as a physiological value. However, according to the disease targeted, the extracellular fluid of the target tissue may change. For example, when the tissue has severe inflammation, the change in the external environment also changes with the degree of inflammation. Therefore, the machine learning model should consider a certain range to adapt to this change. In addition, factors such as age and past medical history can also be considered, or a model that does not consider these factors as shown in the figure can be constructed. Although this may cause problems, as described in the administration period, not considering these factors can ensure more data volume for the training of the machine learning model.

[0062] The ratio between the "number of participants in the clinical trial" and the "number of people in whom the drug efficacy (or side effects) is confirmed" can also be used as one of the outputs. Here, the "drug efficacy" directly refers to the remission rate of the disease. However, simply limiting it to this value cannot comprehensively grasp the overall effect of the pharmaceutical compound on the body. Therefore, the output values also include indirect data obtained from routine blood tests and the like conducted in clinical trials, and these data can serve as indirect indicators for indicating disease remission. For example, taking rheumatoid arthritis as an example, in addition to the data of "remission rate of rheumatoid arthritis", the test results of inflammatory markers such as C-reactive protein (CRP) that are usually detected are also included as an item of "CRP reduction rate". In addition, the output forms of side effects also include, for example: "incidence of headache after administration: 80%", "rash: 50%", "probability of abnormal values of liver and gallbladder markers: 20%", etc. These numerical indicators are used to indicate the possible side effects per unit number of people.

[0063] In this way, even if a drug used in "rheumatoid arthritis" is not used in "systemic lupus erythematosus", through the commonality of the item of CRP reduction, it can still be interpreted that both are drugs for "inhibiting inflammation". For side effects, as shown in the aforementioned examples, subjective items such as "headache" can be used, or objective items such as blood markers like kidney markers can also be used. It should be noted that for a hypertensive drug like "blood pressure reduction", this belongs to the drug efficacy, while for anaphylactic shock (i.e., allergic reaction) it is a side effect. Therefore, the item of "blood pressure reduction" is not unified, but is separately set up as two items: "blood pressure reduction" (drug efficacy) and "blood pressure reduction" (side effect).

[0064] Due to the large variety of drug efficacy / side effect items, there is a risk that some drug data items in the clinical database may have missing values. Therefore, in order to fill in the missing values, the drug data in the clinical database will be input into the currently trained machine learning model (the machine learning model in the estimation unit), the correlation coefficient between drugs will be calculated based on all or part of the output values, and a similarity index between drugs will be generated based on this. In this way, when a missing item appears in a certain drug (compound), the missing value can be supplemented with the data of other drugs (compounds) that have a high correlation with this drug (compound) in the clinical database.

[0065] Based on the above description, the prediction device 1 (and its prediction method, prediction program), starting from the chemical structure of the candidate drug, predicts the efficacy or side effects corresponding to the chemical structure of the candidate drug through machine learning simulation, or predicts both the efficacy and side effects simultaneously. That is, the prediction device 1 (and its prediction method, prediction program) learns chemical information (chemical substance information of the drug), biological information (cellular-level biological information of the drug on cells), and clinical information (pharmacological information of the drug on the living body), enabling the prediction of its chemical properties, cellular-level effects, and efficacy and side effects in the living body solely based on the chemical structure information of the drug. Further, since pharmacokinetic values are also calculated simultaneously during this process, the dosing regimen of the candidate drug can also be predicted. The dosing regimen includes the route of administration, dosing frequency, and dosage. In addition, for diseases that currently have no cure, the efficacy can still be speculated to a certain extent through the side effects of the candidate drug, as well as information such as cellular-level biological information and blood biomarkers. Therefore, without actually synthesizing the candidate drug, by predicting the chemical structure of the candidate drug that may have efficacy, and further predicting the efficacy, side effects, and dosing regimen of the candidate drug, the screening work in new drug development (drug creation) can be simplified, and important references can be provided for the selection of the synthesis direction of the candidate drug.

[0066] An example of the drug creation process for predicting the efficacy and side effects of unknown drugs using the prediction device 1 and the trained machine learning model is as follows. First, randomly select compounds to create a list of candidate drugs covering a variety of compounds. During this process, through existing methods for evaluating synthesis difficulty, etc., compounds that are difficult to synthesize are excluded from List - 1. Subsequently, input all types of compounds into the training network of the machine learning model to obtain the remission rate (efficacy) and side effects of each compound for treating the desired disease. Compare the obtained efficacy and side effects with the benchmark values. Drugs with efficacy lower than the benchmark value or side effects exceeding the benchmark value are excluded, and the remaining drugs are used as candidates to generate List - 2. Next, generate a large number of drugs with partial changes to the chemical structure of each drug in List - 2, exclude drugs that are difficult to synthesize, predict the efficacy and side effects of the remaining drugs, exclude drugs that do not meet the benchmark values, and generate List - 3 based on the remaining drugs. Perform partial structural changes on the newly added compounds in List - 3 and conduct the same evaluation. The operation of changing the partial structure of the compound and evaluating it will be repeated continuously to optimize the efficacy and side effects until there is no significant improvement in the efficacy and side effects of any drug in the list. At this time, the repetitive process ends, and finally, the drugs in the list are sorted, and the top drugs are selected as the final candidate drugs.

[0067] In the administration route of the medicine, information such as 0 (intravenous) should be input preferentially. However, regarding the dose, its value has not been determined. In this regard, it is necessary to adjust the dose in a timely manner and verify how the efficacy and side effects change with the dose, so as to determine the dose range in reality. Only when the efficacy exceeds the reference value and the side effects are lower than the reference value will the medicine be retained as a medicine candidate in the list. In addition, for the administration route, it can be adjusted according to actual needs. For example, in addition to intravenous administration, if oral medicine is needed, etc., the conditions of the administration route can be modified according to specific goals, and the efficacy and side effects can be evaluated.

[0068] When predicting the efficacy and side effects of an unknown medicine by using the machine learning model trained by the prediction device 1 of the present embodiment, multiple candidate medicines can be sorted according to the severity of the efficacy and side effects. In this way, among multiple candidate medicines, it is easy to judge which medicines should be selected for each phase of the test. Since the efficacy and side effects of candidate medicines can be predicted in advance, a priority order can be assigned to the candidate medicines, thereby improving the efficiency of planning and management of pharmaceutical trials and the like.

[0069] Next, in combination with Figure 8 and Figure 9 's flowchart, the prediction method and prediction program of the prediction device 1 in the computer 10 will be explained. The prediction method is executed by the arithmetic element 11 of the computer based on the prediction program. The prediction program enables Figure 1 , Figure 2 's computer 10 to execute functions such as a crawling function, an acquisition function, a preliminary estimation function, an estimation function, an output function, etc. It should be noted that since each function repeats the description of the aforementioned prediction device 1, it will not be described in detail here.

[0070] According to Figure 8 's flowchart, the processing of the arithmetic element 11 of the computer 10 includes various steps such as a crawling step (S100), an acquisition step (S110), a presumption step (S130), and an output step (S140). Of course, the various steps required for the operation of the computer 10 itself are also naturally included. In Figure 8 's shown flowchart, the crawling step (S100) is not essential. As mentioned above, when the user of the prediction device 1 manually inputs information or data, the crawling step (S100) can be omitted.

[0071] The crawling function can obtain at least any one of the chemical substance information of the drug, the biological information at the cellular level of the drug on cells, or the pharmacological information of the drug from websites on the Internet (S100: crawling step). The obtaining function obtains the chemical substance information and the pharmacological information of the drug (S110: obtaining step). In addition, the obtaining function can also obtain the chemical substance information of the drug, the pharmacological information of the drug, and the biological information of the drug at the cellular level. The estimating function estimates the speculative information of the drug by performing machine learning using the chemical substance information and the pharmacological information (S130: estimating step). In addition, the estimating function can also estimate the speculative information of the drug by performing machine learning using the chemical substance information, the pharmacological information, and the biological information at the cellular level. The output function predicts and outputs the efficacy and side effects of the drug on the living body based on the speculative information through retraining of the machine learning model (S140: output step).

[0072] Regarding another implementation of the prediction method and prediction program of the computer 10 in the prediction device 1, in Figure 9 the flowchart shown, the processing of the arithmetic element 11 of the computer 10 includes multiple steps such as a crawling step (S100), an obtaining step (S110), a preliminary estimation step (S120), an estimation step (S130), and an output step (S140). Of course, each step required for the operation of the computer 10 itself is naturally included. In this implementation, a preliminary estimation function is provided. The preliminary estimation function pre-estimates the chemical structure of the drug based on the characteristic information (S120: preliminary estimation step). Even in Figure 9 the flowchart shown, the processing of the crawling step (S100) is not necessary. As described above, when the user of the prediction device 1 manually inputs information or data, the crawling step (S100) can be omitted.

[0073] The computer program of the present invention described above can be recorded on a record medium readable by a processor, and a "non-temporary tangible computer-readable storage medium" can also be used as the record medium, such as an optical disc, a card, a semiconductor memory, a programmable logic circuit, etc.

[0074] In addition, the above computer program can be implemented using script languages such as ActionScript and JavaScript (registered trademark), object-oriented programming languages such as Objective-C and Java (registered trademark), and markup languages such as HTML5. Explanation of reference numerals 1 Prediction device 10 Computer 11 Arithmetic element 12 ROM 13 RAM 14 Storage unit 100 Crawling section 110 Acquisition section 120 Preliminary estimation section 130 Estimation section 140 Output section J1 Chemical substance information of the drug J2 Biological information at the cellular level of the drug on cells J3 Pharmacological information of the drug on the organism K1 Efficacy and side effects of the drug K2 Administration regimen

Claims

1. A prediction device, characterized in that, Comprising: An acquisition unit that acquires chemical substance information of a medicament and pharmacological information of the medicament; A presumption unit that performs machine learning by using the chemical substance information and the pharmacological information, thereby presuming presumption information of the medicament; An output unit that retrains the machine learning model based on the presumption information, thereby predicting and outputting the efficacy and side effects of the medicament on a living body.

2. The prediction device according to claim 1, characterized in that, The acquisition unit further acquires biological information of the medicament at the cellular level, The presumption unit performs machine learning by using the chemical substance information, the pharmacological information, and the biological information at the cellular level.

3. The prediction device according to claim 1, wherein The presumption unit presumes characteristic information of the medicament based on the chemical structure included in the chemical substance information and the medicament administration information included in the pharmacological information.

4. The prediction device according to claim 1, characterized in that, The presumption unit presumes pharmacokinetic information of the medicament based on the chemical structure included in the chemical substance information and the medicament administration information included in the pharmacological information.

5. The prediction device according to claim 2, wherein The presumption unit presumes the biological information at the cellular level based on the chemical structure included in the chemical substance information and the medicament administration information included in the pharmacological information.

6. The prediction device according to claim 3, wherein It further includes a pre-presumption unit for pre-presuming the chemical structure of the medicament based on the characteristic information of the medicament.

7. The prediction device according to claim 1, wherein, When predicting the presumption information, the presumption unit also retrains the machine learning model by using other presumption information.

8. The prediction device according to claim 1, characterized in that, The machine learning is a neural network and an autoencoder is adopted.

9. The prediction device according to claim 3, characterized in that The chemical substance information of the medicament includes chemical structure information and the characteristic information. The chemical structure information is a feature quantity based on the chemical structure of the medicament, and the characteristic information is a feature quantity based on the chemical properties and physical properties of the medicament.

10. The prediction device according to claim 2, characterized in that, The biological information at the cellular level is a feature quantity based on the behavior information when the medicament is administered to cultured cells.

11. The prediction device according to claim 3, characterized in that The pharmacological information includes feature quantities based on the effects of the medicament on a living body and its pharmacokinetic characteristics in vivo, and medicament administration information, where the medicament administration information is a feature quantity based on the medicament administration method and the living body information at the time of administration.

12. The prediction device according to claim 1, characterized in that, When predicting and outputting the efficacy and side effects of the medicament on a living body, the output unit also predicts and outputs the dosing regimen of the medicament.

13. The prediction device according to claim 2, characterized in that The acquisition unit includes a crawling unit for acquiring at least one of the chemical substance information of the medicament, the biological information of the medicament at the cellular level, or the pharmacological information of the medicament from a WEB website on the Internet.

14. A prediction method, characterized in that, A computer executes the following steps: An acquisition step of acquiring chemical substance information of a medicament and pharmacological information of the medicament; A presumption step of performing machine learning by using the chemical substance information and the pharmacological information, thereby presuming presumption information of the medicament; An output step of retrains the machine learning model based on the presumption information, thereby predicting and outputting the efficacy and side effects of the medicament on a living body.

15. The prediction method according to claim 14, characterized in that, The acquisition step further includes acquiring biological information of the medicament at the cellular level, The presumption step uses the chemical substance information, the pharmacological information, and the biological information at the cellular level to perform machine learning.

16. A prediction program, characterized in that, The following functions are executed by a computer: Acquisition function, which acquires the chemical substance information of the medicament and the pharmacological information of the medicament; Presumption function, which performs machine learning by using the chemical substance information and the pharmacological information to presume the presumed information of the medicament; Output function, which retrains the machine learning model based on the presumed information, thereby predicting and outputting the efficacy and side effects of the medicament on the living body.

17. The prediction program according to claim 16, wherein, The acquisition function further includes acquiring the biological information at the cellular level of the medicament, The presumption function performs machine learning by using the chemical substance information, the pharmacological information, and the biological information at the cellular level.

Citation Information

Patent Citations

  • Connectivity prediction method, apparatus, program, recording medium, and production method of machine learning algorithm

    JP2019028879A

  • Deep layer learning method and apparatus of chemical compound characteristic prediction using artificial chemical compound data, as well as chemical compound characteristic prediction method and apparatus

    JP2020009203A