Target molecule prediction method

An AI-optimized mathematical model predicts target molecules in biochemical pathways, addressing inefficiencies in existing screening methods by accurately identifying therapeutic agents for diseases through iterative parameter adjustments.

JP7726486B2Active Publication Date: 2025-08-20INSTITUTE OF SCIENCE TOKYO +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022530548
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-08
Filing Date
2021-06-07
Publication Date
2025-08-20
Estimated Expiration
2041-06-07

AI Technical Summary

Technical Problem

Current methods for screening therapeutic drugs targeting biochemical reaction pathways like the complement pathway are inefficient due to the complexity of the complement pathway and difficulty in experimental reproducibility, lacking a practical AI-based screening method.

Method used

A mathematical model optimized using artificial intelligence to predict target molecules in biochemical reaction pathways by adjusting parameters through machine learning, involving iterative searches and surrogate modeling to minimize differences between estimated and actual accumulation values.

Benefits of technology

Accurately predicts the targets of unknown compounds by optimizing parameters, enabling effective screening of therapeutic agents for diseases beyond complement-related diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007726486000006
    Figure 0007726486000006
  • Figure 0007726486000007
    Figure 0007726486000007
  • Figure 0007726486000008
    Figure 0007726486000008
Patent Text Reader

Abstract

The present invention provides a method for building a mathematical model including optimized parameters, said method including: (1) a step in which, for a substance or a cell having a known target molecule in a biochemical reaction pathway, individual parameters of mathematical models for each of target molecules are optimized by subjecting an artificial intelligence model including mathematical models for each of the target molecules to learning in which a data group indicating the estimated abundance of a reaction product at each of elapsed times from when the substance or cell is brought into contact with a target cell, a data group indicating the actual measured abundances, and a target molecule of the substance or cell are used as learning data; and (2) a step in which, for a substance or cell of which the target molecule is unknown, individual parameters of mathematical models for each of target molecules are optimized by subjecting an artificial intelligence model including mathematical models to learning in which weighting is performed for each of the target molecules and a data group indicating the estimated abundance of a reaction product at each of elapsed times and a data group indicating the actual measured abundances are used as learning data, and said mathematical models include each of the parameters optimized in step (1) and correspond to each of the combinations of all of the molecules that can be targeted by the substance or cell.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a mathematical model for predicting the amount of product accumulation in a biochemical reaction pathway, and a method for constructing an artificial intelligence model having the mathematical model. The present invention also relates to a method for predicting a target molecule in a biochemical reaction pathway of a test substance or a test cell using the constructed mathematical model or artificial intelligence model. The present invention also relates to an apparatus and a program for carrying out the method for predicting a target molecule.

[0002] [Background of the invention] Complement is not only an important effector of innate immunity, but also plays a role in adaptive immune responses, inflammation, blood coagulation, and tumor progression. Complement activation promotes foreign body clearance through opsonization, production of anaphylatoxins (C5a and C3a), and formation of the membrane attack complex (MAC), which is its final product. While complement activation functions as a biological defense, it can also be toxic in pathological conditions. For example, in the kidney, complement deposition in renal tissue can cause multimembraneous glomerulonephritis (MPGN), poststreptococcal acute glomerulonephritis (PSAGN), and lupus nephritis. Complement activation also leads to the formation of the membrane attack complex (MAC), which creates structural pores in the cell membrane. This results in osmotic fluid movement and cation influx, leading to cell death. Thus, because abnormalities in complement activation cause various diseases, compounds targeting the complement pathway are potential candidates for therapeutic or preventive drugs for complement-related diseases. Clinically available therapeutic drugs targeting the complement pathway include an antibody (eculizumab) against the complement component C5; however, the administration route of this antibody is intravenous (IV), which is invasive, and it is very expensive (current cost: $600,000 per person per year). To address these issues, screening for therapeutic drugs for diseases involving complement overactivation has focused on hemolytic activity. However, due to the complexity of the complement pathway and the difficulty of experimental reproducibility, screening for therapeutic drugs has been unsuccessful.

[0003] In recent years, with the advancement of the Fourth Industrial Revolution, the rapid increase in IoT devices, the practical application of deep learning technology, and improvements in hardware performance such as central processing units (CPUs) and graphics processing units (GPUs) have led to growing expectations for the use of artificial intelligence (AI) in pharmaceutical research (i.e., AI drug discovery) and materials development. In particular, big data drug discovery, genomics, drug repositioning, and materials informatics are attracting attention as future themes for pharmaceutical research and materials development. However, no practical method for using AI to screen therapeutic drugs targeting biochemical reaction pathways such as the complement pathway is known. Summary of the Invention [Problem to be solved by the invention]

[0004] Therefore, an object of the present invention is to provide a mathematical model capable of predicting target molecules, such as test substances, in biochemical reaction pathways such as the complement pathway, enabling screening of therapeutic agents that target the pathways, and to provide a method for predicting target molecules, such as test substances, using the model. Another object of the present invention is to provide an apparatus and a program for carrying out the method for predicting target molecules. [Means for solving the problem]

[0005] When constructing the mathematical model, the inventors envisioned that rather than building a mathematical model from scratch, they could construct a mathematical model with higher accuracy in predicting the target pathway by optimizing the parameters (coefficients) of a known mathematical model that predicts the accumulation of each reaction product in a biochemical reaction pathway through machine learning using artificial intelligence. The inventors conceived of optimizing the mathematical model by the following steps: first, initial parameter values were set in the existing mathematical model. Next, for a compound with a known target pathway, initial conditions such as compound concentration were substituted into the mathematical model corresponding to the target pathway to create an estimated graph with time on the horizontal axis and the amount of complement factor accumulation on the vertical axis. A function was set to represent the sum of the differences between the estimated accumulation calculated from the estimated graph and the actual measured value at each time point, and the parameters of the mathematical model were adjusted to minimize this function. Using the adjusted parameters, the same operation was performed on another compound with a known target pathway to adjust the parameters. This series of parameter searches based on adjustments and iterative searches were performed using artificial intelligence (e.g., surrogate modeling, Bayesian optimization, etc.) to increase efficiency (first step).

[0006] Furthermore, the inventors came up with the idea of further optimizing each adjusted parameter of the mathematical model by the following steps: inputting all possible combinations of pathways into the mathematical model whose parameters have been adjusted in the first step, and creating an estimated graph with time on the horizontal axis and complement factor accumulation on the vertical axis. For each estimated graph, the parameters of the mathematical model whose parameters have been adjusted in the first step are readjusted so that the difference from the actual measured values of each complement pathway factor for an unknown target compound is minimized (with certain limitations). From among all the readjusted mathematical models, the one with the smallest error from the actual measured values is deemed to be the correct mathematical model, and the parameters of this mathematical model are adopted. Using the mathematical model with the adopted parameters, the mathematical parameters are adjusted using AI repeatedly using the same procedure based on data from different compounds, thereby fine-tuning the parameters adjusted in the first step and obtaining an optimized mathematical model with the final parameters as constants rather than variables. The inventors continued their research based on these ideas and have completed the present invention.

[0007] That is, in order to solve the above problems, the present invention provides the following [1] to

[17] . [1-1] (1) A step of optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model having a mathematical model for each target molecule using, as learning data, a data group indicating the estimated abundance of a reaction product for each elapsed time after contacting a substance or cell with a target cell, the substance or cell having a known target molecule in a biochemical reaction pathway, the data group indicating the actual abundance, and the target molecule of the substance or cell; and (2) A method for constructing a mathematical model containing optimized parameters, comprising the step of: optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model having the mathematical model using a group of data indicating the estimated abundance of reaction products over time, weighted for each target molecule, and a group of data indicating the actual abundance, as learning data for a substance or cell whose target molecule is unknown; wherein the mathematical model corresponds to each combination of all molecules that may be targeted by the substance or cell, and includes each parameter optimized in step (1). [1-2] (1) A step of optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model having a mathematical model for each target molecule using, as learning data, a data group indicating the estimated abundance of a reaction product for each elapsed time after contacting a substance or cell with a target cell, the substance or cell having a known target molecule in a biochemical reaction pathway, the data group indicating the actual abundance, and the target molecule of the substance or cell; and (2) A method for constructing an artificial intelligence model having a mathematical model containing optimized parameters, the method including a step of: optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model having a mathematical model using, as training data, a group of data indicating the estimated abundance of reaction products for each elapsed time, weighted for each target molecule, and a group of data indicating the actual abundance, for a substance or cell whose target molecule is unknown; wherein the mathematical model corresponds to each combination of all molecules that may be targeted by the substance or cell, and includes each parameter optimized in step (1). [2] The step (1) (I) inputting a group of data indicating the estimated abundance of a reaction product for each elapsed time after contacting a substance or cell whose target molecule in a biochemical reaction pathway is known with a target cell, a group of data indicating the actually measured abundance, and the target molecule of each substance or cell into an artificial intelligence model having a mathematical model for each target molecule including parameters for which initial values are set, searching for optimal values for parameters of the mathematical model corresponding to the target molecule of the substance or cell using the artificial intelligence model, and updating the parameters to the optimal values; (II) adjusting each parameter of the mathematical model for each target molecule by sequentially repeating step (I) for each substance or cell using a data group indicating the estimated abundance of the reaction product for each elapsed time for a plurality of other substances or cells, a data group indicating the actually measured abundance, and the target molecule (provided that the artificial intelligence model is replaced with an artificial intelligence model having a mathematical model in which each parameter has been updated to an optimal value in step (I)); The method according to [1-1] or [1-2], comprising: [3] The step (2) (III) a step of inputting a group of data indicating the estimated abundance of a reaction product for each elapsed time weighted for each target molecule and a group of data indicating the actually measured abundance for a plurality of substances or cells whose target molecules are unknown into an artificial intelligence model having a mathematical model, searching for an optimal value for each parameter of the mathematical model for each target molecule using the artificial intelligence model, and updating the parameters to the optimal value, wherein the mathematical model corresponds to each combination of molecules that may be targeted by each substance or cell, and includes each parameter optimized in step (1); (IV) determining, from among the mathematical models for each target molecule in which each parameter has been updated to an optimal value in step (III), the mathematical model that can best explain the measured abundance is the mathematical model corresponding to the target molecule of each substance or cell; The method according to any one of [1-1] to [2], comprising: [4] The method according to [3], further comprising the step of (V) verifying the validity of the weighting in step (III). [5] (VI) updating the parameters optimized in step (1) to the optimized parameters of the mathematical model determined in step (IV); and (VII) a step of adjusting each parameter of the mathematical model for each target molecule by sequentially repeating steps (III)-(VI) for each substance or cell using a group of data indicating the estimated abundance of the reaction product for each elapsed time weighted for each target molecule and a group of data indicating the actually measured abundance, whereby the artificial intelligence model in step (III) is replaced with an artificial intelligence model having a mathematical model in which the parameters have been updated to optimal values in step (VI); The method according to [3] or [4], comprising: [6] The method according to any one of [2] to [5], wherein the data group indicating the estimated abundance in step (I) is obtained by substituting the concentration of the substance or cell in the culture medium and the abundance of the reaction product at the time of contact of the substance or cell as initial values into a mathematical model in which initial values are set for all parameters. [7] The method according to any one of [1] to [6], wherein the artificial intelligence model in step (1) is a model set with the objective function of minimizing the sum of the squares of the differences between the estimated abundance of the reaction product per elapsed time and the actually measured abundance. [8] The data group indicating the estimated abundance in step (III) is input to a mathematical model including each parameter for each target molecule optimized in step (1), and the concentration of each substance or cell in the culture medium. The method according to any one of [3] to [7], wherein the amount of the reaction product present at the time of contact of the substance or cell with the target cell is assigned as an initial value. [9] The method according to any one of [3] to [8], wherein the search in step (III) is carried out within a range set as a search range, which is a parameter range estimated after contact of a substance or cell with a target cell from the parameter set for each target molecule optimized in step (1).

[10] The method according to any one of [3] to [9], wherein step (IV) is a step of determining that the mathematical model for which the sum of the squares of the differences between the estimated abundances of the reaction products for each weighted elapsed time and the actual abundances is the smallest is the mathematical model corresponding to the target molecule of the substance or cell.

[11] The method according to any one of [2] to

[10] , wherein the search in the step (I) is carried out by surrogate modeling.

[12] The method according to any one of [3] to

[11] , wherein the weighting of the estimated abundance in the step (III) is set based on structural data of the substance.

[13] The method according to any one of [1-1] to

[12] , wherein the biochemical reaction pathway is the complement pathway.

[14] The method according to any one of [1] to

[13] , wherein the substance is selected from the group consisting of a low molecular weight compound, a nucleic acid, a peptide, and an antibody.

[15] (A) inputting a group of data indicating the concentration of each test substance or test cell in the culture medium, the amount of reaction product at the time of contact of the substance or cell with the target cell, and the measured amount of reaction product per elapsed time into an artificial intelligence model having a mathematical model constructed by a method described in any one of [1-1] to

[14] , and calculating the estimated amount of reaction product per elapsed time for each target molecule using the artificial intelligence model; (B) calculating the probability that the test substance or test cell targets each target molecule by comparing the estimated abundance of each target molecule calculated in step (A) with the actual abundance of each target molecule.

[16] A computer program for predicting a target molecule in a biochemical reaction pathway of a test substance or a test cell, comprising: Computer, (P1) a receiving unit that receives at least a group of data indicating the concentration of each test substance or test cell in the culture medium, the amount of reaction product present at the time of contact of the test substance or test cell with the target cell, and the measured amount present at each elapsed time; (P2) a first calculation unit that inputs the input data group into an artificial intelligence model having a mathematical model constructed by the method described in any one of [1] to

[14] , and calculates the estimated abundance of the reaction product of each target molecule per elapsed time; (P3) a second calculation unit that calculates the probability that each target molecule is targeted by the test substance or test cell by comparing the estimated abundance of each target molecule obtained by the first calculation unit with the actual abundance; and (P4) The program for causing the second calculation unit to function as an output unit that outputs the calculation result obtained by the second calculation unit.

[17] 1. A device for predicting a target molecule in a biochemical reaction pathway of a test substance or a test cell, comprising: (M1) a receiving unit that receives at least a group of data indicating the concentration of each test substance or test cell in the culture medium, the amount of reaction product at the time of contact of the substance or cell with the target cell, and the measured amount at each elapsed time; (M2) a first calculation unit that inputs the input data group into an artificial intelligence model having a mathematical model constructed by the method described in any one of [1] to

[14] and calculates the estimated abundance of the reaction product of each target molecule per elapsed time; (M3) a second calculation unit that calculates the probability that the test substance or cell targets each target molecule by comparing the estimated abundance of each target molecule obtained by the first calculation unit with the actual abundance; (M4) an output unit that outputs the calculation result obtained by the second calculation unit; An apparatus having: [Effects of the Invention]

[0008] According to the present invention, by accumulating data on known library compounds and optimizing parameters, it is possible to accurately predict the target of an unknown compound. If the cellular reaction and its time series can be confirmed, it will also be possible to predict the target of a compound for diseases other than complement-related diseases. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 shows an overview of a method for constructing a mathematical model in which each parameter is optimized according to the present invention. [Figure 2] Figure 2 is a block diagram showing a preferred example of the configuration of the device of the present invention. The communication relationship between the device of the present invention and external devices (the sender of the input data group and the destination of the prediction result) is indicated by dashed lines. In the example of Figure 2, the sender of the input data group is also the destination of the prediction result. [Figure 3] FIG. 3 is a diagram showing the flow of the program of the present invention.

[0010] [Detailed Description of the Invention] -1. How to build a mathematical model or an artificial intelligence model- The present invention provides a method for constructing a mathematical model having optimized parameters by using artificial intelligence to predict the amount of product accumulation in a biochemical reaction pathway, and a method for constructing an artificial intelligence model having the constructed mathematical model (hereinafter, this method may be referred to as the "model construction method of the present invention" as it encompasses both a method for constructing a mathematical model and a method for constructing an artificial intelligence model.) Naturally, the method for constructing an artificial intelligence model corresponds to a method for producing an artificial intelligence model, which is a product.

[0011] The model construction method of the present invention includes: (1) For a substance or cell whose target molecule in a biochemical reaction pathway is known (hereinafter also referred to as a "known target substance or cell"), the method includes the step of optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model having a mathematical model for each target molecule using a group of data indicating the estimated abundance of the reaction product for each elapsed time (typically, the elapsed time is divided into specific intervals and data for each interval is used; the same applies below) after the substance or cell is brought into contact with the target cell, a group of data indicating the actually measured abundance, and the target molecule of the substance or cell as training data (hereinafter also referred to as "training data for step (1)").

[0012] Furthermore, in the model construction method of the present invention, after the step (1), (2) The method may include a step of optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model using a set of data indicating the estimated abundance of reaction products over time, weighted for each target molecule, and a set of data indicating the actual abundance, as training data (hereinafter also referred to as "training data for step (2)") for substances or cells whose target molecules in the pathway are unknown (hereinafter also referred to as "unknown target substances or cells"). The mathematical model in step (2) may include a step of corresponding to each combination of all molecules that may be targeted by the known target substances or cells, and including each parameter optimized in step (1).

[0013] As used herein, "comprise(s)" or "comprising" means the inclusion, but not limitation, of the elements that follow the phrase, and thus implies the inclusion of the elements that follow the phrase, but not the exclusion of any other elements.

[0014] Biochemical reaction pathways that can be used in the model construction method of the present invention include, for example, the complement pathway, coagulation pathway, platelet activation pathway, antigen presentation pathway, central carbon metabolism pathway, fatty acid metabolism pathway, amino acid metabolism pathway, nucleotide metabolism pathway, drug metabolism pathway, gene expression control pathway, RNA regulation pathway, translation control pathway, ubiquitin-proteasome pathway, neurotransmission pathway, calcium signaling pathway, cAMP signaling pathway, G protein-coupled receptor signaling pathway, phosphorylation signaling pathway, and cell adhesion and junction signaling pathway, with the complement pathway being preferred. Furthermore, as used herein, "target molecule in a biochemical reaction pathway" refers to a protein or protein complex targeted by a known target substance or cell that promotes the production or degradation of each product in the biochemical reaction pathway or inhibits said promotion. "Target molecule" encompasses not only a single molecule but also a combination of multiple target molecules (i.e., a test substance or test cell targets multiple molecules in the same biochemical reaction pathway). Additionally, "targeted" means that the function of the target molecule in a biochemical reaction pathway is promoted or inhibited by a substance or cell that is brought into contact with the target cell.

[0015] Examples of the complement pathway include C1q, which binds to target proteins or surfaces and initiates complement reactions; C1r and C1s, which are serine proteases that cleave and activate C4 and C2; C2 and C4, which form the C4b2a complex and activate C3; and C3, which forms the C3b4b2a complex and becomes the C3C5 convertase. Examples of products in the biochemical reaction pathway include complement factors in the complement pathway, such as C2, C2b, C3(H2O) / Bb complex, C3a, C3b, C3b / Bb complex, C4, C4a, C4b, C4b / C2a complex, C3b / C3b / Bb complex, C3b / C4b / C2a complex, C5a, C5b, and membrane attack complex (MAC), with MAC, C3a, and C5a being preferred.

[0016] The known target substance or cell used in step (1) above is not particularly limited, but examples thereof include cell extracts, cell culture supernatants, microbial fermentation products, extracts derived from marine organisms, plant extracts, purified or crude proteins (e.g., antibodies, antibody fragments, etc.), peptides, non-peptide compounds, synthetic low-molecular-weight compounds, and natural compounds (e.g., nucleic acids, genes, etc.). Examples of cell types include differentiated cells such as neurons, oligodendrocytes, erythrocytes, monocytes (e.g., lymphocytes (NK cells, B cells, T cells, monocytes, dendritic cells, etc.)), granulocytes (e.g., eosinophils, neutrophils, basophils), megakaryocytes, epithelial cells (e.g., retinal pigment epithelial cells, etc.), endothelial cells (e.g., vascular endothelial cells, hepatic sinusoidal endothelial cells, etc.), muscle cells, fibroblasts (e.g., skin cells, etc.), hair cells, hepatocytes, gastric mucosal cells, intestinal cells, spleen cells, pancreatic cells (e.g., exocrine pancreatic cells, etc.), brain cells, lung cells, kidney cells, and adipocytes, as well as immature cells thereof.

[0017] The target cells used in the present invention are not particularly limited as long as they can measure the amount of a product present in a biochemical reaction pathway, for example, the amount of accumulated product, and may be cultured cells or cells isolated from a living organism. From the viewpoint of suppressing variability among cells, the cells are preferably cultured cells, for example, cells induced to differentiate from stem cells. Examples of such cell types include the same cells as those listed above for the known target substance or cell.

[0018] The method for producing (inducing differentiation of) the cells from stem cells can be appropriately selected from known methods depending on the type of target cells. For example, methods for producing vascular endothelial cells include those described in Yamashita J. et al., Nature, 408(6808):92-96 (2000) and Narazaki G. et al., Circulation, 29;118(5):498-506 (2008). Methods for producing nerve cells include those described in Yan Y. et al., Stem Cells Trans Med, 2:862-870 (2013), Kondo T. et al., Cell Stem Cell, 12:487-496 (2013), WO2015 / 020234, and Doi D. et al., Stem Cell Reports, 2(3):337-50 (2014). Methods for producing oligodendrocytes include those described in Kawabata S. et al., Stem Cell Reports, 6(1):1-8 (2016), etc.; methods for producing hepatocytes include, for example, the methods described in Cai J. et al., Hepatology, 45:1229-1239 (2007) and Cai J. et al., Hepatology, 45:1229-1239 (2007); methods for producing retinal pigment epithelial cells include, for example, the methods described in Yan Y. et al., Stem Cells Trans Med, 2:862-870 (2013) and WO2015 / 053375; methods for producing hematopoietic progenitor cells include, for example, the methods described in WO2013 / 075222, WO2016 / 076415, and Liu S. et al., Cytotherapy, 17:344-358 (2015); and methods for producing erythrocytes or erythroid progenitor cells include, for example, Miharada K. et al., Nat. Biotechnol., 24:1255-1256 (2006), Kurita R. et al., PLoS One, 8:e59890 (2013), and the like; methods for producing megakaryocytes or platelets include, for example, the methods described in Yamamizu K. et al., J. Cell Biol., 189:325-338 (2010) and Laflamme M. et al., Nat. Biotechnol., 25:1015-1024 (2007); methods for producing T cells include, for example, the methods described in WO2016 / 076415 and WO2017 / 221975; and methods for producing skeletal muscle cells include, for example, the methods described in WO2013 / 073246, Uchimura T. et al., Stem Cell Research, 25:98-106 (2017), and Shoji E. et al., Science Reports, 5:12831 (2015), and methods for producing cardiomyocytes include, for example, the method described in Shimoji K. et al., Cell Stem Cell, 6:227-237 (2010).

[0019] Specifically, to produce vascular endothelial cells from stem cells, for example, stem cells are cultured in a medium containing a serum substitute (e.g., B-27), BMP4, and a GSK-3 inhibitor (e.g., CHIR99021) to induce mesodermal cells, and then the cells are cultured in a medium containing VEGF and folskolin. To produce hepatocytes from stem cells, stem cells are cultured in a medium containing a Wnt protein (e.g., Wnt3a) and activin A to induce endodermal cells, and then the cells are cultured in a medium containing a serum substitute (B-27) and FGF2 to produce hepatocytes.

[0020] An example of the "stem cell" is a pluripotent stem cell. As used herein, the term "pluripotent stem cell" refers to a stem cell that can differentiate into tissues and cells with various different morphologies and functions in the body and has the ability to differentiate into cells of any of the three germ layers (endoderm, mesoderm, and ectoderm). Pluripotent stem cells include, but are not limited to, induced pluripotent stem cells (sometimes referred to herein as "iPS cells"), embryonic stem cells (ES cells), embryonic stem cells derived from cloned embryos obtained by nuclear transfer (nuclear transfer embryonic stem cells: ntES cells), pluripotent germ stem cells, and embryonic germ stem cells (EG cells). As used herein, the term "pluripotent stem cell" refers to a stem cell that has the ability to differentiate into cells of a limited number of lineages. Examples of "pluripotent stem cells" include dental pulp stem cells, stem cells derived from oral mucosa, hair follicle stem cells, cultured fibroblasts, and somatic stem cells derived from bone marrow stem cells. Preferred pluripotent stem cells are ES cells and iPS cells. When the pluripotent stem cells are ES cells or any cells derived from human embryos, the cells may be cells produced by destroying the embryo or cells produced without destroying the embryo, but preferably cells produced without destroying the embryo. The stem cells are preferably derived from mammals (e.g., mice, rats, hamsters, guinea pigs, dogs, monkeys, orangutans, chimpanzees, and humans), and more preferably derived from humans.

[0021] "Induced pluripotent stem cells" refer to cells obtained by reprogramming mammalian somatic cells or undifferentiated stem cells by introducing specific factors (nuclear reprogramming factors). Currently, there are various types of "induced pluripotent stem cells," including iPS cells established by Yamanaka et al. by introducing four factors, Oct3 / 4, Sox2, Klf4, and c-Myc, into mouse fibroblasts (Takahashi K, Yamanaka S., Cell, (2006) 126: 663-676); human-derived iPS cells established by introducing the same four factors into human fibroblasts (Takahashi K, Yamanaka S., et al. Cell, (2007) 131: 861-872); Nanog-iPS cells established by selecting cells using Nanog expression as an indicator after introducing the above four factors (Okita, K., Ichisaka, T., and Yamanaka, S. (2007). Nature 448, 313-317); and iPS cells created using a method that does not include c-Myc (Nakagawa M, Yamanaka S., et al. al. Nature Biotechnology, (2008) 26, 101-106), and iPS cells established by introducing six factors using a virus-free method (Okita K et al. Nat. Methods 2011 May;8(5):409-12, Okita K et al. Stem Cells. 31(3):458-66.) can also be used. Other examples that can be used include induced pluripotent stem cells established by introducing four factors, OCT3 / 4, SOX2, NANOG, and LIN28, as developed by Thomson et al. (Yu J., Thomson JA. et al., Science (2007) 318: 1917-1920), induced pluripotent stem cells developed by Daley et al. (Park IH, Daley GQ. et al., Nature (2007) 451: 141-146), and induced pluripotent stem cells developed by Sakurada et al. (JP Patent Publication No. 2008-307007). In addition, all published papers (e.g., Shi Y., Ding S., et al., Cell Stem Cell, (2008) Vol 3, Issue 5, 568-574; Kim JB., Scholer HR., et al., Nature, (2008) 454, 646-650; Huangfu D., Melton DA., et al., Nature Biotechnology, (2008) 26, No 7, Any of the induced pluripotent stem cells known in the art and described in the Japanese Patent Laid-Open Nos. 2008-307007, 2008-283972, US2008-2336610, US2009-047263, WO2007 / 069666, WO2008 / 118220, WO2008 / 124133, WO2008 / 151058, WO2009 / 006930, WO2009 / 006997, and WO2009 / 007852 can be used. Various iPS cell lines established by the NIH, RIKEN, Kyoto University, and other institutions can be used as induced pluripotent cell lines. For example, human iPS cell lines include RIKEN's HiPS-RIKEN-1A strain, HiPS-RIKEN-2A strain, HiPS-RIKEN-12A strain, and Nips-B2 strain, and Kyoto University's 253G1 strain, 201B7 strain, 409B2 strain, 454E2 strain, 606A1 strain, 610B1 strain, and 648A1 strain. In addition, iPS cells may be derived from patients with a disease.

[0022] ES cells are stem cells that are established from the inner cell mass of early mammalian embryos (for example, blastocysts) such as humans and mice, and have the ability to proliferate through pluripotency and self-renewal. ES cells were discovered in mice in 1981 (MJ Evans and MH Kaufman (1981), Nature 292:154-156), and subsequently, ES cell lines were established in humans, monkeys, and other primates (JA Thomson et al. (1998), Science 282:1145-1147; JA Thomson et al. (1995), Proc. Natl. Acad. Sci. USA, 92:7844-7848; JA Thomson et al. (1996), Biol. Reprod., 55:254-259; JA Thomson and VS Marshall (1998), Curr. Top. Dev. Biol., 38:133-165). ES cells can be established by extracting the inner cell mass from the blastocyst of a fertilized egg of a target animal and culturing the inner cell mass on a fibroblast feeder. Methods for establishing and maintaining human and monkey ES cells are described, for example, in US Pat. No. 5,843,780; Thomson JA, et al. (1995), Proc. Natl. Acad. Sci. USA 92:7844-7848; Thomson JA, et al. (1998), Science. 282:1145-1147; Suemori H. et al. (2006), Biochem. Biophys. Res. Commun., 345:926-932; Ueno M. et al. (2006), Proc. Natl. Acad. Sci. USA 103:9554-9559; Suemori H. et al. (2001), Dev. Dyn., 222:273-279; Kawasaki H. et al. (2002), Proc. Natl. Acad. Sci. USA, 99:1580-1585; Klimanskaya I. et al. (2006), Nature. 444:481-485, etc.Alternatively, ES cells can be established using only a single blastomere from an embryo at the cleavage stage before the blastocyst stage (Chung Y. et al. (2008), Cell Stem Cell 2: 113-117), or can be established using developmentally arrested embryos (Zhang X. et al. (2006), Stem Cells 24: 2669-2676). Regarding "ES cells," various mouse ES cell lines established by inGenious targeting laboratory, Inc., RIKEN (Riken), and other institutions are available, while various human ES cell lines established by the University of Wisconsin, NIH, RIKEN, Kyoto University, National Center for Child Health and Development, Cellartis, and other institutions are available. For example, human ES cell lines that can be used include CHB-1 to CHB-12, RUES1, RUES2, and HUES1 to HUES28 strains distributed by ESI Bio, H1 and H9 strains distributed by WiCell Research, and KhES-1, KhES-2, KhES-3, KhES-4, KhES-5, SSES1, SSES2, and SSES3 strains distributed by RIKEN.

[0023] nt ES cells are ES cells derived from cloned embryos produced by nuclear transfer technology and have almost the same properties as ES cells derived from fertilized eggs (Wakayama T. et al. (2001), Science, 292:740-743; S. Wakayama et al. (2005), Biol. Reprod., 72:932-936; Byrne J. et al. (2007), Nature, 450:497-502). Specifically, nt ES (nuclear transfer ES) cells are established from the inner cell mass of blastocysts derived from cloned embryos obtained by replacing the nucleus of an unfertilized egg with the nucleus of a somatic cell. To generate nt ES cells, nuclear transfer technology (Cibelli JB et al. (1998), Nature Biotechnol., 16:642-646) is combined with ES cell generation technology (mentioned above) (Wakayama Sayaka et al. (2008), Experimental Medicine, Vol. 26, No. 5 (Special Issue), pp. 47-52). In nuclear transfer, the nucleus of a somatic cell is injected into an enucleated unfertilized mammalian egg, and the egg can be reprogrammed by culturing it for several hours.

[0024] Pluripotent germline stem cells (GS cells) are pluripotent stem cells derived from germline stem cells (GS cells). Similar to embryonic stem cells, these cells can be induced to differentiate into various cell lineages. For example, when transplanted into mouse blastocysts, chimeric mice can be produced (Kanatsu-Shinohara M. et al. (2003) Biol. Reprod., 69:612-616; Shinohara K. et al. (2004), Cell, 119:1001-1012). They are capable of self-renewal in culture medium containing glial cell line-derived neurotrophic factor (GDNF). Furthermore, germline stem cells can be obtained by repeated passage under culture conditions similar to those for embryonic stem cells (Takebayashi M. et al. (2008), Experimental Medicine, Vol. 26, No. 5 (Special Issue), pp. 41-46, Yodosha, Tokyo, Japan).

[0025] EG cells are derived from primordial germ cells (PGCs) during the fetal stage and possess pluripotency similar to that of ES cells. They can be established by culturing PGCs in the presence of substances such as LIF, bFGF, and stem cell factor (Matsui Y. et al. (1992), Cell, 70:841-847; JL Resnick et al. (1992), Nature, 359:550-551).

[0026] In addition, target cells may be cells that constitute organoids. Such organoids include, for example, three-dimensional structures made only from vascular endothelial cells, three-dimensional structures made from vascular endothelial cells and hepatocytes (characterized by having a vascular network made of vascular endothelial cells between hepatocytes), etc. These organoids are generally made by known culture methods (for example, Nature Cell Biology 18, 246-254 (2016)). Specifically, as the method for making organoids, for example, the method described in Nahmias Y. et al., Tissue Eng., 12(6), 2006, pp.1627-1638, etc. can be mentioned.

[0027] The vascular endothelial cells used to prepare the organoids may be hemogenic endothelial cells (HECs) or non-hemogenic endothelial cells (non-HECs). HECs are endothelial cells capable of producing hematopoietic stem cells (having hematopoietic potential), and are also called blood cell-producing endothelial cells. Alternatively, either HECs or non-HECs may be used alone, or both HECs and non-HECs may be used, or their progenitor cells may be used, or any combination of HECs, non-HECs, and their progenitor cells may be used. Examples of progenitor cells of vascular endothelial cells include cells present in the differentiation process from Flk-1 (CD309, KDR)-positive vascular endothelial cell progenitors (e.g., lateral plate mesoderm cells) to HEC cells (see Cell Reports 2, 553-567, 2012).

[0028] The hepatocytes used to produce such organoids may be differentiated hepatocytes (differentiated hepatocytes), or cells that have committed to hepatocyte differentiation but have not yet differentiated into hepatocytes (undifferentiated hepatocytes), so-called hepatic progenitor cells (e.g., hepatic endoderm cells). Differentiated hepatocytes or undifferentiated hepatocytes may be cells collected from a living body (isolated from the liver in the living body), or may be cells obtained by differentiating pluripotent stem cells such as ES cells and iPS cells, hepatic progenitor cells, or other cells capable of differentiating into hepatocytes. Cells capable of differentiating into hepatocytes can be produced, for example, according to K. Si-Taiyeb, et al. Hepatology, 51 (1): 297-305 (2010), T. Touboul, et al. Hepatology. 51 (5): 1754-65. (2010).

[0029] In addition, mesenchymal stem cells may be used in the preparation of the organoids. Such mesenchymal stem cells may be differentiated cells (differentiated mesenchymal cells) or cells that have been determined to differentiate into mesenchymal cells but have not yet differentiated into mesenchymal cells (undifferentiated mesenchymal cells), so-called mesenchymal stem cells. Terms used by those skilled in the art, such as mesenchymal stem cells, mesenchymal progenitor cells, and mesenchymal cells (R. Peters, et al. PLoS One. 30;5(12):e15689.(2010)), refer to the "mesenchymal cells" in this specification.

[0030] The ratio of the number of vascular endothelial cells to hepatocytes used to prepare the organoids can be adjusted within an appropriate range, for example, so that the components contained in the collected culture supernatant are the desired ones. In one embodiment of the present invention, the ratio of the number of vascular endothelial cells to hepatocytes (vascular endothelial cells:hepatocytes) is typically 1:0.1 to 5, preferably 1:0.1 to 2. When mesenchymal stem cells are used, the ratio can also be adjusted appropriately, and in a preferred embodiment, the ratio of the number of cells (vascular endothelial cells:hepatocytes:stem cells) is 7:10:1.

[0031] When preparing organoids, it is preferable to use an organoid medium that is a mixture of vascular endothelial cell medium and hepatocyte medium (culture medium) at an appropriate ratio (for example, 1:1). Examples of vascular endothelial cell medium include DMEM / F-12 (Gibco), Stempro-34 SFM (Gibco), Essential 6 Medium (Gibco), Essential 8 Medium (Gibco), EGM (Lonza), BulletKit (Lonza), EGM-2 (Lonza), BulletKit (Lonza), EGM-2 MV (Lonza), VascuLife EnGS Comp Kit (LCT), Human Endothelial-SFM Basal Growth Medium (Invitrogen), and human microvascular endothelial cell growth medium (TOYOBO). The culture medium for vascular endothelial cells contains B27 Supplements (GIBCO), BMP4 (bone morphogenetic protein 4), GSKβ inhibitors (e.g., CHIR99021), VEGF (vascular endothelial growth factor), FGF2 (fibroblast growth factor (also known as bFGF (basic fibroblast growth factor))), folskolin, SCF (stem cell factor), TGFβ receptor inhibitors (e.g., SB431542), Flt-3L (Fms-related tyrosine kinase 3 ligand), IL-3 (interleukin 3), and IL-6 (interleukin 6), TPO (thrombopoietin), hEGF (recombinant human epidermal growth factor), hydrocortisone, ascorbic acid, IGF1, FBS (fetal bovine serum), antibiotics (e.g., gentamicin, amphotericin B), heparin, L-glutamine, phenol red, BBE, etc. The amount of these additives to be added can be determined appropriately by one skilled in the art with reference to typical culture conditions for culturing vascular endothelial cells.

[0032] Examples of media for hepatocytes include RPMI (Fujifilm) and HCM (Lonza). Hepatocyte media may contain one or more additives selected from Wnt3a, activin A, BMP4, FGF2, FBS (fetal bovine serum), HGF (hepatocyte growth factor), oncostatin M (OSM), dexamethasone (Dex), etc. Those skilled in the art can appropriately determine the amounts of these additives to be added with reference to typical culture conditions for culturing hepatocytes. Alternatively, media for hepatocytes may include media for hepatocytes containing at least one selected from ascorbic acid, BSA-FAF, insulin, hydrocortisone, and GA-1000; HCM BulletKit (Lonza) from which hEGF (recombinant human epidermal growth factor) has been removed; RPMI1640 (Sigma-Aldrich) supplemented with 1% B27 Supplements (GIBCO) and 10 ng / mL hHGF (Sigma-Aldrich); or a 1:1 mixture of GM BulletKit (Lonza) and HCM BulletKit (Lonza) from which hEGF (recombinant human epidermal growth factor) has been removed, to which dexamethasone, oncostatin M, and HGF have been added.

[0033] The data set indicating the estimated abundance used in step (1) can be obtained, for example, by substituting the concentration of the target known substance or cell in the culture medium and the abundance of the reaction product at the time of contact of the substance or cell as initial values into a mathematical model in which initial values are set for all parameters.

[0034] Examples of mathematical models used in step (1) for the complement pathway include, but are not limited to, those described in Sagar A et al., PLoS One, (2017) 12(11):e0187373, Zewde N et al., PloS One, (2016) 11(3):e0152337, Hirayama H et al., Biosystems, (1996) 39(3):173-185, Korotaevskiy AA et al., Math Biosci, (2009) 222(2):127-143, and Liu B et al., PLoS Comput Biol, (2011)7(1):e1001059. Examples of coagulation pathways include, but are not limited to, models disclosed in Hockin MF et al., J Biol Chem, (2002) 277(21):18322-18333, Chatterjee MS et al., PLoS Comput Biol, (2010) 6(9):e1000950, Nayak S et al., CPT Pharmacometrics Syst Pharmacol, (2015) 4(7):396-405, and http: / / biomodels.caltech.edu / content / model-of-the-month?year=2019&month=04. Examples of fatty acid metabolic pathways include, but are not limited to, Wallstab C et al., FEBS J, (2017) 284(19):3245-3261, Patt AC et al., Math Biosci, (2015) 262:167-181, Adiels M et al., J Lipid Res, (2005) 46:58-67, Shorten PR & Upreti GC, Biochim Biophys Acta, (2005) 1736:94-108, etc. Examples of combinations of central carbon metabolic pathways and fatty acid metabolic pathways include, but are not limited to, Dean JT et al., Biophys J., (2010) 98(8):1385-1395, etc.Examples of phosphorylation signaling pathways include, but are not limited to, those described in Sulaimanov N et al., Wiley Interdiscip Rev Syst Biol Med., (2017) 9(4):e1379 and the models cited therein (listed in Table 1 with reference to the literature). Examples of drug metabolism pathways include, but are not limited to, Toth A et al., PLoS One, (2015) 10(2):e0115533, Stamatelos SK et al., BMC Syst Biol, (2011) 5:16, Stamatelos SK et al., J Theor Biol, (2013) 317:244-256, Suzuki T et al., Trends Pharmacol Sci, (2013) 34(6):340-346, Zhang Q et al., PLoS computational biology, (2007) 3:e24, Zhang Q et al., Toxicology and applied pharmacology, (2009) 237:345-356, and the like. The mathematical model used in step (1) may be a model listed on a mathematical model database site (e.g., BioModels (http: / / biomodels.caltech.edu / )). In addition, it is preferable to adopt, as the initial values of the mathematical model, those described in the literature disclosing the above-mentioned mathematical model.

[0035] [Table 1]

[0036] Contact between a known target substance or cells and target cells can be achieved by adding the known target substance or cells to the culture medium for the target cells, or by transferring or seeding the target cells into a medium to which the known target substance or cells have been added. When the cells are organoids, the above-mentioned "cell culture medium" can be interpreted as "organoid culture medium." Furthermore, the known target substance or cells may be further contacted with a substance that promotes or inhibits the production of reaction products during, before, or after contact with the target cells. The present inventors have previously found that contacting cells with a complement source promotes the formation of MAC in the cells. Therefore, substances that promote the production of the above-mentioned reaction products include, but are not limited to, complement sources in the complement pathway, and can be selected appropriately depending on the type of biochemical reaction pathway targeted. Furthermore, when targeting the complement pathway, the known target substance or cells may be further contacted with a complement activator. Examples of such complement activators include LPS (Lipopolysaccharide) and zymosan, a polysaccharide found in yeast cell walls.

[0037] Examples of the complement source include complement contained in serum (e.g., human serum, bovine serum, horse serum, sheep serum, goat serum, pig serum, llama serum, dog serum, chicken serum, donkey serum, cat serum, rabbit serum, guinea pig serum, hamster serum, rat serum, mouse serum, etc.), complement contained in the culture supernatant of liver or vascular endothelial cells (including organoids of these organs), or each complement component (e.g., C1q, C1r, C1s, C2, C4, C3, C3a, C5s, C3b, C5a, C5b6789, and combinations thereof). Furthermore, the present inventors have previously confirmed that artificially produced organoids secrete complement. Therefore, the culture supernatant of the organoids can also be used as a complement source. Furthermore, the inventors have confirmed that combining multiple types of complement sources promotes MAC formation more than combining each source alone, and therefore it is also preferable to combine multiple types of complement or complement sources (e.g., combining serum with organoid culture supernatant).

[0038] Examples of media used for cell culture include BME medium, BGJb medium, CMRL 1066 medium, Glasgow MEM medium, Improved MEM (IMEM) medium, Improved MDM (IMDM) medium, Medium 199 medium, Eagle MEM medium, αMEM medium, DMEM medium (High glucose, Low glucose), DMEM / F12 medium, Ham's medium, RPMI 1640 medium, Fischer's medium, and mixed media thereof. In addition, the above-mentioned medium for vascular endothelial cells, medium for hepatocytes, medium for organoids, or mixed media thereof may also be used.

[0039] In addition to the above-mentioned components, the medium may contain, as needed, amino acids, L-glutamine, GlutaMAX (product name), non-essential amino acids, vitamins, antibiotics (e.g., Antibiotic-Antimycotic (sometimes referred to as AA in this specification), penicillin, streptomycin, or a mixture thereof), antibacterial agents (e.g., amphotericin B), antioxidants, pyruvic acid, buffers, inorganic salts, and the like.

[0040] The culture period is not particularly limited, but typically ranges from 10 minutes to 7 days, preferably 0.5 to 96 hours, and more preferably 4 to 48 hours. The culture temperature is also not particularly limited, but is preferably 30 to 40°C (e.g., 37°C). The carbon dioxide concentration in the culture vessel is, for example, about 5%.

[0041] The actual abundance of the reaction product can be measured using, but is not limited to, immunological assays using antibodies (e.g., anti-CD59 antibodies capable of detecting MAC, anti-LDH antibodies, etc.) over time, specifically, ELISA, immunostaining, Western blotting, flow cytometry, fluorescent imaging, etc. To quantify expression levels based on the results of immunostaining, Western blotting, etc., the obtained image data can be digitized using image processing software (e.g., ImageJ, etc.). Furthermore, if the reaction product has enzymatic activity, the abundance of the reaction product can be measured using the enzymatic activity as an index, as needed, using a kit, etc.

[0042] Optimization of each parameter of the mathematical model for each target molecule in step (1) can be carried out, for example, by the following steps. (I) inputting the learning data for step (1) into an artificial intelligence model (hereinafter also referred to as a "first artificial intelligence model") having a mathematical model for each target molecule including parameters for which initial values are set, searching for optimal values of parameters of the mathematical model corresponding to the target molecule of the known target substance or cell using the artificial intelligence model, and updating the parameters to the optimal values; (II) A process of adjusting each parameter of the mathematical model for each target molecule by sequentially repeating step (I) (where the artificial intelligence model is replaced with a first artificial intelligence model in which each parameter has been updated to an optimal value in step (I)) for each known target substance or cell, using a group of data indicating the estimated accumulation amount of reaction products over time for a plurality of other known target substances or cells, a group of data indicating the actual abundance, and the target molecule as learning data.

[0043] In the above step (I), the search for the optimal parameter values can be carried out, for example, by using an artificial intelligence model set with an objective function of minimizing the sum of the squares of the differences between the estimated abundance and the actually measured abundance of the reaction product per elapsed time, and minimizing the objective function by utilizing or combining a genetic algorithm, a gradient method (e.g., steepest descent method, stochastic gradient descent method, etc.), etc. The search for the optimal parameter values may be performed at high speed by approximating the objective function using surrogate modeling based on Bayesian optimization.

[0044] The target unknown substance or cell used in step (2) above is not particularly limited, but examples thereof include cell extracts, cell culture supernatants, microbial fermentation products, extracts derived from marine organisms, plant extracts, purified or crude proteins (e.g., antibodies, antibody fragments, etc.), peptides, non-peptide compounds, synthetic low-molecular-weight compounds, and natural compounds (e.g., nucleic acids, genes, etc.). Examples of cell types include differentiated cells such as neurons, oligodendrocytes, erythrocytes, monocytes (e.g., lymphocytes (NK cells, B cells, T cells, monocytes, dendritic cells, etc.)), granulocytes (e.g., eosinophils, neutrophils, basophils), megakaryocytes, epithelial cells (e.g., retinal pigment epithelial cells, etc.), endothelial cells (e.g., vascular endothelial cells, hepatic sinusoidal endothelial cells, etc.), muscle cells, fibroblasts (e.g., skin cells, etc.), hair cells, hepatocytes, gastric mucosal cells, intestinal cells, spleen cells, pancreatic cells (e.g., exocrine pancreatic cells, etc.), brain cells, lung cells, kidney cells, and adipocytes, as well as immature cells thereof.

[0045] As described above, the artificial intelligence model used in step (2) (hereinafter also referred to as the "second artificial intelligence model") has a mathematical model that corresponds to each combination of all molecules that may be targeted by the target unknown substance or cell, and includes each parameter optimized in step (1). Here, the number of all the combinations of molecules is, for example, given that the number of target pathways is m and the maximum number of targets of biochemical reaction pathways that a compound can target is k,

[0046]

number

[0047] where n=0 means that the compound does not target any of the complement pathways. If necessary, the number of combinations can be increased by selecting an unknown target substance or a cell inhibition mode (e.g., reversible inhibition, competitive inhibition, uncompetitive inhibition, mixed inhibition, noncompetitive inhibition, allosteric inhibition, etc.).

[0048] In step (2), a set of data indicating the estimated abundance of reaction products over time, weighted for each target molecule, and a set of data indicating the actually measured abundance are input as training data to the second artificial intelligence model, thereby optimizing each parameter of the mathematical model for each target molecule. By inputting the training data into the artificial intelligence model, an estimated graph is typically created, with the horizontal axis representing time and the vertical axis representing the accumulated amount of reaction products multiplied by the weighting (there will be as many estimated graphs as there are combinations of all the molecules mentioned above).

[0049] Optimization of each parameter of the mathematical model for each target molecule in step (2) can be carried out, for example, by the following steps. (III) inputting the learning data for step (2) into a second artificial intelligence model, searching for optimal values for each parameter of a mathematical model for each target molecule using the second artificial intelligence model, and updating the parameters to the optimal values; and (IV) A step of determining, from among the mathematical models for each target molecule in which each parameter has been updated to an optimal value in step (III), the mathematical model that can best explain the measured abundance is the mathematical model corresponding to the target molecule of each substance or cell.

[0050] The weighting in step (2) can be determined based on, for example, the similarity of the target unknown substance or cell, such as the similarity of the substance's structural data or the similarity of the cell's properties. Typically, a higher weight is assigned to a substance with a higher similarity, and a lower weight is assigned to a substance with a lower similarity. By assigning weights based on factors other than the target molecule, such as structural data, it is possible to prevent divergence in parameter search. The similarity of the substance's structural data can be determined, for example, such that the more common skeletons or the more complex the skeletons shared, the higher the similarity. Examples of such skeletons include, but are not limited to, functional groups, hydrocarbon skeletons (e.g., alkanes, alkenes, alkynes, etc.), aromatic hydrocarbons (e.g., benzene, naphthalene, etc.), and cyclic hydrocarbons in the case of low-molecular-weight compounds, and amino acid sequences, α-helix structures, β-structures, etc. in the case of peptides such as proteins. More specifically, the similarity of structural data of small molecular weight compounds can be determined, for example, by a method using the Topological Fragment Spectra (TFS) representation of chemical structures (Takahashi Y, et al., Advances in Molecular Similarity, (1998) 2:93-104) or a method using the KCOMBU (Cemical compound COMparison using Build-Up algorithm) (Kawabata T, et al., J Chem Inf Model, (2011) 51:1775-1787). Peptide similarity can be determined using the BLAST or DSSP algorithm (Kabsch W, et al., Biopolymers, (1983) 22(12):2577-637).

[0051] The weighting values may be directly set by the above method, or may be adjusted by machine learning using variables with the above values as initial values. For example, the weighting values can be adjusted by using a mathematical model having the parameters adjusted in step (1) or (2) as constants, setting the initial values set by the above method, and then training the model to increase the probability of selecting the correct target pathway. This step can improve the accuracy of predicting the target pathway for compounds with unknown target pathways. Step (2) may also include (V) a step of verifying the validity of the weighting in step (III). The validity of the weighting can be verified, for example, by cross-validation using a comparison between a group of data indicating the actual abundance of the unknown target substance or reaction product after contact between cells and target cells and a group of data indicating the estimated abundance of the reaction product by a mathematical model including the weighting values.

[0052] The data set indicating the estimated abundance in step (III) can be obtained, for example, by substituting the concentration of each substance or cell in the culture medium and the abundance of the reaction product at the time of contact of the substance or cell with the target cell as initial values into a mathematical model including each parameter for each target molecule optimized in step (1).

[0053] To prevent divergence in the parameter search, it is also preferable to limit the search range of the parameter range. Therefore, in one embodiment of the present invention, a method is provided in which the search in step (III) is performed within a range set as a search range of parameters estimated after contact of a substance or cell with a target cell from the parameter set for each target molecule optimized in step (1). The search range of such a parameter range can be set, for example, by previously narrowing down global solutions using a genetic algorithm or the like.

[0054] In the above step (III), the search for the optimal parameter values can be carried out, for example, by using an artificial intelligence model set with an objective function of minimizing the sum of the squares of the differences between the estimated abundance and the actually measured abundance of the reaction product per elapsed time, and by minimizing the objective function by utilizing or combining a genetic algorithm or a gradient method (e.g., steepest descent method, stochastic gradient descent method, etc.). The search for the optimal parameter values may be carried out at high speed by approximating the objective function using surrogate modeling based on Bayesian optimization. This speed-up makes it possible to perform searches using the genetic algorithm or gradient method tens of thousands of times or even millions of times, thereby further improving accuracy.

[0055] The actual amount of the reaction product can be measured in the same manner as described in step (1) above. In addition, the method for contacting the unknown target substance or cells with the target cells and the method for culturing the cells can be similar to the method for contacting the known target substance or cells with the target cells.

[0056] After step (IV) or (V), the method may include the steps of: (VI) updating the parameters optimized in step (1) to the optimized parameters of the mathematical model determined in step (IV); and (VII) adjusting each parameter of the mathematical model for each target molecule by sequentially repeating steps (III)-(VI) above (where the artificial intelligence model in step (III) is replaced with an artificial intelligence model having a mathematical model whose parameters have been updated to optimal values in step (VI)) for each substance or cell using a group of data indicating the estimated abundance of reaction products for each elapsed time weighted for each target molecule and a group of data indicating the actual abundance for a plurality of other unknown target substances or cells.

[0057] The mathematical model that can best explain the measured abundance in step (IV) may be, for example, the one that has the smallest sum of squares of the differences between the weighted estimated abundance of the reaction product for each elapsed time and the measured abundance. Thus, in one embodiment of the present invention, step (IV) may be a step of determining that the mathematical model that has the smallest sum of squares of the differences between the weighted estimated abundance of the reaction product for each elapsed time and the measured abundance is the mathematical model corresponding to the target molecule of the substance or cell.

[0058] The known target substance, unknown target substance, complement source, and the like used in the present invention can be produced by existing synthetic methods, for example, in the case of low molecular weight compounds. Furthermore, in the case of peptides such as proteins, they can be prepared by known peptide synthesis methods (e.g., solid-phase synthesis, liquid-phase synthesis, etc.) or methods for isolating them from living organisms, or commercially available products can be used. Alternatively, proteins can be obtained by introducing a protein-encoding gene into host cells such as Escherichia coli and producing the protein. Protein-encoding genes or nucleic acids used in the present invention can be obtained by using genomic DNA extracted from cells containing the gene, or by preparing cDNA from extracted mRNA and amplifying the nucleic acid of the desired length by PCR using the DNA as a template, or by chemical synthesis using a commercially available automated DNA / RNA synthesizer. Proteins obtained in this manner can be purified and isolated by known purification methods, such as solvent extraction, distillation, column chromatography, liquid chromatography, recrystallization, or a combination thereof.

[0059] The construction method of the present invention will be described in detail below using specific examples, but the scope of the present invention is not limited to this construction method. Process (1) Taking the complement pathway as an example, we use a mathematical model reported in Liu B et al., PLoS Comput Biol, (2011) 7(1):e1001059, which describes the complement pathway, which includes 42 substances, using ordinary differential equations for 42 reactions with 85 parameters. The differential equations are specifically disclosed in Text S1 of the above document, but can generally be expressed by the following formula:

[0060]

number

[0061] [t is the elapsed time, x i are variables in a system of ordinary differential equations that describe the concentration levels of individual biomolecular species, r i is species x i is the number of reactions associated with x i If is a reactant in the jth reaction, then c j =-1(C j =+1), n ij denotes the stoichiometric coefficient, g j is the formula g j =p α x a or g j =p α x a x b (Law of Mass Action) or g j =p α x a x b / (p β +x a ) (Michaelis-Menten law) (where the parameter p is the rate constant of the biochemical reaction and is a rational function of a, b ∈ {1, 2, ..., n} and α, β ∈ {1, 2, ..., m}, describing the kinetics of the corresponding reaction.)

[0062] The initial values of the product abundance and parameters are entered as values described in the aforementioned literature (listed in Tables 2 and 3, respectively, with reference to the literature). Note that * in the tables indicates that the parameter is known.

[0063] [Table 2]

[0064] [Table 3]

[0065] The artificial intelligence model for step (1) is constructed by approximating the ordinary differential equations included in the mathematical model as layers of a deep neural network based on the method of Raissi M et al., arXiv:1708.07469. Human iPS cell-derived vascular endothelial cells, which had been complement-activated with zymosan, were contacted with known targets selected from the target library and cultured for 24 hours. During this time, the culture supernatant was collected every six hours, and the reaction products (C3, C5, MAC) in the supernatant were quantified to obtain training data for step (1). Bayesian optimization was performed using the objective function, which minimizes the sum of the squared differences between the estimated abundance output by the artificial intelligence model for step (1) and the training data for step (1). This process was repeated for each known target to obtain an artificial intelligence model parameter set for each target molecule, and this parameter set was input as the initial parameter value for the artificial intelligence model group for step (2).

[0066] Process (2) The target unknown substance selected from the target library was contacted with human iPS cell-derived vascular endothelial cells that had been complement-activated with zymosan using the same procedure as in step (1), to obtain training data for step (2). The structural information of the target unknown substance was compared with the structural information of each known target substance using an algorithm that uses a combinatorial method to calculate an approximate solution for the correspondence between substances based on the maximum common substructure (MCS) of Kawabata, T. (2011) J. Chem. Info. Model. 51, 1775-1787. Similarity was calculated by comparing the structural information of the target unknown substance with the structural information of each known target substance, and a weighting value ranging from 0 to 1 was set depending on the degree of similarity. The similarity was calculated using Flexible Transformation of 3D Conformer by Pairwise MCS Comparison, which is available from the web resource KCOMBU (http: / / strcomp.protein.osaka-u.ac.jp / kcombu / ). To adjust and verify the validity of the weighting values, an artificial intelligence model for step (2) including the weighting values is used, and only the weighting values are specified as the optimization target as a learning option for Bayesian optimization, and optimization is performed using the same procedure as in step (1). The adjusted weighting values are again input into the artificial intelligence model for step (2), and the parameters included in all pathways, excluding the weighting values, are set as the optimization targets, and optimization is performed using the same procedure as in step (1). The objective function error when using the optimal solution finally obtained is then output. This process is repeated for each unknown target substance, and the parameters and weighting values of the mathematical model included in the artificial intelligence model with the smallest objective function error are adopted as the optimal solution. From these weighting values, the probability that the unknown target substance will target each molecule is estimated.

[0067] Numerical calculations such as model construction based on the above process and implementation of optimization algorithms are implemented on MATLAB (Mathworks).

[0068] -2. Methods for predicting target molecules- A target molecule can be predicted using an artificial intelligence model having a mathematical model constructed by the model construction method of the present invention described in 1 above (hereinafter also referred to as the "artificial intelligence model of the present invention"). Therefore, in another embodiment of the present invention, (A) inputting a data group (hereinafter also referred to as "input data group") showing the concentration of each test substance or test cell in the culture medium, the amount of reaction product of the substance or cell at the time of contact with the target cell, and the measured amount of reaction product per elapsed time into the artificial intelligence model of the present invention, and calculating the estimated amount of reaction product of each target molecule per elapsed time using the artificial intelligence model; and (B) determining, from the estimated abundances of the reaction products of each target molecule per elapsed time calculated in step (A), the target molecule that best explains the actually measured abundances as the target of the test substance or test cell. The present invention provides a method for predicting a target molecule in a biochemical reaction pathway of a test substance or a test cell (hereinafter, sometimes referred to as the "prediction method of the present invention"), which comprises:

[0069] The test substance or test cell used in step (A) is not particularly limited, but examples thereof include cell extracts, cell culture supernatants, microbial fermentation products, extracts derived from marine organisms, plant extracts, purified or crude proteins (e.g., antibodies, antibody fragments, etc.), peptides, non-peptide compounds, synthetic low molecular weight compounds, and natural compounds (e.g., nucleic acids, genes, etc.). Examples of cells include differentiated cells such as nerve cells, oligodendrocytes, erythrocytes, mononuclear cells (e.g., lymphocytes (NK cells, B cells, T cells, monocytes, dendritic cells, etc.)), granulocytes (e.g., eosinophils, neutrophils, basophils), megakaryocytes, epithelial cells (e.g., retinal pigment epithelial cells, etc.), endothelial cells (e.g., vascular endothelial cells, hepatic sinusoidal endothelial cells, etc.), muscle cells, fibroblasts (e.g., skin cells, etc.), hair cells, hepatocytes, gastric mucosal cells, intestinal cells, spleen cells, pancreatic cells (e.g., exocrine pancreatic cells, etc.), brain cells, lung cells, kidney cells, and adipocytes, as well as immature cells thereof.

[0070] Test substances or test cells can also be obtained using any of the many approaches to combinatorial library methods known in the art, including (1) biological libraries, (2) synthetic library methods using deconvolution, (3) "one-bead one-compound" library methods, and (4) synthetic library methods using affinity chromatography selection. While the biological library method using affinity chromatography selection is limited to peptide libraries, the other four approaches can be applied to small molecule compound libraries of peptides, non-peptide oligomers, or compounds (Lam (1997) Anticancer Drug Des. 12:145-67). Examples of methods for the synthesis of molecular libraries can be found in the art (DeWitt et al. (1993) Proc. Natl. Acad. Sci. USA 90:6909-13; Erb et al. (1994) Proc. Natl. Acad. Sci. USA 91:11422-6; Zuckermann et al. (1994) J. Med. Chem. 37:2678-85; Cho et al. (1993) Science 261:1303-5; Carell et al. (1994) Angew. Chem. Int. Ed. Engl. 33:2059; Carell et al. (1994) Angew. Chem. Int. Ed. Engl. 33:2061; Gallop et al. (1994) J. Med. Chem. 37:1233-51).Compound libraries can be prepared in solution (see Houghten (1992) Bio / Techniques 13:412-21) or on beads (Lam (1991) Nature 354:82-4), chips (Fodor (1993) Nature 364:555-6), bacteria (U.S. Pat. No. 5,223,409), spores (U.S. Pat. Nos. 5,571,698, 5,403,484, and 5,223,409), plasmids (Cull et al. (1992) Proc. Natl. Acad. Sci. USA 89:1865-9), or phage (Scott and Smith (1990) Science 249:386-90; Devlin (1990) Science 249:404-6; Cwirla et al. al. (1990) Proc. Natl. Acad. Sci. USA 87:6378-82; Felici (1991) J. Mol. Biol. 222:301-10; U.S. Patent Application No. 2002103360).

[0071] The input data group used in step (A) can be obtained by a method similar to the method for acquiring (measuring) the learning data described in 1. The type of target cells, type of culture medium and culture method, and method for contacting the test substance or test cells with the target cells can also be the same as those described in 1.

[0072] In step (B), for example, the target molecule can be determined based on the probability that the test substance or test cell targets each target molecule by comparing the estimated abundance of each target molecule calculated in step (A) with the actual abundance. For example, the sum of the squares of the differences between the estimated abundance of the reaction product for each elapsed time calculated in step (A) and the actual abundance can be calculated for each target molecule, and this value can be input into a softmax function, with the output value being regarded as the probability for each target molecule. The target molecule with the highest probability can then be determined to be the target of the test substance or test cell.

[0073] It is also possible to readjust each parameter of the mathematical model possessed by the artificial intelligence model by using the judgment results obtained by the prediction method of the present invention as learning data and re-training the artificial intelligence model of the present invention as described in the model construction method of the present invention.

[0074] -3. Device for predicting target molecules- Next, the configuration of an apparatus for carrying out the prediction method of the present invention (hereinafter, sometimes referred to as "the apparatus of the present invention") will be specifically shown. Fig. 2 is a block diagram showing a schematic configuration of an example of the device of the present invention. As shown in Fig. 2, the device is configured to have at least a receiving unit M1, a first calculation unit M2, a second calculation unit M3, and an output unit M4.

[0075] The device may be constructed by combining electronic circuits, electric circuits, and independent processing devices, but a preferred embodiment is one in which the above-mentioned units (reception unit M1, first calculation unit M2, second calculation unit M3, output unit M4, etc.) are configured by a computer and a computer program executed thereon. Below, each unit will be explained using an example in which the device is configured by the computer and the computer program. Hereinafter, the device (the computer and the computer program executed thereon) will be explained simply as the "computer."

[0076] The device may be configured as a desktop, notebook, or tablet computer used individually by a user, or as a server computer accessed by one or more terminal computers via a communication path such as a local area network (LAN) or the Internet. The server computer is preferred because it allows multiple terminal computers to use a single common program or common learning data. The computer as the device is appropriately equipped with input devices and display screens necessary for operation, as well as interfaces and peripheral devices necessary for exchanging data with external devices. In the following description, the device is a server computer accessed from an external computer via the Internet, but is not limited to this.

[0077] The receiving unit M1 is a part for receiving input data groups about a test substance or test cell from an external input device or an external computer. The method for acquiring the input data groups is as described in the model construction method and prediction method of the present invention.

[0078] The external device for inputting the input data group to the reception unit M1 is typically an external computer connected to the device so as to be accessible via the Internet, but may also be a memory (USB memory) connected to the device via an interface such as USB, or various data reading devices. The reception unit M1 is configured to accept the input data group from an external computer or various external devices and transfer the input data group to the first calculation unit M2. The reception unit M1 is also configured to transfer the measured abundance for each elapsed time included in the input data group.

[0079] The first calculation unit M2 processes the input data set delivered from the receiving unit M1 and executes step (A) of the prediction method of the present invention to calculate the estimated abundance of each target molecule's reaction product per elapsed time. Specifically, the first calculation unit M2 uses the artificial intelligence model of the present invention as a subroutine, inputs the input data set into the artificial intelligence model, and receives the estimated abundance of each target molecule's reaction product per elapsed time as a calculation result from the artificial intelligence model. The calculation process performed by the artificial intelligence model is as described in the model construction method and prediction method of the present invention. The estimated abundance of each target molecule's reaction product per elapsed time obtained from the artificial intelligence model is delivered to the second calculation unit M3.

[0080] The second calculation unit M3 is a device component that executes step (B) of the prediction method of the present invention. Specifically, it performs calculations to determine the probability that each target molecule is targeted by the test substance or cell by comparing the estimated abundance of each target molecule delivered from the first calculation unit M2 with the measured abundance delivered from the receiving unit M1. The calculation process performed by the second calculation unit M3 is as described in the model construction method and prediction method of the present invention. Specifically, for example, the sum of the squares of the differences between the estimated abundance of the reaction product for each elapsed time delivered from M2 and the measured abundance delivered from M1 is calculated for each target molecule, and this value is input into a softmax function. The output value can be considered as the probability for each target molecule. The calculation results performed by the second calculation unit M3 are delivered to the output unit M4.

[0081] The output unit M4 is a part that outputs the calculation results of the second calculation unit (e.g., the probability that each target molecule will be targeted by the test substance or cell, the predicted results of the target molecule (the target molecule with the highest probability), etc.) to an external device such as an external computer.

[0082] The method of outputting the calculation results in the second calculation unit may be, for example, to simply output each probability in parallel in a table format for each target molecule, with the target molecule with the highest probability being marked, or to output the results in a table format in which the target molecules are sorted in order of decreasing probability, or to output the results in the form of electronic data that can be read into software or the like and displayed in the above manner.

[0083] The sender of the input data group to the reception unit M1 is not particularly limited, and may be a computer that can communicate with the device of the present invention, such as a terminal computer used by researchers, or a server computer on a LAN that receives the input data group from the terminal computer and transmits the input data group to the device of the present invention.

[0084] On the other hand, the predetermined output destination to which the prediction results are output is not particularly limited, and may be a predetermined output destination such as various image display devices, various print output devices, a computer that can communicate with the device of the present invention, etc. In a typical network system in a research department, the sender of the input data group is often also the output destination of the prediction results (for example, when an input data group is sent from a researcher's terminal computer and the prediction results are sent back to the researcher's terminal computer), but the sender of the input data group and the destination of the prediction results may be different from each other when viewed from the device of the present invention.

[0085] -4. Computer program for predicting target molecules- Each of the components constituting the device of the present invention (such as the above-mentioned receiving unit M1, first calculation unit M2, second calculation unit M3, and output unit M4) may be constructed by combining electronic circuits, electric circuits, and independent processing devices, but a preferred embodiment is to configure each of these components using a computer and a program executed on the computer (hereinafter sometimes referred to as the "program of the present invention").

[0086] As shown in the flow chart of Figure 3, the program of the present invention is basically the same as the flow of the prediction method of the present invention described above, and is configured to have a reception unit P1 that causes a computer to function as reception means, a first calculation unit P2 that causes a computer to function as first calculation means, a second calculation unit P3 that causes a computer to function as second calculation means, and an output unit P4 that functions as output means for the prediction result output by the second calculation unit P3.

[0087] The receiving unit P1 is a program portion for receiving input data groups about test substances or test cells from an external input device or an external computer. The method for acquiring the input data groups is as described in the model building method and prediction method of the present invention. The input data groups can each be configured to directly receive input data groups sent by an operator. Alternatively, the input data groups can be configured to access their own storage device (SSD, HDD, DVD-ROM, CD-ROM, etc.) or an external database to import data based on the data name input by the operator.

[0088] The external device for inputting the input data group to the reception unit P1 is typically an external computer connected to the device so as to be accessible via the Internet, but may also be a memory (USB memory) connected to the device via an interface such as USB, or various data reading devices. The reception unit P1 is configured to accept the input data group from an external computer or various external devices and transfer the input data group to the first calculation unit P2. The reception unit P1 is also configured to transfer the measured abundance for each elapsed time included in the input data group.

[0089] The first calculation unit P2 is a program portion that processes the input data group delivered from the receiving unit P1 and executes step (A) of the prediction method of the present invention described above to calculate the estimated abundance of reaction products for each target molecule per elapsed time. Specifically, the first calculation unit P2 uses the artificial intelligence model of the present invention as a subroutine, inputs the input data group into the artificial intelligence model, and receives the estimated abundance of reaction products for each target molecule per elapsed time as a calculation result from the artificial intelligence model. The calculation process performed by the artificial intelligence model is as described in the model construction method and prediction method of the present invention. The estimated abundance of reaction products for each target molecule per elapsed time obtained from the artificial intelligence model is delivered to the second calculation unit P3.

[0090] The second calculation unit P3 is a program portion that executes step (B) of the prediction method of the present invention described above. The calculation process performed by the second calculation unit P2 is as described in the model construction method and prediction method of the present invention. Specifically, for example, the sum of the squares of the differences between the estimated abundance of the reaction product for each elapsed time delivered from P2 and the actual abundance delivered from P1 is calculated for each target molecule, and this value is input into a softmax function, and the output value can be considered as the probability for each target molecule. The calculation results performed by the second calculation unit P3 are delivered to the output unit P4.

[0091] The output unit P4 is a program part that outputs the above-mentioned probability for each target molecule calculated by the second calculation unit P3. The destination and method of outputting the calculation results of the second calculation unit are as explained in the output unit M4. [Industrial Applicability]

[0092] The artificial intelligence model of the present invention can calculate the estimated abundance of reaction products for each target molecule over time by inputting various input data groups, and can predict the target molecule of a test substance or cell based on the calculated data. This allows for accurate prediction of target molecules in complex biochemical reaction pathways such as the complement pathway in screening for disease therapeutic drugs, contributing to highly reliable drug discovery screening.

[0093] This application is based on patent application No. 2020-099489 filed in Japan (filing date: June 8, 2020), the contents of which are incorporated in their entirety herein. [Explanation of symbols]

[0094] 1. Target molecule prediction device 2. Trained AI model 3. Source of input data (external computer, etc.)

Claims

1. (1) A step of optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model having a mathematical model for each target molecule using, as learning data, a data group indicating the estimated abundance of a reaction product for each elapsed time after contacting a substance or cell with a target cell, the substance or cell having a known target molecule in a biochemical reaction pathway, the data group indicating the actual abundance, and the target molecule of the substance or cell; and (2) A method for constructing a mathematical model containing optimized parameters, the method including a step of optimizing each parameter of the mathematical model for each target molecule by training an artificial intelligence model having the mathematical model using a group of data indicating the estimated abundance of reaction products for each elapsed time, weighted for each target molecule, and a group of data indicating the actual abundance, as training data for a substance or cell whose target molecule is unknown, wherein the mathematical model corresponds to each combination of all molecules that may be targeted by the substance or cell, and includes each parameter optimized in step (1).

2. The step (1) (I) inputting a data group showing the estimated abundance of a reaction product for each elapsed time after contacting a substance or cell whose target molecule in a biochemical reaction pathway is known with a target cell, a data group showing the actually measured abundance, and the target molecule of each substance or cell into an artificial intelligence model having a mathematical model for each target molecule including parameters for which initial values are set, and using the artificial intelligence model to search for optimal values for the parameters of the mathematical model corresponding to the target molecule of the substance or cell, and updating the parameters to the optimal values; (II) a step of adjusting each parameter of the mathematical model for each target molecule by sequentially repeating the step (I) for each substance or cell using a data group indicating the estimated abundance of the reaction product for each elapsed time for a plurality of other substances or cells, a data group indicating the actually measured abundance, and the target molecule (provided that the artificial intelligence model is replaced with an artificial intelligence model having a mathematical model in which each parameter has been updated to an optimal value in the step (I)); The method of claim 1 , comprising:

3. The step (2) (III) a step of inputting a group of data indicating the estimated abundance of a reaction product for each elapsed time weighted for each target molecule and a group of data indicating the actually measured abundance for a plurality of substances or cells whose target molecules are unknown into an artificial intelligence model having a mathematical model, searching for an optimal value for each parameter of the mathematical model for each target molecule using the artificial intelligence model, and updating the parameters to the optimal value, wherein the mathematical model corresponds to each combination of molecules that may be targeted by each substance or cell, and includes each parameter optimized in step (1); (IV) determining, from among the mathematical models for each target molecule in which each parameter has been updated to an optimal value in the step (III), the mathematical model that can best explain the actually measured abundance is the mathematical model corresponding to the target molecule of each substance or cell; 3. The method of claim 1 or 2, comprising:

4. The method of claim 3, further comprising the step of: (V) verifying the validity of the weighting in step (III).

5. (VI) updating the parameters optimized in step (1) to the optimized parameters of the mathematical model determined in step (IV); and (VII) a step of adjusting each parameter of the mathematical model for each target molecule by sequentially repeating steps (III)-(VI) for each substance or cell using a data group indicating the estimated abundance of the reaction product for each elapsed time weighted for each target molecule and a data group indicating the actually measured abundance, the steps (III)-(VI) (wherein the artificial intelligence model in step (III) is replaced with an artificial intelligence model having a mathematical model whose parameters have been updated to optimal values in step (VI)).

5. The method of claim 3 or 4, comprising:

6. The method according to any one of claims 2 to 5, wherein the group of data indicating the estimated abundance in step (I) is obtained by substituting the concentration of the substance or cell in the medium and the abundance of the reaction product at the time of contact of the substance or cell as initial values into a mathematical model in which initial values are set for all parameters.

7. The method according to any one of claims 1 to 6, wherein the artificial intelligence model in step (1) is a model set to minimize the sum of squares of the difference between the estimated abundance and the actually measured abundance of the reaction product for each elapsed time as an objective function.

8. The data group indicating the estimated abundance in step (III) is input to a mathematical model including each parameter for each target molecule optimized in step (1), and the concentration of each substance or cell in the culture medium. The method according to any one of claims 3 to 7, wherein the amount of the reaction product present at the time of contact of the substance or cell with the target cell is assigned as an initial value.

9. The method according to any one of claims 3 to 8, wherein the search in step (III) is performed within a range set as a search range, which is a parameter range estimated after contact of a substance or cell with a target cell from the parameter set for each target molecule optimized in step (1).

10. The method according to any one of claims 3 to 9, wherein step (IV) is a step of determining that a mathematical model for which the sum of the squares of the differences between the weighted estimated abundances of the reaction products for each elapsed time and the actual measured abundances is the smallest is the mathematical model corresponding to the target molecule of the substance or cell.

11. The method according to any one of claims 2 to 10, wherein the search in step (I) is carried out by surrogate modeling.

12. The method according to any one of claims 3 to 11, wherein the weighting of the estimated abundance in step (III) is set based on structural data of the substance.

13. The method according to any one of claims 1 to 12, wherein the biochemical reaction pathway is the complement pathway.

14. The method according to any one of claims 1 to 13, wherein the substance is selected from the group consisting of a low molecular weight compound, a nucleic acid, a peptide, and an antibody.

15. (A) inputting a group of data indicating the concentration of each test substance or test cell in the culture medium, the amount of reaction product at the time of contact of the substance or cell with the target cell, and the amount actually measured over time into an artificial intelligence model having a mathematical model constructed by the method according to any one of claims 1 to 14, and calculating the estimated amount of reaction product for each target molecule over time using the artificial intelligence model; (B) A method for predicting a target molecule in a biochemical reaction pathway of a test substance or a test cell, comprising a step of calculating the probability that the test substance or test cell targets each target molecule by comparing the estimated abundance of each target molecule calculated in step (A) with the actual abundance.

16. A computer program for predicting a target molecule in a biochemical reaction pathway of a test substance or a test cell, comprising: Computer, (P1) a receiving unit that receives at least a group of data indicating the concentration of each test substance or test cell in a culture medium, the amount of a reaction product present at the time of contact of the test substance or test cell with a target cell, and the amount actually measured at each elapsed time; (P2) a first calculation unit that inputs the input data group into an artificial intelligence model having a mathematical model constructed by the method according to any one of claims 1 to 14, and calculates the estimated abundance of the reaction product of each target molecule per elapsed time; (P3) a second calculation unit that calculates the probability that each target molecule is targeted by the test substance or test cell by comparing the estimated abundance of each target molecule obtained by the first calculation unit with the actual abundance; and (P4) An output unit that outputs the calculation result obtained by the second calculation unit The program for causing the computer to function as a

17. 1. A device for predicting a target molecule in a biochemical reaction pathway of a test substance or a test cell, comprising: (M1) a receiving unit that receives at least a group of data indicating the concentration of each test substance or test cell in the culture medium, the amount of reaction product at the time of contact of the substance or cell with the target cell, and the measured amount at each elapsed time; (M2) a first calculation unit that inputs the input data group into an artificial intelligence model having a mathematical model constructed by the method according to any one of claims 1 to 14, and calculates the estimated abundance of the reaction product for each target molecule per elapsed time; (M3) a second calculation unit that calculates the probability that the test substance or cell targets each target molecule by comparing the estimated abundance of each target molecule obtained by the first calculation unit with the actual abundance; (M4) an output unit that outputs the calculation result obtained by the second calculation unit; An apparatus having:

Citation Information

Patent Citations

  • Cell-based assays by matching and their use

    JP2015519876A

  • Evaluation of tgf-β cell signaling pathway activity using mathematical modeling of target gene expression

    JP2018503354A

  • Systems and methods for patient-specific prediction of drug response from cell-based genomics

    JP2018527644A

  • Method and apparatus for estimating biochemical reaction pathway of intracellular molecule group using membrane potential time series

    JP2019154299A

  • Apparatus and methods for assessing an ability of an organism(s) to metabilize toxid compounds

    US20200043567A1