Importance determination method, importance determination apparatus, and computer program
Patent Information
- Application Number
- JP2022141468
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-09-03
AI Technical Summary
Current methods for evaluating cell maturation from stem cells, such as iPS cells and ES cells, rely on single evaluation parameters, making it difficult to comprehensively assess multiple experimental results and determine the maturity of cells.
An importance determination method using machine learning to construct an estimation model that outputs the importance of experimental conditions and results, enabling comprehensive evaluation of cell maturation by calculating SHAP values and utilizing models like SHAP and XGBoost to prioritize experimental conditions for maturation.
Enables the acquisition of knowledge about cell maturation by determining the degree of importance of experimental conditions, allowing for mutual and comprehensive evaluation of multiple experiments, thereby optimizing the maturation process.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to methods of determining the importance, and in particular to determining the importance of each of one or more experimental conditions in an experiment on the maturation of immature cells. [Background technology]
[0002] Traditionally, the determination of experimental conditions for cell culture has been largely based on the experience and intuition of researchers. In recent years, there have been efforts to quantify such experience and intuition. In addition, many experimental conditions and experimental results have been quantified, and large amounts of data have been handled, making it nearly impossible for humans to judge the correlation between each condition.
[0003] On the other hand, machine learning can be used to check the quality of cells from data on experimental conditions and experimental results, and optimize the experimental conditions. For example, Patent Document 1 (JP 2019-193587 A) discloses a technology for optimizing culture conditions and environments using a decision tree, with the aim of controlling the culture environment to realize culture conditions according to the culture purpose.
[0004] In the field of biology, especially in in vitro evaluation using cells, evaluation focusing on one evaluation parameter of a phenomenon occurring in a living body is common. However, in a living body, various phenomena occur due to the complex changes of various parameters. For this reason, the current in vitro evaluation focusing on one parameter is considered to be insufficient as an evaluation that mimics a living body. Therefore, whether or not there is a response to a drug is determined using the results of multiple parameters obtained in an in vitro evaluation and machine learning. For example, Non-Patent Document 1 (EK Lee, et al, Stem Cell Reports, 2017, 9, 1560-1570, November 14, 2017) discloses a technology that, when a drug is added to cardiomyocytes, the drug response is determined from all parameters obtained using machine learning, rather than determining the drug response using one parameter as in the past.
[0005] It is known that cells differentiated from stem cells (iPS cells and ES cells) are generally immature. It is said that these cells are functionally and structurally closer to fetal cells than adult cells in the case of human cells. Research on maturing stem cell-derived cells is being conducted worldwide. For example, regarding iPS cardiomyocytes, Non-Patent Document 2 (K. Ronaldson - Bouchard et al., Nature, 2018, 556, 239, April 12, 2018) discloses a technique for co-culturing iPS cardiomyocytes and fibroblasts, constructing a three-dimensional structure, and then performing electrical stimulation to mature the cells. It is known that iPS cardiomyocytes can be matured by co-culturing with other cells, three-dimensional culture, and electrical stimulation, and attempts are being made to mature the cells by simultaneously using these three techniques.
[0006] In Non-Patent Document 2, many experiments such as fluorescent immunostaining, electron microscope observation, action potential, contractile force, Ca transient, and gene expression level are performed as maturation evaluation experiments. However, there are problems with such maturation research. Mature cells and immature cells are compared, and one evaluation parameter from each experiment is extracted to determine the presence or absence of maturation. For this reason, when multiple evaluation experiments are performed in one study, the results of multiple experiments cannot be evaluated mutually and comprehensively. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] JP 2019-193587 A [Non-patent literature]
[0008] [Non-Patent Document 1] EK Lee, et al, Stem Cell Reports, 2017, 9, 1560-1570, November 14, 2017 [Non-Patent Document 2] K. Ronaldson - Bouchard et al., Nature, 2018, 556, 239, April 12, 2018 Summary of the Invention [Problem to be solved by the invention]
[0009] Cells differentiated from stem cells (iPS cells and ES cells) are generally known to be immature. In the case of human cells, these cells are said to be closer in function and structure to fetal cells than to adult cells. Research into maturing these stem cell-derived cells is being conducted around the world. However, there are several challenges in this maturation research. One of the biggest challenges is that mature and immature cells are compared, and a single evaluation parameter from each experiment is extracted to determine whether or not the cells have matured. For this reason, when multiple evaluation experiments are performed within a single study, the results of the multiple experiments cannot be evaluated mutually and comprehensively.
[0010] For example, the problem with Patent Document 1 is that the culture environment cannot be controlled from multiple experimental result data. Only one experimental result data is used, which is not suitable for maturation experiments that require mutual and comprehensive evaluation of multiple experiments. In other words, in the research on the above cells, it is of great significance to obtain knowledge for cell maturation.
[0011] The present invention has been devised in view of the above circumstances, and an object of the present invention is to provide a technique that enables knowledge regarding cell maturation to be obtained. [Means for solving the problem]
[0012] The present disclosure provides an importance determination method for determining experimental conditions necessary for maturation of immature cells, which includes the steps of: constructing an estimation model that is subjected to a learning process by inputting one or more experimental conditions and one or more experimental results, and inputting one or more experimental conditions and one or more experimental results into the estimation model, thereby obtaining a calculation result of the importance of each of the one or more experimental conditions for each of the one or more experimental results. Effect of the Invention
[0013] According to the importance determination method of the present disclosure, when estimation data including one or more experimental conditions and one or more experimental results for an experiment for maturing cells are input, the importance of each of the one or more experimental results for each of the one or more experimental conditions is obtained, thereby making it possible to obtain knowledge for maturing cells. [Brief description of the drawings]
[0014] [Figure 1] 1 shows a configuration of an importance determination device 1. [Diagram 2] 2 is a block diagram showing functions realized by a CPU 12 and a storage 14 of the importance determination device 1. FIG. [Diagram 3] 2 is a flowchart of a process performed by the importance determination device 1. [Figure 4] FIG. 10 is a diagram illustrating an example of a relationship between experiment condition data and experiment result data. [Diagram 5] FIG. 13 is a diagram showing an example of an array of SHAP values calculated for each of a plurality of types of experimental conditions. [Figure 6] FIG. 1 is a diagram showing a schematic correlation between experimental conditions and SHAP values. [Figure 7] FIG. 13 is a diagram for explaining a method for calculating a total value of SHAP values under each experimental condition. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the following drawings, the same or corresponding parts are designated by the same reference numerals, and the description thereof will not be repeated.
[0016] [Device configuration] 1 shows a configuration of an importance determination device 1. The importance determination device 1 is realized, for example, by a general-purpose computer. The importance determination device 1 includes a display 10, an input device 11, a CPU (Central Processing Unit) 12, a memory 13, a storage 14, and a data bus 15. The display 10, the input device 11, the CPU 12, the memory 13, and the storage 14 are interconnected via the data bus 15. The input device 11 is realized, for example, by a keyboard and / or a mouse. The CPU 12 uses the memory 13 as a primary memory and executes a program stored in the storage 14. The storage 14 stores various data necessary for the execution of the program and various data representing information on the execution results of the program.
[0017] FIG. 2 is a block diagram showing functions realized by the CPU 12 and storage 14 of the importance determination device 1. As shown in FIG.
[0018] The storage 14 stores teacher data 141, a trained model 142, and an operating program 143. The operating program 143 is a program for learning and using an estimation model. In this specification, an estimation model that has been subjected to a learning process is also referred to as a "trained model 142." Each function of the CPU 12 shown in FIG. 2 is realized by the CPU 12 executing the operating program 143.
[0019] The importance determination device 1 is configured to be able to communicate with the worker terminal 2. The communication may be wired or wireless.
[0020] The importance determination device 1 receives experimental conditions and experimental results from the operator terminal 2 or the input device 11. The CPU 12 accepts the input of the experimental conditions and experimental results in the experimental condition and result acquisition unit 121. The experimental conditions are experimental conditions for an experiment for maturing cells. The experimental results are experimental results obtained in an experiment for maturing cells. The experimental condition and result acquisition unit 121 saves all input data in the teacher data 141 of the storage 14 and inputs them to the data selection unit 122. The experimental condition and result acquisition unit 121 further saves the experimental results for each pair of the experimental conditions and the experimental results, and the "importance" assigned to each pair of the experimental conditions and the experimental results, as teacher data 141 in the storage 14. The importance represents the degree of importance of one or more types of experimental conditions for each of one or more types of experimental results constituting the pair of the experimental conditions and the experimental results. The importance saved as the teacher data 141 may be given in advance to each pair of the experimental conditions and the experimental results by the operator. Moreover, the importance may be the importance output from importance calculation section 125 for a pair of experimental results.
[0021] A portion of the input data is selected in the data selection unit 122, and the selected data is input to the model construction unit 123 as data for constructing a model.
[0022] The selection ratio of data in the data selection unit 122 may be set arbitrarily. In one implementation example, 30% of the data of all input models is selected as data to be input to the model construction unit 123. The selection method may be specified by the type of estimation model. More specifically, when the estimation model is a regression model, data is selected randomly, and when the estimation model is a classification model, it is desirable to select at least one sample of data to be classified as a pair. That is, when the estimation model is a classification model, when classifying pairs of experimental conditions and experimental results into ◯ and ×, data may be selected so that at least one set of experimental condition and experimental result data to be classified into “◯” and at least one set of experimental condition and experimental result data to be classified into “×” are included as data to be input to the model construction unit 123. An example of classification is whether or not the experimental condition includes the addition of a given drug. In this example, the experimental condition is given the classification “◯” when it represents the addition of drug A, and given the classification “×” when it represents the non-addition.
[0023] The model construction unit 123 performs a learning process of the estimation model using the data input from the data selection unit 122 (selected data) or the teacher data 141 (all data input to the experimental condition / result acquisition unit 121 and all data of importance input from the importance calculation unit 125). In the learning process, the input data is "data of experimental conditions" and "data of experimental results", and the output data is an estimation model. By this learning process, the estimation model outputs "data of experimental results" estimated by inputting "data of experimental conditions". By the learning process, a trained model 142 is created. The created trained model 142 is stored in a storage. In one implementation example, the trained model 142 is stored in the storage 14 as an estimation model and one or more parameters (identified as a result of learning) of the estimation model.
[0024] The model acquisition unit 124 acquires the trained model 142. In one implementation example, acquiring the trained model 142 means reading out parameters of an estimation model. The importance calculation unit 125 calculates the importance using the trained model 142 acquired in the model acquisition unit 124 and the data acquired in the data selection unit 122.
[0025] The calculated importance is scored in a scoring calculation unit 126, and information for determining experimental conditions correlated with maturity is output in a determination unit 127. The calculated data is also output to teacher data 141. The determined information is output to the operator terminal via a display.
[0026] [Processing flow] Fig. 3 is a flowchart of the process executed by the importance determination device 1. The process of Fig. 3 is realized by the CPU 12 executing the operation program 143. Hereinafter, the flow of the process relating to importance determination will be described with reference to Fig. 3.
[0027] In step S10, the importance determination device 1 sets the experimental conditions. The experimental conditions are acquired in step S10, for example, by accepting input of the experimental condition data from the worker terminal 2. In step S12, the importance determination device 1 acquires the results of the experiment (experimental results) according to the experimental conditions set in step S10. The experimental results are acquired in step S12, for example, by accepting input of the experimental result data from the worker terminal 2.
[0028] In step S 14 , the importance determination device 1 inputs the data on the experiment conditions and the experiment results input from the operator terminal 2 to the experiment condition / result acquisition unit 121 .
[0029] In step S16, the importance determination device 1 determines whether or not the trained model 142 has been constructed. In one implementation example, if the trained model 142 is stored in the storage 14, it determines that the trained model 142 has been constructed, and if the trained model 142 is not stored in the storage 14, it determines that the trained model 142 has not been constructed. If the importance determination device 1 determines that the trained model 142 has been constructed (YES in step S16), it proceeds to control step S22, and if not (NO in step S16), it proceeds to control step S18.
[0030] In step S18, the importance determination device 1 causes the data selection unit 122 to select data for learning the estimation model. The data of the experimental conditions and the experimental results used for learning the estimation model is learning data different from the data of the experimental conditions and the experimental results input in step S14 used for calculating the importance (data input to the trained model 142 in step S23 described later).
[0031] In step S20, the importance determination device 1 causes the model construction unit 123 to carry out a learning process for the estimation model.
[0032] In step S22, the importance determination device 1 causes the model acquisition unit 124 to acquire the trained model 142.
[0033] In step S23, the importance determination device 1 functions as the importance calculation unit 125 to input the "data on experimental conditions and experimental results" input in step S14 to the trained model 142. In this specification, the data input in step S23 to the trained model 142 out of the "data on experimental conditions and experimental results" input in step S14 is also referred to as estimation data.
[0034] In step S24, the importance determination device 1 causes the importance calculation unit 125 to calculate, in the trained model 142, the importance of each of one or more experimental conditions included in the "data of experimental conditions and experimental results" for each of one or more experimental results constituting the "data of experimental conditions and experimental results", and obtains the calculation results.
[0035] In step S25, the importance determination device 1 inputs the importance acquired in step S24 to the teacher data 141. This causes the teacher data (importance) of each of the one or more types of experimental conditions included in the "experimental condition / experimental result data" to be saved.
[0036] In step S26, the importance determination device 1 causes the scoring calculation unit 126 to perform scoring for the experimental conditions included in the "experimental condition-experiment result data". In one implementation example, the scoring is a calculation of the sum of the importance of one or more types of experimental conditions. Then, in step S26, the importance determination device 1 outputs a predetermined number of types of experimental conditions that are ranked high as a result of the scoring as experimental conditions necessary for maturation. The output may be a notification to the operator terminal 2, or may be displayed on the display 10. The number of types of experimental conditions to be output may be "3", as shown as condition x8, condition x2, and condition x4 in FIG. 5 described later. However, such a number of types is merely an example, and may be appropriately set for each situation to which the technology of this embodiment is applied.
[0037] In step S28, the determination unit 127 is caused to determine the experimental conditions necessary for maturation, and the experimental conditions correlated with maturation are listed in ascending or descending order, and the result is notified to the operator terminal. In one implementation example, the SHAP value is used to determine the experimental conditions necessary for maturation. More specifically, if the SHAP value of a certain experimental condition is equal to or greater than a given threshold, the experimental condition is identified as an experimental condition necessary for maturation.
[0038] [Scoring] In this embodiment, the importance of the experimental results is scored in the following manner. First, immature cells are matured using any number of experimental conditions (x1 to xi (i=1, 2, 3, . . .)), and j experimental results (y1 to y j (j=1,2,3…)), we use machine learning to find one experimental result y j The importance of each experimental condition x1 to xi g(y j ) (j=1,2,3...) and calculate the importance of each experimental condition g(y j ) can be summed to score the importance of multiple experimental results.
[0039] [Estimation model and its training] The estimation model and its training are explained.
[0040] In this embodiment, linear regression, logistic regression, support vector machine (SVM), decision tree, random forest, deep learning (neural network), naive Bayes, k-means, principal component analysis (PCA), LightGBM, XGBoost, Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Generative Adversarial Networks (GAN), etc. can be used as estimation models.
[0041] Among them, decision tree methods such as decision tree, random forest, LightBGM, and XGBoost that can check the classification conditions are preferable. Furthermore, LightBGM and XGBoost, which are gradient boosting decision tree methods that can handle missing values, are preferable as estimation models used in this embodiment because they enable learning even when some types of values of the experimental conditions are missing in the set of experimental conditions and experimental results used for learning.
[0042] In addition, it is preferable that the machine learning model uses regression or classification appropriately so as to increase the prediction accuracy.
[0043] Furthermore, in order to calculate the importance from the estimated model calculated by the above method, it is desirable to use SHapley Additive exPlanations (SHAP) or Permutation Importance, which can quantify the correlation between the experimental conditions and the experimental results. Among them, SHAP can calculate the positive and negative contribution rates of each variable in a prediction model that tends to be a black box, and can be expressed as a value called the SHAP value (Shapley Value).
[0044] By calculating the SHAP value for each type of experimental condition using one or more types of experimental conditions and one or more types of experimental results for maturation assessment, important experimental conditions for maturation can be predicted from the set of one or more types of extracted experimental results. Then, by adding up the SHAP values calculated for each of one or more types of experimental conditions for each of one or more types of experimental results for maturation assessment, it is possible to score the set of one or more types of experimental conditions, making it possible to comprehensively assess maturation using a set of multiple experimental results, and further making it possible to predict experimental conditions that are correlated with that maturation. Furthermore, by using multiple machine learning methods, it is possible to improve accuracy and understanding.
[0045] As an example of using multiple machine learning techniques, an example of using SHAP and XGBoost will be described below.
[0046] FIG. 4 is a diagram showing an example of a relationship between experimental condition data and experimental result data. i (i=1,2,3...) to perform a maturation operation on n immature cells, and the experimental results y1 to y j When (j=1,2,3...) is obtained, create a table of experimental condition data and experimental result data as shown in Figure 4.
[0047] Figure 4 shows that i types of experimental conditions and j types of experimental results are registered for each of n samples. For the first sample, the value of the experimental condition x1 is shown as "x1-1" and the value of the experimental result y1 is shown as "y1-1". For the nth sample, the value of the experimental condition x1 is shown as "x1-n" and the value of the experimental result y1 is shown as "y1-n".
[0048] Then, using SHAP and XGBoost, the absolute value of the SHAP value corresponding to each of the j types of experimental results for each of the i types of experimental conditions is calculated.
[0049] FIG. 5 is a diagram showing an example of an array of SHAP values calculated for each of a plurality of experimental conditions. In FIG. 5, the absolute values of the SHAP values are shown in descending order. In the example shown in FIG. 5, the experimental result y j It has been shown that the experimental conditions with the highest correlation with (j=1,2,3...) are x8, x2, x4, in that order.
[0050] Fig. 6 is a diagram showing a schematic diagram of the correlation between the experimental conditions and the SHAP value. Fig. 6 shows the correlation between each experimental condition and the SHAP value. In Fig. 6, when the numerical value (feature value) of the experimental condition corresponding to the SHAP value is high, the spot is shown as a circle. When the numerical value of the experimental condition corresponding to the SHAP value is low, the spot is shown as a triangle.
[0051] An example of a feature value for an experimental condition is whether or not a given drug was added. If the experimental condition indicates that the given drug was added, the feature value is high. If the experimental condition indicates that the given drug was not added, the feature value is low.
[0052] Another example of the numerical value of the experimental condition is the number of days of culture. If the number of days of culture is longer than the given number of days, the numerical value is considered to be high. If the number of days of culture is less than the given number of days, the numerical value is considered to be low.
[0053] Yet another example of the numerical value of the experimental condition represents the culture method. If the culture method is three-dimensional culture, the numerical value is considered to be high. If the culture method is two-dimensional culture, the numerical value is considered to be low.
[0054] In the example of Fig. 6, the condition with a higher value (feature value) in condition x8 has a higher SHAP value, and the condition with a lower value in condition x2 has a higher SHAP value. By being provided with information such as that shown in Fig. 6, the operator can determine whether or not to set each experimental condition to a higher value in order to mature the cells.
[0055] For example, the following specific examples are assumed for the conditions x8, x2, and x4 shown in FIG.
[0056] Condition x8=Number of culture days Condition x2 = Addition or non-addition of drug A Condition x4 = Culture method (2D culture / 3D culture) In this case, the operator arbitrarily corresponds the longest culture period of 14 days to the above numerical value being "high," the addition of drug A to the above numerical value being "high," and three-dimensional culture to the above numerical value being "high."
[0057] In Figure 6, among the three experimental conditions, condition x8, condition x2, and condition x4, the SHAP value was the largest in condition x8, followed by condition x2, and the smallest in condition x4.
[0058] The order of magnitude of the SHAP values indicates the order of the degree of contribution of each experimental condition to cell maturation. More specifically, among the three experimental conditions, "number of days of culture" has the greatest influence on maturation, followed by "addition or non-addition of drug A," and then "culture method."
[0059] The information presented in FIG. 6 further indicates the direction for cell maturation for each experimental condition.
[0060] More specifically, from FIG. 6, it can be understood whether the value (feature value) of each experimental condition is positively or negatively correlated with the SHAP value.
[0061] In the example of Figure 6, the longer the culture period, the more positive the SHAP value. Therefore, Figure 6 shows the possibility that the longer the culture period, the more likely it is that the longer the culture period will contribute to cell maturation.
[0062] On the other hand, in the example of Figure 6, the SHAP value and positive values are higher for "low = no addition" when comparing the presence or absence of addition of drug A. Therefore, Figure 6 shows the possibility that not adding drug A contributes to cell maturation.
[0063] FIG. 7 is a diagram for explaining a method for calculating the sum of the SHAP values under each experimental condition. j The sum of the SHAP values of each is shown as a bar graph.
[0064] In FIG. 7, for example, the sum of the SHAP values for condition X1 is calculated by dividing the results y1 to y j In other words, the sum of the SHAP values for each experimental condition is calculated for the condition X1 of each experimental condition by subtracting the results y1 to y j The calculation of the sum of the SHAP values of the results y1 to y j However, the scoring method is not limited to addition, and other methods such as multiplication may also be used.
[0065] In the scoring table in Figure 7, the experimental conditions with a high combined SHAP value are the experimental conditions necessary for maturation. Furthermore, by checking the positive and negative correlations of each experimental condition using Figure 6, it is possible to optimize the experimental conditions necessary for maturation (whether to set each experimental condition to correspond to a high numerical value or a low numerical value for cell maturation).
[0066] [Specific examples of determining experimental conditions for cell maturation] Furthermore, we specifically illustrate the maturation of iPS cardiomyocytes.
[0067] The experimental conditions for maturing iPS cardiomyocytes include the number of days in culture, cell lot, cell diameter, survival rate at seeding, type of medium, type of added factors, two-dimensional / three-dimensional culture, and medium volume. The experimental results include the expression levels of genes such as αMHC, βMHC, KCNJ2, and SERCA2 by PCR, contractile force, parameters of Ca transient waveform (beat rate, amplitude, peak width (PWD10-90), each slope (rising slope, falling slope), waveform area, etc.), and responsiveness to drugs (E-4031, isoproterenol, propranolol, etc.). In this case, when SHAP and XGBoost are used, the scoring value increases with the type of added factors and the number of days in culture, which are conditions for maturing iPS cardiomyocytes in previous studies.
[0068] The immature cells may be derived from any animal species, living organisms, or induced pluripotent stem cells. The animal species may be any of mammals, reptiles, fish, and birds, including humans, mice, and rats. Primary cells collected from living organisms, and cells differentiated from induced pluripotent stem cells, including iPS cells and ES cells, may also be used. The cell species may be any cell that constitutes a living organism, including nerve cells, cardiomyocytes, hepatocytes, fibroblasts, epithelial cells, vascular endothelial cells, α cells, β cells, δ cells, and T cells.
[0069] The experimental conditions used for machine learning may be any of the cell state, culture environment, operation conditions, etc. In the case of cells derived from a living organism, the cell state may be the environment in which the collected organism grew, the sex, the operator at the time of collection, the time of collection, etc. In the case of cells derived from induced pluripotent stem cells, the conditions such as the animal species, race, age, and sex from which the cells used to create the induced pluripotent stem cells were collected, and the type and amount of reagents used to create the induced pluripotent stem cells, the lot of the reagent, the operation procedure, the operator, etc. may be used as experimental conditions. In addition, the type and amount of reagents used to induce differentiation from induced pluripotent stem cells into each cell, the lot of the reagent, the operation procedure, the operator, etc. may be used as experimental conditions. Furthermore, in any cell, the survival rate, the average diameter of the cells, the expression level of various biomarkers, etc. may be used as experimental conditions.
[0070] The culture environment and operation conditions may include the amount of medium, the interval between medium changes, the components of the medium, the types of factors added to the medium, the culture temperature, the carbon dioxide concentration, the temperature, humidity, and time of day during operation.
[0071] The experimental results are generally those of experiments that determine maturation. In the case of PCR or Western blotting, the expression level of RNA or protein may be used to determine maturation. Various experimental results such as membrane potential, extracellular potential, Ca transient, mechanical response, metabolic activity, and response to drugs may also be used.
[0072] According to the importance determination method implemented by the importance determination device of the present embodiment described above, a calculation result of the importance of one or more experimental results included in the estimation data is obtained for one or more experimental conditions included in the estimation data, which can provide useful knowledge about cell maturation that requires mutual and comprehensive evaluation of experimental conditions in multiple experiments.
[0073] The embodiments disclosed herein should be considered to be illustrative and not restrictive in all respects. The scope of the present invention is defined by the claims, not by the description of the embodiments described above, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0074] 1 Importance determination device, 2 worker terminal, 10 display, 11 input device, 13 memory, 14 storage, 15 data bus, 121 result acquisition unit, 122 data selection unit, 123 model construction unit, 124 model acquisition unit, 125 importance calculation unit, 126 scoring calculation unit, 127 judgment unit, 141 teacher data, 142 trained model, 143 operating program.
Claims
1. 1. A method for determining the experimental conditions necessary for maturation of immature cells, comprising: receiving input of estimation data including one or more experimental conditions and one or more experimental results for an experiment to mature cells; constructing an estimation model that has been subjected to a learning process by inputting one or more types of experimental conditions and one or more types of experimental results, so that the estimation model outputs one or more types of experimental results in response to the input of one or more types of experimental conditions; and obtaining a calculation result of the importance of each of the one or more experimental conditions for each of the one or more experimental results by inputting one or more experimental conditions and one or more experimental results into the estimation model.
2. The importance determination method according to claim 1 , wherein the estimation model is a regression model or a classification model.
3. 3. The method for determining importance according to claim 1, wherein the estimation model is a model using a decision tree, principal component analysis, random forest, support vector machine, linear discriminant analysis, quadratic discriminant analysis, binomial logistic regression, stochastic descendant descent, Adaboost, artificial neural network, naive Bayes method, Gaussian process, or nearest neighbor method.
4. The importance determination method according to claim 1 or 2, wherein the estimation model is a model using a gradient boosting decision tree.
5. The method for quantifying the importance is SHAP or Permutation Importance.
3. The importance determination method according to claim 1 or 2.
6. The importance determination method described in claim 1 or claim 2 further comprises a step of summing the importance of one or more types of experimental conditions contained in the quantified estimation data, and outputting a predetermined number of types of experimental conditions with the highest summed value as experimental conditions necessary for maturation.
7. The step of receiving the input includes receiving an input of the estimation data from an operator terminal; The importance determination method according to claim 6 , wherein the outputting step includes outputting experimental conditions required for the maturation to the operator terminal.
8. 3. The importance determination method according to claim 1, wherein the immature cells are mammalian cells.
9. 3. The method for determining importance according to claim 1, wherein the immature cells are cells derived from a living organism or from induced pluripotent stem cells.
10. 3. The importance determination method according to claim 1, wherein the immature cells are nerve, heart, liver, pancreas, kidney, or immune cells.
11. one or more processors; a memory; 3. An importance determination device, wherein the memory causes the one or more processors to perform the importance determination method according to claim 1 when the memory is executed by the one or more processors.
12. A computer program that, when executed by a computer, causes the computer to implement the importance determination method according to claim 1 or 2.