Computer system, doping element search method, computer-implemented method, and computer program
Patent Information
- Application Number
- PCT/JP2026/009275
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-03-11
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026009275_01102026_PF_FP_ABST
Abstract
Description
Computer system, dopant element search method, computer-implemented method, and computer program
[0001] The present disclosure relates to a computer system, a dopant element search method, a computer-implemented method, and a computer program.
[0002] Conventionally, methods using machine learning (ML) have achieved results in many fields including the material science field. For example, conventionally, photoelectrode materials have been designed by selecting dopants based on experimental trial and error and evaluating their performance. In contrast, it is known that an ML model using a regression algorithm can predict the performance of a selected dopant and select a suitable dopant based on the prediction result.
[0003] Machine Learning Guided Dopant Selection for Metal Oxide-Based Photoelectrochemical Water Splitting: The Case Study of Fe2O3 and CuO, Zhiliang Wang, Yuang Gu, Lingxia Zheng, Jingwei Hou, Huajun Zheng, Shijing Sun, and Lianzhou Wang, Adv. Mater. 2022, 34, 2106776
[0004] A model using regression analysis is constructed by combining explanatory variables specified by a large number of descriptors such as elemental feature quantities, for example. On the other hand, when the number of explanatory variables increases, overfitting occurs, which reduces prediction accuracy. For this reason, selection of explanatory variables greatly affects the prediction accuracy of a regression model and the interpretability of the regression model.
[0005] Therefore, there is a demand for a technique that appropriately selects explanatory variables in a regression model and improves the accuracy of prediction results obtained by the regression model.
[0006] The technology of the present disclosure may include a processor that performs the following operations: generating a first regression model learned using a plurality of explanatory variables; evaluating the first regression model to obtain standard regression coefficients for each of the explanatory variables; and generating a second regression model learned using explanatory variables, the explanatory variables having standard regression coefficients whose absolute value is greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each of the standard regression coefficients.
[0007] One aspect of this disclosure is a doping element search method, which allows for the search of elements to be doped into the hematite photoelectrode based on the performance of the hematite photoelectrode obtained using the second regression model.
[0008] Other aspects of this disclosure may include a computer implementation method performed by one or more computers, which involves generating a first regression model learned using a plurality of explanatory variables, evaluating the first regression model to obtain standard regression coefficients for each of the explanatory variables, and generating a second regression model learned using the explanatory variables, wherein the standard regression coefficients have an absolute value greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each of the standard regression coefficients.
[0009] Another aspect of this disclosure, a computer program, can cause the computer to perform operations including generating a first regression model learned using a plurality of explanatory variables, evaluating the first regression model to obtain standard regression coefficients for each of the explanatory variables, and generating a second regression model learned using explanatory variables whose standard regression coefficients have an absolute value greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each of the standard regression coefficients.
[0010] Further details will be described in the embodiments below.
[0011] Figure 1 is a diagram of the computer system configuration in the embodiment. Figure 2 is a flowchart explaining the model generation process. Figure 3 is a diagram showing an example of the relationship between each explanatory variable and the standard regression coefficient. Figure 4 is a histogram of the standard regression coefficients for each explanatory variable. Figure 5 is a flowchart explaining the overview of the doped element search method. Figure 6 is a graph showing the relationship between the hyperparameter n and the prediction accuracy. Figure 7 is a diagram showing the regression coefficients of the explanatory variables used in the second regression model. Figure 8 is a scatter plot of the predicted photocurrent density and the actual photocurrent density of the first and second regression models, where the horizontal axis of (a) shows the predicted photocurrent density obtained from the first regression model, and the horizontal axis of (b) shows the predicted photocurrent density obtained from the second regression model. Figure 9 is a diagram showing the regression coefficients of the explanatory variables used in the second regression model. Figure 10 is a scatter plot showing the relationship between the predicted logS and the actual logS of the first and second regression models. Figure 11 is a scatter plot showing the relationship between the predicted logS and the actual logS of the nonlinear regression models generated by one-stage and two-stage machine learning. Figure 12 is a conceptual diagram showing the procedure for determining the optimal data from candidate data using Bayesian optimization with a GPR model and acquisition function. Figure 13 shows the number of searches when searching for hematite materials using a two-stage regression model. Figure 14 shows the number of searches when searching for hematite materials using a two-stage regression model. Figure 15 shows the number of searches when searching for hematite materials using a one-stage regression model as a comparative example. Figure 16 shows the number of searches when searching for hematite materials using a one-stage regression model as a comparative example. Figure 17 shows the number of searches when searching for organic molecular materials using a two-stage regression model. Figure 18 shows the number of searches when searching for organic molecular materials using a two-stage regression model. Figure 19 shows the number of searches when searching for organic molecular materials using a one-stage regression model as a comparative example. Figure 20 shows the number of searches when searching for organic molecular materials using a one-stage regression model as a comparative example.
[0012] <1. Overview of the computer system, doped element search method, computer implementation method, and computer program>
[0013] (1) The computer system according to the embodiment may include a processor that performs the following operations: generate a first regression model learned using a plurality of explanatory variables; evaluate the first regression model to obtain standard regression coefficients for each of the explanatory variables; and generate a second regression model learned using explanatory variables whose standard regression coefficients have an absolute value greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each of the standard regression coefficients. In this case, explanatory variables are suitably selected in the regression model to improve the accuracy of the prediction results obtained by the regression model.
[0014] (2) The value σ may be the standard deviation of the standard regression coefficient of each of the explanatory variables.
[0015] (3) The standard regression coefficient for each explanatory variable may be the average of a plurality of standard regression coefficients obtained by performing evaluations multiple times on the first regression model.
[0016] (4) Generating the second regression model may include accepting the parameter n and repeatedly evaluating the performance of the second regression model.
[0017] (5) Generating the first regression model may include machine learning using LASSO regression.
[0018] (6) Generating the second regression model may include machine learning using LASSO regression.
[0019] (7) Generating the second regression model may include machine learning using Gaussian process regression.
[0020] (8) The first regression model and the second regression model may include generating a regression model in which the properties of the material are the explanatory variables and the performance of the material is the dependent variable.
[0021] (9) The material may be a photocatalyst used in hematite photoelectrodes.
[0022] (10) The first regression model and the second regression model may include a dependent variable which is a value representing the photoelectrochemical performance of the hematite photoelectrode, and explanatory variables which include elemental characteristics of the elements doped into the hematite photoelectrode and spectral data of the hematite photoelectrode.
[0023] (11) The doping element search method according to the embodiment can search for elements to dope the hematite photoelectrode based on the performance of the hematite photoelectrode obtained using the second regression model.
[0024] (12) The first regression model and the second regression model may include generating a regression model in which the composition or physical properties of the inorganic material are used as explanatory variables and the evaluated physical properties of the inorganic material are used as the dependent variable.
[0025] (13) The first regression model and the second regression model may include generating a regression model in which the molecular structure or physical properties of the organic material are used as explanatory variables and the evaluated physical properties of the organic material are used as the dependent variable.
[0026] (14) The operation may further include inputting descriptors corresponding to each of the multiple candidate data into the second regression model, obtaining predicted values and confidence levels of the prediction for each of the multiple candidate data from the second regression model, calculating an acquisition function based on the predicted values and confidence levels of the prediction, and selecting the optimal data from the multiple candidate data based on the acquisition function.
[0027] (15) The second regression model may be a nonlinear regression model.
[0028] (16) A computer implementation method according to an embodiment is a computer implementation method executed by one or more computers, which may include generating a first regression model learned using a plurality of explanatory variables, evaluating the first regression model to obtain standard regression coefficients for each of the explanatory variables, and generating a second regression model learned using explanatory variables, the explanatory variables having standard regression coefficients whose absolute value is greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each of the standard regression coefficients.
[0029] (17) The computer program according to the embodiment can cause the computer to perform operations including generating a first regression model learned using a plurality of explanatory variables, evaluating the first regression model to obtain standard regression coefficients for each of the explanatory variables, and generating a second regression model learned using explanatory variables, the explanatory variables having standard regression coefficients whose absolute value is greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each of the standard regression coefficients.
[0030] A computer program according to the embodiment may be configured to cause a computer to execute the computer implementation method. Furthermore, a computer program according to the embodiment may cause a computer to function as the system. The computer program may be recorded on a computer-readable, non-temporary recording medium.
[0031] <2. Examples of computer systems, doped element search methods, computer implementation methods, and computer programs>
[0032] The embodiments will be described in more detail below with reference to the drawings.
[0033] The computer system, computer implementation method, and computer program in this embodiment are techniques for suppressing overfitting and improving the predictive performance of an object by reducing the number of explanatory variables to, for example, about one-tenth or more, by performing a two-stage regression analysis on a regression model that has a vast number of explanatory variables (e.g., more than 100 types) using regression analysis.
[0034] First, the computer system in this embodiment will be described. Figure 1 is a diagram showing the configuration of the computer system 1 in this embodiment.
[0035] Computer system 1 is implemented by one or more computers. Computer system 1 includes a processor 10, a memory 20 connected to the processor 10, an input device 30, and an output device 40.
[0036] The processor 10 is, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or another type of processor. The memory 20 may include, for example, a primary storage device and a secondary storage device. The primary storage device is, for example, RAM (Random Access Memory). The secondary storage device is, for example, an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The memory 20 stores a program 21 executed by the processor 10 and descriptor data 22. The descriptor data 22 is data relating to descriptors used in the model generation process 11, which will be described later.
[0037] The processor 10 reads and executes the program 21 stored in the memory 20. The program 21 contains instructions to cause the computer to perform various processes in order to operate the computer as the computer system 1 according to the embodiment (details will be described later).
[0038] The input device 30 is a keyboard, mouse, touch panel, etc., that receives user input and transmits signals based on that user input to the processor 10. The output device 40 is a display that displays the display data provided by the processor 10.
[0039] Next, the model generation process 11 executed by the computer system 1 will be described. Figure 2 is a flowchart illustrating the model generation process 11. The following model generation process 11 will be explained using an example where LASSO (Least Absolute Shrinkage and Selection Operator) regression is used as the regression algorithm. However, the regression algorithm is not limited to LASSO regression; other regression algorithms (e.g., ridge regression, ElasticNet, principal component regression, partial least squares regression, support vector regression, Gaussian process regression) may be used as long as the objective of the invention is achieved.
[0040] In step S1, a first regression model is generated by performing a first-stage LASSO regression. As is well known in LASSO regression, a penalty (L1 regularization term) is added to the sum of the absolute values of regression coefficients, whereby the regression coefficients of unnecessary variables can be set to zero. This allows reduction of explanatory variables with small contribution degrees. The explanatory variables of the first regression model consist of pre-prepared descriptors, and include, for example, one hundred or more types of explanatory variables. The descriptors relate, for example, to feature values of an element doped into a photoelectrode or analysis data of the photoelectrode (material properties, physical properties). The objective variable is a photocurrent density (material performance) that can quantitatively evaluate the photoelectrochemical performance (PEC performance) of the photoelectrode. The first regression model is subjected to machine learning using these explanatory variables, and a trained first regression model is generated (constructed).
[0041] In step S2, the trained first regression model is evaluated, and regression coefficients are acquired. This is performed to generate a regression model using explanatory variables with large contribution degrees from among a plurality of explanatory variables, that is, explanatory variables having large absolute values of regression coefficients. For example, the evaluation is performed by repeating 10-fold nested cross-validation (Nested CV) 10 times on the first regression model. Since the value of a regression coefficient changes depending on the data division method, 100 different regression coefficients are obtained for each explanatory variable.
[0042] In step S3, an average value of standard regression coefficients of 100 regression coefficients for each acquired explanatory variable is obtained. Further, a standard deviation σ is obtained from the obtained average value of standard regression coefficients (hereinafter, the average value of "standard regression coefficients" is simply referred to as "standard regression coefficient").
[0043] Here, FIG. 3 is a diagram showing an example of the relationship between each explanatory variable and the standard regression coefficient. In the graph of FIG. 3, the horizontal axis represents the explanatory variable and the vertical axis represents the standard regression coefficient. FIG. 4 is a histogram of the standard regression coefficients of each explanatory variable. In the graph of FIG. 4, the horizontal axis represents the standard regression coefficient and the vertical axis represents the number of explanatory variables. The absolute value of the standard regression coefficient can indicate the contribution of the explanatory variable to the objective variable. From FIG. 3, it can be seen that most of the standard regression coefficients of the explanatory variables are zero or close to zero. From the obtained standard regression coefficients, a distribution as shown in FIG. 3 and FIG. 4 is obtained, and the standard deviation σ of the standard regression coefficients is obtained.
[0044] In step S4, a second regression model is generated by performing the second-stage LASSO regression. In the second-stage LASSO regression, machine learning is performed using, among the explanatory variables used in the first regression model, explanatory variables whose absolute value of the standard regression coefficient is equal to or greater than the threshold t=±nσ (see formula (1) in FIG. 4).
[0045] Accordingly, in addition to explanatory variables with zero contribution (standard regression coefficient of 0) that can be originally rejected in the first-stage LASSO regression, explanatory variables whose contribution is close to zero can also be rejected. Therefore, in the second regression model, explanatory variables that can be regarded as substantially not contributing to the objective variable are rejected, and the number of explanatory variables is reduced, whereby both suppression of overfitting and interpretability of the regression model can be achieved.
[0046] Here, the hyperparameter n determines the number of explanatory variables that directly affect prediction accuracy, and it is difficult to determine it automatically. Therefore, by setting an appropriate n while evaluating the second regression model with the above-mentioned Nested CV, it becomes possible to select explanatory variables (descriptors) that provide a regression model with high generalization performance. Note that performance evaluation is performed using, for example, the coefficient of determination R 2 , mean absolute error (MAE) and root mean square error (RMSE) as performance indicators.
[0047] Next, a procedure for searching for a doped element including the above-described model generation process 11 will be described together with specific examples.
[0048] The computer system 1, dope element search method, computer implementation method, and computer program in this embodiment can be used as a material determination (search) method based on predictive performance obtained from a regression model. For example, a regression model is constructed with elemental characteristics of multiple elements doped into the hematite photoelectrode and spectral data of the hematite photoelectrode as explanatory variables, and photocurrent density that can quantitatively evaluate PEC performance as the objective variable. This regression model is applicable to predict the performance of the hematite photoelectrode and to search for multiple elements to dope into the hematite photoelectrode. In the following description, the computer system 1, dope element search method, computer implementation method, and computer program in this embodiment will be applied to the performance prediction of the hematite photoelectrode described above.
[0049] Figure 5 is a flowchart illustrating the outline of the doped element search method.
[0050] First, hematite samples were synthesized and photoelectrodes were prepared (step S11). Specifically, hematite photoelectrodes doped with multiple elements were synthesized, and three of each of 97 different samples containing two or three (or more) elements were prepared. The photoelectrodes were fabricated using the spin coating method.
[0051] Next, analytical data was collected for the fabricated photoelectrodes (step S12). Specifically, the actual composition of hematite particles containing the dopant was measured for all fabricated photoelectrode samples using an X-ray fluorescence (XRF) spectrometer. In addition, X-ray diffraction (XRD) patterns, UV-vis transmittance (UV-vis) spectra, Raman spectra, and current-voltage (I-V) curves were measured for the photoelectrode samples.
[0052] Based on the collected analytical data, PEC performance data is collected (step S13). The PEC performance data is quantified, for example, from the photocurrent I (mA) in the 1.6V vs. Reversible Hydrogen Electrode (RHE) curve extracted from the Linear Sweep Voltammetry (LSV) curve.
[0053] Next, as a preprocessing step for generating the regression model, the explanatory variables are prepared (step S14). The regression model has, for example, elemental features and Raman spectra as explanatory variables. The regression model has the photocurrent density obtained from the photocurrent as the dependent variable.
[0054] Specifically, descriptors for elemental characteristics used as explanatory variables were generated using the Python® library XenonPy, based on the empirical formula determined by XRF. Elemental descriptors for valence, dopant concentration, ionic radius, and M-O bond formation enthalpy (per mol-O and per mol-M) were also generated as original descriptors. Furthermore, necessary processing was performed on the analysis data, including XRD patterns, UV-vis spectra, and Raman spectra, to generate descriptors. A total of 244 types of descriptors were generated.
[0055] For these 244 types of descriptors, descriptors with more than 80% identical values and all elements (or their absolute values) are 10 -7 Descriptors that were less than a certain value were removed. After data preprocessing, the number of explanatory variables was reduced from 244 to 177. In addition, all descriptors were standardized to become explanatory variables.
[0056] Next, a regression model is constructed using the explanatory variables obtained in the preprocessing step (step S15). The regression model is constructed by performing a two-stage regression analysis using the procedure described in the model generation process 11 in Figure 2. The constructed regression model is evaluated each time, for example, using nested cross-validation as described in step S3 in Figure 2, and finally a regression model with high predictive accuracy is generated.
[0057] Here, Figure 6 is a graph showing the relationship between the hyperparameter n and prediction accuracy. Prediction accuracy is the coefficient of determination R 2 The model was evaluated using RMSE and MAE. In the example in Figure 6, underfitting occurs when the parameter n is large (greater than 1.0), i.e., when the absolute value of the threshold t is large and there are few explanatory variables. When the parameter n is small (less than 1.0), i.e., when the absolute value of the threshold t is small and there are many explanatory variables, overfitting occurs, resulting in a decrease in generalization performance. From this, it can be seen that manually determining an appropriate parameter n is important.
[0058] As a result of determining an appropriate parameter n, the number of explanatory variables decreased from 177 to 16. Figure 7 shows the regression coefficients of the explanatory variables used in the second regression model. The bar graph shows the mean value, and the error bars show the standard deviation of the regression coefficient of each explanatory variable at 100. From Figure 7, it can be seen that the original descriptors, such as valence, contribute more to the prediction of photoelectrode performance than the descriptors generated using XenonPy. These original descriptors are thought to be closely related to changes in material properties.
[0059] Figure 8 shows scatter plots of predicted and actual photocurrent densities for the first and second regression models, respectively. The horizontal axis in (a) represents the predicted photocurrent density obtained from the first regression model, and the horizontal axis in (b) represents the predicted photocurrent density obtained from the second regression model. Comparing the scatter plots from the first and second regression models, it can be seen that the second regression model shows less dispersion and improved prediction accuracy. Also, the coefficient of determination R for the first regression model in Figure 8(a) is shown. 2 The RMSE and MAE are 0.594, 0.156, and 0.118, respectively, and the coefficient of determination of the second regression model R in Figure 8(b) is shown. 2 The RMSE and MAE were 0.760, 0.120, and 0.089, respectively. In other words, the coefficient of determination for the second regression model was R 2The coefficient of photocurrent density increased, while RMSE and MAE decreased. This suggests that a two-stage regression approach, which involves variable selection using LASSO regression in the first stage followed by LASSO regression in the second stage, improves the predictive performance of the target variable, photocurrent density.
[0060] In step S16, data analysis based on the regression model is performed. For example, the photocurrent can be predicted for a vast number of hypothetical samples using the regression model (second regression model), and a suitable composition can be determined.
[0061] Furthermore, as will be discussed later, the explanatory variables selected in a regression model can also be used as a method for selecting variables for a nonlinear model.
[0062] This can reduce the enormous amount of time previously required to find the optimal dopant combination using material discovery methods based on researchers' experience, which was a major challenge.
[0063] The above describes an example of material discovery using a regression model constructed by a two-stage regression analysis, focusing on the elemental composition doped into hematite photoelectrodes. However, the configuration and application method of the regression model can be appropriately modified depending on the type of material and the target of the discovery. Below, other embodiments using the computer system, computer implementation method, and computer program constructed by the computer program of this disclosure will be described.
[0064] In the embodiments described above, an example was explained in which the computer system, computer implementation method, and computer program of this disclosure were applied to an inorganic material, such as a hematite sample, to predict the performance of a hematite photoelectrode. However, the computer system, computer implementation method, and computer program of this disclosure are not limited to inorganic materials but may also be applied to organic materials.
[0065] For example, the first and second regression models can be constructed by machine learning, inputting molecular descriptors generated based on molecular structure, for example, as explanatory variables, and logS, the common logarithm of the solubility of organic molecules in water, as the dependent variable. Molecular descriptors are generated based on SMILES (Simplified Molecular Input Line Entry System), which represents molecular structure information assigned to each organic molecule, and are descriptors that numerically represent the physical properties and structural characteristics of molecules. These molecular descriptors include multiple types of descriptors that reflect the physical properties of molecules, such as features related to the hydrophobicity of molecules, molecular size, degree of branching, aromatic ring structure, functional group composition, and atomic surface area.
[0066] Figure 9 shows the regression coefficients of the explanatory variables used in the second regression model. The bar graph shows the standard regression coefficient for each explanatory variable. The second regression model in Figure 9, like in Figure 7, is a model generated by performing machine learning using LASSO regression in the first and second stages. There were 209 explanatory variables (descriptors) used as input for the first stage of LASSO regression.
[0067] Figure 9 shows that by applying the second stage of LASSO regression, 54 explanatory variables effective for predicting logS have been selected from the 209 descriptors used initially.
[0068] Figure 10 is a scatter plot showing the relationship between the predicted logS and the actual logS for the first and second regression models. Figure 10(a) shows the predicted values from the regression model (first regression model) generated by a single-stage machine learning process using LASSO regression. Figure 10(b) shows the predicted values from the regression model (second regression model) generated by a two-stage machine learning process using LASSO regression.
[0069] Comparing the first and second regression models, the second regression model shows less dispersion and improved prediction accuracy. Furthermore, in the second regression model, the coefficient of determination R... 2The coefficient of ion increased, while RMSE and MAE decreased. This suggests that even in regression models used to predict the properties of organic materials, using a second regression model improves prediction performance.
[0070] Furthermore, in the above-described embodiment, a two-stage model generation was performed using a linear regression model. However, if the number of explanatory variables can be reduced by excluding explanatory variables with small contributions in the first stage, and an optimal regression model for performance prediction using the selected explanatory variables can be generated in the second stage, then a nonlinear regression model may be generated in both the first and second stages. Alternatively, a linear regression model may be generated in the first stage and a nonlinear regression model in the second stage.
[0071] For example, in the first stage, LASSO regression, which is effective in reducing explanatory variables, may be performed to generate a linear first regression model. In the second stage, Gaussian process regression (GPR), which is effective for subsequent material exploration (details will be described later), may be performed to generate a nonlinear second regression model.
[0072] Here, Figure 11 is a scatter plot showing the relationship between the predicted logS and the actual logS of nonlinear regression models generated by one-stage and two-stage machine learning. Figure 11(a) shows the predicted values from a regression model generated by one-stage machine learning using GPR (one-stage regression model). That is, it shows the predicted values from a nonlinear regression model that does not perform the reduction of explanatory variables by the first stage of machine learning. Figure 11(b) shows the predicted values from a regression model generated by two-stage machine learning using LASSO regression in the first stage and GPR in the second stage (two-stage regression model).
[0073] Comparing the one-stage regression model with the two-stage regression model, it can be seen that the two-stage regression model has less dispersion and improved prediction accuracy. This suggests that prediction accuracy is improved by generating a nonlinear model as a two-stage regression model using the first regression model generated using LASSO regression in the first stage. In addition, values that behave like outliers in the one-stage regression model, specifically indicated by dotted circles in Figure 11(a), show reduced error in the two-stage regression model.
[0074] Furthermore, comparing the two-stage regression model in Figure 11(b) with the second regression model generated using LASSO regression in the first and second stages, as shown in Figure 10(b), it can be said that the second regression model generated using GPR in the second stage improves prediction accuracy.
[0075] From the above, it can be said that the computer system, computer implementation method, and computer program according to the present invention can be applied not only to the generation of linear regression models generated by two-stage machine learning, but also to the generation of nonlinear regression models.
[0076] Furthermore, the second regression model generated in this way is not limited to linear or nonlinear models; it can be any type of regression model. By using a second regression model that can output the predicted value along with the reliability of the prediction for each candidate material, it becomes possible to select the optimal data from among the candidate data. Specifically, by using the output of the second regression model to evaluate multiple candidate data associated with the explanatory variable, it is possible to select the data that can maximize or minimize the dependent variable from among the candidate data.
[0077] One method for evaluating and selecting candidate data is to use an acquisition function that takes the predicted values and confidence levels of the predictions, which are the outputs of the second regression model, as inputs. Since the value of the acquisition function serves as an index for evaluating candidate data characterized by the explanatory variables, it can be used to select data from the candidate data that maximizes / minimizes the dependent variable.
[0078] For example, if candidate data is identified by the composition (in the case of inorganic materials) or molecular structure (in the case of organic materials) of candidate materials in material discovery (hereinafter, composition and molecular structure may be referred to as "composition, etc."), the explanatory variables (descriptors) obtained from this composition, etc. are input into the second regression model to obtain the predicted value and the confidence level of the prediction, which are the outputs of the second regression model. An acquisition function is calculated using this predicted value and confidence level of the prediction as input, and based on the acquisition function, candidate data (optimal data) that can maximize the dependent variable of the second regression model can be selected from among multiple candidate data characterized by the explanatory variables.
[0079] For example, when using a GPR model as the second regression model, the second regression model outputs predicted values and prediction confidence for the candidate data. In this way, the acquisition function can be calculated using the output of the second regression model, which can simultaneously output predicted values and prediction confidence. Bayesian optimization can be used as an example of an optimization method to select candidate data based on the calculated acquisition function. Here, Figure 12 is a conceptual diagram showing the procedure for determining the optimal data from among the candidate data using Bayesian optimization with a GPR model and acquisition function.
[0080] First, candidate data representing candidate materials is prepared (step S21). The candidate data may be, for example, data that uses the composition of the candidate material as a descriptor, or data in which the physical properties of the candidate material are associated with the composition as a descriptor. A predetermined number of candidate data are selected from the candidate data as initial data (step S22). The initial data is used to generate the model in the model generation process 11 shown in Figure 2. The method of selecting the initial data is not particularly limited, and various methods can be used, such as random selection, use of existing evaluated materials, or selection based on design experimental methods.
[0081] Next, the model generation process 11 is executed using the initial data to generate a second regression model (e.g., a GPR model) (step S23). Descriptors of each candidate data other than the initial data (unevaluated candidate data) are input as explanatory variables of the second regression model. The second regression model outputs predicted values and confidence levels of the predictions for the target variable (e.g., photocurrent density of hematite photoelectrode, solubility of organic material) as evaluated physical properties of the candidate materials.
[0082] Next, the output of the second regression model is input to the acquisition function, and the acquisition function is calculated, which outputs an evaluation value for each candidate data (step S24). The acquisition function may be, for example, Expected Improvement (EI) or Upper Confidence Bound (UCB). The output of the acquisition function can be a value that evaluates the target variable of each candidate data output from the second regression model. That is, if the evaluation value of the acquisition function is high, it indicates that the candidate data associated with the descriptor input to the second regression model is data of candidate material with a high priority to be evaluated.
[0083] Once the acquisition function is calculated for all candidate data, the candidate data with the highest evaluation is obtained as the optimal data (step S25). The optimal data obtained here can be selected from the candidate data as the data for the candidate material that has the highest predicted value of the evaluated physical property (the dependent variable of the second regression model), and can be actually synthesized, prototyped, and evaluated.
[0084] Furthermore, if the material properties of the candidate data selected as optimal data are actually evaluated, the corresponding candidate data is updated using the evaluation values obtained from the evaluation. The updated candidate data can be used as evaluated data for retraining in the model generation process 11 (step S26).
[0085] To verify the effectiveness of this material discovery method, simulations were performed on multi-element materials and organic molecular materials doped into hematite photoelectrodes. In the simulations, candidate data was created using a dataset with known evaluation properties (target variables), and the effectiveness of the material discovery was evaluated by the number of searches. The search was a process that repeated steps S21 to S26 in Figure 12. The number of searches was defined as the number of searches until candidate data with pre-set evaluation properties was selected as the optimal data.
[0086] Figures 13 and 14 show the number of searches when searching for hematite materials using a two-stage regression model. Figures 15 and 16 show the number of searches when searching for hematite materials using a one-stage regression model, as a comparative example. Figures 17 and 18 show the number of searches when searching for organic molecular materials using a two-stage regression model. Figures 19 and 20 show the number of searches when searching for organic molecular materials using a one-stage regression model, as a comparative example.
[0087] A two-stage regression model is a regression model generated by two-stage machine learning, using LASSO regression in the first stage and GPR in the second stage. A one-stage regression model is a regression model generated by one-stage machine learning using GPR.
[0088] Candidate data consists of data for which the evaluation property values as the objective variable are known. Initial data was randomly selected from candidate data having physical property values that are 1 / 5 or less of the maximum value of all candidate data. Initial data can be determined using methods such as the D-optimal criterion or LHS (Latin Hypercube Sampling). 1000 sets of initial data were prepared, and 1000 material searches were performed.
[0089] The acquisition functions used were Expected Improvement (EI) and Upper Confidence Bound (UCB). Two types of weighting coefficients (high parameters) were used for UCB (UCB_1.0 and UCB_5.0). The weighting coefficients are values used to adjust how much confidence in the predictions is considered when applying them to the predicted values obtained from the regression model.
[0090] In the search for hematite materials, elemental features (XenonPy) were used as explanatory variables for the regression models in Figures 13 and 15, as descriptors for the physical properties of the hematite material. For the regression models in Figures 14 and 16, elemental ratios were used as explanatory variables for the composition of the hematite material. The number of search iterations was defined as the number of iterations required until candidate data updated by 400% relative to the maximum value of the photocurrent density (dependent variable) in the initial data was selected as the optimal data.
[0091] Figures 13 to 16 show that in both cases where physical properties and composition were used as explanatory variables, the number of iterations was significantly reduced when a two-stage regression model was used. The fact that there was no difference between the cases where physical properties were used and the cases where composition was used as explanatory variables is presumed to be because the photocurrent density of the hematite photoelectrode is strongly influenced by the composition.
[0092] In the search for organic molecular materials, for the regression models in Figures 17 and 19, descriptors relating to the molecular properties of the organic molecular materials (RDKit Descriptors) were used as explanatory variables. For the regression models in Figures 18 and 20, descriptors relating to the molecular structure of the organic molecular materials (MorganFP) were used as explanatory variables. The number of searches was defined as the number of iterations required until the candidate data with the maximum known solubility (dependent variable) was selected as the optimal data.
[0093] Figures 17 to 20 show that in all cases where molecular properties and molecular structure were used as explanatory variables, the number of searches was clearly reduced when a two-stage regression model was used.
[0094] Furthermore, from the comparison of Figures 13 and 14, and Figures 17 and 18, it can be seen that the descriptor can be either the composition of the material or a descriptor relating to its physical properties. When using composition, the material composition is directly obtained as an optimization result. When using physical properties, the material is ultimately identified by selecting candidate materials associated with the descriptor.
[0095] In other words, when candidate material searches use composition and other parameters directly as explanatory variables, the interpretability of the search results may decrease, making it difficult to understand the factors contributing to material performance. In contrast, by associating descriptors representing material properties with the composition and other parameters of candidate materials, it becomes possible to generate regression models using property descriptors. This configuration allows for searches while analyzing the relationship between material performance and each property. Furthermore, by using this search process, evaluation results obtained using property descriptors are output as candidate materials associated with composition and other parameters, making it possible to directly apply the search results to the next material evaluation process.
[0096] The present invention is not limited to the above embodiments, and various modifications are possible.
[0097] 1: Computer system 10: Processor 11: Model generation process 20: Memory 21: Program 22: Descriptor data 30: Input device 40: Output device
Claims
1. A computer system comprising a processor that performs the following operations: generating a first regression model learned using multiple explanatory variables; evaluating the first regression model to obtain standard regression coefficients for each explanatory variable; and generating a second regression model learned using explanatory variables, the explanatory variables having standard regression coefficients whose absolute value is greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each standard regression coefficient.
2. The computer system according to claim 1, wherein the value σ is the standard deviation of the standard regression coefficients of each of the explanatory variables.
3. The computer system according to claim 2, wherein the standard regression coefficient for each of the explanatory variables is the mean of a plurality of standard regression coefficients obtained by performing evaluations multiple times on the first regression model.
4. The computer system according to claim 1, wherein generating the second regression model includes receiving the parameter n and repeatedly evaluating the performance of the second regression model.
5. The computer system according to claim 1, wherein generating the first regression model includes machine learning using LASSO regression.
6. The computer system according to claim 5, wherein generating the second regression model includes machine learning using LASSO regression.
7. The computer system according to claim 5, wherein generating the second regression model includes machine learning by Gaussian process regression.
8. The computer system according to claim 1, wherein the first regression model and the second regression model generate regression models in which the properties of the material are the explanatory variables and the performance of the material is the dependent variable.
9. The computer system according to claim 8, wherein the material is a photocatalyst used in a hematite photoelectrode.
10. The computer system according to claim 9, wherein the first regression model and the second regression model include the objective variable, which is a value representing the photoelectrochemical performance of the hematite photoelectrode, and the explanatory variables, which include elemental characteristics of the elements doped into the hematite photoelectrode and spectral data of the hematite photoelectrode.
11. A method for searching for elements to dope a hematite photoelectrode based on the performance of the hematite photoelectrode obtained using the second regression model described in claim 10.
12. The computer system according to claim 1, wherein the first regression model and the second regression model generate regression models in which the composition or physical properties of the inorganic material are used as explanatory variables and the evaluated physical property values of the inorganic material are used as the objective variable.
13. The computer system according to claim 1, wherein the first regression model and the second regression model generate regression models in which the molecular structure or physical properties of the organic material are used as explanatory variables and the evaluated physical properties of the organic material are used as dependent variables.
14. The computer system according to claim 1, further comprising: inputting descriptors corresponding to each of a plurality of candidate data into the second regression model; obtaining predicted values and confidence levels of the prediction for each of the plurality of candidate data from the second regression model; calculating an acquisition function based on the predicted values and confidence levels of the prediction; and selecting the optimal data from the plurality of candidate data based on the acquisition function.
15. The computer system according to claim 14, wherein the second regression model is a nonlinear regression model.
16. A computer implementation method executed by one or more computers, comprising: generating a first regression model learned using a plurality of explanatory variables; evaluating the first regression model to obtain standard regression coefficients for each of the explanatory variables; and generating a second regression model learned using explanatory variables, among the explanatory variables, having standard regression coefficients whose absolute value is greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each of the standard regression coefficients.
17. A computer program that causes a computer to perform the following operations: generate a first regression model learned using multiple explanatory variables; evaluate the first regression model to obtain standard regression coefficients for each explanatory variable; and generate a second regression model learned using explanatory variables whose standard regression coefficients have an absolute value greater than a threshold t = ±nσ obtained from a parameter n consisting of a positive number and a value σ that depends on each standard regression coefficient.