Prediction device, learning device, and method and program therefor
By employing a multi-fidelity machine learning model that distinguishes between datasets of varying reliability, the method addresses the accuracy issues in material property prediction, achieving precise and cost-effective predictions using a combination of experimental, simulation, and literature data.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-04-02
AI Technical Summary
Existing prediction methods for material physical properties using machine learning do not adequately distinguish between experimental and simulation data, leading to insufficient prediction accuracy due to the lower reliability of simulation data compared to experimental data.
A machine learning model is developed that considers the fidelity of different datasets, including experimental, simulation, and literature data, to accurately predict material properties by learning the relationship between descriptors and physical properties, using a multi-fidelity approach such as delta-learning or Gaussian process regression.
This method enables accurate and efficient prediction of material properties, even with limited reliable data, by effectively utilizing datasets of varying fidelity, thereby improving prediction accuracy and reducing the need for extensive experimental data collection.
Smart Images

Figure JP2025033391_02042026_PF_FP_ABST
Abstract
Description
Prediction device, learning device, their methods, and programs
[0001] The present disclosure relates to a prediction device, a learning device, their methods, and programs.
[0002] Techniques for predicting the physical properties of substances based on machine learning are known. For example, Patent Document 1 discloses a data processing method for creating a prediction module that predicts a value related to a feature amount by inputting values of a plurality of explanatory variables, from an original dataset including experimental data and simulation data.
[0003] Japanese Patent Application Laid-Open No. 2021-22276
[0004] However, in the prior art, there is room for improving the prediction accuracy. For example, although Patent Document 1 corrects the calculated value in the simulation data based on the relationship between the experimental data and the simulation data, it does not distinguish between the experimental data and the simulation data when creating the prediction module.
[0005] One aspect of the present disclosure aims to accurately predict the physical properties of substances.
[0006] The present disclosure has the following configuration.
[0007] <1> A descriptor calculation unit configured to calculate a descriptor of the target substance based on the identification information of the target substance, and a physical property prediction unit configured to predict the physical properties of the target substance by inputting the descriptor of the target substance into a machine learning model that has learned the relationship between the descriptor and a predetermined physical property using a plurality of datasets with different set faithfulness degrees.
[0008] <2> The prediction device according to <1> above, wherein the plurality of datasets include at least two of experimental data, simulation data, or literature data.
[0009] <3> The prediction device according to <2> above, wherein the faithfulness degree is set based on the acquisition method of the dataset.
[0010] <4> The prediction device according to <3> above, wherein the experimental data is set to a higher fidelity than the simulation data and the literature data, and the simulation data is set to a higher fidelity than the literature data.
[0011] <5> The prediction device according to <2> above, wherein the fidelity is set based on the similarity between the datasets.
[0012] <6> The prediction device described in <5> above, wherein the higher the correlation coefficient between the multiple datasets and a reference dataset, the higher the fidelity set.
[0013] <7> The prediction device according to <5> above, wherein the fidelity is set to be higher the smaller the difference between the multiple datasets and a reference dataset.
[0014] <8> The prediction device according to any one of <1> to <7> above, wherein the machine learning model includes a first model trained using a dataset with a first fidelity set, and a second model trained on the difference between physical property values included in a dataset with a second fidelity set and predicted physical property values by the first model.
[0015] <9> The prediction device according to any one of <1> to <7> above, wherein the machine learning model is a Gaussian process regression model trained using a dataset obtained by combining a dataset with a first fidelity setting and a dataset with a second fidelity setting.
[0016] <10> The prediction device according to any one of <1> to <7> above, wherein the machine learning model includes a first model trained using a dataset with a first fidelity set, and a second model trained using a dataset with a second fidelity set, to which predicted values of physical properties by the first model are added.
[0017] <11> A learning device comprising: a fidelity setting unit configured to set different fidelity levels for each of a plurality of datasets, each containing descriptors of a target substance and values indicating predetermined physical properties of the target substance; and a model generation unit configured to generate a machine learning model by learning the relationship between the descriptors and the physical properties using the plurality of datasets on which the fidelity levels have been set.
[0018] <12> A prediction method that includes the steps of: a computer calculating a descriptor of a target substance based on identification information of the target substance; and a machine learning model that has learned the relationship between the descriptor and predetermined physical properties using a plurality of datasets with different fidelity settings, by inputting the descriptor of the target substance into the machine learning model.
[0019] <13> A learning method comprising: a procedure for a computer to set different fidelity levels for each of a plurality of datasets, each containing descriptors of a target substance and values indicating predetermined physical properties of the target substance; and a procedure for generating a machine learning model by learning the relationship between the descriptors and the physical properties using the plurality of datasets on which the fidelity levels have been set.
[0020] <14> A program for causing a computer to perform the following steps: a procedure for calculating a descriptor of a target substance based on identification information of the target substance; and a procedure for predicting the physical properties of a target substance by inputting the descriptor of the target substance into a machine learning model that has learned the relationship between the descriptor and predetermined physical properties using multiple datasets with different fidelity settings.
[0021] <15> A program for causing a computer to perform the following steps: setting different fidelity levels for each of a plurality of datasets, each containing descriptors of a target substance and values indicating predetermined physical properties of the target substance; and generating a machine learning model by learning the relationship between the descriptors and the physical properties using the plurality of datasets with the fidelity levels set.
[0022] According to one aspect of this disclosure, the physical properties of a material can be predicted with high accuracy.
[0023] Figure 1 is a block diagram showing an example of the overall configuration of a physical property prediction system. Figure 2 is a block diagram showing an example of a computer. Figure 3 is a block diagram showing an example of the functional configuration of a physical property prediction system. Figure 4 is a diagram showing an example of a differential learning model. Figure 5 is a diagram showing an example of training data for a differential learning model. Figure 6 is a flowchart showing an example of the learning process according to the first embodiment. Figure 7 is a flowchart showing an example of the model generation process according to the first embodiment. Figure 8 is a flowchart showing an example of the prediction process according to the first embodiment. Figure 9 is a flowchart showing an example of the physical property prediction process according to the first embodiment. Figure 10 is a diagram showing an example of a Gaussian process regression model. Figure 11 is a diagram showing an example of training data for a Gaussian process regression model. Figure 12 is a diagram showing an example of a low-fidelity feature model. Figure 13 is a diagram showing an example of training data for a low-fidelity feature model. Figure 14 is a flowchart showing an example of the model generation process according to the third embodiment. Figure 15 is a flowchart showing an example of the physical property prediction process according to the third embodiment.
[0024] Hereinafter, embodiments of this disclosure will be described with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0025] [First Embodiment] The first embodiment of this disclosure is an example of an information processing system for predicting the physical properties of a substance. Hereinafter, the information processing system according to this embodiment will be referred to as the "physical property prediction system." The substance to be predicted will be referred to as the "substance to be predicted."
[0026] In this embodiment, the material to be predicted can be any material. Furthermore, the physical properties to be predicted can be any physical properties relating to the material to be predicted. For example, the physical properties to be predicted may include properties relating to the electrical properties of the material. For example, the physical properties relating to electrical properties may include the dielectric constant.
[0027] Traditionally, techniques for predicting the physical properties of materials using methods such as machine learning have been established. In machine learning, it is known that prediction accuracy improves as the amount of training data increases. On the other hand, in material property prediction, experiments are often required to obtain the correct physical property values to be included in the training data. Therefore, collecting a large amount of training data requires considerable cost. To obtain the correct physical property values at a low cost, physical property values calculated by simulations such as quantum chemical calculations or molecular dynamics calculations are sometimes added to the training data.
[0028] However, material properties calculated through simulation are less reliable than those measured experimentally. If a machine learning model is generated based on training data containing a mixture of experimentally measured and simulated material properties, the prediction accuracy may not improve sufficiently.
[0029] This embodiment aims to accurately predict the physical properties of a material. To this end, this embodiment predicts the physical properties of a target material by inputting the descriptor of the target material into a machine learning model that has learned the relationship between descriptors and physical properties using multiple datasets with different fidelity settings.
[0030] In one respect, this embodiment uses a machine learning model that has learned the relationship between descriptors and physical properties while considering the fidelity of the dataset, thus enabling accurate prediction of the physical properties of materials. In another respect, this embodiment enables efficient prediction of the physical properties of materials because, even if there is only a small amount of highly reliable data such as experimental data, the model can be appropriately trained by adding less reliable data such as simulation data or literature data.
[0031] <Overall Configuration> The overall configuration of the physical property prediction system according to this embodiment will be described with reference to Figure 1. Figure 1 is a block diagram showing an example of the overall configuration of the physical property prediction system.
[0032] As shown in Figure 1, the physical property prediction system 1000 includes a learning device 10, a prediction device 20, and a terminal device 30. The learning device 10, the prediction device 20, and the terminal device 30 are connected via a communication network N to enable data communication. The communication network N may be, for example, a LAN (Local Area Network), a VPN (Virtual Private Network), or the Internet.
[0033] The learning device 10 is an example of an information processing device such as a personal computer, workstation, or server that learns a prediction model. The prediction model is an example of a machine learning model that predicts predetermined physical properties. The predetermined physical properties may include, for example, dielectric constant. The predetermined physical properties are not limited to these, and any physical properties may be used as the target of prediction.
[0034] The predictive model may be a machine learning model that has learned the relationship between the descriptor of the substance to be trained and the predetermined physical properties of that substance. Hereafter, the substance to be trained will also be called the "training target substance." The training target substance can be any substance, just like the target substance.
[0035] The learning device 10 may train a predictive model using multiple datasets. The learning device 10 may set different fidelity levels for each of the multiple datasets. Each dataset may include identification information of the target substance and physical property values indicating the physical properties of the substance. The learning device 10 may learn the relationship between descriptors and physical properties while taking into account the fidelity levels set for each dataset.
[0036] The prediction model may be a multi-fidelity model. A multi-fidelity model may be constructed using methods such as delta-learning, Gaussian process regression (co-Kriging), and low-fidelity features, as an example. In this embodiment, an example of constructing a prediction model using a multi-fidelity model constructed by delta-learning will be described.
[0037] The learning device 10 may calculate a descriptor of a substance based on the identification information of the substance to be learned. The learning device 10 may generate training data for training a prediction model based on the descriptor and physical properties of the substance to be learned. The learning device 10 may generate a prediction model by learning the relationship between the descriptor and physical properties based on the training data. The learning device 10 may output the trained prediction model to the prediction device 20.
[0038] The prediction device 20 is an example of an information processing device such as a personal computer, workstation, or server that predicts the physical properties of a target substance. The prediction device 20 may predict the physical properties of the target substance based on a trained prediction model. The trained prediction model may be generated by the learning device 10. The prediction device 20 may calculate a descriptor of the target substance based on the identification information of the target substance. The prediction device 20 may predict the physical properties of the target substance by inputting the descriptor of the target substance into the trained prediction model.
[0039] Terminal device 30 is an example of an information processing terminal such as a personal computer, smartphone, or tablet terminal operated by a user of the physical property prediction system 1000. Terminal device 30 may accept input of identification information of the substance to be predicted and transmit it to the prediction device 20. Terminal device 30 may receive prediction results from the prediction device 20 and present them to the user.
[0040] The overall configuration of the physical property prediction system 1000 shown in Figure 1 is just one example, and various system configurations are possible depending on the application and purpose. For example, one or more of the learning device 10, prediction device 20, and terminal device 30 may be included in multiple units of the physical property prediction system 1000. For example, the learning device 10 or prediction device 20 may be implemented using multiple computers, or as a cloud computing service. For example, the physical property prediction system 1000 may be implemented using a standalone computer. The classification of devices such as the learning device 10, prediction device 20, and terminal device 30 shown in Figure 1 is just one example.
[0041] <Hardware Configuration> The hardware configuration of the physical property prediction system 1000 will be explained with reference to Figure 2. The learning device 10, prediction device 20, and terminal device 30 are implemented, for example, by a computer. Figure 2 is a block diagram showing an example of the computer's hardware configuration.
[0042] As shown in Figure 2, the computer 500 includes a CPU (Central Processing Unit) 501, ROM (Read Only Memory) 502, RAM (Random Access Memory) 503, HDD (Hard Disk Drive) 504, input device 505, display device 506, communication interface 507, and external interface 508. The CPU 501, ROM 502, and RAM 503 form what is known as a computer. Each piece of hardware in the computer 500 is interconnected via a bus line 509. The input device 505 and display device 506 may also be used by connecting them to the external interface 508.
[0043] The CPU 501 is a computing device that reads programs and data from a storage device such as ROM 502 or HDD 504 onto RAM 503 and executes processing to realize the overall control and functions of the computer 500. The computer 500 may have a GPU (Graphics Processing Unit) in addition to or instead of the CPU 501.
[0044] ROM 502 is an example of a non-volatile semiconductor memory (storage device) that can retain programs and data even when the power is turned off. ROM 502 functions as the main memory, storing various programs and data necessary for the CPU 501 to execute the various programs installed on HDD 504. Specifically, ROM 502 stores boot programs such as BIOS (Basic Input Output System) and EFI (Extensible Firmware Interface) that are executed when the computer 500 starts up, as well as data such as OS (Operating System) settings and network settings.
[0045] The RAM 503 is an example of a volatile semiconductor memory (storage device) in which programs and data are erased when the power is turned off. The RAM 503 is, for example, a DRAM (Dynamic Random Access Memory) or a SRAM (Static Random Access Memory). The RAM 503 provides a work area where various programs installed in the HDD 504 are expanded when executed by the CPU 501.
[0046] The HDD 504 is an example of a non-volatile storage device that stores programs and data. Programs and data stored in the HDD 504 include an OS, which is basic software for controlling the entire computer 500, and applications that provide various functions on the OS. Note that the computer 500 may use a storage device (e.g., SSD: Solid State Drive) that uses a flash memory as a storage medium instead of the HDD 504.
[0047] The input device 505 is a touch panel used by the user to input various signals, operation keys and buttons, a keyboard and a mouse, a microphone for inputting audio data such as voice, etc.
[0048] The display device 506 is composed of a display such as a liquid crystal or an organic EL (Electro-Luminescence) that displays a screen, a speaker that outputs audio data such as voice, etc.
[0049] The communication I / F 507 is an interface for connecting to a communication network and enabling the computer 500 to perform data communication.
[0050] The external I / F 508 is an interface with an external device. Examples of external devices include a drive device 510.
[0051] The drive device 510 is a device for setting the recording medium 511. The recording medium 511 here includes media that record information optically, electrically, or magnetically, such as CD-ROMs, flexible disks, and magneto-optical disks. The recording medium 511 may also include semiconductor memory that records information electrically, such as ROMs and flash memory. This allows the computer 500 to read and / or write to the recording medium 511 via the external interface 508.
[0052] The various programs to be installed on the HDD 504 are installed, for example, when the distributed recording medium 511 is set in a drive device 510 connected to an external I / F 508, and the various programs recorded on the recording medium 511 are read by the drive device 510. Alternatively, the various programs to be installed on the HDD 504 may be installed by downloading them via the communication I / F 507 from the communication network N or another network different from the communication network N.
[0053] <Functional Configuration> The functional configuration of the physical property prediction system 1000 will be explained with reference to Figure 3. Figure 3 is a block diagram showing an example of the functional configuration of the physical property prediction system.
[0054] ≪Learning Device≫ As shown in Figure 3, the learning device 10 includes a data storage unit 101, a fidelity setting unit 110, a descriptor calculation unit 120, a learning data generation unit 130, and a model generation unit 140. The learning device 10 functions as the data storage unit 101, fidelity setting unit 110, descriptor calculation unit 120, learning data generation unit 130, and model generation unit 140 when a pre-installed learning program is executed.
[0055] For example, the data storage unit 101 is implemented by the HDD 504 shown in Figure 2. For example, the fidelity setting unit 110, the descriptor calculation unit 120, the learning data generation unit 130, and the model generation unit 140 are implemented by a process in which a program loaded from the HDD 504 shown in Figure 2 onto the RAM 503 is executed by the CPU 501.
[0056] Multiple datasets are stored in the data storage unit 101. Multiple datasets may be pre-stored in the data storage unit 101. The datasets may include identification information of the learning target substance and physical property values of the learning target substance.
[0057] Identification information is information that can identify a substance. Examples of identification information may include the compound name, structural formula, SMILES (Simplified Molecular Input Line Entry System) information, ECFP (Extended Connectivity Circular Fingerprints) information, etc. Identification information is not limited to these; any information that can identify a substance can be used.
[0058] The dataset may include, for example, experimental data, simulation data, or literature data. The data storage unit 101 may store at least two of the experimental data, simulation data, or literature data. In this embodiment, an example is described in which experimental data, simulation data, and literature data are pre-stored in the data storage unit 101.
[0059] Experimental data may include experimental values of the physical properties of the target substance measured experimentally, as physical property values of the substance. Multiple datasets may include multiple experimental data. Multiple experimental data may be from multiple experiments using different experimental methods. The experimental methods may include, for example, experiments using existing substances, or experiments using novel substances synthesized from multiple substances.
[0060] Simulation data may include calculated values obtained by simulation as physical properties of the target substance. Multiple datasets may include multiple simulation data. Multiple simulation data may be simulation data using different simulation methods. For example, the simulation method may be quantum chemical calculation or molecular dynamics calculation. Simulation data obtained by quantum chemical calculation may include multiple simulation data with different numbers of basis function parameters. Simulation data obtained by molecular dynamics calculation may include multiple simulation data with different total number of atoms in the system.
[0061] Literature data may include literature values described in publicly available literature as physical properties of the target substance. These literature values may be experimental or calculated. Multiple datasets may include multiple literature data. These multiple literature data may be from different literatures.
[0062] The fidelity setting unit 110 sets the fidelity level for the dataset. The fidelity setting unit 110 may set different fidelity levels for each of the multiple datasets. The fidelity setting unit 110 may also set the fidelity level for each of the multiple datasets read from the data storage unit 101.
[0063] Fidelity can be any information that can be ranked. Fidelity can be any numerical value. Fidelity can be a numerical value where a higher value indicates a higher rank. Fidelity can also be a numerical value where a lower value indicates a higher rank. In this embodiment, fidelity is defined as a numerical value where a higher value indicates a higher rank. For example, if there are three datasets, fidelity 3 indicates the highest fidelity, fidelity 2 indicates the second highest fidelity, and fidelity 1 indicates the lowest fidelity.
[0064] The fidelity setting unit 110 may set the fidelity of a dataset based on the method of acquiring the dataset. The fidelity setting unit 110 may also set the fidelity of each dataset in the order of experimental data, simulation data, and literature data, with the fidelity decreasing in that order. For example, the fidelity setting unit 110 may set the fidelity to 3 for experimental data, to 2 for simulation data, and to 1 for literature data.
[0065] The fidelity setting unit 110 may set fidelity according to the experimental method if the dataset contains multiple experimental data. For example, the fidelity setting unit 110 may set a relatively high fidelity for experimental data measured in experiments using existing substances. Conversely, the fidelity setting unit 110 may set a relatively low fidelity for experimental data measured in experiments using novel substances synthesized from multiple substances.
[0066] The fidelity setting unit 110 may set a fidelity level according to the simulation method if the dataset contains multiple simulation data sets. For example, if the dataset contains multiple simulation data sets from quantum chemical calculations, the fidelity setting unit 110 may set a lower fidelity level as the number of basis function parameters decreases. Also, if the dataset contains multiple simulation data sets from molecular dynamics calculations, the fidelity setting unit 110 may set a lower fidelity level as the total number of atoms in the system decreases. For example, when simulating the behavior of molecules on a substrate, a higher fidelity level may be set as the size of the substrate or the number of molecules on the substrate increases.
[0067] The fidelity setting unit 110 may set fidelity for a dataset based on the similarity between datasets. The fidelity setting unit 110 may set a higher fidelity the higher the correlation coefficient with a reference dataset. The fidelity setting unit 110 may also set a higher fidelity the smaller the difference with a reference dataset. The difference between datasets may, for example, be the sum of the absolute differences of each data point included in the dataset.
[0068] The reference dataset may be predetermined. The reference dataset may also be set based on the method of data acquisition. For example, the reference dataset may be experimental data.
[0069] The descriptor calculation unit 120 calculates the descriptor of the target substance. The descriptor calculation unit 120 may calculate the descriptor of the target substance based on the identification information of the target substance. The descriptor calculation unit 120 may calculate the descriptor of the target substance based on the identification information of the target substance included in the dataset. The descriptor calculation unit 120 may calculate the descriptor of each target substance included in each of the multiple datasets.
[0070] The descriptor calculation unit 120 may calculate descriptors according to the physical properties to be predicted. For example, if the physical property to be predicted is dielectric constant, the descriptor calculation unit 120 may calculate, as an example, the amount of carboxylic acid, water solubility, the logarithm of the octanol / water partition coefficient (logP), etc.
[0071] The descriptor calculation unit 120 may calculate descriptors according to the fidelity set for the dataset. For example, the descriptor calculation unit 120 may calculate logP and water solubility for the learning target substances included in the dataset with fidelity 1, and calculate carboxylic acid content and water solubility for the learning target substances included in the dataset with fidelity 2.
[0072] The learning data generation unit 130 generates learning data. The learning data is electronic data used to train a prediction model. The learning data generation unit 130 may generate learning data that includes identification information of the target substance, descriptors of the target substance, and physical properties of the target substance. The identification information of the learning data may be identification information included in the dataset. The descriptors of the learning data may be descriptors calculated by the descriptor calculation unit 120. The physical properties of the learning data may be physical properties included in the dataset. In other words, the learning data generation unit 130 may generate learning data by adding descriptors calculated by the descriptor calculation unit 120 to the dataset stored in the data storage unit 101.
[0073] The learning data generation unit 130 may generate multiple learning data sets. For example, the learning data generation unit 130 may generate learning data for each fidelity level based on a dataset for each fidelity level. The learning data generation unit 130 may also generate a single learning data set by combining the learning data for each fidelity level. The learning data generation unit 130 may store the generated learning data in the data storage unit 101.
[0074] The model generation unit 140 generates a prediction model. The model generation unit 140 may generate a prediction model based on the training data generated by the training data generation unit 130. The model generation unit 140 may generate a prediction model based on the training data read from the data storage unit 101. The model generation unit 140 may also generate a prediction model by learning the relationship between descriptors and physical properties based on the descriptors and physical property values included in the training data.
[0075] The model generation unit 140 outputs a trained prediction model. The model generation unit 140 may also transmit the trained prediction model to the prediction device 20. The trained prediction model may be stored in the model storage unit 201 of the prediction device 20.
[0076] <Prediction Device> As shown in Figure 3, the prediction device 20 comprises a model storage unit 201, an information acquisition unit 210, a descriptor calculation unit 220, and a physical property prediction unit 230. The prediction device 20 functions as the model storage unit 201, information acquisition unit 210, descriptor calculation unit 220, and physical property prediction unit 230 when a pre-installed prediction program is executed.
[0077] For example, the model storage unit 201 is implemented by the HDD 504 shown in Figure 2. For example, the information acquisition unit 210, the descriptor calculation unit 220, and the physical property prediction unit 230 are implemented by a process in which a program loaded from the HDD 504 shown in Figure 2 onto the RAM 503 is executed by the CPU 501.
[0078] The model storage unit 201 stores a trained prediction model. The model storage unit 201 may already have a pre-trained prediction model stored in it. The prediction model stored in the model storage unit 201 may be generated by the learning device 10.
[0079] The information acquisition unit 210 acquires identification information of the substance to be predicted. The information acquisition unit 210 may also receive identification information of the substance to be predicted transmitted by the terminal device 30. The information acquisition unit 210 may also accept input of identification information of the substance to be predicted via the input device 505 of the prediction device 20.
[0080] The descriptor calculation unit 220 calculates the descriptor of the target substance. The descriptor calculation unit 220 may calculate the descriptor of the target substance based on the identification information of the target substance. The descriptor calculation unit 220 may calculate the descriptor of the target substance based on the identification information of the target substance acquired by the information acquisition unit 210. The descriptor calculation unit 220 may calculate the descriptor in the same manner as the descriptor calculation unit 120 of the learning device 10.
[0081] The physical property prediction unit 230 predicts the physical properties of the target substance. The physical property prediction unit 230 may predict the physical properties of the target substance based on a prediction model generated by the learning device 10. The physical property prediction unit 230 may predict the physical properties of the target substance based on a prediction model read from the model storage unit 201. The physical property prediction unit 230 may predict the physical properties of the target substance based on descriptors calculated by the descriptor calculation unit 220. The physical property prediction unit 230 may predict the physical properties of the target substance by inputting the descriptors of the target substance into the prediction model.
[0082] The physical property prediction unit 230 may output the predicted physical properties. The physical property prediction unit 230 may transmit the predicted physical properties to the terminal device 30. The physical property prediction unit 230 may present the predicted physical properties to the user. The physical property prediction unit 230 may display the predicted results on the display device 506 of the prediction device 20. The physical property prediction unit 230 may output the information used for the prediction along with the predicted physical properties. The information used for the prediction may, as an example, include at least one of the identification information of the substance to be predicted or the descriptor of the substance to be predicted.
[0083] ≪Model Structure≫ The prediction model according to the first embodiment consists of a high-fidelity model constructed by differential learning (hereinafter also referred to as the "differential learning model"). The model structure of the differential learning model will be explained with reference to Figure 4. Figure 4 is a diagram showing an example of a differential learning model.
[0084] Figure 4 shows an example of a differential learning model constructed using three datasets D1 to D3. Dataset D1 is a dataset with a fidelity of 1 (e.g., literature data). Dataset D2 is a dataset with a fidelity of 2 (e.g., simulation data). Dataset D3 is a dataset with a fidelity of 3 (e.g., experimental data).
[0085] As shown in Figure 4, the differential learning model includes three prediction models M1 to M3. Prediction model M1 is a model that predicts material properties based on material descriptors. Prediction model M1 takes material descriptors as input and outputs predicted values for the material properties of that material. Prediction model M1 is learned based on material property values (material property value 1) included in dataset D1.
[0086] Predictive model M2 is a model that predicts the difference between the prediction result based on dataset D1 with fidelity 1 and the prediction result based on dataset D2 with fidelity 2. Predictive model M2 takes a material descriptor as input and outputs a predicted value of the difference between the prediction result based on dataset D1 and the prediction result based on dataset D2. Predictive model M2 is learned based on the difference (physical property 2 - predicted value 1) between the predicted value of predictive model M1 (predicted value 1) and the physical property values (physical property 2) included in dataset D2. More specifically, when a material descriptor is input, predictive model M2 is learned so that the output difference 21 approaches the difference (physical property 2 - predicted value 1).
[0087] Predictive model M3 is a model that predicts the difference between the prediction result based on dataset D2 with fidelity 2 and the prediction result based on dataset D3 with fidelity 3. Predictive model M3 takes a material descriptor as input and outputs a predicted value of the difference between the prediction result based on dataset D2 and the prediction result based on dataset D3. Predictive model M3 is learned based on the sum of the predicted value of predictive model M1 (predictive value 1) and the predicted value of predictive model M2 (difference 21) (predictive value 2) and the difference between that and the physical property value (physical property value 3) included in dataset D3 (physical property value 3 - predictive value 2). More specifically, when a material descriptor is input, predictive model M3 is learned so that the output difference 32 approaches the difference (physical property value 3 - predictive value 2).
[0088] In material property prediction using a differential learning model, the sum of the predicted value from prediction model M1 (predicted value 1), the predicted value from prediction model M2 (difference 21), and the predicted value from prediction model M3 (difference 32) (predicted value 1 + difference 21 + difference 32) is calculated. Figure 4 shows an example of constructing a differential learning model using three datasets, but when using four or more datasets, prediction models M4 and below can be added to predict the difference in the prediction results, similar to prediction model M3.
[0089] The predictive model M1 that predicts material properties and the predictive models M2 and M3 that predict differences may be the same type of machine learning model or different types of machine learning models. For example, both predictive models may be random forests. As another example, one predictive model may be a random forest and the other predictive model may be a Gaussian process regression.
[0090] In this embodiment, prediction model M1 is an example of a first model. Prediction models M2 and M3 are examples of a second model.
[0091] ≪Training Data≫ Figure 5 shows an example of training data for a differential learning model. Figure 5 shows examples of high-fidelity datasets (e.g., experimental data) and low-fidelity datasets (e.g., simulation data), as well as examples of high-fidelity training data and low-fidelity training data.
[0092] As shown in Figure 5, each dataset includes, as data items, SMILES information, which is an example of identification information, and dielectric constant, which is an example of a physical property value. The SMILES information in a fidelity 1 dataset and a fidelity 2 dataset may or may not overlap (i.e., there may be SMILES information present in both datasets, or SMILES information present in only one). Note that high-fidelity datasets (e.g., experimental data) often have a limited amount of data that can be collected. Conversely, low-fidelity datasets (e.g., simulation data) often have a large amount of data that can be collected. However, the number of data points in each dataset can be arbitrary.
[0093] Each training data set includes, as data items, SMILES information, which is an example of identification information; carboxylic acid amount, logP, water solubility, etc., which are examples of descriptors; and dielectric constant, which is an example of a physical property value. In other words, training data is information to which descriptors have been added to the dataset. In this embodiment, high-fidelity training data and low-fidelity training data may or may not have overlapping descriptors (i.e., there may be descriptors present in both, or descriptors present in only one).
[0094] <Processing Procedure> The material property prediction method performed by the material property prediction system 1000 will be explained with reference to Figures 6 to 9. The material property prediction method includes a learning process performed by the learning device 10 (see Figure 6) and a prediction process performed by the prediction device 20 (see Figure 8).
[0095] ≪Learning Process≫ Figure 6 is a flowchart showing an example of the learning process according to the first embodiment. The learning process is the process of learning a predictive model. The learning process is performed by the learning device 10.
[0096] In step S1, the fidelity setting unit 110 of the learning device 10 reads multiple datasets from the data storage unit 101. The fidelity setting unit 110 sets different fidelity levels for each of the multiple datasets. The fidelity setting unit 110 sends information indicating the fidelity level set for each dataset to the learning data generation unit 130.
[0097] In step S2, the descriptor calculation unit 120 of the learning device 10 reads multiple datasets from the data storage unit 101. For each of the multiple datasets, the descriptor calculation unit 120 obtains identification information of the learning target substances contained in that dataset. Based on the identification information obtained from the datasets, the descriptor calculation unit 120 calculates descriptors for the learning target substances. The descriptor calculation unit 120 sends the descriptors of the learning target substances contained in each dataset to the learning data generation unit 130.
[0098] In step S3, the learning data generation unit 130 of the learning device 10 receives information indicating the fidelity of each dataset from the fidelity setting unit 110. The learning data generation unit 130 also receives descriptors of the learning target substances included in each dataset from the descriptor calculation unit 120.
[0099] The training data generation unit 130 generates training data for each fidelity level by adding descriptors of the target substance to each dataset. The training data generation unit 130 sends the training data for each fidelity level to the model generation unit 140.
[0100] In step S4, the model generation unit 140 of the learning device 10 receives learning data for each fidelity level from the learning data generation unit 130. Based on the learning data for each fidelity level, the model generation unit 140 learns the relationship between descriptors and physical properties of the target substance to be learned, thereby generating a trained predictive model. In this embodiment, the model generation unit 140 constructs a predictive model by differential learning by executing a model generation process (see Figure 7).
[0101] In step S5, the model generation unit 140 of the learning device 10 outputs the trained prediction model generated in step S4. The learning device 10 transmits the trained prediction model output from the model generation unit 140 to the prediction device 20.
[0102] The prediction device 20 receives a trained prediction model from the learning device 10. The prediction device 20 stores the received trained prediction model in the model storage unit 201.
[0103] ≪Model Generation Process≫ Figure 7 is a flowchart showing an example of the model generation process according to the first embodiment. The model generation process corresponds to step S4 of the learning process (see Figure 6).
[0104] In step S11, the model generation unit 140 generates a predictive model with fidelity 1 based on training data with fidelity 1. The model generation unit 140 may also generate a predictive model with fidelity 1 by learning the relationship between descriptors and physical properties included in the training data with fidelity 1.
[0105] The process from step S12 to step S15 is repeated for each integer f between 1 and F, where F is the maximum fidelity value. The integer f is initially set to 1, and 1 is added to it each time the process from step S12 to step S15 is executed.
[0106] In step S12, the model generation unit 140 predicts the physical properties based on a prediction model with fidelity 1. The model generation unit 140 may also predict the physical properties of the material to be learned by inputting descriptors included in the training data with fidelity f into the prediction model with fidelity 1. The model generation unit 140 acquires the predicted values (predicted values with fidelity 1) output by the prediction model with fidelity 1.
[0107] Step S13 is repeatedly performed for each integer i between 2 and f-1. The integer i is initially set to 2, and 1 is added to it each time step S13 is performed.
[0108] In step S13, the model generation unit 140 predicts the difference i(i-1) based on the prediction model for the difference i(i-1). For example, if F=3 and f=3, the model generation unit 140 may predict the difference 21 based on the prediction model for the difference 21. Note that if f=1, f-1=0, so step S13 is not executed.
[0109] In step S14, the model generation unit 140 calculates the difference f(f-1). The model generation unit 140 may calculate the difference f(f-1) based on the physical properties included in the training data for fidelity f, the predicted value for fidelity 1 obtained in step S12, and the difference predicted in step S13. As an example, the model generation unit 140 may calculate the difference f(f-1) by subtracting the sum of the predicted value for fidelity 1 and the difference i(i-1) from the physical properties included in the training data.
[0110] For example, if F = 3 and f = 2, the model generation unit 140 may calculate the difference 21 by subtracting the predicted value for fidelity 1 from the physical property value for fidelity 2. Alternatively, if F = 3 and f = 3, the model generation unit 140 may calculate the difference 32 by subtracting the sum of the predicted value for fidelity 1 and the difference 21 from the physical property value for fidelity 3.
[0111] In step S15, the model generation unit 140 generates a prediction model for the difference f(f-1). The model generation unit 140 may generate a prediction model for the difference f(f-1) based on the difference f(f-1) calculated in step S14.
[0112] The model generation unit 140 generates a prediction model with fidelity 1 and a prediction model with differences 21 to F(F-1) by executing the model generation process shown in Figure 7. The model generation unit 140 outputs the prediction model with fidelity 1 and the prediction model with differences 21 to F(F-1) together as a prediction model composed of a difference learning model.
[0113] In the explanation above for Figure 7, the process was described as proceeding to train the prediction model for difference i(i-1) after the prediction model for difference f(f-1) has been trained. However, it is also possible to train the prediction model for difference f(f-1) in parallel with training the prediction model for difference i(i-1).
[0114] <Prediction Processing> Figure 8 is a flowchart showing an example of prediction processing according to the first embodiment. Prediction processing is the process of predicting physical properties based on a trained prediction model. Prediction processing is performed by the prediction device 20.
[0115] In step S21, the user of the physical property prediction system 1000 inputs identification information of the substance to be predicted into the terminal device 30. The terminal device 30 transmits the identification information of the substance to be predicted to the prediction device 20.
[0116] The prediction device 20 receives identification information of the target substance from the terminal device 30. The information acquisition unit 210 of the prediction device 20 acquires the identification information of the target substance received by the prediction device 20. The information acquisition unit 210 sends the acquired identification information of the target substance to the descriptor calculation unit 220.
[0117] In step S22, the descriptor calculation unit 220 of the prediction device 20 receives identification information of the substance to be predicted from the information acquisition unit 210. Based on the identification information of the substance to be predicted, the descriptor calculation unit 220 calculates the descriptor of the substance to be predicted. The descriptor calculation unit 220 sends the descriptor of the substance to be predicted to the physical property prediction unit 230.
[0118] In step S23, the physical property prediction unit 230 of the prediction device 20 receives a descriptor of the material to be predicted from the descriptor calculation unit 220. The physical property prediction unit 230 reads a learned prediction model from the model storage unit 201. The physical property prediction unit 230 inputs the descriptor of the material to be predicted into the read prediction model. The prediction model predicts predetermined physical properties based on the input descriptor of the material to be predicted and outputs predicted values for the predetermined physical properties. The physical property prediction unit 230 acquires the predicted values output from the prediction model. In this embodiment, the physical property prediction unit 230 predicts predetermined physical properties based on the differential learning model by executing a physical property prediction process (see Figure 9).
[0119] In step S24, the physical property prediction unit 230 of the prediction device 20 outputs the physical property prediction results. The physical property prediction results include the predicted values of the physical properties obtained in step S23. The physical property prediction results may also include the information used to predict the physical properties. As an example, the information used for prediction may include at least one of the identification information of the substance to be predicted or the descriptor of the substance to be predicted.
[0120] The prediction device 20 transmits the prediction results of the physical properties to the terminal device 30. The terminal device 30 receives the prediction results of the physical properties from the prediction device 20. The terminal device 30 presents the received prediction results of the physical properties to the user. The terminal device 30 may display the predicted values included in the prediction results on the display device 506. The terminal device 30 may also display the information used for the prediction along with the predicted values.
[0121] Users of the physical property prediction system 1000 may refer to the predicted physical property results displayed on the display device 506 of the terminal device 30 and use them for material development. For example, a user may use the predicted physical property values of each candidate substance to select a substance from among several candidate substances that has properties that meet a target as a raw material for a specific product.
[0122] <<Physical Property Prediction Processing>> Figure 9 is a flowchart showing an example of physical property prediction processing according to the first embodiment. The physical property prediction processing corresponds to step S23 of the prediction processing (see Figure 8).
[0123] In step S31, the physical property prediction unit 230 predicts the physical properties of the target substance based on a prediction model with fidelity of 1. The physical property prediction unit 230 may also predict the physical properties of the target substance by inputting the descriptor of the target substance into the prediction model with fidelity of 1. The physical property prediction unit 230 acquires the predicted values (predicted values with fidelity of 1) output by the prediction model with fidelity of 1.
[0124] Step S32 is repeatedly performed for each integer f between 2 and F, where F is the maximum fidelity value. The integer f is initially set to 2, and 1 is added to it each time step S32 is performed.
[0125] In step S32, the material property prediction unit 230 predicts the difference f(f-1) based on the prediction model for the difference f(f-1). For example, if F=3, the material property prediction unit 230 predicts the difference 21 based on the prediction model for the difference 21 when f=2, and predicts the difference 32 based on the prediction model for the difference 32 when f=3.
[0126] In step S33, the physical property prediction unit 230 calculates predicted values for the physical properties of the target substance. The physical property prediction unit 230 may calculate the predicted values for the physical properties of the target substance based on the predicted value with fidelity 1 obtained in step S31 and the difference f(f-1) predicted in step S32. The physical property prediction unit 230 may also calculate the sum of the predicted value with fidelity 1 and the difference f(f-1) as the predicted value for the physical properties of the target substance. The physical property prediction unit 230 outputs a prediction result that includes the predicted value for the physical properties of the target substance.
[0127] [Second Embodiment] In the first embodiment, an example was described in which a prediction model is constructed using a high-fidelity model built by differential learning (Delta-Learning). In the second embodiment, an example is described in which a prediction model is constructed using a high-fidelity model built by Gaussian process regression (Co-Kriging).
[0128] The following describes the physical property prediction system 1000 according to the second embodiment, focusing on the differences from the first embodiment. Unless otherwise specified, the physical property prediction system 1000 according to the second embodiment may be configured in the same way as the first embodiment.
[0129] ≪Model Structure≫ The prediction model according to the second embodiment consists of a high-fidelity model constructed by Gaussian process regression (hereinafter also referred to as the "Gaussian process regression model"). The model structure of the Gaussian process regression model will be explained with reference to Figure 10. Figure 10 is a diagram showing an example of the Gaussian process regression model.
[0130] Figure 10 shows an example of a Gaussian process regression model constructed using three datasets D1 to D3. As shown in Figure 10, the Gaussian process regression model includes one predictive model M1.
[0131] The predictive model M1 is generated by performing Gaussian process regression on the training data D4. The training data D4 is a dataset created by combining datasets D1 to D3. Additionally, a data item indicating fidelity is added to the training data D4.
[0132] In material property prediction using a Gaussian process regression model, the predicted values of the prediction model M1 are calculated. Figure 10 shows an example of constructing a Gaussian process regression model using three datasets, but when using four or more datasets, a training data D4 can be generated by combining all the datasets, and Gaussian process regression can be performed on the training data D4.
[0133] ≪Training Data≫ Figure 11 shows an example of training data for a Gaussian process regression model. Figure 11 shows an example of a high-fidelity dataset (e.g., experimental data), a low-fidelity dataset (e.g., simulation data), and an example of training data created by combining multiple training datasets.
[0134] As shown in Figure 11, each dataset is the same as in the first embodiment. That is, each dataset includes SMILES information and dielectric constant, and the SMILES information may or may not be duplicated.
[0135] The training data includes, as data items, SMILES information as an example of identification information, carboxylic acid content, water solubility, etc. as examples of descriptors, fidelity, and dielectric constant as an example of a physical property value. In other words, the training data is information obtained by adding descriptors and fidelity to the dataset and combining them. In this embodiment, the descriptors calculated from the dataset with high fidelity are the same as the descriptors calculated from the dataset with low fidelity.
[0136] ≪Learning Device≫ The learning device 10 according to this embodiment includes a data storage unit 101, a fidelity setting unit 110, a descriptor calculation unit 120, a learning data generation unit 130, and a model generation unit 140. The processing content of the learning data generation unit 130 and the model generation unit 140 in the learning device 10 according to this embodiment differs from that of the first embodiment.
[0137] In this embodiment, the learning data generation unit 130 generates learning data by combining the datasets stored in the data storage unit 101. When combining the datasets, the learning data generation unit 130 adds the fidelity set by the fidelity setting unit 110 and the descriptors calculated by the descriptor calculation unit 120 to each dataset.
[0138] In this embodiment, the model generation unit 140 generates a predictive model based on training data which is a combination of datasets stored in the data storage unit 101. Specifically, the model generation unit 140 generates a predictive model by performing Gaussian process regression on training data which includes identification information, descriptors, fidelity, and physical properties.
[0139] ≪Prediction Device≫ The prediction device 20 according to this embodiment includes a model storage unit 201, an information acquisition unit 210, a descriptor calculation unit 220, and a physical property prediction unit 230. The processing content of the physical property prediction unit 230 in the prediction device 20 according to this embodiment differs from that of the first embodiment.
[0140] The physical property prediction unit 230 according to this embodiment predicts the physical properties of a target substance by inputting the identification information and descriptor of the target substance and the fidelity into the prediction model according to this embodiment. The fidelity may be a fixed value. The fidelity may also be the value that shows the highest fidelity among the fidelity values set by the fidelity setting unit 110.
[0141] [Third Embodiment] In the first embodiment, an example was described in which a prediction model is constructed using a high-fidelity model built by differential learning (Delta-Learning). In the second embodiment, an example was described in which a prediction model is constructed using a high-fidelity model built by Gaussian process regression (Co-Kriging). In the third embodiment, an example was described in which a prediction model is constructed using a high-fidelity model built using low-fidelity features.
[0142] The following describes the physical property prediction system 1000 according to the third embodiment, focusing on the differences from the first embodiment. Unless otherwise specified, the physical property prediction system 1000 according to the third embodiment may be configured in the same way as the first embodiment.
[0143] ≪Model Structure≫ The prediction model according to the third embodiment consists of a high-fidelity model (hereinafter also referred to as the "low-fidelity feature model") constructed using low-fidelity features. The model structure of the low-fidelity feature model will be explained with reference to Figure 12. Figure 12 is a diagram showing an example of the low-fidelity feature model.
[0144] Figure 12 shows an example of a low-fidelity feature model constructed using three datasets D1 to D3. As shown in Figure 12, the low-fidelity feature model includes three predictive models M1 to M3.
[0145] Predictive model M1 is a model that predicts material properties based on material descriptors. Predictive model M1 is trained based on material property values (material property value 1) included in dataset D1.
[0146] Predictive model M2 is a model that predicts material properties based on material descriptors. Predictive model M2 is trained based on material property values (material property values 2) included in dataset D2.
[0147] Predictive model M3 is a model that predicts material properties based on material descriptors. Predictive model M3 is trained based on material property values (material property value 3) included in dataset D3. However, predictive model M3 includes the predicted values from predictive model M1 (predicted value 1) and predictive model M2 (predicted value 2) as explanatory variables. In other words, a low-fidelity feature model includes the predicted values from the low-fidelity predictive model as explanatory variables in the highest-fidelity predictive model.
[0148] In predicting material properties using a low-fidelity feature model, the descriptor of the material to be predicted, the predicted value from prediction model M1 (predicted value 1), and the predicted value from prediction model M2 (predicted value 2) are input to prediction model M3. The predicted value from prediction model M3 becomes the prediction result of the low-fidelity feature model. Figure 12 shows an example of constructing a low-fidelity feature model using three datasets, but when using four or more datasets, similar to prediction model M3, a prediction model that predicts material properties can be generated based on training data that adds the predicted value from the low-fidelity prediction model to the material property values from the dataset with the highest fidelity.
[0149] The predictive models M1 to M3 that predict material properties may be of the same type of machine learning model, or they may be of different types. For example, all predictive models may be random forests. As another example, one or more predictive models may be random forests, and the remaining predictive models may be Gaussian process regressions.
[0150] In this embodiment, prediction models M1 and M2 are examples of the first model. Prediction model M3 is an example of the second model.
[0151] ≪Training Data≫ Figure 13 shows an example of training data for a low-fidelity feature model. Figure 13 shows an example of a high-fidelity dataset (e.g., experimental data) and a low-fidelity dataset (e.g., simulation data), as well as an example of high-fidelity training data and a low-fidelity training data.
[0152] As shown in Figure 13, each data set is the same as in the first embodiment. That is, each data set includes SMILES information and dielectric constant, and the SMILES information may or may not be duplicated.
[0153] Each training data set includes, as data items, SMILES information, which is an example of identification information; descriptors, such as carboxylic acid amount, logP, and water solubility; and dielectric constant, which is an example of a physical property value. Furthermore, high-fidelity training data includes predicted values (low-fidelity predicted values) from a prediction model that has been trained on low-fidelity training data. In other words, training data is information to which predicted values from a low-fidelity prediction model and descriptors have been added to the dataset. In this embodiment, as in the first embodiment, the descriptors of high-fidelity training data and low-fidelity training data may or may not overlap.
[0154] <Processing Procedure> The material property prediction method executed by the material property prediction system 1000 will be described with reference to Figures 14 and 15. The material property prediction method according to the third embodiment differs from the first embodiment in the model generation process of the learning process (step S4 in Figure 6) and the material property prediction process of the prediction process (step S23 in Figure 8).
[0155] ≪Model Generation Process≫ Figure 14 is a flowchart showing an example of the model generation process according to the third embodiment.
[0156] The process from step S41 to step S42 is repeated for each integer f between 1 and F-1, where F is the maximum fidelity value. The integer f is initially set to 1, and 1 is added to it each time the process from step S41 to step S42 is executed.
[0157] In step S41, the model generation unit 140 generates a predictive model with fidelity f based on the training data with fidelity f. The model generation unit 140 may also generate a predictive model with fidelity f by learning the relationship between descriptors and physical properties included in the training data with fidelity f.
[0158] In step S42, the model generation unit 140 predicts the physical properties based on a prediction model with fidelity f. The model generation unit 140 may also predict the physical properties of the material to be learned by inputting descriptors included in the learning data with fidelity F into the prediction model with fidelity f. The model generation unit 140 obtains the predicted values (predicted values with fidelity f) output by the prediction model with fidelity f.
[0159] In step S43, the model generation unit 140 updates the training data for fidelity F. The model generation unit 140 may update the training data for fidelity F based on the predicted value of fidelity f predicted in step S42. The model generation unit 140 may also update the training data for fidelity F by adding the predicted value of fidelity f as an explanatory variable to the training data for fidelity F.
[0160] In step S44, the model generation unit 140 generates a predictive model with fidelity F based on the training data with fidelity F. The model generation unit 140 may also generate a predictive model with fidelity F based on the training data updated in step S43. The model generation unit 140 may also generate a predictive model with fidelity F by learning the relationship between descriptors and predicted values included in the training data with fidelity F and the physical properties.
[0161] The model generation unit 140 generates prediction models with fidelity levels 1 to F by executing the model generation process shown in Figure 14. The model generation unit 140 then outputs the prediction models with fidelity levels 1 to F as a prediction model composed of low-fidelity feature models.
[0162] <<Material Property Prediction Processing>> Figure 15 is a flowchart showing an example of material property prediction processing according to the third embodiment.
[0163] Step S51 is repeatedly performed for each integer f between 1 and F-1, where F is the maximum fidelity value. The integer f is initially set to 1, and 1 is added to it each time step S51 is performed.
[0164] In step S51, the physical property prediction unit 230 predicts the physical properties of the target substance based on a prediction model with fidelity f. The physical property prediction unit 230 may also predict the physical properties of the target substance by inputting the descriptor of the target substance into the prediction model with fidelity f. The physical property prediction unit 230 obtains the predicted values (predicted values with fidelity f) output by the prediction model with fidelity f.
[0165] In step S52, the physical property prediction unit 230 generates explanatory variables for fidelity F. The physical property prediction unit 230 may generate explanatory variables for fidelity F based on the predicted value of fidelity f predicted in step S51. The physical property prediction unit 230 may also generate explanatory variables for fidelity F by adding the predicted value of fidelity f to the descriptor of the material to be predicted.
[0166] In step S53, the physical property prediction unit 230 predicts the physical properties of the target substance based on a prediction model with fidelity F. The physical property prediction unit 230 may also predict the physical properties of the target substance by inputting the explanatory variables with fidelity F generated in step S52 into the prediction model with fidelity F. The physical property prediction unit 230 acquires the predicted values output by the prediction model with fidelity F as the predicted values of the physical properties of the target substance. The physical property prediction unit 230 outputs a prediction result that includes the predicted values of the physical properties of the target substance.
[0167] <Effects of the Embodiment> The prediction device 20 according to this embodiment predicts the physical properties of a target substance by inputting the descriptor of the target substance into a machine learning model that has learned the relationship between descriptors and physical properties using multiple datasets with different fidelity settings.
[0168] In one respect, this embodiment uses a machine learning model that has learned the relationship between descriptors and physical properties while considering the fidelity of the dataset, thus enabling accurate prediction of the physical properties of materials. In another respect, this embodiment enables efficient prediction of the physical properties of materials because, even if there is only a small amount of highly reliable data such as experimental data, the model can be appropriately trained by adding less reliable data such as simulation data or literature data.
[0169] Multiple datasets may include at least two of experimental data, simulation data, or literature data. In one respect, according to this embodiment, even when experimental data is scarce, the predictive accuracy of the machine learning model can be improved by adding simulation data or literature data.
[0170] The fidelity may be set based on how the dataset was acquired. Experimental data may be assigned a higher fidelity than simulation data and literature data. Simulation data may be assigned a higher fidelity than literature data. In one respect, according to this embodiment, an appropriate fidelity can be set depending on how the datasets were acquired.
[0171] The fidelity may be set based on the similarity between datasets. Multiple datasets may be assigned higher fidelity the higher their correlation coefficient with a reference dataset. Alternatively, multiple datasets may be assigned higher fidelity the smaller their difference with a reference dataset. In one respect, this embodiment allows for setting an appropriate fidelity level based on the similarity between datasets.
[0172] The machine learning model may include a first model trained using a dataset with a first fidelity setting, and a second model trained on the difference between the physical property values included in a dataset with a second fidelity setting and the predicted physical property values by the first model. In one aspect, according to this embodiment, the physical properties of a material can be predicted with high accuracy based on a multi-fidelity model constructed by differential learning.
[0173] The machine learning model may be a Gaussian process regression model trained using a dataset that combines a dataset with a first fidelity setting and a dataset with a second fidelity setting. In one respect, according to this embodiment, the physical properties of a material can be predicted with high accuracy based on a multi-fidelity model constructed by Gaussian process regression.
[0174] The machine learning model may include a first model trained using a dataset with a first fidelity setting, and a second model trained using a dataset with a second fidelity setting, to which the predicted physical properties from the first model are added. In one aspect, according to this embodiment, the physical properties of a material can be predicted with high accuracy based on a multi-fidelity model constructed from low-fidelity features.
[0175] [Supplement] Each function of the embodiments described above can be realized by one or more processing circuits. Hereinafter, "processing circuit" in this specification includes processors programmed to execute each function by software, such as CPUs (Central Processing Units) or GPUs (Graphics Processing Units) implemented by electronic circuits, as well as devices such as ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), FPGAs (Field Programmable Gate Arrays), and conventional circuit modules designed to execute each function described above.
[0176] While embodiments of the present disclosure have been described in detail above, the embodiments disclosed herein are illustrative and not restrictive in all respects. The embodiments can be modified and improved in various ways without departing from the scope and spirit of the appended claims. The features described in the above embodiments can be combined in any way that is not inconsistent with other configurations.
[0177] This application claims priority to Japanese Patent Application No. 2024-166438, filed with the Japan Patent Office on September 25, 2024, which is incorporated herein by reference to its entire contents.
[0178] 10: Learning device 20: Prediction device 30: Terminal device 101: Data storage unit 110: Fidelity setting unit 120: Descriptor calculation unit 130: Learning data generation unit 140: Model generation unit 201: Model storage unit 210: Information acquisition unit 220: Descriptor calculation unit 230: Physical property prediction unit 1000: Physical property prediction system
Claims
1. A prediction device comprising: a descriptor calculation unit configured to calculate a descriptor of a target substance based on identification information of the target substance; and a property prediction unit configured to predict the physical properties of a target substance by inputting the descriptor of the target substance into a machine learning model that has learned the relationship between the descriptor and predetermined physical properties using a plurality of datasets with different fidelity settings.
2. The prediction device according to claim 1, wherein the plurality of datasets include at least two of experimental data, simulation data, or literature data.
3. The prediction device according to claim 2, wherein the fidelity is set based on the method for acquiring the dataset.
4. The prediction device according to claim 3, wherein the experimental data is set to a higher fidelity than the simulation data and the literature data, and the simulation data is set to a higher fidelity than the literature data.
5. The prediction device according to claim 2, wherein the fidelity is set based on the similarity between the datasets.
6. The prediction device according to claim 5, wherein the fidelity is set to be higher the higher the correlation coefficient between the plurality of datasets and a reference dataset.
7. The prediction device according to claim 5, wherein the fidelity is set higher the smaller the difference between the plurality of datasets and a reference dataset.
8. The prediction device according to any one of claims 1 to 7, wherein the machine learning model includes a first model trained using a dataset with a first fidelity set, and a second model trained on the difference between physical property values included in a dataset with a second fidelity set and predicted physical property values by the first model.
9. The prediction device according to any one of claims 1 to 7, wherein the machine learning model is a Gaussian process regression model trained using a dataset obtained by combining a dataset with a first fidelity setting and a dataset with a second fidelity setting.
10. The prediction device according to any one of claims 1 to 7, wherein the machine learning model includes a first model trained using a dataset with a first fidelity setting, and a second model trained using a dataset with a second fidelity setting to which predicted values of physical properties by the first model have been added.
11. A learning device comprising: a fidelity setting unit configured to set different fidelity levels for each of a plurality of datasets containing descriptors of a target substance and values indicating predetermined physical properties of the target substance; and a model generation unit configured to generate a machine learning model by learning the relationship between the descriptors and the physical properties using the plurality of datasets on which the fidelity levels have been set.
12. A prediction method comprising: a procedure in which a computer calculates a descriptor of a target substance based on identification information of the target substance; and a procedure in which the descriptor of the target substance is input into a machine learning model that has learned the relationship between the descriptor and predetermined physical properties using multiple datasets with different fidelity settings, thereby predicting the physical properties of the target substance.
13. A learning method comprising: a procedure for a computer to set different fidelity levels for each of a plurality of datasets, each containing descriptors of a target substance and values indicating predetermined physical properties of the target substance; and a procedure for generating a machine learning model by learning the relationship between the descriptors and the physical properties using the plurality of datasets on which the fidelity levels have been set.
14. A program for causing a computer to perform the following steps: a procedure for calculating a descriptor of a target substance based on identification information of the target substance; and a procedure for predicting the physical properties of a target substance by inputting the descriptor of the target substance into a machine learning model that has learned the relationship between the descriptor and predetermined physical properties using multiple datasets with different fidelity settings.
15. A program for causing a computer to perform the following steps: setting different fidelity levels for each of a plurality of datasets containing descriptors of a target substance and values representing predetermined physical properties of the target substance; and generating a machine learning model by learning the relationship between the descriptors and the physical properties using the plurality of datasets with the fidelity levels set.
Citation Information
Patent Citations
Learning device, physical property prediction device, learning program, and physical property prediction program
WO2024116642A1