Evaluation device, inference device, evaluation method, program, and non-transitory computer-readable medium

The evaluation device assesses NNP models' accuracy and generalization performance by comparing inference results with quantum chemical calculations, addressing the challenge of high-precision inference across diverse chemical structures.

JP7702279B2Active Publication Date: 2025-07-03PREFERRED NETWORKS INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021098302
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-11
Publication Date
2025-07-03
Estimated Expiration
2041-06-11

AI Technical Summary

Technical Problem

Existing neural network potentials (NNPs) face challenges in achieving high-precision inference while maintaining generalization performance, with no appropriate index for evaluating both accuracy and generalization performance across various chemical structures.

Method used

An evaluation device that compares inference results from a trained NNP model with quantum chemical calculations using a validation dataset to assess accuracy and generalization performance across multiple domains, using metrics like mean absolute error (MAE) to determine model quality.

Benefits of technology

Enables accurate evaluation of NNP models' generalization and accuracy, ensuring they can reliably infer various chemical structures by setting thresholds for MAE, thereby improving model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702279000001
    Figure 0007702279000001
  • Figure 0007702279000002
    Figure 0007702279000002
  • Figure 0007702279000003
    Figure 0007702279000003
Patent Text Reader

Abstract

To evaluate generalization performance of an NNP that learns a first principle calculation of various chemical structures.SOLUTION: An evaluation device includes: an inferring part for inputting an atomic structure to an already trained model for inferring a physical property value of an atomic structure from the atomic structure and successively propagating it, and acquiring an inferred result; and an evaluation part for comparing the inferred result with the physical property values acquired by executing a quantum chemical calculation to the atomic structure and evaluating the inferred result. The inferring part acquires each inferred result for the atomic structures that belong to multiple domains. The evaluation part evaluates each of the atomic structures that belong to the multiple domains by using the physical property values. Datasets of the atomic structures and the physical property values that the inferring part and the evaluation part use for inference and evaluation are not used for training the already trained model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an evaluation device, an inference device, an evaluation method, a program, and a non-transitory computer-readable medium.

Background Art

[0002] In order to obtain the energy in molecules, crystals, etc., it is necessary to calculate by a method such as first-principles calculation. Research on NNP (Neural Network Potential) has been conducted as a method for realizing inference based on the results obtained by this first-principles calculation.

[0003] However, as a common problem in neural network models, the NNP model with a higher learning degree to achieve high-precision inference is vulnerable to the extrapolation problem. As a result, when trained to make inferences with high accuracy, the generalization performance decreases, and conversely, in the training of a model with enhanced generalization performance, the accuracy problem becomes significant. For this reason, an index for evaluating both the generalization performance and accuracy of NNP is required, but no appropriate such index has been presented currently. Regarding NNP that has learned only a limited domain such as a specific combination of elements, methods for evaluating accuracy and generalization performance have been reported, but there is no report on a method for accurately inferring various chemical structures for NNP that has learned the results obtained by first-principles calculations of various chemical structures.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

[0005] The present disclosure provides an evaluation device that evaluates the generalization performance and accuracy of an NNP that has learned various chemical structures. [Means for Solving the Problems]

[0006] According to one embodiment, the evaluation device includes an inference unit that inputs an atomic structure to a trained model that infers physical property values of the atomic structure from the atomic structure and performs forward propagation to obtain an inference result, and an evaluation unit that compares and evaluates the inference result with the physical property values obtained by performing quantum chemical calculations on the atomic structure. The inference unit obtains each of the inference results for the atomic structures belonging to a plurality of domains. The evaluation unit evaluates using the physical property values for each of the atomic structures belonging to the plurality of domains. The dataset of the atomic structure and the physical property values used by the inference unit and the evaluation unit for inference and evaluation is a dataset that has not been used for training the trained model.

[0007] The trained model evaluated by this evaluation device may be provided as a model that executes inference in an inference device.

[0008] This evaluation device may be provided in the training device, and the training device may execute the training of the model while evaluating the accuracy and generalization performance of the model to be trained based on the evaluation of the evaluation device.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The drawings and the description of the embodiments are shown as examples and do not limit the present invention.

[0011] (First Embodiment) FIG. 1 is a block diagram schematically showing an evaluation device according to an embodiment. The evaluation device 1 includes an input unit 100, a storage unit 102, an inference unit 104, an evaluation unit 106, and an output unit 108. The evaluation device 1 is a device that evaluates the generalization performance and accuracy of the trained model NN. In addition to the configuration shown in FIG. 1, configurations necessary for the operation of the evaluation device 1 are appropriately provided although not shown.

[0012] The input unit 100 is an interface that receives data input in the evaluation device 1. The evaluation device 1 receives the input of data necessary for evaluation via the input unit 100. Further, the evaluation device 1 may receive the input of data regarding the trained model NN to be evaluated via the input unit 100, for example, data regarding hyperparameters and parameters, etc.

[0013] The storage unit 102 stores data necessary for evaluation in the evaluation device 1. The storage unit 102 stores, for example, information input from the input unit 100, a program for operating the evaluation device 1, and in addition to the evaluation results, intermediate results necessary during the calculation and information such as parameters regarding the trained model NN.

[0014] The inference unit 104 inputs the data regarding the atomic structure input from the input unit 100 into the trained model NN and performs forward propagation to output an inference result. This inference result is a physical property value in the input atomic structure. As an example, this physical property value is a quantity including energy (potential). The inference unit 104 outputs the inference result output from the trained model NN to the storage unit 102 or the evaluation unit 106.

[0015] The trained model NN is a neural network model that outputs a physical property value when an atomic structure is input. This trained model NN is, for example, a neural network model used for NNP and is a model for inferring energy from an atomic structure. The trained model NN is a model trained by an appropriate machine learning method using an atomic structure and an energy value associated with the atomic structure.

[0016] The training of the trained model NN is executed using the atomic structure and the physical property values obtained by first-principles calculations for the atomic structure as a training dataset, but does not prevent the use of other first-principles methods. This first-principles calculation may be, for example, a calculation based on DFT (Density Function Theory), etc. In this case, energy is cited as a physical property value as described above. From this, the trained model NN is formed as a model that infers the calculation results of first-principles calculations (or quantum chemical calculations), DFT, etc. when an atomic structure is input.

[0017] The inference unit 104 performs inference using a validation dataset, which is a dataset other than the dataset used for this training (the training dataset and the dataset for cross-validation checks in training). The validation dataset is obtained, for example, by performing DFT calculations on the atomic structure. As another example, data such as a database that is already known and that has not been used for training can also be used. For example, data from public databases such as PubChem, Material Project, and ICSD, or data from public databases such as complex structure databases can be obtained and used as the validation dataset.

[0018] The evaluation unit 106 evaluates the trained model NN by comparing the result inferred by the inference unit 104 with the physical property values in the validation dataset. Specifically, it evaluates using the inference result obtained by inputting a certain atomic structure to the inference unit 104 and the physical property values associated with the atomic structure. The evaluation unit 106 compares the inference result with the physical property values, for example, calculates the absolute error, and evaluates the trained model NN.

[0019] The validation dataset contains a plurality of data belonging to a plurality of domains respectively. Here, a domain refers to a chemical region related to the atomic structure. For example, examples of domains include crystal structure, amorphous structure, diatomic molecule, molecular structure, surface structure, and adsorption structure, but are not limited thereto.

[0020] To improve the evaluation accuracy of generalization performance, for example, for amorphous structures generated by 2-atom molecules, neighboring 2-molecules, large molecules, adsorption structures, MD (Molecular Dynamics) simulations, etc., energies can be calculated by methods such as DFT and incorporated as verification datasets.

[0021] When the domain is a crystal structure, data can be collected from the Material Project or by DFT calculations or the like.

[0022] When the domain is an amorphous structure, after creating an amorphous state using NNP for the above crystal structure, data can be collected by DFT calculations or the like.

[0023] When the domain is a 2-atom molecule, molecular structure, or neighboring 2-molecule, the interatomic distance of the 2 atoms can be varied in various ways and data can be collected by DFT calculations or the like.

[0024] When the domain is a surface structure, a surface structure can be generated using a function that cuts out the surface from the crystal structure, and data can be collected by DFT calculations or the like.

[0025] When the domain is an adsorption structure, a state where a molecule is adsorbed on the surface structure can be generated, and data can be collected by DFT calculations or the like.

[0026] In the above data collection, various appropriate parameter settings in, for example, VASP (registered trademark), Gaussian (registered trademark), etc., such as appropriate calculation methods and software, can be used.

[0027] As an example, when performing DFT calculations, functionals such as ωB3LYP, B3P86, ωB97X, ωB97X-D, APFD, HSE, TPSSh, M06-2X, PBE, rPBE, revPBE, PBEsol, and PBE0 may be used, but it does not prevent the use of other functionals. By using such functionals, the energy values in the chemical structures listed above can be calculated. In addition, a Hubbard correction term that is a function of the occupancy matrix n or the density matrix ρ can also be added to these energy calculations. Also, the basis functions used can be included in the domain information.

[0028] The inference unit 104 obtains an inference result from the atomic structure of the verification dataset including the datasets belonging to these multiple domains, and the evaluation unit 106 evaluates the trained model NN by comparing the verification dataset with the inference result. As evaluation metrics, energy, force, energy difference, adsorption energy, all interatomic distances and all atomic angles, or lattice constants, etc. can be used. By using these metrics, they can also be used as metrics for comprehensively evaluating each mode. Using these metrics, the inference result is compared with the verification dataset.

[0029] As a non-limiting example, the evaluation unit 106 may evaluate the mean absolute error (MAE) for each domain as the evaluation value. As other examples, for example, the mean square error (MSE), the root mean square error (RMSE), and the coefficient of determination (R2) can be used.

[0030] The evaluation unit 106 evaluates the trained model NN based on the evaluation values obtained for each domain. For example, a threshold value of the evaluation value is defined, and when the evaluation value of each domain is below (or above depending on the evaluation value) the threshold value, the trained model NN may be evaluated as having generalization performance.

[0031] The output unit 108 outputs the evaluation result of the evaluation unit 106. For example, when a trained model NN has generalization performance, a result indicating that it has generalization performance may be output. As another example, the evaluation value calculated by the evaluation unit 106 may be output.

[0032] FIG. 2 is a flowchart showing the processing of the evaluation apparatus 1 according to an embodiment.

[0033] First, the evaluation apparatus 1 receives data and a verification data set regarding the trained model NN to be evaluated via the input unit 100 (S100).

[0034] The inference unit 104 acquires a plurality of atomic structures belonging to a certain domain from the verification data set, and acquires an inference value by the trained model NN for each of the acquired data (S102). When the trained model is a model based on NNP, the inference unit 104 acquires an inference value of energy. Further, in order to improve the accuracy, an inference value of force may also be acquired.

[0035] The evaluation unit 106 compares the inference value for the atomic structure with the physical property value in the verification data set for the atomic structure, and acquires an evaluation value (S104). The evaluation unit 106 calculates, for example, the MAE between the inference values for a plurality of atomic structures belonging to the domain and the evaluation values. This calculated MAE is used as the evaluation value of this domain of the trained model NN.

[0036] The evaluation unit 106 determines whether an evaluation value can be acquired in the domain to be determined (S106). If there is a domain that is the target of determination but for which an evaluation value cannot be acquired (S106: NO), acquisition of the evaluation value for the domain is executed (S102 to S104).

[0037] If there is no domain for which an evaluation value cannot be acquired in the domain to be determined (S106: YES), the evaluation unit 106 evaluates the trained model NN using the evaluation values of the acquired domains (S108).

[0038] The evaluation unit 106 outputs the evaluation result via the output unit 108, and the evaluation device 1 ends the process (S110). By checking this output result, it becomes possible to evaluate the accuracy and generalization performance of the trained model NN.

[0039] In the above process, the processes of S102 to S104 can be processed in parallel. For example, for a plurality of verification data sets belonging to the same domain, the inference values can be obtained in parallel, and then the evaluation values can be obtained. Furthermore, operations in different domains may be executed in parallel. For this execution, for example, an accelerator such as a GPU (Graphics Processing Unit) can be used.

[0040] When using energy as the evaluation value, the evaluation unit 106 may evaluate the trained model NN based on whether the MAE in each domain is less than a predetermined threshold, for example, less than 0.05 eV. The evaluation of the trained model NN may be output as a pass when the MAE is less than the predetermined threshold in all domains to be evaluated. As another example, a score for the generalization performance may be calculated based on the MAE in each domain and the score may be output. As still another example, the MAE of each domain may be output as the evaluation value.

[0041] In the case of energy, although it was set to less than 0.05 eV as an example not limited above, this value is an index indicating that this degree of accuracy of error from the DFT calculation is required to calculate the ease of reaction progress (activation energy). Desirably, the predetermined threshold is 0.03 eV, and more desirably, 0.02 eV. For example, when the error is 0.05 eV or more, it becomes difficult to accurately estimate the superiority or inferiority of chemical reactions and various physical properties. Therefore, the predetermined threshold may be set to 0.05 eV.

[0042] In addition, although the MAE was used above, the variance or standard deviation of the MAE may be calculated and this variance or the like may be used as an evaluation value. For example, based on this standard deviation, if the value of the MAE for each domain is less than 3σ, it may be considered a pass. Desirably, it may be 2.5σ, and more desirably, 2σ. Also, the standard deviation itself may be compared with a predetermined threshold value.

[0043] The evaluation unit 106 can also use the interatomic distance as an evaluation value. This evaluation value can be used, for example, when the domain can define lattice constants such as a molecular structure, a surface structure, and an adsorption structure. In this case, the trained model NN is, for example, a model that infers the total interatomic distance when the structure is optimized by DFT calculation. The inference unit 104 infers the total interatomic distance using this trained model NN, and the evaluation unit 106 obtains the total interatomic distance from the verification dataset, takes the difference between the inferred value and this total interatomic distance as an error, and calculates the MAE from this error.

[0044] Also, the evaluation unit 106 can use the lattice constant as an evaluation value. This evaluation value can be used, for example, when the domain includes a crystal structure. In this case, the trained model NN outputs, as an inferred value, the lattice constant obtained by optimizing the lattice structure by DFT calculation, for example. The evaluation unit 106 obtains the lattice constant from the verification dataset, takes the difference between the result inferred by the inference unit 104 as an error, and calculates the MAE.

[0045] In these cases, for example, it may be used as an evaluation that the value of the MAE is within a predetermined distance and the MAE is within 3σ.

[0046] The trained model NN may be a neural network model that can input calculation conditions determined for each domain together with the atomic structure. This trained model NN may be a model that infers the energy or the like when DFT calculation is performed under the calculation conditions based on the calculation conditions input together with the atomic structure.

[0047] FIG. 3 is a diagram showing an example of a combination when the trained model NN can input DFT calculation conditions. The domain corresponds to the above domain, and the calculation conditions show, as an example, the exchange-correlation functional used, etc.

[0048] The domains respectively show the structures of molecule, molecule, crystal, crystal, amorphous, amorphous, surface, surface, two atoms, adsorption, adsorption in order from the top. The description of the calculation conditions defines the functional, basis function, etc.

[0049] When the trained model NN can input such conditions, the evaluation device 1 can obtain a dataset that matches each calculation condition from the verification dataset in the evaluation of each domain, and within this dataset, perform inference by the inference unit 104 and evaluation by the evaluation unit 106.

[0050] As described above, according to the present embodiment, the accuracy of the trained model NN can be judged by the MAE itself, and between domains, it is possible to evaluate the generalization performance by using the evaluation value using this MAE.

[0051] (Second Embodiment) In the first embodiment, it was assumed that the MAE was calculated for each domain, but the calculation of the MAE can be further subdivided. For example, the MAE may be obtained for each atomic structure having a predetermined element among the atomic structures belonging to the domain.

[0052] Specifically, first, a domain is specified. In this domain, for each element, a dataset is extracted from the verification dataset, and inference is performed using this dataset to calculate the MAE.

[0053] For example, from the verification dataset, an atomic structure containing hydrogen atoms in the domain is extracted, and the inference unit 104 infers the energy value. The evaluation unit 106 calculates the MAE of the atomic structure having hydrogen atoms in this domain using the difference between this inferred value and the energy value in the corresponding verification dataset. Similarly, the MAE of helium atoms, lithium atoms, ··· is calculated. For all elements included in the verification dataset, the MAE may be calculated, or a target element may be determined in advance, and the MAE for the target element may be calculated.

[0054] The evaluation unit 106 executes the evaluation of the domain by statistically processing a plurality of MAEs calculated for each domain. For example, as described above, processing using variance or standard deviation or the like may be performed, or evaluation such as simply comparing the average value of the MAE with a predetermined threshold value may be performed. As another example, each of the obtained MAEs may be compared with a predetermined threshold value, for example, 0.02 eV or the like.

[0055] This comparison can also evaluate the generalization performance by evaluating the accuracy in each atomic structure as the granularity becomes smaller and further using the evaluation value for each domain. Therefore, according to the present embodiment, it is possible to realize the evaluation of individual accuracy and generalization performance.

[0056] (Third Embodiment) The technical scope of the present disclosure naturally extends to an inference device 2 that performs inference using a trained model NN evaluated by the above evaluation device 1.

[0057] FIG. 4 is a block diagram schematically showing an inference device 2 according to an embodiment. The inference device 2 includes an input unit 200, a storage unit 202, an inference unit 204, and an output unit 206. The inference device 2 is, for example, a device that infers the energy in an arbitrary atomic structure in a plurality of domains by NNP. Since the overall configuration is the same as that of the evaluation device 1, some detailed descriptions are omitted.

[0058] The input unit 200 includes an interface for inputting the atomic structure to be inferred. The storage unit 202 stores data necessary for the operation of the inference device 2. The inference unit 204 infers energy from the atomic structure using the trained model NN. The output unit 206 outputs the result inferred by the inference unit 204.

[0059] Here, the trained model NN may be a model evaluated to have high accuracy and generalization performance in the evaluation in each of the above-described embodiments. For example, in the evaluation device 1, it may be a model that has passed the test.

[0060] Since such an inference device 2 guarantees the accuracy and generalization performance of the trained model NN at a certain level, appropriate energy can be obtained from the atomic structure.

[0061] As shown in FIG. 3, when the trained model NN is formed as a model that can input calculation conditions, the inference device 2 also acquires appropriate calculation conditions via the input unit 200. By inputting the calculation conditions together with the atomic structure to the trained model NN by the inference unit 204, the inference device 2 can output an inference result of energy with high accuracy.

[0062] (Fourth Embodiment) The above-described evaluation device 1 can also be incorporated into the training device.

[0063] FIG. 5 is a block diagram schematically showing a training device according to an embodiment. The training device 3 includes an input unit 300, a storage unit 302, an optimization unit 304, an evaluation device 1, and an output unit 306. Similar to the third embodiment, the overall configuration is the same as that of the evaluation device 1, and thus some detailed descriptions are omitted.

[0064] The input unit 300 includes an interface for inputting data and the like necessary for training. The storage unit 302 stores data necessary for the operation of the training device 3.

[0065] The optimization unit 304 trains the model NN1 based on any appropriate machine learning method. Here, the optimization unit 304 uses a training dataset different from the aforementioned verification dataset to execute the training of the model NN1.

[0066] The evaluation device 1 evaluates the model NN1 trained by the optimization unit 304 using the evaluation method shown in the aforementioned embodiments. That is, the evaluation device 1 evaluates the accuracy and generalization performance of the model NN1. Based on the evaluation results, it is determined whether the optimization of the model NN1 in the training device 3 has been completed. If the optimization is insufficient, the optimization unit 304 repeats the training of the model NN1 to optimize the model NN1 again.

[0067] This training may be repeated until the evaluation in the evaluation device 1 is passed.

[0068] The output unit 306 outputs the parameters of the model NN1 optimized by training to the outside or the storage unit 302 and ends the process.

[0069] In this way, the evaluation device 1 may be configured to be incorporated into the training device 3. By incorporating the evaluation device 1 into the training device 3, it becomes possible to train a model that realizes a predetermined accuracy and a predetermined generalization performance in training.

[0070] All of the above trained models may be, for example, in the concept including a model distilled by a general method after being trained as described.

[0071] Some or all of each device (evaluation device 1, inference device 2, or training device 3) in the foregoing embodiments may be configured by hardware, or may be configured by information processing of software (program) executed by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like. When configured by information processing of software, software that realizes at least some functions of each device in the foregoing embodiments is stored in a non-transitory storage medium (non-transitory computer-readable medium) such as a flexible disk, a CD-ROM (Compact Disc-Read Only Memory), or a USB (Universal Serial Bus) memory, and the information processing of the software may be executed by having it read into a computer. Further, the software may be downloaded via a communication network. Furthermore, the information processing may be executed by hardware when the software is implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0072] The type of storage medium for storing the software is not limited. The storage medium is not limited to removable ones such as magnetic disks or optical disks, and may be a fixed-type storage medium such as a hard disk or a memory. Further, the storage medium may be provided inside the computer or outside the computer.

[0073] FIG. 6 is a block diagram showing an example of the hardware configuration of each device (evaluation device 1, inference device 2, or training device 3) in the foregoing embodiments. Each device may be realized as a computer 7 including, as an example, a processor 71, a main storage device 72 (memory), an auxiliary storage device 73 (memory), a network interface 74, and a device interface 75, which are connected via a bus 76.

[0074] The computer 7 in FIG. 6 includes one of each component, but may include a plurality of the same components. Also, in FIG. 6, one computer 7 is shown, but software may be installed on a plurality of computers, and each of the plurality of computers may execute the same or different parts of the software processing. In this case, it may be in the form of distributed computing in which each computer communicates via a network interface 74 or the like to execute processing. That is, each device (evaluation device 1, inference device 2, or training device 3) in the above-described embodiment may be configured as a system in which one or a plurality of computers execute instructions stored in one or a plurality of storage devices to realize functions. Further, it may be configured such that information transmitted from a terminal is processed by one or a plurality of computers provided on the cloud, and the processing result is transmitted to the terminal.

[0075] The various operations of each device (evaluation device 1, inference device 2, or training device 3) in the above-described embodiment may be executed in parallel using one or a plurality of processors or using a plurality of computers via a network. Also, the various operations may be allocated to a plurality of arithmetic cores in the processor and executed in parallel. Also, part or all of the processing, means, etc. of the present disclosure may be executed by at least one of a processor and a storage device provided on the cloud that can communicate with the computer 7 via a network. Thus, each device in the above-described embodiment may be in the form of parallel computing by one or a plurality of computers.

[0076] The processor 71 may be an electronic circuit (processing circuit, Processing circuit, Processing circuitry, CPU, GPU, FPGA, or ASIC, etc.) that includes a control device and an arithmetic device of a computer. Further, the processor 71 may be a semiconductor device or the like that includes a dedicated processing circuit. The processor 71 is not limited to an electronic circuit using electronic logic elements, and may be realized by an optical circuit using optical logic elements. Further, the processor 71 may include an arithmetic function based on quantum computing.

[0077] The processor 71 can perform arithmetic processing based on data and software (programs) input from each device and the like of the internal configuration of the computer 7, and output the arithmetic results and control signals to each device and the like. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, and the like.

[0078] Each device (evaluation device 1, inference device 2, or training device 3) in the above-described embodiments may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When using a plurality of electronic circuits, each electronic circuit may communicate wired or wirelessly.

[0079] The main memory device 72 is a storage device that stores instructions executed by the processor 71 and various data, etc., and the information stored in the main memory device 72 is read by the processor 71. The auxiliary storage device 73 is a storage device other than the main memory device 72. These storage devices mean any electronic components capable of storing electronic information, and may be semiconductor memories. The semiconductor memory may be either a volatile memory or a non-volatile memory. The storage device for storing various data in each of the devices (evaluation device 1, inference device 2, or training device 3) in the above-described embodiments may be realized by the main memory device 72 or the auxiliary storage device 73, or may be realized by a built-in memory built into the processor 71. For example, the storage units 102, 202, and 302 in the above-described embodiments may be realized by the main memory device 72 or the auxiliary storage device 73.

[0080] A plurality of processors may be connected (coupled) to one storage device (memory), or a single processor may be connected. A plurality of storage devices (memories) may be connected (coupled) to one processor. When each of the devices (evaluation device 1, inference device 2, or training device 3) in the above-described embodiments is composed of at least one storage device (memory) and a plurality of processors connected (coupled) to this at least one storage device (memory), the configuration may include that at least one of the plurality of processors is connected (coupled) to at least one storage device (memory). Also, this configuration may be realized by the storage devices (memories) and processors included in a plurality of computers. Furthermore, the configuration may include a configuration in which the storage device (memory) is integrated with the processor (for example, a cache memory including an L1 cache and an L2 cache).

[0081] The network interface 74 is an interface for connecting to a communication network 8, either wirelessly or wired. The network interface 74 may use an appropriate interface such as one that conforms to an existing communication standard. Information exchange may be performed between the external device 9A connected via the communication network 8 and the network interface 74. Note that the communication network 8 may be any one of a WAN (Wide Area Network), LAN (Local Area Network), PAN (Personal Area Network), or a combination thereof, as long as information exchange is performed between the computer 7 and the external device 9A. Examples of a WAN include the Internet, etc., examples of a LAN include IEEE802.11, Ethernet (registered trademark), etc., and examples of a PAN include Bluetooth (registered trademark), NFC (Near Field Communication), etc.

[0082] The device interface 75 is an interface such as USB that directly connects to the external device 9B.

[0083] The external device 9A is a device connected to the computer 7 via a network. The external device 9B is a device directly connected to the computer 7.

[0084] The external device 9A or the external device 9B may be, for example, an input device. The input device is a device such as a camera, microphone, motion capture, various sensors, etc., a keyboard, mouse, or touch panel, etc., and provides the acquired information to the computer 7. It may also be a device equipped with an input unit, memory, and processor, such as a personal computer, tablet terminal, or smartphone.

[0085] Further, the external device 9A or the external device 9B may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound or the like. Further, it may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0086] Further, the external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage such as an HDD.

[0087] Further, the external device 9A or the external device 9B may be a device having some functions of the components of each device (evaluation device 1, inference device 2, or training device 3) in the above-described embodiment. That is, the computer 7 may transmit or receive some or all of the processing results of the external device 9A or the external device 9B.

[0088] In this specification (including the claims), when an expression such as "at least one (one side) of a, b, and c" or "at least one (one side) of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, a - b, a - c, b - c, or a - b - c. Also, for any element, multiple instances may be included, such as a - a, a - b - b, a - a - b - b - c - c, etc. Further, it also includes adding other elements other than the enumerated elements (a, b, and c), such as having d like a - b - c - d.

[0089] In this specification (including the claims), when expressions such as "using data as input / based on data / according to data / in response to data" (including similar expressions) are used, unless otherwise specified, it includes cases where various data themselves are used as input, and cases where some processing has been performed on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) are used as input. Also, when it is described that some result is obtained "based on data / according to data / in response to data", it includes cases where the result is obtained based only on the said data, and may also include cases where the result is obtained under the influence of other data, factors, conditions, and / or states other than the said data. Further, when it is described that "data is output", unless otherwise specified, it includes cases where various data themselves are used as output, and cases where some processing has been performed on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) are output.

[0090] In this specification (including the claims), when the terms "connected" and "coupled" are used, they are intended as non - limiting terms that include any of direct connection / coupling, indirect connection / coupling, electrical connection / coupling, communicative connection / coupling, operative connection / coupling, physical connection / coupling, etc. The terms should be appropriately interpreted according to the context in which they are used, but connection / coupling forms that are not intentionally or naturally excluded should be interpreted non - limitatively as being included in the terms.

[0091] In this specification (including the claims), when the expression "A is configured to B" is used, the physical structure of element A has a configuration capable of performing operation B, and it may include that the permanent or temporary setting / configuration of element A is set to actually perform operation B. For example, when element A is a general-purpose processor, the processor has a hardware configuration capable of performing operation B, and it may be set to actually perform operation B by a permanent or temporary program (instruction) setting. Also, when element A is a dedicated processor or a dedicated arithmetic circuit, etc., regardless of whether control instructions and data are actually attached, the circuit structure of the processor may be implemented to actually perform operation B.

[0092] In this specification (including the claims), when terms meaning containment or possession (e.g., "comprising / including" and "having", etc.) are used, they are intended as open-ended terms, including cases where they contain or possess things other than the object indicated by the object of the term. When the object of these terms meaning containment or possession does not specify a quantity or is an expression suggesting a singular number (an expression with "a" or "an" as an article), the expression should be interpreted as not being limited to a specific number.

[0093] In this specification (including the claims), even if an expression such as "one or more" or "at least one" is used in one place and an expression that does not specify a quantity or implies a singular number (an expression with "a" or "an" as an article) is used in another place, the latter expression is not intended to mean "one". Generally, an expression that does not specify a quantity or implies a singular number (an expression with "a" or "an" as an article) should be construed as not necessarily being limited to a specific number.

[0094] In this specification, if it is described that a specific effect (advantage / result) is obtained for a specific configuration of a certain embodiment, unless there are other reasons, it should be understood that the same effect can also be obtained for one or more other embodiments having the same configuration. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, etc., and the effect is not necessarily obtained by the configuration. The effect is only obtained by the configuration described in the embodiment when various factors, conditions, and / or states, etc., are satisfied, and in the invention according to the claim that defines the configuration or a similar configuration, the effect is not necessarily obtained.

[0095] In this specification (including the claims), when terms such as "maximize" are used, it includes obtaining a global maximum value, obtaining an approximation of the global maximum value, obtaining a local maximum value, and obtaining an approximation of the local maximum value, and should be appropriately interpreted according to the context in which the term is used. Also included is obtaining an approximation of these maximum values probabilistically or heuristically. Similarly, when terms such as "minimize" are used, it includes obtaining a global minimum value, obtaining an approximation of the global minimum value, obtaining a local minimum value, and obtaining an approximation of the local minimum value, and should be appropriately interpreted according to the context in which the term is used. Also included is obtaining an approximation of these minimum values probabilistically or heuristically. Similarly, when terms such as "optimize" are used, it includes obtaining a global optimum value, obtaining an approximation of the global optimum value, obtaining a local optimum value, and obtaining an approximation of the local optimum value, and should be appropriately interpreted according to the context in which the term is used. Also included is obtaining an approximation of these optimum values probabilistically or heuristically.

[0096] In this specification (including the claims), when a plurality of hardware performs a predetermined process, each piece of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Also, some of the hardware may perform a part of the predetermined process and another piece of hardware may perform the remainder of the predetermined process. In this specification (including the claims), when an expression such as "one or more pieces of hardware perform a first process and the one or more pieces of hardware perform a second process" is used, the hardware that performs the first process and the hardware that performs the second process may be the same or different. That is, it is sufficient that the hardware that performs the first process and the hardware that performs the second process are included in the one or more pieces of hardware. Note that the hardware may include an electronic circuit or a device including an electronic circuit, etc.

[0097] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the above-described individual embodiments. Various additions, changes, replacements, and partial deletions are possible without departing from the conceptual ideas and spirits of the present invention derived from the content defined in the claims and their equivalents. For example, in all the embodiments described above, when numerical values or mathematical formulas are used for explanation, they are shown as examples and are not limited thereto. Also, the order of each operation in the embodiments is shown as an example and is not limited thereto.

Explanation of Signs

[0098] 1: Evaluation device, 100: Input unit, 102: Storage unit, 104: Inference unit, 106: Evaluation unit, 108: Output unit, 2: Inference device, 200: Input unit, 202: Storage unit, 204: Inference unit, 206: Output unit, 3: Training device, 300: Input unit, 302: Storage unit, 304: Optimization unit, 306: Output unit

Claims

1. An inference unit that inputs an atomic structure into a trained model, which is a neural network model for inferring physical property values of the atomic structure from the atomic structure, and obtains an inference result; An evaluation unit that evaluates the trained model based on the inference result and the physical property value related to the atomic structure; Comprising: The inference unit obtains each inference result for the atomic structure belonging to a plurality of domains indicating a plurality of chemical regions related to the atomic structure; The evaluation unit calculates an evaluation value using the physical property value and the inference result for each of the atomic structures belonging to the plurality of domains, and evaluates the trained model; The dataset of the atomic structure and the physical property value used by the inference unit and the evaluation unit for inference and evaluation is a dataset not used for training the trained model; An evaluation device.

2. Evaluating the trained model for each atomic structure containing each element; The evaluation device according to Claim 1.

3. The physical property value is obtained by first-principles calculation. The evaluation device according to Claim 1 or Claim 2.

4. The trained model is a neural network model used for NNP (Neural Network Potential). The evaluation device according to any one of Claims 1 to 3.

5. The plurality of domains include at least one of the domains of crystal structure, amorphous structure, diatomic molecule, molecular structure, surface structure or adsorption structure. The evaluation device according to any one of Claims 1 to 4.

6. The evaluation unit calculates the mean absolute error between the inference result for a plurality of the atomic structures and the physical property value as the evaluation value, and performs the evaluation. The evaluation device according to any one of Claims 1 to 5.

7. The evaluation unit calculates the mean absolute error for each domain as the evaluation value, and performs the evaluation. The evaluation device according to Claim 6.

8. The evaluation unit calculates the variance or standard deviation of the mean absolute error for each domain as the evaluation value, and performs the evaluation. The evaluation device according to Claim 7.

9. The inference result includes at least any one of energy, force, energy difference, adsorption energy, total interatomic distance, total atomic angle or lattice constant. The evaluation device according to any one of Claims 1 to 8.

10. The evaluation unit performs evaluation based on whether the evaluation value is less than a predetermined threshold value. The evaluation apparatus according to any one of Claims 1 to 9.

11. Using the evaluation apparatus according to any one of Claims 1 to 10 for evaluation and using a trained model that has obtained a predetermined evaluation value to infer physical property values from an atomic structure. Inference apparatus.

12. Training the model based on the evaluation of the model using the evaluation apparatus according to any one of Claims 1 to 10. Training apparatus.

13. A computer inputs an atomic structure into a trained model that is a neural network model to obtain an inference result, performs evaluation of the trained model based on the inference result and the physical property value related to the atomic structure, An evaluation method, wherein the trained model is a model for inferring the physical property value from the atomic structure, the computer obtains each of the inference results for the atomic structures belonging to a plurality of domains indicating a plurality of chemical regions related to the atomic structure, calculates an evaluation value using the physical property value and the inference result for each of the atomic structures belonging to the plurality of domains, and evaluates the trained model, the dataset of the atomic structure and the physical property value used for inference and evaluation is a dataset not used for training the trained model. Evaluation method.

14. Causing a computer to input an atomic structure into a trained model that is a neural network model to obtain an inference result, to perform evaluation of the trained model based on the inference result and the physical property value related to the atomic structure, A program for causing the above, wherein the trained model is a model for inferring the physical property value from the atomic structure, causing the computer to obtain each of the inference results for the atomic structures belonging to a plurality of domains indicating a plurality of chemical regions related to the atomic structure, to calculate an evaluation value using the physical property value and the inference result for each of the atomic structures belonging to the plurality of domains, and to evaluate the trained model, causing the above to be executed, the dataset of the atomic structure and the physical property value used for inference and evaluation is a dataset not used for training the trained model. Program.

15. Causing a computer to input an atomic structure into a trained model that is a neural network model to obtain an inference result, Based on the inference result and the physical property value related to the atomic structure, evaluate the trained model. A program for causing the above to be executed, The trained model is a model for inferring the physical property value from the atomic structure, On the computer, For the atomic structures belonging to a plurality of domains indicating a plurality of chemical regions related to the atomic structure, obtain each of the inference results, Calculate an evaluation value using the physical property value and the inference result for each of the atomic structures belonging to the plurality of domains, and evaluate the trained model. Cause the above to be executed, The dataset of the atomic structure and the physical property value used for inference and evaluation is a dataset not used for training the trained model. A non-transitory computer-readable medium storing the program.

Citation Information

Patent Citations

  • Crystal form estimating device, crystal form estimating method, neural network manufacturing method, and program

    JP2020166706A

  • Calculation method of surface energy of crystalline material using atomic neural network

    KR102260838B1