Inference device, inference method, and training device

The inference device enhances NNP accuracy by using a two-step process with a first trained model for generalizability and a second model for correction, addressing the limitations of existing NNPs in achieving high accuracy and versatility across diverse atomic structures.

JP2025122318APending Publication Date: 2025-08-21PREFERRED NETWORKS INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024017689
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-08
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing neural network potentials (NNPs) face challenges in achieving high accuracy for various substances while maintaining versatility due to the high cost of generating accurate training data and the limitations of using only stable structures for learning, which can lead to inaccuracies when applied to unstable structures.

Method used

An inference device that utilizes a first trained model to generate an initial physical property value and a second trained model to correct it using intermediate layer outputs, combining methods like DFT and Coupled-Cluster Singles-and-Doubles to enhance accuracy without increasing computational cost.

Benefits of technology

The device achieves highly accurate physical property values with the same versatility as NNPs by generating energy values with higher precision through a two-step process, utilizing a first trained model for generalizability and a second model for correction, ensuring high accuracy across various atomic structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025122318000001_ABST
    Figure 2025122318000001_ABST
Patent Text Reader

Abstract

To provide an inference device which generates a physical property value with high accuracy while maintaining the same high versatility as a neural network potential.SOLUTION: An inference device includes at least one memory and at least one processor. The at least one processor inputs an atomic structure into a first learned model which has performed learning by learning data calculated with a first method so as to generate a first physical property value corresponding to the atomic structure, and inputs an output from an intermediate layer of the first learned model into a second learned model so as to generate a third physical property value.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD Embodiments of the present disclosure relate to an inference device, an inference method, and a training device. [Background technology]

[0002] A neural network (hereafter referred to as Neural Network Potential (NNP)) is known that predicts the overall energy of an atomic state and the force acting on each atom. NNP can output energy and / or force in an extremely short time compared to simulations of electronic states such as Density Functional Theory (DFT).

[0003] NNP can perform general-purpose, highly accurate calculations for various substances (wide range of elements and structures). For some organic molecules, the required accuracy is very high in practical use, so there is a demand for models that can predict more accurately than NNP.

[0004] In addition, NNPs are sometimes constructed for small molecular systems using training data with higher accuracy than the training data used in NNPs. However, the cost of generating high-accuracy training data is high, and it is difficult to generate a sufficient amount of data to train NNPs. For this reason, training data consisting of only stable structures of a small number of molecules is used for learning. However, such learning can sometimes raise questions about the accuracy when used for unstable structures. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] So Takamoto, Chikashi Shinagawa, Daisuke Motoki, Kosuke Nakago, Wenwen Li, Iori Kurata, Taku Watanabe, Yoshihiro Yayama, Hiroki Iriguchi, Yusuke Asano, Tasuku Onodera, Takafumi Ishii, Takao Kudo, Hideki Ono, Ryohto Sawada, Ryuichiro Ishitani, Marc Ong, Taiki Yamaguchi, Toshiki Kataoka, Akihide Hayashi, Nontawat Charoenphakdee, Takeshi Ibuka, “Towards universal neural network potential for material discovery applicable to arbitrary combinations of 45 elements” Nature Communications volume 13, Article number: 2991 (2022), URL: https: / / www.nature.com / articles / s41467-022-30687-9 [Non-patent document 2] Justin S. Smith, Roman Zubatyuk, Benjamin Nebgen, Nicholas Lubbers, Kipton Barros, Adrian E. Roitberg, Olexandr Isayev, Sergei Tretiak, “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules” Scientific Data. 2020; 7: 134.Published online 2020 May 1. doi: 10.1038 / s41597-020-0473-z,URL:https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC7195467 / Summary of the Invention [Problem to be solved by the invention]

[0006] The problem that the present disclosure aims to solve is to generate highly accurate physical property values ​​while maintaining the same high versatility as neural network potentials. [Means for solving the problem]

[0007] An inference device according to an embodiment includes at least one memory and at least one processor, which inputs an atomic structure to a first trained model trained using training data calculated by a first method to generate a first physical property value corresponding to the atomic structure, and inputs an output from an intermediate layer of the first trained model to a second trained model to generate a third physical property value. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram illustrating an example of a hardware configuration of an inference device according to an embodiment. [Figure 2]FIG. 2 is a diagram illustrating an example of functional blocks in a processor according to the embodiment. [Figure 3] FIG. 3 is a flowchart illustrating an example of a procedure of the energy inference process according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of functional blocks in a processor in the learning device according to the embodiment. [Figure 5] FIG. 5 is a flowchart illustrating an example of a procedure of a learning process according to the embodiment. [Figure 6] FIG. 6 is a diagram showing an example of an initial energy value, a high-accuracy energy value, and a difference (correction value) between the initial energy value and the high-accuracy energy value according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, the embodiments will be described in detail with reference to the drawings.

[0010] (Embodiment) FIG. 1 is a block diagram showing an example of the hardware configuration of an inference device 1 according to this embodiment. As shown in FIG. 1, the inference device 1 may be connected to an external device 9A via a communication network 5. The inference device 1 may also include an external device 9B connected via a device interface 39. The inference device 1 may receive a notation representing the structure of a substance composed of multiple atoms input by a user. The substance is, for example, a molecule. Note that the substance is not limited to molecules and may be various crystals, etc. The notation may be, for example, SMILES (Simplified Molecular Input Line Entry System) notation related to the substance and input by a user. SMILES notation represents, for example, information about a specific molecule (information about atoms and how they are connected) according to certain rules. For example, SMILES notation represents granular information, such as methane having one carbon (C) connected to four hydrogens (H). Note that the notation is not limited to SMILES notation and may be any other known notation as long as it can uniquely identify the substance. For the sake of concreteness, it is assumed below that information input by a user via an input device (to be described later) is information that complies with the SMILES notation (hereinafter referred to as SMILES information).

[0011] The inference device 1 includes a computer 30 and an external device 9B connected to the computer 30 via a device interface 39. The computer 30 includes, as an example, a processor 31, a main storage device (memory) 33, an auxiliary storage device (memory) 35, a network interface 37, and a device interface 39. The inference device 1 may be realized as a computer 30 in which the processor 31, the main storage device 33, the auxiliary storage device 35, the network interface 37, and the device interface 39 are connected via a bus 41.

[0012] Although the computer 30 shown in FIG. 1 includes one of each component, it may also include multiple of the same component. Furthermore, while FIG. 1 shows a single computer 30, the software may be installed on multiple computers, with each of the multiple computers executing the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates via a network interface 37 or the like to execute processing. In other words, the inference device 1 in this embodiment may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize various functions described below. Furthermore, information transmitted from a terminal may be processed by one or more computers provided on the cloud, and the processing results may be transmitted to a terminal such as a display device (display unit) corresponding to the external device 9B.

[0013] The various calculations of the inference device 1 in this embodiment may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processor cores within a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. disclosed herein may be executed by at least one of a processor and a storage device provided on a cloud that can communicate with the computer 30 via a network. Thus, the various functions described below in this embodiment may be implemented in the form of parallel computing using one or more computers.

[0014] The processor 31 may be an electronic circuit (such as a processing circuit, processing circuitry, CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit)) including a control device and an arithmetic device of the computer 30. The processor 31 may also be a semiconductor device including a dedicated processing circuit. The processor 31 is not limited to an electronic circuit using electronic logic elements, but may also be realized by an optical circuit using optical logic elements. The processor 31 may also include an arithmetic function based on quantum computing.

[0015] The processor 31 performs arithmetic processing based on data and software (programs) input from each device, etc., configured internally of the computer 30, and can output the arithmetic results and control signals to each device, etc. The processor 31 may control each component constituting the computer 30 by executing the OS (Operating System) of the computer 30, applications, etc.

[0016] The inference device 1 in this embodiment may be realized by one or more processors 31. Here, the processor 31 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other via wire or wirelessly.

[0017] The main memory device 33 is a memory device that stores instructions executed by the processor 31 and various data, and information stored in the main memory device 33 is read by the processor 31. The auxiliary memory device 35 is a memory device other than the main memory device 33. Note that these memory devices refer to any electronic component that can store electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory device for saving various data used in the inference device 1 according to this embodiment may be realized by the main memory device 33 or the auxiliary memory device 35, or may be realized by an internal memory built into the processor 31. For example, the memory unit in this embodiment may be realized by the main memory device 33 or the auxiliary memory device 35.

[0018] Multiple processors may be connected (coupled) to one storage device (memory), or a single processor 31 may be connected. Multiple storage devices (memories) may be connected (coupled) to one processor. When the inference device 1 in this embodiment is configured with at least one storage device (memory) and multiple processors connected (coupled) to this at least one storage device (memory), it may include a configuration in which at least one of the multiple processors is connected (coupled) to at least one storage device (memory). This configuration may also be realized by storage devices (memories) and processors 31 included in multiple computers. Furthermore, it may include a configuration in which the storage device (memory) is integrated with the processor 31 (for example, a cache memory including an L1 cache and an L2 cache).

[0019] The network interface 37 is an interface for connecting to the communication network 5 wirelessly or via a wire. The network interface 37 may be an appropriate interface, such as one that conforms to an existing communication standard. Information may be exchanged with an external device 9A connected via the communication network 5 through the network interface 37. The communication network 5 may be any one of a WAN (Wide Area Network), a LAN (Local Area Network), a PAN (Personal Area Network), etc., or a combination thereof, as long as information is exchanged between the computer 30 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.

[0020] The device interface 39 is an interface such as a USB (Universal Serial Bus) that directly connects to an output device such as a display device, an input device, and an external device 9 B. The output device may also have a speaker that outputs sound and the like.

[0021] The external device 9A is a device connected to the computer 30 via a network. The external device 9B is a device connected directly to the computer 30.

[0022] The external device 9A or the external device 9B may be, for example, an input device (input unit). The input device is, for example, a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the computer 30. The external device 9A or the external device 9B may also be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0023] Furthermore, the external device 9A or the external device 9B may be, for example, an output device (output unit). The output device may be, for example, a display device (display unit) such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound or the like. Furthermore, the external device 9A or the external device 9B may be a device that includes an output device, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0024] Furthermore, the external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage such as an HDD.

[0025] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of the inference device 1 in this embodiment. In other words, the computer 30 may transmit or receive some or all of the processing results of the external device 9A or the external device 9B.

[0026] FIG. 2 is a diagram illustrating an example of functional blocks in the processor 31. The processor 31 has, for example, a first energy value generation unit 311, a second energy value generation unit 313, and a third energy value generation unit 315 as functions realized by the processor 31. Hereinafter, each generation unit will be described as generating an energy value. However, instead of an energy value, the generation unit may generate a physical property value such as a band gap, a dipole moment, an ionization energy, an electronegativity, or an excitation spectrum. The functions realized by the first energy value generation unit 311, the second energy value generation unit 313, and the third energy value generation unit 315 are stored as programs in, for example, the main storage device 33 or the auxiliary storage device 35. The processor 31 can realize the functions related to the first energy value generation unit 311, the second energy value generation unit 313, and the third energy value generation unit 315 by reading and executing the programs stored in the main storage device 33 or the auxiliary storage device 35.

[0027] The first energy value generation unit 311 may generate a three-dimensional atomic structure based on SMILES information (hereinafter referred to as SMILES information) input via an input device. The atomic structure corresponds to an atomic arrangement in which multiple atoms of a substance are three-dimensionally arranged, as expressed in SMILES notation. For example, the first energy value generation unit 311 generates the atomic structure by inputting SMILES notation into a neural network (hereinafter referred to as a neural network potential (NNP)). Since known techniques can be used as appropriate for the process of generating the atomic structure based on SMILES information, a description thereof will be omitted. The NNP is highly versatile and can generate accurate energy values ​​for various atomic structures. The NNP may be referred to as a first trained model. In other words, the first trained model may be realized by the NNP. The NNP may generate training data using, for example, PBE, which is a DFT method, and train the NNP. For example, the first trained model may be trained in advance using training data calculated using a first method, such as DFT including PBE. That is, the first technique is DFT. Note that the first trained model is not limited to NNP, and other trained neural networks may be used.

[0028] The first energy value generation unit 311 inputs the atomic structure into a first trained model trained using training data calculated by the first method, and generates a first physical property value corresponding to the atomic structure. In this case, the first energy value generation unit 311 may be referred to as a first physical property value generation unit. The first physical property value is, for example, an energy value. For example, the first energy value generation unit 311 may input the generated atomic structure into the first trained model and generate a first energy value corresponding to the atomic structure. The first energy value generation unit 311 stores the generated first energy value, for example, in the main storage device 33 or the auxiliary storage device 35. The first trained model is trained in advance and stored, for example, in the main storage device 33 or the auxiliary storage device 35. Since known techniques can be used as appropriate to generate the first energy value using the first trained model using the atomic structure, a description thereof will be omitted.

[0029] Furthermore, the first energy value generation unit 311 may output intermediate data from one of multiple intermediate layers in the first trained model. The one intermediate layer may be, for example, the last intermediate layer among the multiple intermediate layers. Note that the one intermediate layer is not limited to the last intermediate layer and may be a predetermined intermediate layer. Furthermore, instead of only one intermediate layer, multiple intermediate layers selected from the multiple intermediate layers may be used. The first energy value generation unit 311 stores the intermediate data in, for example, the main storage device 33 or the auxiliary storage device 35. Note that the output and storage of the intermediate data may be achieved by the second energy value generation unit 313.

[0030] The second energy value generation unit 313 inputs an output from an intermediate layer of the first trained model into the second trained model to generate a second physical property value. That is, the second trained model generates a second physical property value relating to the difference between a physical property value corresponding to an atomic structure calculated by a second method different from the first method and the first physical property value. The second physical property value is, for example, an energy value. The second method is, for example, CCSD (Coupled-Cluster Singles-and-Doubles). The second energy value generation unit 313 may be referred to as a second physical property value generation unit. For example, the second energy value generation unit 313 may input intermediate data into the second trained model to generate a second energy value. The second trained model is trained in advance and stored, for example, in the main storage device 33 or the auxiliary storage device 35. The second trained model may be configured, for example, by a neural network having multiple fully connected intermediate layers (hereinafter referred to as a fully connected neural network). The second trained model may infer the difference between the first energy value and an energy value based on structural optimization and energy calculation with higher accuracy than the method for generating training data for the first trained model (hereinafter referred to as the energy difference). The energy difference corresponds to a correction value for correcting the output from the first trained model to a highly accurate energy value. That is, the second trained model corresponds to, for example, a fully connected neural network that takes intermediate data as input and outputs an energy difference. Note that the second trained model may be a neural network that outputs highly accurate energy itself. Furthermore, the input to the second trained model may not only be the output of the intermediate layer, but may also be the output of the first trained model, atomic structure information, etc., alone or in combination. Furthermore, the energy difference may be information related to the difference, and may be not only the energy difference but also various coefficients or fixed values. For example, the output of the second trained model may be any information that can convert the output of the first trained model to a third energy value, and may be a coefficient by which the output of the first trained model is multiplied or divided, or a fixed value different from the difference by which the output of the first trained model is added or subtracted from the output of the first trained model.

[0031] The second trained model is not limited to the fully connected neural network. For example, the second trained model may have a structure up to an intermediate layer in the first trained model and a neural network connected after the intermediate layer. Specifically, the second trained model may have a structure up to one of the multiple intermediate layers in the first trained model and a fully connected neural network connected after the one intermediate layer. In this case, the second energy value generation unit 313 inputs the generated atomic structure into the second trained model and outputs the energy difference. In this case, there is no need to output (extract) intermediate data from the first trained model or temporarily store the intermediate data. The second energy value generation unit 313 associates the generated second energy value with the generated atomic structure and stores it in, for example, the main memory device 33 or the auxiliary memory device 35.

[0032] The third energy value generation unit 315 generates the third physical property value based on the second physical property value and the first physical property value. As a result, the inference device 1 inputs the output from the intermediate layer of the first trained model into the second trained model to generate the third physical property value. That is, the inference device 1 inputs the first physical property value into the second trained model in addition to the output from the intermediate layer of the first trained model to output the third physical property value. The third physical property value is, for example, an energy value. The third energy value generation unit 315 may be referred to as a third physical property value generation unit. For example, the third energy value generation unit 315 may generate the third energy value by adding the first energy value and the second energy value. The third energy value is an energy value corresponding to the atomic structure generated by the first trained model. The third energy value corresponds to an energy value obtained by correcting the first energy value. That is, the third physical property value is a physical property value with relatively higher accuracy than the first physical property value. For example, the third energy value is an energy value with relatively higher accuracy than the first energy value. The third energy value generating unit 315 stores the generated third energy value in the main storage device 33 or the auxiliary storage device 35, for example, in association with the generated atomic structure.

[0033] The configuration of the inference device 1 has been described above. The process of determining an energy value by the inference device 1 (hereinafter referred to as the energy inference process) will now be described with reference to FIG. 3. Hereinafter, the first physical property value, the second physical property value, and the third physical property value will be described as the first energy value, the second energy value, and the third energy value, but are not limited to this. For example, the first physical property value, the second physical property value, and the third physical property value may be any one of a band gap, a dipole moment, an ionization energy, an electronegativity, or an excitation spectrum.

[0034] FIG. 3 is a flowchart illustrating an example of a procedure for the energy inference process.

[0035] (Energy inference processing) (Step S301) The first energy value generation unit 311 may input the SMILES information into the first trained model to generate an atomic structure. The first energy value generation unit 311 may input the generated atomic structure into the first trained model to generate a first energy value. The first energy value generation unit 311 may also extract intermediate data from one hidden layer. The first energy value generation unit 312 stores the generated first energy value and intermediate data in association with the generated atomic structure, for example, in the main storage device 33 or the auxiliary storage device 35.

[0036] (Step S302) The second energy value generation unit 313 may input intermediate data to the second trained model to generate the second energy value. The second energy value generation unit 313 stores the generated second energy value in, for example, the main storage device 33 or the auxiliary storage device 35 in association with the generated atomic structure.

[0037] (Step S303) The third energy value generation unit 315 may generate the third energy value by adding the first energy value and the second energy value. The third energy value generation unit 315 associates the generated third energy value with the atomic structure and stores the third energy value in, for example, the main storage device 33 or the auxiliary storage device 35.

[0038] In view of the above, the inference device 1 according to the present embodiment may input an atomic structure to a first trained model trained using training data calculated by the first method to generate a first physical property value corresponding to the atomic structure, and input an output from an intermediate layer of the first trained model to a second trained model to generate a third physical property value. Furthermore, the third physical property value in the inference device 1 according to the embodiment is a physical property value with relatively higher accuracy than the first physical property value. Furthermore, the second trained model in the inference device 1 according to the embodiment may be a neural network. The second trained model in the inference device 1 according to the embodiment may have a structure up to the intermediate layer in the first trained model and a neural network connected after the intermediate layer.

[0039] As a result, the inference device 1 according to this embodiment uses a first trained model that has high versatility (generalizability) similar to that of NNP, and a second trained model that can generate a difference between NNP and a highly accurate energy value (a correction value of the energy value output by NNP) based on structural optimization and energy calculation with higher accuracy than the method of generating training data for the first trained model. This makes it possible to generate energy values ​​(third energy values) for various atomic structures with high versatility and high accuracy without changing the data distribution of the first trained model. As a result, the inference device 1 according to this embodiment can generate highly accurate energy values ​​with the same computational cost as a neural network potential while maintaining the same high generalizability as a neural network potential.

[0040] Hereinafter, a process of learning (training) a neural network to be learned will be described with respect to the second trained model according to this embodiment. The neural network to be learned (hereinafter referred to as the learning target NN) will be described as a fully connected multi-layer neural network (Deep Neural Network), but is not limited to this. The learning target NN may also be realized by a single-layer fully connected neural network. Learning of the learning target NN may be performed by a learning device (also referred to as a training device) using teacher data (also referred to as correct answer data) and training data corresponding to the teacher data and input to the learning target NN. The learning device is realized, for example, by the processor 31 shown in FIG. 1 .

[0041] FIG. 4 is a diagram illustrating an example of functional blocks in the processor 31 of the learning device. The learning device may also be referred to as a training device. The processor 31 may have, for example, a collection unit 316, a structure generation unit 317, a learning energy generation unit 319, and a learning unit 321 as functions realized by the processor 31. The functions realized by the collection unit 316, the structure generation unit 317, the learning energy generation unit 319, and the learning unit 321 are stored as programs in, for example, the main storage device 33 or the auxiliary storage device 35. The processor 31 can realize the functions related to the collection unit 316, the structure generation unit 317, the learning energy generation unit 319, and the learning unit 321 by reading and executing the programs stored in, for example, the main storage device 33 or the auxiliary storage device 35. The processing contents realized by the collection unit 316, the structure generation unit 317, the learning energy generation unit 319, and the learning unit 321 will be described later with reference to FIG. 5.

[0042] The teacher data and training data used for learning correspond to learning data. Learning the target NN using the teacher data and training data corresponds to adjusting multiple parameters included in a fully connected multi-layer (or single-layer) neural network, for example, by backpropagation.

[0043] The second trained model is generated by training the neural network to be trained using a difference between a first intermediate output value output from an intermediate layer of the first trained model by inputting a first atomic structure generated by optimizing an initial atomic structure using the first trained model into the first trained model, an initial physical property value generated by inputting the first atomic structure into the first trained model, and a high-precision physical property value calculated by applying a second atomic structure generated based on the initial atomic structure by structural optimization with higher precision than the first trained model to a method for calculating high-precision physical property values ​​using the first trained model. For example, the second trained model is generated by training a target NN using a first intermediate output value output from one hidden layer by inputting a first atomic structure generated by optimizing an initial atomic structure using the first trained model into the first trained model, and the difference between an initial energy value generated by inputting the first atomic structure into the first trained model and a high-precision energy value calculated by applying a second atomic structure generated based on the initial atomic structure through structure optimization with higher precision than the method for generating training data for the first trained model to a method for energy calculation with higher precision than the method for generating training data for the first trained model. Additionally, the second trained model may be generated by inputting an unstable atomic structure into the first trained model, inputting the second intermediate output value output from the hidden layer of the first trained model into the target NN, and further training the target NN so that the spatial derivative of the output value becomes zero. The input for training the second trained model may be not only the output from the hidden layer, but also the output from the first trained model, atomic structure information, etc., alone or in combination. The second trained model may be trained to output high-precision energy itself. Furthermore, the difference during training may be information related to the difference, and may be not only an energy difference but also various coefficients or fixed values. For example, the output of the second trained model during training may be any information that can convert the output of the first trained model into a third energy value, and may be a coefficient by which the output of the first trained model is multiplied or divided, or a fixed value different from the difference to be added to or subtracted from the output of the first trained model.

[0044] The following describes a process (hereinafter referred to as a learning process) that includes a process procedure for generating learning data and a learning procedure for a learning object NN. Fig. 5 is a flowchart showing an example of the procedure of the learning process.

[0045] (Learning process) (Step S501) The collection unit 316 may collect multiple initial atomic structures of the substance to be learned. The multiple initial atomic structures are, for example, data used in training for generating the first learned model. For example, the collection unit 316 collects the multiple initial atomic structures from a predetermined database. The collection unit 316 stores the multiple initial atomic structures in, for example, the main storage device 33 or the auxiliary storage device 35.

[0046] (Step S502) The structure generation unit 317 may input a plurality of initial atomic structures into the first trained model, perform structure optimization, and generate a plurality of first atomic structures. The structure generation unit 317 may store the plurality of first atomic structures in the main storage device 33 or the auxiliary storage device 35. The plurality of first atomic structures is, for example, stable structure data of small organic molecules. Note that the plurality of first atomic structures is not limited to stable structure data of small organic molecules, and may be data of other systems and / or other stable structures.

[0047] (Step S503) The training energy generation unit 319 may input the plurality of first atomic structures into the first trained model, respectively, and generate a plurality of initial energy values. The initial energy values ​​may be initial physical property values. In this case, the training energy generation unit 319 may be referred to as a training physical property value generation unit. The training energy generation unit 319 associates the plurality of initial energy values ​​with the plurality of first atomic structures and stores them in, for example, the main storage device 33 or the auxiliary storage device 35.

[0048] (Step S504) The structure generation unit 317 may generate multiple second atomic structures by applying a structural optimization method with higher accuracy than the method for generating training data for the first trained model to each of the multiple initial atomic structures. A structural optimization method with higher accuracy than the method for generating training data for the first trained model is, for example, first-principles calculation (a calculation method based on quantum mechanics that does not rely on experimental values ​​other than fundamental physical constants). That is, the structure generation unit 317 may generate multiple second atomic structures by applying the first-principles calculation to each of the multiple initial atomic structures. The first-principles calculation is, for example, DFT-ωB97XD / 6-31G(d) using density functional theory (DFT). Note that the first-principles calculation is not limited to DFT-ωB97XD / 6-31G(d), and any known structural optimization algorithm with higher accuracy than the method for generating training data for the first trained model can be applied. The structure generation unit 317 stores the multiple second atomic structures in, for example, the main storage device 33 or the auxiliary storage device 35.

[0049] (Step S505) The training energy generation unit 319 may calculate multiple high-precision physical property values ​​by applying each of the multiple second atomic structures to a method for calculating physical property values ​​with higher accuracy than the method for generating training data for the first trained model. For example, the training energy generation unit 319 may calculate multiple high-precision energy values ​​by applying each of the multiple second atomic structures to a method for calculating energy with higher accuracy than the method for generating training data for the first trained model. An example of a method for calculating energy (high-precision physical property values) with higher accuracy than the method for generating training data for the first trained model is DLPNO-CCSD(T) / cc-pv(t,q)z (extrapolation method). Note that the method for calculating energy with higher accuracy than the method for generating training data for the first trained model is not limited to DLPNO-CCSD(T) / cc-pv(t,q)z, and any known method for calculating energy with higher accuracy than the method for generating training data for the first trained model can be applied. Furthermore, when generating band gap energy as a physical property value other than energy, DFT may be used as a method for generating training data for the first trained model, and RPA (Random Phase Approximation) may be used as a method with higher accuracy than the first trained model. When generating dipole moment, ionization energy, or electronegativity, DFT may be used as a method for generating training data for the first trained model, and CCSD(T) may be used as a method with higher accuracy than the first trained model. When generating excitation spectra, semi-empirical molecular orbital method may be used as a method for generating training data for the first trained model, and CAS-SCF may be used as a method with higher accuracy than the first trained model. The training energy generation unit 319 stores multiple high-accuracy energy values, for example, in the main memory device 33 or the auxiliary memory device 35.

[0050] (Step S506) The training energy generation unit 319 may calculate a difference between the initial energy value and the high-precision energy value for each of the first atomic structures. The difference corresponds to a correction value for correcting the initial energy value to the high-precision energy value. The training energy generation unit 319 stores the high-precision energy values ​​in, for example, the main storage device 33 or the auxiliary storage device 35.

[0051] Figure 6 shows the initial energy value E based on the first trained model. NNP (X NNP,stable ) and the energy value (high-precision energy value) E based on structural optimization and energy calculation with higher accuracy than the method for generating training data for the first trained model. CCSD(T) (X ωB97XD,stable ) and the difference (correction value) E for each of the plurality of first atomic structures CORR (X NNP,stable ) is a diagram showing an example of X shown in FIG. ωB97XD,stable indicates the stable atomic structure calculated by high-precision calculation of geometry optimization for the initial atomic structure. X NNP,stable indicates the stable structure of the atom generated by the first trained model such as NNP for the initial atomic structure. As shown in Figure 6, the high-precision energy value E CCSD(T) (X ωB97XD,stable ) is the initial energy value E NNP (X NNP,stable ) compared to the energy value E CCSD(T) is low, so the initial energy value E NNP (X NNP,stable ) shows that it is more accurate.

[0052] Also, as shown in Figure 6, the correction value E CORR (X NNP,stable ) is the initial energy value E NNP (X NNP,stable ) and high-precision energy value E CCSD(T) (X ωB97XD,stable ) and the difference (E CCSD(T) (X ωB97XD,stable )-E NNP (X NNP,stable In the learning process, the stable structure X of the atom generated by the first trained model is used for the learning target NN. NNP,stable Enter the correction value E CORR (X NNP,stable ) may be trained to output

[0053] (Step S507) The learning energy generating unit 319 may associate a plurality of first atomic structures with a plurality of differences, which correspond to teacher data during training of the learning target NN.

[0054] (Step S508) The collection unit 316 may collect unstable structures of multiple atoms used in learning about the first trained model. An unstable atomic structure refers to an atom in an intermediate position that is not a stable structure, such as when a molecular bond is stretched or a chemical reaction is occurring. The collection unit 316 stores the unstable atomic structures in, for example, the main memory device 33 or the auxiliary memory device 35.

[0055] (Step S509) The learning energy generation unit 319 may associate "energy gradient = 0" with unstable structures of multiple atoms. "Energy gradient = 0" corresponds to force = 0. This corresponds to teacher data for further training for "energy gradient = 0."

[0056] (Step S510) The learning unit 321 may input each of the multiple first atomic structures into the first trained model and output a first intermediate output value from one intermediate layer. The learning unit 321 may generate a first training dataset by associating multiple first intermediate output values ​​(training data) with multiple differences (teaching data) based on the multiple first atomic structures. The learning unit 321 stores the generated first training dataset in, for example, the main storage device 33 or the auxiliary storage device 35.

[0057] Furthermore, the learning unit 321 may input unstable structures of multiple atoms into the first trained model and output second intermediate output values ​​from one intermediate layer. The learning unit 321 may generate a second training dataset by associating multiple second intermediate output values ​​(training data) with "energy gradient = 0." The learning unit 321 stores the generated second training dataset in, for example, the main storage device 33 or the auxiliary storage device 35.

[0058] (Step S511) The learning unit 321 uses the first learning data set to train the neural network to be trained. That is, the learning unit 321 may train the neural network to be trained using a plurality of first intermediate output values ​​(training data) and a plurality of differences (teaching data). A known method can be applied to the learning process for the neural network to be trained, and therefore a description thereof will be omitted. In this step, the neural network to be trained uses an output value output from one intermediate layer of the first trained model as input, and the initial energy value E NNP (X NNP,stable ) and high-precision energy value E CCSD(T) (X ωB97XD,stable ) and the difference (E CCSD(T) (X ωB97XD,stable )-E NNP (X NNP,stable The connection weights between the nodes may be learned to output an energy value indicative of the

[0059] The learning unit 321 may use the first learning data set to train the neural network to be trained. That is, the learning unit 321 may use a plurality of first intermediate output values ​​(training data) and a plurality of differences (teaching data) to train the neural network to be trained. A known method can be applied to the learning process for the neural network to be trained, so a description thereof will be omitted. In this step, in the neural network to be trained, an output value output from one intermediate layer of the first trained model is used as input, and an initial energy value E NNP (X NNP,stable ) and high-precision energy value E CCSD(T) (X ωB97XD,stable ) and the difference (E CCSD(T) (X ωB97XD,stable )-E NNP (X NNP,stable The connection weights between the nodes may be learned to output an energy value indicative of the

[0060] (Step S512) The learning unit 321 may use the second learning data set to learn the neural network to be trained in step S512. That is, the learning unit 321 may further train the neural network to be trained by inputting each of the multiple second intermediate output values ​​(training data) to the neural network to be trained so that the spatial derivative of the output value becomes zero. Since a known method can be applied to the learning process in this step, a description thereof will be omitted. A second trained model may be generated by this step.

[0061] Based on the above, the training device according to this embodiment trains a first trained model that generates a first physical property value corresponding to an input atomic structure from training data calculated by a first method, and trains a second trained model that uses an output from an intermediate layer of the first trained model as an input to the second trained model and generates a third physical property value based on the second physical property value output from the second trained model and the first physical property value. For example, the second trained model according to this embodiment may be generated by training a neural network to be trained using a first intermediate output value output from the intermediate layer when the first atomic structure generated by optimizing an initial atomic structure using the first trained model is input to the first trained model, and a difference between an initial physical property value generated by inputting the first atomic structure into the first trained model and a high-precision physical property value calculated by applying a second atomic structure generated based on the initial atomic structure through structural optimization with higher precision than the method for generating training data for the first trained model to a method for calculating physical properties with higher precision than the first trained model. Furthermore, the second trained model in this embodiment may be generated by inputting an unstable atomic structure into the first trained model, inputting the second intermediate output value output from the intermediate layer of the first trained model into the neural network to be trained, and further performing training on the neural network to be trained so that the spatial derivative of the output value becomes zero.

[0062] For these reasons, the second trained model according to this embodiment can generate highly accurate energy values ​​for stable atomic structures with the same amount of calculation (calculation speed) as the first trained model, i.e., without increasing the calculation cost. Furthermore, the second trained model according to this embodiment can generate second energy values ​​by training using the second training dataset so as not to distort the distribution of the first energy values ​​output by the first trained model.

[0063] When the technical idea of ​​the embodiments is realized as an inference method, the inference method inputs an atomic structure to a first trained model trained using training data generated by a first technique, generates a first physical property value corresponding to the atomic structure, inputs output from an intermediate layer of the first trained model to a second trained model, and generates a third physical property value based on the output second physical property value and the first physical property value. For example, the inference method of the embodiments may input an atomic structure to a first trained model implemented using a neural network potential, generates a first energy value corresponding to the atomic structure, input intermediate data corresponding to output from one of multiple intermediate layers in the first trained model to a second trained model that infers the difference between the first energy value and an energy value based on structural optimization and energy calculation with higher accuracy than the first trained model, generates a second energy value, and generates a third energy value by adding the first energy value and the second energy value. The procedure and effects of the energy inference process of the inference method are similar to those described in the embodiments, and therefore will not be described again.

[0064] When the technical idea in the embodiment is realized by an inference program, the inference program may cause a computer to input an atomic structure into a first trained model realized by a neural network potential, generate a first energy value corresponding to the atomic structure, input intermediate data corresponding to the output from one of multiple intermediate layers in the first trained model into a second trained model that infers the difference between the first energy value and an energy value based on structural optimization and energy calculation with higher accuracy than a method of generating training data for the first trained model, generate a second energy value, and generate a third energy value by adding the first energy value and the second energy value.

[0065] For example, the energy inference process can be realized by installing the inference program in a computer in various analytical devices or analysis servers that analyze the energy and / or force of an atomic structure composed of multiple atoms and expanding the program in memory. In this case, the program that can cause a computer to execute the inference method can also be stored and distributed on a storage medium such as a magnetic disk (such as a hard disk), an optical disk (such as a CD-ROM or DVD), or a semiconductor memory. The procedure and effect of the energy inference process using the inference program are the same as those in the embodiment, so a description thereof will be omitted.

[0066] Some or all of the devices in the above-described embodiments may be configured as hardware, or may be configured as information processing software (programs) executed by a CPU, a GPU, or the like. In the case of software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored on a non-transitory storage medium (non-transitory computer-readable medium) such as a flexible disk, a CD-ROM (Compact Disc-Read Only Memory), or a USB memory, and the software information processing may be executed by loading the software into the computer 30. The software may also be downloaded via the communication network 5. Furthermore, the software may be implemented in a circuit such as an ASIC or FPGA, so that the information processing is executed by hardware.

[0067] The type of storage medium that stores the software is not limited. The storage medium is not limited to removable media such as magnetic disks or optical disks, but may be fixed storage media such as hard disks or memory. The storage medium may be provided inside the computer or outside the computer.

[0068] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.

[0069] In this specification (including the claims), when expressions such as "using data as input / based on / according to / in response to" (including similar expressions) are used, unless otherwise specified, this includes cases where various data itself is used as input, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is used as input. Furthermore, when a statement is made that a result is obtained "based on / according to / in response to data," this includes cases where the result is obtained based solely on the data in question, as well as cases where the result is obtained as a result of being influenced by other data, factors, conditions, and / or states other than the data in question. Furthermore, when a statement is made that "data is output," this includes cases where various data itself is used as output, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is output, unless otherwise specified.

[0070] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that encompass any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately according to the context in which they are used, but any form of connection / coupling that is not intentionally or naturally excluded should be interpreted as being included in these terms without limitation.

[0071] In this specification (including the claims), the expression "A configured to B" may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.

[0072] When used in this specification (including the claims), terms implying containing or possessing (e.g., "comprising / including" and "having") are intended to be open-ended terms that include containing or possessing things other than the object designated by the object of the term. When the object of such a term implies no quantity or a singular number (e.g., an article such as "a" or "an"), the expression should be construed as not being limited to a specific number.

[0073] In this specification (including the claims), even if expressions such as "one or more" or "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.

[0074] In this specification, when a particular advantage / result is described as being obtained from a particular configuration of an embodiment, it should be understood that the same advantage / result can also be obtained from one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, etc., and that the effect is not necessarily obtained by the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or states, etc. are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same or a similar configuration.

[0075] When used in this specification (including the claims), terms such as "maximize" include finding a global maximum, finding an approximation of a global maximum, finding a local maximum, and finding an approximation of a local maximum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these maxima probabilistically or heuristically. Similarly, when used in this specification (including the claims), terms such as "minimize" include finding a global minimum, finding an approximation of a global minimum, finding a local minimum, and finding an approximation of a local minimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these minima probabilistically or heuristically. Similarly, when used in this specification (including the claims), terms such as "optimize" include finding a global optimum, finding an approximation of a global optimum, finding a local optimum, and finding an approximation of a local optimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these optima probabilistically or heuristically.

[0076] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that hardware may include an electronic circuit or a device including an electronic circuit.

[0077] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices (memories) may store only a portion of the data, or may store the entire data.

[0078] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention derived from the content defined in the claims and their equivalents. For example, in all of the above-described embodiments, when numerical values ​​or formulas are used in the explanation, they are shown as examples and are not limited to these. Furthermore, the order of each operation in the embodiments is shown as an example and is not limited to these. [Explanation of symbols]

[0079] 1 Reasoning device 5. Communication Network 9A external device 9B External device 30 Computer 31 processors 33 Main memory 35 Auxiliary storage device 37 Network Interface 39 Device Interfaces 41 Bus 311 First energy value generation unit 313 Second energy value generation unit 315 Third energy value generation unit 316 Collection Department 317 Structure generation part 319 Learning Energy Generation Unit 321 Learning Department

Claims

1. at least one memory; at least one processor, The at least one processor inputting an atomic structure into a first trained model trained using training data calculated by a first method, and generating a first physical property value corresponding to the atomic structure; An output from an intermediate layer of the first trained model is input to a second trained model to generate a third physical property value. Reasoning device.

2. The second trained model generates a second physical property value relating to a difference between a physical property value corresponding to the atomic structure calculated by a second method different from the first method and the first physical property value; generating the third physical property value based on the second physical property value and the first physical property value; The inference device according to claim 1 .

3. inputting the first physical property value to the second trained model in addition to the output from the intermediate layer, and outputting the third physical property value; The inference device according to claim 1 .

4. the third physical property value is a physical property value having relatively higher accuracy than the first physical property value; The inference device according to claim 1 .

5. The second trained model is a neural network. The inference device according to claim 1 .

6. The second trained model has a structure up to an intermediate layer in the first trained model and a neural network connected after the intermediate layer. The inference device according to claim 1 .

7. The inference device of claim 1 , wherein the first trained model is a neural network potential.

8. The second trained model is trained to estimate a difference between an output of the first trained model and an output of a second method different from the first method. The inference device according to claim 1 .

9. the first physical property value and the third physical property value are energy values; An inference device according to any one of claims 1 to 8.

10. The first technique is DFT. An inference device according to any one of claims 1 to 8.

11. The second method, which is different from the first method, is CCSD. An inference device according to any one of claims 1 to 8.

12. The first physical property value and the third physical property value are any one of a band gap, a dipole moment, an ionization energy, an electronegativity, and an excitation spectrum. An inference device according to any one of claims 1 to 8.

13. The second trained model is a first intermediate output value output from the intermediate layer when a first atomic structure generated by optimizing an initial atomic structure using the first trained model is input to the first trained model; and Using a difference between an initial physical property value generated by inputting the first atomic structure into the first trained model and a high-precision physical property value calculated by applying a second atomic structure generated based on the initial atomic structure by structural optimization with higher precision than the first trained model to a method for calculating a high-precision physical property value than the first trained model, Generated by training the neural network to be learned, An inference device according to any one of claims 1 to 8.

14. The second trained model is The neural network to be trained is further trained by inputting an unstable atomic structure into the first trained model, inputting a second intermediate output value output from the intermediate layer into the neural network to be trained, and then training the neural network to be trained so that the spatial derivative of the output value becomes zero. The inference device of claim 13.

15. inputting an atomic structure into a first trained model trained using training data generated by a first method, and generating a first physical property value corresponding to the atomic structure; an output from an intermediate layer of the first trained model is input to a second trained model, and a third physical property value is generated based on the output second physical property value and the first physical property value; Reasoning method.

16. at least one memory; at least one processor, The at least one processor training a first trained model that generates a first physical property value corresponding to an atomic structure when an atomic structure is input from the training data calculated by the first method; training a second trained model that generates a third physical property value based on the second physical property value output from the second trained model and the first physical property value, using an output from an intermediate layer of the first trained model as an input to the second trained model; training equipment.