Estimation device, estimation method, training device, training method, program, and model generation method
The estimation device uses neural networks to efficiently handle multiple atomic types and states, addressing the limitations of DFT and deep learning models in quantum chemistry calculations, enabling accurate material property predictions and searches.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2026-03-04
AI Technical Summary
Quantum chemistry calculations, such as DFT, are time-consuming and difficult to apply to comprehensive material searches, while deep learning models struggle to handle multiple atomic types and different states like molecules and crystals simultaneously.
An estimation device and method that utilizes neural networks to extract atomic features in a latent space, process graph information, and predict physical properties of molecules and crystals, allowing for efficient handling of various atomic types and states.
Enables accurate and efficient estimation of physical properties of materials, supporting high-speed material searches and structural relaxations, even with increased atomic types, by leveraging neural networks for feature extraction and prediction.
Smart Images

Figure 0007823928000002 
Figure 0007823928000003 
Figure 0007823928000004
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an estimation device, a training device, an estimation method, and a training method. [Background technology]
[0002] Quantum chemistry calculations, such as first-principles calculations such as DFT (Density Functional Theory), are relatively reliable and easy to interpret because they calculate physical properties such as the energy of electronic systems from a chemical background. However, they are time-consuming to calculate and are difficult to apply to comprehensive material searches, so they are currently used for analysis to understand the properties of discovered materials. In response to this, the development of material property prediction models using deep learning technology has progressed rapidly in recent years.
[0003] However, as mentioned above, DFT requires long calculation times. On the other hand, while models using deep learning technology can predict physical properties, existing models that allow coordinate input have difficulty in increasing the number of atomic types and in simultaneously handling different states such as molecules and crystals, as well as their coexistence. DISCLOSURE OF THE INVENTION
[0004] One embodiment provides an estimation device and method that improve the accuracy of estimating physical property values of a material system, and a training device and method therefor.
[0005] According to one embodiment, the estimation device includes one or more memories and one or more processors, which input vectors related to atoms to a first network that extracts features of the atoms in a latent space from the vectors related to the atoms, and estimate the features of the atoms in the latent space via the first network. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a schematic block diagram of an estimation device according to an embodiment. [Figure 2]FIG. 2 is a schematic diagram of an atom feature acquisition unit according to an embodiment. [Figure 3] FIG. 10 is a diagram showing an example of setting coordinates of molecules and the like according to an embodiment. [Figure 4] FIG. 10 is a diagram showing an example of obtaining graph data of molecules or the like according to an embodiment. [Figure 5] FIG. 10 is a diagram showing an example of graph data according to an embodiment. [Figure 6] 10 is a flowchart showing processing of an estimation device according to an embodiment. [Figure 7] FIG. 1 is a schematic block diagram of a training device according to one embodiment. [Figure 8] FIG. 1 is a schematic diagram of a configuration for training an atom feature acquisition unit according to an embodiment. [Figure 9] FIG. 4 is a diagram showing an example of teacher data of physical property values according to an embodiment. [Figure 10] FIG. 10 is a diagram showing a state in which the physical property values of atoms are trained according to an embodiment. [Figure 11] FIG. 2 is a schematic block diagram of a structural feature extraction unit according to an embodiment. [Figure 12] 10 is a flowchart illustrating an overall training process according to one embodiment. [Figure 13] 10 is a flowchart showing a process for training a first network according to one embodiment. [Figure 14] FIG. 10 is a diagram showing an example of physical property values output from a first network according to an embodiment. [Figure 15] 10 is a flowchart illustrating the process of training the second, third, and fourth networks according to one embodiment. [Figure 16] FIG. 10 is a diagram showing an example of output of physical property values according to an embodiment. [Figure 17] 1 illustrates an example implementation of an estimation device or a training device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0007] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes embodiments of the present invention with reference to the drawings. The drawings and the description of the embodiments are given by way of example only and are not intended to limit the present invention.
[0008] [Estimation device] 1 is a block diagram showing the functions of an estimation device 1 according to this embodiment. The estimation device 1 of this embodiment estimates and outputs physical property values of an estimation target, which is a molecule or the like (hereinafter, the term "molecule or the like" refers to a single atom, molecule, or crystal), based on information such as the type of atom, coordinate information, and boundary condition information. This estimation device 1 includes an input unit 10, a storage unit 12, an atomic feature acquisition unit 14, an input information configuration unit 16, a structural feature extraction unit 18, a physical property prediction unit 20, and an output unit 22.
[0009] Necessary information such as the type and coordinates of atoms, which are information to be estimated, such as molecules, and boundary conditions, is input to the estimation device 1 via the input unit 10. In this embodiment, for example, information on the type, coordinates, and boundary conditions of atoms is described as being input, but the present invention is not limited to this and any information that defines the structure of a substance whose physical property values are to be estimated may be used.
[0010] The coordinates of an atom are, for example, three-dimensional coordinates of the atom in absolute space or the like. For example, they may be coordinates in a coordinate system using a translationally invariant and rotationally invariant coordinate system. However, they are not limited to this, and any coordinate system that can appropriately express the structure of an atom in an object such as a molecule to be estimated may be used. By inputting the coordinates of the atom, it is possible to define the relative position of the atom in the molecule or the like.
[0011] Boundary conditions, for example, when obtaining physical property values for a crystal to be estimated, are input as coordinates of atoms within a unit cell or a supercell in which unit cells are repeatedly arranged. In this case, boundary conditions are set for cases where the input atoms are at the boundary surface with the vacuum, or where the same atomic arrangement is repeated nearby. For example, when a molecule is brought close to a catalyst crystal, a boundary condition may be assumed in which the crystal surface in contact with the molecule is the boundary with the vacuum, and the crystal structure is otherwise continuous. In this way, the estimation device 1 can estimate not only physical property values related to molecules, but also physical properties related to crystals and physical property values related to both crystals and molecules.
[0012] The storage unit 12 stores information required for estimation. For example, data used for estimation input via the input unit 10 may be temporarily stored in the storage unit 12. Parameters required for each unit, such as parameters required to form a neural network provided in each unit, may also be stored. Furthermore, when the estimation device 1 specifically realizes information processing by software using hardware resources, programs, executable files, etc. required for this software may also be stored.
[0013] The atomic feature acquisition unit 14 generates a quantity indicating the feature of the atom. The quantity indicating the feature of the atom may be expressed in a one-dimensional vector format, for example. The atomic feature acquisition unit 14 includes a neural network (first network) such as an MLP (Multilayer Perceptron) that converts a one-hot vector indicating the atom into a vector in a latent space when it is input, and outputs this vector in the latent space as the feature of the atom.
[0014] Alternatively, the atomic feature acquisition unit 14 may receive other information, such as a tensor or vector, representing an atom, instead of a one-hot vector. The other information, such as a one-hot vector, tensor, or vector, may be, for example, a code representing the atom of interest, or similar information. In this case, the input layer of the neural network may be formed as a layer having a dimension different from that using a one-hot vector.
[0015] The atomic feature acquisition unit 14 may generate a feature for each estimation, or as another example, may store the estimation results in the storage unit 12. For example, features may be stored in the storage unit 12 for frequently used atoms such as hydrogen atoms, carbon atoms, and oxygen atoms, and features may be generated for other atoms for each estimation.
[0016] When the input atomic coordinates, boundary conditions, and atomic features generated by the atomic feature acquisition unit 14 or features for distinguishing similar atoms are input, the input information construction unit 16 converts the structure of the molecule, etc. into a graph format and makes it compatible with the input of a network that processes graphs provided in the structural feature extraction unit 18.
[0017] The structural feature extraction unit 18 extracts structural features from the graph information generated by the input information configuration unit 16. This structural feature extraction unit 18 includes a graph-based neural network such as a GNN (Graph Neural Network) or a GCN (Graph Convolutional Network).
[0018] The physical property value prediction unit 20 predicts and outputs physical property values from the structural features of the estimation target, such as molecules, extracted by the structural feature extraction unit 18. This physical property value prediction unit 20 includes a neural network, such as an MLP. The characteristics of the neural network provided may differ depending on the physical property value to be obtained. For this reason, a plurality of different neural networks may be prepared, and one of them may be selected according to the physical property value to be obtained.
[0019] The output unit 22 outputs the estimated physical property values. Here, output is a concept that includes both outputting to the outside of the estimation device 1 via an interface and outputting to the inside of the estimation device 1, such as the storage unit 12.
[0020] Each component will be described in more detail.
[0021] (Atom Feature Acquisition Unit 14) As described above, the atomic feature acquisition unit 14 includes a neural network that outputs a vector in a latent space when a one-hot vector representing an atom is input. The one-hot vector representing an atom is, for example, a one-hot vector representing information about atomic nuclei. More specifically, it is, for example, a one-hot vector obtained by converting the number of protons, the number of neutrons, and the number of electrons. For example, by inputting the number of protons and the number of neutrons, it is also possible to acquire features of isotopes. For example, by inputting the number of protons and the number of electrons, it is also possible to acquire features of ions.
[0022] The input data may include information other than the above. For example, information such as atomic number, group, period, block, and half-life between isotopes in the periodic table may be provided as an input in addition to the one-hot vector. Also, the one-hot vector and another input may be combined as a one-hot vector in the atomic feature acquisition unit 14. For example, a discrete value may be stored in a one-hot vector, and a quantity (scalar, vector, tensor, etc.) representing that quantity may be added as the input.
[0023] The one-hot vector may be generated separately by the user. As another example, a one-hot vector generation unit may be provided separately that receives input such as an atomic name, atomic number, or other ID indicating an atom, and generates a one-hot vector by referencing a database or the like from this information in the atomic feature acquisition unit 14. Note that when continuous values are also given as input, an input vector generation unit may be further provided that generates a vector separate from the one-hot vector.
[0024] The neural network (first network) provided in the atomic feature acquisition unit 14 may be, for example, the encoder portion of a model trained by a neural network that forms an encoder and a decoder. The encoder and decoder may be configured, for example, by a Variational Encoder Decoder that imparts variance to the encoder output, similar to a VAE (Variational Autoencoder). An example in which a Variational Encoder Decoder is used will be described below, but the network is not limited to a Variational Encoder Decoder, and any model such as a neural network that can appropriately acquire vectors in a latent space for atomic features, i.e., feature quantities.
[0025] 2 is a diagram showing the concept of the atomic feature acquisition unit 14. The atomic feature acquisition unit 14 includes, for example, a one-hot vector generation unit 140 and an encoder 142. The encoder 142 and a decoder described later are part of the network using the above-mentioned Variational Encoder Decoder. Although the encoder 142 is shown, other networks, computing units, etc. for outputting features may be inserted after the encoder 142.
[0026] The one-hot vector generation unit 140 generates a one-hot vector from a variable indicating an atom. When a value to be converted into a one-hot vector, such as the number of protons, is input, the one-hot vector generation unit 140 generates a one-hot vector using the input data.
[0027] When the input data is an indirect value such as an atomic number or an atomic name, the one-hot vector generation unit 140 generates a one-hot vector by acquiring a value such as the number of protons from, for example, a database inside or outside the estimation device 1. In this way, the one-hot vector generation unit 140 performs appropriate processing based on the input data.
[0028] In this way, when input information to be converted into a one-hot vector is directly input, the one-hot vector generation unit 140 converts each of the variables into a format suitable for a one-hot vector and generates a one-hot vector. On the other hand, when only the atomic number is input, the one-hot vector generation unit 140 may automatically obtain data required for conversion into a one-hot vector from the input data and generate a one-hot vector based on the obtained data.
[0029] Although the above description has been given using one-hot vectors for input, this is merely an example, and the present embodiment is not limited to this. For example, it is also possible to input vectors, matrices, tensors, etc. that do not use one-hot vectors.
[0030] If a one-hot vector is stored in the memory unit 12, it may be obtained from the memory unit 12, or if a user prepares a one-hot vector separately and inputs it to the estimation device 1, the one-hot vector generation unit 140 is not a required component.
[0031] The one-hot vector is input to the encoder 142. The encoder 142 generates a vector z μ and the vector z μ A vector σ indicating the variance of 2 The output is a vector z sampled from this output result. For example, during training, this vector z μ Reconstructing atomic features from
[0032] The atomic feature acquisition unit 14 calculates the generated vector z μis output to the input information construction unit 16. It is also possible to use the reparametrization trick, which is used as one method of VAE. In this case, the vector z may be obtained as follows using a vector ε of random values. The symbol odot (a dot in a circle) indicates the product of each element of a vector.
number
[0033] As another example, z with no variance may be output as an atomic feature.
[0034] As will be described later, the first network is trained as a network that includes an encoder that extracts features when an atomic one-hot vector or the like is input, and a decoder that outputs physical property values from the features. By using an appropriately trained atomic feature acquisition unit 14, it becomes possible for the network to extract information necessary for predicting physical property values of molecules, etc., without the user having to select it.
[0035] Using such an encoder and decoder is advantageous compared to directly inputting physical property values, as it allows for more information to be utilized even if the physical property values required for all atoms are unknown. Furthermore, because the mapping is performed within a continuous latent space, atoms with similar properties are mapped closer to each other in the latent space, while atoms with different properties are mapped farther apart, allowing for interpolation between atoms. Therefore, it is possible to output results by interpolating between atoms without including all atoms in the training data, and it is possible to generate features that can output highly accurate physical property values even when the training data for some atoms is insufficient.
[0036] In this way, the atomic feature acquisition unit 14 is configured to include, for example, a neural network (first network) capable of extracting features that can decode the physical property values of each atom. 2It is also possible to convert the dimension of a one-hot vector of the order of 1 to 16 into a feature vector of about 16 dimensions. In this way, the first network is configured with a neural network whose output dimension is smaller than the input dimension.
[0037] (Input information configuration unit 16) The input information construction unit 16 generates a graph relating to the atomic arrangement and connections in a molecule, etc., based on the input data and the data generated by the atomic feature acquisition unit 14. The input information construction unit 16 considers the boundary conditions along with the structure of the input molecule, etc., determines whether there are adjacent atoms, and if there are adjacent atoms, determines their coordinates.
[0038] For example, in the case of a single molecule, the input information construction unit 16 generates a graph using the atomic coordinates indicated in the input as neighboring atoms. In the case of a crystal, for example, the coordinates of atoms within the unit lattice are determined from the input atomic coordinates, and for atoms located on the periphery of the unit lattice, the coordinates of outer neighboring atoms are determined from the repeating pattern of the unit lattice. In the case of an interface in the crystal, for example, neighboring atoms are determined without applying the repeating pattern to the interface side.
[0039] 3 is a diagram showing an example of coordinate setting according to this embodiment. For example, when generating a graph of only molecule M, the graph is generated from the types of the three atoms that make up molecule M and their relative coordinates.
[0040] For example, when generating a graph of only a crystal that has repetition and has an interface I, the graph is generated assuming the crystal's unit lattice C as a repetition C1 to the right, a repetition C2 to the left, a repetition C3 to the bottom, a repetition C4 to the bottom left, a repetition C5 to the bottom right, and so on, assuming neighboring atoms for each atom. In the figure, the dotted line indicates the interface I, the unit lattice indicated by the dashed line indicates the input crystal structure, and the area indicated by the dashed-dotted line indicates the area assuming the repetition of the crystal's unit lattice C. In other words, the graph is generated assuming neighboring atoms for each atom that makes up the crystal within a range that does not exceed the interface I.
[0041] When estimating the physical properties of a molecule acting on a crystal, such as in a catalyst, a graph is generated by calculating the coordinates of adjacent atoms from each atom constituting the molecule and adjacent atoms from the lattice atoms constituting the crystal, assuming a repetition that takes into account the molecule M and the interface I of the crystal described above.
[0042] Note that, since there is a limit to the size of the graph to be input, for example, the interface I, unit lattice C, and repetition of unit lattice C may be set so that molecule M is at the center. In other words, the graph may be generated by performing an appropriate number of repetitions of unit lattice C to obtain coordinates. To generate the graph, for example, the unit lattice C closest to molecule M is set as the center, and the unit lattice C is assumed to be repeated up, down, left, and right so as not to exceed the number of atoms that can be expressed in the graph within a range that does not exceed the interface, and the coordinates of each adjacent atom are obtained.
[0043] 3, one unit cell C of a crystal having an interface I for one molecule M is input, but this is not limiting. For example, there may be a plurality of molecules M, or a plurality of crystals.
[0044] The input information construction unit 16 may also calculate the distance between two atoms constructed as described above, and the angle between three atoms when a certain atom is used as a vertex. These distances and angles are calculated based on the relative coordinates of each atom. The angle is obtained, for example, using the dot product of vectors or the law of cosines. For example, calculations may be performed for all combinations of atoms, or the input information construction unit 16 may determine a cutoff radius Rc, search for other atoms within the cutoff radius Rc for each atom, and calculate the combinations of atoms present within this cutoff radius Rc.
[0045] An index may be assigned to each of the constituent atoms, and the calculated results may be stored together with the combination of the indexes in the storage unit 12. When calculating, the structural feature extraction unit 18 may read these values from the storage unit 12 at the timing when they are to be used, or the input information construction unit 16 may output these values to the structural feature extraction unit 18.
[0046] Although the diagram is shown in two dimensions for ease of understanding, molecules and the like exist in three-dimensional space, and therefore the repetition conditions may be applied to both the front and back sides of the drawing.
[0047] In this way, the input information configuration unit 16 generates a graph to be input to the neural network from the information on the input molecules and the like and the characteristics of each atom generated by the atom characteristic acquisition unit 14.
[0048] (Structural feature extraction unit 18) As described above, the structural feature extraction unit 18 of this embodiment includes a neural network that receives graph information and outputs features related to the structure of the graph. The input graph features may include angle information.
[0049] The structural feature extraction unit 18 is designed to maintain an output that is invariant to, for example, the substitution of atoms of the same kind in the input graph, and the translation and rotation of the input structure. This is because the physical properties of real materials do not depend on these quantities. For example, by defining the angles between adjacent atoms and three atoms as shown below, it is possible to input graph information that satisfies these conditions.
[0050] First, for example, the structural feature extraction unit 18 determines the maximum number of neighboring atoms Nn and the cutoff radius Rc, and acquires neighboring atoms for the atom A of interest (atom of interest). By setting the cutoff radius Rc, it is possible to eliminate atoms whose mutual influence is negligible and to prevent too many atoms from being extracted as neighboring atoms. In addition, by performing graph convolution multiple times, it is possible to incorporate the influence of atoms outside the cutoff radius.
[0051] If the number of neighboring atoms is less than the maximum number of neighboring atoms Nn, dummy atoms of the same type as atom A are randomly placed far enough away from atom A than the cutoff radius Rc. If the number of neighboring atoms is greater than the maximum number of neighboring atoms Nn, for example, Nn atoms are selected in order of closest distance from atom A and are used as neighboring atom candidates. Considering such neighboring atoms, the combination of three atoms is as follows: Nn There are C2 ways. For example, if Nn=12, 12 C2 = 66 ways.
[0052] The cutoff radius Rc is related to the interaction distance of the physical phenomenon you want to reproduce. In the case of a densely packed system such as a crystal, the cutoff radius Rc is 4 to 8 × 10 -8 Using cm makes it possible to ensure sufficient accuracy in many cases. On the other hand, when considering interactions between a crystal surface and a molecule, or between molecules, the two are not structurally connected, so repeated graph convolution cannot take into account the influence of distant atoms, and so the cutoff radius becomes the maximum direct interaction distance. Even in this case, the cutoff radius Rc is set to 8 × 10 -8 It can be applied by considering cm~ and starting the initial shape from that distance.
[0053] The maximum number of neighboring atoms Nn is set to about 12 from the viewpoint of calculation efficiency, but is not limited to this. It is also possible to consider the influence of atoms within the cutoff radius Rc that are not selected as Nn neighboring atoms by repeating graph convolution.
[0054] For one atom of interest, the input set is, for example, a concatenation of the characteristics of the atom, the characteristics of the two adjacent atoms, the distance between the atom and the two adjacent atoms, and the angle between the atom and the two adjacent atoms. The atom's characteristics are used as node characteristics, and the distance and angle are used as edge characteristics. The edge characteristics can be used as they are, or they can be processed as desired. For example, they can be binned to a specific width, or a Gaussian filter can be applied.
[0055] FIG. 4 is a diagram for explaining an example of how to obtain graph data. Consider the atom of interest as atom A. As in FIG. 3, it is shown in two dimensions, but more precisely, the atoms exist in three-dimensional space. In the following explanation, it is assumed that the candidate neighbor atoms for atom A are atoms B, C, D, E, and F. However, the number of atoms is determined by Nn, and the candidate neighbor atoms vary depending on the structure and state of the molecule, etc., so it is not limited to this. For example, if atoms G, H, etc. are also present, the following feature extraction, etc., is similarly performed within a range not exceeding Nn.
[0056] The dotted arrow pointing from atom A indicates the cutoff radius Rc. The dotted circle indicates the range from atom A to the cutoff radius Rc. The neighboring atoms of atom A are searched within this dotted circle. If the maximum number of neighboring atoms Nn is 5 or more, the five neighboring atoms of atom A are determined to be atoms B, C, D, E, and F. In this way, edge data is generated not only for atoms that are connected in the structural formula, but also for atoms that are not connected in the structural formula within the range formed by the cutoff radius Rc.
[0057] The structural feature extraction unit 18 extracts combinations of atoms to obtain angle data with atom A as the vertex. Hereinafter, a combination of atoms A, B, and C will be referred to as ABC. There are 5C2=10 combinations for atom A: ABC, ABD, ABE, ABF, ACD, ACE, ACF, ADE, ADF, and AEF. The structural feature extraction unit 18 may, for example, assign an index to each of these. The index may focus on atom A only, or may be assigned uniquely taking into account multiple atoms or all atoms. By assigning an index in this way, it becomes possible to uniquely specify the combination of the atom of interest and its adjacent atoms.
[0058] For example, suppose the index of the combination of A, B, and C is 0. Graph data in which the combination of adjacent atoms is atom B and atom C, that is, graph data with index 0, is generated for atom B and atom C, respectively.
[0059] For example, for atom A, which is the atom of interest, atom B is the first adjacent atom and atom C is the second adjacent atom. As data related to the first adjacent atoms, the structural feature extraction unit 18 combines information on the characteristics of atom A, the characteristics of atom B, the distance between atoms A and B, and the angle between atoms B, A, and C. As data related to the second adjacent atoms, it combines information on the characteristics of atom A, the characteristics of atom C, the distance between atoms A and B, and the angle between atoms C, A, and B.
[0060] The interatomic distances and angles between three atoms may be calculated by the input information composing unit 16, or if the input information composing unit 16 has not calculated them, they may be calculated by the structural feature extraction unit 18. The distances and angles can be calculated using a method similar to that described for the input information composing unit 16. Furthermore, the timing of calculation may be dynamically changed, such that if the number of atoms is greater than a predetermined number, the structural feature extraction unit 18 calculates them, and if the number of atoms is less than the predetermined number, the input information composing unit 16 calculates them. In this case, which calculation is to be performed may be determined based on the status of resources such as memory and processor.
[0061] Hereinafter, when focusing on atom A, the features of atom A will be referred to as the node features of atom A. In the above case, the data on the node features of atom A is redundant, so they may be stored together. For example, the graph data of index 0 may be configured with information on the node features of atom A, the features of atom B, the distance between atoms A and B, the angles between atoms B, A, and C, the features of atom C, the distance between atoms A and C, and the angles between atoms C, A, and B.
[0062] The distance between atoms A and B and the angles between atoms B, A, and C are collectively referred to as the edge feature of atom B, and similarly, the distance between atoms A and C and the angles between atoms C, A, and B are collectively referred to as the edge feature of atom C. Because edge features contain angle information, they are quantities that differ depending on the atoms they are paired with. For example, the edge feature of atom B when its neighboring atoms are B and C will have a different value than the edge feature of atom B when its neighboring atoms are B and D.
[0063] The structural feature extraction unit 18 generates data on all combinations of two adjacent atoms for all atoms in the same manner as the graph data for atom A described above.
[0064] FIG. 5 shows an example of graph data generated by the structural feature extraction unit 18.
[0065] For the node features of atom A, which is the first atom or atom of interest, atomic features and edge features are generated for each combination of neighboring atoms that exist within the cutoff radius Rc from atom A. The horizontal connections in the diagram may be linked by indexes, for example. Just as neighboring atoms of atom A, the first atom of interest, were selected and features were obtained, features are also obtained for the combinations of second, third, and higher neighboring atoms for atoms B, C, etc., as second, third, and higher atoms of interest, respectively.
[0066] In this way, node features, as well as atomic features and edge features related to adjacent atoms are obtained for all atoms. As a result, the feature of the target atom is a tensor of (n_site, site_dim), the feature of the adjacent atoms is a tensor of (n_site, site_dim, n_nbr_comb, 2), and the edge feature is a tensor of (n_site, edge_dim, n_nbr_comb, 2). Note that n_site is the number of atoms, site_dim is the dimension of the vector indicating the atomic feature, and n_nbr_comb is the number of combinations of adjacent atoms for the target atom (= Nn C2), edge_dim is the dimension of the edge feature. The neighboring atom feature and edge feature are obtained for each neighboring atom by selecting two neighboring atoms for the target atom, so they are tensors with twice the dimensions of (n_site, site_dim, n_nbr_comb) and (n_site, edge_dim, n_nbr_comb), respectively.
[0067] The structural feature extraction unit 18 includes a neural network that updates and outputs atomic features and edge features upon receiving this data. That is, the structural feature extraction unit 18 includes a graph data acquisition unit that acquires data related to a graph, and a neural network that updates upon receiving this graph data. This neural network includes a second network that outputs (n_site, site_dim)-dimensional node features from input data having the dimensions (n_site, site_dim + edge_dim + site_dim, n_nbr_comb, 2), and a third network that outputs (n_site, edge_dim, n_nbr_comb, 2)-dimensional edge features.
[0068] The second network comprises a network that reduces the dimension to a (n_site, site_dim, n_nbr_comb, 1)-dimensional tensor when a tensor having the characteristics of two neighboring atoms of a target atom is input, and a network that reduces the dimension to a (n_site, site_dim, 1, 1)-dimensional tensor when a tensor having the characteristics of the reduced-dimensional neighboring atoms of the target atom is input.
[0069] The first-stage network of the second network converts the features of each neighboring atom of atom A (the atom of interest), where atoms B and C are neighboring atoms, into features for the combination of neighboring atoms B and C for atom A (the atom of interest). This network makes it possible to extract the features of the combination of neighboring atoms. For atom A (the first atom of interest), all combinations of neighboring atoms are converted into these features. Furthermore, for atom B (the second atom of interest), ..., features are converted in the same way for all combinations of neighboring atoms. This network converts the tensor indicating the features of neighboring atoms from (n_site, site_dim, n_nbr_comb, 2) dimensions to (n_site, site_dim, n_nbr_comb, 1) dimensions.
[0070] The second-stage network of the second network extracts node features of atom A that incorporate the features of neighboring atoms from the combination of atoms B and C, the combination of atoms B and D, ..., and the combination of atoms E and F for atom A. This network makes it possible to extract node features that take into account the combinations of neighboring atoms for the atom of interest. Furthermore, for atoms B, ..., node features are extracted in the same way, taking into account all combinations of neighboring atoms. This network converts the output of the second-stage network from (n_site, site_dim, n_nbr_comb, 1) dimensions to (n_site, site_dim, 1, 1), which is the same dimension as the node feature.
[0071] The structural feature extraction unit 18 of this embodiment updates the node features based on the output of the second network. For example, the output of the second network and the node features are added together, and an updated node feature (hereinafter referred to as the updated node feature) is obtained through an activation function such as tanh(). This processing does not need to be provided separately from the second network in the structural feature extraction unit 18; this addition and activation function processing may be provided as an output layer of the second network. Similarly to the third network described below, the second network can eliminate information that may be unnecessary for the finally obtained physical property values.
[0072] The third network is a network that receives edge features and outputs updated edge features (hereinafter referred to as updated edge features). The third network converts a (n_site, edge_dim, n_nbr_comb, 2)-dimensional tensor into a (n_site, edge_dim, n_nbr_comb, 2)-dimensional tensor. For example, by using a gate or the like, information that is unnecessary for the physical property values that are ultimately desired to be obtained is reduced. The third network having this function is generated by training the parameters using a training device described below. In addition to the above, the third network may further include a second-stage network having the same input and output dimensions.
[0073] The structural feature extraction unit 18 of this embodiment updates the edge feature based on the output of the third network. For example, the output of the third network and the edge feature are added together, and an updated edge feature is obtained through an activation function such as tanh(). Furthermore, when multiple features are extracted for the same edge, the average value of these may be calculated to create a single edge feature. These processes do not need to be provided in the structural feature extraction unit 18 separately from the third network; the addition and activation function processes may be provided as part of the output layer of the third network.
[0074] Each of the second and third networks may be formed by a neural network that appropriately uses, for example, a convolution layer, batch normalization, pooling, gate processing, activation functions, etc. Without being limited to the above, it may be formed by an MLP, etc. Furthermore, for example, it may be a network having an input layer that can further input a tensor obtained by squaring each element of an input tensor.
[0075] As another example, the second network and the third network may be formed as a single network rather than as separate networks. In this case, when a node feature, a feature of an adjacent atom, and an edge feature are input, the second network and the third network are formed as a network that outputs updated node features and edge features according to the above example.
[0076] In this way, the structural feature extraction unit 18 generates data regarding the nodes and edges of the graph taking into account adjacent atoms based on the input information constructed by the input information construction unit 16, and updates the generated data to update the node features and edge features of each atom. The updated node features are node features that take into account adjacent atoms. The updated edge features are edge features obtained by deleting information that may be redundant information regarding the physical property values to be obtained from the generated edge features.
[0077] (Physical property prediction unit 20) As described above, the physical property prediction unit 20 of this embodiment includes a neural network (fourth network) such as an MLP that predicts and outputs a physical property when it receives input of features relating to the structure of molecules, etc., such as update node features and update edge features. The update node features and update edge features may not only be input as they are, but may also be processed in accordance with the physical property value to be obtained as described below.
[0078] The network used for this property prediction may be changed depending on the nature of the property to be predicted. For example, if energy is to be obtained, the features for each node are input to the same fourth network, the obtained output is the energy of each atom, and the sum of these is output as the total energy value.
[0079] When predicting properties between predetermined atoms, the updated edge features are input to the fourth network to predict the desired physical property value.
[0080] When predicting a physical property value determined from the entire input, the average, sum, etc. of the update node features is calculated, and this calculated value is input to the fourth network to predict the physical property value.
[0081] In this way, the fourth network may be configured as a different network for the desired physical property value, and in this case, at least one of the second network and the third network may be formed as a neural network that extracts features used to acquire the physical property value.
[0082] As another example, the fourth network may be formed as a neural network that outputs multiple physical property values at the same time, and in this case, at least one of the second network and the third network may be formed as a neural network that extracts features used to obtain the multiple physical property values.
[0083] In this way, the second network, the third network, and the fourth network may be formed as neural networks with different parameters and layer shapes depending on the physical property values to be obtained, and trained based on each physical property value.
[0084] The physical property value prediction unit 20 appropriately processes the output from the fourth network based on the desired physical property value and outputs it. For example, when calculating the total energy, if the energy for each atom is obtained by the fourth network, these energies are summed and output. In other examples, the value output by the fourth network is similarly processed appropriately for the desired physical property value, and the resulting output value is used.
[0085] The quantity output by the physical property value prediction unit 20 is output to the outside or inside of the estimation device 1 via the output unit 22.
[0086] 6 is a flowchart showing the flow of processing by the estimation device 1 according to this embodiment. The overall processing by the estimation device 1 will be described using this flowchart. Detailed descriptions of each step are given above.
[0087] First, the estimation device 1 of this embodiment accepts input of data via the input unit 10 (S100). The input information includes boundary conditions of molecules, etc., structural information of molecules, etc., and information of atoms constituting the molecules, etc. The boundary conditions of molecules, etc. and structural information of molecules, etc. may be specified by relative coordinates of atoms, for example.
[0088] Next, the atomic feature acquisition unit 14 generates features of each atom constituting the molecule, etc., from the input information on the atoms used in the molecule, etc. (S102). As described above, various atomic features may be generated in advance by the atomic feature acquisition unit 14 and stored in the storage unit 12, etc. In this case, they may be read from the storage unit 12 based on the type of atom used. The atomic feature acquisition unit 14 acquires atomic features by inputting atomic information into a trained neural network provided within itself.
[0089] Next, the input information constructing unit 16 constructs information for generating graph information of molecules, etc. from the input boundary conditions, coordinates, and atomic characteristics (S104). For example, as shown in the example of Fig. 3, the input information constructing unit 16 generates information describing the structure of molecules, etc.
[0090] Next, the structural feature extraction unit 18 extracts structural features (S106). The extraction of structural features is performed by two processes: a process for generating node features and edge features for each atom of a molecule, etc., and a process for updating the node features and edge features. The edge features include information on the angle formed by two adjacent atoms with the atom of interest as the vertex. The generated node features and edge features are extracted as updated node features and updated edge features, respectively, via a trained neural network.
[0091] Next, the physical property value prediction unit 20 predicts the physical property value from the updated node features and updated edge features (S108). The physical property value prediction unit 20 outputs information from the updated node features and updated edge features via the trained neural network, and predicts the physical property value based on this output information.
[0092] Next, the estimation device 1 outputs the estimated physical property values to the outside or inside of the estimation device 1 via the output unit 22 (S110). As a result, it becomes possible to estimate and output physical property values based on information including information on the characteristics of atoms in the latent space and angle information between adjacent atoms in a molecule or the like taking into account boundary conditions.
[0093] As described above, according to this embodiment, graph data including node features including atomic features and edge features including angle information between two adjacent atoms is used based on boundary conditions, the atomic arrangement in a molecule, etc., and the extracted atomic features, and updated node features and edge features including the features of adjacent atoms are extracted, and the physical property values are estimated using these extraction results, thereby enabling highly accurate estimation of physical property values. Because atomic features are extracted in this way, the same estimation device 1 can be easily applied even when the number of types of atoms is increased.
[0094] In this embodiment, the output is obtained by combining differentiable calculations. In other words, the output estimation result can be traced back to information on each atom. For example, when the total energy P of the input structure is estimated, the force acting on each atom can be calculated by calculating the derivative of the input coordinates for the estimated total energy P. This differentiation can be performed without any problems because a neural network is used and, as will be described later, other calculations are also performed using differentiable calculations. By obtaining the force acting on each atom in this manner, structural relaxation using this force can be performed at high speed. Furthermore, for example, it is possible to calculate energy using coordinates as input and replace DFT calculations with N-th-order automatic differentiation. Similarly, differential calculations expressed in Hamiltonians, etc., can also be easily obtained from the output of the estimation device 1, allowing for faster analysis of various physical properties.
[0095] By using this estimation device 1, for example, it becomes possible to search for materials having desired physical property values from various molecules, more specifically, molecules having various structures, molecules having various atoms, etc. For example, it is also possible to search for catalysts that are highly reactive with a certain compound.
[0096] [Training device] The training device according to this embodiment trains the aforementioned estimation device 1. In particular, the training device trains the neural networks provided in the atomic feature acquisition unit 14, the structural feature extraction unit 18, and the physical property value prediction unit 20 of the estimation device 1. In this specification, training refers to generating a model that has a structure such as a neural network and is capable of producing an appropriate output for an input.
[0097] 7 is an example of a block diagram of the training device 2 according to this embodiment. In addition to the atomic feature acquisition unit 14, input information configuration unit 16, structural feature extraction unit 18, and physical property value prediction unit 20 provided in the estimation device 1, the training device 2 also includes an error calculation unit 24 and a parameter update unit 26. The input unit 10, storage unit 12, and output unit 22 may be common to the estimation device 1 or may be unique to the training device 2. Detailed descriptions of components that are the same as those in the estimation device 1 will be omitted.
[0098] The flow indicated by the solid lines is the processing for forward propagation, and the flow indicated by the dashed lines is the processing for backward propagation.
[0099] Training data is input to the training device 2 via the input unit 10. The training data includes input data and output data that serves as teacher data.
[0100] The error calculation unit 24 calculates the error between the training data and the output from each neural network in the atomic feature acquisition unit 14, the structural feature extraction unit 18, and the physical property prediction unit 20. The method of calculating the error for each neural network is not limited to the same calculation, and may be appropriately selected based on the parameters to be updated or the network configuration.
[0101] The parameter update unit 26 back-propagates the error in each neural network and updates the parameters of the neural network based on the error calculated by the error calculation unit 24. The parameter update unit 26 may compare all neural networks with the training data, or may update the parameters for each neural network using the training data.
[0102] Each module of the above-described estimation device 1 can be formed by differentiable calculations. Therefore, it is possible to calculate gradients in the order of the physical property value prediction unit 20, the structural feature extraction unit 18, the input information configuration unit 16, and the atomic feature acquisition unit 14, and errors can be appropriately backpropagated in places other than the neural network.
[0103] For example, if you want to estimate the total energy as a physical property, (x i , y i , z i ) is the coordinate (relative coordinate) of the i-th atom, A is the atomic characteristic, and the total energy P = Σ i F i (x i , y i , z i , A i ) In this case, dP / dx i Since differential values such as these can be defined for all atoms, it becomes possible to backpropagate errors from the output to the calculation of the atomic characteristics in the input.
[0104] As another example, each module may be individually optimized. For example, the first network provided in the atom feature acquisition unit 14 may be generated by optimizing a neural network that can extract physical property values from one-hot vectors using atom identifiers and physical property values. The optimization of each network will be described below.
[0105] (Atom Feature Acquisition Unit 14) The first network of the atom feature acquisition unit 14 can be trained to output a feature value when, for example, an atom identifier or a one-hot vector is input. As described above, this neural network may utilize, for example, a variational encoder / decoder based on a VAE.
[0106] 8 shows an example of the formation of a network used to train the first network 146. For example, the first network 146 may use the encoder 142 portion of a variational encoder / decoder that includes an encoder 142 and a decoder 144.
[0107] The encoder 142 is a neural network that outputs features in the latent space for each type of atom, and is the first network used in the estimation device 1.
[0108] The decoder 144 is a neural network that outputs physical property values when it receives the vector in the latent space output by the encoder 142. In this way, by connecting the decoder 144 after the encoder 142 and performing supervised learning, it becomes possible to train the encoder 142.
[0109] As described above, one-hot vectors representing the properties of atoms are input to the first network 146. As described above, this may include a one-hot vector generation unit 140 that generates one-hot vectors when an atomic number, an atomic name, or a value indicating the properties of each atom is input.
[0110] The data used as the training data may be, for example, various physical property values. These physical property values may be obtained from, for example, a scientific chronology.
[0111] 9 is a table showing an example of physical property values. For example, the atomic properties listed in this table are used as training data for the output of the decoder 144.
[0112] In the table, values in parentheses were calculated using the method described within the parentheses. The ionic radii are calculated for the first to fourth coordinations. For example, for oxygen, the ionic radii are calculated for coordinations 2, 3, 4, and 6, respectively.
[0113] When a one-hot vector representing an atom is input to a neural network equipped with encoder 142 and decoder 144 shown in Figure 8, optimization is performed so that the properties shown in Figure 9 are output. In this optimization, error calculation unit 24 calculates the loss between the output value and training data, and parameter update unit 26 performs backpropagation based on this loss, finds a gradient, and updates the parameters. By performing optimization, encoder 142 functions as a network that outputs a vector in a latent space from the one-hot vector, and decoder 144 functions as a network that outputs a physical property value from the vector in this latent space.
[0114] The parameters are updated using, for example, a variational encoder / decoder. As described above, the reparameterization trick may also be used.
[0115] After the optimization is completed, the neural network that forms the encoder 142 is designated as the first network 146, and the parameters for this encoder 142 are obtained. The output value is, for example, z μ It can be a vector of variance σ 2 Alternatively, the value may be one that takes into account z μ and σ 2 and outputs both z μ and σ 2 When random numbers are used, a fixed random number table may be used, for example, so that the process can be backpropagated.
[0116] The physical property values of atoms shown in the table of FIG. 9 are just an example, and it is not necessary to use all of these physical property values, and physical property values other than those shown in this table may also be used.
[0117] When using various physical property values, certain physical property values may not exist depending on the type of atom. For example, for a hydrogen atom, the second ionization energy does not exist. In such a case, for example, network optimization may be performed assuming that this value does not exist. In this way, even if there is a value that does not exist, it is possible to generate a neural network that outputs physical property values. In this way, even if it is not possible to input all physical property values, the atomic feature acquisition unit 14 according to this embodiment can generate atomic features.
[0118] Furthermore, by generating the first network 146 in this way, one-hot vectors are mapped into a continuous space, so atoms with similar properties are located close to each other in the latent space, while atoms with significantly different properties are located farther apart in the latent space. Therefore, for atoms in between, even if their properties are not present in the training data, interpolation can be used to output results. Furthermore, even if there is insufficient training data for some atoms, it is possible to estimate their features.
[0119] The atomic feature vectors extracted in this way can also be input to the estimation device 1. Even if the amount of learning data for some atoms is insufficient or missing during training of the estimation device 1, estimation can be performed by interpolating interatomic features. This also makes it possible to reduce the amount of data required for training.
[0120] Figure 10 shows some examples of features extracted by this encoder 142 decoded by the decoder 144. The solid lines indicate the values of the training data, and the output values of the decoder 144 are shown with a variance for the atomic number. The variance indicates the output values input to the decoder 144 with a variance for the feature vector based on the features and variance output by the encoder 142.
[0121] From top to bottom, the diagram shows examples of covalent radii using Pyykko's method, van der Waals radii using UFF, and second ionization energies. The horizontal axis is atomic number, and the vertical axis is in the appropriate units for each.
[0122] The graph of the sharing radius shows that good values are output for the training data.
[0123] It can be seen that good values are output for the van der Waals radius and second ionization energy relative to the training data. The values deviate when the atomic number exceeds 100, but this is because the values are not currently available as training data, and training is being carried out without any training data. This results in a large variance in the data, but a certain degree of value is output. Furthermore, as mentioned above, the second ionization energy of the hydrogen atom does not exist, but it can be seen that it is output as an interpolated value.
[0124] In this way, it can be seen that by using training data for the output of the decoder 144, the encoder 142 can accurately acquire features in the latent space.
[0125] (Structural feature extraction unit 18) Next, the training of the second and third networks of the structural feature extraction unit 18 will be described.
[0126] 11 is a diagram extracting a portion related to the neural network of the structural feature extraction unit 18. The structural feature extraction unit 18 of this embodiment includes a graph data extraction unit 180, a second network 182, and a third network 184.
[0127] The graph data extraction unit 180 extracts graph data such as node features and edge features from input data on the structure of molecules, etc. This extraction does not require training if it is performed using a rule-based method that allows inversion.
[0128] However, a neural network may also be used to extract graph data, and in this case, it is possible to train the second network 182, the third network 184, and the fourth network of the physical property value prediction unit 20 together as a network.
[0129] When the second network 182 receives the feature (node feature) of the atom of interest and the feature of the neighboring atom output by the graph data extraction unit 180, it updates and outputs the node feature. This updating may be performed, for example, by converting from an (n_site, site_dim, n_nbr_comb, 2)-dimensional tensor to an (n_site, site_dim, n_nbr_comb, 1)-dimensional tensor by sequentially applying a convolutional layer, batch normalization, dividing the data into gates and other data and an activation function, pooling, and batch normalization, then converting from an (n_site, site_dim, n_nbr_comb, 1)-dimensional tensor to an (n_site, site_dim, 1, 1)-dimensional tensor by sequentially applying a convolutional layer, batch normalization, dividing the data into gates and other data and an activation function, pooling, and batch normalization, and finally calculating the sum of the input node feature and this output to update the node feature via the activation function.
[0130] When the third network 184 receives the features of neighboring atoms and the edge features output by the graph data extraction unit 180, it updates and outputs the edge features. This updating may be performed, for example, by a neural network that sequentially applies a convolutional layer, batch normalization, and activation functions, pooling, and batch normalization to separate the data into gates and other data, then sequentially applies a convolutional layer, batch normalization, and activation functions, pooling, and batch normalization to separate the data into gates and other data, and finally calculates the sum of the input edge feature and this output to update the edge feature via the activation function. Regarding the edge feature, for example, a tensor with the same dimensions (n_site, site_dim, n_nbr_comb, 2) as the input is output.
[0131] In a neural network formed in this way, the processing in each layer is differentiable, so that backpropagation of errors can be performed from the output to the input. Note that the above-described network configuration is shown as an example, and the present invention is not limited to this. Any configuration may be used as long as the node features can be updated to appropriately reflect the features of adjacent atoms, and the operations in each layer are substantially differentiable. "Substantially differentiable" refers to cases where the operations are approximately differentiable in addition to cases where the operations are differentiable.
[0132] The error calculation unit 24 calculates an error based on the update node features backpropagated by the parameter update unit 26 from the physical property value prediction unit 20 and the update node features output by the second network 182. The parameter update unit 26 updates the parameters of the second network 182 using this error.
[0133] Similarly, the error calculation unit 24 calculates an error based on the updated edge feature backpropagated by the parameter update unit 26 from the physical property value prediction unit 20 and the updated edge feature output by the third network 184. The parameter update unit 26 updates the parameters of the third network 184 using this error.
[0134] In this way, the neural network provided in the structural feature extraction unit 18 is trained together with the training of the parameters of the neural network provided in the physical property value prediction unit 20 .
[0135] (Physical property prediction unit 20) The fourth network included in the physical property value prediction unit 20 outputs a physical property value when it receives the updated node feature and the updated edge feature output by the structural feature extraction unit 18. The fourth network has a structure such as an MLP, for example.
[0136] The fourth network can be trained using the same method as that used for training a normal MLP, etc. Examples of the loss used include the mean absolute error (MAE) and the mean square error (MSE). As described above, the error is backpropagated to the input of the structural feature extraction unit 18, thereby training the second network, the third network, and the fourth network.
[0137] The fourth network may have different forms depending on the physical property values to be obtained (output). That is, the output values of the second network, the third network, and the fourth network may be different based on the physical property values to be obtained. Therefore, the fourth network may be formed in a form that can be appropriately obtained based on the physical property values to be obtained, and may be trained.
[0138] In this case, the parameters of the second and third networks may be set to initial values that have already been trained or optimized to obtain other physical property values. Also, multiple physical property values to be output as the fourth network may be set, and in this case, multiple physical property values may be simultaneously used as training data to perform training.
[0139] As another example, the first network may also be trained by backpropagating to the atomic feature acquisition unit 14. Furthermore, instead of training the first network in combination with other networks from the beginning of training to the fourth network, transfer learning may be performed by training the first network in advance using the above-described training method for the atomic feature acquisition unit 14 (for example, a variational encoder / decoder using the reparametrization trick), and then backpropagating from the fourth network to the third network, the second network, and then to the first network. This makes it possible to easily obtain an estimation device that can obtain desired estimation results.
[0140] The estimation device 1 equipped with the neural network thus obtained is capable of backpropagation from output to input. In other words, it is possible to differentiate the output data with respect to the input variables. This makes it possible to know, for example, how the physical property values output by the fourth network change by changing the coordinates of the input atoms. For example, if the output physical property value is potential, the positional derivative is the force acting on each atom. This can also be used to perform optimization to minimize the energy of the input structure to be estimated.
[0141] Although the details of each neural network described above are trained as described above, any commonly known training method may be used for overall training. For example, any appropriate learning method, such as loss function, batch normalization, training termination condition, activation function, optimization method, batch learning, mini-batch learning, or online learning, may be used.
[0142] FIG. 12 is a flowchart illustrating the overall training process.
[0143] The training device 2 first trains the first network (S200).
[0144] Next, the training device 2 trains the second network, the third network, and the fourth network (S210). At this timing, as described above, the first network may also be trained.
[0145] When training is completed, the training device 2 outputs the parameters of each trained network via the output unit 22. Here, the output of parameters is a concept that encompasses outputting parameters to the outside of the training device 2 as well as outputting parameters to the inside, such as storing parameters in the memory unit 12 within the training device 2.
[0146] FIG. 13 is a flowchart showing the process of training the first network (S200 in FIG. 12).
[0147] First, the training device 2 receives input of data to be used for training via the input unit 10 (S2000). The input data is stored, for example, in the storage unit 12 as needed. The data required for training the first network are vectors corresponding to atoms, information required to generate one-hot vectors in this embodiment, and quantities indicating the properties of the atoms corresponding to the atoms (for example, the amount of substance of the atoms). The quantities indicating the properties of the atoms are, for example, those shown in FIG. 9. Alternatively, the one-hot vectors corresponding to the atoms themselves may be input.
[0148] Next, the training device 2 generates one-hot vectors (S2002). This process is not required if one-hot vectors are input in S2000. In other cases, one-hot vectors corresponding to atoms are generated based on information that is converted into one-hot vectors, such as the number of protons.
[0149] Next, the training device 2 forward propagates the generated or input one-hot vector to the neural network shown in Fig. 8 (S2004). The one-hot vector corresponding to the atom is converted into a physical property value via the encoder 142 and the decoder 144.
[0150] Next, the error calculation unit 24 calculates the error between the physical property value output from the decoder 144 and the physical property value acquired from the scientific chronology or the like (S2006).
[0151] Next, the parameter update unit 26 backpropagates the calculated error and updates the parameters (S2008). The error backpropagation is performed up to the one-hot vector, that is, the input of the encoder.
[0152] Next, the parameter update unit 26 determines whether the training has ended (S2010). This determination is made based on a predetermined training end condition, such as completion of a predetermined number of epochs, ensuring a predetermined accuracy, etc. Note that the training may be batch learning or mini-batch learning, but is not limited to these.
[0153] If the training is not completed (S2010: NO), the processes from S2004 to S2008 are repeated. In the case of mini-batch learning, the data to be used may be changed and the processes may be repeated.
[0154] If the training is completed (S2010: YES), the training device 2 outputs the parameters via the output unit 22 (S2012) and ends the process. Note that the output may be only the parameters related to the encoder 142, i.e., the parameters related to the first network 146, or the parameters of the decoder 144 may also be output. 2 It is converted from a one-hot vector with dimensions of order 1 to a vector representing features in a latent space of, for example, 16 dimensions.
[0155] FIG. 14 shows the results of estimating the energy of a molecule, etc., by the structural feature extraction unit 18 and the physical property prediction unit 20 trained using the output of the first network according to this embodiment as input, and the results of estimating the energy of the same molecule, etc., by the structural feature extraction unit 18 and the physical property prediction unit 20 according to this embodiment trained using the output related to atomic features of a comparative example (CGCNN: Crystal Graph Convolutional Networks, https: / / arxiv.org / abs / 1710.10324v2) as input.
[0156] The graph on the left is for a comparative example, and the graph on the right is for the first network of this embodiment. In these graphs, the horizontal axis shows the values obtained by DFT, and the vertical axis shows the values estimated by each method. In other words, ideally, all values would be on the diagonal line from the bottom left to the top right, and the greater the variation, the lower the accuracy.
[0157] These figures show that, compared to the comparative example, the deviation from the diagonal line is smaller, and more accurate physical property values are output, i.e., more accurate atomic features (vectors in the latent space) can be obtained. The MAE for each is 0.031 for this embodiment and 0.045 for the comparative example.
[0158] Next, an example of the process for training the second to fourth networks will be described. Fig. 15 is a flowchart showing an example of the process for training the second, third, and fourth networks (S210 in Fig. 12).
[0159] First, the training device 2 acquires the characteristics of the atoms (S2100). This acquisition may be performed each time using the first network, or the characteristics of each atom estimated by the first network may be stored in advance in the storage unit 12, and this data may be read out.
[0160] Next, the training device 2 converts the atomic features into graph data via the graph data extraction unit 180 of the structural feature extraction unit 18, and inputs this graph data to the second network and the third network. The updated node features and updated edge features obtained by forward propagation are processed if necessary and input to the fourth network, thereby forward propagating the fourth network (S2102).
[0161] Next, the error calculation unit 24 calculates the error between the output of the fourth network and the teacher data (S2104).
[0162] Next, the parameter update unit 26 back-propagates the error calculated by the error calculation unit 24 to update the parameters (S2106).
[0163] Next, the parameter update unit 26 determines whether the training has ended (S2108), and if it has not ended (S2108: NO), it repeats the processes from S2102 to S2106, and if it has ended, it outputs the optimized parameters (S2110) and ends the process.
[0164] When training the first network using transfer learning, the process of FIG. 15 is performed after the process of FIG. 13. When performing the process of FIG. 15, the data acquired in S2100 is one-hot vector data. Then, in S2102, forward propagation is performed through the first network, the second network, the third network, and the fourth network. Necessary processes, for example, processes executed by the input information configuration unit 16, are also appropriately performed. Then, the processes of S2104 and S2106 are performed to optimize the parameters. The one-hot vector and the back-propagated error are used for updating on the input side. In this way, by re-training the first network, it is also possible to optimize the latent space vector acquired in the first network based on the physical property value that is ultimately desired to be acquired.
[0165] 16 shows an example of values estimated by this embodiment and the above-mentioned comparative example for several physical properties. The left side shows the comparative example, and the right side shows the present embodiment. The horizontal and vertical axes are the same as those in FIG. 14.
[0166] As can be seen from this figure, the dispersion of values is smaller in this embodiment than in the comparative example, and it is clear that the physical property values can be estimated in a manner close to the results of DFT.
[0167] As described above, the training device 2 according to this embodiment can acquire the characteristics of atomic properties (physical property values) as a low-dimensional vector, and furthermore, by converting the acquired atomic characteristics into graph data containing angle information and inputting it into a neural network, it is possible to perform highly accurate estimation of the physical property values of molecules, etc., by machine learning.
[0168] In this training, the architecture for feature extraction and property prediction is common, so the amount of training data can be reduced when increasing the number of atomic types.In addition, since the input data only needs to include atomic coordinates and the coordinates of adjacent atoms of each atom, it can be applied to various forms such as molecules and crystals.
[0169] The estimation device 1 trained by such a training device 2 can quickly estimate physical property values such as energy of a system that receives as input any atomic arrangement, such as a molecule, a crystal, two molecules, a molecule and a crystal, or a crystal interface. Furthermore, since these physical property values can be position-differentiated, it becomes possible to easily calculate the force acting on each atom. For example, while energy has previously required an enormous amount of calculation time to calculate various physical property values using first-principles calculations, it is now possible to quickly calculate this energy by forward propagating the trained network.
[0170] As a result, for example, it is possible to optimize the structure to minimize the energy, and by linking with a simulation tool, it is possible to speed up the calculation of the properties of various substances based on this energy and differentiated forces. Furthermore, for example, for molecules with a changed atomic arrangement, it is possible to quickly estimate the energy by simply changing the input coordinates and inputting them into the estimation device 1, without having to perform complex energy calculations again. As a result, it is easy to search for materials over a wide range using simulations.
[0171] A part or all of each device (estimation device 1 or training device 2) in the above-described embodiments may be configured as hardware, or may be configured as software (program) information processing executed by a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), etc. In the case of software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored on a non-transitory storage medium (non-transitory computer-readable medium) such as a flexible disk, a CD-ROM (Compact Disc-Read Only Memory), or a USB (Universal Serial Bus) memory, and the software information processing may be executed by reading the software. Alternatively, the software may be downloaded via a communication network. Furthermore, the software may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), thereby executing the information processing by hardware.
[0172] The type of storage medium that stores the software is not limited. The storage medium is not limited to removable media such as magnetic disks or optical disks, but may be fixed storage media such as hard disks or memory. The storage medium may be provided inside the computer or outside the computer.
[0173] 17 is a block diagram showing an example of the hardware configuration of each device (estimation device 1 or training device 2) in the above-mentioned embodiment. Each device may be realized as a computer 7 including a processor 71, a main memory device 72, an auxiliary memory device 73, a network interface 74, and a device interface 75, which are connected via a bus 76.
[0174] Although the computer 7 in FIG. 17 includes one of each component, it may also include multiple of the same component. Although FIG. 17 shows one computer 7, the software may be installed on multiple computers, and each of the multiple computers may execute the same or different parts of the software. In this case, a distributed computing configuration may be used in which the computers communicate with each other via a network interface 74 or the like to execute the processing. That is, each device (estimation device 1 or training device 2) in the above-described embodiment may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize its functions. Furthermore, the system may be configured such that information transmitted from a terminal is processed by one or more computers provided on a cloud, and the processing results are transmitted to the terminal.
[0175] The various calculations of each device (estimation device 1 or training device 2) in the above-described embodiments may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processing cores within a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. of the present disclosure may be executed by at least one of a processor and a storage device provided on a cloud that can communicate with computer 7 via a network. Thus, each device in the above-described embodiments may be implemented in the form of parallel computing using one or more computers.
[0176] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) including a computer control device and arithmetic device. The processor 71 may also be a semiconductor device including a dedicated processing circuit. The processor 71 is not limited to an electronic circuit using electronic logic elements, but may also be realized by an optical circuit using optical logic elements. The processor 71 may also include an arithmetic function based on quantum computing.
[0177] The processor 71 performs arithmetic processing based on data and software (programs) input from each device, etc. configured inside the computer 7, and can output the arithmetic results and control signals to each device, etc. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, etc.
[0178] Each device (estimation device 1 and / or training device 2) in the above-described embodiments may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or devices. When multiple electronic circuits are used, the respective electronic circuits may communicate with each other via wire or wirelessly.
[0179] The main memory device 72 is a memory device that stores instructions executed by the processor 71, various data, etc., and information stored in the main memory device 72 is read by the processor 71. The auxiliary memory device 73 is a memory device other than the main memory device 72. Note that these memory devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory device for saving various data in each device (estimation device 1 or training device 2) in the above-mentioned embodiments may be realized by the main memory device 72 or the auxiliary memory device 73, or may be realized by an internal memory built into the processor 71. For example, the memory unit 12 in the above-mentioned embodiment may be implemented in the main memory device 72 or the auxiliary memory device 73.
[0180] Multiple processors may be connected (coupled) to one storage device (memory), or a single processor may be connected. Multiple storage devices (memories) may be connected (coupled) to one processor. When each device (estimation device 1 or training device 2) in the above-described embodiments is configured with at least one storage device (memory) and multiple processors connected (coupled) to this at least one storage device (memory), a configuration may be included in which at least one of the multiple processors is connected (coupled) to at least one storage device (memory). This configuration may also be realized by storage devices (memories) and processors included in multiple computers. Furthermore, a configuration in which a storage device (memory) is integrated with a processor (for example, a cache memory including an L1 cache and an L2 cache) may be included.
[0181] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wire. The network interface 74 may be one that conforms to an existing communication standard. The network interface 74 may exchange information with an external device 9A connected via the communication network 8.
[0182] The external device 9A may include, for example, a camera, a motion capture device, an output destination device, an external sensor, or an input source device. The external device 9A may also include an external storage device (memory), for example, a network storage. The external device 9A may also be a device having some of the functions of the components of each device (the estimation device 1 or the training device 2) in the above-described embodiments. The computer 7 may receive some or all of the processing results via the communication network 8, such as a cloud service, or may transmit them to an external device.
[0183] The device interface 75 is an interface such as a USB that directly connects to the external device 9B. The external device 9B may be an external storage medium or a storage device (memory). The storage unit 12 in the above-described embodiment may be realized by the external device 9B.
[0184] The external device 9B may be an output device. The output device may be, for example, a display device for displaying images or a device for outputting audio or the like. Examples of the output device include, but are not limited to, output destination devices such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), an organic EL (Electro Luminescence) panel, a speaker, a personal computer, a tablet terminal, or a smartphone. The external device 9B may also be an input device. The input device includes a device such as a keyboard, a mouse, a touch panel, or a microphone, and provides information input by these devices to the computer 7.
[0185] In this specification (including the claims), the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d as in abcd.
[0186] In this specification (including the claims), expressions such as "using data as input / based on / according to / in response to" (including similar expressions) include cases where various data itself is used as input, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is used as input, unless otherwise specified. Furthermore, when a statement is made that a result is obtained "based on / according to / in response to data," this statement includes cases where the result is obtained based solely on the data in question, as well as cases where the result is obtained as a result of being influenced by other data, factors, conditions, and / or states other than the data in question. Furthermore, when a statement is made that "data is output," this statement includes cases where various data itself is used as output, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is output, unless otherwise specified.
[0187] In this specification (including the claims), the terms "connected" and "coupled" are intended as open-ended terms that encompass any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, physically connection / coupling, etc. The terms should be interpreted appropriately according to the context in which the terms are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in the terms without limitation.
[0188] In this specification (including the claims), the expression "A configured to B" may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.
[0189] Furthermore, in this specification (including claims), when multiple pieces of hardware of the same type perform predetermined processing, each piece of hardware may perform only a part of the predetermined processing, may perform all of the predetermined processing, or may not perform the predetermined processing in some cases. In other words, when it is stated that "one or more pieces of predetermined hardware perform a first processing, and the hardware performs a second processing," the hardware performing the first processing and the hardware performing the second processing may be the same or different.
[0190] For example, in this specification (including the claims), when multiple processors perform multiple processes, each of the multiple processors may perform only a part of the multiple processes, may perform all of the multiple processes, or in some cases may not perform any of the multiple processes.
[0191] For example, in this specification (including the claims), when multiple memories store data, each of the multiple memories may store only a portion of the data, may store the entire data, or in some cases may not store any of the data.
[0192] In this specification (including the claims), terms implying containing or possessing (e.g., "comprising," "including," "having," etc.) are intended to be open-ended terms that include containing or possessing things other than the object designated by the object of the term. When the object of such a term implies no quantity or a singular number (e.g., an article such as "a" or "an"), the expression should be construed as not being limited to a specific number.
[0193] In this specification (including the claims), even if expressions such as "one or more" or "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.
[0194] In this specification, when a particular advantage / result is described as being obtained from a particular configuration of an embodiment, it should be understood that the same advantage / result can also be obtained from one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, etc., and that the effect is not necessarily obtained by the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or states, etc. are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same or a similar configuration.
[0195] In this specification (including the claims), terms such as "maximize" include finding a global maximum, finding an approximation of a global maximum, finding a local maximum, and finding an approximation of a local maximum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these maxima probabilistically or heuristically. Similarly, terms such as "minimize" include finding a global minimum, finding an approximation of a global minimum, finding a local minimum, and finding an approximation of a local minimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these minima probabilistically or heuristically. Similarly, terms such as "optimize" include finding a global optimum, finding an approximation of a global optimum, finding a local optimum, and finding an approximation of a local optimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these optima probabilistically or heuristically.
[0196] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention derived from the content defined in the claims and their equivalents. For example, in all of the above-described embodiments, the numerical values used in the explanations are shown as examples and are not limited to these. Furthermore, the order of each operation in the embodiments is shown as an example and is not limited to these.
[0197] For example, in the above-described embodiment, the characteristic values are estimated using the characteristics of atoms, but information such as the temperature of the system, pressure, the charge of the entire system, and the spin of the entire system may also be taken into consideration. Such information may be input as a supernode connected to each node, for example. In this case, by forming a neural network that can input the supernode, it becomes possible to output energy values and the like that take into consideration information such as temperature.
[0198] (Addendum) Each of the above-described embodiments can be implemented, for example, using a program as follows. (1) When executed by one or more processors, inputting the vectors relating to atoms to a first network that extracts features of the atoms in a latent space from the vectors; Estimating atomic features in the latent space via the first network; program. (2) When executed by one or more processors, Constructing a target atomic structure based on the input atomic coordinates, atomic characteristics, and boundary conditions; Based on the structure, the distances between the atoms and the angles between the three atoms are obtained; updating the node features and the edge features using the features of the atoms as node features and the distance and the angle as edge features, and estimating the node features and the edge features; program. (3) When executed by the one or more processors, a vector indicating the properties of atoms included in a target is input to the first network according to any one of claims 1 to 7, and characteristics of the atoms in a latent space are extracted; constructing a structure of the target atoms based on the coordinates of the atoms, the extracted atomic features in the latent space, and boundary conditions; inputting the atomic features and the node features based on the structure into the second network according to any one of claims 10 to 12 to obtain the updated node features; inputting the atomic features and the edge features based on the structure into the third network according to any one of claims 13 to 16 to obtain the updated edge features; the acquired updated node features and updated edge features are input to a fourth network that estimates physical properties from node features and edge features, thereby estimating physical properties of the object; program. (4) When executed by one or more processors, inputting the vectors relating to atoms to a first network that extracts features of the atoms in a latent space from the vectors relating to atoms; inputting the characteristics of the atoms in the latent space to a decoder that outputs physical property values of the atoms when the characteristics of the atoms in the latent space are input, and estimating the property values of the atoms; the one or more processors calculate an error between the estimated atomic property value and training data; backpropagating the calculated error to update the first network and the decoder; outputting the parameters of the first network; program. (5) When executed by one or more processors, the method constructs a structure of the target atoms based on the input atomic coordinates, atomic characteristics, and boundary conditions; Based on the structure, the distances between the atoms and the angles between the three atoms are obtained; inputting information based on the atomic features, the distance, and the angle into a second network that uses the atomic features as node features to obtain updated node features, and a third network that uses the distance and the angle as edge features to obtain updated edge features; Calculating an error based on the updated node features and the updated edge features; backpropagating the calculated error to update the second network and the third network; program. (6) When executed by one or more processors, a first network extracts atomic features in a latent space from vectors related to atoms by inputting vectors indicating the properties of atoms included in the target; and extracting atomic features in the latent space. constructing a structure of the target atoms based on the coordinates of the atoms, the extracted atomic features in the latent space, and boundary conditions; Based on the structure, the distances between the atoms and the angles between the three atoms are obtained; The atomic features are used as node features to obtain updated node features; the atomic features and the node features based on the structure are input to a second network to obtain the updated node features; The distance and the angle are used as edge features to obtain updated edge features; the atomic features and the edge features based on the structure are input to a third network to obtain the updated edge features; the acquired updated node features and updated edge features are input to a fourth network that estimates physical properties from node features and edge features, and the physical properties of the object are estimated; Calculating an error from the estimated physical property value of the object and training data; back-propagating the calculated error to the fourth network, the third network, the second network, and the first network, and updating the fourth network, the third network, the second network, and the first network; program. (7) The programs described in (1) to (6) may each be stored in a non-transitory computer-readable medium, and may be configured to cause one or more processors to execute the methods described in (1) to (6) by reading one or more programs described in (1) to (6) stored in the non-transitory computer-readable medium. [Explanation of symbols]
[0199] 1: Estimation device, 10: Input section, 12: Storage part, 14: Atomic feature acquisition unit, 140: One-hot vector generation unit, 142: Encoder, 144: decoder, 146: First Network, 16: Input information configuration unit, 18: Structural feature extraction unit, 180: Graph data extraction unit, 182: Second Network, 184: Third Network, 20: Physical property prediction unit, 22: Output section, 2: Training equipment, 24: Error calculation section, 26: Parameter update section
Claims
1. at least one storage device; at least one processor; The at least one processor generating a feature of at least one atom included in the plurality of atoms by inputting information of the at least one atom into a first neural network; generating a graph based on the information of the plurality of atoms; inputting the graph into a neural network to estimate physical property values of the plurality of atoms; the at least one processor estimates the physical property value by inputting the generated features to the neural network as node features of the graph; The neural network a second neural network that receives at least the node features and outputs updated node features; a third neural network that receives at least the edge features of the graph and outputs updated edge features; a fourth neural network that receives at least the updated node feature and the updated edge feature and outputs the physical property value; Equipped with Estimation device.
2. at least one storage device; at least one processor; The at least one storage device storing the features generated by inputting the atomic information into a first neural network; The at least one processor obtaining the characteristic of at least one atom in a plurality of atoms from the at least one storage device; generating a graph based on the information of the plurality of atoms; inputting the graph into a neural network to estimate physical property values of the plurality of atoms; the at least one processor estimates the physical property value by inputting the acquired features into the neural network as node features of the graph; The neural network a second neural network that receives at least the node features and outputs updated node features; a third neural network that receives at least the edge features of the graph and outputs updated edge features; a fourth neural network that receives at least the updated node feature and the updated edge feature and outputs the physical property value; Equipped with Estimation device.
3. The first neural network is a model trained using physical property values of atoms as training data. The estimation device according to claim 1 or 2.
4. the first neural network is the encoder in a trained neural network forming an encoder and a decoder; The estimation device according to any one of claims 1 to 3.
5. The atomic information input to the first neural network includes at least one of atomic nucleus information, the number of protons, the number of neutrons, the number of electrons, the atomic number, the group, period, block, half-life between isotopes in the periodic table, the atomic name, the atomic number, and the atomic identifier. The estimation device according to any one of claims 1 to 4.
6. The features are vectors in a latent space. The estimation device according to any one of claims 1 to 5.
7. the configurations of the second neural network and the third neural network are determined based on the physical property values to be estimated; The estimation device according to any one of claims 1 to 6.
8. the configuration of the fourth neural network is determined based on the physical property value to be estimated; The estimation device according to any one of claims 1 to 7.
9. the fourth neural network outputs a plurality of the physical property values. The estimation device according to any one of claims 1 to 8.
10. The edge features include information about distances and angles between atoms. The estimation device according to any one of claims 1 to 9.
11. The at least one processor generating the edge feature based on information about a target atom included in the plurality of atoms and information about a plurality of atoms present within a predetermined range from the target atom; The estimation device according to any one of claims 1 to 10.
12. the information on the plurality of atoms includes information on the three-dimensional positional relationships of the plurality of atoms; The estimation device according to any one of claims 1 to 11.
13. The information on the three-dimensional positional relationship includes information on distances and angles between atoms. The estimation device according to claim 12.
14. The physical property value includes at least energy. The estimation device according to any one of claims 1 to 13.
15. the at least one processor calculates the force acting on each atom by positionally differentiating the energy; The estimation device according to claim 14.
16. the at least one processor calculates the force acting on each atom using backpropagation of the neural network; The estimation device according to claim 15.
17. the at least one processor performs optimization to minimize the energy of the structure using the calculated forces acting on each atom. The estimation device according to claim 16.
18. the at least one processor generates the graph based on the types of the plurality of atoms, coordinate information of the plurality of atoms, and boundary conditions. The estimation device according to any one of claims 1 to 17.
19. The neural network is constructed to generate an output that is invariant to translations and rotations of the input structure. The estimation device according to any one of claims 1 to 18.
20. the neural network includes a graph-based configuration; The estimation device according to any one of claims 1 to 19.
21. The second neural network, the third neural network, and the fourth neural network are models trained based on the same error. The estimation device according to any one of claims 1 to 20.
22. The estimation device according to any one of claims 1 to 21 estimates the physical property value. Estimation method.
23. causing the at least one processor to perform the estimation method of claim 22; program.
24. at least one storage device; at least one processor; The at least one processor generating a feature of at least one atom included in the plurality of atoms by inputting information of the at least one atom into a first neural network; generating a graph based on the information of the plurality of atoms; inputting the graph into a neural network to estimate physical property values of the plurality of atoms; Calculating an error between the estimated physical property value and training data; updating parameters of the neural network based on the error; the at least one processor estimates the physical property value by inputting the generated features to the neural network as node features of the graph; The neural network a second neural network that receives at least the node features and outputs updated node features; a third neural network that receives at least the edge features of the graph and outputs updated edge features; a fourth neural network that receives at least the updated node feature and the updated edge feature and outputs the physical property value; Equipped with the at least one processor updates parameters of the second neural network, the third neural network, and the fourth neural network based on the error. training equipment.
25. at least one storage device; at least one processor; The at least one storage device storing the features generated by inputting the atomic information into a first neural network; The at least one processor obtaining the characteristic of at least one atom in a plurality of atoms from the at least one storage device; generating a graph based on the information of the plurality of atoms; inputting the graph into a neural network to estimate physical property values of the plurality of atoms; Calculating an error between the estimated physical property value and training data; updating parameters of the neural network based on the error; the at least one processor estimates the physical property value by inputting the acquired features into the neural network as node features of the graph; The neural network a second neural network that receives at least the node features and outputs updated node features; a third neural network that receives at least the edge features of the graph and outputs updated edge features; a fourth neural network that receives at least the updated node feature and the updated edge feature and outputs the physical property value; Equipped with the at least one processor updates parameters of the second neural network, the third neural network, and the fourth neural network based on the error. training equipment.
26. The edge features include information about distances and angles between atoms.
26. A training device according to claim 24 or claim 25.
27. the at least one processor updates parameters of the first neural network based on the error. A training device according to claim 24 or claim 26 which relies on claim 24.
28. the at least one processor transfer trains the first neural network based on the error; 28. The training device of claim 27.
29. The first neural network is the encoder of a neural network forming an encoder and a decoder, The at least one processor calculating a second error between an output value obtained by inputting training atomic information into the neural network forming the encoder and decoder and second teacher data; updating parameters of neural networks forming the encoder and decoder based on the second error; A training device according to claim 25 or claim 26 which relies on claim 25.
30. The training atom information includes at least one of atomic nucleus information, the number of protons, the number of neutrons, the number of electrons, the atomic number, the group, period, block, and isotope half-life in the periodic table, the atomic name, the atomic number, and the atomic identifier.
30. The training device of claim 29.
31. The features are vectors in a latent space.
31. A training device according to any one of claims 24 to 30.
32. the information on the plurality of atoms includes information on the three-dimensional positional relationships of the plurality of atoms; 32. A training device according to any one of claims 24 to 31.
33. The information on the three-dimensional positional relationship includes information on distances and angles between atoms.
33. The training device of claim 32.
34. the neural network includes a graph-based configuration; 34. A training device according to any one of claims 24 to 33.
35. A training device according to any one of claims 24 to 34, which trains the neural network. Training method.
36. causing the at least one processor to perform the training method of claim 35; program.
37. The training device of claim 24 to claim 34, wherein the neural network is generated. Model generation method.
Citation Information
Patent Citations
Material creation device, and material creation method
JP2018010428A
Crystal form estimating device, crystal form estimating method, neural network manufacturing method, and program
JP2020166706A