Information processing method and information processing apparatus
The neural network method addresses the lack of invariance and equivariance in existing models by calculating adaptive frames for molecular and crystal structures, enhancing the prediction of physical properties and forces by considering atomic interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
Existing machine learning models for processing molecular or crystal structures fail to incorporate atomic positional and directional information while ensuring rotational and translational invariance, neglecting interactions between atoms.
A neural network approach that calculates frames representing coordinate axes at each node, extracting feature quantities with predetermined symmetry, and accounts for interactions between nodes using adaptive frames based on atomic states and interatomic interactions.
Enables calculations that consider interactions between nodes, maintaining invariance and equivariance, improving predictions of physical properties and forces in molecular and crystal structures.
Smart Images

Figure 2026060748000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an information processing method and an information processing apparatus. [Background technology]
[0002] Conventionally, techniques for processing information related to molecular or crystal structures using neural networks are known (see, for example, Non-Patent Documents 1-4). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Yi-Lun Liao, Tess Smidt, "Equiformer: Equivariant Graph Attention Transformer for 3D Atomistic Graphs", 36th Conference on Neural Information Processing Systems (NeurIPS 2022), [Retrieved September 9, 2024], Internet<URL:https: / / openreview.net / forum?id=_efamP7PSjg> [Non-Patent Document 2] Omri Puny, Matan Atzmon, Edward J. Smith, Ishan Misra, Aditya Grover, Heli Ben-Hamu, Yaron Lipman, "Frame Averaging for Invariant and Equivariant Network Design", Conference paper at ICLR 2022, [Retrieved September 9, 2024], Internet<URL:https: / / openreview.net / forum?id=zIUyj55nXR> [Non-Patent Document 3] Alexandre Duval, Victor Schmidt, Alex Hernandez Garcia, Santiago Miret, Fragkiskos D. Malliaros, Yoshua Bengio, David Rolnick, "FAENet: frame averaging equivariant GNN for materials modeling", Proceedings of the 40th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023, [searched on September 9, Reiwa 6], Internet <URL:https: / / dl.acm.org / doi / 10.5555 / 3618408.<3618769>
Non-Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0004] Non-patent documents 1 to 4 above disclose the processing of information related to the physical properties or forces of molecular or crystal structures using machine learning models such as neural network models. When processing information related to the physical properties or forces of molecular or crystal structures using machine learning models, the performance can be improved by incorporating not only distance information between atoms constituting the molecular or crystal structure, but also positional and directional information such as relative position vectors representing the relative positions between atoms within the machine learning model.
[0005] Furthermore, even if a molecular structure or crystal structure is subjected to rotational or translational operations, its physical properties remain unchanged, thus possessing invariance. In addition, when a molecular structure or crystal structure is subjected to rotational or translational operations, the forces acting on that molecular structure or crystal structure change in the same way as the rotational or translational operations, thus possessing similar invariance.
[0006] However, when rotational or translational operations are performed on molecular or crystal structures during computational processing within a machine learning model, there is no unique method for incorporating atomic positional information or directional information representing the positional relationships between atoms into the machine learning model while guaranteeing the aforementioned invariance or similar alteration.
[0007] In this regard, one of the existing approaches disclosed in the above Non-Patent Documents 2-4 is the concept of a "frame." By handling positional information of atoms or directional information representing the positional relationships between atoms on a frame, which is a coordinate system determined by certain rules for molecular or crystal structures, invariance or similar variance is considered in calculations performed within the machine learning model.
[0008] In the technologies disclosed in Non-Patent Documents 2-4 above, various calculations within the neural network model are performed using a predetermined fixed frame for each molecular or crystal structure. However, it is believed that some kind of interaction exists between the multiple atoms constituting the molecular or crystal structure. Therefore, simply using a fixed frame presents the problem that it is not possible to consider any interactions that exist between multiple atoms. Furthermore, there is no known technology that considers the interactions that exist between the nodes constituting a structure, not only for molecular or crystal structures but also for other structures, when setting frames for the nodes.
[0009] This disclosure aims to perform calculations using a neural network that take into account the interactions between nodes when performing calculations on multiple nodes that constitute a predetermined structure. [Means for solving the problem]
[0010] To achieve the above objective, the information processing method relating to this disclosure is an information processing method using a neural network for a structure represented as a set of nodes arranged in space, wherein a computer performs the following steps: receiving input of the state of each node; calculating a frame representing the coordinate axes at each node based on the states between each node; and extracting feature quantities having a predetermined symmetry of the structure from the nodes using the frame.
[0011] Furthermore, the information processing device of the present disclosure is an information processing device that uses a neural network for a structure represented as a set of nodes arranged in space, and comprises: a receiving unit that receives input of the state of each node; a calculation unit that calculates a frame representing the coordinate axes at each node based on the states between each node; and an extraction unit that extracts information having a predetermined symmetry of the structure from the nodes using the frame. [Effects of the Invention]
[0012] According to the information processing method and information processing apparatus of this disclosure, when performing calculations concerning multiple nodes constituting a predetermined structure, it is possible to perform calculations that take into account the interactions that exist between the nodes. [Brief explanation of the drawing]
[0013] [Figure 1] This is a diagram illustrating an embodiment of the present disclosure. [Figure 2] This diagram illustrates invariants and equivariants in molecular or crystalline structures. [Figure 3] This is a diagram to explain the frame. [Figure 4] This is a diagram illustrating a neural network model that processes information about molecular or crystal structure. [Figure 5] This figure illustrates the differences between the prior art and the embodiments of this disclosure. [Figure 6] This is a diagram illustrating an embodiment of the present disclosure. [Figure 7] This is a diagram to explain Transformers. [Figure 8A] This block diagram shows the hardware configuration of the trained model generation device according to this embodiment. [Figure 8B] This is a block diagram showing the schematic configuration of the trained model generation system of this embodiment. [Figure 9A] This is a block diagram showing the hardware configuration of the information processing device according to this embodiment. [Figure 9B] This is a block diagram showing the schematic configuration of the information processing system of this embodiment. [Figure 10] This figure illustrates the trained model generation processing routine executed by the trained model generation device of this embodiment. [Figure 11] This diagram illustrates the information processing routine executed by the information processing device of this embodiment. [Figure 12] This diagram illustrates the information processing routine executed by the information processing device of this embodiment. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of the present disclosure will be described with reference to the drawings. In this embodiment, an information processing device according to the present disclosure will be used as an example. In each drawing, the same or equivalent components and parts are given the same reference numerals. Also, the dimensions and proportions in the drawings are exaggerated for illustrative purposes and may differ from the actual proportions.
[0015] <Summary of this embodiment> Figure 1 is a diagram illustrating an embodiment of the present disclosure. As shown in Figure 1, we consider a case where structural information representing a molecular structure or crystal structure is input to a neural network M, which is an example of a machine learning model, and the output from the neural network M is used. In this case, the output from the neural network M is used, for example, in the following tasks. In this embodiment, the case in which the target structure is a molecular structure or a crystal structure is described as an example, but it is not limited to this. In this embodiment, the physical properties of a structure composed of atoms are inferred using information that has symmetry in the structure.
[0016] (1. Prediction of physical properties independent of coordinate system) For example, the formation energy, total energy, or band gap of a molecular or crystal structure are physical properties that do not depend on the coordinate system of the molecular or crystal structure. Therefore, when predicting invariants that are invariant with respect to rotation or translation of a molecular or crystal structure (hereinafter also simply referred to as invariants), the output from a neural network M can be used. For example, by training a neural network M so that its output is a physical property, it becomes possible to perform a task that predicts that physical property.
[0017] Furthermore, classifications of whether a molecular or crystal structure is harmful to the human body, whether a crystal structure is superconducting, or the space group of a crystal structure are classifications based on physical properties that do not depend on the coordinate system of the molecular or crystal structure (referred to as "invariant classifications" in the table described later). For this reason, when predicting these categories for the entire molecular or crystal structure or for each atom, the output from a neural network M can be used. For example, by training the neural network M so that its output is a space group category, it becomes possible to perform a task that predicts the category of that space group.
[0018] (2. Prediction of physical properties that depend on the coordinate system) For example, the forces acting on each atom constituting a molecular or crystalline structure are equivariant quantities with respect to rotation or translation (hereinafter also simply referred to as equivariates). Therefore, when predicting equivariates with respect to rotation or translation for a molecular or crystalline structure, the output from a neural network M can be used. For example, by training the neural network M so that its output is the force acting on each atom, it becomes possible to perform a task that predicts the force acting on each atom.
[0019] (3. Extracting information to be used as input for other neural networks) For example, if the neural network M shown in Figure 1 is a neural network for information extraction, the output from this neural network M can be used as input to another neural network as a state vector (or feature vector) representing the overall structure of the molecular or crystal structure or the state of each atom. Alternatively, the output from this neural network M can be used as input to another neural network as a state vector (or feature vector) representing geometric information (e.g., forces acting on atoms) about each atom constituting the overall molecular or crystal structure.
[0020] Tables 1 and 2 below show examples of invariants, invariant classifications, and equivariants in molecular or crystalline structures.
[0021] [Table 1]
[0022] [Table 2]
[0023] The "Physical Properties (or Physical Characteristics)" column in Tables 1 and 2 above represents the name of the specific physical property. The "Invariant or Equivalent" column indicates whether the physical property belongs to the category of invariant, invariant, or equivariant. The "Molecular Structure" column indicates whether the physical property is an invariant, invariant, or equivariant when the target structure is a molecular structure. The "Crystal Structure" column indicates whether the physical property is an invariant, invariant, or equivariant when the target structure is a crystal structure. The "Set of Multiple Molecular Structures or Crystal Structures" column represents a set of molecular or crystal structures, and indicates whether the physical property is an invariant, invariant, or equivariant when the target structure is a "set of multiple molecular structures," a "set of at least one molecular structure and at least one crystal structure," or a "set of multiple crystal structures." Note that the classifications of invariants, invariant classifications, and equivalents in Tables 1 and 2 above are based on general interpretations, and other classifications may also exist.
[0024] Figure 2 is a diagram illustrating invariants and equivariants in a molecular or crystal structure. For example, as shown in Figure 2, a rotational operation R and a translational operation t are performed on a molecular or crystal structure, and the position vector p of each atom is given. i is the position vector p' i The change to is represented by the following formula. Note that i is an index used to identify the atoms constituting the molecular or crystalline structure.
[0025]
number
[0026] As shown in Figure 2, when a molecular or crystal structure rotates or translates, the position vectors of each atom change, but the essential properties of the material represented by the molecular or crystal structure remain unchanged. As mentioned above, for example, physical properties such as the energy of a molecular or crystal structure have invariance, meaning they do not change even when the molecular or crystal structure rotates. On the other hand, forces acting on a molecular or crystal structure have the same invariance, meaning they change with the rotation of the molecular or crystal structure.
[0027] As mentioned above, when rotational or translational operations are performed on molecular or crystal structures during computational processing within a machine learning model, it is difficult to guarantee the invariance or similar variance described above. In this regard, as mentioned above, a technique of setting a "frame" for molecular or crystal structures is known.
[0028] Figure 3 is a diagram illustrating the concept of a frame. As shown in Figure 3, when a frame E representing a coordinate system is set for a target molecular structure or crystal structure S, the frame E rotates when the molecular structure or crystal structure S rotates, thus keeping the positional coordinates of the atoms constituting the molecular structure or crystal structure S constant. Therefore, by setting a frame for a molecular structure or crystal structure S, the orientation of the molecular structure or crystal structure S is normalized. Furthermore, when the molecular structure or crystal structure S rotates, the coordinate axes are reset according to certain rules for the molecular structure or crystal structure S. By using frames, it becomes possible to estimate invariants and equivariates.
[0029] For example, if the frame transformation due to rotation is E T (or E -1 Consider the case where a function represented by ) and realized within the neural network is represented by f, and the set of position vectors representing the position coordinates of atoms is represented by P.
[0030] In this case, it is the product of the frame transformation E due to rotation T and the set P of the atomic position vectors. T Consider passing P to the function f inside the neural network. In this case, from the function f inside the neural network, f(E T P) is output. The frame transformation E due to rotation T and the product of the position vector P T P is useful information when estimating invariants because the frame transformation E due to rotation T is performed. For example, inside the function f, the relative positions between atoms based on the coordinates E T P after the frame transformation, the relative positions from the center-of-gravity position, etc. become invariants with respect to rotation and translation, and become useful information. Also, when estimating the covariant, by multiplying the inverse transformation E of the frame by the output f(E T P) from the function f inside the neural network, Ef(E T P) is calculated. This Ef(E T P) reflects the information in the original coordinates, and thus becomes useful information when estimating the covariant (for example, force).
[0031] FIG. 4 is a diagram for explaining a neural network model that processes information related to a molecular structure or a crystal structure. Assume that a state vector sequence x representing the states of N atoms in a molecular structure or a crystal structure and a position vector sequence p representing the three-dimensional positions of the N atoms in the molecular structure or the crystal structure are represented by the following equations. In the crystal structure, although the number of atoms N is infinite, the state vector sequence x of a finite number of atoms in the unit cell, the position vector sequence p representing the three-dimensional position information of those atoms, and the lattice vectors representing the unit repetition directions in the three-dimensional space of the unit cell are used, and the expression is aggregated into the expression by the unit cell, and it is assumed that the states and positions of each atom in the entire crystal structure are represented by the periodic repetition of the unit cell.
[0032] <000065三>
Equation
[0033] In this case, the state vector representing the state of an atom is updated in each layer of the neural network. Specifically, as shown in Figure 4, for example, if the molecular structure or crystal structure is composed of sodium (Na) and chlorine (Cl), the state vectors of each of these multiple sodium (Na) atoms and each of the multiple chlorine (Cl) atoms are repeatedly updated by each layer of the neural network (labeled "Interatomic message passing" in Figure 4). Note that "AE" shown in Figure 4 is a neural network known as AtomEmbedding.
[0034] In the neural network shown in Figure 4, when the state vectors of multiple sodium (Na) and multiple chlorine (Cl) are input to the neural network, the input state vectors are transformed by AE.
[0035] Specifically, the state vector x of each of the multiple sodium Na 0 1,x 0 2,x 0 3,x 0 4 and the respective state vectors x of multiple chlorine atoms Cl 0 5,x 0 6,x 0 7,x 0 8 and are output from AE. And these state vector x 0 1,x 0 2,x 0 3,x 0 4,x 0 5,x 0 6,x 0 7,x 0 8 is input to the next layer, "Interatomic message passing", and the state vector x 0 1,x 0 2,x 0 3,x 0 4,x 0 5,x 0 6,x0 7,x 0 8 is the state vector x 1 1,x 1 2,x 1 3,x 1 4,x 1 5,x 1 6,x 1 7,x 1 It is updated to 8. These updates are repeated L times, and ultimately the state vector x L 1,x L 2,x L 3,x L 4,x L 5,x L 6,x L 7,x L It will be updated to version 8.
[0036] Thus, the state vector x i Initially, the state vector x is a vector that symbolically represents the atomic species of each atom (e.g., sodium or chlorine), but through repeated state updates in each layer of the neural network, it is gradually transformed to represent the atomic state within a given molecular or crystal structure. The update of the state vector x as described above is expressed by the following equation (A).
[0037]
number
[0038] In the above equation, X is a matrix whose column vectors are the state vectors x of each of the N atoms. Also, in the above equation, P is a matrix whose column vectors are the position vectors p of each of the N atoms. Also, in the above equation, x i ' is the updated state vector of atom i. i This is a function implemented within a neural network, where x is the state vector of atom i. i It is a function that outputs '. Also, in the above formula, f i←j This is a function implemented within a neural network, and the state vector x of the atom i of interest.i and the state vector x of other atoms j j This is a function that outputs values related to the interaction between [the two elements]. It also includes weights w. ij is a weight representing the interaction strength between the atom of interest i and other atoms j. Within the neural network, the state vectors of multiple atoms are repeatedly updated according to equation (A) above.
[0039] Note that function f i Inside, the absolute value of the relative position vector |p| represents the distance between atoms i and j. j -p i The information of | is used. The absolute value of the relative position vector |p j -p i The information represented by | is invariant under rotation and translation of molecular or crystal structures. Therefore, the function f i The absolute value of the relative position vector |p j -p i By using the information from |, it becomes possible to describe interatomic interactions in a translationally and rotationally invariant manner.
[0040] The function f in equation (A) above i←j In contrast, when a coordinate transformation using frames is introduced, the above equation (A) can be expressed by the following equation (B).
[0041]
number
[0042] E in equation (B) above T This is a matrix representing the coordinate transformation by the frame. The relative position vector E between two atoms i and j on the coordinate system represented by the frame. T p j -E T p i Since it is rotationally and translationally invariant, the function f inside the neural network model i This can effectively capture more three-dimensional positional information of atoms constituting the molecular or crystalline structure. Note that ET P is information obtained by projecting the position vector of the atom onto frame E, therefore E T P is also referred to as projected information.
[0043] Figure 5 is a diagram illustrating the differences between the prior art and embodiments of the present disclosure. As shown in T1 of Figure 5, conventionally, when processing information about molecular structure or crystal structure, a single fixed frame E fix A predefined frame E was set for the molecular structure or crystal structure. Therefore, for atoms A1, A2, A3, A4 that make up the molecular structure or crystal structure, fix It was set to [this].
[0044] However, as mentioned above, it is believed that some kind of interaction exists between the multiple atoms that make up the molecular or crystalline structure. Therefore, a problem with simply using a fixed frame is that it is not possible to consider any interactions that exist between multiple atoms. Also, depending on the task being addressed, using a fixed frame may not be appropriate in some cases.
[0045] Therefore, in the embodiments of this disclosure, as shown in T2 of Figure 5, individual frames E of atom i suitable for the atomic state or interatomic interactions resulting from the molecular structure or crystal structure. iThe frame is set. Specifically, since the interactions between atoms differ depending on the atoms constituting the molecular or crystal structure, the frame is set adaptively according to the atoms constituting the molecular or crystal structure. For example, in the embodiment of this disclosure, as shown in T2 of Figure 5, frame E1 is set for atom A1, frame E2 is set for atom A2, frame E3 is set for atom A3, and frame E4 is set for atom A4. For example, frame E1 set for atom A1 is calculated according to the weights representing the interaction strength of the state between atom A1 of interest and multiple other atoms A2, A3, A4. This makes it possible to perform calculations that take into account the interactions that exist between atoms when performing calculations on multiple atoms constituting the molecular or crystal structure.
[0046] Figure 6 is a diagram illustrating an embodiment of the present disclosure. Consider the case where the interaction function for updating the state vector of each atom is represented by equation (B) above. In the embodiment of the present disclosure, weight w represents the interaction strength regarding the state between the atom of interest i and several other atoms j1, j2, j3, j4, j5, j6 present around atom i. ij The weight w is calculated within the neural network. ij Frame E reflects this. i Set this for atom i.
[0047] Specifically, as shown in Figure 6, the atom of interest i has a weak interaction with weight w. ij We disregard surrounding atoms with small interaction strengths (for example, atoms j1 and j2 in Figure 6) and set a frame for the atom i of interest. On the other hand, as shown in Figure 6, atoms with strong interaction strengths with the atom i of interest, and weight w ij Surrounding atoms with a large weight (e.g., j3, j4 in Figure 6) are given importance, and a frame is set for the atom of interest i. Although surrounding atoms j1 and j2 in Figure 6 are located near the atom of interest i, their weight w between them and the atom of interest i is considered important. ij Because it is small, it is disregarded when calculating the frame of the atom i of interest.
[0048] In this embodiment, the weight w between the target atom i and a plurality of other atoms j ij is realized by diverting the attention weight of a known neural network, the Transformer. As a result, the weight w between the target atom i and a plurality of other atoms j ij is automatically learned inside the Transformer.
[0049] FIG. 7 is a diagram for explaining the Transformer. As shown in FIG. 7, in this embodiment, the expected value of the state vector x is calculated in the multi-head attention layer MA existing inside the known Transformer TF. The Transformer TF includes an encoder TF1 and a decoder TF2. Nx described near LA in the encoder TF1 represents that the layer LA constituting the encoder TF1 is repeated N times. By repeating this layer LA, the state vector x is gradually updated. Each element (e.g., Add&Norm, etc.) constituting the Transformer TF shown in FIG. 7 is already known, so the description thereof is omitted.
[0050] Specifically, the following formula is executed in the multi-head attention layer MA in the Transformer TF.
[0051]
Equation
[0052] In the above formulas, q i represents the query used in the Transformer, k i represents the key used in the Transformer, and v irepresents the value used in the transformer. Here, i is an index for identifying an atom. And W Q , W K , W V each represents a weight matrix used in the transformer. d K represents the dimension of the query and key, and d V represents the dimension of the value.
[0053] The weight w ij in the above formula (D) is a weight representing the interaction strength of the state between the atom i of interest and a plurality of other atoms j. Also, a ij in the above formula (D) is a coefficient used in the transformer model. k, j are indices for identifying atoms.
[0054] x’ i in the above formula (E) represents the state vector of the atom i of interest. Also, b ij in the above formula (E) is a coefficient used in the transformer model. In this embodiment, according to the following formula (F), b ij in the above formula (E) is calculated.
[0055]
Equation
[0056] As shown in the above formula (F), in this embodiment, the relative position vector p j - p i between the atom i of interest and another atom j, and the frame E i calculated for the atom i of interest (for example, the three-axis vectors e[[ID=5cribed]] i,1 , e i,2 , e i,3Projection information representing the value calculated by the product of ) is input to the function f inside the transformer model, b ij It is calculated as follows.
[0057] In this embodiment, b is calculated ij This is the frame E calculated for the atom i of interest. i The value corresponds to this. Therefore, this b ij By incorporating this into the internal calculations of the transformer, it becomes possible to perform calculations that take into account the interactions between atoms that make up the molecular or crystal structure.
[0058] <Details of this embodiment> The outline of this embodiment is as described above. Details of this embodiment will be described below. Note that matters already explained in the above outline will be explained again below.
[0059] [1.1 Systems that use molecular structure or crystal structure as input] (1. Overall System Configuration) The system of this embodiment receives structural data representing at least one molecular structure or crystal structure. The system then inputs the input information, including information derived from the molecular structure or crystal structure, into a neural network to calculate and output information such as the state of the molecular structure or crystal structure, its physical properties, or the structure updated to a stable form. In addition to the state of a single structural data, the system may also receive structural data for multiple molecular structures or multiple crystal structures, calculate the states or relationships between these materials, and output that information.
[0060] Molecular structure data includes at least information about the position and type of each atom in the structure. Furthermore, molecular structure data may also include information such as the empirical formula or the bonding state between each atomic pair. Molecular structure data typically represents the atomic arrangement in three-dimensional space, and the coordinates are given in Cartesian coordinates or oblique coordinate systems, etc.
[0061] Crystal structure data represents the periodically repeating arrangement of atoms in space and is usually expressed as lattice vectors representing the shape of a unit cell and information about the position and type of atoms within the unit cell. Alternatively, a crystal structure representation that expresses information equivalent to the lattice vectors and the position and type of atoms within the unit cell may be used as crystal structure data. The unit cell structure usually represents the arrangement of atoms in three-dimensional space, with atomic positions given as three-dimensional vectors representing Cartesian coordinates or unit cell coordinates, and lattice vectors given as three three-dimensional vectors representing unit translation quantities. However, when dealing with two-dimensional materials, the structure may be represented by the arrangement of atoms in two-dimensional space, atomic positions as two-dimensional vectors, and lattice vectors as two two-dimensional vectors representing unit translation quantities. Furthermore, information such as the empirical formula or space group may be included in addition to the structural information of the unit cell. In the following, data representing at least one of molecular structure data and crystal structure data will also be referred to as structural information.
[0062] The input information to the neural network may include information other than that derived from the input structural information, depending on the task. For example, it may include the descriptive text attached to the structural information, the name of the substance, and environmental information such as the temperature, orientation, or pressure at which the substance is placed.
[0063] The output information from the neural network can include, depending on the task, physical properties such as band gap or formation energy, the direction or magnitude of forces acting on atoms, the structure updated to a stable form, the chemical reaction state between multiple structures, or the relationship between the input description and the structure. Furthermore, one or more feature vectors abstractly representing the "input information," including structural information, may be output as input to other machine learning models such as external neural networks, SVMs, or t-SNEs. Here, the feature vectors may also include multidimensional tensor quantities.
[0064] Neural network training is performed by calculating the gradient dL / dY between the neural network output (e.g., represented by a tensor Y) and the loss function L defined according to the task, and consequently the gradient dL / dΘ between the neural network parameters Θ and the loss function L, using backpropagation. Then, the parameters Θ are optimized using optimization techniques such as stochastic gradient descent to minimize the loss function L for the training data. If, during the backpropagation process, a module within the neural network contains discrete operations and its gradient diverges or vanishes, an approximate gradient obtained by approximating that module with a continuous function may be partially used.
[0065] The gradient of the output with respect to the loss function, dL / dY, can be given in various ways. A typical example is the "supervised learning" method, where the predicted output Y of the neural network is given by the ground truth Y. * Prepare a loss function L(Y,Y) defined between the predicted value and the correct value, such as the mean squared error or cross-entropy error. * It is calculated as the gradient dL / dY of ). Note that a method called the diffusion model also adds noise N to the ground truth value. t Diffusion process Y, which involves adding steps t =Y * +N1+···N t In contrast, the noisy state Y t From a state of less Y t-1 This can be considered a supervised learning method for predicting the output Y. In the "self-supervised learning" method (sometimes considered an "unsupervised learning" method), the loss function evaluates how far the predicted output Y is from the most plausible state by using laws (such as physical laws) or properties that can be assumed for the predicted output Y. For example, when Y represents an updated structure, the instability of the structure Y can be calculated based on physical laws and given as the loss function L(Y). In addition, there is a quantity O(Y) that can be calculated from Y based on physical laws, etc. (for example, an X-ray diffraction pattern), and the correct value of O obtained from simulations or observations. *Loss function L(O(Y), * It can also be obtained as the gradient dL / dY of ). Furthermore, the relationships between multiple input structures or additional information can be used as a training signal. For example, when multiple pairs of structural information and their X-ray diffraction patterns are given as input information, a neural network can be trained that maps structural information and X-ray diffraction pattern data to the same feature space by outputting the feature vectors of both the structural information and the X-ray diffraction pattern, and defining a loss that brings the distance between the feature vectors of corresponding structures and X-ray diffraction patterns closer, while increasing the distance between the feature vectors of other structures and X-ray diffraction patterns. In addition to X-ray diffraction patterns, text information such as explanatory texts associated with the structure can also be used as information linked to the structural information. Similarly, by using category information assigned to the input structure (e.g., metallic or nonmetallic) to output feature vectors for multiple structures, a loss function can be defined that reduces the distance between feature vectors of structures belonging to the same category while increasing the distance between feature vectors of structures belonging to different categories. This allows for the learning of abstract feature vectors and feature spaces based on the categories of structures, which can then be used for searching for similar structures based on the distance between feature vectors or for visualizing the feature space using t-SNE, etc. In the "reinforcement learning" method, the input information is considered as the state s in reinforcement learning, the neural network as the policy function π, and the output Y as the deterministic action prediction a=π(s) or the stochastic action distribution π(a|s). The reward function r defined for these can then be included in the loss function as L=-r, etc., with its sign reversed. By using reinforcement learning, even if the reward function r is not differentiable with respect to a or π(a|s), a gradient that increases the reward function r (decreases the loss function) can be calculated and used as the gradient dL / dY. Furthermore, the gradient dL / dY may be calculated by a combination of multiple methods, such as supervised learning, self-supervised learning, unsupervised learning, or reinforcement learning.
[0066] (1.2 State Variable Update Operations) Within the neural network described above, we consider an operation that updates state variables related to an input structure using input information that includes information derived from the input structure. In the operation under consideration, the input structure is represented by a set of N nodes, each node being a unit that groups together at least one atom in the structure. The information that each node possesses is provided by at least a sequence of state vectors representing the state of each node and a sequence of position vectors representing the spatial coordinates of each node. The sequence of state vectors and the sequence of position vectors are expressed by the following equations.
[0067]
number
[0068] A neural network is a learnable function f that, at any point in its interior, uses at least some or all of the matrix X representing the state of the node at that moment and the matrix P representing the position of the node. i←j This results in some or all of the state vectors x in matrix X. i Update it as follows:
[0069]
number
[0070] Here, the function f i←j This calculates the effect that node j has on the state of node i as an abstract state vector, and w ij w is a scalar variable that represents its weight. ijFurthermore, the values are to be dynamically calculated in response to the input information using matrices X and P, or variables calculated elsewhere in the neural network. Within the neural network, one or more such inter-node interaction operations are included, and by combining a linear transformation Wx+b (W represents a matrix and b represents a vector) for each state vector x, a known batch normalization or known layer normalization, a known nonlinear function such as a Rectified Linear Unit (ReLU) function or a sigmoid function, and an integration process such as mean pooling or max pooling, the variables related to the input information are repeatedly updated and transformed into a predicted value Y according to the task.
[0071] (1.3 Node Configuration) A node may directly represent an individual atom within a molecular or crystal structure. In particular, the initial state vector x of a node given as input to the network. i As such, a learnable atomic embedding vector corresponding to the atomic number of atom i is used. The atomic embedding vector can be realized using known techniques and is expressed by the following equation (3). Below, ATOMICNUMBER(i) represents the atomic number of atom i, and AtomEmbedding represents the calculation of the vector using atomic embedding. AtomEmbedding can be realized using known techniques.
[0072]
number
[0073] Alternatively, the state vector may be a vector formed by concatenating embedding vectors corresponding to the group or period of atom i in the periodic table. In this case, for example, the state vector x is represented by the following equation (4). Below, GROUP(i) represents the group of atom i, and GroupEmbedding represents the calculation of the vector by embedding the group of atom i. Also, below, PERIOD(i) represents the period of atom i, and PeriodEmbedding represents the calculation of the vector by embedding the period of atom i. GroupEmbedding and PeriodEmbedding are implemented using known techniques.
[0074]
number
[0075] Alternatively, the atomic species corresponding to atom i may be represented by a mixture of multiple atomic species based on site occupancy, and the mixing ratio of each atomic species n is o n When given by , the state vector may be a linear combination of atomic embedding vectors. In this case, for example, the state vector x is expressed by the following equation (5).
[0076]
number
[0077] Furthermore, a node may represent multiple atoms in a structure. In this case, the same atom may be included in multiple nodes, and the number of atoms included in a node does not have to be constant. Within a neural network, the configuration of nodes (e.g., the number of nodes, the atom assignment to each node, or the position coordinates of each node) does not need to be constant, and the configuration of nodes may change through node merging or splitting along the way. For node merging, in addition to merging processes such as average pooling and maximum pooling, known techniques such as merging or splitting based on attention mechanisms can also be used.
[0078] For example, conventional technologies use a node configuration where nodes correspond one-to-one with atoms at the time of input to the neural network. However, they also use a network configuration that converts multiple atom nodes into "neural atom" nodes along the way, and then returns to a node configuration with individual atoms.
[0079] (1.4 State update through interaction manipulation using attention mechanisms) The interaction operation represented by equation (2) above can be realized by the attention mechanism in known transformer-type neural networks. In the attention mechanism in transformers, the state vector sequence X is transformed into a matrix called the query Q, key K, and value V. Note that the following φ Q ,φ K ,φ V This is a function used in Transformers.
[0080]
number
[0081] Note that the above matrix Q is the following vector q i It has as a column vector vector q. i This is a query vector used in Transformers.
[0082]
number
[0083] Furthermore, the above matrix K is the following vector k i It has as a column vector vector k. i This is a key vector used in Transformers.
[0084]
number
[0085] Also, the above matrix V has the following vector v i as its column vector. The vector v i is the value vector used in the transformer.
[0086]
Number
[0087] Here, N is the total number of nodes, d K represents the number of rows of matrices Q and K, and d V represents the number of rows of matrix V. Here, each transformation φ is usually a linear transformation φ Q (X) = W Q X or φ Q (X) = W Q X + b Q etc., and is given to each of the matrices Q, K, and V, but other forms may also be used. Here, W Q is the query weight matrix used in the transformer. Also, b Q is an arbitrary vector.
[0088] Then, using the matrices Q, K, and V, the interaction function between nodes is given as follows. The following Attention i is a function representing the attention mechanism of the transformer.
[0089]
Number
[0090] Here, the weight vector w i used for updating the state vector x i = [w i1 , w i2 , ···, w iN is represented by the following equation (11).
[0091]
number
[0092] In equation (11) above, u(x) is a function that takes an arbitrary real-valued vector as input. Typically, the softmax(x) function is used, which is a known function that normalizes the vector output so that each element is non-negative and the sum is 1. Alternatively, the sigmoid(x) function, which is a known function that normalizes each element of the vector output to a value between 0 and 1, or the following function u(x), which is a known function that normalizes the sum to 1, may be used.
[0093]
number
[0094] Alternatively, one may use an identity function u(x)=x that does not perform any transformations. In equation (11) above, c is a scalar coefficient, and the following values are usually used.
[0095]
number
[0096] Note that in equation (11) above, q T i The part K can be directly x without going through the matrix Q or the transformation φ(Q) or φ(V) to matrix K. T i It is sometimes represented as WK, etc.
[0097] a in equation (11) above i =[a i1 ,a i2 ,···, a iN ] T Each a ij or b in formula (10) above ijThis primarily serves as a known relative position encoding, abstractly representing information such as the relative position between two nodes i and j using scalar and vector values, respectively. These may be calculated from structural or state variables such as matrix P or matrix X, and may include learnable parameters. Furthermore, information other than matrix P or matrix X, such as interatomic bond states related to the input structure, may also be used. In addition, a ij It is also possible to manually set the value for a, for example, when the softmax function or the sigmoid function is used as the function u, ij By assigning a large negative value to a specific j, the inflow of information from j to i can be blocked. This type of masking is useful in known autoregressive networks that use neural networks to predict structure by outputting atoms one by one.
[0098] Furthermore, the operation represented by equation (10) above is a representation that encompasses both operations generally called self-attention and operations generally called cross-attention. For example, when the node configuration represents a single molecular or crystal structure, equation (10) above can be considered as a self-attention operation representing the interaction between nodes within a single structure. On the other hand, when the node configuration represents multiple structures, a ij By manipulating the value of w, the interaction weights when nodes i and j belong to the same structure can be determined. ij By forcing the equation to equal 0, equation (10) can be viewed as a mutual attention operation representing the action from one structure to another.
[0099] [2 Adaptive Frame] (2.1 Frame Overview) For each node coordinate of the input structure, an arbitrary rotation matrix R(R T Consider the case where the operation expressed in equation (12) below is performed by a rigid body transformation using an arbitrary square matrix (R=I) and a translation vector t.
[0100]
number
[0101] Alternatively, when using the following row-column expression (12-1), the rigid body transformation is represented by the following equation (13).
[0102] [Number] (12-1)
[0103] [Number]
[0104] Even when the transformation such as the above equation (12) or the above equation (13) is performed, the state update operation by the above equation (1) is independent of the presence or absence of coordinate transformation, and f i (X, P’) = f i (X, P) is a function f that remains invariant i and we want to find it. A frame is a geometric coordinate transformation technique that is useful in such cases.
[0105] A frame is inherent in the matrix P representing the structure and is an axis [e1, e2, ···, e K (represented by the matrix E ∈ R dP×K ) pointed to. In the function f i , instead of directly using the matrix P representing the input coordinates for the update of the state vector x i , by using the projection information E † P into the space by the axis E as the following equation (14), a transformation invariant to rotation can be obtained. Here, E † is usually the transpose matrix of E, and when E is a square matrix, the inverse matrix E -1 can also be used. Also, if it is a quantity that is originally invariant to rotation and translation such as distance, it can be directly calculated from the matrix P and used.
[0106] For example, when the input coordinates are rotated as P’ = RP, the axis is also rotated as E’ = RE, and E’ † P’ = (RE)[[ID=�7]] † RP = E† R T RP=E † P becomes E † It can be seen that P is rotation-invariant coordinate information.
[0107] Furthermore, if E is invariant with respect to the translation t of P, then f i - Within this, by using information based on relative position (including the distance and angle between two points), f i - This can be made rotationally and translationally invariant. Because,
[0108] JPEG2026060748000035.jpg1144
[0109] At that time,
[0110] JPEG2026060748000036.jpg1069
[0111] And so,
[0112] JPEG2026060748000037.jpg1134
[0113] This is because the components cancel each other out within their relative positions.
[0114] To define E in a translationally invariant manner, it can be determined based on the coordinates obtained by removing the mean from matrix P, or on the relative positions of two points in matrix P.
[0115] (2.2 Conventional framing methods) Conventional frame methods use a fixed (static) global frame E that is independent of the node state X and node i for each position matrix P representing the input structure.
[0116] (2.2.1 Frame averaging) The way frame E is chosen is not unique, and there is a finite set E = {E1, E2, ..., E} of which multiple frames are candidates. |E|When obtained as}, invariance can be guaranteed by calculating for all frame candidates and averaging the results, as shown in equation (15) below. This is called frame averaging. Frame averaging is a known technique (see, for example, Non-Patent Document 3).
[0117]
number
[0118] Furthermore, to reduce the computational load during training, stochastic frame averaging, which randomly selects one frame from the candidate frames and executes it during training, has also been proposed in the past (see, for example, Non-Patent Document 3).
[0119] (2.2.2 Frames using PCA for molecular structure) As a method for calculating frames, a known method involves performing principal component analysis (PCA) on a matrix P representing the atomic coordinates of the molecule to determine a frame E=[e1,e2,e3] using an orthonormal vector set (see, for example, Non-Patent Document 3).
[0120] Specifically, the dimension d of the matrix P representing the coordinates P If = 3, we calculate the 3x3 covariance matrix Σ for matrix P, and then calculate its eigenvalues λ1, λ2, λ3 and the eigenvectors e1, e2, e3 normalized to length 1. Here, even if there is no degeneracy of eigenvalues and the order of e1, e2, e3 can be uniquely arranged using the order of eigenvalues λ1>λ2>λ3, there remains arbitrariness in the sign of the eigenvectors. Therefore, the frame E=[±e1,±e2,±e3] is 2 3 = There are 8 possible candidates. If reflection is excluded from the matrix R representing the rotation to be considered, then the case where det(E)=-1 is excluded, and in this case there are 4 candidates. Regardless of the sign of det(E), if frame averaging or stochastic frame averaging based on equation (15) is performed on all 8 frame candidates, the function f iThis becomes a function invariant with respect to enantiomers of structure P. On the other hand, if we consider four frame candidates by defining det(E) as either positive or negative, then the function f i This function will distinguish between enantiomers of structure P and return different values.
[0121] (2.2.3 Frames using PCA for crystal structures) For crystal structures, a method is known in which PCA is performed on the atomic coordinates within the unit cells of the crystal structure represented as unit cells to determine a frame E=[e1,e2,e3] using an orthonormal vector set (see, for example, Non-Patent Document 2).
[0122] However, PCA has a problem where, if the matrix P representing the structure has high symmetry, eigenvalue degeneracy of the covariance matrix occurs, leaving arbitrariness of rotation with respect to the frame axes. In particular, crystal structures often have such symmetry, and in the case of cubic crystal systems in particular, arbitrariness of rotation remains on all three axes. Furthermore, the choice of unit cells in a crystal structure is generally arbitrary, and the above method has the problem of calculating different results for the same structure that is only represented by a different unit cell representation.
[0123] (2.2.4 Frames using lattice vectors for crystal structure) For crystal structures, a frame method using lattice vectors has been proposed (see, for example, Non-Patent Document 4). This method considers a linear sum of three lattice vectors l1, l2, and l3, which represent the translational directions of a unit cell, with integer coefficients expressed by the following equation (16).
[0124]
number
[0125] Then, from the combinations of coefficients n1, n2, and n3 in equation (16) above, a set of three linearly independent vectors is selected as the frame E=[e1,e2,e3] in ascending order of norm ||e||2. Note that there is no rule specifying which one to choose when there are multiple candidates for e with the same length.
[0126] (2.3 Adaptive frames linked to interactions) The frames proposed so far have been defined directly for the input structure P. On the other hand, considering the applications in which frames are used, their purpose is to effectively utilize geometric features such as the relative positions between nodes when updating the state of nodes through interactions between nodes, as represented by equation (2) above.
[0127] In this case, in the interaction model shown in equation (2) above, if we assume the weight w ij If there is a node j where x is 0, then that node is x i It does not affect the state update. Therefore, within this operation, the existence of such node j should be ignored, and when defining the frame, it is considered that defining the frame for a structure in which the coordinates of such node have been excluded from P will allow for the definition of a coordinate system that better captures the effect of this function.
[0128] Therefore, in the embodiments of this disclosure, when updating the state of each node i, the weight w of the inter-node interaction is applied. i Frame E is such that the coordinates of nodes with larger weights are given more importance, while those with smaller weights are disregarded or ignored. i This is calculated for each node i as follows:
[0129]
number
[0130] In this way, local frames E are adaptively calculated according to the interaction weights. i Using this, the state vector x iThis frame E allows for more effective extraction of the geometric features necessary for updating. i Using this, the state update operation in equation (2) above can be rewritten as follows.
[0131]
number
[0132] Local frame E i Projection information E † i By using P, rotation-invariant features can be extracted, and features that are inherently rotation-invariant, such as distance, can also be calculated and used from the input coordinates P.
[0133] (3. Methods for constructing adaptive frameworks) frame in formula (17) above (i) (w i Here are some implementation examples of the function P). The purpose of this function is to obtain a fixed number (K) axis vectors [e1, e2, ..., e K ](matrix E i The task is to determine and output one or more pairs of (represented by ). Furthermore, the following steps may be commonly used within this function.
[0134] (Non-negative weighting) weight vector w i Each component w ij This assumes a non-negative scalar value, but if given as an arbitrary real number, a transformation will be performed to make the weights non-negative. i ←φ(w i ) may be applied beforehand. As an example of a transformation, take the absolute value for each element w i ←|w ij Examples include:
[0135] (Relative position coordinates) Instead of absolute position coordinates P, use relative position coordinates P that are invariant to translational operations on P. i Convert and use. P iAs an example, it is possible to use relative position coordinates centered on node i, which are expressed by the following equations (19) to (21).
[0136]
number
[0137] Note that in the above equation, p i Instead of subtracting, the mean coordinates of P are expressed by the following equation: - You may use relative position coordinates obtained by subtracting from each node's coordinate p.
[0138]
number
[0139] Alternatively, in the above equation, p i Instead of subtracting, the weighted average coordinate p is expressed by the following formula. - wi You may use relative position coordinates obtained by subtracting from each node's coordinate p.
[0140]
number
[0141] These translation-invariant relative position coordinates are grouped together as P i Represented by E as P i By determining based on E i This is invariant under translational operations on P.
[0142] (Exclusion of target node) P i or w i You may delete the row or column corresponding to node i from the above. In particular, p i Relative position coordinates P centered on this coordinate i Now, let's consider the relative coordinate p corresponding to i. ij Since it becomes a 0 vector, it is excluded.
[0143] (Adding random noise to weights) In the frame selection algorithm, weight w i This includes operations such as selecting an axis based on the magnitude of each value. When multiple nodes have the same weight, it may not be possible to determine the magnitude relationship, therefore the weight w i You may assign a small random number to each element and then randomly order them. The random number is assigned to each w ij Both addition and multiplication are possible. Such operations are equivalent to applying stochastic frame averaging in existing studies. Furthermore, such operations on weights are also effective for the eigenvalue degeneracy problem in weighted PCA.
[0144] (Calculation of multiple frame candidates) If multiple candidate frames can be calculated based on factors such as magnitude and the arbitrariness of positive and negative signs, in addition to probabilistically selecting one frame by methods such as adding random noise to the weights mentioned above, the results of state updates from multiple candidate frames may be averaged, as shown in equation (15) above. This is equivalent to applying frame averaging, as studied in existing research.
[0145] (Order of multiple frame axes) In the frame determination method described below in the embodiments of this disclosure, the frame axis [e1,e2,···,e K A certain order is set for ], but these orders can be rearranged according to certain rules. For example, the order of the axes obtained by the following method can be reversed to [e K ,e K-1 The final output may be [,···,e1].
[0146] (3.1 Weighted PCA Frame) (Step 3.1.1 Calculation of the weighted covariance matrix) Relative position coordinates P centered on node i i We calculate the weighted covariance matrix. The weighted covariance matrix is expressed by the following formula.
[0147]
number
[0148] Note that the weighted covariance matrix Σ is used here. i 1 / Σ for the whole j w ij Multiplying by a coefficient such as 1 / N does not change the final result, so we will omit the explanation of the difference. Also, relative position coordinate P i Alternatively, normalized mutual relative position coordinates may be used, in which each relative position vector is normalized to a length of 1.
[0149] (Step 3.1.2 Obtaining Eigenvectors) Weighted covariance matrix Σ i d P The number of eigenvalues λ1, λ2, ..., λ dP The corresponding eigenvectors e1, e2, ..., e dP We obtain the following. Note that dP is P or Σ i This is the spatial dimension. Furthermore, since the length ||e||2 is arbitrary, each eigenvector e is assumed to be normalized to a constant value such as 1.
[0150] (Step 3.1.2 Rearranging Eigenvectors) The eigenvectors are rearranged according to the relative magnitudes of their eigenvalues. The order of the eigenvalues (ascending or descending) is irrelevant. Coordinate P i When there is high symmetry, the eigenvalues λ1, λ2, ..., λ dP Multiple eigenvalues with the same value may appear within the dataset. In this case, the "random noise addition to weights" described above can be used to randomly determine the order relationship between multiple eigenvalues with equal values. Also, the eigenvalues λ1, λ2, ..., λ dP Random noise can be directly added to the data to perform the ordering.
[0151] (Step 3.1.2 Frame Output) The rearranged eigenvectors e1, e2, ..., e dP Select K eigenvectors from and, while preserving their order, Ei =[±e1,±e2,···,±e K Outputs ]. The number of axes K is 1 ≤ K ≤ d P It can be set to any integer value within the range. Since the sign of each eigenvector e is arbitrary, when choosing K eigenvectors, 2 K Several frame candidates are possible, especially K=d P At that time, 2 dP There are several possible candidates, but in this case, depending on the type of rotation being targeted, det(E i The candidates can also be narrowed down by the sign of ). For example, if reflection is excluded from the target, that is, the function f in equation (18) - i←j When distinguishing between enantiomers, det(E i E such that )>0 i We only need to consider this, and in this case, 2 dP-1 It will become a street.
[0152] (3.2 Orthogonal Maximum Frame) (Step 1: Select the axis according to the maximum weight) weight vector w i Each element w ij From among them, select node j with the maximum weight, P i The relative position coordinate p corresponding to j within this coordinate system ij Select e1.
[0153] (Step 2: Selecting axes according to each weight) Below is the k-th axis e k The decision operation is repeated sequentially for k=2, 3, ..., K. (a) Select node j that maximizes the evaluation function expressed by the following formula.
[0154]
number
[0155] (b)P i The relative position coordinate p corresponding to j within this coordinate system ij Choose this option. (c) k-th axis ek Select according to the following formula.
[0156]
number
[0157] After determining each e in Step 1 and Step 2(c) above, the length may be normalized by setting e ← e / ||e||2, etc. In this case, the expression in Step 2(c) can be simplified as follows.
[0158]
number
[0159] Step 2(a) Evaluation function w' ij The function φ(p ij ,e m ) is vector p ij and e m This is a function that returns a non-negative scalar value such that it is 0 when the two points are parallel and takes its maximum value when they are orthogonal. For example, p ij and e m The angle θ between them ijm The cosine function value of cosθ ijm =(p ij / ||p ij ||2) T (e m / ||e m Using ||2) φ(p ij ,e m ) = 1 - |cosθ ijm It is given as |, etc.
[0160] (Step 3: Output of the frame) The obtained [e1,e2,···,e K ] to frame E i Output as follows. Note that the number of axes K is 1 ≤ K ≤ d P It can be set to any integer within the range K=d P and d P If =3, then the final axis e of the previous step. KThe decision is made using the cross product operation without going through steps 2(a)-(c), e K =e1 × e2 and e K You may give either -e1 × e2 as a candidate. Or, e K You may also give two possible options: ±e1 × e2.
[0161] (3.3 Non-orthogonal maximum frame) (Step 1: Select the axis according to the maximum weight) weight vector w i Each element w ij From among them, select node j with the maximum weight, P i The relative position coordinate p corresponding to j within this coordinate system ij Select e1.
[0162] (Step 2: Selecting axes according to each weight) Below is the k-th axis e k The decision operation is repeated sequentially for k=2, 3, ..., K. (a) Select node j that maximizes the evaluation function expressed by the following formula.
[0163]
number
[0164] (b)P i The relative position coordinate p corresponding to j within this coordinate system ij to e k Select as follows. After determining each e in step 1 and step 2(c) above, the length may be normalized by setting e ← e / ||e||2, etc. Also, the evaluation function w' in step 2(a) ij The function φ(p ij ,e m ) is defined similarly to the case of the orthogonal maximum frame.
[0165] (Step 3: Output of the frame) The obtained [e1,e2,···,e K ] to frame E iIt outputs as follows. Unlike weighted PCA frames and orthogonal maximum frames that determine orthogonal axes, the number of axes K can be set to any integer value of 1 or greater without upper limit.
[0166] (3.4 Greedy Frame) (Step 1: Selecting nodes based on weights) weight vector w i element w ij From among them, select K nodes j1, j2, ..., j in descending order of weight. K Choose this option.
[0167] (Step 2: Selecting an axis based on the selected node) Relative position coordinates corresponding to the selected node [p ij1 ,p ij2 ,···,p ijK ] to P i Obtained from among them, this is the frame axis [e1, e2, ..., e K Select as ] and frame E i The output is generated. Here, the length of each axis vector e may be normalized by setting e ← e / ||e||², etc. Unlike the weighted PCA frame and orthogonal maximum frame which find orthogonal axes, the number of axes K can be set to any integer value of 1 or greater without upper limit.
[0168] (3.5 Mixed frame) Multiple frame selection methods as described above may be combined. For example, M types of frame selection methods may be used, and from the mth method, K m Individual frame axis E i (m) =[e (m) 1,e (m) 2,···,e (m) Km When obtaining ], a frame in which those axes are aligned may be used. This frame is represented by the following formula.
[0169]
number
[0170] Furthermore, if there are clearly overlapping axes among the combinations, they may be consolidated into one. For example, in the three methods of orthogonal maximal frames, non-orthogonal maximal frames, and greedy frames, the weight w is commonly set as e1. ij The relative position coordinates p of node j where the value is maximized. ij The one that is chosen.
[0171] (Designing features across 4 nodes) The key to this method is that in the state update based on the interaction represented by equation (2) above, the projection information E is used as shown in equation (18) above. † i This involves performing calculations that include some or all of the information about P.
[0172] Here, frame E i Whether each axis vector e within the graph is normalized to a length of 1, or whether they are orthogonal to each other, can be arbitrarily set during the design phase.
[0173] Furthermore, the relative position coordinates P defined above i Projection information E using † i P i Projection information E † i This includes the calculation of P, because the relative position coordinates P i The reference coordinate p(p i It can be expressed as in equation (21A) using (or mean coordinates, etc.) and calculated as in equation (21b).
[0174]
number
[0175] Furthermore, projection information E † i For P, the function l(||p) is the length of each coordinate. i A diagonal matrix D with ||2) as its diagonal corners. l(P)=diag(l(||p1||2),l(||p2||2),···,l(||p N Using ||2)), E † i PD l(P) In that case, projection information E † i Let P be included. Using l(r) = 1 / r, PD l(P) This represents the normalized length of each coordinate vector p. Similarly, E † i P i For the diagonal matrix D l(Pi) When multiplied by, it can be expressed as shown in the following equation (21C), and thus the projection information E † i This includes P.
[0176]
number
[0177] When using the attention mechanism of the transformer, this information is the internode feature b in equation (10) above. ij This is used in the calculation of E. Below, these projection information E † i Specific examples of rotation-invariant and translation-invariant features containing P or P are given.
[0178] In common, a G-dimensional Gaussian basis function Γ is used as a function to convert scalar values into vectors of arbitrary dimensions. The Gaussian basis function Γ is defined by the parameters of the mean and standard deviation of G Gaussian bases (represented by the G-dimensional vectors μ and σ, respectively), and the d-th dimension component is expressed as follows.
[0179]
number
[0180] Here, μ and σ may be given as constant parameters, or they may be learned as learnable parameters along with the other parameters.
[0181] (4.1 Relative position information) Coordinates p of node i i Relative position coordinates P centered on this coordinate i And, frame E where each axis vector e is normalized to length 1. i Using and, projection information E † i P i Calculate E † i P i Projection information E corresponding to node j within † i P ij This is the relative position vector p between node i and node j. ij This is the result of transforming it into rotation-invariant frame coordinates. E is a K-dimensional vector. † i P ij The k-th component of is transformed into a vector by the Gaussian basis function Γ, and then linearly transformed and summed up, resulting in the quantity b in equation (23). ij These can be used as translation- and rotation-invariant internode features. Here, W1,···,W K These are learnable matrix parameters.
[0182]
number
[0183] (4.2 Angle Information) Coordinates p of node i i Relative position coordinates P centered on this coordinate i And a diagonal matrix D whose diagonals are the reciprocals of the lengths of each relative position. l(Pi) And, frame E where each axis vector e is normalized to length 1. i Using E † i P i D l(Pi) Calculate E † i P i D l(Pi) Projection information E corresponding to node j within † i p ij / ||p ij ||2 represents the relative position vector between node i and node j and each axis e of the frame. k This represents the value of the cosine function of the angle between the two. These are each transformed into vectors using the Gaussian basis function Γ, and then linearly transformed and summed up to obtain the quantity b in equation (24). ij These can be used as translation- and rotation-invariant internode features. Here, W1,···,W K These are learnable matrix parameters.
[0184]
number
[0185] (4.3 Distance Information) The distance between nodes is inherently invariant under translation and rotation. Therefore, the length of the relative position vector ||p| can be calculated directly from P. j -p i The quantity obtained by transforming ||2 into a vector using the Gaussian basis function Γ, and then performing a linear transformation, is b in the following equation (25). ij These can be used as translation-invariant and rotation-invariant internode features, where W is a learnable matrix parameter.
[0186]
number
[0187] (4.4 Combinations) The three types of feature quantities b mentioned above ij These can also be added together in any combination. In this case, the matrix parameter W or the Gaussian basis parameters μ and σ may be prepared individually for each type.
[0188] [Information processing device of the embodiment] Figure 8A is a block diagram showing the hardware configuration of the trained model generation device 14A according to the embodiment. As shown in Figure 8A, the trained model generation device 14A includes a CPU (Central Processing Unit) 42A, a GPU (Graphics Processing Unit) 43A, memory 44A, GPU memory 45A, storage device 46A, input / output interface 48A, storage medium reader 50A, and communication interface 52A. Each component is connected to the others so as to be able to communicate with each other via a bus 54A.
[0189] The memory device 46A stores a trained model generation program for executing the processes described later. The CPU 42A and GPU 43A, which are examples of processors, are central processing units that execute various programs and control each configuration. Specifically, the CPU 42A reads a program from the memory device 46A and executes the program using memory 44A as a workspace. The CPU 42A performs the various calculations described above according to the program stored in the memory device 46A. The GPU 43A also reads a program from the memory device 46A and executes the program using GPU memory 45A as a workspace. The GPU 45A performs the various calculations described above according to the program stored in the memory device 46A.
[0190] More specifically, GPU43A and GPU memory45A primarily efficiently parallelize various numerical calculations within the model described in the trained model generation program. GPU43A reads the GPU computation program from the program read from storage device 46A and executes the program using GPU memory45A as the workspace.
[0191] Memory 44A and GPU memory 45A are composed of RAM (Random Access Memory) and temporarily store programs and data as a working area. Storage device 46A is composed of ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), etc., and stores various programs including the operating system and various data.
[0192] The I / F48A is an interface for inputting data from and outputting data to external devices. It may also be connected to various input devices, such as keyboards and mice, and output devices, such as displays and printers, for outputting various types of information. A touch panel display may be used as an output device, allowing it to function as an input device as well.
[0193] The storage medium reader 50A reads data stored on various storage media such as CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, Blu-ray disc, and USB (Universal Serial Bus) memory, and writes data to these storage media.
[0194] Communication I / F52A is an interface for communicating with other devices, and standards such as Ethernet®, FDDI, and Wi-Fi® are used.
[0195] The trained model generation device 14A of this embodiment generates a trained model for calculating information about multiple atoms constituting a molecular structure or crystal structure using the method described above.
[0196] The functional configuration of the pre-trained model generation system 10A of this embodiment will now be described. As shown in Figure 8B, the pre-trained model generation system 10A of this embodiment includes an input device 12A, a pre-trained model generation device 14A, and an output device 16A.
[0197] The input device 12A is a device for inputting various information to the trained model generation device 14A. The output device 16A is a device for outputting various information calculated by the trained model generation device 14A.
[0198] Functionally, the trained model generation device 14A includes a learning unit 18A. Furthermore, a data storage unit 30A and a trained model storage unit 32A are provided in a predetermined memory area of the trained model generation device 14A. Each functional configuration is realized when the CPU 42A reads each program stored in the storage device 46A, expands it into memory 14A, and executes it. At this time, some computational processing programs are sent to the GPU 43A for efficient parallel processing.
[0199] First, the trained model generator 14A trains a transformer, which is an example of a neural network. In the following, at least one atom among the multiple atoms constituting the molecular or crystal structure will be treated as a single node.
[0200] The data storage unit 30A stores training data representing molecular or crystal structures. The training data represents molecular or crystal structures and is data corresponding to the machine learning algorithm used to train the transformer.
[0201] For example, when training a transformer using unsupervised machine learning or self-supervised learning, molecular structure data or crystal structure data without correct labels is stored in the data storage unit 30A as training data. When training a transformer using supervised machine learning, molecular structure data or crystal structure data with correct labels is stored in the data storage unit 30A as training data. When training a transformer using reinforcement learning, data representing simulation results, etc., is stored in the data storage unit 30A as training data.
[0202] Furthermore, the data storage unit 30A stores various types of information necessary for the trained model generation device 14A to perform predetermined information processing.
[0203] When the learning unit 18 receives an instruction signal to train the Transformer, it reads the training data stored in the data storage unit 30.
[0204] The learning unit 18 reads the training data stored in the data storage unit 30. Then, the learning unit 18 trains the Transformer using a predetermined machine learning algorithm based on the training data.
[0205] Specifically, the learning unit 18 obtains a trained transformer by training the transformer TF shown in Figure 7. Then, the learning unit 18 stores the trained transformer in the trained model storage unit 32.
[0206] Furthermore, the state vector x is updated in the learned attention mechanism within the multi-head attention layer MA1 of the encoder TF1 of the transformer TF shown in Figure 7, or within the masked multi-head attention layer MA2 of the decoder TF2. Note that each of the above processes is configured to be executable within the multi-head attention layer. Depending on the task, the encoder TF1 or decoder TF2 outputs an invariant, an invariant classification, an equivariant, or a feature. An invariant, an invariant classification, an equivariant, or a feature is an example of information about multiple atoms. Furthermore, information about multiple atoms is an example of information having a predetermined symmetry in the structure. A "predetermined symmetry" is mathematically defined as a group, and in the case of rotational and translational symmetry in N-dimensional space, it is the Euclidean group E(N) (N is a natural number, and so on), and for rotational symmetry, it is the orthogonal group O(N), etc. (if det=1, the prefix Special is added). For example, the symmetry used in this embodiment is SE(3). "Having symmetry" means that the transformations generated by the above groups are either invariant, equivariant, or both.
[0207] [Information processing device of the embodiment] Figure 9A is a block diagram showing the hardware configuration of the information processing device 14B according to this embodiment. As shown in Figure 9A, the information processing device 14B includes a CPU (Central Processing Unit) 42B, a GPU (Graphics Processing Unit) 43B, memory 44B, GPU memory 45B, storage device 46B, input / output interface 48B, storage medium reader 50B, and communication interface 52B. Each component is connected to the others so as to be able to communicate with each other via bus 54B.
[0208] The storage device 46B stores information processing programs for executing the processes described later. The CPU 42B and GPU 43B, which are examples of processors, are central processing units that execute various programs and control various components. Specifically, the CPU 42B reads a program from the storage device 46B and executes the program using memory 44B as a workspace. The CPU 42B performs the various calculations described above according to the program stored in the storage device 46B. The GPU 43B also reads a program from the storage device 46B and executes the program using GPU memory 45B as a workspace. The GPU 45B performs the various calculations described above according to the program stored in the storage device 46B.
[0209] More specifically, GPU43B and GPU memory45B primarily efficiently perform parallel processing of various numerical calculations within models described in information processing programs. GPU43B reads the GPU calculation program from the program read from storage device 46B and executes the program using GPU memory45B as the workspace.
[0210] Memory 44B and GPU memory 45B are composed of RAM (Random Access Memory) and temporarily store programs and data as a working area. Storage device 46B is composed of ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), etc., and stores various programs including the operating system and various data.
[0211] The I / F48B is an interface for inputting data from and outputting data to external devices. It may also be connected to various input devices, such as keyboards and mice, and output devices, such as displays and printers, for outputting various types of information. A touch panel display may also function as an input device by being used as an output device.
[0212] The storage medium reader 50B reads data stored on various storage media such as CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, Blu-ray disc, and USB (Universal Serial Bus) memory, and writes data to these storage media.
[0213] Communication I / F52B is an interface for communicating with other devices, and standards such as Ethernet®, FDDI, and Wi-Fi® are used.
[0214] The information processing device 14B of this embodiment calculates information about multiple atoms constituting a molecular structure or crystal structure using the method described above.
[0215] The functional configuration of the information processing system 10B of this embodiment will now be described. As shown in Figure 9B, the information processing system 10B of this embodiment includes an input device 12B, an information processing device 14B, and an output device 16B.
[0216] The input device 12B is a device for inputting various types of information to the information processing device 14B. The output device 16A is a device for outputting various types of information calculated by the information processing device 14B.
[0217] Functionally, the information processing device 14B includes a processing unit 20B. The processing unit 28B is an example of the receiving unit, calculation unit, and extraction unit of this disclosure. In addition, a data storage unit 30B and a trained model storage unit 32B are provided in a predetermined storage area of the information processing device 14B. Each functional configuration is realized by the CPU 42B reading each program stored in the storage device 46B, expanding it into memory 14B, and executing it.
[0218] When the information processing device 14B receives structural information representing a molecular structure or crystal structure, it calculates the state vector x of at least one atom constituting the molecular structure or crystal structure.
[0219] The data storage unit 30B stores various types of information necessary for the information processing device 14B to perform predetermined information processing. The trained model storage unit 32B stores trained transformers generated by the trained model generation device 14A.
[0220] When the processing unit 20B receives an instruction signal to calculate the state vector x of the nodes constituting the target molecular structure or crystal structure, it reads out the trained encoder and trained decoder, or both, from the trained transformers stored in the trained model storage unit 32B. The following explanation uses the case where the trained encoder is used to update the state vector x as an example, but is not limited to this. The state vector x may be updated using the trained decoder, or both the trained encoder and trained decoder may be used to update the state vector x.
[0221] Then, the processing unit 20B receives the state vector x of each of the multiple nodes as input to the trained encoder.
[0222] A trained encoder accepts the state of each node representing one or more atoms among the multiple atoms constituting a molecular or crystal structure. Specifically, the trained encoder accepts the initial value of the state vector x for each of the multiple nodes.
[0223] Next, the trained encoder performs the processing described above in its multi-head attention layer. Specifically, for each of the multiple nodes, the trained encoder assigns a weight w that represents the interaction strength between the state of node i of interest and the states of multiple other nodes j. ij Based on this, frame E represents the coordinate axes at node i of interest. i The system then calculates the frame E calculated for each of the nodes, and the trained encoder then performs a process to calculate information about multiple atoms.
[0224] More specifically, first, the trained encoder takes the state vector x of each of the multiple nodes. i And, a position vector p representing the spatial position of each of the multiple nodes. i We accept it.
[0225] Next, the trained encoder outputs the state vector x of each of the multiple nodes. i Based on this, weights w represent the interaction strength between the node of interest i and several other nodes j. ij A weight vector w whose components are i To calculate this, a weight vector w is calculated for each node i of interest. i The following is calculated. Specifically, the trained encoder calculates the weight vector w for each node i of interest according to equation (D) above. i This is calculated.
[0226] Next, the trained encoder calculates the state vector x of node i for each of the multiple nodes. i and the state vector x of multiple other nodes j i Weight vector w representing the interaction strength betweeni Based on this, frame E represents the coordinate axes at node i of interest. i This is calculated. As a result, frame E is set for each of the multiple nodes. The trained encoder calculates the frame for each node by performing the process described above (3. Method for constructing adaptive frames).
[0227] (Weighted PCA frame) For example, a trained encoder computes a node-by-node frame by performing the process described in (3.1 Weighted PCA Frames) above. Specifically, when computed a node-by-node frame, the trained encoder calculates the relative position vector p between the node of interest i and several other nodes j for each of the multiple nodes. j -p i Alternatively, a relative coordinate matrix P whose components are vectors whose length has been normalized to 1. i And the weight w between the node of interest i and several other nodes j. ij A weight vector w whose components are i The weight matrix diag(w) is composed of the following elements: i Based on this, by performing Principal Component Analysis (PCA), a frame E represented by one or more axis vectors is obtained. i Calculate.
[0228] (Maximum orthogonal frame) Alternatively, for example, a trained encoder may compute the frame for each node by performing the process described in (3.2 Orthogonal Maximum Frames) above.
[0229] In this case, the trained encoder may calculate mutually orthogonal axis vectors as frames by performing the following calculation process.
[0230] For example, when a trained encoder calculates a frame for each node, it uses the weights w between the node of interest i and several other nodes j. ijFrom among them, select the other node j1 corresponding to the weight with the largest value. Then, the trained encoder outputs the position vector p of the selected other node j1. j1 The node p that we are focusing on. i The relative position vector p represents the difference between and . j1 -p i The normalized vector corresponding to this is set as the first axis vector e1.
[0231] Next, the trained encoder selects another node j2 such that the evaluation function between it and the first axis vector e1 is minimized. For example, the trained encoder selects another node j2 such that the absolute value of the dot product between it and the first axis vector e1 is small and the weight is large. Then, the trained encoder calculates the position vector p of the selected other node j2. j2 and the position vector p of the node of interest i The relative position vector p represents the difference between and . j2 -p i A second axis vector e2 is set that is orthogonal to the first axis vector e1 by selecting the corresponding vector and modifying the selected vector to be orthogonal to the first axis vector e1.
[0232] Then, the trained encoder sets a third axis vector e3 that is orthogonal to the first axis vector e1 and the second axis vector e2, thereby representing a frame E that is composed of one or more axis vectors. i This is calculated. For example, by performing the calculation of the cross product, a third axis vector e3 that is orthogonal to the first axis vector e1 and the second axis vector e2 can be identified.
[0233] (Non-orthogonal maximum frame) Alternatively, for example, a trained encoder may compute the number of frames per node by performing the process described in (3.3 Non-orthogonal Maximum Frames) above.
[0234] In this case, the trained encoder may calculate non-orthogonal axis vectors as frames by performing the following calculation process.
[0235] For example, when a trained encoder calculates a frame for each node, it uses the weights w between the node of interest i and several other nodes j. ij From among them, select the other node j1 corresponding to the weight with the largest value. Then, the trained encoder outputs the position vector p of the selected other node j1. j2 The position vector p of node i of interest i The relative position vector p represents the difference between and . j1 -p i The normalized vector corresponding to this is set as the first axis vector e1.
[0236] Next, the trained encoder selects another node j2 such that the absolute value of the dot product with the first axis vector e1 becomes small and the weight becomes large. Then, the trained encoder calculates the position vector p of the selected other node j2. j2 The node p that we are focusing on. i The relative position vector p represents the difference between and . j2 -p i Select the corresponding vector and set the selected vector as the second axis vector e2.
[0237] Next, the trained encoder selects another node j3 such that the absolute value of the dot product between it and the first axis vector e1 becomes small, the absolute value of the dot product between it and the second axis vector e2 becomes small, and the weight becomes large, and the position vector p of the selected other node j3 is determined. j3 The node p that we are focusing on. i The relative position vector p represents the difference between and . j3 -p i Select the corresponding vector and set the selected vector as the third axis vector e3.
[0238] Alternatively, for example, a trained encoder may compute a per-node frame by performing the process described in (3.4 Greedy Frames) or (3.5 Mixed Frames) above. When performing (3.4 Greedy Frames) above, the trained encoder, when computed a per-node frame, selects multiple other nodes in descending order of weight from the weights between the node of interest and multiple other nodes for each of the multiple nodes. The trained encoder then computes a frame represented by one or more axis vectors by setting each of the vectors corresponding to the relative position vectors representing the difference between the position vectors of the selected multiple other nodes and the node of interest as one or more axis vectors.
[0239] Next, the trained encoder calculates frame E for each of the multiple nodes, for node i of interest. i For this, the position p of node i of interest i and the position p of multiple other nodes j j The relative position vector p represents the difference between the two. j -p i By projecting this, we obtain relationship information between the node of interest i and several other nodes j, which is b. ij The following is calculated. For example, a trained encoder calculates relational information b according to the above equation (F). ij Calculate.
[0240] Then, the trained encoder calculates relational information b for each of the multiple nodes, with respect to the node i of interest. ij And weight lol ij Based on this, the state vector x of node i of interest i The process of updating is repeated. For example, a trained encoder will update the state vector x of node i of interest according to equation (E) above. i Repeat the process of updating.
[0241] Then, the trained encoder updates the node state vector x for each of the multiple nodes. iThe process is executed to output '. For example, the trained encoder uses the updated node state vector x of equation (E) above. i Output '.
[0242] As mentioned above, the trained encoder has a layer configuration that allows it to perform each of the above processes.
[0243] <Operation of the trained model generator 14A> Next, the operation of the trained model generation device 14A in this embodiment will be explained with reference to the figure. When the trained model generation device 14A receives training data, it stores it in the data storage unit 30. Then, when the trained model generation device 14A receives an instruction signal to start the training process, it executes the trained model generation processing routine shown in Figure 10.
[0244] <Trained Model Generation Routine> In step S100, the learning unit 18A acquires multiple training data stored in the data storage unit 30A.
[0245] In step S102, the learning unit 18A trains the Transformer using a known machine learning algorithm based on the training data acquired in step S100.
[0246] In step S104, the learning unit 18A acquires the learned transformers.
[0247] In step S106, the learning unit 18A stores the trained transformer acquired in step S104 into the trained model storage unit 32A.
[0248] Next, when the information processing device 14B receives structural information of the target molecular structure or crystal structure, it executes the information processing routine shown in Figure 11.
[0249] <Information Processing Routine> In step S200, the processing unit 20B receives the input structural information.
[0250] In step S202, the processing unit 20B sets the initial value of the state vector for each node of the molecular structure or crystal structure. For example, the processing unit 20B sets the initial value of the state vector x according to equations (3) to (5) above.
[0251] In step S204, the processing unit 20B reads out the encoder portion of the trained transformer stored in the trained model memory unit 32B.
[0252] In step S206, the processing unit 20B inputs the node-specific state vector x initialized in step S202 to the trained encoder read in step S204.
[0253] In step S208, the processing unit 20B causes the trained encoder to perform processing. The processing in step S208 is performed by the subroutine shown in Figure 12.
[0254] In step S300 of Figure 12, the trained encoder receives the state vector x of each of the multiple nodes.
[0255] In step S302 of Figure 12, the trained encoder uses the state vector x of each of the multiple nodes. i Based on this, weights w represent the interaction strength between the node of interest i and several other nodes j. ij A weight vector w whose components are i Calculate.
[0256] In step S304 of Figure 12, the trained encoder calculates the state vector x of node i of interest for each of the multiple nodes. i and the state vector x of multiple other nodes j j Weight vector w representing the interaction strength between iBased on this, frame E represents the coordinate axes at node i of interest. i This calculates the frame E for each node. For example, by performing each of the processes described above, the frame E for each node is calculated. i Calculate.
[0257] In step S306 of Figure 12, the trained encoder calculates the frame E for each of the multiple nodes for node i of interest. i For this, the position p of node i of interest i and the position p of multiple other nodes j j The relative position vector p represents the difference between the two. j -p i By projecting this, we obtain relationship information between the node of interest i and several other nodes j, which is b. ij The following is calculated. For example, a trained encoder calculates relational information b according to the above equation (F). ij Calculate.
[0258] In step S306 of Figure 12, the trained encoder calculates relational information b for each of the multiple nodes, with respect to the node i of interest. ij And weight lol ij Based on this, the state vector x of node i of interest i The process of updating is repeated. For example, a trained encoder will update the state vector x of node i of interest according to equation (E) above. i The process of updating is repeated. Then, the trained encoder updates the node state vector x for each of the multiple nodes. i The process is executed to output '. For example, the trained encoder uses the updated node state vector x of equation (E) above. i Output '.
[0259] Next, in step S210 of Figure 11, the processing unit 20 obtains the updated state vector x' output in step S208.
[0260] Then, in step S212, the processing unit 20 outputs the updated state vector x'.
[0261] As described above, the updated state vector x' corresponds to the physical properties of the molecular or crystal structure, the forces acting on the molecular or crystal structure, or input to other neural networks. As described above, the updated state vector x' reflects the information about frame E set for each node. Furthermore, frame E set for each node is a coordinate system that takes into account the interactions that exist between atoms. Therefore, when performing calculations on multiple atoms constituting a molecular or crystal structure, it is possible to perform calculations that take into account the interactions that exist between atoms.
[0262] As explained above, when the information processing device 14B uses a neural network to calculate information about multiple atoms constituting a molecular structure or crystal structure, the neural network receives the state of each node representing one or more atoms among the multiple atoms constituting the molecular structure or crystal structure. Next, for each of the multiple nodes, the neural network calculates a frame representing the coordinate axes of the node of interest based on weights representing the interaction strength between the state of the node of interest and the states of the other nodes. Then, the neural network performs a process to calculate information about the multiple atoms based on the frames calculated for each of the multiple nodes. This makes it possible to perform calculations that take into account the interactions that exist between atoms when performing calculations about multiple atoms constituting a molecular structure or crystal structure.
[0263] When using machine learning models such as neural networks to predict physical properties or forces on molecular or crystal structures, it is expected that performance can be improved by incorporating not only interatomic distance information but also positional and directional information such as relative position vectors within the machine learning model. However, it is not obvious how a machine learning model can incorporate atomic position or orientation information while guaranteeing invariance or similar invariance with respect to rotation or translation operations on the coordinate system of the input structure.
[0264] In the embodiments of this disclosure, instead of using a predetermined fixed frame for each input structure, the frame is adaptively determined according to the target task or the inference status of atomic states within the machine learning model, and the machine learning model is trained accordingly. By adaptively learning the frame according to the inference state of each atom within each prediction task or machine learning model, rather than directly defining a fixed frame for the input structure, a frame that follows the operation described by the machine learning model can be realized, thereby improving the accuracy of prediction tasks. Furthermore, by using this in various neural networks that process molecular structures or crystal structures as input, more advanced information about the structure can be transmitted to subsequent networks, and improved performance can be expected in tasks such as structure generation.
[0265] Furthermore, the technology disclosed herein is not limited to the embodiments described above, and various modifications and applications are possible without departing from the gist of this disclosure.
[0266] [Differentiation] In the above embodiment, the use of a transformer as an example of a neural network was described, but the invention is not limited to this. A layer capable of performing the above-described processing may be constructed using a neural network different from that of a transformer, and the above-described processing may be performed using such a layer.
[0267] Furthermore, the molecular structure or crystal structure of the above embodiment may be, for example, an enantiomer. In this case, since an enantiomer corresponds to a reflection of a certain molecular structure or crystal structure in a mirror, it is possible to design the neural network to distinguish between enantiomers or not, depending on the intended application. For example, if frame E is a square matrix, and frame E is calculated such that the sign of det(E) is always constant, the neural network will distinguish between enantiomers. Conversely, if frame E is calculated such that the sign of det(E) changes randomly, the neural network will not distinguish between enantiomers. Similarly, by converting the above-mentioned signed projection information to absolute values, it is possible to avoid distinguishing between enantiomers.
[0268] Furthermore, when estimating these variables, the neural network calculates and updates rotational and translationally invariant feature vectors for each atom using interatomic relative distances and projection information from frames. From these, spatial geometric quantities such as forces are regressed and output based on an invariant coordinate system, and then estimated by converting the frames inversely back to the orientation of the input coordinate system. Here, the frames used in the inverse conversion of the output values are the frames E calculated for each atom. i Alternatively, a conventional fixed frame may be used. Furthermore, instead of inverse frame transformation, in the output layer, for example, v i =Σ j w ij (p j -p i v is a weighted sum of geometric quantities such as relative position vectors calculated from the input structure, as shown above. i By outputting this, it is possible to estimate the same variables in accordance with the orientation of the input structure. Here, the weight w ij This is calculated based on the state vectors of atoms, which are computed and updated within the neural network.
[0269] Furthermore, although the above embodiment describes the case where the structure represented as a set of nodes arranged in space is a molecular structure or a crystal structure, it is not limited to these. For example, each node of the target structure may consist of an atom, molecule, gene, human, mobile body, robot, tangible object, fluid, and physical entity in information space, or a combination of two or more of these. In this case, the above information processing device performs information processing using a neural network on the structure represented as a set of nodes arranged in space. Specifically, the information processing device receives input of the state of each node, calculates a frame representing the coordinate axes at each node based on the states between each node, and extracts information having a predetermined symmetry of the structure from the nodes using the frame. In this case, the frame is calculated by principal component analysis or the like. For example, the frame is calculated by the method exemplified in 3.1 to 3.5 above. The frame may be an orthogonal or non-orthogonal system. Having symmetry means either being invariant or equivariant with respect to symmetry. Furthermore, symmetry includes any or a combination of rotation, translation, and inversion in space.
[0270] Furthermore, as mentioned above, the above embodiment was described using the case where a trained encoder is used to calculate the state vector x, but it is not limited to this. The state vector x may be calculated using a trained decoder, or it may be calculated using both a trained encoder and a trained decoder.
[0271] Furthermore, in the above embodiment, each process executed by the CPU or GPU after reading the software (program) may be executed by various processors other than the CPU or GPU. Examples of such processors include PLDs (Programmable Logic Devices) such as FPGAs (Field-Programmable Gate Arrays) whose circuit configuration can be changed after manufacturing, and dedicated electrical circuits that are processors with circuit configurations specifically designed to execute specific processes, such as ASICs (Application Specific Integrated Circuits). Each process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (for example, multiple FPGAs, and a combination of a CPU or GPU and FPGAs). More specifically, the hardware structure of these various processors is an electrical circuit that combines circuit elements such as semiconductor elements.
[0272] Furthermore, although the above embodiment describes a configuration in which each program is pre-stored (installed) on a storage device, the invention is not limited to this configuration. Programs may be provided in a form stored on a storage medium such as a CD-ROM, DVD-ROM, Blu-ray disc, or USB memory. Programs may also be provided in a form that can be downloaded from an external device via a network. [Examples]
[0273] In this embodiment, a simulation experiment is conducted regarding the proposed method described above. In this embodiment, the task of predicting physical properties is simulated when the proposed method is applied to a lightweight neural network based on GeoFormer disclosed in the following reference. Specifically, in this embodiment, a task similar to the one shown in Table 1 of the following reference is simulated. The following reference uses the QM9 dataset.
[0274] References: "Geometric Transformer with Interatomic Positional Encoding", <https: / / proceedings.neurips.cc / paper_files / paper / 2023 / hash / aee2f03ecb2b2c1ea55a43946b651cfd-Abstract-Conference.html>
[0275] Table 3 below shows the simulation results of this embodiment. In the table below, μ represents the dipole moment of the molecule, and its unit is [mD]. G represents the free energy of the molecule, and its unit is [meV]. The values shown in the table below are the mean absolute errors between the correct values and the predicted values of the physical properties. The "GeoFormer (lightweight version, no frame)" section below shows the errors when calculating the dipole moment and free energy, which are examples of molecular physical properties, using the lightweight version of GeoFormer, a neural network disclosed in the above reference. The "PCA (existing fixed frame)" section below shows the errors when calculating the dipole moment and free energy, which are examples of molecular physical properties, using PCA to set a fixed frame for the lightweight version of GeoFormer. The "Orthogonal Maximum (proposed)" section below shows the errors when calculating the dipole moment and free energy, which are examples of molecular physical properties, using the proposed frame (a frame formed by orthogonal axes) for the lightweight version of GeoFormer. Furthermore, the "Non-orthogonal Maximum (Proposed)" item below represents the error when calculating the dipole moment and free energy, which are examples of molecular properties, using the proposed frame (a frame formed by non-orthogonal axes) in the lightweight version of GeoFormer.
[0276] As shown in the table below, the proposed method of this embodiment has a smaller error and can estimate material properties with greater accuracy than conventional methods. [Table 3]
[0277] (Note) The following is an addendum regarding the nature of this disclosure.
[0278] (Note 1) In a neural network-based information processing method for a structure represented as a set of nodes arranged in space, A step of receiving input on the state of each node, A step of calculating a frame representing the coordinate axes at each node based on the state between each node, The steps include: extracting information having a predetermined structural symmetry from the node using the frame; An information processing method in which a computer performs a process having [a certain characteristic]. (Note 2) Each of the aforementioned nodes consists of one or more combinations of atoms, molecules, genes, humans, mobile bodies, robots, tangible objects, fluids, and physical entities in information space. The information processing method described in Appendix 1. (Note 3) The aforementioned frame is calculated using principal component analysis. The information processing method described in Appendix 1. (Note 4) The aforementioned frame uses an orthogonal system. The information processing method described in Appendix 1. (Note 5) The aforementioned frame uses a non-orthogonal system. The information processing method described in Appendix 1. (Note 6) Having the aforementioned symmetry means that the symmetry is either invariant or equivariant. The information processing method described in Appendix 1. (Note 7) The aforementioned symmetry includes any or a combination of rotation, translation, and inversion in the space. The information processing method described in Appendix 1. (Note 8) Using the aforementioned symmetry information, the physical properties of a structure composed of atoms are inferred. The information processing method described in Appendix 1. (Note 9) The aforementioned physical properties include at least one of the following: formation energy, total energy, zero-point energy, enthalpy, free energy, internal energy, specific heat, total magnetization, absolute magnetization, HOMO (Highest Occupied Molecular Orbital) energy level, LUMO (Lowest Unoccupied Molecular Orbital) energy level, HOMO-LUMO gap, band gap, superconducting transition temperature, bulk modulus, shear modulus, energy at which a set of molecular or crystal structures bond, physical properties representing the classification of magnetism, whether or not it is a ferroelectric, whether or not it is a pyroelectric, whether or not it is a piezoelectric, whether or not it is superconducting, whether or not a set of molecular or crystal structures undergoes a predetermined chemical reaction, force acting on atoms, variable of atomic position to a stable state structure, dipole moment, and atomic position in each structure when a set of molecular or crystal structures bond. The information processing method described in Appendix 8. (Note 10) The aforementioned structure is a molecular structure or a crystalline structure. When using the aforementioned neural network to calculate information about multiple atoms constituting a molecular structure or crystal structure, The neural network receives the state of each node representing one or more atoms among the multiple atoms constituting the molecular structure or crystal structure. The neural network calculates a frame representing the coordinate axes of each of the multiple nodes based on weights representing the interaction strength between the state of the node of interest and the states of the other nodes, The neural network performs a process such that it calculates the symmetry information based on the frames calculated for each of the plurality of nodes. The information processing method described in Appendix 1. (Note 11) The aforementioned symmetry information is the state of the node, The aforementioned neural network is The spatial position of each of the multiple nodes is further accepted. Based on the state of each of the multiple nodes, a weight is calculated that represents the strength of the interaction between the node of interest and the other nodes. For each of the multiple nodes, relative position information representing the difference between the position of the node of interest and the positions of the other nodes is calculated by projecting relative position information, which represents the difference between the position of the node of interest and the positions of the other nodes, onto the frame calculated for the node of interest. For each of the multiple nodes, the state of the node of interest is repeatedly updated based on the relationship information and weight calculated for the node of interest. For each of the multiple nodes, perform a process that outputs the updated state of the node. The information processing method described in Appendix 10. (Note 12) The state for each of the plurality of nodes is a state vector representing the state of the node, The spatial position of each of the plurality of nodes is a position vector representing the position of the node, The weight representing the interaction strength regarding the state between the node of interest and multiple other nodes is a weight vector whose components are the weights between the node of interest and multiple other nodes. The aforementioned neural network is For each of the multiple nodes, the relationship information between the node of interest and the multiple other nodes is calculated by projecting a relative position vector, which represents the difference between the position vector of the node of interest and the position vectors of the other nodes, onto the frame set for the node of interest. For each of the multiple nodes, the state vector of the node of interest is repeatedly updated based on the relationship information and weights calculated for the node of interest. Output the updated state vector of the node of interest for each of the multiple nodes. The information processing method described in Appendix 11. (Note 13) The neural network calculates the frame for each node, For each of the multiple nodes, the frame represented by one or more axis vectors is calculated by performing principal component analysis based on a relative coordinate matrix whose components are the relative position vectors between the node of interest and the multiple other nodes, and a weight matrix whose components are the weight vectors between the node of interest and the multiple other nodes. The information processing method described in Appendix 12. (Note 14) The neural network calculates the frame for each node, For each of the multiple nodes, the frame represented by one or more axis vectors is calculated by performing principal component analysis based on a relative coordinate matrix whose components are vectors normalized to 1, which are the lengths of the relative position vectors between the node of interest and the other multiple nodes, and a weight matrix whose components are weight vectors, which are the weights between the node of interest and the other multiple nodes. The information processing method described in Appendix 12. (Note 15) When the neural network calculates the frame for each node, From among the weights between the node of interest and a plurality of other nodes, select the other node corresponding to the weight with the largest weight, and set the normalized vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the node of interest as the first axis vector. Select another node such that the absolute value of the dot product between it and the first axis vector becomes smaller and the weight becomes larger, select a vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the position vector of the node of interest, and set a second axis vector orthogonal to the first axis vector by modifying the selected vector so that it is orthogonal to the first axis vector. By setting a third axis vector orthogonal to the first axis vector and the second axis vector, the frame represented by one or more axis vectors is calculated. The information processing method described in Appendix 12. (Note 16) When the neural network calculates the frame for each node, From among the weights between the node of interest and a plurality of other nodes, select the other node corresponding to the weight with the largest weight, and set the normalized vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the position vector of the node of interest as the first axis vector. Select another node such that the absolute value of the dot product with the first axis vector becomes smaller and the weight becomes larger, select a vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the node of interest, and set the selected vector as the second axis vector. Select another node such that the absolute value of the dot product with the first axis vector becomes smaller, the absolute value of the dot product with the second axis vector becomes smaller, and the weight becomes larger; select a vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the position vector of the node of interest; and set the selected vector as the third axis vector. The information processing method described in Appendix 12. (Note 17) The neural network calculates the frame for each node, For each of the multiple nodes, select multiple other nodes in descending order of their weights from the weights between the node of interest and the multiple other nodes, and set each of the vectors corresponding to the relative position vector representing the difference between the position vectors of the selected multiple other nodes and the node of interest as one or more axis vectors, thereby calculating the frame represented by the one or more axis vectors. The information processing method described in Appendix 12. (Note 18) An information processing device using a neural network for a structure represented as a set of nodes arranged in space, A reception unit that receives input on the status of each node, A calculation unit that calculates a frame representing the coordinate axes at each node based on the state between each node, An extraction unit that extracts information having a predetermined structural symmetry from the node using the frame, An information processing device having (Note 19) A program that causes a computer to perform information processing using a neural network on a structure represented as a set of nodes arranged in space, It accepts input of the state of each node, Based on the state between each node, calculate a frame representing the coordinate axes at each node. Using the aforementioned frame, information having a predetermined structural symmetry is extracted from the nodes. A program that causes a computer to perform a process. [Explanation of Symbols]
[0279] 14A Information Processing Device 14B Pre-trained model generator 18A Learning Department 20B Processing Unit 30A, 30B Data Storage Unit 32A, 32B Pre-trained model memory
Claims
1. In a neural network-based information processing method for a structure represented as a set of nodes arranged in space, A step of receiving input on the state of each node, A step of calculating a frame representing the coordinate axes at each node based on the state between each node, The steps include: extracting information having a predetermined structural symmetry from the node using the frame; An information processing method in which a computer performs a process having [a certain characteristic].
2. Each of the aforementioned nodes consists of one or more combinations of atoms, molecules, genes, humans, mobile bodies, robots, tangible objects, fluids, and physical entities in information space. The information processing method according to claim 1.
3. The aforementioned frame is calculated using principal component analysis. The information processing method according to claim 1.
4. The aforementioned frame uses an orthogonal system. The information processing method according to claim 1.
5. The aforementioned frame uses a non-orthogonal system. The information processing method according to claim 1.
6. Having the aforementioned symmetry means that the symmetry is either invariant or equivariant. The information processing method according to claim 1.
7. The aforementioned symmetry includes any or a combination of rotation, translation, and inversion in the space. The information processing method according to claim 1.
8. Using the aforementioned symmetry information, the physical properties of a structure composed of atoms are inferred. The information processing method according to claim 1.
9. The aforementioned physical properties include at least one of the following: formation energy, total energy, zero-point energy, enthalpy, free energy, internal energy, specific heat, total magnetization, absolute magnetization, HOMO (Highest Occupied Molecular Orbital) energy level, LUMO (Lowest Unoccupied Molecular Orbital) energy level, HOMO-LUMO gap, band gap, superconducting transition temperature, bulk modulus, shear modulus, energy at which a set of molecular or crystal structures bond, physical properties representing the classification of magnetism, whether or not it is a ferroelectric, whether or not it is a pyroelectric, whether or not it is a piezoelectric, whether or not it is superconducting, whether or not a set of molecular or crystal structures undergoes a predetermined chemical reaction, force acting on atoms, variable of atomic position to a stable state structure, dipole moment, and atomic position in each structure when a set of molecular or crystal structures bond. The information processing method according to claim 8.
10. The aforementioned structure is a molecular structure or a crystalline structure. When using the aforementioned neural network to calculate information about multiple atoms constituting a molecular structure or crystal structure, The neural network receives the state of each node representing one or more atoms among the multiple atoms constituting the molecular structure or crystal structure. The neural network calculates a frame representing the coordinate axes of each of the multiple nodes based on weights representing the interaction strength between the state of the node of interest and the states of the other nodes, The neural network performs a process such that it calculates the symmetry information based on the frames calculated for each of the plurality of nodes. The information processing method according to claim 1.
11. The aforementioned symmetry information is the state of the node, The aforementioned neural network is The spatial position of each of the multiple nodes is further accepted. Based on the state of each of the multiple nodes, a weight is calculated that represents the strength of the interaction between the node of interest and the other nodes. For each of the multiple nodes, relative position information representing the difference between the position of the node of interest and the positions of the other nodes is calculated by projecting relative position information, which represents the difference between the position of the node of interest and the positions of the other nodes, onto the frame calculated for the node of interest. For each of the multiple nodes, the state of the node of interest is repeatedly updated based on the relationship information and weight calculated for the node of interest. For each of the multiple nodes, perform a process that outputs the updated state of the node. The information processing method according to claim 10.
12. The state for each of the plurality of nodes is a state vector representing the state of the node, The spatial position of each of the plurality of nodes is a position vector representing the position of the node, The weight representing the interaction strength regarding the state between the node of interest and multiple other nodes is a weight vector whose components are the weights between the node of interest and multiple other nodes. The aforementioned neural network is For each of the multiple nodes, the relationship information between the node of interest and the multiple other nodes is calculated by projecting a relative position vector, which represents the difference between the position vector of the node of interest and the position vectors of the other nodes, onto the frame set for the node of interest. For each of the multiple nodes, the state vector of the node of interest is repeatedly updated based on the relationship information and weights calculated for the node of interest. Output the updated state vector of the node of interest for each of the multiple nodes. The information processing method according to claim 11.
13. The neural network calculates the frame for each node, For each of the multiple nodes, the frame represented by one or more axis vectors is calculated by performing principal component analysis based on a relative coordinate matrix whose components are the relative position vectors between the node of interest and the multiple other nodes, and a weight matrix whose components are the weight vectors between the node of interest and the multiple other nodes. The information processing method according to claim 12.
14. The neural network calculates the frame for each node, For each of the multiple nodes, the frame represented by one or more axis vectors is calculated by performing principal component analysis based on a relative coordinate matrix whose components are vectors normalized to 1, which are the lengths of the relative position vectors between the node of interest and the other multiple nodes, and a weight matrix whose components are weight vectors, which are the weights between the node of interest and the other multiple nodes. The information processing method according to claim 12.
15. When the neural network calculates the frame for each node, From among the weights between the node of interest and a plurality of other nodes, select the other node corresponding to the weight with the largest weight, and set the normalized vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the node of interest as the first axis vector. Select another node such that the absolute value of the dot product between it and the first axis vector becomes smaller and the weight becomes larger, select a vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the position vector of the node of interest, and set a second axis vector orthogonal to the first axis vector by modifying the selected vector so that it is orthogonal to the first axis vector. By setting a third axis vector orthogonal to the first axis vector and the second axis vector, the frame represented by one or more axis vectors is calculated. The information processing method according to claim 12.
16. When the neural network calculates the frame for each node, From among the weights between the node of interest and a plurality of other nodes, select the other node corresponding to the weight with the largest weight, and set the normalized vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the position vector of the node of interest as the first axis vector. Select another node such that the absolute value of the dot product with the first axis vector becomes smaller and the weight becomes larger, select a vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the node of interest, and set the selected vector as the second axis vector. Select another node such that the absolute value of the dot product with the first axis vector becomes smaller, the absolute value of the dot product with the second axis vector becomes smaller, and the weight becomes larger; select a vector corresponding to the relative position vector representing the difference between the position vector of the selected other node and the position vector of the node of interest; and set the selected vector as the third axis vector. The information processing method according to claim 12.
17. The neural network calculates the frame for each node, For each of the multiple nodes, the frame represented by the one or more axis vectors is calculated by selecting multiple other nodes in descending order of their weights from the weights between the node of interest and the multiple other nodes, and setting each of the vectors corresponding to the relative position vector representing the difference between the position vectors of the selected multiple other nodes and the node of interest as one or more axis vectors. The information processing method according to claim 12.
18. An information processing device using a neural network for a structure represented as a set of nodes arranged in space, A reception unit that receives input on the status of each node, A calculation unit that calculates a frame representing the coordinate axes at each node based on the state between each node, An extraction unit that extracts information having a predetermined structural symmetry from the node using the frame, An information processing device having