Inference device, training device, and inference method
A neural network model using transformation and denoising fields addresses the complexity of generating organic reaction pathways, offering accurate and efficient simulation of chemical reactions by learning from datasets to predict activation barriers.
Patent Information
- Application Number
- PCT/JP2025/001399
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
Generating accurate chemical reaction pathways and transition states in molecular simulations is challenging due to the complexity of three-dimensional atomic structures, especially for organic reactions, where existing methods struggle to provide precise minimum energy pathways or transition states without expert intervention.
A neural network-based approach that utilizes a transformation guidance field and a denoising field to generate reaction paths by learning from datasets, allowing for fast and versatile generation of organic reaction pathways directly from initial states to final states, incorporating features like bond changes and dihedral angle rotations.
Enables the generation of reaction paths that closely resemble actual pathways, facilitating the prediction of activation barriers and accelerating reaction simulations by providing a rough sketch of complex organic reactions with high accuracy and versatility.
Smart Images

Figure JP2025001399_24072025_PF_FP_ABST
Abstract
Description
Inference device, training device, and inference method
[0001] The present disclosure relates to a reasoning device, a training device, and a reasoning method.
[0002] Traditionally, mapping chemical reaction pathways and their corresponding activation barriers has been an important challenge in molecular simulation.
[0003] Due to the inherent complexity of 3D atomic structures, generating an initial guess for a chemical reaction pathway can be difficult without expert tuning. Chemical reaction pathways are sometimes generated using generative models, such as machine learning potentials. However, these models are generally not very accurate. Therefore, it is difficult to provide accurate minimum-energy pathways or transition states where the forces are nearly zero in the actual potential.
[0004] Zheng, S. et al. Towards Predicting Equilibrium Distributions for Molecular Systems with Deep Learning. arXiv [physics. chem-ph] 2023Triplett, L. ; Lu, J. Diffusion Methods for Generating Transition Paths. arXiv [physics. comp-ph] 2023
[0005] The objective of the present disclosure is to generate a route closer to an optimized reaction route than conventional techniques, to generate a complex route with fewer instructions, or to generate a variety of routes at higher speeds, but is not limited to these. Note that the objectives corresponding to the effects of the configurations shown in the embodiments described below can also be positioned as other objectives.
[0006] An example of an inference device disclosed herein includes at least one memory and at least one processor, and the at least one processor repeatedly outputs information necessary for determining a reaction path from the trained model each time atomic information is input into the trained model, and determines the reaction path based on the information necessary for determining the reaction path.
[0007] A training device as an example of the present disclosure includes at least one memory and at least one processor, and the at least one processor inputs training data to a neural network to be learned, thereby outputting inference data to be inferred by the neural network and uncertainty data indicating the uncertainty of the inference data, calculates a loss function including an exponent of a multivariate normal distribution having a difference between ground truth data corresponding to the training data and the inference data and a covariance matrix having the uncertainty data as elements, and a determinant of the covariance matrix, and trains the neural network and the uncertainty data so as to minimize the value of the loss function.
[0008] FIG. 1 is a model image of a trained flow according to an embodiment. FIG. 2 is a diagram showing an example of input / output to a trained model according to an embodiment. FIG. 3 is a diagram showing an example of interaction of an e3-attention network according to an embodiment. FIG. 4 is a diagram showing an example of a readout unit in FIG. 2. FIG. 5 is a diagram showing the transition of the average norm of the difference vector between the prediction of the transformation guidance field and the denoising field and the training data during learning about transition1x according to an embodiment. FIG. 6 is a diagram showing an example of the results of generating a reaction path with various parameters for the same reaction path as in FIG. 1 according to an embodiment. FIG. 7 is a diagram showing the results of generating a reaction path with various parameters for the same reaction path as in FIG. 1 according to an embodiment. 8 H 9FIG. 8 is a diagram showing an example of the results of generating a cyclization reaction of O with various parameters. FIG. 8 is a diagram showing an example of a rotation reaction of polyethylene used to verify the limits of generalization performance, according to an embodiment. FIG. 9 is a table showing an example of the results of generating a reaction path, according to an embodiment. FIG. 10 is a diagram showing a table summarizing whether the final state is one in which the central CC bond has rotated compared to the initial state, according to an embodiment. FIG. 11 is a table showing whether the final product, when generated under various conditions, is as specified, does not include rotation of bonds other than the specified bonds, and does not involve bond changes along the way, according to an embodiment. FIG. 12 is a diagram showing an example of a rotation reaction of polyethylene used to verify the limits of generalization performance, according to an embodiment. FIG. 9 is a table showing an example of the results of generating a reaction path, according to an embodiment. FIG. 10 is a diagram showing a table summarizing whether the final state is one in which the central CC bond has rotated compared to the initial state, according to an embodiment. FIG. 11 is a table showing whether the final product is as specified, does not include rotation of bonds other than the specified bonds along the way, and does not involve bond changes along the way, according to an embodiment. 8 H 9 FIG. 13 is a diagram showing an example of the results of conditional generation for a Diels alder reaction with O. FIG. 13 is a diagram showing an example of the results of conditional generation performed by applying bond changes to test data 833 data separated from training data, according to an embodiment. FIG. 14 is a diagram showing an example of the results of optimizing those results obtained with classifier-free guidance for w = 8.0 that satisfy given conditions at the PFP v4.0.0 level using the string method, and comparing the activation barriers of the corresponding data in the test dataset. FIG. 15 is a plot of the types of final state structures obtained when random generation is performed 16 times from a certain initial state, according to an embodiment. FIG. 16 is a block diagram showing an example of the hardware configuration of an information processing device that realizes an inference device according to an embodiment. FIG. 17 is a diagram showing an example of functional blocks realized by one or more processors, according to an embodiment. FIG. 18 is a flowchart showing an example of the procedure for reaction path inference processing, according to an embodiment. Fig. 19 is a diagram showing a comparative example showing reaction path 1 generated by a known scan, and an example of a reaction path generated by this embodiment. Fig. 20 is a diagram showing an example of energy at the coordinates of an inferred atom in a reaction path according to an embodiment. Fig. 21 is a diagram showing an example of functional blocks of a training device implemented by one or more processors according to an embodiment.
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The configuration of the embodiment described below, and the actions and effects brought about by the configuration, are merely examples and are not limited to the following description. First, an example of the present embodiment will be described, and then a specific example of the processing of the present embodiment will be described.
[0010] Mapping chemical reaction pathways and their corresponding activation barriers is a key challenge in molecular simulation. Given the inherent complexity of three-dimensional atomic structures, even generating initial guesses for these pathways can be challenging without expert craftsmanship. This embodiment introduces an innovative approach using neural networks to generate initial guesses for reaction pathways based on learning from a database of initial states and transition pathways. The proposed method begins by inputting the coordinates of the initial state and then applies incremental modifications to the structure. This iterative process generates the predicted reaction pathways and the coordinates of the final states. This method is fast because it does not require the on-the-fly calculation of actual potential energy surfaces. The application of this geometry-based method extends to complex reaction pathways exhibited by organic reactions. Training was performed on the Transition1x dataset of organic reaction pathways. The results showed that the generated reactions have a significant similarity to a test set of chemical reaction pathways. The flexibility of the method allows for generating reactions that conform to predefined criteria or in a random manner.
[0011] 1. Introduction Computational chemistry techniques are constantly improving our understanding of chemical reactions. Notably, the intersection of machine learning and computational chemistry has shown great potential for materials exploration based on atomistic energetics. Recent advances in machine learning have successfully accelerated computational chemistry. The emergence of machine learning potentials has significantly accelerated molecular dynamics, including chemical reactions. However, because chemical reactions are inherently rare events, the acceleration provided by machine learning potentials alone is insufficient to track chemical reactions within a feasible timeframe. Therefore, even with the improved speed of potentials, simply implementing simple sampling methods such as standard molecular dynamics or Monte Carlo methods is insufficient. Further development of sampling techniques is necessary.
[0012] In recent years, generative models that are invariant to translation, rotation, and permutation have emerged, and sampling methods themselves have made remarkable progress. Focusing on sampling methods using machine learning in particular, methods that convert simple distributions, such as Gaussian distributions, into complex distributions that the data should follow have become mainstream. In particular, methods such as normalized flow, diffusion models, and flow matching have been well studied for the purpose of sampling molecular structures. A common feature of these methods is that they repeatedly convert simple distributions, such as Gaussian distributions, into the desired distribution.
[0013] For example, E(n) equivariant normalizing flow samples the coordinates of each atom constituting a molecule from a Gaussian distribution, and learns the direction for generating an actual molecule using normalized flow. GeoDiff similarly samples the coordinates of each atom from a Gaussian distribution to learn the direction for generating a molecular structure, but uses a diffusion model to define this direction. DiffDock and Torsional Diffusion introduce a coordinate system using intramolecular dihedral angles and translational degrees of freedom, and apply a diffusion model within that coordinate system. Equivariant flow matching converts the coordinates of each atom from a Gaussian distribution. In this case, flow matching is used to define the movement direction of the atom. DiG (Distributional Grapher) determines the movement direction using a diffusion model, but learns using various coordinates for each system. Furthermore, while CDVAE generates bulk systems, the use of VAE-conditioned flows allows for the prediction of the energy of the resulting structure with a low-cost function prior to the time-consuming generation process.
[0014] In addition to the task of generating stable structures, the task of generating reaction paths (RPs) is also in high demand. Reaction paths (RPs) and transition states (TSs) provide important information about chemical reactions. A transition state is the highest energy point on a minimum energy path (MEP). The height of a transition state is an important parameter that determines the rate of a chemical reaction.
[0015] In recent years, in addition to models that generate stable structures, methods capable of sampling RP have also been proposed. Diffusion Methods for Generating Transition Paths discretize reaction paths and use a diffusion model in a space whose size is the product of the number of degrees of freedom of the structure and the number of discretized images. On the other hand, the Boltzmann Generator and Distributional Grapher (DiG) directly interpolate two points on a Gaussian distribution, generate structures from each point, and obtain paths connecting different basins. In particular, DiG has been trained on many systems using Grapher, and is expected to have general applications. These methods have the ability to smoothly interpolate between two basins. However, because they do not use MEP information in their training, it is unclear whether the generated paths are close to MEP.
[0016] A future application of reaction path generation is the off-lattice extension of Kinetic Monte Carlo (KMC). For example, recent attempts have been made to combine KMC with reinforcement learning.
[0017] To generate reaction paths and accelerate KMC, it is necessary to rapidly enumerate the initial state (IS) to the final state (FS) and activation barriers. Reaction generation methods have been developed to handle small molecules on solids or solid surfaces. However, no model capable of generating organic reactions in a continuous space has been proposed. Therefore, to our knowledge, the most effective method applicable to organic reactions and potentially capable of accelerating reaction simulations is temperature-accelerated dynamics (TAD), which samples high-temperature molecular dynamics (MD). One reason for this situation is that reactions in organic materials follow curvilinear reaction paths in which various atomic degrees of freedom are interrelated. This complex degree of freedom makes it difficult to treat organic materials in a 3N-dimensional space.
[0018] Generation using a generation model is generally not very accurate, so it is difficult to provide accurate MEPs or TSs such that the force is almost zero at the actual potential. However, if the rough shape of the reaction path can be given, it should be possible to predict the activation barrier by optimizing the reaction path using CI-NEB or by using a model that predicts the activation energy from the rough shape of the reaction path. Therefore, in order to speed up reaction simulations of organic compounds, a method for quickly generating approximate reaction paths is needed.
[0019] In this embodiment, we developed a model that can generate appropriate organic reaction pathways based on simple instructions or completely randomly. This is a groundbreaking method for dealing with the complex degrees of freedom of organic molecules. The proposed method takes the structure of the IS and an arbitrary reaction type as input, gradually modifies the structure from the IS along the reaction pathway, and simultaneously obtains an approximate reaction pathway and FS. To achieve this, we introduced two fields: a transformation guidance field and a denoising field. By modifying the structure according to these two fields, a reaction pathway connecting the IS and FS is obtained. This method is fast and can learn and generate reaction pathways directly from a dataset. It is also highly versatile, and we were able to train a generalized model for transition1x data. Furthermore, we were able to generate reactions for molecules with a larger number of atoms than the transition1x data.
[0020] 2. Method 2.1 Training target Let N be the number of atoms. In general, there are a large number of reaction paths, but the i-th reaction path is x RP,i Here, s is a parameter and is defined to satisfy the following formula (1). At this time, the minimum value of s is 0, and the maximum value is x RP,i The length of the reaction path is L i Let's say. Any coordinate x to x RP,iConsider the foot of a perpendicular line dropped to the point. In this case, the parameter s is defined by the following equation (2): matches.
[0021] In the proposed method, two types of fields are defined. The first field is the transformation guidance field. The transformation guidance field is the tangent vector of the reaction path facing the FS side at the foot of the perpendicular line dropped onto the reaction path to be followed. The transformation guidance field is defined as follows:
[0022] The second field is the denoising field. The denoising field is a perpendicular vector that is dropped onto the reaction path to be followed and points in the direction of the reaction path. The denoising field is defined as follows:
[0023] The machine learning model is t t,i (x) and t d,i (x) is learned. Instead of having information about the index of the training data, a condition vector given by the feature c is received as input. The feature c is, for example, a condition vector indicating the feature related to the connection desired by the user. t,i (x) is approximated by a machine learning model as y t,i (x, c). d,i (x) is approximated by a machine learning model as y d,i Let (x, c) be the t,i (x) is the reaction path differentiated by the parameter. Therefore, starting from the initial state as shown in the following equation (5), t t,i By integrating with respect to (x), the reaction path can be obtained.
[0024] Therefore, it is expected that a reaction path can be generated by learning the path from IS through RP to FS. t,iIf a reaction path is generated using only the above, it will deviate from the actual reaction path due to inference errors, approximation errors, etc. Fields at positions away from the reaction path are not learned, and there is no physical meaning in moving along the tangent to the reaction path at positions away from the reaction path. Therefore, once it deviates from the reaction path, it deviates from the reaction path at an accelerated rate. To solve this problem, score matching methods and diffusion models learn to what position, even in the vicinity of where actual data appears, it can return to the area where the data actually appears. Similarly, in reaction path learning, if a position deviating from the actual reaction path is reached, it must return in the direction of the reaction path. For this purpose, a linear combination of a transformation guidance field and a denoising field is used for generation, as shown in the following equation (6).
[0025] Also, unlike a general diffusion model, s is a parameter corresponding to the length of the reaction path, and the maximum value of s is unknown at the time of generation. Therefore, t is used as a variable to determine whether or not the reaction is in the middle of the reaction path. f Define t f is an integer scalar output that takes the value 0 or 1. f,i is approximated by a machine learning model f Define (x, c). f (x, c) are binary (two-component) outputs, and the end of production is determined based on which of the two outputs (two components) is larger. For ease of understanding, the reaction path is modeled as a curve in two-dimensional space in Figure 1.
[0026] Figure 1 is a model image of the learned flow. The IS is considered to be at the center of Figure 1. The three curves coming out of the IS are considered to be the optimized reaction path. The end of the reaction path on the non-IS side is FS. The dotted lines extending from each point throughout Figure 1 are the denoising field (t d ), and the dashed line is the transformation guidance field (t t )
[0027] In Figure 1, the center of the figure shows IS. Three curves coming out of IS each show a different reaction path, and these three reaction paths are x RP,i , i∈[1,3]. When such a reaction path exists, the field shown in Figure 1 is used as a teacher for learning by selecting the reaction path that has the shortest denoising field at each coordinate x. In Figure 1, the dashed line represents the transformation guidance field, and the dotted line represents the denoising field.
[0028] 2.2 Related Work The model used here is heavily influenced by flow matching, diffusion models, and reinforcement learning, with many modifications. Both the transformation guidance field and the denoising field can be interpreted in the context of flow matching and diffusion models, respectively. We also discuss the differences in the problem setting compared to imitation learning.
[0029] 2.2.1 Relationship between the transformation guidance field and existing generative models The transformation guidance field can be interpreted as a type of flow matching. In general diffusion models and flow matching, the concept of "time" is used to smoothly connect a simple distribution and a generated distribution. For example, one example is a Gaussian distribution at t = 0 and a Boltzmann distribution followed by molecules at t = 1. However, the transformation guidance field has elements that make it different from general diffusion models and flow matching.
[0030] It can also be interpreted that the transformation guidance field uses a distribution localized near IS instead of a simple distribution, and the product distribution uses a distribution localized around FS. In flow matching and diffusion models, the path during generation itself is not important, and only the distribution of the final state is important.
[0031] In contrast, in the transformation guidance field, the path itself is also important, and all structures obtained in the generation process are used to construct the reaction path. Also, in this problem setting, a vector field independent of time is learned, in contrast to the diffusion model and flow matching, in which the vector field learned changes with time. The path that transitions from IS along the learned vector field becomes the reaction path. Since the vector field that does not include time is the prediction target, the number of steps until generation is complete is not uniquely determined, unlike general flow matching and diffusion models. Therefore, the stopping condition was also predicted using equation (32) described below.
[0032] 2.2.2 Relationship between the denoising field and existing generative models The denoising field is related to denoising score matching in the case where all points on the reaction path are data distributions. Although it cannot be proven when the continuity of the reaction path has a negative effect, if the reaction path can be interpreted as a set of discrete points and there is only one point on the reaction path close to a certain point x, it can be proven that the newton step of the log likelihood of the distribution after perturbation in the denoising field and denoising score matching is the same. Let the sth discretized point on the i-th reaction path be x. RP,i,s In denoising score matching, the distribution spread around the data point is expressed by the following equation (7).
[0033] where: is the probability density distribution of the normal distribution expressed by the following equation (8). The gradient of the logarithm of equation (7) with respect to x is given by the following equation (9). Further, the second derivative of the logarithm of equation (7) with respect to x is given by the following equation (10).
[0034] Here, a Soft-Nearest function expressed in the form of the following equation (11) is introduced. As shown below: Define Then, in the limit of σ→0, the Soft-Nearest function satisfies the following equation: It converges to the value at the point on the reaction path closest to x.
[0035] Due to its nature, the denoising field coincides with the σ→0 limit of the following equation (12), and the following equation (13) holds. At any σ, Newton step and t d,i To find out how far apart (x, σ) is, t d,i The difference between (x, σ) multiplied by the Hessian of the logarithmic probability density and the gradient of the logarithmic probability density is defined as in the following equation (14). Equation (14) becomes the following equation (15). In equation (15), the contents of the parentheses asymptotically approach 0 at O(exp(-1 / σ)) as σ → 0. Therefore, as σ → 0, R i (x, σ) asymptotically approaches 0, and t d,i It can be shown that (x) is a Newton step of the logarithmic probability density when σ → 0. Therefore, in many cases, the denoising field term becomes a value close to the newton step when σ = 1.
[0036] 2.2.3 Relationship with Imitation Learning Reaction path prediction can also be considered a sequential decision-making problem that predicts coordinates along the MEP at each time. One approach is imitation learning (behavior cloning), which performs supervised learning from a correct action sequence. Regarding imitation learning (behavior cloning), an algorithm has been proposed to address the issue of not being able to make predictions if the trajectory deviates from the correct one during inference, resulting in a state that was not observed during training. The denoising field proposed in this embodiment also has a similar effect in that it has the function of learning the direction of the perpendicular to the correct trajectory and returning it to a point on the correct trajectory. However, our denoising field learns a score that is the gradient of the log-likelihood by removing noise. The gradient of the log-likelihood is expressed as p in the denominator, as shown in Equation (9). i Because of (x, σ), the score can be large even in low regions, and it is expected that the probability can be returned to a high region.
[0037] In known methods, a model (policy) under training is used to generate training data. On the other hand, in the experiment of this embodiment, training data is generated by adding random noise to the correct answer, and it is worth considering generating training data using a different method. However, since known methods and reinforcement learning frameworks assume discrete time, the direction from the deviation point to the next step point on the correct answer trajectory can be trivially defined, and a model that directly predicts this direction was trained. On the other hand, in this embodiment, t t,i Since the reaction path needs to be obtained by integrating with respect to (x), the definition of the next step in continuous time is not trivial. Therefore, in this study, we modeled it as two fields: a transformation guidance field and a denoising field.
[0038] 3. Training 3.1 Notation represents the tensor product of e3nn. represents concatenate. represents the operation of taking the product of two features and taking the sum in the feature direction. i A symbol with one subscript, such as c, indicates that it is the symbol for the i-th node (i-th atom). ij A symbol with two subscripts, such as The values written in bold are e(3)-equivariant quantities. is normalized represents teeth represents the norm of . All variables are written in e3nn notation according to the transformation rule to which they belong. As a particularly important note, the 0e component is an e3-invariant scalar component, and the 1o component is an e3-times equivariant vector component. Also, writing 128 x 0e indicates that the feature is composed of 128 e3-invariant scalar components. "FNN" indicates an operation that applies a fully connected neural network to the 0e component of each atom. silu was used as the activation function. resblock was used for transformations with the same number of inputs and outputs.
[0039] 3.1 Neural network architecture Let i and j be atom indices. The overall structure of the model is shown below. The formal input of the model is shown in equation (16) below. Here, x is the coordinate of the current structure, x IS is the coordinate of IS. Z i is the atomic number. ij is a feature vector that specifies the response to be generated. In this experiment, c only deals with edge features.
[0040] To ensure translational and rotational homogeneity, only relative coordinates were used. The input and output using relative coordinates are shown in Figure 2 and the following equation (17). Here, r is the relative coordinate between atoms. There are three types of relative coordinates: 1. The relative coordinate x between atoms in the IS structure IS,i -x IS,j 2. Relative coordinates x between atoms in the current structurei -x j 3. Relative coordinate x between atoms in IS and the current structure i -x IS,i These together are ij It is written as x i -x IS,i For the part, x IS,i From x i Only the edges to which information points were used.
[0041] The (i, j) pairs were determined with reference to a known method. This known method is a method of making the attention edge sparse in the Transformer. In this known method, supernodes were prepared in addition to normal nodes. Edges are connected between nodes that are close to each other, between all nodes and supernodes, and between pairs of randomly selected nodes. These three types of connections are called window, global, and random. In this embodiment, atoms within a distance of 12 Å were selected and connected up to the 128th nearest neighbor (window). Also, 32 atoms from atoms within 30 Å were randomly connected (random). Also, two supernodes were provided, and connections were made between the supernodes and all atoms (global). Since supernodes do not have coordinates, the relative coordinates of the edges between supernodes are expressed as r ij = 0.
[0042] In formula (17), c ijis composed of a value of 0 or 1 indicating whether the condition is met. The first condition indicates whether the reaction breaks a bond (i.e., whether atoms move farther apart). The second condition indicates whether the reaction forms a bond (i.e., whether atoms move closer together). The third condition indicates whether the reaction rotates the dihedral angle about the bond by more than 105°. The third condition was set to zero if either the first or second condition was met. The fourth condition indicates whether the first three conditions are used as model input. If the fourth condition is set to 0, the first through third conditions are set to 0, and the fourth condition is set to 1 for conditional generation and 0 for unconditional generation. The fifth condition indicates whether each interatomic bond is an edge connecting the IS to the current structure. The sixth condition indicates whether the edge is a window edge. The seventh condition indicates whether the edge is an edge connection. The eighth condition indicates whether this connection is an atom-to-supernode connection. The ninth condition indicates whether this connection is a supernode-to-supernode connection.
[0043] The entire model is shown in Figure 2. i A one-hot vector is used to embed Z i The embedding of is further transformed and normalized for each node by the neural network, and then treated as a node feature with only 0e components, which becomes the input of the first interaction block. the relative distance and normalized relative coordinates Divide into and embed each. We use sinusoidal embedding to embed The spherical harmonic function e3nn is used for , where the 1x0e component is always 1 and the 1x1o component is normalized is.
[0044] The embedding of the relative distance is a scalar feature of the edge, and c ij Therefore, c ijAfter concatenation, the result of conversion by FNN is treated as edge scalar feature S ij The normalized relative coordinates are embedded into the edge vector feature v ij is treated as an input to the interaction block.
[0045] The interaction was carried out five times, during which S ij and v ij was fixed. On the other hand, the node features were different values each time. ij Since the tensor product of these is included in the interaction block, the node features obtain higher-order tensor features for each interaction.
[0046] 3.3 Interaction Figure 3 shows the interaction of the e3-attention network. The interaction part of Figure 2 is shown in Figure 3(a). e3-attention is implemented using three edge features (q ij , k ij , v ij ) is generated using the three edge feature blocks introduced in Section 3.4. i , q ij , k ij , v ij Attention is obtained from the following equations (18) and (19). In equation (18), "·" indicates the dot product, which is obtained by multiplying each element by the feature and then adding them together. As a result, w ij becomes 1×0e. n' obtained from equation (19) i n of the inputs of the interaction block i and normalize to obtain the overall output. Here, normalization is expressed by equation (20). Normalizing in this way maintains rotational equivalence.
[0047] 3.4 Edge Feat In Section 3.3, edge features are constructed using the edge feature block. Figure 3(b) shows the edge feature block. First, to represent interactions, node features are i and j It is expressed as: n i , n j , S ij are concatenated to form a single edge tensor feature. An o3 linear transform is then performed and an FNN is applied. Using this result as input, the tensor product of multiple edge vector features is calculated to create the output. This tensor product ensures that the feature contains high-rank components.
[0048] 3.5 Readout (ReadOut) The readout section in Figure 2 is shown in Figure 4. There are four outputs. Output y fin indicates the stopping condition and has a size of 2x0e. Output y std,i is an output to allow for the uncertainty of predictions that varies from atom to atom, and its size is 1 × 0e. tan,i and y prp,i are the outputs of the directional prediction. Each has a size of 1×0. Using inner product decoding (Section 3.6), y tan,i and y prp,i The size of y is determined. tan,i and y prp,i Inner product decoding was also used for the direction of y, but the 1o component of the feature was used instead of the constant array. fin To obtain the y, the 0e features of all atoms were transformed using NN, and logsumexp was applied to aggregate the features of each atom into molecular features. The molecular features were transformed using FNN, and y fin was obtained.
[0049] 3.6 Inner product decoding Prediction vector y prp and y tan The norm of the scalar y stdThe features were decoded using an inner product. After applying the FNN transform to the 0e components of the features, softmax was applied to the features. The inner product was obtained using an array of equally spaced numbers with the same dimensions as the features.
[0050] y prp and y tan The sequence used to predict the norm of y is an equally spaced sequence from -2 Å to 2 Å in 0.1 Å increments. std The sequences used to predict the nuclei are evenly spaced from 0.1 Å to 1.0 Å in 0.1 Å increments.
[0051] 3.7 Dataset (Detaset) Transition1x was used. When verifying the results, the generated results may be re-optimized using the String method. In this case, a fast machine learning potential, PFP (Preferred Potential), is used to enable rapid optimization. Therefore, the training data used was obtained from RP optimization of Transition1x using PFP, and this was used as the learning dataset. The version of PFP used was Crystal U0 mode v4.0.0. After re-optimization, some paths split into multiple paths, and 10,074 paths were generated. Among the IS and FS of these paths, molecules with an energy difference within 0.05 eV and a distance between the most displaced atoms within 0.1 Å were considered identical, and identity was determined for all IS and FS. Then, only reactions involving bond changes or dihedral angle rotation were extracted, and among the reactions sharing the same IS and FS, only the reaction with the lowest activation barrier was extracted. While this process eliminated some paths, all paths were duplicated twice in a round trip. As a result, 11,801 reactions were treated as unique paths.
[0052] Next, the data was divided into train, valid, and test datasets. When dividing, molecules with the same composition were divided into the same group. Therefore, it was guaranteed that the same data did not exist in the train and valid. 90% of the total composition was treated as train data, 5% was treated as validation data, and the remaining 5% was treated as test data.
[0053] 3.8 Sampling The coordinates of the data that are the direct learning target of the NN are sampled along the reaction path. One reaction path is selected, and a structure with noise added to the IS is used as the initial value. The data is sampled by moving in the direction of the reaction path like flow matching, approaching the reaction path, and adding noise. Here, dw is the derivative of the Wiener process. g is sampled uniformly from the range [0.0, 0.2] each time a reaction path is selected. This process samples the structure around the reaction path. The resulting distribution is a tube-like distribution around the reaction path, and the distribution perpendicular to the path is a χ 2 The distribution is close to the distribution obtained by multiplying the data sampled from the distribution by g / (2^(1 / 2)). fin To facilitate the learning of fin = 0 data and t fin The ratio of data for = 1 was selected to be 1:1. To learn unconditional generation as well, training for unconditional generation was performed with a probability of 0.3, and for conditional generation with a probability of 0.7.
[0054] 3.9 Training t and y d For learning, we referred to score matching and flow matching methods. However, for ease of learning, we predicted the standard deviation of the output and defined a loss using the standard deviation. First, the multidimensional Gaussian distribution centered on t is expressed as follows: The negative logarithm of equation (22) was taken, and the following equation (23) was used as the loss. Regarding the denoising field and the translation guidance field, the loss was defined as in the following equations (24) and (25). Here, Σ is originally a 3N x 3N matrix, but in this case, all elements except the diagonal element are set to 0. The variance is predicted for each atom, and is a different value for each atom, but within the same atom, it is output so that it has the same variance in the x, y, and z directions. By adding standard deviation to learning, it is no longer necessary to fit the output to inputs that are difficult to learn, and the intention is to make it possible to further match the output to inputs that are easy to learn. y f The cross-entropy error expressed by the following equation (26) was used for fitting. f,i (x, c)] ifeat is the binary (two-component) output y f,i i in (x, c) feat It is the th element.
[0055] The loss during training was a linear combination of the losses related to the transformation guidance field, the denoising field, and the stopping condition. d Ross and y t Ross and y f The loss coefficients for all are set to the same. and FIG. 5 shows the average values of train data and valid data.
[0056] Figure 5 shows the transition of the average norm of the difference vector between the prediction and training data of the transformation guidance field and the denoising field during learning about transition1x. train-t is the train data for the transformation guidance field. valid-t is the validation data for the transformation guidance field. train-d is the train data for the denoising field in the train data. valid-d is the validation data for the denoising field. The horizontal axis in Figure 5 is the number of training rounds. The vertical axis in FIG. 5 represents the values of the train data and validation data.
[0057] As shown in FIG. 5, as the learning progresses, the train data and validation data become smaller. The loss quickly decreases and the value becomes almost constant, whereas continued to fall against the train data. While the value of the train data continues to decrease, it stops decreasing for the valid data. This is considered to be a typical example of overfitting. Because overfitting is occurring, it suggests that even if the dataset is for molecules with similar numbers of atoms, increasing the amount of reaction data may help improve performance.
[0058] 4. Results 4.1 Definition of the field used for generation Equation (21) was used during learning, but the output y t , y d Using the above, a reaction path can be generated using the following equation (27): where x is the coordinate, c is an arbitrary condition vector, yt is the predicted value of the transformation guidance, yd is the predicted value of the denoising, g is a constant, dw is the Wiener process, and dt and α are constants (e.g., the step size of the numerical solution). For molecules that are significantly larger than the size of the training data, t and y d In some cases, the direction of the vectors was reversed, and the generation did not proceed when using equation (27). In such cases, the following equation (28) was used for generation. Here, y d T is y t y orthogonalized to d and is defined as the following equation (29). For the conditional generation, the classifier free guidance shown in the following equation (30) was used. In orthogonalization using classifier-free guidance, Equation (29) is applied to the output with and without conditions, and then Equation (30) is applied, and then orthogonalization is performed again using Equation (29). Here, y(x, c) is the output vector when there is a condition, and y(x, 0) is the output vector when there is no condition. Also, the termination condition is y f The determination was made using y f is a binary (two-component) output, and generation ends when the first element becomes greater than the zeroth element.
[0059] 4.2 Importance of the denoising field To confirm the importance of the denoising field, we generated a denoising field with several different sizes. First, to verify the importance of the denoising field on a two-dimensional toy model, we used the same hypothetical reaction path as in Figure 1 and generated it using equation (21). α∈{0.0, 1.0}, g∈{0.0, 0.4} were used. The results are shown in Figure 6. α and g are constants.
[0060] Figure 6 shows an example of the results of generating a reaction path using various parameters for the same reaction path as in Figure 1. In Figure 6 (a), α = 0.0 and g = 0.0. In Figure 6 (b), α = 0.0 and g = 0.4. In Figure 6 (c), α = 1.0 and g = 0.0. In Figure 6 (d), α = 1.0 and g = 0.4.
[0061] Even when there is no noise at g = 0.0, the path (dashed line) deviates from the true path (solid line) when α = 0.0, whereas at α = 1.0 there is less deviation from the true path (solid line). At g = 0.4, the importance of the denoising field increases further, and at α = 0.0 the path passes through a location far removed from the true path (solid line), whereas at α = 1.0 the distribution wraps around the true path (solid line).
[0062] Furthermore, reaction paths for real molecules were generated using various parameters. The initial states and conditions of the reaction paths were determined using the C 8 H 9 One of the O routes was selected from the routes re-optimized with PFP v4.0.0. Generation was performed using α∈{0.0, 1.0}, g∈{0.0001, 0.01}. Generation was performed 16 times for each parameter, and the RMSD with the final state when the YARP route was optimized with PFP was calculated, and the average and standard deviation of the RMSD were calculated. The generation results are shown in Figure 7.
[0063] FIG. 7 shows C 8 H 9 7A and 7B are diagrams showing examples of results obtained by generating the cyclization reaction of O with various parameters. In (a) of FIG. 7, α = 0.0 and g = 0.0001. In (b) of FIG. 7, α = 0.0 and g = 0.01. In (c) of FIG. 7, α = 1.0 and g = 0.0001. In (d) of FIG. 7, α = 1.0 and g = 0.01. In FIG. 7, the numbers written under each generated result are the average and standard deviation of the RMSD between the original reaction path included in YARP and the final state optimized by PFP.
[0064] As shown in Figure 7, when the noise was large, g = 0.01, the shape was distorted to the point where the original structure was barely retained when α = 0.0, and it was beautiful when α = 1.0. Although it was slightly distorted compared to the structure when α = 1.0 and g = 0.0001, it was qualitatively the same structure. Even when the noise was small, various bonds were significantly distorted when α = 0.0 and g = 0.0001, where no denoising field was used.
[0065] 4.3 Field definition used for generation and generalization performance limit We generated reaction paths for central dihedral angle rotation reactions of polyethylene of various lengths by changing the definition of the denoising field, the denoising coefficient, and the magnitude of the classifier-free guidance (Equation (30)). Figure 8 shows an example of a rotation reaction of polyethylene used to verify the generalization performance limit. Equations (27) and (28) were used for generation.
[0066] The parameters used were dt = 0.1 and α∈{0.1, 1.0}. The value of α was chosen to be 1.0, which theoretically corresponds to the newton step, and 0.1, which is a smaller value. The classifier-free guidance used was w∈{0.0, 1.0, 2.0, 4.0, 8.0, 16.0}. Conditional generation was performed for polyethylenes with an even number of carbon atoms (n) between 2 and 16, with the dihedral angle rotation of the central CC bond. Generation sometimes worked well and sometimes did not, depending on the conditions, formulas, and parameters used. First, we investigated whether the generation time would be too long. We checked whether the time until completion was within 400 steps.
[0067] The reasons for the long generation time are as follows: (1) y t and y d (1) The atoms continue to change position after the bond dissociates. (2) The reaction path to be generated is too long to be expressed in 400 steps.
[0068] In the case of (1), the reaction often does not proceed, and the molecule oscillates while roughly maintaining the same structure. However, there are also cases where the reaction proceeds little by little while oscillating. In the case of (2), after the molecule splits into two, it continues to move without terminating. (3) does not occur when attempting to generate a CC rotation correctly, but it does occur when attempting to generate a longer, multi-step reaction path. The results are shown in Figure 9.
[0069] Figure 9 is a table showing the results. It shows whether the termination condition was met within 400 steps when generated under various conditions. The check marks indicate parameter sets that met the termination condition within 400 steps. In each table, the horizontal axis (w) represents the strength of classifier-free guidance (Equation (30)). The vertical axis (n) represents the number of carbon atoms in polyethylene. The tables in the upper left and upper right of Figure 9 were generated using Equation (27). The tables in the lower left and lower right of Figure 9 were generated using Equation (28). In Figure 9, the tables in the upper left and lower left used α = 0.1. The tables in the upper right and lower right used α = 1.0. Here, α is the denoising coefficient.
[0070] First, it was noticeable that there were many cases where convergence was not achieved when generating equation (27) using α = 1.0. Even systems with n < 7, where the number of carbon atoms is sufficiently small and expected to be well-trained, sometimes failed to converge. When these cases were examined, unintended hydrogen transfer occurred when n = 2, w = 1.0, or n = 6, w = 16.0, and even after the hydrogen transfer was complete, the convergence conditions were not met and oscillation continued. When n = 6, w = 1.0, or n = 8, w = 8.0, 2.0, 1.0, or 0.0, the reaction did not proceed at all and oscillations occurred with a structure close to the initial state. Similarly, when equation (27) was used, convergence tended to be easier when α = 0.1 was used. Even in this case, convergence was often not achieved when n = 16, but convergence was achieved in many other cases.
[0071] The reason why the reaction does not proceed at all and remains in the initial state is that y t and y d It is possible that the two equations may show contradictory directions. Therefore, we used equation (28) for generation. The results of generation using equation (28) are shown in the table at the bottom of Figure 9. In this case, the convergence conditions were met in many examples even when α = 1.0, and the number of examples that converged even when α = 0.1 was greater than with equation (27).
[0072] We also checked whether the results were as specified and summarized them in Figures 10 and 11. Figure 10 shows a table summarizing whether the final state was a rotation of the central CC bond compared to the initial state. Figure 11 shows a table in which check marks are placed only for irrelevant bond angles that did not change significantly during the reaction pathway.
[0073] Figure 10 is a table showing whether the final product was as specified when generated under various conditions. Checkmarks in Figure 10 indicate pairs that were generated correctly. In each table, the horizontal axis (w) represents the strength of classifier-free guidance (Equation 30). The vertical axis (n) represents the number of carbon atoms in polyethylene. The tables in the upper left and upper right of Figure 10 were generated using Equation (27). The tables in the lower left and lower right of Figure 10 were generated using Equation (28). The tables in the upper left and lower left of Figure 10 used α = 0.1. The tables in the upper right and lower right of Figure 10 used α = 1.0. Here, α is the denoising coefficient.
[0074] Figure 11 shows a table that examines whether the final product, generated under various conditions, is as specified, does not include any rotations other than the specified bonds, and does not involve any bond changes. In Figure 11, check marks indicate pairs that were generated correctly. In each table, the horizontal axis (w) represents the strength of classifier-free guidance (Equation 30). The vertical axis (n) represents the number of carbon atoms in polyethylene. The tables in the upper left and upper right of Figure 11 are generated using Equation (27). The tables in the lower left and lower right of Figure 11 are generated using Equation (28). The tables in the upper left and lower left of Figure 11 use α = 0.1. The tables in the upper right and lower right of Figure 11 use α = 1.0. Here, α is the denoising coefficient.
[0075] When w = 0.0, which does not use classifier-free guidance, the reaction was derived ignoring the specified conditions, and rotation of the central CC bond was not derived. When using formula (27) with α = 0.1, the results were clear: the larger w was, the easier it was to satisfy the conditions. Furthermore, the final state satisfied the given conditions up to n = 12, which exceeded the size of the training data. On the other hand, when α = 1.0 was used with formula (27), there were very few examples of generation that satisfied the conditions. When formula (28) was used, even when α = 1.0, there were more examples that satisfied the conditions than formula (27). In summary, when formula (27) was used with α = 0.1, the generation results of conditional generation for small molecules behaved in an easy-to-handle manner. On the other hand, when using a trained model in practical use, it is desirable to complete generation in finite steps even for data that exceeds the training data. For such purposes, formula (28) can be used.
[0076] 4.4 Generation example Figure 12 shows an example of the results of conditional generation with bond changes. Equation (27) without orthogonalization was used for generation. α was set to 0.1. Also, classifier-free guidance with w = 4 was used (Equation (30)). The molecule was C included in the YARP dataset. 8 H 9 The reaction of O and a typical Diels Alder reaction were selected. The data used for the train was Transition 1x, which only included up to C7. Nevertheless, C 8 H 9 We were able to generate a reaction in this molecule, O.
[0077] FIG. 12 shows the C 8 H 9This figure shows an example of the conditional generation results for the Diels Alder reaction with O. In (1) in FIG. 12, c is determined so that the third atom and the eighth atom bond. In (2) in FIG. 12, c is determined so that the seventh atom and the seventeenth atom and the zeroth atom and the eleventh atom bond, and the seventeenth atom and the eleventh atom dissociate. In (3) in FIG. 12, the seventeenth atom and the seventh atom separate, and the seventeenth atom and the ninth atom bond. In (4) in FIG. 12, the zeroth atom and the second atom bond. In (5) in FIG. 12, the eleventh atom and the third atom, and the tenth atom and the third atom bond.
[0078] The first reaction in Figure 12 involves dihedral angle rotation and bond formation. Creating the initial values for such a reaction path without a neural network requires meticulous manipulation of the dihedral angle rotation and bond formation, which is extremely labor-intensive. Using this neural network, this reaction can be created in just a few tens of seconds, simply by specifying the bonds. Furthermore, the first reaction in Figure 12 involves dihedral angle rotation, and the reaction coordinate is significantly curved from the Cartesian coordinate system. Nevertheless, a qualitatively correct reaction path was generated. The other reactions included in Figure 12 include a wide variety of reactions. The second reaction in Figure 12 is a hydrogen addition reaction to a double bond. The third reaction in Figure 12 is a hydrogen transfer reaction. The fourth reaction in Figure 12 is a bimolecular bonding reaction. The fifth reaction in Figure 12 is a Diels-Alder reaction. These reactions were also generated qualitatively correctly.
[0079] 4.5 Conditional Generation for Diverse Data Conditional generation was performed by changing the bonds of the test data 833 separated from the training data. Equation (27) without orthogonalization was used with α = 0.1, and equation (30) was used for conditional generation. The larger w was, the more reactions proceeded that satisfied the conditions for the data. The results are shown in Figure 13.
[0080] In Fig. 13, w is a parameter of classifier-free guidance. In Fig. 13, number is the number of reaction paths whose generation results satisfy the generation conditions. In Fig. 13, ratio is the number of successful paths divided by the number of test data.
[0081] Among the results obtained using the classifier-free guidance with w = 8.0, those that met the given conditions were optimized using the string method at the PFP v4.0.0 level, and the activation barriers of the corresponding data in the test dataset were compared. In this comparison, if the output is a qualitatively identical pathway, it is expected that the same activation barriers will be obtained. The results are shown in Figure 14(a). That is, Figure 14(a) shows an example of a comparison of the activation barriers of reactions in the test dataset and the activation barriers of the generated optimization results. Furthermore, Figure 14(b) shows a histogram of the energy difference between the generated and optimized RP and the test RP. That is, Figure 14(b) shows a histogram of the energy difference between the test data and the optimization results.
[0082] Most reactions show the same activation barrier as the test dataset and are expected to follow qualitatively similar pathways. On the other hand, some reactions follow pathways completely different from those in the test dataset, resulting in qualitatively different activation barriers. In general, even if the bond changes are the same, qualitatively different RPs exist. Therefore, it is difficult to generate reactions that are exactly the same as the test data. In this experiment, we sometimes obtained reactions with higher activation barriers than the test data, but we also sometimes obtained pathways with lower barriers than the test data. Therefore, it is believed that the model obtained here has learned well to the extent that can be confirmed in this experiment.
[0083] 4.6 Random Generation White noise was added during generation. Equation (27) was used for generation. The structures were optimized before and after generation, and bond changes and rotations were analyzed. The number of times different changes occurred when generated from the same initial state was counted. Reaction paths were randomly generated 16 times from the IS of the test dataset, and bond and rotation changes between the IS and FS were analyzed. The number of reactions resulting in different changes generated from each IS was counted. Here, the number 16 times is an artificial cutoff that was arbitrarily selected. Due to the large number of test datasets, only 16 generations per dataset were performed to reduce the computational cost of the experiment. dt was set to 0.1, g to 0.1, and α to 0.1. The histogram of the results is shown in Figure 15.
[0084] Figure 15 is a plot of the types of final-state structures obtained when random generation was performed 16 times from a certain initial state. The horizontal axis of Figure 15 represents the types of final-state structures obtained, and the vertical axis of Figure 15 represents the number of initial states that gave the number of final-state structures.
[0085] In Figure 15, the values on the horizontal axis are most often 1 or 2, which indicates that there are many cases where the same reaction is obtained no matter how many times the IS is generated. However, there are also many cases where the values on the horizontal axis are higher. In these cases, many different reactions can be obtained from the same IS.
[0086] 5. Conclusions A machine learning model for reaction path generation has been proposed. This model allows for the generation of a rough sketch of the entire reaction path with a small number of neural network evaluations. This model can handle reaction paths in 3N-dimensional space and can generate complex reactions such as organic reactions. This model learned and generalized the results of Transition1x, generating reactions similar to the test data. This model can be used for both conditional and random generation. In this experiment, conditional generation utilizing bond changes was performed. By appropriately using classifier-free guidance, a reaction path was generated that closely reproduced the bond changes in the test data. Furthermore, by incorporating randomness into the generation, a wide variety of reactions were generated from each initial state.
[0087] An example of this embodiment has been described above. A specific example of an inference device that realizes the above processing according to this embodiment will now be described. The inference device is more generally realized by an information processing device.
[0088] 16 is a block diagram showing an example of the hardware configuration of an information processing device 1 that realizes an inference device according to this embodiment. As shown in Fig. 16, the information processing device 1 may be connected to an external device 9A via a communication network 5. The information processing device 1 may also include an external device 9B connected via a device interface 39. The information processing device 1 receives input of the arrangement (coordinates) of multiple atoms related to the generation of a compound composed of multiple atoms input by a user.
[0089] The information processing device 1 includes a computer 30 and an external device 9B connected to the computer 30 via a device interface 39. The computer 30 includes, for example, a processor 31, a main storage device (memory) 33, an auxiliary storage device (memory) 35, a network interface 37, and a device interface 39. The information processing device 1 may be realized as the computer 30 in which the processor 31, the main storage device 33, the auxiliary storage device 35, the network interface 37, and the device interface 39 are connected via a bus 41.
[0090] Although the computer 30 shown in FIG. 16 includes one of each component, it may also include multiple of the same component. Also, while FIG. 16 shows one computer 30, the software may be installed on multiple computers, with each of the multiple computers executing the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates via a network interface 37 or the like to execute the processing. In other words, the information processing device 1 in this embodiment may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize various functions described below. Furthermore, information transmitted from a terminal may be processed by one or more computers on a cloud, and the processing results may be transmitted to a terminal such as a display device (display unit) corresponding to the external device 9B.
[0091] Various computations of the information processing device 1 in this embodiment may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, various computations may be distributed to multiple processor cores within a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. disclosed herein may be executed by at least one of a processor and a storage device provided on a cloud that can communicate with the computer 30 via a network. Thus, various functions described below in this embodiment may be implemented in the form of parallel computing using one or more computers.
[0092] The processor 31 may be an electronic circuit (such as a processing circuit, processing circuitry, CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit)) that includes a control device and an arithmetic device of the computer 30.
[0093] The processor 31 may also be a semiconductor device or the like including a dedicated processing circuit. The processor 31 is not limited to an electronic circuit using electronic logic elements, but may also be realized by an optical circuit using optical logic elements. The processor 31 may also include a calculation function based on quantum computing.
[0094] The processor 31 performs arithmetic processing based on data and software (programs) input from each device, etc., configured internally of the computer 30, and can output the arithmetic results and control signals to each device, etc. The processor 31 may control each component constituting the computer 30 by executing the OS (Operating System) of the computer 30, applications, etc.
[0095] The information processing device 1 in this embodiment may be realized by one or more processors 31. Here, the processor 31 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other via wire or wirelessly.
[0096] The main memory device 33 is a memory device that stores instructions executed by the processor 31 and various data, and information stored in the main memory device 33 is read by the processor 31. The auxiliary memory device 35 is a memory device other than the main memory device 33. Note that these memory devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory device for saving various data used in the information processing device 1 according to this embodiment may be realized by the main memory device 33 or the auxiliary memory device 35, or may be realized by an internal memory built into the processor 31. For example, the memory unit in this embodiment may be realized by the main memory device 33 or the auxiliary memory device 35.
[0097] The main storage device 33 or the auxiliary storage device 35 stores the trained model. The trained model may be generated by learning an e3-equivalent GNN or Attention to improve generalization, for example. The trained model includes coordinates of atoms in an initial state before a reaction of a plurality of atoms in compound generation, and coordinates r of atoms before inferring a reaction path in compound generation. ij and the atomic number of the atom, Z i The coordinates of an atom in the initial state and the atomic number of the atom may be input to the trained model, or the coordinates and atomic number of the atom updated from the coordinates of the atom in the initial state may be input. The coordinates of an atom and the atomic number of the atom are also referred to as atomic information. The coordinates of an atom may be relative coordinates, information capable of expressing relative coordinates, absolute coordinates, or information capable of expressing absolute coordinates, and the atomic number of an atom may be the type of atom.
[0098] The trained model includes a feature vector c that specifies the reaction to be generated. ij may be further input. In this case, the trained model corresponds to a neural network for conditional generation. For conditional generation, a known classifier-free guidance may be included. ij is a feature vector related to the reaction pathway, and is also called a feature vector.ij is a vector indicating a feature quantity generated in an intermediate layer of a classifier that is configured, for example, by a neural network that classifies reaction pathways. ij may be any condition vector related to a reaction path corresponding to bond separation / dissociation, rotation, etc., and may be expressed as a one-hot vector. ij may be a vector related to the latent space of the VAE, or may be a feature of the Large Language Model (LLM).
[0099] The main storage device 33 or the auxiliary storage device 35 stores the output from the trained model. The output from the trained model is, for example, a tangent vector y tan and the perpendicular vector y prp and y, which indicates the uncertainty of the output from the trained model. std and the stopping condition y for stopping the determination of the reaction path fin In addition, the tangent vector y tan and the perpendicular vector y prp and y, which indicates the uncertainty of the output from the trained model. std and the stopping condition y for stopping the determination of the reaction path fin is also referred to as the information necessary for determining the reaction path. tan is a vector that indicates a tangent to the solid line reaction path, i.e., the reaction path toward the final state of the reaction in the production of the compound, as shown by the dashed line in Figure 1. prp is the tangent vector y tan The vector is a vector along the foot of a perpendicular line from the starting point of the reaction to the reaction path. Note that these vectors are information on the starting point, direction, and length, and may be information that can be used to construct a vector.
[0100] y output from the trained model std is data indicating the uncertainty of the output from the trained model. That is, y std The inverse of y indicates the reliability of the output from the trained model. stdare, for example, the diagonal and / or off-diagonal components of a covariance matrix relating to the difference between output data output from a neural network to be trained when training data is input into the neural network to be trained, and correct answer data corresponding to the training data. The diagonal and / or off-diagonal components of the covariance matrix relate to, for example, the difference between output data output from a neural network to be trained when training data is input into the neural network to be trained, and correct answer data corresponding to the training data. The trained model is generated by training the neural network and the diagonal and / or off-diagonal components so as to minimize a loss function having the number of atoms, the diagonal and / or off-diagonal components of the covariance matrix, and the difference.
[0101] Stopping condition y fin is, for example, a two-component vector, which corresponds to the value when solving a classification problem in a neural network. fin corresponds to information regarding the termination of reaction path determination (hereinafter referred to as termination information). The main storage device 33 or the auxiliary storage device 35 also stores predetermined conditions to be compared with the termination information. The termination information and the predetermined conditions are used to determine whether to execute the repeated process of inputting data to the trained model and determining the coordinates of atoms at the inference destination of the reaction path.
[0102] The trained model (which may be referred to as a first trained model or a path inference trained model) may further output energy corresponding to the coordinates of the atom of the inference target. The main storage device 33 or the auxiliary storage device 35 may further store another trained model (which may be referred to as a second trained model or an energy estimation trained model) different from the trained model. The other trained model may include coordinates of atoms in an initial state before a reaction of multiple atoms in the production of a compound, and coordinates r of atoms before the inference of a reaction path in the production of a compound. ij and the atomic number of the atom, Z i At this time, the main storage device 33 or the auxiliary storage device 35 stores the output from the other trained model. The output from the other trained model is the energy corresponding to the coordinates of the atom to be inferred.
[0103] The main memory 33 or the auxiliary memory 35 stores the tangent vector y tan and the perpendicular vector y prp and the initial state before the reaction path is inferred, an algorithm (hereinafter referred to as a coordinate inference algorithm) for determining the coordinates of the atom to be inferred in the reaction path is stored. The coordinate inference algorithm is an algorithm for solving a differential equation such as Equation (27) or Equation (28). As the coordinate inference algorithm, for example, a known numerical solution method such as the Euler method or the Runge-Kutta method can be applied. Therefore, a description of the coordinate inference algorithm will be omitted. The tangent vector y tan and the perpendicular vector y prp The coordinates of the updated atom to be inferred may be determined by inputting the tangent vector y tan and the perpendicular vector y prp The coordinates of the atoms may be input to the trained model, and updated coordinates of the atoms may be output.
[0104] Multiple processors may be connected (coupled) to one storage device (e.g., memory), or a single processor 31 may be connected. Multiple storage devices (memories) may be connected (coupled) to one processor. When the information processing device 1 in this embodiment is configured with at least one storage device (memory) and multiple processors connected (coupled) to this at least one storage device (memory), it may include a configuration in which at least one of the multiple processors is connected (coupled) to at least one storage device (memory). This configuration may also be realized by storage devices (memories) and processors 31 included in multiple computers. Furthermore, it may include a configuration in which the storage device (memory) is integrated with the processor 31 (e.g., a cache memory including an L1 cache and an L2 cache).
[0105] The network interface 37 is an interface for connecting to the communication network 5 wirelessly or via a wire. An appropriate interface, such as one that complies with an existing communication standard, may be used as the network interface 37. Information may be exchanged with an external device 9A connected via the communication network 5 via the network interface 37.
[0106] The communication network 5 may be any one of a WAN (Wide Area Network), a LAN (Local Area Network), a PAN (Personal Area Network), etc., or a combination thereof, as long as information is exchanged between the computer 30 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.
[0107] The device interface 39 is an interface such as a USB (Universal Serial Bus) that directly connects to an output device such as a display device, an input device, and the external device 9 B. The output device may also have a speaker that outputs sound and the like.
[0108] The external device 9A is a device connected to the computer 30 via a network, while the external device 9B is a device directly connected to the computer 30.
[0109] The external device 9A or the external device 9B may be, for example, an input device (input unit). The input device is, for example, a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the computer 30. The external device 9A or the external device 9B may also be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0110] Furthermore, the external device 9A or the external device 9B may be, for example, an output device (output unit). The output device may be, for example, a display device (display unit) such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound, etc. Furthermore, the external device 9A or the external device 9B may be a device that includes an output device, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0111] The external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage device such as an HDD.
[0112] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of the information processing device 1 in this embodiment. In other words, the computer 30 may transmit or receive some or all of the processing results of the external device 9A or the external device 9B.
[0113] 17 is a diagram showing an example of functional blocks realized by one or more processors 31. The processor 31 has, for example, an acquisition unit 311, a calculation unit 313, a determination unit 315, and a judgment unit 317 as functions realized by the processor 31. The functions realized by the acquisition unit 311, the calculation unit 313, the determination unit 315, and the judgment unit 317 are each stored as a program in, for example, the main storage device 33 or the auxiliary storage device 35. The processor 31 can realize the functions related to the acquisition unit 311, the calculation unit 313, the determination unit 315, and the judgment unit 317 by reading and executing each program stored in the main storage device 33 or the auxiliary storage device 35.
[0114] The acquisition unit 311 acquires information about atoms (and / or molecules) related to the generation of a compound. For example, the acquisition unit 311 acquires, from an input device, a list of atoms (and / or molecules) input (specified or selected) by a user via the input device. This list is a list of atoms (and / or molecules) that will be reactants (hereinafter referred to as a reaction target list). The reaction target list may further include data on the coordinates of the atoms (and / or molecules) before they react to form the compound. Note that acquisition of the reaction target list by the acquisition unit 311 is not limited to acquisition from an input device; it may also be acquired from a terminal device corresponding to the external device 9A via the communication network 5, or from an external memory corresponding to the external device 9B via the device interface 39.
[0115] The calculation unit 313 calculates the coordinates of atoms in the initial state before the reaction of a plurality of atoms in the production of a compound, and the coordinates r of atoms before the inference of the reaction path in the production of the compound. ij and atomic number Z i The calculation unit 313 inputs the feature vector c ij The calculation unit 313 may further input the tangent vector y output from the trained model. tan and the perpendicular vector y prp and y, which indicates the uncertainty of the output from the trained model. std and the stopping condition (stopping information) y fin The calculation unit 313 stores the output from the trained model in the main storage device 33 or the auxiliary storage device 35.
[0116] When the first trained model or the pathway inference trained model outputs energy corresponding to the coordinates of the atom of the inference target, the calculation unit 313 associates the energy output from the first trained model or the pathway inference trained model with the coordinates of the atom of the inference target and stores them in the main storage device 33 or the auxiliary storage device 35. The calculation unit 313 also stores in the second trained model or the energy estimation trained model the coordinates of the atoms in the initial state before the reaction of multiple atoms in the production of a compound and the coordinates r of the atoms before the inference of the reaction pathway in the production of a compound. ij and the atomic number of the atom, Z i In this case, the calculation unit 313 associates the energy output from the second trained model or the energy estimation trained model with the coordinates of the atom of the inference target, and stores the energy in the main storage device 33 or the auxiliary storage device 35.
[0117] The determination unit 315 determines the tangent vector y tan and the perpendicular vector y prp and the coordinates of the atom in the initial state, the determination unit 315 determines the coordinates of the atom to be inferred in the reaction path. For example, the determination unit 315 reads out a coordinate inference algorithm from the main storage device 33 or the auxiliary storage device 35. Next, the determination unit 315 solves a differential equation such as Equation (27) or Equation (28) using the coordinate inference algorithm. In this way, the determination unit 315 determines the coordinates of the atom to be inferred in the reaction path. In addition, for example, the determination unit 315 calculates the tangent vector y tan and the coordinates of the atoms in the initial state, the coordinates of the inferred updated atoms in the reaction pathway may be determined, and the normal vector y prp The determination unit 315 may determine the updated coordinates of the inference target atom in the reaction path based on the coordinates of the atom in the initial state and the coordinates of the atom in the initial state. The determination unit 315 stores the determined coordinates of the inference target atom in the main storage device 33 or the auxiliary storage device 35.
[0118] The determination unit 317 reads out, from the main storage device 33 or the auxiliary storage device 35, predetermined conditions to be compared with the stop information. The determination unit 317 compares the stop information with the predetermined conditions, and determines whether the stop information meets the predetermined conditions. If the stop information meets the predetermined conditions, the determination unit 317 stops input to the trained model. At this time, the determination unit 317 stops the process of determining the coordinates of the atom at the inference destination of the reaction path, which is performed by the determination unit 315 described below. In other words, if the stop information meets the predetermined conditions, the determination unit 317 stops the repeated process of inputting to the trained model and determining the coordinates of the atom at the inference destination of the reaction path.
[0119] Furthermore, if the stop information has not reached a predetermined condition, the judgment unit 317 executes a repeated process of inputting the information to the trained model. At this time, the determination unit 315 also executes a repeated process of determining the coordinates of the atom at the inference destination of the reaction path. During the repeated process, the trained model may be input with the coordinates of the atom in the initial state and the coordinates of the updated atom, or may be input with the coordinates of the updated atom without inputting the coordinates of the atom in the initial state, or may be input with the coordinates of the atom in the initial state and the coordinates of multiple atoms in the case of multiple updates, or may be input with the coordinates of the atom in the initial state but not the coordinates of the atom in the case of multiple updates, or may be input with the coordinates of the atom in the initial state and the coordinates of one or more latest atoms among the multiple updates.
[0120] The above describes the configuration of the information processing device (inference device) 1. Below, the procedure for the process of inferring a reaction path (hereinafter referred to as reaction path inference process) executed by the information processing device 1 will be described with reference to Fig. 18. Fig. 18 is a flowchart showing an example of the procedure for the reaction path inference process.
[0121] (Reaction Path Inference Processing) (Step S1) The acquisition unit 311 acquires atomic information related to the generation of a compound. The acquisition of atomic information is realized by inputting or selecting a reaction target list to an input device in response to a user instruction, acquiring a reaction target list from a terminal device via the communication network 5, or acquiring a reaction target list from an external memory. The acquisition unit 311 stores the acquired reaction target list in the main storage device 33 or the auxiliary storage device 35. Prior to executing step S2, the calculation unit 313 reads out a trained model (first trained model or path inference trained model) from the main storage device 33 or the auxiliary storage device 35. Note that, prior to executing step S2, the calculation unit 313 may further read out a second trained model (energy estimation trained model) from the main storage device 33 or the auxiliary storage device 35.
[0122] (Step S2) The calculation unit 313 calculates the coordinates of atoms in the initial state before the reaction of multiple atoms in the compound production and the coordinates r of atoms before the reaction path in the compound production is inferred. ij and the atomic number of the atom, Z i and are input to the trained model (first trained model or pathway inference trained model). At this time, the calculation unit 313 inputs a feature vector c ij In addition to the first trained model, the calculation unit 313 may further input the coordinates of the atoms in the initial state and the coordinates r of the atoms before inference into the second trained model. ij and the atomic number of the atom, Z i You can also enter:
[0123] (Step S3) From the trained model, the tangent vector y tan and the perpendicular vector y prp and the stopping condition (stopping information) y fin At this time, the calculation unit 313 outputs the tangent vector y tan and the perpendicular vector y prp and the stopping condition (stopping information) y fin The calculation unit 313 obtains y from the trained model. std y stdThe trained model that outputs the tangent vector y tan and the perpendicular vector y prp and the stopping condition (stopping information) y fin In this case, the calculation unit 313 may use the output (the tangent vector y tan , perpendicular vector y prp , stop condition (stop information) y fin ) which shows the uncertainty of std Or 1 / y, which indicates certainty std may be displayed on the display as auxiliary information to the user.
[0124] In addition, when the first trained model or the path inference trained model outputs energy corresponding to the coordinates of the inferred atom, the calculation unit 313 acquires the energy output from the first trained model or the path inference trained model. The calculation unit 313 may also acquire energy corresponding to the coordinates of the inferred atom from the second trained model or the energy estimation trained model. The calculation unit 313 stores the outputs from the various trained models in the main memory device 33 or the auxiliary memory device 35.
[0125] (Step S4) The determination unit 315 determines the tangent vector y tan and the perpendicular vector y prp and the coordinates of the atom in the initial state, the determination unit 315 determines the coordinates of the atom to be inferred in the reaction path. The determination unit 315 determines the coordinates of the atom to be inferred in the reaction path by solving a differential equation such as Equation (27) or Equation (28) using a coordinate inference algorithm. The determination unit 315 stores the determined coordinates of the atom to be inferred in the main storage device 33 or the auxiliary storage device 35.
[0126] (Step S5) The determination unit 317 reads from the main storage device 33 or the auxiliary storage device 35 a predetermined condition to be compared with the stop information. The determination unit 317 compares the stop information with the predetermined condition and determines whether the stop information meets the predetermined condition. The stop information is, for example, a two-component vector (A, B). The predetermined condition is that the first component A of the two-component vector is smaller than the second component B (A<B). In this case, since the stop condition does not meet the predetermined condition (No in step S5), the processing from step S2 onward is repeated. Furthermore, if the first component A of the two-component vector exceeds the second component B (A>B), the stop condition meets the predetermined condition (Yes in step S5), and the processing of step S6 is executed. Note that if the first component A and the second component B are equal (A=B), whether or not the repeated processing is required is preset according to the implementation.
[0127] (Step S6) The determination unit 315 determines the reaction path of the compound using the coordinates of the atoms in the initial state of the reaction path and the coordinates of multiple inference targets. The determination unit 315 stores the determined reaction path in the main memory device 33 or the auxiliary memory device 35.
[0128] 19 shows an example of a comparative example CE showing a reaction path RP1 generated by a known scan, and an example of a reaction path RP2 generated by this embodiment. For ease of explanation, in FIG. 19, the vertical and horizontal axes represent coordinates (e.g., relative coordinates) of atoms to be reacted. The hatching in FIG. 19 indicates dark hatching at low energy and light hatching at high energy.
[0129] As shown in FIG. 19 , in Comparative Example CE, for example, a reaction path is scanned along the vertical direction while being constrained in the vertical direction. Then, the constraint conditions along the vertical direction are released, and the reaction path RP1 shown in FIG. 19 is obtained. Next, in the Comparative Example, the reaction path is determined by optimizing the reaction path using, for example, the Nudged Elastic Band method. That is, in the Comparative Example, the reaction path is determined by multiple optimization methods, such as scanning and the Nudged Elastic Band method. On the other hand, in the present embodiment EM, the reaction path RP2 shown in FIG. 19 can be determined by repeatedly solving the differential calculus formula (Equation 27) or (Equation 28) through the processing of steps S2 to S5 as shown in FIG. 18 .
[0130] In this embodiment, for example, energy can be output in step S3 for the coordinates of the inferred atom. Fig. 20 is a diagram showing an example of energy at the coordinates of the inferred atom in a reaction path. As shown in Fig. 20, in this embodiment, energy according to the reaction path can be calculated.
[0131] Based on the above, the information processing device (inference device) 1 according to this embodiment inputs the coordinates of atoms in the initial state before a reaction of multiple atoms in the production of a compound, the coordinates of the atoms before inference of the reaction path in the production of the compound, and the atomic numbers of the atoms into a trained model, outputs a tangent vector of the reaction path toward the final state of the reaction in the production of the compound and a perpendicular vector along the foot of the perpendicular line from the starting point of the tangent vector to the reaction path from the start point of the tangent vector, and determines the coordinates of the atom to be inferred on the reaction path based on the tangent vector, the perpendicular vector, and the initial state before the reaction. Furthermore, the trained model in the information processing device (inference device) 1 according to this embodiment further outputs information regarding the termination of the inference of the reaction path, and the information processing device (inference device) 1 according to this embodiment repeats the input to the trained model and the determination of the coordinates of the atom to be inferred until the information regarding the termination satisfies a predetermined condition.
[0132] Furthermore, the information processing device (inference device) 1 according to this embodiment calculates a feature vector c ijis further input to the trained model. Furthermore, the information processing device (inference device) 1 according to this embodiment sequentially determines the coordinates of the atom to be inferred by using a linear combination of the tangent vector and the normal vector and the initial state. Furthermore, the information processing device (inference device) 1 according to this embodiment further determines the coordinates of the atom to be inferred by using a Wiener process of multiple atoms in the generation of a compound.
[0133] For these reasons, the information processing device (inference device) 1 according to this embodiment can accurately infer a reaction path related to the production of a compound in a short time and at a low computational cost compared to Comparative Example CE. As a result, the information processing device (inference device) 1 according to this embodiment can generate reaction paths for a variety of reactions in a realistic time frame by utilizing high-speed calculations and pruning for the inference of a reaction path related to the production of a compound. That is, the information processing device (inference device) 1 according to this embodiment can generate a path closer to a more optimized reaction path because a trained model formed by a neural network can learn a direction that is likely to be a reaction path. Additionally, the information processing device (inference device) 1 according to this embodiment can generate complex paths with fewer instructions, and can generate a variety of paths by using the derivative (randomness) of the Wiener process of the atom in the calculation of the atom of the inference strategy.
[0134] Furthermore, the trained model in the information processing device (inference device) 1 according to this embodiment further outputs energy corresponding to the coordinates of the atom at the inference target. Furthermore, the information processing device (inference device) 1 according to this embodiment inputs the coordinates of the atom at the start of the reaction, the coordinates of the atom before inference, and the atomic number of the atom to another trained model, and outputs energy corresponding to the coordinates of the atom at the inference target from the other trained model. As a result, the information processing device (inference device) 1 according to this embodiment can calculate energy, for example, as shown in FIG. 20 , according to the coordinates of the atom during the inference of the reaction path. As a result, the information processing device (inference device) 1 according to this embodiment can estimate the reaction speed in the production of a compound.
[0135] The inference device 1 has been described above. The generation of a trained model used in the inference device 1 will now be described. The neural network to be trained will be referred to as the training target NN hereinafter. Training of the training target NN is performed using a training device. The hardware configuration of the training device is similar to that shown in FIG. 16 , and therefore will not be described here.
[0136] The training device inputs training data into a learning target NN, and outputs inference data of the inference target by the learning target NN and uncertainty data indicating the uncertainty of the inference data. The training device calculates the difference between the correct answer data corresponding to the training data and the inference data. The training device calculates a loss function including the calculated difference and the uncertainty data. The loss function includes, for example, an exponent of a multivariate normal distribution (multivariate Gaussian distribution) having a covariance matrix with uncertainty data as elements and the difference, and a determinant of the covariance matrix. The training device trains the learning target NN and the uncertainty data so as to minimize the value of the loss function. Note that when a stopping condition y fin If , the trainer may use the cross-entropy error to train the target NN.
[0137] FIG. 21 is a diagram showing an example of functional blocks of a training device implemented by one or more processors 31. Below, functions implemented by the training device will be described using the functional blocks shown in FIG. 21. The processor 31 in the training device has an acquisition unit 321 and a learning unit 323. For the sake of concreteness, the inference data will be defined as the tangent vector y tan and the perpendicular vector y prp and the uncertainty data is y std The following description will be given assuming that:
[0138] The acquisition unit 321 acquires the learning data from, for example, an external device 9A such as an external memory or a server device, or from an external device 9B corresponding to an input device. The learning data includes training data and correct answer data corresponding to the training data. The training data is data input to the learning target neural network. The training data is, for example, information on atoms (and / or molecules) related to the production of compounds. The training data includes a feature vector c ij The correct answer data may have, for example, a tangent vector y tan and the perpendicular vector y prp In addition, when energy is output from the learning target neural network, the correct answer data may have energy. The learning data is set in advance by an input device or by predetermined numerical calculations, etc.
[0139] The learning unit 323 inputs training data to the learning target NN. As a result, the learning unit 323 obtains output from the learning target NN (hereinafter referred to as learning output data). The learning unit 323 calculates the difference between the learning output data and the correct answer data. Furthermore, the learning unit 323 compares the calculated difference with the uncertainty data (y std ) is calculated. The loss function is, for example, calculated using the uncertainty data (y std ) as elements, and the exponent of a multivariate normal distribution (multivariate Gaussian distribution) having the difference, and the determinant |Σ| of the covariance matrix. In this embodiment, the loss function used for training the learning target NN is: fin is a linear combination with equation (26) which shows the cross-entropy error of
[0140] Equation (23) is derived from a multidimensional Gaussian distribution. The diagonal elements of the covariance matrix Σ included in the loss function are the tangent vector y tan and the perpendicular vector y prp and the elements of the vector corresponding to y stdThe features of each term on the right side of formula (23) will be explained. The first term on the right side of formula (23) is a constant term that depends on the number of atoms to be generated in the compound. The second term on the right side of formula (23) indicates that if the covariance matrix Σ is small, the loss (error) in the loss function will be small. The numerator in the third term on the right side of formula (23) indicates that if the difference is small, the loss (error) in the loss function will be small. The denominator in the third term on the right side of formula (23) indicates that if the covariance matrix Σ is large, the loss (error) in the loss function will be small. For example, when the difference is large, the element y of the covariance matrix Σ std Also, when the difference is small, the element y of the covariance matrix Σ std may be small. Also, the element y of the covariance matrix Σ std When is small, the contribution of the difference to the loss function (Equation (23)) becomes large.
[0141] The learning unit 323 trains the learning target NN by applying backpropagation using the loss function to the learning target NN. For example, when training the learning target NN to minimize the loss functions of Equation (24) and Equation (25), a large y std The learning target NN is trained (learned) so as to output a small y std In this way, the learning unit 323 generates a trained model from the learning target NN.
[0142] Based on the above, an information processing device (training device) 1 according to this embodiment inputs training data into a neural network to be trained, outputs inference data to be inferred by the neural network and uncertainty data indicating the uncertainty of the inference data, calculates a loss function including an exponent of a multivariate normal distribution having a difference between the ground truth data corresponding to the training data and the inference data and a covariance matrix having the uncertainty data as elements, and a determinant of the covariance matrix, and trains the neural network and the uncertainty data so as to minimize the value of the loss function. According to this training device 1, by learning the uncertainty data and incorporating it into the loss function, it is possible to more accurately train the neural network to be trained and improve the performance of the trained model.
[0143] When the technical idea of the embodiment is realized by an inference method, in the inference method, at least one computer inputs the coordinates of the atoms in the initial state before the reaction of multiple atoms in the production of a compound, the coordinates of the atoms before the inference of the reaction path in the production of the compound, and the atomic numbers of the atoms into a trained model, outputs a tangent vector of the reaction path toward the final state of the reaction in the production of the compound and a perpendicular vector along the foot of the perpendicular line from the starting point of the tangent vector to the reaction path from the starting point of the tangent vector, and determines the coordinates of the atom to be inferred on the reaction path based on the tangent vector, the perpendicular vector, and the initial state. The procedure and effects of the reaction path inference process related to the inference method are the same as those described in the embodiment, so explanation is omitted.
[0144] Some or all of the devices in the above-described embodiments may be configured as hardware, or may be configured as information processing software (programs) executed by a CPU, GPU, or the like. In the case of software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored on a non-transitory storage medium (non-transitory computer-readable medium) such as a flexible disk, CD-ROM (Compact Disc-Read Only Memory), or USB memory, and the software information processing may be executed by loading the software into the computer 30. The software may also be downloaded via the communication network 5. Furthermore, the software may be implemented in a circuit such as an ASIC or FPGA, so that the information processing is executed by hardware.
[0145] The type of storage medium that stores the software is not limited. The storage medium is not limited to removable media such as magnetic disks or optical disks, but may be fixed storage media such as hard disks or memory. The storage medium may be provided inside the computer or outside the computer.
[0146] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, a-b, a-c, bc, or a-bc. It may also include multiple instances of any element, such as a-a, a-bb-b, a-a-bb-cc-c, etc. Furthermore, it also includes the addition of elements other than the listed elements (a, b, and c), such as having d, as in a-bc-d.
[0147] In this specification (including the claims), when expressions such as "using data as input / based on / according to / in response to" (including similar expressions) are used, unless otherwise specified, this includes cases where various data itself is used as input, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is used as input. Furthermore, when it is stated that a result is obtained "based on / according to / in response to data," this includes cases where the result is obtained based solely on the data in question, as well as cases where the result is obtained in response to other data, factors, conditions, and / or states other than the data in question. Furthermore, when it is stated that "data is output," unless otherwise specified, this includes cases where various data itself is used as output, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is output.
[0148] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that include any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately depending on the context in which they are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in these terms without any restriction.
[0149] In this specification (including the claims), when the expression "A configured to B" is used, it may include a situation in which the physical structure of element A has a configuration capable of performing operation B, and a permanent or temporary setting / configuration of element A is configured / set to actually perform operation B. For example, when element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Furthermore, if element A is a dedicated processor or dedicated arithmetic circuit, etc., it is sufficient that the circuit structure of the processor is implemented to actually execute operation B, regardless of whether control instructions and data are actually attached.
[0150] When used in this specification (including the claims), terms implying containing or possessing (e.g., "comprising," "including," "having," etc.) are intended to be open-ended terms that include cases in which something other than the object indicated by the object of the term is contained or possessed. When the object of such a term implies no quantity or a singular number (e.g., an article such as "a" or "an"), the expression should be construed as not being limited to a specific number.
[0151] In this specification (including the claims), although expressions such as "one or more" or "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.
[0152] In this specification, when a particular advantage / result is described as being obtained with respect to a particular configuration of a certain embodiment, it should be understood that the same advantage / result can also be obtained with one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, etc., and that the effect is not necessarily obtained with the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or states, etc. are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same configuration or a similar configuration.
[0153] When terms such as "maximize" are used in this specification (including the claims), they include finding a global maximum, finding an approximation of a global maximum, finding a local maximum, and finding an approximation of a local maximum, and should be interpreted appropriately according to the context in which the term is used. They also include probabilistic or heuristic approximations of these maxima. Similarly, when terms such as "minimize" are used, they include finding a global minimum, finding an approximation of a global minimum, finding a local minimum, and finding an approximation of a local minimum, and should be interpreted appropriately according to the context in which the term is used. They also include probabilistic or heuristic approximations of these minima. Similarly, when terms such as "optimize" are used, they include finding a global optimum, finding an approximation of a global optimum, finding a local optimum, and finding an approximation of a local optimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these optimum values probabilistically or heuristically.
[0154] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that the hardware may include an electronic circuit or a device including an electronic circuit.
[0155] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices (memories) may store only a portion of the data, or may store the entire data.
[0156] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention derived from the content defined in the claims and their equivalents. For example, in all of the above-described embodiments, when numerical values or formulas are used in the explanation, they are shown as examples and are not limited to these. Furthermore, the order of each operation in the embodiments is shown as an example and is not limited to these.
[0157] With respect to the above embodiments, the following supplementary notes are disclosed as aspects and optional features of the invention. (Supplementary Note 1) An inference device comprising: at least one memory; and at least one processor, wherein the at least one processor repeatedly outputs information necessary for determining a reaction path from the trained model each time atomic information is input to the trained model, and determines the reaction path based on the information necessary for determining the reaction path. (Supplementary Note 2) The inference device according to Supplementary Note 1, wherein the information necessary for determining the reaction path is information capable of constituting at least a vector. (Supplementary Note 3) The inference device according to Supplementary Note 1, wherein the information necessary for determining the reaction path includes at least a tangent vector of the reaction path. (Supplementary Note 4) The inference device according to Supplementary Note 3, wherein the information necessary for determining the reaction path further includes at least a normal vector of the reaction path. (Supplementary Note 5) The inference device according to Supplementary Note 4, wherein the normal vector is a vector along a foot of a normal extending from a starting point of the tangent vector to the reaction path. (Supplementary Note 6) The inference device according to any one of Supplements 1 to 5, wherein the trained model further outputs information regarding the termination of reaction path determination, and the processor repeats inputting information about the atoms to the trained model and outputting information necessary for determining the reaction path from the trained model until the information regarding the termination of reaction path determination satisfies a predetermined condition. (Supplementary Note 7) The inference device according to any one of Supplements 1 to 5, wherein the processor further inputs a feature vector related to the reaction path to the trained model. (Supplementary Note 8) The inference device according to any one of Supplements 1 to 5, wherein the atomic information includes at least the coordinates of the atoms and the atomic number of the atoms. (Supplementary Note 9) The inference device according to Supplementary Note 4 or Supplementary Note 5, wherein the processor sequentially determines the reaction path using at least a linear combination of the tangent vector and the normal vector. (Supplementary Note 10) The inference device according to any one of Supplements 1 to 5, wherein the processor further determines the reaction path using a Wiener process of the atoms.(Supplementary Note 11) The inference device according to any one of Supplementary Notes 1 to 5, wherein the trained model further outputs energy corresponding to the coordinates of the atoms. (Supplementary Note 12) An inference device comprising: at least one memory; and at least one processor, wherein the at least one processor inputs training data to a neural network to be trained, and outputs at least a portion of a covariance matrix relating to a difference between output data output from the neural network and ground truth data corresponding to the training data, and generates the covariance matrix by training the neural network and at least the portion of the covariance matrix so as to minimize a loss function having at least the portion of the covariance matrix and the difference. (Supplementary Note 13) A training device comprising: at least one memory; and at least one processor; wherein the at least one processor inputs training data to a neural network to be trained, thereby outputting inference data to be inferred by the neural network and uncertainty data indicating the uncertainty of the inference data, calculates a loss function including an exponent of a multivariate normal distribution having a difference between ground truth data corresponding to the training data and the inference data and a covariance matrix having the uncertainty data as elements, and a determinant of the covariance matrix, and trains the neural network and the uncertainty data so as to minimize a value of the loss function. (Supplementary Note 14) An inference method comprising: repeating, each time atomic information is input to a trained model, outputting information necessary for determining a reaction path from the trained model, and determining the reaction path based on the information necessary for determining the reaction path.(Supplementary Note 15) An inference device comprising: at least one memory; and at least one processor, wherein the at least one processor inputs a first coordinate of an atom to a trained model, the trained model outputs first information for moving the atom from the first coordinate to a second coordinate, obtains a second coordinate based on the first information, inputs the second coordinate to the trained model, the trained model outputs second information for moving the atom from the second coordinate to a third coordinate, obtains the third coordinate based on the second information, and determines a reaction path of the atom. (Supplementary Note 16) The inference device according to Supplementary Note 15, wherein the processor inputs the first coordinate together with the second coordinate to the trained model. (Supplementary Note 17) The inference device according to Supplementary Note 15, wherein the processor inputs the third coordinate and the first coordinate to a trained model, the trained model outputs third information for moving the atom from the third coordinate to a fourth coordinate, and obtains the fourth coordinate based on the third information, and determines a reaction path of the atom. (Supplementary Note 18) The inference device according to Supplementary Note 17, wherein the processor inputs the second coordinate together with the third coordinate and the first coordinate to a trained model. (Supplementary Note 19) A training device comprising: at least one memory; and at least one processor, wherein the at least one processor trains by repeatedly outputting information necessary for determining a reaction path from the trained model each time atomic information is input to the trained model.
[0158] REFERENCE SIGNS LIST 1 Information processing device 5 Communication network 9A External device 9B External device 30 Computer 31 Processor 33 Main memory device 35 Auxiliary memory device 37 Network interface 39 Device interface 41 Bus 311 Acquisition unit 313 Calculation unit 315 Decision unit 317 Judgment unit 321 Acquisition unit 323 Learning unit
Claims
1. An inference device comprising at least one memory and at least one processor, wherein the at least one processor repeatedly outputs information necessary for determining a reaction path from the learned model each time information of an atom is input to the learned model, and determines the reaction path based on the information necessary for determining the reaction path.
2. The inference device according to claim 1, wherein the information necessary for determining the reaction path is at least information capable of constituting a vector.
3. The inference device according to claim 1, wherein the information necessary for determining the reaction path at least includes a tangent vector of the reaction path.
4. The inference device according to claim 3, wherein the information necessary for determining the reaction path further includes at least a perpendicular vector of the reaction path.
5. The inference device according to claim 4, wherein the perpendicular vector is a vector along the foot of the perpendicular from the starting point of the tangent vector to the reaction path.
6. The learned model further outputs information regarding the termination of the determination of the reaction path, and the processor repeats the input of the information of the atom to the learned model and the output of the information necessary for determining the reaction path from the learned model until the information regarding the termination of the determination of the reaction path satisfies a predetermined condition. The inference device according to any one of claims 1 to 5.
7. The processor further inputs a feature vector regarding the reaction path to the learned model. The inference device according to any one of claims 1 to 5.
8. The information of the atom at least includes the coordinates of the atom and the atomic number of the atom. The inference device according to any one of claims 1 to 5.
9. The processor sequentially determines the reaction path using at least a linear combination of the tangent vector and the perpendicular vector. The inference device according to claim 4 or 5.
10. The processor further uses the Wiener process of the atom to determine the reaction path. The inference device according to any one of claims 1 to 5.
11. The learned model further outputs the energy corresponding to the coordinates of the atom. The inference device according to any one of claims 1 to 5.
12. An inference device comprising at least one memory and at least one processor, wherein the at least one processor inputs training data into a neural network to be learned and outputs at least a part of a covariance matrix regarding a difference between output data output from the neural network and correct data corresponding to the training data, and is generated by training the neural network and at least the part of the covariance matrix so as to minimize a loss function having at least the part of the covariance matrix and the difference.
13. A training device comprising at least one memory and at least one processor, wherein the at least one processor inputs training data into a neural network to be learned, outputs inference data to be inferred by the neural network and uncertainty data indicating uncertainty of the inference data, calculates a loss function including an exponent of a multivariate normal distribution having a covariance matrix having as elements a difference between correct data corresponding to the training data and the inference data and the uncertainty data, and a determinant of the covariance matrix, and trains the neural network and the uncertainty data so as to minimize a value by the loss function.
14. An inference method of repeatedly outputting information necessary for determining a reaction path from a learned model every time information of an atom is input into the learned model, and determining the reaction path based on the information necessary for determining the reaction path.
15. An inference device comprising at least one memory and at least one processor, wherein the at least one processor inputs a first coordinate of an atom into a learned model, the learned model outputs first information for moving the atom from the first coordinate to a second coordinate, acquires the second coordinate based on the first information, inputs the second coordinate into the learned model, the learned model outputs second information for moving the atom from the second coordinate to a third coordinate, acquires the third coordinate based on the second information, and determines a reaction path of the atom.
16. The inference device according to claim 15, wherein the processor inputs the first coordinate into the learned model together with the second coordinate.
17. The processor inputs the third coordinate and the first coordinate into a learned model, the learned model outputs third information for moving the atom from the third coordinate to a fourth coordinate, based on the third information, a fourth coordinate is acquired, and a reaction path of the atom is determined, the inference device according to claim 15.
18. The processor inputs the second coordinate into the learned model together with the third coordinate and the first coordinate, the inference device according to claim 17.
19. A training device comprising at least one memory and at least one processor, wherein the at least one processor trains to repeatedly output, from the learned model, information necessary for determining a reaction path each time information of an atom is input into the learned model.
Citation Information
Patent Citations
KR20230161867A