Message-passing graph neural networks with vector-scalar message passing and run-time geometric operations
The MPGNN system addresses computational inefficiencies in molecular property prediction by integrating vector-scalar interactions and runtime geometry, achieving superior accuracy and scalability in molecular dynamics simulations.
Patent Information
- Application Number
- JP2025503493
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-11-05
AI Technical Summary
Existing computational methods for predicting molecular properties, such as those used in computational chemistry, face challenges including high computational cost, data dependency, and inefficiency in capturing geometric information, limiting their application to large molecules and molecular dynamics simulations.
A message-passing graph neural network (MPGNN) system that incorporates a vector-scalar interactive message-passing mechanism and runtime geometry computation to efficiently encode and update geometric relationships, reducing computational complexity and improving accuracy in molecular property predictions.
The MPGNN system achieves high accuracy in predicting molecular properties with reduced computational cost, outperforming state-of-the-art models on various datasets, including MD17 and QM9, and demonstrates scalability to large molecules like proteins, while providing efficient molecular dynamics simulations.
Smart Images

Figure 2025536172000001_ABST
Abstract
Description
[Background technology]
[0001] background
[0001] In the field of computational chemistry, computer-based techniques have been developed to predict molecular properties through computer simulation. These molecular properties are of great interest in a variety of fields because they can have a wide-ranging impact on the appearance and function of molecules or materials. For example, in the field of drug design, changes in molecular properties can affect the efficacy of drugs. In the field of drug discovery, molecular properties can affect the potential for therapeutic use of naturally occurring substances. In the field of quantum chemistry, quantum mechanical calculations of the electronic contributions to the physical and chemical properties of molecules and materials are a fundamental area of exploration. As discussed below, there remains opportunity for improvement in computational methods for predicting molecular properties that can be applied beyond the field of computational chemistry. Summary of the Invention
[0002] overview
[0002] To address the problems discussed herein, a computerized system and method are provided. In one aspect, the computerized system includes a processor that receives a molecular graph in a message-passing graph neural network (MPGNN) and generates a scalar embedding representing features of the nodes and edges of the graph and a vector embedding representing geometric relationships of the graph. The system processes the scalar embedding through a vector-scalar interactive message-passing mechanism of a message-passing sub-block of the MPGNN, generates scalar information from the scalar embedding, and passes it to an embedding space including the vector embedding. The system updates the vector embedding based on the embedding space including the scalar information and the vector embedding. The system updates the scalar embedding based on a runtime geometry calculation of the geometric relationships encoded in the vector embedding. The system computes an updated molecular graph based on the updated scalars and the vector embedding, and outputs a target molecular property value based on the updated molecular graph.
[0003]
[0003] This Summary is provided to introduce select concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Moreover, the claimed subject matter is not limited to implementations that solve any or all of the disadvantages noted in any part of this disclosure. [Brief explanation of the drawings]
[0004] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] 1 is a schematic diagram of a computing system including a message-passing graph neural network (MPGNN). [Figure 2]
[0005] 2 shows a detailed view of the MPGNN of the computing system of FIG. 1. [Figure 3]
[0006] 2 shows a schematic diagram of the runtime angle calculations of directional unit vectors and relational position bond angle vectors, as well as the runtime dihedral angle calculations performed by the MPGNN of the computing system of FIG. 1 . [Figure 4]
[0007] Figure 1 shows the interatomic distance distribution of the system and MD simulations driven by density functional theory (DFT). [Figure 5]
[0008] The energy landscape of chignolin and an evaluation of the energy predictions from the system, equivariant transformation (ET) model, and molecular mechanics (MM) model are shown in Figure 1 . [Figure 6]
[0009] Figure 1 shows the system visualization and model interpretability. [Figure 7]
[0010] 2 is a table illustrating the performance of the system of FIG. 1 , showing the mean absolute error (MAE) of energies (kcal / mol) and forces (kcal / mol / °A) for seven small organic molecules in MD17. [Figure 8]
[0011] 2 is a table showing the performance of the system of FIG. 1 , showing the MAE of energies (kcal / mol) and forces (kcal / mol / °A) of 10 small organic molecules in rMD17. [Figure 9A]
[0012] 1 shows an overall schematic diagram of the scalar and vector embedding workflow of the Vector-Scalar Interactive Message Passing (ViS-MP) mechanism of the system in FIG. [Figure 9B]
[0013] 9B shows a detailed schematic diagram of the edge fusion graph attention module of the vector-scalar interactive message passing mechanism of FIG. 9A. [Figure 10]
[0014] For the seven molecules in MD17, we show a schematic representation of the system in Figure 1 and the potential energy surfaces (PES) explored by molecular dynamics simulations driven by DFT. [Figure 11]
[0015] 2 shows the system of FIG. 1 and a graph of the time consumption of a molecular dynamics simulation driven by DFT. [Figure 12]
[0016] 2 is a table showing the performance of the system of FIG. 1 , showing the MAE of 12 molecular features with respect to QM9. [Figure 13]
[0017] 1 shows an ablation study of the system in FIG. 1 for aspirin in the MD17 dataset. [Figure 14]
[0018] 1 illustrates a flowchart of a computerized method according to an exemplary implementation of the present disclosure. [Figure 15] 1 shows a flowchart of a computerized method according to an exemplary implementation of the present disclosure. [Figure 16] 1 shows a flowchart of a computerized method according to an exemplary implementation of the present disclosure. [Figure 17]
[0019] 1 illustrates an exemplary computing environment in which embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0005] Detailed Description
[0020] Computer-based techniques have been developed to predict molecular properties through computer simulations. For example, density functional theory (DFT) is a powerful and widely used quantum physics computational technique that can often accurately predict various molecular properties, such as molecular energies and forces, molecular geometry, and so on. However, DFT is time-consuming and computationally intensive, often requiring up to several hours on conventional processors for a single model of a simple molecule. For many complex systems, computing an exact DFT solution is impractical with current hardware. This currently represents a barrier to predicting molecular properties.
[0006]
[0021] Recently, neural network models have been developed for application in the field of molecular dynamics simulation. Although the accuracy of these models has improved as discussed below, these models generally suffer from the drawback of high computational cost. Therefore, there are challenges in widely applying such neural network models to molecular dynamics simulation.
[0007]
[0022] Molecular dynamics (MD) models calculate the potential energy and resulting interatomic forces of each atom in a molecular system as it changes its physical position during simulation time, thereby describing the kinetic and thermodynamic properties of the molecular system. MD is widely used in the fields of physics, chemistry, biology, and medicine. Ab initio MD simulations, such as those driven by DFT, can accurately calculate energies and forces, but due to their high computational cost, the application of such techniques is limited to large molecular systems and long simulation times, as discussed above. In contrast, classical MD simulations using experimental force fields can achieve fast simulation results for large systems, but suffer from the drawback of being unable to capture quantum effects due to electron motion, and force field parameters calculated within such simulations generally cannot be reused.
[0008]
[0023] In recent years, deep learning (DL) has demonstrated its powerful ability to learn from raw data without the need for handcrafted features in many fields, and DL models for computing the potential energy of molecular systems have attracted increasing attention. However, an inherent drawback of deep learning is its requirement for large amounts of data, which poses a barrier to its wider application in more scenarios. To mitigate the data dependency of DL models for computing potential energy, a subfield called geometric deep learning (GDL) can incorporate symmetry inductive biases into the design of neural networks. Here, symmetry describes the conservation of physical laws, i.e., physical properties that remain invariant despite transformations, such as translations or rotations, performed on the underlying data. Due to these limitations, GDL does not require data augmentation and can only be scaled to limited data scenarios.
[0009]
[0024] Within GDL, equivariant graph neural networks (EGNNs) have been proposed to model molecular geometry. One type of EGNN for computing energy potentials realizes equivariance based on group representation theory and utilizes high-order geometric tensors. While this approach exploits geometric information, certain operations, such as the Clebsch-Gordan product (CG product), typically lead to extremely large computational overhead at scale beyond their limits, significantly hindering the practical application of this approach to large molecules. To alleviate this and improve the modeling of directional information, another approach utilizes scalarized angle representations via vector embeddings or dot products of the vector embeddings themselves, thereby capturing equivariance in the model. In such models, the angular and dihedral information of the molecular system is explicitly encoded, thereby reducing computational cost to some extent by avoiding CG product operations. However, even these models suffer from relatively high computational costs and leave room for improvement in terms of accuracy. Furthermore, these models suffer from the drawback of not being robust to various conformations of the molecular system during molecular dynamics simulations. These models appear unable to effectively utilize such geometric information in message passing, which may limit their performance.
[0010]
[0025] To address these issues, a computing system configured to run a message-passing graph neural network (MPGNN) is provided, which exploits directional information to achieve high accuracy at low computational cost. As discussed below, embodiments of a computing system with MPGNN described herein outperformed state-of-the-art methods for molecules in the Molecular Dynamics 17 (MD17) and revised rMD17 datasets and achieved superior prediction scores for 11 of 12 quantum properties on the Quantum Machine 9 (QM9) dataset. In addition, simulation results demonstrate that the embodiments described herein offer the potential technical advantage of being able to scale to protein molecules containing hundreds of atoms while achieving ab initio accuracy (i.e., a level of accuracy approaching the ground truth data on which the model is trained) without molecular segmentation. Evaluations and case studies discussed below demonstrate that these embodiments may be able to efficiently explore the conformational space of small and large molecules alike, while providing reasonable interpretability for mapping geometric representations to molecular structures.
[0011]
[0026] At a high level, MPGNN implements a runtime geometry computation (RGC) function to extract and encode angle and dihedral information with linear computational complexity, significantly speeding up model training and inference and reducing memory consumption. Additionally, it incorporates a vector-scalar interactive message-passing mechanism to effectively utilize geometric information by equivariantly combining vector and scalar hidden representations. By incorporating these two modules, MPGNN can achieve high computational efficiency while fully utilizing geometric information. Comprehensive benchmark evaluations show that the MPGNN described herein outperforms state-of-the-art algorithms for all molecules in the MD17 and revised MD17 datasets and demonstrates superior performance on the QM9 dataset, demonstrating the powerful power of molecular geometric representation. Next, we discuss ab initio molecular dynamics (AIMD) simulations of each molecule in MD17 driven by an MPGNN trained on only 0.7% of the data. The close agreement between the interatomic distance distributions and the explored potential energy surfaces between AIMD and quantum simulations demonstrates that the disclosed MPGNN is capable of performing data-efficient, high-fidelity simulations. To further explore the scalability of MPGNN for large molecules, an all-atom MD dataset of the simplest protein, chignolin, at the DFT level was constructed, consisting of 9,543 distinct conformations of a 166-atom protein derived from replica-exchange molecular dynamics and calculated by DFT. This is believed to be the first MD dataset at the DFT level for a real all-atom protein. When evaluated on this dataset, the disclosed MPGNN also achieved superior performance compared to other deep learning models in predicting potential energy and experimental force fields. Additionally, MPGNN was shown to exhibit reasonable interpretability for mapping geometric representations to molecular structures.
[0012]
[0027] 1 shows a computing system 10 including one or more processors 12 operatively coupled via a data bus 14 to associated memory 16, data storage 18, display 20, and communication interface 22. The one or more processors 12 are configured to execute a message passing graph neural network (MPGNN) 24 during a training phase to train the MPGNN 24 on a training dataset 26, and to execute the trained MPGNN 24 during an inference phase to thereby generate inference results 28. Both the training dataset and the inference results may be stored in data storage 16.
[0013]
[0028] During the training phase, i.e., prior to the inference phase, one or more processors are configured to train on a training dataset 26 that includes a plurality of molecular graphs 30 for different conformational geometries of a molecular system 29 and respective ground truth values of target molecular properties 38 for each molecular graph 30. As discussed below in connection with FIG. 12 , a variety of target molecular properties 38 can be used. For example, the target molecular properties 38 can be energy parameters, force parameters, or dipole moments. In one example, the target molecular properties 38 can be predicted in parallel by MPGNN 24. The ground truth values are typically computed via density functional theory (DFT), although other types of numerical solvers may be used if desired.
[0014]
[0029] As shown in Figure 1, MPGNN 24 includes an embedding block 32, one or more MPGNN processing blocks 34, and an output block 36. A detailed view of the internal structure of each MPGNN processing block 34 is shown in Figure 2. Turning briefly to Figure 2, each MPGNN processing block 34 includes a respective message passing sub-block 40 and a respective update sub-block 42. The message passing sub-blocks 40 are configured to provide a vector-scalar interactive message passing mechanism 44.
[0015]
[0030] Returning to FIG. 1 , during the inference phase, one or more processors 12 are configured to execute the trained MPGNN 24 and provide inference-time input to the MPGNN. To this end, the one or more processors are further configured to receive a molecular graph 30 of the molecular system 29 as input to the MPGNN at inference time. The molecular graph 30 input at inference time is typically the same type of molecule that the MPGNN 24 was trained on during the training phase, but has a different conformational geometry than the training examples contained in the training dataset 26. The molecular graph 30 typically includes nodes N connected by edges E. The nodes N represent atoms, and the edges E represent inter-atomic bonds in the molecular system 29.
[0016]
[0031] The one or more processors 12 are further configured to process the molecular graph 30 using the embedding block 32, as described in more detail below, to thereby generate scalar embeddings 46 that encode scalar information describing characteristics of the nodes N and edges E, and vector embeddings 48 that represent geometric relationships between the nodes and edges of the molecular graph.
[0017]
[0032] Referring briefly to Figure 2, the scalar embedding 46 includes scalar node embeddings 46B, 46C and scalar edge embedding 46A. The scalar node embeddings 46B, 46C encode the type of atom represented by each node in the molecular graph 30. The scalar target node embedding 46C (h i ) encodes such information about target node i, and the scalar source node embedding 46B(h j ) encodes such information for each source node j connected (i.e., molecularly bonded) to target node i by an edge ij. Atom types can be shown, for example, as chemical symbols or another encoding scheme can be used. Scalar edge embedding 46A(f ij) encodes the interatomic distances represented by each edge E in the molecular graph 30. Other scalar information, such as bond types, functional group labels, etc., may also be included as needed. These are exemplary types of scalar information included in these embeddings, and it will be understood that other scalar information may also be encoded in the scalar embedding 46. The vector embedding 48 encodes geometric information including a directional unit vector 48A for each node N and a relative position bond angle vector 48B for each of multiple pairs of nodes connected by a bond in the molecular graph (i.e., an edge in the molecular graph 30). It will be understood that other geometric information may also be included in the vector embedding 48. The geometric information describes the angles of molecular bonds between atoms, and the scalar information describes the distances between atoms and the types of atoms included in the molecular graph 30.
[0018]
[0033] During inference, initial values for the molecular graph are provided as input, and then each MPGNN processing block 34 updates a prediction of changes in the molecular graph's geometry over discrete time. Thus, it will be appreciated that processing a molecular graph using an embedding block involves encoding initial values for scalar node embeddings, scalar edge embeddings, and vector embeddings of the molecular graph via the embedding block. These initial values are then manipulated and updated by MPGNN processing blocks 34 arranged in successive layers during inference.
[0019]
[0034] To that end, and referring to Figure 2, the one or more processors 12 are further configured to process the scalar embedding 46 via a vector-scalar interactive message passing mechanism 44 of the message passing sub-block 40 of the MPGNN 24, thereby generating and passing scalar information from the scalar embedding 46 to an embedding space that includes a vector embedding 48. The vector-scalar interactive message passing mechanism 44 includes a scalar message function 44A and a vector message function 44B, and the embedding space is a latent space in memory that is accessible by both of these functions 44A, 44B.
[0020]
[0035] Referring to FIG. 2 , the one or more processors 12 are further configured to update the vector embeddings 48 based on the scalar information from the scalar embeddings 46 and an embedding space that includes the vector embeddings 48 via a vector embedding update function 52 of the update sub-block 42, thereby generating updated vector embeddings 48A.
[0021]
[0036] The one or more processors 12 are further configured to update the scalar embedding 46 (including the scalar target node embedding 46C and the scalar edge embedding 46A) based on runtime geometry calculations of the geometric relationships encoded in the vector embedding 48 (more specifically, based on its directional unit vector 48A) via a scalar edge embedding update function 54 and a scalar node embedding update function 56 of the update sub-block 42, thereby generating an updated scalar embedding (including an updated scalar edge embedding 46A' and an updated scalar target node embedding 46C') for the target node, and the scalar source node embedding 46B is updated in another pass through the update sub-block if the source node is the target node. The runtime geometry calculations are performed by a runtime geometry calculation function 50 of the update sub-block 42, further details of which are provided below. The runtime geometry calculation function 50 includes angle operation logic 50A and dihedral operation logic 50B. As shown in Figure 3, the angle operation logic 50A is configured to calculate a directional unit vector 48A for the target node and each relational position bond angle vector 48B for the edge connecting the target node to the source node. It will be appreciated that the directional unit vectors 48A are equivariant vector representations of the change in position and therefore energy of the target node i. Typically, they are only calculated for source nodes that are within a threshold distance of the target node, i.e., source nodes that are within the neighborhood of the target node. As further shown in Figure 3, the dihedral operation logic 50B is configured to calculate the dihedral angle θ between two edges based on orthogonally projecting the directional unit vectors 48A for each node i, j along the respective edge onto nodes m, n.
[0022]
[0037] Returning to FIG. 1 , the one or more processors 12 are further configured, via update sub-block 42, to compute an updated molecular graph 30A based on the updated scalar embeddings 46A′, 46C′ and updated vector embeddings 48′ for each node in the molecular graph 30A by independently considering each node N as a target node.
[0023]
[0038] Returning to FIG. 1, the one or more processors 12 are further configured to output, via an output block 36, the value of a target molecular property 38 of the molecular system 29 determined based on the updated molecular graph 30A.
[0024]
[0039] As shown, the value of the target molecular property 38, along with the updated molecular graph 30A, can be output to data storage 16 as inference results 28, as well as to a downstream program 39 for further processing. As an example, the downstream program 39 can be a molecular dynamics simulation program for use in a molecular dynamics simulation 39A over multiple time steps. For example, the potential energy of each node in the updated molecular graph 30A can be calculated, and the potential energies of all nodes can be summed to calculate the total potential energy as the target molecular property 38. Similarly, as described below, atomic forces along each edge can be calculated based on the potential energy at each node.
[0025]
[0040] The one or more processors 12 are configured to process the scalar embeddings 46 via the vector-scalar interactive message passing mechanism 44 of the message passing sub-block 40, at least in part, in the MPGNN processing block loop, by: for each of a plurality of target nodes in the molecular graph 30, for each MPGNN processing block 34, generating an edgewise scalar message 58 for each edge connected to the target node via the scalar message function 44A of the message passing sub-block; passing the edgewise scalar message 58 to the vector message function 44B of the message passing sub-block 40; and generating an aggregated vector message 60 via the vector message function 44B based on the vector embeddings 48 and the edgewise scalar messages 58 from each of the edges connected to the target node, where the vector embeddings 48 and the edgewise scalar messages 58 are aggregated together for the target node via the vector scalar aggregation function 62. Thus, the one or more processors 12 may be configured to aggregate the vector embeddings and scalar embeddings for each target node across source nodes connected to the target node, respectively, in the MPGNN processing block loop before updating the vector embeddings and scalar embeddings to generate respective aggregated scalar messages and aggregated vector messages for the target node. The scalar message encoding information includes or is based on one or more (typically all) of the following: scalar node embeddings for the target node and nearby (i.e., connected, within a threshold distance) source nodes, scalar edge embeddings for the target node, and computed attention scores from a trained graph attention network for one or more scalar node embeddings and scalar edge embeddings for the target node.
[0026]
[0041] The one or more processors are configured to generate the scalar message by, at least in part, fusing the scalar node embedding and the scalar edge embedding, thereby generating a fused scalar embedding. The fusion may be achieved by concatenation, a Hadamard product, adding a learnable bias term, or other suitable techniques. Furthermore, the computation of the attention score may be based on the fused scalar embedding via a nonlinear activation function or other suitable activation function.
[0027]
[0042] In the MPGNN processing block loop, the one or more processors may be configured to update the vector embeddings via the update sub-block by updating the vector embeddings for the target node via the update sub-block based, at least in part, on the aggregated scalar messages and the aggregated vector messages for the target node.
[0028]
[0043] In the MPGNN processing block loop, the one or more processors are configured to update the scalar embeddings via the update sub-block, at least in part, by performing runtime geometry calculations to compute runtime values of relative position bond angle vectors, directional unit vectors, and dihedral angles for each target node via the update sub-block, computing updated scalar node embeddings for the target nodes based on the computed relative position bond angle vectors, aggregated scalar messages for the target nodes, and scalar node embeddings for the target nodes via the update sub-block, and computing updated scalar edge embeddings based on the computed dihedral angles and scalar edge embeddings via the update sub-block. The one or more processors may be configured to compute an updated molecular graph based on the updated scalar edge embeddings, updated scalar node embeddings, and updated vector embeddings.
[0029]
[0044] The following paragraphs provide additional implementation details for a specific embodiment of MPGNN24, referred to as a Vector Scalar Interactive Graph Neural Network (ViSNet), which can be implemented by one or more processors 12 of the computing system 10 described above. In the following paragraphs, reference letters are used sparingly to avoid confusion with mathematical symbols for embeddings, vectors, functions, and their output parameters, generally with reference to Figures 1 and 2. References to ViSNet below should be understood to refer to a specific implementation of MPGNN24, and similar terminology is used to describe similar components as described above. At a high level, ViSNet embeds the 3D structure of molecules, extracts geometric information through a series of ViSNet blocks, and outputs energy and forces through an output block. The energy and forces predicted by ViSNet can be used to drive molecular dynamics simulations. Referring to Figure 2, as briefly discussed above, the ViSNet block consists of two sub-blocks: a message sub-block and an update sub-block. These two sub-blocks implement a vector-scalar interactive message passing mechanism (one of its implementations is referred to as ViS-MP in this specification) that implements Equations 5 to 9. Specifically, the message sub-block first calculates the message function φ m Through the scalar message m ij and vector messages
number
number
number
number
number
number
number
number
number
number
[0030]
[0045] ViSNet is a multi-purpose GDL model for predicting energy potentials. It can predict potential energy, atomic forces, and various quantum chemical properties by taking atomic coordinates and atomic numbers as input. As shown in Figure 1, the ViSNet model consists of an embedding block, multiple stacked ViSNet blocks, and a subsequent output block. The atomic numbers and coordinates are fed into the embedding block, and then into the ViSNet block, which extracts and encodes a geometric representation. The geometric representation is then used to predict molecular properties through the output block. Notably, ViSNet is an energy-conserving potential model, i.e., the predicted atomic forces are derived from the negative gradient of the potential energy with respect to the atomic coordinates.
[0031]
[0046] As shown in Figure 2 and briefly discussed above, each ViSNet block consists of a message sub-block and an update sub-block. These blocks work together as part of a vector-scalar interactive message passing mechanism called ViS-MP. The rich geometric information embedded in the messages passed through ViS-MP is extracted by an RGC function with linear complexity. The RGC function and ViS-MP are described in detail in the following paragraphs.
[0032]
[0047] Turning now to RGC functions, the success of classical force fields indicates that geometric features such as interatomic distances, angles, and dihedra are useful in determining the total potential energy of a molecule. Explicit extraction of invariant geometric representations in prior methods often suffers from the drawback of consuming large amounts of time and memory during model training and inference. When considering atoms, the computational complexity is O(N) as a square of N in the number of neighboring atoms for the calculation of angular information. 2 ), and for the dihedron it scales as N cubed O(N 3 ) is the order of . To alleviate this problem, an RGC function has been proposed that uses an equivariant vector representation called the directional unit vector for each node to preserve its geometric information. RGC directly calculates the geometric information from the directional unit vector, which is simply a single summation of the vector from target node i to the neighboring source node j. Therefore, the computational complexity can be reduced to O(N).
[0033]
[0048] Considering the substructure of an exemplary molecule with four atoms shown in FIG. 3, the angle information of target node i is given by the vector
number
number
number
number
number
number
number
number
[0034]
[0049] Runtime angle calculations as well as vectors
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0035]
[0050] Moving now to the details of ViS-MP, in order to make effective use of geometric information and strengthen the interaction between scalars and vectors, a vector-scalar interactive message passing mechanism (ViS-MP) is designed for intersection nodes and edges for angles and dihedra, respectively. The following operations are performed by ViS-MP:
number
number
number
number
number
number
number
number
number
number
number
number
number
[0036]
[0051] Here, we provide a more complete explanation of how ViS-MP works with reference to Figure 9A. Figure 9A shows the general workflow of scalar and vector embedding. Specifically, one ViSNet block consists of two modules: i) Scalar2Vec, which appends scalar embeddings to vectors, and ii) Vec2Scalar, which refines the scalar embeddings constructed by the RGC function. The input of Scalar2Vec is the node embedding h i , edge embedding f ij , directional unit vector
number
number
number
number
number
[0037]
[0052] In summary, geometric features are extracted by taking dot products with the outputs of the RGC functions, and the scalar and vector embeddings are periodically updated with each other in ViS-MP to learn a comprehensive geometric representation from the molecular graph.
[0038]
[0053] ViSNet can be used to perform accurate quantum chemical property predictions. As evidence, ViSNet has been evaluated for energy, force, and other molecular property predictions on several popular benchmark datasets, including MD17, revised MD17 (referred to as "rMD17"), and QM9. MD17 consists of MD trajectories for seven organic small molecules, with the number of conformations for each molecular dataset ranging from 133,700 to 993,237. The rMD17 dataset reproduces MD17 with higher accuracy. QM9 consists of 12 quantum chemical properties for 133,385 organic small molecules with up to nine heavy atoms. ViSNet has been compared to the results of other state-of-the-art algorithms for molecular property prediction, including the kernel-based algorithms FCHL19 and GAP; the directional information-based algorithms SchNet, ANI, PhysNet, EGNN, ACE, DimeNet / DimeNet++, GemNet, PaiNN, and ET; and the group representation theory-based algorithms UNiTE and NequIP. Details of training ViSNet for each benchmark are provided below.
[0039]
[0054] It is noteworthy that, when comparing all molecules, ViSNet outperformed the algorithms, achieving the lowest mean absolute error (MAE) for both predicted energy and force, as shown in Table 1 in Figure 7 and Table 2 in Figure 8. Table 1 shows the mean absolute error (MAE) for energy (kcal / mol) and force (kcal / mol / °A) for seven organic small molecules in MD17 compared to state-of-the-art algorithms. The best results in each category are highlighted in bold. Table 2 shows the mean absolute error (MAE) for energy (kcal / mol) and force (kcal / mol / °A) for 10 organic small molecules in rMD17 compared to state-of-the-art algorithms. The best results in each category are highlighted in bold. ViSNet was trained using only 950 samples per molecule type, with another 50 samples used for the validation set, yet it still outperformed kernel-based algorithms by a large margin. This indicates that ViSNet's equivariant model design efficiently captures geometric information, thus significantly reducing the need for a large number of training samples. On the one hand, compared to PaiNN and ET, ViSNet incorporates more directional geometric information through the RGC function strategy, which appears to contribute to its improved performance. On the other hand, considering that Gem-Net incorporates angle and dihedral information, ViSNet's superior performance to Gem-Net indicates that ViS-MP can better utilize geometric information during message passing and achieve better results.
[0040]
[0055] Furthermore, ViSNet also achieved superior performance in quantum chemical property prediction in QM9. Extended Data Table 1 in Figure 12 shows the mean absolute error (MAE) for 12 molecular properties in QM9 compared to state-of-the-art algorithms. The best results in each category are highlighted in bold. As shown, ViSNet outperformed the algorithms when comparing 11 of the 12 chemical properties and achieved comparable results for the remaining properties.
[0041]
[0056] Molecular dynamics simulations are one useful application of the potential energy and atomic forces predicted by ViSNet. To evaluate ViSNet as a potential energy prediction model for ab initio molecular dynamics simulations, we created an instance of ViSNet in the ASE simulation framework, trained on only 0.7% of the samples available in MD17 (i.e., 950 samples for model training), and performed ab initio MD simulations for all seven organic molecules. In this analysis, all simulations were performed under the Berendsen thermostat with a time step of τ = 0.5 femtoseconds, with other settings identical to those used in the MD17 dataset.
[0042]
[0057] Figure 4 shows the interatomic distance distributions of molecular dynamics simulations driven by ViSNet and DFT. (a) shows the atom density at a radius r centered on an arbitrary atom. The interatomic distance distribution h(r) is defined as the ensemble average of the atom densities at radius r. Figures 2(b)–(h) show that the distribution obtained from ViSNet is very close to the distribution generated by DFT. The interatomic distance distribution is defined as the ensemble average of the atom densities. (b)–(h) show a comparison of the interatomic distance distributions between ViSNet and DFT simulations for all seven organic molecules in MD17. The interatomic distance distributions shown are derived from AIMD simulations as a potential energy model for ViSNet and ab initio molecular dynamics simulations at the DFT level for each of the seven molecules. ViSNet results are shown using solid lines, and DFT results are shown using dashed lines. As can be seen, the curves obtained from the DFT and ViSNet results are nearly indistinguishable. The structures of the corresponding molecules are shown in the upper right corner of each panel (b)–(h) of Figure 4 .
[0043]
[0058] Figure 10 compares the potential energy surfaces sampled by ViSNet and DFT for these molecules. Figure 10 shows the potential energy surfaces (PES) explored by molecular dynamics simulations driven by ViSNet and DFT for seven molecules in MD17. The potential energy surfaces were plotted to allow for the inspection of conformational ensembles from the simulations driven by DFT and ViSNet. Snapshots of each simulation trajectory were aligned to the initial structure of the corresponding molecule in MD17. Principal component analysis (PCA) was then applied to the coordinates of each conformation. PC1 and PC2 were set as the two axes. The potential energy values were calculated as ΔG(x,y) = k B It was calculated as Tlng(x,y). B is the Boltzmann constant, T is the temperature of the system, and g(x,y) represents the normalized joint probability distribution. The minimum energy value was set to zero. 100 bins were applied on both the x-axis and y-axis to generate the landscape. In each panel, the left and right subfigures show the PES of the same molecule obtained by DFT and ViSNet, respectively.
[0044]
[0059] The consistent potential energy surface shown in Figure 10 suggests that ViSNet can successfully recover the kinetic properties and conformational space from the simulation trajectories, demonstrating its usefulness for practical molecular dynamics simulations. Furthermore, compared to the prohibitive computational cost of DFT, ViSNet dramatically reduces the computational time by 2-3 orders of magnitude, as shown in Figures 11 and 13 (Extended Table 2). These results demonstrate that ViSNet can serve as a potential energy prediction model for high-fidelity molecular dynamics simulations with a small number of training samples and at a much lower computational cost.
[0045]
[0060] Figure 11 shows the time required to run molecular dynamics simulations driven by ViSNet and DFT. Simulations were performed using the ASE framework for seven molecules in MD17. The average time required for each simulation step is recorded in seconds. The y-axis is a base 10 logarithmic scale. As shown, the ViSNet operation takes significantly less time per step on average than the DFT calculation.
[0046]
[0061] Figure 5 shows the visualization of the chignolin energy landscape and the evaluation of energy predictions by ViSNet, equivariant transformation (ET), and molecular mechanics (MM) methods. In (a), the energy landscape of chignolin sampled by REMD is shown. The x-axis of the landscape is the distance between the main chain O on Y2 and the main chain N on G6, and the y-axis is the distance between the main chain O on E4 and the main chain N on T7. Six representative structures were then selected for visualization. Each structure is animated, with residues depicted as stick figures. Histograms show the mean absolute error (MAE) between the energy differences predicted / calculated by ViSNet, ET, and MM and the ground truth calculated by DFT for the corresponding structures. In (b)–(d), we show the energy correlation in the test dataset between the ground truth calculated by DFT and predictions by ViSNet, ET, and molecular mechanics. Each panel shows the corresponding distribution of energy predictions or calculations and ground truth.
[0047]
[0062] By applying ViSNet to real proteins, we can explore its scalability from small organic molecules to large biomolecules. The time complexity of DFT is roughly O(N^3) as the number of atoms. 3Considering the scaling in the order of σ, we employed the simplest protein, chignolin, with 166 atoms, to construct a DFT-level MD dataset for model training and evaluation. To generate the data, we performed 80-ns replica-exchange molecular dynamics (REMD) simulations to sample various folded and unfolded states of chignolin. As a result, 9,543 representative conformations were collected, and nuclear energies and forces were calculated using the Gaussian16 software package. This is believed to be the first DFT-level MD dataset for a real all-atom protein. The data generation process is described in detail below. The chignolin dataset was split into training, validation, and test sets in an 8:1:1 ratio. ViSNet and its comparison models were trained with default settings on a Tesla V100 GPU, achieving the best performance in the evaluations detailed below, including ET, NequIP, and GemNet, on the chignolin dataset. During model training, GemNet failed due to insufficient GPU memory despite setting the batch size to 1, and NequIP suffered from underfitting with default hyperparameters on the Signorin dataset. ViSNet and ET were successfully trained and compared with molecular mechanics (MM). DFT results were used as ground truth. Figure 5(a) shows the free energy landscape of Signorin sampled by REMD, and d Y2-G6 (distance between the main chain O on Y2 and the main chain N on G6) and d E4-T7The energy concentration basin on the left indicates the folded state, and the energy scattering basin on the right indicates the unfolded state. Six representative structures were selected in the low potential energy region in both the folded and unfolded states, and several intermediate states were selected at high potential energies. The energy predictions for the six representative structures were visualized, and ViSNet significantly outperformed ET and MM trained on the same dataset using experimental force fields. Figures 5(b)–(d) show the correlation between the predicted energies by ViSNet, ET, and MM and the ground truth values obtained by DFT for the conformations in the test set. ViSNet achieved the lowest MAE and the highest R2 score. These results suggest that ViSNet has the ability to scale to real proteins with a small training set and can achieve excellent accuracy and efficiency.
[0048]
[0063] To further explore where ViSNet's performance improvement comes from, a comprehensive ablation study was conducted. Specifically, we excluded runtime angle computation logic (without A), runtime dihedral computation logic (without D), and both (without A&D) from the ablation study of ViSNet performance to evaluate the usefulness of each part. Furthermore, several model variations were designed using different message-passing mechanisms based on ViS-MP for scalar and vector interactions. For example, ViSNet-N was designed to directly aggregate dihedral information to intersection nodes, while ViSNet-T was designed to leverage a different form of dihedral computation. The results of the ablation study are shown in Extended Table 2 of Figure 13. Specifically, Extended Table 2 shows the results of the ablation study of ViSNet for aspirin in the MD17 dataset. Extended Table 2 lists the results of ViSNet, its variations without runtime angle calculation (without A), without runtime dihedral calculation (without D), and without both (without A&D), as well as their comparison results for the message-passing variations ViSNet-N and ViSNet-T. The best results are shown in bold. Based on these results, we can conclude that both types of directional geometric information are useful, and that dihedral information contributes somewhat to the final performance. Furthermore, the significant performance degradation from ViSNet-N and ViSNet-T further validates the effectiveness of the ViS-MP mechanism.
[0049]
[0064] Here, we discuss the visualization and interpretability of ViSNet on molecular structures. To illustrate how incorporating geometric features such as angles and dihedra improves ViSNet's expressiveness, we present a model of ViSNet's interpretability by mapping the angle and dihedral representations derived from the dot products of the model's directional unit vectors to the atoms and bonds of the molecular structure. This bridges the gap between ViSNet's geometric representation and molecular structures. Dot products of directional unit vectors extracted from 50 aspirin samples on the validation set.
number
[0050]
[0065] Figures 6(a) and (b) show the clustering results of the embedding after taking the dot product of the directional unit vectors. Figures 6(d) and (c) show the unmapped clustering results for the chemical structure of aspirin, with Figure 6(c) indexing the atoms of aspirin. Interestingly, the embeddings for the cross nodes and edges could be clearly clustered into several clusters, shown in different colors. For example, the carbon atom C 11 and carbon atom C 12 possess different positions and connect to different atoms, but have similar substructures ({C -1 -O2O3C6} and {C -2 -O1O4C 13}), so their dot product
number
[0051]
[0066] Figure 6 shows a visualization to evaluate the model interpretability of ViSNet.
number
number
number
number
[0052]
[0067]
number
[0053]
[0068] Now, moving on to the detailed operation of ViSNet, ViSNet derives molecular properties (e.g., energy) from the current state of atoms.
number
number
number
number
[0054]
[0069] In the embedding block, ViSNet expands direct node and edge embeddings with neighboring nodes. ViSNet first finds the atomic chemical symbol z i and compute the edge representation of edges within a cutoff distance through a radial basis function (RBF). It will be appreciated that RBF cuts off edges beyond a threshold distance, typically expressed in Angstroms, for computational efficiency. Then, we consider atom i, its one-hop neighbor j, and the directly connected edges e within the cutoff distance. ijThe initial embedding of
number
number
number
[0055]
[0070] N(i) denotes the set of one-hop neighbors of node i, and j is one of the neighbors.
number
number
number
number
[0056]
[0071] Referring to FIG. 9A, the Scalar2Vec module calculates vector embeddings using both scalar messages (Equation 5) derived from node and edge scalar embeddings and vector messages (Equation 6) with inherent geometric information.
number
[0057]
[0072] Turning briefly to Figure 9B, a detailed diagram of the edge fusion graph attention module of Figure 9A is shown. The edge fusion graph attention module implements a graph attention network that combines features of a graph neural network with one or more attention layers. As shown, the edge fusion graph attention module includes a multi-head attention mechanism that takes as inputs a scalar target node embedding 46B (as the Q value of the attention mechanism), a scalar source node embedding 46C (as both the K and V values of the attention mechanism), and a scalar edge embedding 46A (as both the K and V values of the attention mechanism), along with the output of a RBF, which is a cosine cutoff of the scalar value of the magnitude of the radial vector along edge ij. The fused representation is passed through a nonlinear activation function, such as a sigmoidal linear unit (SiLU) function, as shown in Equation 11. In addition, the value (V) in the attention mechanism of the graph attention network implemented by the edge fusion graph attention module is fused by edge features before multiplying the attention score weighted by the cosine cutoff, as shown in Equation 12.
number
number
[0058]
[0073] Then, returning to FIG. 9A, the calculated
number
number
number
number
number
[0059]
[0074] Continuing with Figure 9A, in the Vec2Scalar module, node embedding
number
number
number
number
number
number
[0060]
[0075] In the output block, the scalar and vector embeddings of the nodes are updated with multiple gate-equivariant blocks,
number
number
number
[0061]
[0076] In QM9, the molecular dipole is calculated as follows:
number
number
number
number
number
[0062]
[0077] Here, we describe the design of the chignolin dataset. The initial structure for replica exchange molecular dynamics (REMD) simulations was derived from the Protein Data Bank (PDBID: 5AWL). Water molecules in the crystal structure were removed. The FF19SB force field was then applied to describe the interatomic interactions of chignolin with a generalized Born implicit solvent model. The solvent model uses a second modification of the Bondi van der Waals radius set. The Amber20 program CHIR_RST was applied to create a chiral suppression file during the REMD simulation to maintain chiral properties at high temperatures. The system underwent a minimization process during startup, including 500 steepest descent and 500 conjugate gradient cycles. After energy minimization, the system underwent a 200-ps equilibration run at 300, 400, 500, 600, 700, 800, 900, and 1000 K with random initial speeds. The final equilibrated structure was used for REMD simulation at the corresponding temperature. In the production step, each replica was operated for 2 ps and then exchanged to a neighboring temperature. This exchange was performed 5,000 times in each production step, resulting in eight replica temperatures, resulting in a total simulation time of 80 ns. The sampling interval for each simulation trajectory was 0.4 ps, and therefore there were 200,000 points in the trajectory. 10,000 points were uniformly selected from the REMD trajectory to generate a Gaussian16 input file. The potential energy and interatomic forces for each conformation were calculated using the M06-2X function and the 6-31G function. * The integration grid was set to ultra-fine accuracy.
[0063]
[0078] Finally, 9,543 SCF-converged conformations with total potential energies and atomic forces were selected from the chignolin dataset. The distribution of total energies ranged from -2,831,076.155 kcal / mol to -2,830,477.983 kcal / mol, and several representative conformations are shown in Supplementary Figure 2. Note that the total energies do not show a normal distribution, with two peaks corresponding to the folded and unfolded states of chignolin, which increases the difficulty of model training in this dataset.
[0064]
[0079] Regarding the data partitioning scheme, for the QM9 dataset, we followed previous research and randomly split the dataset into 110,000 samples as the training set, 10,000 samples as the validation set, and the rest as the test set. To evaluate the effectiveness of ViSNet on simulated data, ViSNet was trained on MD17 and rMD17 with a limited data set consisting of only 950 uniformly sampled conformations for model training and only 50 conformations for validation.
[0065]
[0080] Furthermore, the entire Signolin dataset was randomly divided into 80%, 10%, and 10% as the training, validation, and test datasets. Six representative conformations were selected from the test set for illustrative purposes.
[0066]
[0081] Regarding the experimental setup, for the QM9 dataset, a batch size of 32 and a learning rate of 1e-4 were used for all features. The mean squared error (MSE) loss was used to train the model. For the molecular dynamics datasets including MD17, rMD17, and chignolin, a combined MSE loss was utilized for energy and force prediction. The weights for the energy loss were set to 0.05 for MD17 and rMD17 and 0.2 for chignolin. The weights for the force loss were set to 0.95 for MD17 and rMD17 and 0.8 for chignolin. The batch size was set to 4, and the learning rates were chosen from 2e-4, 3e-4, and 4e-4 for the different molecules. The cutoff was set to 5 for small molecules in QM9, MD17, and rMD17, but was changed to 4 for chignolin to reduce the number of edges in the molecular graph. Learning rate decay was used when the validation loss no longer decreased. Patience was set to 15 epochs for QM9 and 30 epochs for MD17, rMD17, and Signorin. The learning rate decay coefficient was set to 0.8 for these models. An early stopping strategy was employed to prevent overfitting. The ViSNet model trained on the molecular dynamics dataset had 9 hidden layers, and the embedding dimension was set to 256. For the QM9 dataset, a larger model was used, i.e., the embedding dimension was changed to 512. Experiments were performed on an NVIDIA® 32G-V100 GPU.
[0067]
[0082] 14-16 show a flowchart of a computerized method 100 according to an exemplary implementation of the present disclosure. 100 can be implemented by the hardware and software of the computing system 10 described above, or by other suitable hardware and software.
[0068]
[0083] At 102, the method 100 includes, during a training stage prior to the inference stage, training a message-passing graph neural network (MPGNN) on a training dataset including a plurality of molecular graphs for different conformational geometries of the molecular system and respective ground truth values for a target molecular property for each molecular graph. As shown, the target molecular property may be an energy parameter 102A, a force parameter 102B, or a dipole moment 102C. As discussed above, other parameters are also contemplated.
[0069]
[0084] At 104, the method 100 includes executing a message passing graph neural network (MPGNN) via one or more processors of a computing device. At 106, the method 100 further includes receiving a molecular graph of the molecular system as input to the MPGNN. The molecular graph typically includes nodes connected by edges, where the nodes represent atoms and the edges represent bonds between atoms of the molecular system.
[0070]
[0085] At 108, the method 100 includes processing the molecular graph using MPGNN to generate a scalar embedding that encodes scalar information describing characteristics of the nodes and edges and a vector embedding that represents the geometric relationships between the nodes and edges of the molecular graph. As shown at 110, the scalar embedding may include a scalar node embedding and a scalar edge embedding. The scalar node embedding can encode the type of atom represented by each node, and the scalar edge embedding can encode the interatomic distance represented by each edge. The vector embedding can encode geometric information including the directional unit vector of each node and the relative position bond angle vector of each of multiple node pairs in the molecular graph. At 112, processing the molecular graph at 108 is shown to include encoding initial values for the scalar node embedding, scalar edge embedding, and vector embedding of the molecular graph.
[0071]
[0086] Continuing from step 112 of Figure 14 through step 114 of Figure 15, an MPGNN processing block loop is shown beginning at step 114 and continuing to step 138 of Figure 16, with method 100 looping back until the last MPGNN processing block is reached. Thus, method 100 may further perform steps 114-138 in the MPGNN processing block loop for each of a plurality of target nodes in the molecular graph, for each MPGNN processing block.
[0072]
[0087] At 114, method 100 includes processing the scalar embedding via a vector-scalar interactive message passing mechanism of the MPGNN, thereby generating scalar information from the scalar embedding and passing it to an embedding space that includes the vector embedding. Substeps of the processing at 114 are shown at 116-122. At 116, it is shown that processing the scalar embedding via a vector-scalar interactive message passing mechanism at 114 is accomplished, at least in part, by generating a scalar message via a scalar message function of the MPGNN, where the encoding information of the scalar message is based on one or more of a scalar node embedding for the target node and neighboring source nodes, a scalar edge embedding for the target node, and a computed attention score from a trained graph attention network for one or more scalar node embeddings and scalar edge embeddings for the target node. In 118, it is shown that generating scalar messages can be achieved, at least in part, by fusing scalar node embeddings and scalar edge embeddings, thereby generating a fused scalar embedding, where the fusion is achieved by concatenation, Hadamard products, or the addition of a learnable bias term, and the computation of attention scores is based on the fused scalar embedding via a nonlinear activation function. Other techniques for fusing embeddings and other activation functions that preserve isovariance can be applied as well.
[0073]
[0088] Processing the scalar embedding at 114 may include passing the scalar message to a vector message function of MPGNN at 120 and generating a vector message via the vector message function based on the vector embedding and the scalar message at 122.
[0074]
[0089] At 124, in the MPGNN processing block loop, the method 100 further includes aggregating the vector embeddings and scalar embeddings, respectively, for each target node across source nodes connected to the target node to generate an aggregated scalar message and an aggregated vector message, respectively, for the target node before updating the vector embeddings and scalar embeddings.
[0075]
[0090] At 126, method 100 includes updating the vector embeddings based on an embedding space that includes the vector embeddings and the scalar information from the scalar embeddings. At 128, updating the vector embeddings is accomplished by updating the vector embeddings for the target nodes based, at least in part, on the aggregated scalar messages and the aggregated vector messages for the target nodes.
[0076]
[0091] At 134, method 100 includes updating the scalar embedding based on runtime geometry computations of the geometric relationships encoded in the vector embedding. At 132, updating the scalar embedding is accomplished at least in part by performing runtime geometry computations to compute runtime values of a relative position bond angle vector, a direction unit vector, and a dihedral angle for the target node, computing an updated scalar node embedding for the target node based on the computed relative position bond angle vector, the aggregated scalar message for the target node, and the scalar node embedding for the target node, and computing an updated scalar edge embedding based on the computed dihedral angle and scalar edge embedding.
[0077]
[0092] At 134, the method 100 includes computing an updated molecular graph based on the updated scalar embedding and the updated vector embedding for each node. At 136, an updated molecular graph is computed based on the updated scalar edge embedding, the updated scalar node embedding, and the updated vector embedding.
[0078]
[0093] At 138, the method 100 includes determining whether the last MPGNN processing block is complete, and if so, the method proceeds to step 140. Otherwise, the method 100 loops back to step 114 of Figure 15. At 140, the method 100 includes outputting the value of the target molecular property of the molecular system determined based on the updated molecular graph. At 142, the method may include processing the output value for the target molecular property and / or the output updated molecular graph using a downstream program, such as a molecular dynamics program, as discussed above.
[0079]
[0094] The systems and methods described herein have demonstrated technical advantages over state-of-the-art models, as discussed above, in terms of improved accuracy while reducing computational costs.
[0080]
[0095] In some embodiments, the methods and processes described herein may be coupled to a computing system of one or more computing devices. In particular, such methods and processes may be implemented as a computer application program or service, an application programming interface (API), a library, and / or other computer program product.
[0081]
[0096] 17 illustrates a schematic representation of a non-limiting embodiment of a computing system 600 capable of implementing one or more of the methods and processes described above. The computing system 600 is shown in a simplified manner. The computing system 600 may embody the computer system 10 described above and shown in FIG. 1. The computing system 600 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones), and / or other computing devices, as well as wearable computing devices such as smart watches and head-mounted augmented reality devices.
[0082]
[0097] Computing system 600 includes a logic processor 602, volatile memory 604, and non-volatile storage 606. Computing system 600 may optionally include a display subsystem 608, an input subsystem 610, a communication subsystem 612, and / or other components not shown in FIG.
[0083]
[0098] Logical processor 602 includes one or more physical devices configured to execute instructions. For example, a logical processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical entities. Such instructions may be implemented to perform a task, implement a data type, transform the state of one or more components, achieve a technical effect, or otherwise arrive at a desired result.
[0084]
[0099] A logical processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, a logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of logical processor 602 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, individual components of a logical processor may be distributed among two or more separate devices, be remotely located, and / or be configured for cooperative processing. Aspects of a logical processor may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. In such cases, it will be understood that these virtualized aspects may execute on different physical logical processors of various different machines.
[0085]
[0100] Non-volatile storage 606 includes one or more physical devices configured to hold instructions executable by the logical processor to implement the methods and processes described herein. When such methods and processes are implemented, the state of non-volatile storage 606 may change, for example, to hold different data.
[0086]
[0101] The non-volatile storage 606 may include removable and / or internal physical devices. The non-volatile storage 606 may include optical memory (e.g., CD, DVD, HD-DVD, Blu-ray Disc, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, FLASH memory, etc.), and / or magnetic memory (e.g., hard disk drives, floppy disk drives, tape drives, MRAM, etc.), or other mass storage technologies. The non-volatile storage 606 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It will be understood that the non-volatile storage 606 is configured to retain instructions even when power to the non-volatile storage 606 is interrupted.
[0087]
[0102] Volatile memory 604 may include physical devices including random access memory. Volatile memory 604 is typically utilized by logical processor 602 to temporarily store information during the processing of software instructions. It will be understood that volatile memory 604 typically does not continue to store instructions once power to volatile memory 604 is interrupted.
[0088]
[0103] Aspects of the logic processor 602, volatile memory 604, and non-volatile storage 606 may be integrated together into one or more hardware logic components, which may include, for example, a field programmable gate array (FPGA), a program and application specific integrated circuit (PASIC / ASIC), a program and application specific standard product (PSSP / ASSP), a system on a chip (SOC), or a complex programmable logic device (CPLD).
[0089]
[0104] The terms “module,” “program,” and “engine” may be used to describe aspects of computing system 600 that are typically implemented in software by a processor to perform a particular function using a portion of volatile memory, which function includes transformative operations that specifically configure the processor to perform a particular function. Thus, a module, program, or engine may be instantiated via logical processor 602, which executes instructions held by non-volatile storage 606 using a portion of volatile memory 604. It will be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms “module,” “program,” and “engine” may encompass individual executable files, data files, libraries, drivers, scripts, database records, etc., or groups thereof.
[0090]
[0105] If included, the display subsystem 608 can be used to present a visual representation of the data maintained by the non-volatile storage device 606. The visual representation can take the form of a graphical user interface (GUI). As the methods and processes described herein modify the data maintained by the non-volatile storage device, and thus transform the state of the non-volatile storage device, the state of the display subsystem 608 can likewise be transformed to visually represent the changes to the underlying data. The display subsystem 608 can include one or more display devices utilizing virtually any type of technology. Such display devices can be combined with the logic processor 602, the volatile memory 604, and / or the non-volatile storage device 606 in a shared enclosure, or they can be peripheral display devices.
[0091]
[0106] If included, input subsystem 610 may include or interface with one or more user input devices, such as a keyboard, mouse, touchscreen, or game controller. In some embodiments, the input subsystem may include or interface with selected natural user input (NUI) components. Such components may be integrated or peripheral, and transduction and / or processing of input actions may be handled on-board or off-board. Exemplary NUI components may include microphones for speech and / or voice recognition, infrared cameras, color cameras, stereo cameras, and / or depth cameras for machine vision and / or gesture recognition, head trackers, eye trackers, accelerometers and / or gyroscopes for motion detection and / or intent recognition, electric field sensing components for assessing brain activity, and / or any other suitable sensors. If included, communications subsystem 612 may be configured to communicatively couple the various computing devices described herein with each other and with other devices. The communications subsystem 612 may include wired and / or wireless communications devices compatible with one or more different communications protocols. By way of non-limiting example, the communications subsystem may be configured for communication over a wired or wireless local or wide area network, such as a wireless telephone network or an HDMI with a Wi-Fi connection. In some embodiments, the communications subsystem may enable the computing system 600 to send and / or receive messages to and / or from other devices over a network, such as the Internet.
[0092]
[0107] The following paragraphs discuss several aspects of the present disclosure. According to one aspect of the present disclosure, a computing system is provided. The system may include one or more processors, the one or more processors configured to execute, during an inference stage, a message-passing graph neural network (MPGNN) including an embedding block, one or more MPGNN processing blocks, each including a respective message-passing sub-block and a respective update sub-block, and an output block, the message-passing sub-blocks being configured to provide a vector-scalar interactive message-passing mechanism. The processor may be further configured to receive a molecular graph of the molecular system as input to the MPGNN, the molecular graph including nodes connected by edges, the nodes representing atoms, and the edges representing inter-atomic bonds of the molecular system. The processor may be further configured to process the molecular graph using the embedding block, thereby generating a scalar embedding that encodes scalar information describing characteristics of the nodes and edges, and a vector embedding that represents geometric relationships between the nodes and edges of the molecular graph. The processor may be further configured to process the scalar embedding via a vector-scalar interactive message passing mechanism of the message passing sub-block of the MPGNN, thereby generating scalar information from the scalar embedding and passing it to an embedding space including the vector embedding. The processor may be further configured via an update sub-block to update the vector embedding based on the scalar information from the scalar embedding and the embedding space including the vector embedding. The processor may be further configured via the update sub-block to update the scalar embedding based on runtime geometry computation of the geometric relationships encoded in the vector embedding. The processor may be further configured via the update sub-block to compute an updated molecular graph based on the updated scalar embedding and the updated vector embedding for each node.The processor may be further configured to output, via an output block, the value of the target molecular property of the molecular system determined based on the updated molecular graph.
[0093]
[0108] According to this aspect, the scalar embedding may include a scalar node embedding and a scalar edge embedding, where the scalar node embedding encodes the type of atom represented by each node, the scalar edge embedding encodes the interatomic distance represented by each edge, and the vector embedding encodes geometric information including the directional unit vector of each node and the relative position bond angle vector of each of multiple node pairs in the molecular graph, and processing the molecular graph using the embedding block may include encoding initial values for the scalar node embedding, scalar edge embedding, and vector embedding of the molecular graph via the embedding block.
[0094]
[0109] According to this aspect, the processor may be configured to process the scalar embeddings via a vector-scalar interactive message passing mechanism of the message passing sub-block by, at least in part, in the MPGNN processing block loop, generating, for each of a plurality of target nodes in the molecular graph, a scalar message via a scalar message function of the message passing sub-block for each MPGNN processing block, where the encoding information of the scalar message is based on one or more of a scalar node embedding for the target node and neighboring source nodes, a scalar edge embedding for the target node, and attention scores computed from a trained graph attention network for one or more scalar node embeddings and scalar edge embeddings for the target node; passing the scalar message to a vector message function of the message passing sub-block; and generating, via the vector message function, a vector message based on the vector embedding and the scalar message.
[0095]
[0110] According to this aspect, the processor may be configured to generate the scalar message at least in part by fusing a scalar node embedding and a scalar edge embedding, thereby generating a fused scalar embedding, where the fusion may be achieved by concatenation, a Hadamard product, or the addition of a learnable bias term, and where the computation of the attention score is based on the fused scalar embedding via a non-linear activation function.
[0096]
[0111] According to this aspect, in the MPGNN processing block loop, the processor may be configured to aggregate the vector embeddings and scalar embeddings for each target node, respectively, across source nodes connected to the target node, before updating the vector embeddings and scalar embeddings, to generate respective aggregated scalar messages and aggregated vector messages for the target node.
[0097]
[0112] According to this aspect, in the MPGNN processing block loop, the processor may be configured to update the vector embeddings via the update sub-block by updating the vector embeddings for the target node via the update sub-block based, at least in part, on the aggregated scalar messages and the aggregated vector messages for the target node.
[0098]
[0113] According to this aspect, in the MPGNN processing block loop, the processor may be configured to update, via the update sub-block, the scalar embeddings by, at least in part, performing, via the update sub-block, runtime geometry calculations to compute runtime values of relative position bond angle vectors, directional unit vectors, and dihedral angles for each target node; computing, via the update sub-block, updated scalar node embeddings for the target nodes based on the computed relative position bond angle vectors, aggregated scalar messages for the target nodes, and scalar node embeddings for the target nodes; and computing, via the update sub-block, updated scalar edge embeddings based on the computed dihedral angles and scalar edge embeddings.
[0099]
[0114] According to this aspect, the processor may be configured to compute an updated molecular graph based on the updated scalar edge embedding, the updated scalar node embedding, and the updated vector embedding.
[0100]
[0115] According to this aspect, the processor may be further configured to train the MPGNN during a training stage, prior to the inference stage, on a training dataset comprising a plurality of molecular graphs for different conformational geometries of the molecular system and, for each molecular graph, a respective ground truth value for the target molecular property.
[0101]
[0116] According to this aspect, the ground truth values can be computed via density functional theory.
[0102]
[0117] According to this aspect, the target molecular property may be an energy parameter, a force parameter, or a dipole moment.
[0103]
[0118] According to this aspect, the values of the target molecular properties can be output to a molecular dynamics simulation program for use in the molecular dynamics simulation.
[0104]
[0119] According to another aspect of the present disclosure, a computerized method is provided. The computerized method may include executing a message-passing graph neural network (MPGNN) via one or more processors of a computing device. The computerized method may further include receiving as input to the MPGNN a molecular graph of the molecular system, the molecular graph including nodes connected by edges, the nodes representing atoms, and the edges representing interatomic bonds of the molecular system. The computerized method may further include processing the molecular graph using the MPGNN to generate a scalar embedding that encodes scalar information describing characteristics of the nodes and edges, and a vector embedding that represents geometric relationships between the nodes and edges of the molecular graph. The computerized method may further include processing the scalar embedding via the MPGNN's vector-scalar interactive message-passing mechanism to generate scalar information from the scalar embedding and pass it to an embedding space including the vector embedding. The computerized method may further include updating the vector embedding based on the scalar information from the scalar embedding and the embedding space including the vector embedding. The computerized method may further include updating the scalar embedding based on a run-time geometry calculation of the geometric relationships encoded in the vector embedding. The computerized method may further include computing an updated molecular graph based on the updated scalar embedding and the updated vector embedding for each node. The computerized method may further include outputting values of the target molecular property of the molecular system determined based on the updated molecular graph.
[0105]
[0120] According to this aspect, the scalar embedding may include a scalar node embedding and a scalar edge embedding, where the scalar node embedding encodes the type of atom represented by each node, the scalar edge embedding encodes the interatomic distance represented by each edge, and the vector embedding encodes geometric information including a directional unit vector for each node and a relative position bond angle vector for each of multiple node pairs in the molecular graph, and processing the molecular graph may include encoding initial values for the scalar node embedding, scalar edge embedding, and vector embedding of the molecular graph.
[0106]
[0121] According to this aspect, processing scalar embeddings via a vector-scalar interactive message passing mechanism can be achieved, at least in part, by generating, in an MPGNN processing block loop, for each of a plurality of target nodes in the molecular graph, a scalar message via a scalar message function of the MPGNN for each MPGNN processing block, wherein the encoding information of the scalar message is based on one or more of a scalar node embedding for the target node and neighboring source nodes, a scalar edge embedding for the target node, and a computed attention score from a trained graph attention network for one or more scalar node embeddings and scalar edge embeddings for the target node; passing the scalar message to the vector message function of the MPGNN; and generating a vector message via the vector message function based on the vector embedding and the scalar message.
[0107]
[0122] According to this aspect, generating the scalar message may be achieved, at least in part, by fusing the scalar node embedding and the scalar edge embedding, thereby generating a fused scalar embedding; the fusion may be achieved by concatenation, a Hadamard product, or adding a learnable bias term; and computing the attention score is based on the fused scalar embedding via a non-linear activation function; and in the MPGNN processing block loop, the computerized method may further include, before updating the vector embedding and the scalar embedding, aggregating the vector embedding and the scalar embedding for each target node across source nodes connected to the target node to generate respective aggregated scalar messages and aggregated vector messages for the target node.
[0108]
[0123] According to this aspect, in the MPGNN processing block loop, updating the vector embeddings can be accomplished at least in part by updating the vector embeddings for the target nodes based on the aggregated scalar messages and aggregated vector messages for the target nodes; updating the scalar embeddings can be accomplished at least in part by performing runtime geometry calculations to compute runtime values of relative position bond angle vectors, directional unit vectors, and dihedral angles for each target node; computing updated scalar node embeddings for the target nodes based on the computed relative position bond angle vectors, the aggregated scalar messages for the target nodes, and the scalar node embeddings for the target nodes; and computing updated scalar edge embeddings based on the computed dihedral angles and scalar edge embeddings; and an updated molecular graph can be computed based on the updated scalar edge embeddings, the updated scalar node embeddings, and the updated vector embeddings.
[0109]
[0124] According to this aspect, the computerized method may further include, during a training stage prior to the inference stage, training the MPGNN on a training dataset including a plurality of molecular graphs for different conformational geometries of the molecular system and, for each molecular graph, a respective ground truth value for a target molecular property, where the target molecular property may be an energy parameter, a force parameter, or a dipole moment.
[0110]
[0125] According to another aspect of the present disclosure, a computing system is provided. The system may include one or more processors, the one or more processors being configured to execute, during an inference stage, a message-passing graph neural network (MPGNN) including an embedding block, one or more MPGNN processing blocks, each including a respective message-passing sub-block and a respective update sub-block, and an output block, where the message-passing sub-blocks are configured to provide a vector-scalar interactive message-passing mechanism. The processor may be further configured to receive a molecular graph of the molecular system as input to the MPGNN, the molecular graph including nodes connected by edges, the nodes representing atoms, and the edges representing inter-atomic bonds of the molecular system. The processor may be further configured to process the molecular graph using the embedding block to thereby generate a scalar embedding that encodes scalar information describing characteristics of the nodes and edges, and a vector embedding that represents geometric relationships between the nodes and edges of the molecular graph. The processor may be further configured to update the vector embedding via the update sub-block based on the scalar information from the scalar embedding and the vector embedding. The processor may be further configured, via an update sub-block, to update the scalar embedding based on runtime geometry computation of the geometric relationships encoded in the vector embedding. The processor may be further configured, via an update sub-block, to compute an updated molecular graph based on the updated scalar embedding and the updated vector embedding for each node. The processor may be further configured, via an output block, to output values of target molecular properties of the molecular system determined based on the updated molecular graph.
[0111]
[0126] It will be understood that the configurations and / or techniques described herein are exemplary in nature and are susceptible to many variations, and therefore, these specific embodiments or examples should not be considered limiting. The particular routines or methods described herein may represent one or more of any number of processing strategies. As such, the various acts shown and / or described may be performed in the order shown and / or described, in other orders, in parallel, or omitted. Similarly, the order of the processes described above may be changed.
[0112]
[0127] The subject matter of this disclosure includes all novel and non-obvious combinations and subcombinations of the various processes, systems and configurations, and other features, functions, acts and / or properties disclosed herein, and any and all equivalents thereof.
Claims
1. 1. A computing system including one or more processors, the one or more processors: During the inference stage, Implementing a message passing graph neural network (MPGNN) including an embedding block, one or more MPGNN processing blocks each including a respective message passing sub-block and a respective update sub-block, and an output block, wherein the message passing sub-blocks are configured to include a vector-scalar interactive message passing mechanism; receiving a molecular graph of a molecular system as input to the MPGNN, the molecular graph including nodes connected by edges, the nodes representing atoms and the edges representing bonds between atoms of the molecular system; processing the molecular graph using the embedding block to generate scalar embeddings that encode scalar information describing characteristics of the nodes and edges, and vector embeddings that represent geometric relationships between the nodes and edges of the molecular graph; processing the scalar embedding through the vector-scalar interactive message passing mechanism of the message passing sub-block of the MPGNN, thereby generating the scalar information from the scalar embedding and passing it to an embedding space containing the vector embedding; updating the vector embeddings based on the embedding space including the scalar information from the scalar embeddings and the vector embeddings via the update sub-blocks; updating the scalar embedding based on run-time geometry computation of the geometric relationships encoded in the vector embedding via the update sub-blocks; computing, via the update sub-block, an updated molecular graph based on the updated scalar embedding and updated vector embedding for each node; outputting, via the output block, the value of the target molecular property of the molecular system determined based on the updated molecular graph; 1. A computing system configured to:
2. the scalar embedding includes a scalar node embedding and a scalar edge embedding, the scalar node embedding encodes the type of atom represented by each node, the scalar edge embedding encodes the interatomic distance represented by each edge, and the vector embedding encodes geometric information including a directional unit vector for each node and a relative position bond angle vector for each of a plurality of node pairs in the molecular graph; processing the molecular graph using the embedding block includes encoding, via the embedding block, initial values for the scalar node embedding, the scalar edge embedding, and the vector embedding for the molecular graph; The computing system of claim 1 .
3. The processor, at least in part, In the MPGNN processing block loop, for each of a plurality of target nodes in the molecular graph, for each MPGNN processing block: generating a scalar message via a scalar message function of the message passing sub-block, wherein encoding information of the scalar message is based on one or more of the scalar node embeddings for the target node and neighboring source nodes, the scalar edge embedding for the target node, and attention scores computed from a trained graph attention network for the one or more scalar node embeddings and scalar edge embeddings for the target node; passing the scalar message to a vector message function of the message passing sub-block; generating a vector message based on the vector embedding and the scalar message via the vector message function; and processing the scalar embedding via the vector-scalar interactive message passing mechanism of the message passing sub-block by performing The computing system of claim 2 .
4. The processor, at least in part, fusing the scalar node embedding and the scalar edge embedding, thereby generating a fused scalar embedding. wherein the fusion is achieved by concatenation, Hadamard product, or addition of a learnable bias term, and the computation of the attention score is based on the fused scalar embedding via a non-linear activation function. The computing system of claim 3 .
5. In the MPGNN processing block loop, the processor: configured to aggregate the vector embeddings and scalar embeddings for each target node, respectively, across the source nodes connected to the target node before updating the vector embeddings and scalar embeddings, to generate a respective aggregated scalar message and an aggregated vector message for the target node. The computing system of claim 3 .
6. In the MPGNN processing block loop, the processor at least partially: updating the vector embedding for the target node via the update sub-block based on the aggregated scalar messages and the aggregated vector messages for the target node; configured to update the vector embedding via the update sub-block by The computing system of claim 5 .
7. In the MPGNN processing block loop, the processor at least partially: performing, via the update sub-block, the runtime geometry calculations to compute runtime values of the relative position bond angle vector, the direction unit vector, and a dihedral angle for each target node; computing, via the update sub-block, an updated scalar node embedding for the target node based on the computed relative position bond angle vector, the aggregated scalar message for the target node, and the scalar node embedding for the target node; computing, via the update sub-block, an updated scalar edge embedding based on the computed dihedral angles and the scalar edge embedding; 7. The computing system of claim 6, configured to update the scalar embedding via the update sub-block by performing:
8. 8. The computing system of claim 7, wherein the processor is configured to compute the updated molecular graph based on the updated scalar edge embedding, the updated scalar node embedding, and the updated vector embedding.
9. the processor: During a training phase prior to the inference phase, training the MPGNN on a training dataset comprising a plurality of molecular graphs for different conformational geometries of the molecular system and respective ground truth values for the target molecular property for each molecular graph; The computing system of claim 1 further configured to:
10. The computing system of claim 9 , wherein the ground truth values are computed via density functional theory.
11. The computing system of claim 1 , wherein the target molecular property is an energy parameter, a force parameter, or a dipole moment.
12. The computing system of claim 1 , wherein the values of the target molecular properties are output to a molecular dynamics simulation program for use in a molecular dynamics simulation.
13. Executing a message passing graph neural network (MPGNN) via one or more processors of a computing device; receiving a molecular graph of a molecular system as input to the MPGNN, the molecular graph including nodes connected by edges, the nodes representing atoms and the edges representing bonds between atoms of the molecular system; processing the molecular graph using the MPGNN to generate scalar embeddings that encode scalar information describing characteristics of the nodes and edges, and vector embeddings that represent geometric relationships between the nodes and edges of the molecular graph; processing the scalar embedding via a vector-scalar interactive message passing mechanism of the MPGNN, thereby generating the scalar information from the scalar embedding and passing it to an embedding space containing the vector embedding; updating the vector embeddings based on the scalar information from the scalar embeddings and the embedding space including the vector embeddings; updating the scalar embedding based on run-time geometry computations of the geometric relationships encoded in the vector embedding; computing an updated molecular graph based on the updated scalar embedding and the updated vector embedding for each node; outputting a value of the target molecular property of the molecular system determined based on the updated molecular graph; A computerized method comprising:
14. the scalar embedding includes a scalar node embedding and a scalar edge embedding, the scalar node embedding encodes the type of atom represented by each node, the scalar edge embedding encodes the interatomic distance represented by each edge, and the vector embedding encodes geometric information including a directional unit vector for each node and a relative position bond angle vector for each of a plurality of node pairs in the molecular graph; processing the molecular graph includes encoding initial values for the scalar node embedding, the scalar edge embedding, and the vector embedding for the molecular graph; 14. The computerized method of claim 13.
15. Processing the scalar embedding via the vector-scalar interactive message passing mechanism includes, at least in part, In the MPGNN processing block loop, for each of a plurality of target nodes in the molecular graph, for each MPGNN processing block: generating a scalar message via a scalar message function of the MPGNN, wherein encoding information of the scalar message is based on one or more of the scalar node embeddings for the target node and neighboring source nodes, the scalar edge embeddings for the target node, and attention scores computed from a trained graph attention network for the one or more scalar node embeddings and scalar edge embeddings for the target node; Passing the scalar message to a vector message function of the MPGNN; generating a vector message based on the vector embedding and the scalar message via the vector message function; This is achieved by 15. The computerized method of claim 14.