Efficient scaling of neural network interatomic potential predictions on cpu clusters

By using graph neural network (GNN) molecular dynamics simulations parallelized across multiple CPUs, the problems of computational resource consumption and accuracy in large-scale atomic systems are solved, achieving efficient molecular dynamics simulations on CPUs, which is suitable for designing new materials.

CN114360655BActive Publication Date: 2026-03-31ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing molecular dynamics simulation methods struggle to balance computational resource consumption and accuracy, especially in large-scale atomic systems. Traditional methods are computationally expensive and lack sufficient accuracy, making efficient parallelization on CPUs difficult.

Method used

Molecular dynamics simulations are performed using a graph neural network (GNN) that is parallelized across multiple CPUs. By adjusting the neighbor distance and truncating the summation, force prediction is optimized, reducing computational load and storage requirements, and improving accuracy.

Benefits of technology

It enables efficient molecular dynamics simulations on CPUs, improving the computational speed and accuracy of large-scale atomic systems, and is suitable for designing new materials such as fuel cell units and batteries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360655B_ABST
    Figure CN114360655B_ABST
Patent Text Reader

Abstract

Efficient scaling of neural network interatomic potential predictions on CPU clusters is provided. Element simulations are described using a machine learning system that is parallelized across multiple processors. A multi-element system is partitioned into a plurality of partitions, each partition including a subset of real elements contained within the partition and empty elements outside of the partition that influence the real elements. For each processor of the multiple processors, a force vector is predicted for the real elements within the multi-element system by passing back through a graph neural network (GNN) having a plurality of layers and parallelized across the multiple processors, the prediction including separately adjusting neighbor distances for each of the plurality of layers of the GNN. Physical phenomena are described based on the force vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various aspects of this disclosure typically involve scaling molecular dynamics simulations using neural network force fields (NNFFs) across multiple CPUs, each with limited knowledge of all atoms in the system. Background Technology

[0002] Molecular dynamics (MD) is a computational materials science approach used to simulate the motion of atoms in material systems under realistic operating pressures and temperatures. Methods exist for calculating the underlying atomic forces used in simulating atomic motion. One approach is ab-initial quantum mechanics. This method is highly accurate but also extremely expensive due to the enormous amount of computational resources required. While other methods exist that consume fewer computational resources, they do not offer the same level of accuracy. Summary of the Invention

[0003] A computational approach for element simulation using a machine learning system parallelized across multiple processors is described, based on one or more illustrative examples. The multi-element system is divided into partitions, each containing a subset of real elements within the partition and ghost elements outside the partition that influence the real elements. For each of the multiple processors, a graph neural network (GNN) with multiple layers and parallelized across the processors is used to predict force vectors for the real elements within the multi-element system. This prediction involves adjusting neighbor distances for each of the multiple layers of the GNN. Physical phenomena are described based on these force vectors.

[0004] A computational system for elemental simulation using a machine learning system parallelized across multiple processors is described, based on one or more illustrative examples. The system includes multiple processing nodes, each node comprising a memory storing instructions for a GNN algorithm of molecular dynamics (MD) software and a processor programmed to execute the instructions to perform operations. These operations include: operating on one of multiple partitions, each partition comprising a subset of real elements contained within the partition and empty elements outside the partition that affect the real elements; predicting force vectors for the subset of real elements contained within the partition by backward traversing a GNN having multiple layers and parallelized across multiple processing nodes, the prediction operation comprising adjusting neighbor distances for each of the multiple layers of the GNN; and describing physical phenomena based on combinations of force vectors from the multiple processing nodes.

[0005] According to one or more illustrative examples, a non-transitory computer-readable medium includes instructions to be executed by multiple processing nodes of a parallelized machine learning system. When executed, the instructions cause the system to perform operations including: operating on one of a plurality of partitions, each partition including a subset of real elements contained within the partition and empty elements outside the partition affecting the real elements; predicting force vectors for the subset of real elements contained within the partition by causing backward traversal of a GNN having multiple layers and parallelized across multiple processing nodes, the prediction operation including adjusting neighbor distances for each of the multiple layers of the GNN; and describing physical phenomena based on a combination of force vectors from the multiple processing nodes. Attached Figure Description

[0006] Figure 1 An example of a GNN force field with direct force prediction is illustrated;

[0007] Figure 2 The illustration shows an example architecture of the autogradient graphical neural network (GNN) scheme;

[0008] Figure 3 The illustration shows an example partitioning scheme with local atoms and empty atoms with neighbor enumeration;

[0009] Figure 4 The illustration shows an example of noise as a function of distance from the central atom i;

[0010] Figure 5 The illustration shows example diagrams of noise from distant atoms using various methods;

[0011] Figure 6 The illustration shows an example of reducing the number of neighbors (M) in each layer of a neural network;

[0012] Figure 7 The diagram illustrates an example relationship of memory requirements per hardware (GPU or CPU node) for the GNNFF algorithm in several different parallel CPU configurations on GPUs and CPUs.

[0013] Figure 8 The illustration shows examples of molecular dynamics simulation speed versus the number of atoms in the simulation for different GPU and parallel CPU hardware configurations;

[0014] Figure 9 The illustration shows an example process of element simulation using a machine learning system that is parallelized across multiple CPUs; and

[0015] Figure 10 The illustration shows an example compute node used when performing element simulation in a machine learning system that is parallelized across multiple CPUs. Detailed Implementation

[0016] Embodiments of this disclosure are described herein. However, it is to be understood that the disclosed embodiments are merely examples, and other embodiments may take various forms and alternative forms. The figures are not necessarily to scale; some features may be enlarged or minimized to show details of specific components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching those skilled in the art to employ the embodiments in different ways. As will be understood by those skilled in the art, various features illustrated and described with reference to any of the figures may be combined with features illustrated in one or more other figures to produce embodiments not explicitly illustrated or described. The combinations of illustrated features provide representative embodiments of typical applications. However, various combinations and modifications of features consistent with the teachings of this disclosure may be desired for particular applications or implementations.

[0017] Molecular dynamics (MD) methods are beneficial for studying physical phenomena, such as, but not limited to, ion transport, chemical reactions, and bulk and surface degradation of materials in systems such as devices or functional materials. Non-limiting examples of such material systems include fuel cell units, surface coatings, batteries, water desalination, and water filtration.

[0018] Methods exist for calculating the underlying atomic forces used in simulating atomic motion. For example, ab initio quantum mechanics can be accurate, but it can also be computationally expensive, as applying this method requires a huge amount of computing resources.

[0019] Neural networks have been used to fit and predict quantum mechanical energies. These methods are known as neural network force fields (NNFFs). They use quantum mechanical energy to predict the negative derivative of the energy with respect to the atomic orientation (atomic force). However, these methods are computationally expensive. Given the above, the desired approach is a computational method for calculating atomic forces that provides a sufficient level of accuracy while consuming a reasonable amount of computational resources.

[0020] Molecular dynamics uses atomic orientation (and possibly charge, bond, or other structural information) to calculate interatomic forces for each atom, which is then used to modify atomic velocities in simulations. The resulting atomic trajectories are used to describe physical phenomena such as, but not limited to, ion transport motions in batteries (e.g., lithium-ion batteries) and fuel cell units (e.g., fuel cell unit electrolytes), chemical reactions during the degradation of bulk and surface materials, phase transitions in solid-state materials, molecular binding, and protein folding, for example, in drug design, bioscience, and biochemical design.

[0021] Material properties are determined by their atoms and their interactions. For many properties, length and time scales of interest can be obtained through MD simulations, which are simulations that predict the motion of individual atoms. These simulations involve the following setups: (i) obtaining atomic orientations; (ii) calculating forces based on atomic orientations; (iii) updating velocities based on the calculated forces; (iv) updating atomic orientations based on velocities; and (v) repeating the simulations as needed.

[0022] For computational power, a trade-off may occur between accuracy and computational speed. The most accurate methods of atomic forces consider the electronic structure of the system, such as ab initio density functional theory (DFT). These simulations are known as ab initio molecular dynamics (AIMD). These methods are typically on the scale of N. 3 Where N is the number of atoms in the system, although the size of the tight-binding method can be as small as N, and the size of the coupled clustering method can be as small as N. 7 That's huge. Nevertheless, these calculations are often expensive, thus leaving room for expectations of more efficient computational methods. Due to memory and computational load, such calculations are currently limited to simulations of about 500 atoms within about 1 nanosecond.

[0023] On the other hand, classical molecular dynamics models completely ignore electrons, replacing their role with generalized interatomic force generation functions. These are typically simple functions, such as Lennard-Jones, Buckingham, and Morse, which approximate binding strength and distance, but often fail to adequately describe material properties. Some groups have attempted to increase the complexity of these models (e.g., ReaxFF and charge-balanced schemes) to balance accuracy and computational cost.

[0024] Machine learning of interatomic potentials can be used as a solution that offers first-principles accuracy with faster computation time. Deep neural networks are a particularly popular choice because their large parameter space allows for fitting a wide range of data. A typical training scheme is as follows: (i) defining the neural network architecture; (ii) generating first-principles data via DFT computation, etc.; (iii) training the neural network parameters on the first-principles data; and (iv) using the trained neural network to predict forces in MD simulations.

[0025] Regarding aspect (i), various architectural choices can be made in the definition of a neural network. In one example, the neural network can account for Behler-Parinello features hard-coded into the model. In another example, a graph convolutional network or graph neural network (GNN) can be implemented, where each atom and bond is an element of the graph, and messages are passed between them. As yet another example, a hypergraph transformer model can be utilized, where atoms and bonds are elements of the graph, and messages are passed through a transformer “attention” architecture that weights closer atoms relative to more distant atoms.

[0026] Regarding aspect (iv), various options are also possible in the methods for calculating the force. In the example, a direct prediction model can be used, where the force is predicted rather than necessarily the negative gradient of the scalar energy. As another possibility, a self-gradient model can be used, where the model predicts the total energy and a self-gradient (automatic differentiation) algorithm is used to calculate the force as the negative gradient of the energy relative to the atom orientation.

[0027] Figure 1 An example graph neural network (GNN) 100 with direct force prediction is described. In the illustrated example GNN 100, atom orientations are mapped onto a graph with nodes (v) and edges (e). Key features such as atom chemical type and distance are then extracted (rotation-invariant) and sent to a message passing routine. Unit vectors are also extracted and weighted to ensure the final force is rotationally covariant (i.e., changing the atom by rotation will change the force accordingly). The force predictor is coupled to a thermostat to ensure that a lack of energy does not lead to heat buildup. Therefore, the thermostat is configured to appropriately increase or decrease the kinetic energy.

[0028] Figure 2 The illustration shows an example architecture 200 of a self-gradient GNN scheme incorporating self-gradient forces. In such a method, the input to the neural network is the orientation of each atom. r 0 And any other relevant information. The outputs of the final layer are summed to produce molecular energy. E The prediction. Then the derivative can be calculated. , where its negative value It can be used to act on atoms i Predicting forces on the surface. Assume the total energy of the neural network is a set of orientations. A smooth function guarantees the derivative. Energy saving.

[0029] Here, the forward arrow conveys the prediction of the force, and the backward arrow conveys the auto-gradient function used to predict the force, i.e., the gradient of the energy relative to the absolute orientation. The derivative of each layer is iteratively computed by moving backward through the network to obtain the value for each... i calculate ; and return force prediction .

[0030] GNNs can be trained as deep neural networks. For example, a loss function can be formed. L To predict with force The network weights are compared with the ground truth weights on the labeled dataset, and then gradient descent can be used to optimize the network weights with respect to that loss. Because gradient descent is performed on gradients, this requires calculating higher-order derivatives; therefore, the time per training iteration will be approximately twice that of a feedforward neural network.

[0031] This is a kind of comparison Figure 1 The method shown is relatively more expensive because it requires one derivative for prediction and two derivatives for training, unlike the direct force model, which does not require a derivative for prediction but does require one derivative for training. However, it is more accurate because the thermostat of the GNN is not strictly necessary, and the force is guaranteed to be the derivative of the scalar energy field, except for rotational covariance.

[0032] Turning to aspect (iv), further optimizations can be implemented using trained neural networks to predict forces in MD simulations. For relatively small simulations (<500 atoms), force prediction using GPUs is feasible, just as it is for training. However, for larger systems, neural network predictions for many atoms are unsuitable for GPUs due to memory limitations. Typical hardware architectures have one CPU per node, or perhaps several GPUs per node for state clusters, but the ability to predict hundreds of atoms in a single batch limits the scalability of the network. Therefore, it is desirable to define a method to predict forces acting on the CPU more efficiently. Deep neural networks are typically slow on CPUs because the critical operations of matrix multiplication are not efficiently parallelized. Nevertheless, GPU batch sizes are typically 10–50 atoms (depending on the network size), so for >100–500 atoms, even if using a CPU would reduce speed by 10 times, a parallelized CPU would be more efficient than a single GPU.

[0033] Even for classical potentials, scaling MD simulations across multiple CPUs is a well-studied problem. One such solution is the Large-Scale Atom / Molecular Massive Parallel Simulator (LAMMPS). LAMMPS operates by distributing atoms across multiple CPUs by assigning ownership of any atom within a certain volume in physical space (e.g., a 2×2×3 nm cube) to each CPU. Instead of retaining information about every atom in the system, a CPU operates only on information about the atoms within its corresponding volume, plus information about those atoms adjacent to that volume. LAMMPS utilizes various algorithms to identify the atomic neighbors of each CPU (called a neighbor list). These atomic neighbors include ghost atoms, which refer to atoms owned by different CPUs but close enough that the CPU itself needs information about their orientation. With this information, the force can be predicted individually on each CPU, with minimal communication between CPUs only when an atom moves from ownership (i.e., computational domain) to ownership of another CPU.

[0034] This disclosure relates to one or more methods and hardware for implementing these methods, enabling molecular dynamics simulations on a CPU. These simulations can be used to describe physical phenomena, allowing for the design of new materials, such as those used in fuel cell units, batteries, sensors, etc. As mentioned above, the simulations can utilize techniques such as GNNs. These methods involve a trade-off in accuracy; however, the training process itself typically introduces an error of approximately 100 meV / A in force predictions, and the DFT method itself introduces an error of approximately 10–100 meV / A in force predictions, depending on the system in question.

[0035] In an illustrative example, the deep neural network uses network quantization, which refers to the method of storing floating-point numbers in a low-bit-width form. This makes the network more computationally efficient on the CPU. This is feasible for direct force models, but not for autogradient schemes because the low-bit-width form is not implemented, and inaccuracies will accumulate when gradients are applied.

[0036] In another illustrative example, a deep neural network incorporates a neighbor architecture to increase compatibility with LAMMPS. A typical deep network contains multiple message-passing routines between atoms, such as convolutions, sigmoids, transformer matrices, etc. For the purposes of this example, let's assume we have three iterations, for example, three layers in the network, such as... Figure 2 As shown. Recall that the LAMMPS algorithm divides atoms among CPUs, where each CPU has a subset of atoms and additionally includes information about a certain number of empty atoms necessary to predict the forces on its local atoms. Here is an example.

[0037] Figure 3 An exemplary partitioning scheme 300 with local atoms and empty atoms enumerated by neighbors is illustrated. Figure 3 As shown, the real atoms in the top left corner are divided into four groups, each to be executed on a separate processor. Each group includes information about local atoms (i.e., those atoms inside the box) and empty atoms (i.e., those atoms that are close enough to the local atoms to significantly influence the local atoms with forces above the system error range). An image of a specific processor is shown on the right. The processor knows the real atoms in its box, as well as its first, second, and third-level neighbors. Therefore, in this example, all neighbors up to the third-level neighbor of each local atom are available to the processor; other empty atoms can be safely ignored.

[0038] Figure 4 The diagram illustrates noise as a distance from the central atom. i Example diagram 400 for a function of distance. For example... Figure 4 As shown, the number of empty atoms is poorly scaled. If the typical neighbor distance of a network is, for example, 5 Å (a relatively small amount, assuming the nearest neighbor bond length is only about 1-2 Å), then the radius of an empty atom could, for example, include all atoms extending outwards to 15 Å (three layers), which increases the number of atoms processed per processor. A key aspect of fast CPU prediction is keeping the number of atoms as small as possible. This is because each CPU has a limited amount of processing speed and memory, both of which scale linearly with the number of atoms whose properties need to be computed.

[0039] In addition, such as Figure 4 and Figure 5 As described, the numerical noise in force prediction is high for distant neighbors. Specifically, the autogradient version of the force field (such as...) Figure 2 (As shown) Calculate the energy of each atom And its absolute orientation with respect to the orientation of each atom. Differentiation. Therefore, the action on the central atom... i The total force on is If every orientation in the system is accurately identified, then the scheme will be as accurate as a neural network can achieve. However, if only the three empty atoms mentioned earlier are included, due to their limited knowledge of their own environment, Ej This will be inaccurate. Furthermore, if we use six layers of empty atoms, allowing us to fully understand the atoms... j In such an environment, computing costs will skyrocket.

[0040] Figure 5Figure 500 illustrates example noise from distant atoms using various methods. This example, showing noise from distant atoms, illustrates this second point. In a single system where the orientation of each atom is precisely known, the force can be seen to decay with distance, as expected. However, in LAMMPS implementations, which consist of limited knowledge per processor, noise accumulates for distant atoms. This unexpected effect is not described in the literature and gives the naive expression... Significant obstacles were presented.

[0041] To address these two issues, two approaches are provided for NN architectures. These two approaches can be used to improve the scaling of NN atomic potential prediction for CPU clusters and are the focus of the remainder of this disclosure.

[0042] First, instead of summing the same number of neighbors at each layer of the network, the neighbor distances are adjusted. In one implementation, such as... Figure 6 As shown in Example 600, the network has only 4 neighbor atoms in the first layer, 8 in the second layer, and 16 in the third layer. In other implementations, scaling can be implemented using distance; for example, the network's first layer includes bonds within 3 Å, the second layer within 5 Å, and the third layer within 7 Å. This has minimal impact on network accuracy but significant impact on prediction speed for two reasons. First, prediction speed involves a size of... Matrix multiplication, where M It is the number of neighbors included in the layer, and N This is the number of atoms in the batch. Therefore, M The reduction in [the distance between atoms] will improve computational efficiency superlinearly. Second, the distance between empty atoms does not need to be as large because the farthest atoms that need to transfer information to the local processor are significantly closer.

[0043] The second approach to this solution is to truncate the sum. Instead of including each atom in the summation. j Only included in atoms i Or contains atoms i Within a certain cutoff section of the processor j The value of . As some non-limiting examples, the cutoff part can be based on distance (e.g., 3 Å, 5 Å, 7 Å, etc.) and / or the number of neighbors (e.g., 16 nearest neighbors). Thus, noise from distant atoms is truncated in the model, which significantly improves accuracy when parallelized across multiple processors.

[0044] Figure 7Example 700 illustrates the memory requirements per hardware (GPU or CPU node) for the GNNFF algorithm in several different parallel CPU configurations. For a CPU with large memory nodes, 20 cores / node are used. For a CPU with regular memory nodes, 2 cores / node are used. The dashed line indicates the hard cutoff of the maximum available memory in a given GPU / CPU node.

[0045] Therefore, scalability of the parallel CPU algorithm used for GNNFF can be observed. This scalability is demonstrated by tests on LiPS superionic conductor material systems. Figure 7 As can be seen, for GPUs, the memory requirements of the algorithm scale approximately linearly (n=1), but for parallel CPUs, they scale only sublinearly. n <1, approximately The increase is approximately due to the additional memory overhead caused by the use of empty atoms in each individual computation core. Scaling is performed, although the scaling is still based on the increasing number of atoms / nodes, it does not scale linearly with the number of atoms. Once the number of computational cores increases, it facilitates... n The scaling is expected to be linear.

[0046] Figure 8 The illustration shows examples of the speed of MD simulations in Molecular Dynamics 800 compared to the number of atoms for different GPU and parallel CPU hardware configurations. For a CPU with large memory nodes, 20 cores / node were used. For a CPU with regular memory nodes, 2 cores / node were used. The dashed line indicates the approximate maximum number of atoms we can process in the different available hardware configurations. Logically, the simulation speed is affected by the number of atoms. While MD speed decreases with increasing number of atoms (n=-1), for the range of atoms of interest in large-scale MD applications, the speed decrease for MD processed by parallel CPUs is slower than other techniques.

[0047] Figure 9 The illustration depicts an example process 900 for elemental simulation using a machine learning system parallelized across multiple CPUs. For instance, the computational system for elemental simulation could utilize multiple processing nodes, each including a memory storing instructions for a GNN algorithm of molecular dynamics (MD) software and a processor programmed to execute those instructions. These instructions would enable the system to perform the operations of process 900 according to the method described above.

[0048] At operation 902, the multi-element system is divided into multiple partitions. Each partition includes a subset of the real elements contained within the partition and empty elements outside the partition that affect the real elements. (Refer to the above.) Figure 3 An example partition is shown.

[0049] At operation 904, for each of the multiple processors, is the real-element prediction force vector for one of the partitions of the multi-element system. Prediction can be performed by making the prediction traverse backwards through a GNN with multiple layers. Prediction can include adjusting the neighbor distances for each of the multiple layers of the GNN. Adjusting the neighbor distances for each of the multiple layers can include considering an increasing amount of neighbor elements as the depth of the layer within the GNN increases. For example, a first layer of the multiple layers might consider edges within a first distance of the real element, and a deeper second layer of the multiple layers might consider edges within a second distance of the real element, where the second distance is greater than the first distance. In another example, a first layer of the multiple layers might consider a first number of nearest neighbors closest to the real element, and a deeper second layer of the multiple layers might consider a second number of nearest neighbors closest to the real element, where the second number is greater than the first number.

[0050] At operation 906, physical phenomena are described based on force vectors. For example, the resultant force across the elements of the entire system can be used to describe ion transport, chemical reactions, and / or bulk and surface degradation in a material system (such as a device or functional material). Non-limiting examples of such material systems include fuel cell units, surface coatings, batteries, water desalination, and water filtration. After operation 906, process 900 ends.

[0051] One or more embodiments of the GNN algorithm and / or method use, for example Figure 10 The computing platform 1000 shown in the diagram is implemented using a computing platform such as computing platform 1000. Computing platform 1000 may include memory 1002, processor 1004, and non-volatile storage device 1006. Processor 1004 may include one or more devices selected from a high-performance computing (HPC) system, including high-performance cores, microprocessors, microcontrollers, digital signal processors, microcomputers, central processing units (CPUs), graphics processing units (GPUs), tensor processing units (TPUs), field-programmable gate arrays, programmable logic devices, state machines, logic circuits, analog circuits, digital circuits, or any other device that manipulates signals (analog or digital signals) based on computer-executable instructions residing in memory 1002. Memory 1002 may include a single memory device or multiple memory devices, including but not limited to random access memory (RAM), volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, cache memory, or any other device capable of storing information. The non-volatile storage device 1006 may include one or more persistent data storage devices, such as hard disk drives, optical drives, tape drives, non-volatile solid-state devices, cloud storage, or any other device capable of persistently storing information.

[0052] Processor 1004 may be configured to read into memory 1002 and execute computer-executable instructions residing in GNN software module 1008 of non-volatile storage device 1006 and embodying one or more embodiments of GNN algorithms and / or methods. Processor 1004 may be further configured to read into memory 1002 and execute computer-executable instructions residing in MD software module 1010 (such as LAMMPS) of non-volatile storage device 1006 and embodying MD algorithms and / or methods. Software modules 1008 and 1010 may include operating systems and applications. Software modules 1008 and 1010 may be compiled or interpreted from computer programs created using a wide variety of programming languages ​​and / or techniques, which are not limited and include, individually or in combination, Java, C, C++, C#, Objective C, Fortran, Pascal, JavaScript, Python, Perl, and PL / SQL. In one embodiment, PyTorch, as a package of the Python programming language, may be used to implement code for one or more embodiments of GNN. In another embodiment, PyTorch XLA or TensorFlow (both packages of the Python programming language) can be used to implement the code for one or more embodiments of the GNN. The code framework can be based on Crystal Graph Convolutional Neural Network (CGCNN) code, which is available under a license from MIT in Cambridge, Massachusetts.

[0053] When executed by processor 1004, the computer-executable instructions of GNN software module 1008 and MD software module 1010 enable computing platform 1000 to implement one or more of the GNN algorithms and / or methods and MD algorithms and / or methods disclosed herein. Non-volatile storage device 1006 may also include GNN data 1012 and MD data 1014 supporting the functionality, features, and processes of one or more embodiments described herein.

[0054] Program code embodying the algorithms and / or methods described herein can be distributed individually or collectively as a program product in a wide variety of different forms. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to perform aspects of one or more embodiments. Computer-readable storage media, inherently non-transitory, can include tangible media that are volatile and non-volatile, as well as removable and non-removable, implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media may further include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, portable compact disc read-only memory (CD-ROM) or other optical storage devices, magnetic tape, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be read by a computer. Computer-readable program instructions can be downloaded from the computer-readable storage medium to a computer, another type of programmable data processing device, or another device, or downloaded via a network to an external computer or external storage device.

[0055] Computer-readable program instructions stored in a computer-readable medium can be used to direct a computer, other type of programmable data processing apparatus, or other device to operate in a particular manner, causing the instructions stored in the computer-readable medium to produce an article of art including instructions that implement the functions, actions, and / or operations specified in a flowchart or diagram. In some alternative embodiments, the functions, actions, and / or operations specified in the flowcharts and diagrams can be reordered, processed sequentially, and / or processed concurrently, consistent with one or more embodiments. Furthermore, any of the flowcharts and / or diagrams may include more or fewer nodes or blocks than those consistently illustrated in one or more embodiments.

[0056] While the entire invention has been described through various embodiments, and these embodiments have been described in considerable detail, the applicant does not intend to limit the scope of the appended claims or restrict them in any way to such details. Additional advantages and modifications will readily occur to those skilled in the art. Therefore, the invention is not limited in its broader aspects to the specific details, representative apparatuses and methods, and the illustrative examples shown and described. Consequently, deviations from such details may be made without departing from the spirit or scope of the overall inventive concept.

Claims

1. A computing method for element simulation using a machine learning system parallelized across multiple processors, the method comprising: partitioning a multi-element system into a plurality of partitions, each partition including a subset of real elements contained within the partition and empty elements outside the partition that affect the real elements; for each of the multiple processors, predicting force vectors for the real elements within the multi-element system by a graph neural network (GNN) having a plurality of layers and parallelized across the multiple processors, the predicting including separately adjusting neighbor distances for each of the plurality of layers of the GNN; and describing a physical phenomenon based on the force vectors.

2. The method of claim 1, wherein, Adjusting the neighbor distances for each of the plurality of layers includes increasing an amount of neighbor elements as depth of layers within the GNN increases.

3. The method of claim 2, wherein, A first layer of the plurality of layers has edges within a first distance from the real elements, and a second, deeper layer of the plurality of layers has edges within a second distance from the real elements, the second distance being greater than the first distance.

4. The method of claim 3, wherein, The elements are atoms, and the edges are bonds between the atoms.

5. The method of claim 2, wherein, A first layer of the plurality of layers has a first number of nearest neighbors closest to the real elements, and a second, deeper layer of the plurality of layers has a second number of nearest neighbors closest to the real elements, the second number being greater than the first number.

6. The method of claim 1, further comprising: Summing forces for each of the real elements is truncated to include only those elements within a predefined cutoff from the respective real element.

7. The method of claim 6, wherein, The predefined cutoff is a predefined distance from the respective real element.

8. The method of claim 6, wherein, The predefined cutoff is a maximum number of neighbors closest to the respective real element.

9. The method of claim 1, further comprising: The GNN is trained as a deep neural network using a loss function formed to compare the force predictions to ground truth forces on a labeled dataset such that network weights of the GNN are optimized relative to the loss function using gradient descent.

10. A computing system for element simulation using a machine learning system parallelized across multiple processors, the system comprising: a plurality of processing nodes, each node including a memory storing instructions of a GNN algorithm of molecular dynamics (MD) software and a processor programmed to execute the instructions to perform operations including: operating on one of a plurality of partitions, each partition including a subset of real elements contained within the partition and empty elements outside the partition that affect the real elements; predicting force vectors for the subset of real elements contained within the partition by having the GNN pass backward through the GNN having a plurality of layers and parallelized across the plurality of processing nodes, the predicting including separately adjusting neighbor distances for each of the plurality of layers of the GNN; and describing a physical phenomenon based on a combination of the force vectors from the plurality of processing nodes.

11. The system of claim 10, wherein, Adjusting the neighbor distances for each of the plurality of layers includes increasing an amount of neighbor elements as depth of layers within the GNN increases.

12. The system of claim 11, wherein, A first layer of the plurality of layers has edges within a first distance from the real elements, and a second, deeper layer of the plurality of layers has edges within a second distance from the real elements, the second distance being greater than the first distance.

13. The system of claim 12, wherein, The elements are atoms, and the edges are bonds between the atoms.

14. The system of claim 11, wherein, A first layer of the plurality of layers has a first number of nearest neighbors closest to the real elements, and a second, deeper layer of the plurality of layers has a second number of nearest neighbors closest to the real elements, the second number being greater than the first number.

15. The system of claim 10, further comprising: truncating the sum of forces for each real element to include only those elements within a predefined cutoff of the respective real element.

16. The system of claim 15, wherein, The predefined cutoff is a predefined distance from the respective real element.

17. The system of claim 15, wherein, The predefined cutoff is a maximum number of nearest neighbors to the respective real element.

18. A non-transitory computer-readable medium comprising instructions that, when executed by a plurality of processing nodes of a parallelized machine learning system, cause the system to perform operations comprising: operating on one of a plurality of partitions, each partition comprising a subset of real elements contained within the partition and empty elements outside the partition that affect the real elements; predicting force vectors for the subset of real elements contained within the partition by having the GNN pass backwards through a plurality of layers and parallelized across the plurality of processing nodes, the predicting operation comprising separately adjusting a neighbor distance for each of the plurality of layers of the GNN; and describing the physical phenomenon based on a combination of the force vectors from the plurality of processing nodes.

19. The medium of claim 18, wherein, adjusting the neighbor distance for each of the plurality of layers comprises increasing a quantity of neighbor elements as depth of layers within the GNN increases, wherein one or more of the following are performed: a first layer of the plurality of layers has edges within a first distance from the real elements and a second, deeper layer of the plurality of layers has edges within a second distance from the real elements, the second distance being greater than the first distance; or a first layer of the plurality of layers has a first number of nearest neighbors closest to the real elements and a second, deeper layer of the plurality of layers has a second number of nearest neighbors closest to the real elements, the second number being greater than the first number.

20. The medium of claim 18, further comprising: truncating the sum of forces for each real element to include only those elements within a predefined cutoff of the respective real element, wherein the predefined cutoff is one or more of: a predefined distance from the respective real element or a maximum number of nearest neighbors to the respective real element.

Citation Information

Patent Citations

  • Neural network force field computational algorithms for molecular dynamics computer simulations

    CN111241655A

  • System and method for simulating the time-dependent behaviour of atomic and / or molecular systems subject to static or dynamic fields

    US20080147360A1