AUTOENCODER-BASED MACHINE-LEARNED INTERATOMAR POTENTIALS FOR SCALABLE, ADVANCED HAMILTONIAN

An autoencoder-based method addresses scalability and long-range effects in interatomic potentials, enabling accurate simulation of atomic systems with discontinuities and transitions, enhancing molecular dynamics applications.

DE102025133604A1Pending Publication Date: 2026-02-26ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025133604
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-23
Filing Date
2025-08-22
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing machine learning models for interatomic potentials face challenges in scalability, inability to account for long-range effects, and failure to capture discontinuities and transitions, limiting their application in large-scale simulations like molecular dynamics.

Method used

The use of an autoencoder with a bounded latent space to learn auxiliary properties of atomic systems, combined with a machine learning model, enables scalable and accurate determination of total energy by discretizing states to handle both short-range and long-range effects, as well as discontinuities and transitions.

Benefits of technology

This approach allows for precise simulation of atomic systems, including bond breaking and magnetic transitions, while maintaining computational efficiency and stability, making it suitable for large-scale simulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods for a machine learning network that train and then execute both an autoencoder and a machine learning model within a context of machine-learned interatomic potentials are disclosed. The system described here is configured to embed atomic positions and species of a given atomic system and apply these to an autoencoder to learn an auxiliary property, and to a machine learning model to learn local energies. The auxiliary property is then used to generate a Hamiltonian auxiliary description. By combining both the Hamiltonian auxiliary description and the local energies, properties such as the total energy of the atomic system are determined.By processing machine learning through both an autoencoder and a machine learning model, such methods ensure that long-range and short-range effects are taken into account, while also conveniently enabling realistic discontinuities and / or transitions within the potential energy surface.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present disclosure relates to training and executing a combination of an autoencoder and a machine learning model to determine long-range and short-range energies of a given atomic system. BACKGROUND OF THE INVENTION

[0002] Various machine learning techniques have proven useful in predicting interaction energies, forces, and other properties of different atomic systems. However, there is a non-trivial balancing act involved in using a given machine learning model for large atomic systems while also considering short-range effects and transitions (e.g., a magnetic transition). Providing scalable machine learning techniques for implementation in iterative processes, such as molecular dynamics, remains a challenge. SUMMARY

[0003] In contrast to previous implementations of machine-learned interatomic potentials (MLIPs), the present disclosure uses both an autoencoder and a machine learning model to determine properties of an atomic system, such as the total energy within the system. The autoencoder ensures that the given MLIP architecture is scalable and can be implemented in larger simulations, such as molecular dynamics, with an accuracy similar to that of ab initio methods. The combination of accuracy and scalability is at least partly due to the restriction of the autoencoder's latent space, such that the autoencoder has a fixed dimension. The output is therefore a set of discretized states.These discretized states can then be used to determine an additional Hamiltonian description of a given atomic system, which can be efficiently scaled to large systems. When combined with local energies learned using a deep neural network, one or more Gaussian processes, or another type of MLIP-based model, the resulting total energy accounts for both large and short-range effects within the system and conveniently allows for discontinuities and / or transitions if they are relevant to the particular atomic system and simulation environment. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates a system for training and using a machine learning model, such as a neural network, according to some embodiments. Fig. Figure 2 illustrates a computer-implemented method for training and using a machine learning model, such as a neural network, according to some embodiments. Fig. Figure 3 illustrates a high-level workflow diagram for machine-learned interatomic potentials of a given atomic system according to some embodiments. Fig. 4 illustrates an extension of the in Fig. 3 introduced high-level workflow diagram, where backpropagation is used to determine related forces of the given atomic system, according to some embodiments. Fig. Figure 5A illustrates a workflow diagram for an autoencoder-supported energy determination in connection with machine-learned interatomic potentials of an atomic system according to some embodiments. Fig. 5B further illustrates the in Fig. 5A introduced workflow diagram, further detailing the advantages of using an autoencoder according to some embodiments. Fig. Figure 6 illustrates another workflow diagram for autoencoder-assisted energy determination in connection with machine-learned interatomic potentials of an atomic system according to some embodiments. Fig. Figure 7 illustrates an exemplary implementation of applying an autoencoder-based energy determination to a scenario of dissolving salt in water according to some embodiments. Fig. Figure 8 is a flowchart illustrating a process of executing an autoencoder to learn an auxiliary property of an atomic system and applying the learned auxiliary property to determine long-range and short-range effects of the atomic system according to some embodiments. Fig. Figure 9 is a flowchart illustrating iterative processing based on molecular dynamics according to some embodiments. DETAILED DESCRIPTION

[0004] Embodiments of the present disclosure are described herein. It is understood, however, that the disclosed embodiments are merely examples and that other embodiments may take different and alternative forms. The figures are not necessarily to scale; some features may be enlarged or reduced to show details of certain components. Therefore, the specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching a person skilled in the art how to use the embodiments in various ways.As the person skilled in the art will understand, various features illustrated and described with reference to any of the figures can be combined with features illustrated in one or more other figures to create embodiments not expressly illustrated or described. The combinations of illustrated features provide representative embodiments for a typical application. However, various combinations and modifications of the features, consistent with the teachings of this disclosure, may be desirable for certain applications or implementations.

[0005] “A”, “an”, and “the”, as used here, refer to both singular and plural referents unless the context clearly indicates otherwise. For example, “a processor” programmed to perform various functions refers to a processor programmed to perform each individual function, or to more than one processor programmed together to perform each of the different functions.

[0006] Applications of machine-learned interatomic potentials (MLIP) are vast and diverse. However, prior to the development of the present disclosure, past implementations of MLIP either (1) limited the ability to scale while including interactions beyond a limited number of neighboring atoms, (2) limited the ability to learn long-range effects, and / or (3) limited the ability to account for discontinuities and / or transitions. When using MLIP to compute energy and forces, it is important to have the flexibility to integrate all three of these capabilities depending on a given type of atomic system and for a more complete and comprehensive analysis.The following few paragraphs describe the context for each of these challenges faced by past implementations of MLIP, followed by an explanation of how the present disclosure overcomes the need to prioritize one of these effects at the expense of one or more of the other effects.

[0007] Regarding the ability to scale beyond a limited number of neighboring atoms, previous implementations of MLIP could not overcome the difficulty of scaling deep learning networks, especially in very large atomic systems. Historically, deep neural networks used a cutoff distance for neighboring atoms r. c, whereby information in the deep neural network was only passed between atoms within the defined cutoff distance. A further limitation was that common message-passing deep neural networks, such as Nequip, applied this cutoff distance at every layer of the deep neural network, so that a given deep neural network with N layers would have an effective cutoff of NT. c and has a number of effective neighboring atoms, which are known as (No. c ) 3 scaled. This scaling of type (No. c ) 3 This makes it virtually impossible to divide the atoms in a given atomic system across different processors during a specific production run, which is typically of interest in large-scale simulation techniques such as molecular dynamics. Furthermore, the scaling of type (No. c ) 3Computationally, this is cubically expensive, since the number of atoms increases with the cube of N. This lack of scalability could not be overcome even with techniques like Allegro, which convert energy into energy per atom E. i = ∑ j∈Ni E ij (N i ) divide, where N i the set of all atoms in the neighborhood of i (i.e., within the cutoff) and no others, and where E ij is an effective pairwise energy corresponding to two atoms i and j.

[0008] Regarding the ability to learn long-range effects such as electrostatics and delocalized electrons (e.g., magnetic conductors), previous implementations of MLIP, such as Allegro, which is limited to purely local energies, could not overcome this difficulty. Energies and forces associated with action at a distance can be learned through interactions within an atomic system that extend beyond a suitable cutoff radius r. c The outflow, which is typically several angstroms, is strongly affected. This has led either to an enormous increase in the value of r. cThis, in turn, led to an increase in instabilities, computation time, and / or memory requirements, or to the complete neglect of long-range effects, which again resulted in a significant loss of accuracy in the given simulation. By focusing on interactions within the neighborhood of a central atom i, the analysis of long-range effects (e.g., the interaction between the central atom i and another atom outside the cutoff radius r) becomes less critical. c ) lost.

[0009] Even when an auxiliary network is used to learn electrostatic point charges using density function theory (DFT) datasets, which can then be fed into a known procedure for calculating long-range electrostatic forces and energies (e.g., an Ewald summation), a comprehensive method for integrating short-range and long-range effects into a given simulation type has been lacking. Other attempts have included fitting point charges to DFT-based charges, such as by using Hirshfeld or Mulliken charge partitioning schemes, or deriving effective charge values ​​from other quantities, such as fitting only to the total energy while neglecting local energies.

[0010] None of these attempts addressed the problem of enabling scalability while simultaneously ensuring the stable inclusion of long-range effects. A significant drawback of such methods in current technology is that atomic charges can fluctuate frequently and considerably, making the entire potential energy surface extremely sensitive to the initial simulation configuration as well as to small perturbations of atomic positions during the simulation. Furthermore, overall charge neutrality must be guaranteed at all times, e.g., Σ i q i = Q total , where Q totalThe total charge of the atomic system is zero and sums to zero. Enforcing the neutrality constraint means that charge updates cannot occur locally without considering all atoms in the atomic system. Taken together, the sensitivity of charge values ​​to precise atomic configurations and the need to enforce charge neutrality have thus far made it impossible to partition the atomic system into purely local components and efficiently perform molecular dynamics simulations.

[0011] Regarding the ability to account for discontinuities and / or transitions, previous applications of MLIP did not effectively capture discontinuities in the potential energy surface. Often, there are segments of the potential energy surface that are smooth with respect to atomic position, and other segments corresponding to a transition (e.g., a magnetic transition, bond breaking, charge transfer) where there should be an abrupt change in the potential energy surface. Previous applications of MLIP failed to target a transition that would make the potential energy surface discontinuous while also allowing the potential energy surface to be continuously differentiable and adequately smooth (and therefore stable) in each region due to the abrupt change of the given transition.

[0012] To address these challenges, the present disclosure employs an autoencoder with a bounded latent space to learn one or more auxiliary properties of an atomic system according to some embodiments. By defining the latent space based on a hyperparameter or other dimension-based scheme that allows the autoencoder to map the atomic positions and species of an atomic system, the resulting learned auxiliary property is defined by a finite number of discrete states (e.g., charge states, oxidation states, magnetic states, or some other atomic property) that can be used to construct an auxiliary Hamiltonian that is analytical and therefore readily scalable to longer distances or regions than a conventional MLIP method.In parallel, the atomic positions and species of the atomic system can also be used as input to a machine learning model, such as a deep neural network, to learn local energies. The combination of both Hamilton's auxiliary description and the learned local energies enables the specific MLIP architectures described here to determine the total energy of a given atomic system with high precision, thereby addressing the three challenges that the scientific community described above has faced.

[0013] In particular, the present disclosure provides a scalable solution because the autoencoder ensures that mapping produces a finite, discretized set of states while still being configured to benefit from the application of large models, such as a deep neural network. Furthermore, both long-range and short-range effects are properly accommodated for the use of the combined architecture of an autoencoder and a machine learning model. Additionally, discontinuities and / or transitions are more accurately described using the present disclosure because the discretized set of states allows for abrupt changes in the auxiliary property learned by the autoencoder, so the methods and systems described herein better simulate bond breaking or a magnetic phase transition to a spin glass, etc.

[0014] The following description continues with a general introduction to machine learning techniques relevant to the machine learning methods for interatomic potentials described here. Next, various embodiments of autoencoder- and machine learning model-based architectures are discussed. This disclosure then demonstrates the versatility of the methods and systems described here for use in determining macro- and micro-level properties of various molecular compositions and in their implementation in larger simulations, such as molecular dynamics (MD).

[0015] Fig. Figure 1 illustrates a System 100 for training and using a neural network, such as a deep neural network. It is understood that, while the following paragraphs refer to Fig. 1 and Fig. The two given exemplary embodiments refer to a deep neural network, additional embodiments of Fig. 1 and Fig. 2 can be applied to any other type of neural network-based or non-neural network-based machine learning model (e.g., Gaussian processes) that is configured to be designed, trained, and optimized for various machine-learned interatomic potential applications.

[0016] Furthermore, and in relation to the description presented here, a “deep” learning model, such as a deep neural network, can be defined as having multiple hidden layers (e.g., one, two, or ten hidden layers) between an input layer and an output layer of the model. A deep learning model can also be used to describe a machine learning model configured to learn complex patterns and representations based on training and / or validation datasets used as inputs to the deep learning model. Additional embodiments relating to such types of machine learning models are described here with reference to Machine Learning Model 210, Network 306, Deep Neural Network 406, Network 518, Learning 618, and Block 810.

[0017] In some embodiments, the system 100 may include an input interface for accessing training data 102 for the neural network. For example, as in Fig. Figure 1 illustrates that the input interface is formed by a data storage interface 104, which can access the training data 102 from a data storage device 106. For example, the data storage interface 104 can be a storage interface or a persistent storage interface, such as a hard disk or SSD interface, but also a personal, local, or wide area network interface, such as a Bluetooth, ZigBee, or Wi-Fi interface, or an Ethernet or fiber optic interface. The data storage device 106 can be internal data storage of the system 100, such as a hard disk or SSD, but also external data storage, such as network-accessible data storage.

[0018] In some embodiments, the data storage 106 may further comprise a data representation 108 of an untrained version of the model (e.g., a version of the machine learning model that still needs to be trained), which the system 100 can access from the data storage 106. It is understood, however, that the training data 102 and the data representation 108 of the untrained neural network can each also be accessed from another data storage, e.g., via another subsystem of the data storage interface 104. Each subsystem can be of a type as described above for the data storage interface 104. In other embodiments, the data representation 108 of the untrained neural network can be generated internally by the system 100 based on design parameters for the neural network and therefore does not need to be explicitly stored on the data storage 106.System 100 can further include a processor subsystem 110, which can be configured to provide an iterative function during the operation of System 100 as a replacement for a stack of layers of the neural network to be trained. Here, the respective layers of the stack being replaced can have shared weights and can receive as input an output from a previous layer, or, for a first layer of the stack, an initial activation and part of the stack's input. Processor subsystem 110 can also be configured to iteratively train the neural network using the training data 102 (thus, for example, generating updated versions of the machine learning model with respect to an initial "untrained" version of the model).Here, an iteration of the training by processor subsystem 110 can include a forward propagation part and a backward propagation part. Processor subsystem 110 can be configured to perform the forward propagation part by determining, among other operations that define the forward propagation part to be executed, an equilibrium point of the iterative function at which the iterative function converges to a fixed point. Determining the equilibrium point involves using a numerical root-finding algorithm to find a root solution for the iterative function minus its input, and by providing the equilibrium point as a substitute for an output of the stack of layers in the neural network.System 100 can further include an output interface for outputting a data representation 112 of the trained neural network, where this data can also be referred to as trained model data 112. For example, as in . Fig. Figure 1 illustrates that the output interface is formed by the data storage interface 104, wherein the interface in these embodiments is an input / output (“I / O”) interface through which the trained model data 112 can be stored in the data storage 106. For example, the data representation 108, which defines the “untrained” neural network, can be at least partially replaced during or after training by the data representation 112 of the trained neural network by adjusting the parameters of the neural network, such as weights, hyperparameters, and other types of neural network parameters, to reflect the training on the training data 102. This is also shown in Fig. Figure 1 illustrates this by reference numerals 108 and 112, which refer to the same data set on data storage 106. In other embodiments, data representation 112 can be stored separately from data representation 108, which defines the “untrained” neural network. In some embodiments, the output interface can be separate from the data storage interface 104, but can generally be of the type described above for data storage interface 104.

[0019] Fig. Figure 2 illustrates a computer-implemented method for training and using a neural network according to some embodiments. The system 200 can include at least one computing system 202. The computing system 202 can include at least one processor 204 operatively connected to a memory unit 208. The processor 204 can include one or more integrated circuits implementing the functionality of a central processing unit (CPU) 206 and, in some embodiments, a graphics processing unit (GPU). The CPU 206 can be a commercially available processing unit implementing an instruction set such as one of the x86, ARM, Power, or MIPS instruction set families. During operation, the CPU 206 can execute stored program instructions retrieved from the memory unit 208.The stored program instructions can include software that controls the operation of the CPU 206 to perform the process described here. In some examples, the processor 204 can be a system-on-a-chip (SoC) that integrates the functionality of the CPU 206, the memory unit 208, a network interface, and input / output interfaces into a single integrated device. The computing system 202 can implement an operating system to manage various aspects of the process.

[0020] The memory unit 208 can include volatile and non-volatile memory for storing instructions and data. The non-volatile memory can include solid-state memory such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is disabled or loses electrical power. The volatile memory can include static and dynamic random-access memory (RAM) that stores program instructions and data. For example, the memory unit 208 can store a machine learning model 210 or algorithm, a training dataset 212 for the machine learning model 210 (e.g., density function theory (DFT) training datasets), a raw source dataset 214, an autoencoder, etc.

[0021] The computing system 202 may include a network interface device 220 configured to provide communications with external systems and devices. For example, the network interface device 220 may include a wired and / or wireless Ethernet interface, as defined by the 802.11 standard family of the Institute of Electrical and Electronics Engineers (IEEE). The network interface device 220 may include a cellular communication interface for communicating with a cellular network (e.g., 3G, 4G, 5G). The network interface device 220 may also be configured to provide a communication interface to an external network 222 or a cloud.

[0022] The external network 222 can be referred to as the World Wide Web or the Internet. The external network 222 can establish a standard communication protocol between computing devices. The external network 222 enables the easy exchange of information and data between computing devices and networks. One or more servers 224 can communicate with the external network 222.

[0023] The computer system 202 can include an input / output (I / O) interface 218, which can be configured to provide digital and / or analog inputs and outputs. The I / O interface 218 can include additional serial interfaces for communicating with external devices (e.g., a Universal Serial Bus (USB) interface).

[0024] The computing system 202 may include a human-machine interface (HMI) device 216, which may include any device that enables the system 200 to receive control inputs. Examples of input devices may include human interface inputs such as keyboards, mice, touchscreens, speech input devices, and other similar devices. The computing system 202 may include a display device 226. The computing system 202 may include hardware and software for outputting graphic and text information to the display device 226. The display device 226 may include an electronic display screen, projector, printer, or other suitable device for displaying information to a user or operator.The computing system 202 can also be configured to enable interaction with a remote HMI and remote display devices via the network interface device 220.

[0025] System 200 can be implemented using one or more computing systems. While the example represents a single computing system 202 that implements all of the described features, it is intended that different features and functions can be implemented separately and communicate with each other through multiple computing units. The specific system architecture chosen can depend on a variety of factors.

[0026] System 200 can implement a machine learning algorithm 210 configured to analyze the raw source dataset 214. The raw source dataset 214 can contain raw or unprocessed sensor data that may be representative of an input dataset for a machine learning system. The raw source dataset 214 can include DFT training datasets and / or any other atomic descriptors relating to atomic positions and atomic species of various systems. In some examples, the machine learning algorithm 210 can be a neural network algorithm designed to perform a predetermined function. For example, the neural network algorithm can be configured within a context of machine-learned interatomic potentials to learn local energies of a system.

[0027] The computer system 200 can store a training dataset 212 for the machine learning algorithm 210. The training dataset 212 can represent a set of previously constructed data for training the machine learning algorithm 210. The training dataset 212 can be used by the machine learning algorithm 210 to learn weighting factors associated with a neural network algorithm. The training dataset 212 can include a set of source data exhibiting corresponding results or outcomes that the machine learning algorithm 210 attempts to duplicate through the learning process. In a context of machine-learned interatomic potentials, the machine learning algorithm 210 can predict energies and / or other atomic properties of a given atomic system.

[0028] The machine learning algorithm 210 can be operated in a learning mode using the training dataset 212 as input. The machine learning algorithm 210 can be executed over a number of iterations using the data from the training dataset 212. With each iteration, the machine learning algorithm 210 can update internal weighting factors based on the results obtained. For example, the machine learning algorithm 210 can compare output results (e.g., annotations) with those included in the training dataset 212. Since the training dataset 212 contains the expected results, the machine learning algorithm 210 can determine when performance is acceptable. After the machine learning algorithm 210 has reached a predetermined performance level (e.g.,(With 100% agreement with the results associated with the training dataset 212), the machine learning algorithm 210 can be run using data that is not in the training dataset 212. The trained machine learning algorithm 210 can be applied to new datasets to generate annotated data.

[0029] The machine learning algorithm 210 can be configured to identify a specific feature in the raw source data 214. The raw source data 214 can include a variety of instances or an input dataset for which annotation results are desired. The machine learning algorithm 210 can be programmed to process the raw source data 214 to identify the presence of the specific features. The machine learning algorithm 210 can be configured to identify a feature in the raw source data 214 as a predetermined feature (e.g., an atomic system comprising water molecules has evidence of hydrogen and oxygen). The raw source data 214 can be derived from a variety of sources. For example, the raw source data 214 can be actual input data collected by a machine learning system. The raw source data 214 can be machine-generated for testing the system.As an example, the raw source data may include 214 DFT training datasets relating to different concentrations of salt dissolved in water.

[0030] In the example, the machine learning algorithm 210 can then process the raw source data 214 and output a value of predicted local energies. A machine learning algorithm 210 can generate a confidence level or confidence factor for each output it produces. For example, a confidence value exceeding a predetermined high-confidence threshold can indicate that the machine learning algorithm 210 is confident that the identified feature corresponds to the specified feature. A confidence value less than a low-confidence threshold can indicate that the machine learning algorithm 210 has some uncertainty about the presence of the specified feature.

[0031] Fig. Figure 3 illustrates a high-level workflow diagram for MLIP of a given atomic system according to some embodiments.

[0032] In some embodiments, MLIP, as used and described here, can be used to define a set of atomic positions, {r→i}, and corresponding atomic species, {Z i}, of a given atomic system to a scalar energy, E, as in Fig. 3 shown, and, in extension, to additional properties such as forces, {F→i}, as in Fig. 4 shown. In some embodiments, this can be considered equivalent to learning the potential energy surface of the atomic system.

[0033] As applied here, atomic species can refer to the atomic number, the isotope, an elemental description, or any other property used to distinguish different atomic identities via simulation.

[0034] As shown in Process 300, atomic positions and atomic species can be referred to as atomic descriptors 302 or as an atomic description 302. For example, in some embodiments where Process 300 resembles a process flow that is fulfilled for a customer, the customer may provide a requirement to determine the total energy of a given atomic system and then provide an atomic description 302 for the computing system that performs the calculation. Fig. The 3 procedures shown are carried out. It is equally understood that the atomic positions and atomic species within atomic descriptors 302 refer to data which, when compiled, provide a simulated atomic structure of the atomic system based on the provided atomic positions and atomic species.

[0035] Embedding 304 refers to the conversion of atomic positions and atomic species into inputs for the interatomic potentials, denoted as “{V}” in the figure. These inputs for the interatomic potentials are also referred to here as atomic descriptors. The embeddings can be designed to be invariant or covariant with respect to certain symmetry groups of the atomic system or physics, such as translation, rotation, exchange of atoms, or various crystal symmetries. Such an embedding is further characterized by the Fig. 5A, Fig. 5B and Fig. 6 discussed here. The embedding 304 is then provided as input for the “network” 306, in which the learning of one or more properties about the atomic system takes place. Fig. 3 and Fig. Figure 4 serves to illustrate a general process flow of MLIP, while the Fig. 5A, Fig. 5B and Fig. 6. Illustrate the combined use of an autoencoder and a machine learning model during the learning phase. 306.

[0036] The one or more properties learned during learning stage 306 are then used to calculate a total energy of the atomic system, also known as output 308 in Fig. 3 is designated.

[0037] The Process 300 can be applied to various computational simulations, such as those that reconstruct structures from experimental data, molecular dynamics simulations, methods for finding a specific atomic configuration for an atomic system, atomic Monte Carlo or large canonical Monte Carlo simulations, and even simulations that identify probable reactions and / or transition states.

[0038] Fig. 4 illustrates an extension of the in Fig. 3 introduced high-level workflow diagram, where backpropagation is used to determine related forces of the given atomic system, according to some embodiments.

[0039] Similar to the one that was in Fig. As introduced in 3, process 400 represents a set of atomic positions and atomic species that are embedded into atomic descriptors during embedding stage 404. Then, one or more machine learning models are applied during learning stage 406 to determine the total energy of an atomic system, as shown by output 408.

[0040] As additionally in Fig. As shown in Figure 4, learning 406 refers to a deep neural network. A deep neural network typically has multiple layers that are learned using backpropagation, with the network weights being optimized to fit a loss function such as L = |E pred - E actual | 2 to minimize, and where R actualThe training data is extracted from a higher-accuracy method, such as DFT. Backpropagation can also be described as backpropagation through the network model, where the network weights are adjusted based on the model's error rate. An example of a technique that uses backpropagation is PyTorch's AutoGrad functionality.

[0041] In some embodiments, backpropagation can also be used to predict forces acting in Fig. 4 are designated as forces 410, as the derivative of the total energy with respect to the atomic positions. The deep neural network can then be further trained on forces, using a loss function such as L=∑i|F→i,pred−F→i,actual|2 and again highly precise forces, such as those of DFT, are used. According to some embodiments, the loss function can also focus on the stress tensor or a combination of two or more of the above.

[0042] Certain embodiments, which are in Fig. The embodiments illustrated in Figure 4 refer to learning 406 as implemented using a deep neural network (e.g., Nequip, Allegro) to determine both total energy and forces using backpropagation. However, other embodiments of process 400 may refer to learning 406 as implemented using one or more Gaussian processes (e.g., FLARE). As used here, one or more Gaussian processes can refer to the use of one or more such processes which, when combined, implement a model. The model may then be referred to as a Gaussian process model or a Gaussian-based model.

[0043] Depending on a specific implementation of processes 300 and 400 for a given problem provided by a customer, it is understood that learning 306 and 406 can be adapted to refer to a deep neural network or to one or more Gaussian processes. Furthermore, Gaussian processes and deep neural networks can be defined as the two main classes of MLIP. Therefore, workflow diagrams illustrated in all figures and their corresponding text herein are intended to refer to MLIP implementations that include those using either Gaussian processes or deep neural networks, depending on the given implementation of this disclosure.

[0044] Furthermore, an exemplary embodiment of a molecular dynamics algorithm that implements the in Fig. The approach type shown in section 4 is used, in addition to the one shown in the diagram. Fig. 9 illustrates.

[0045] Fig. 5A and Fig. Figure 5B illustrates a workflow diagram for autoencoder-assisted energy determination in connection with machine-learned interatomic potentials of an atomic system according to some embodiments.

[0046] As introduced above, and to ensure that the one or more models included in Learning 306 or Learning 406 are configured to capture significant topology and / or charge transitions, while also eliminating spurious noise when no such transition occurs, Learning 306 or Learning 406 can resemble a combination of an Autoencoder 506 and a Learning Network 518. As introduced above, Learning Network 518 can resemble a deep neural network or one or more Gaussian processes configured to be combined into a Gaussian-based model.

[0047] Furthermore, Autoencoder 506 can be defined as having a latent space with limited dimension, so that it can then be configured to learn atomic system-specific transitions.

[0048] In some embodiments, the autoencoder 506 can resemble a PyTorch module with parameters that are learned during a training stage. The parameters can then be fixed and used to predict one or more auxiliary properties 508, such as charge, during a given iteration of process 500.

[0049] The autoencoder can also be defined as an autoencoder with a bounded latent space of dimension D, which in Fig. 5A is referred to as the latent space representation (or representation of latent space). The value of D can correspond to the number of expected states or it can be a hyperparameter selected during a training phase. Furthermore, the determination of the dimension D can be additionally influenced by the complexity of the given atomic system or by the nature of the auxiliary property to be learned.

[0050] In some embodiments, autoencoder 506 may resemble a variation autoencoder, a regularized autoencoder, a sparse autoencoder, or any other type of artificial neural network that maps the atomic descriptors 502 by a latent space of limited dimension 506 to predict an auxiliary property 508.

[0051] As in Fig. As shown in Figure 5A, the output of autoencoder 506 is a learned auxiliary property defined by a finite set of discrete states. This is also referred to in the figure as auxiliary states 508. The discretized set of states is equal to the number of dimensions chosen for the latent space of autoencoder 506. In some embodiments, this can also be referred to as the dimension of the hyperparameter. For example, the learned auxiliary property 508 can relate to atomic charges, where the atomic charges change only when an essential transformation of the local environment in which the atomic system is simulated is detected.

[0052] In some embodiments, the auxiliary property 508 can be one or more charge states ({q i}), oxidation states, magnetic states ({m→i}) or another atomic property that is specifically relevant to the atomic system under investigation. Additionally, within decoder 528, there can be a transformation between the restricted latent space and the floating-point physical property (charge state, oxidation state, magnetic state, etc.) of the auxiliary Hamiltonian.

[0053] One or more of the auxiliary properties are then used to determine a Hamiltonian auxiliary description of the atomic system. As shown in the figure, the Hamiltonian auxiliary description includes the analysis of both short-range and long-range effects, while also taking into account discontinuities and / or transitions relating to the given atomic system. An example of such a transition is given in Fig. Figure 5B illustrates and is further discussed below. Furthermore, the illustrated embodiments in Fig. 5A presents a Hamiltonian auxiliary description that describes the energies of the atomic system. However, other embodiments of Fig. 5A To map a Hamiltonian auxiliary description that describes the forces of the atomic system.

[0054] In parallel, embedded atomic descriptors are also provided to learning 518, as indicated by atomic descriptors 514 and embedding 516. It is understood that in the Fig. In the embodiments illustrated in 5A, the sequence illustrated by 502, 504, 506, 508 and 510 can be carried out in parallel to the sequence illustrated by 514, 516, 518 and 520, or in a sequential order. Fig. The 6 illustrated embodiments also demonstrate a sequential order of process 600 and are further discussed below.

[0055] As introduced above, learning 518 can resemble a form of deep neural network or Gaussian-based model configured for machine-learned interatomic potentials and for learning local energies 520 of an atomic system. Following the examples given above, if and when a target property to be learned is related to energy, Hamilton's auxiliary description 510 can describe energies of the atomic system, and learning 518 can be configured to learn local energies so that a total energy 512 for the atomic system can be determined. In other embodiments, if the target property to be learned is directly related to forces, Hamilton's auxiliary description 510 can describe forces of the atomic system, and learning 518 can be configured to learn forces directly so that a total related force 512 for the atomic system can be determined.

[0056] After learning both the auxiliary Hamiltonian description and the local energies of the atomic system, the total energy of the atomic system can then be determined. In some embodiments, the determination of the total energy is based on the strictly local energies learned during learning 518 and on the analytical auxiliary Hamiltonian description 510 determined via the autoencoder 506. This ensures that the determination of the total energy takes long-range effects into account. In some embodiments, the auxiliary Hamiltonian description may resemble an Ewald summation of the charges per atom or a magnetic Hamiltonian description, using learned parameters J(r). ij ) contains.

[0057] Referring again to atomic descriptors 502 and 514, these are specifically local atomic descriptors, since they involve an analysis of atoms j in the neighborhood of a central atom i. Thus, learning 518, when performed using an MLIP-based model such as Allegro, yields local energies E i = ∑ j∈Ni E ij out of.

[0058] Another element relevant to the architectures shown in sections 500 and 600 is the configuration and enforcement of charge neutrality. For example, the 506 autoencoder can be configured to enforce charge neutrality when it runs. When the autoencoder receives a signal that charge neutrality should be followed, the auxiliary property is learned, which also corresponds to charge neutrality.

[0059] In some embodiments and for electrostatic applications, the loss function of the autoencoder could be generalized to include charge neutrality for the auxiliary states in the latent space in addition to the usual loss. This restricts the autoencoder from learning a latent space as a function of input coordinates and atom types, taking into account the total charge of the atomic system (e.g., normally zero).

[0060] During the dynamics, there can be an additional step regarding charge neutrality to ensure that Σ i q i = Q totalInstead of having highly fluctuating charges and therefore requiring frequent reassessment of global neutrality, the present disclosure is configured such that charges remain approximately constant in most steps, as a consequence of the finite number of discrete states and thus a correspondingly small number of possible changes. Furthermore, the computational system is configured to derive when the auxiliary states have changed as a function of the input coordinates and species (e.g., via a list of auxiliary variables and a product of the autoencoder), so that recalculation during the dynamics is only permitted when absolutely necessary. This is analogous to how, in modern scalable dynamics (e.g., LAMMPS), the atom neighbor list is not computed at every time step. This is also in Fig. 9 shown.

[0061] In some embodiments, a customer may specify whether or not to enforce charge neutrality. In other embodiments, the architectures illustrated in 500 and 600 may be configured to determine whether charge neutrality should be followed, for example, based on incoming information from the customer indicating whether or not this is an electrostatic issue.

[0062] In some scenarios relating to specific atomic systems, such as when dealing with solids in liquids, charge neutrality is not enforced. In a scenario where salt is dissolved in water (see also the description that refers here to Fig. (referring to point 7), excess charge in the area under investigation does not correspond to any reasonable interpretation or simulation of salt dissolved in water, and thus charge neutrality is enforced. In yet another example, some scenarios involving point charges would require the autoencoder to conform to charge neutrality, while scenarios involving polarizability or magnetic moment might not require such a constraint.

[0063] Forces can be described as the sum of F→j,local=∑idEi,local / dr→j and the aid- F→j,aux=dEaux / dr→j can be calculated. Since the autoencoder is designed to keep the auxiliary properties approximately constant, this latter derivative can be approximated as an analytical term, e.g., for electrostatics. Fj≈qiqj / rij2.

[0064] The output of the network can be energy, forces, stresses, polarizability, point charges, an electrostatic field, a magnetic moment and / or any other atomic system-wide and / or atom-specific property that can be used in a simulation such as molecular dynamics, structure or conformation search, Monte Carlo simulation or any other atomistic simulation.

[0065] As in Fig. As shown in Figure 5B, the autoencoder 506 can resemble the one illustrated with the encoder 524, the latent space representation 526, and the decoder 528. The latent space representation 526 resembles a lower-dimensional space with respect to the dimensions of the encoder 524 and the decoder 526, configured to map atomic descriptors to one or more auxiliary properties of the given atomic system. Furthermore, the autoencoder 506 can resemble an artificial neural network or any other feed-forward network configured to learn auxiliary properties of atomic systems.

[0066] The encoder 524 is used to encode the inputs relating to a combination of the atomic positions and atomic species 502, which have been transformed into a certain embedding 504 of the given atomic system. The latent space representation (or representation of latent space) 526 then refers to a restricted latent space with a certain dimension D, in which core features, information, and / or dependencies of the input data are processed and then passed to the decoder 528. The decoder 528 then has the task of generating the auxiliary property based on the core features learned within the restricted latent space.

[0067] As in Fig. Figure 5B illustrates that auxiliary states 508, learned via the autoencoder 506, enable discontinuities and / or transitions within the potential energy surface. In the Fig. In the example shown in Figure 5B, the vertical dashed line within line 532 illustrates bond breaking, as there is an abrupt change in atomic charge with increasing bond length. Line 530 illustrates the same type of transition without having learned the auxiliary states 508 using the autoencoder 506. The abrupt transition is lost, and the auxiliary property is much more volatile, leading to instabilities in the multidimensional potential energy surface (PES).

[0068] Fig. Figure 6 illustrates another workflow diagram for autoencoder-assisted energy determination in connection with machine-learned interatomic potentials of an atomic system according to some embodiments.

[0069] Similar to what happened through the 500 process in Fig. 5A and Fig. As illustrated in Figure 5B, Process 600 demonstrates the use of atomic positions and species to determine the total energy of a given atomic system. Atomic descriptors 602 are embedded as inputs for the interatomic potentials during Embedding 604 and subsequently provided to the autoencoder 606. As with Process 500 above, the autoencoder 606 is configured to provide a one-dimensional representation of latent space.

[0070] As in Fig. As shown in Figure 6, the output of the autoencoder 606 is a learned auxiliary property defined by a finite set of discrete states. This is also referred to in the figure as auxiliary states 608. The discretized set of states is equal to the number of dimensions chosen for the latent space of the autoencoder 606. In some embodiments, the discretized set of states within the representation of the latent space 526 is then further transformed into a set of auxiliary states 608 by a decoder 528.

[0071] One or more auxiliary properties are then used to determine a Hamiltonian auxiliary description of the atomic system. As shown in the figure, the Hamiltonian auxiliary description includes the analysis of both short-range and long-range effects, while also taking into account discontinuities and / or transitions relating to the given atomic system.

[0072] In parallel, embedded atomic descriptors are also provided to learning 618, as indicated by atomic descriptors 614 and embedding 616. As in Fig. As shown in Figure 6, the sequence of 602, 604, 606, and 608 can be performed in parallel with the embedding of the atomic description, which is shown with the sequence of 614 and 616. However, before the model is executed during learning 618, auxiliary states 608 are also provided as inputs. This can be advantageous in certain embodiments, whereby providing the machine learning model with both embedded atomic descriptors and auxiliary states ensures a more robust model for learning local energies.

[0073] Similar to the one that was in Fig. 5A and Fig. As illustrated in Figure 5B, both the specific Hamiltonian auxiliary description 610 and the learned local energies 620 are used to determine the total energy 612 of the atomic system, and in some embodiments can also be used to determine forces, mechanical stresses or any other physically meaningful value.

[0074] Fig. Figure 7 illustrates an exemplary implementation of applying an autoencoder-based energy determination to a scenario of dissolving salt in water according to some embodiments.

[0075] In the given scenario, which is in Fig. As shown in Figure 7, process 700 describes the overall process of receiving a request from a customer to determine a property of an atomic system that simulates salt dissolving in water. As shown in Block 702, a problem received from the customer in this scenario requests the determination of the total energy of a given system in which a specific concentration of NaCl is dissolved in water.

[0076] In some embodiments, the customer may additionally provide the atomic positions, as shown by the crystal lattice of NaCl and the water molecules in block 702 of the figure, and the atomic species, understood to be sodium, chlorine, oxygen and hydrogen, from the problem statement, i.e. the problem description.

[0077] In other embodiments, the customer merely provides the requirement to determine the total energy of the given atomic system, and methods and systems such as those described here are configured to determine the atomic positions and species.

[0078] The atomic description, which includes the atomic positions and species, is then fed to a computing system with an architecture such as that in Fig. 5A, Fig. 5B or Fig. The 6 shown are provided. Block 704 thus comprises embedding the atomic positions and species in atomic descriptors, providing the atomic descriptors to an autoencoder, learning a given auxiliary property based on these atomic descriptors, and determining a resulting Hamiltonian auxiliary description of the atomic system. Block 704 also includes providing the embedded atomic description to a deep neural network or a Gaussian-based model to determine local energies of the given atomic system.

[0079] In some embodiments and as in Fig. As illustrated in Figure 6, the learned auxiliary properties of the autoencoder can be provided as an input to the deep neural network or the Gaussian-based model, in addition to the embedded atomic descriptors that are provided as input.

[0080] Hamilton's auxiliary description and the learned local energies can then be used to determine the customer's requirement, such as the total energy of the atomic system. As shown in Block 706, the contribution of Hamilton's auxiliary description can be performed via an Ewald summation.

[0081] The results from Block 706 can then be provided to the customer according to some embodiments. Block 704 can also include two or more iterations of an autoencoder-assisted energy determination, and thus the customer can receive information about the results of one or more of these iterations.

[0082] The given scenario 700 also demonstrates the versatility of the autoencoder. For example, the autoencoder can be provided with rotational and / or exchange symmetries, depending on the embedding that is completed. Then, in a first example, if the latent space with bounded dimension of the autoencoder is set to have a dimension of 4, the autoencoder can be configured to learn that screened charges of sodium ions are always +0.9, that chloride ions are always -0.9, that the hydrogen of a water molecule has a partial charge of 0.5, and that oxygen has a partial charge of -1.0.

[0083] In a second example, if the latent space with bounded dimension of the autoencoder is set to have a dimension of 5, 6, 7, 8, 9 or 10, the autoencoder can additionally be configured to learn distinctions between water, protons, hydroxyls and hydronium.

[0084] Furthermore, the autoencoder is also versatile by including or excluding the consideration of charge neutrality. For example, if the autoencoder is provided with information indicating that there are 5 sodium ions and 5 chloride ions, and if the dimension of the latent space is set to 2, then the autoencoder learns that the total charge prediction is zero.

[0085] Fig. Figure 8 is a flowchart illustrating a process of executing an autoencoder to learn an auxiliary property of an atomic system and applying the learned auxiliary property to determine long-range and short-range effects of the atomic system according to some embodiments.

[0086] Process 800 is provided within the context of receiving a specific problem statement from a customer and learning specific properties of a given atomic system. In some embodiments, Process 800 can be performed on one or more processors of a computing system, such as the one that, with respect to Fig. 1 and Fig. 2 is described. A user interface can also be provided via such a computing system so that one or more customers can submit information regarding requirements for MLIP-based simulations to be executed on the computing system and receive results of these requirements.

[0087] As shown in Block 802, a customer can provide an atomic description of an atomic system, where the atomic description includes at least atomic positions and atomic species known to exist within the atomic system. The problem statement can further define specific goals, such as an interest in using machine learning to determine a Hamiltonian auxiliary description of the atomic system, and / or an interest in a description of the total energy of the atomic system, etc.

[0088] In Block 804, the atomic positions and species are embedded in atomic descriptors, which can be invariant or covariant with respect to certain symmetries of the system. These embeddings are then provided to an autoencoder in Block 806, where the atomic descriptors are mapped to a latent space of bounded dimension by the autoencoder in order to learn an auxiliary property of the atomic system. This auxiliary property corresponds to a discretized set of states, such as charge states, oxidation states, or magnetic states.

[0089] In Block 808, the learned auxiliary property is used to generate an auxiliary Hamiltonian description of the given atomic system. Such an auxiliary Hamiltonian description is configured to describe both long-range and short-range effects of the atomic system, resulting from the use of atomic positions and species, and from the application of an atomic cutoff radius r. c does not need to be determined.

[0090] In Block 810, the embedded atomic descriptors are also provided, either concurrently or sequentially, to a machine learning model to determine local energies of the atomic system. In some embodiments, the machine learning model may resemble a deep neural network, a Gaussian-based model, or any other MLIP-based model configured to learn local energies of an atomic system. Furthermore, the learned auxiliary property described in Block 806 may also be provided as an input to the machine learning model in some embodiments.

[0091] In Block 812, the total energy of the atomic system is determined using Hamilton's auxiliary description and the learned local energies.

[0092] Following the determination of the total energy of the atomic system, the process can be repeated 800 times. The number of iterations beyond the initial iteration can be determined based on the specific properties to be learned by the autoencoder and the machine learning model, or based on the complexity of the given problem.

[0093] Once the total energy has been determined and / or convergence within a given threshold has been achieved, the results of the autoencoder-supported energy determination are made available to the customer via a customer interface.

[0094] In some embodiments, Process 800 can resemble a subprocess within a larger context. For example, once a Hamiltonian auxiliary description is determined using Process 800, the Hamiltonian auxiliary description can then be used as input to another technique to determine a ground state of the atomic system. In another example, once a total energy of the atomic system is determined using Process 800, backpropagation can be applied to the autoencoder-based energy determination architecture, such as those described in Fig. 5A, Fig. 5B and Fig. Figure 6 shows how other properties, such as related forces of the atomic system, can also be determined. In yet another example, energy and force calculations using Process 800 can be iteratively integrated into a molecular dynamics (MD) simulation.

[0095] Fig. Figure 9 is a flowchart illustrating iterative, molecular dynamics-based processing according to some embodiments. As indicated by the arrow between update positions 910 and atom positions 902, process 900 can be computed more than once. Furthermore, and as also shown by "Update Partitioning," "Update Neighbor List," and "Update Loads," one or more steps within process 900 can be performed during each iteration, during every second iteration, during each "N ten Iteration, etc., will be executed. The following paragraphs illustrate an example execution form in which process 900 is executed for the first time.

[0096] In Block 902 the atomic positions {r→} (The atomic positions) for the respective atoms within the given atomic system are listed. As introduced above, a problem statement provided by a customer may contain an atomic description, for example, information about atomic positions and atomic species within the atomic system of interest to the problem.

[0097] In block 904, these atomic positions are partitioned across two or more processors, such as two or more processors 204, which have access to memory 208 that stores the MLIP model. In some embodiments, the processors 204 are configured to partition respective atomic positions, so that adjacent atoms are partitioned onto the same processor. This can reduce the amount of Message Passing Interface (MPI) communication that needs to occur during and / or between iterations of process 900. As indicated by the arrows, this partitioning may not need to be repartitioned at each time step of the simulation.

[0098] In block 906, an atomic neighbor list is created from the set of all atoms j within the neighborhood of atom i. For example, the neighbor list can include all atoms within a cutoff radius r. ccontained within atom i. As introduced above, atomic positions can be found within atomic descriptors. {r→j,Zj} can be defined, for example, in the atomic descriptors 502, 514, 602, and 614 for each atom j within the atomic system. As indicated by the arrows, this neighbor list may not need to be regenerated at every time step of the simulation.

[0099] In block 908, a load list can be created based on an autoencoder, for example in the Fig. 5A, Fig. 5B and Fig. The architectures described in section 6 are shown. As indicated by the arrows, these charges may not need to be regenerated at every time step of the simulation, as they can be relatively stable. In some embodiments, an additional charge rebalancing step is performed here to ensure that Σ i q i = Q total is, where Q totalThe total allocated load of the system is.

[0100] In block 910, the total energy of the atomic system is measured. E({r→j,Zj}), determined using an autoencoder-based energy determination, as for example in the Fig. 5A, Fig. 5B and Fig. Six architectures are described here. As introduced above, the energy calculation provides the atomic positions, species, charges, and / or other properties and makes them available to both an autoencoder and a machine learning model. The auxiliary Hamiltonian description and the local energies are each learned and then applied to calculate a total energy of the atomic system. Additionally, the forces can also be calculated. F→i({r→j,Zj}), can be calculated, for example by backpropagation and using the values ​​of the local energies learned by the neural network 518 or 618.

[0101] In block 912, the calculated energies and forces are used as input to update the atomic positions, for example by using an integration scheme such as leapfrog integration.

[0102] As shown by the arrow between blocks 912 and 902, MPI then communicates results between each of the processors used in partitioning to implement updated atomic positions and / or updated total energy based on that particular iteration of process 900.

[0103] As introduced above, process 900 can be iterated more than once, and blocks 904, 906, 908, and / or 910 can be updated during at least some of the subsequent iterations of process 900. For example, after determining the updated atomic positions via blocks 912 and 902, the partitioning of the atomic positions to the respective processors can be updated before proceeding with the creation of neighbor lists in block 906. In some embodiments, the partitioning can be updated every N iterations, where N can be a value such as 1000. In another example, after partitioning the atomic positions in block 904, the atomic neighbor list can be updated in block 906. In some embodiments, the neighbor lists can all be updated every M iterations, where M can be a value such as 100.In another example, the loads can be updated after the neighboring lists in block 906 have been created. Similarly, the loads may or may not be updated with each iteration, depending on the integration performed to update positions 912 during the previous iteration.

[0104] It is understood that Process 900 can be repeated any given number of times according to convergence criteria that have been set, time limits, computing power limitations, or any other number of implementation and / or customer-specific criteria.

[0105] While exemplary embodiments have been described above, it is not intended that these embodiments describe all possible forms encompassed by the claims. The terms used in the description are descriptive and not limiting, and it is understood that various modifications may be made without departing from the fundamental idea and scope of the disclosure. As previously described, the features of different embodiments can be combined to form further embodiments of the invention that may not be expressly described or illustrated.While various embodiments may be described as offering advantages or being preferred over other embodiments or implementations of the prior art with respect to one or more desired properties, the person skilled in the art recognizes that one or more features or properties may be compromised to achieve desired overall system attributes, which depend on the specific application and implementation. These attributes may include, but are not limited to, cost, strength, durability, life-cycle costs, marketability, appearance, packaging, size, ease of maintenance, weight, manufacturability, ease of assembly, etc.Therefore, insofar as embodiments are described as less desirable than other embodiments or implementations of the prior art with regard to one or more properties, these embodiments are not outside the scope of disclosure and may be desirable for certain applications.

Claims

[1] Computer-implemented method for executing a machine learning network for machine-learned interatomic potentials, comprising: Receiving data that provides an atomic description of an atomic system, wherein the atomic description includes atomic positions and atomic species of respective atoms in the atomic system; Embedding the data specifying the atomic positions and atomic species into atomic descriptors; Mapping the embedded atomic descriptors through a latent space with limited dimension of an autoencoder to learn a discretized set of auxiliary states of the atomic system; Generating a Hamiltonian auxiliary description based on the learned auxiliary states; Providing the embedded atomic descriptors and the learned auxiliary states as inputs to a machine learning model; Executing the machine learning model to learn local energies of the atomic system; and Output of a total energy of the atomic system based on Hamilton's auxiliary description and on the learned local energies. [2] Computer-implemented method according to claim 1, further comprising: Determine, based on the total energy output of the atomic system, related forces of the atomic system by backpropagation; and Spending of related forces. [3] Computer-implemented method according to claim 2, wherein: the total energy output and the related forces output of the atomic system are provided for integration within a given iteration of a molecular dynamics simulation; and The procedure further includes: Receiving a notification that during a subsequent iteration of the molecular dynamics simulation, one or more of the atomic positions were updated with respect to the atomic positions within the atomic description; Embedding the updated atomic positions and atomic species into updated atomic descriptors; and Re-mapping the updated embedded atomic descriptors and re-running the machine learning model to output an updated total energy of the atomic system. [4] Computer-implemented method according to claim 1, wherein the method further comprises: Determine, based on a learned type of auxiliary states, to enforce charge neutrality of the atomic system during mapping of the embedded atomic descriptors through the latent space with bounded dimension of the autoencoder; and Providing a specification of the required load neutrality for the autoencoder. [5] Computer-implemented method according to claim 1, wherein the method further comprises: Determine, based on a learned type of auxiliary states, not to enforce charge neutrality of the atomic system during mapping of the embedded atomic descriptors by the latent space with limited dimension of the autoencoder; and Providing a statement regarding the non-restriction of load neutrality for the autoencoder. [6] Computer-implemented method according to claim 1, further comprising: additionally receiving a specification of a type of auxiliary state to be learned; and Determining a dimension for the latent space with bounded dimension to be applied, at least partially based on the complexity of the atomic system or the type of auxiliary states. [7] Computer-implemented method according to claim 1, wherein the discretized set of auxiliary states is one or more of: charge states; Oxidation states; or magnetic states. [8] Computer-implemented method according to claim 1, wherein the autoencoder is a variation autoencoder, a regularized autoencoder or a sparse autoencoder. [9] Computer-implemented method according to claim 1, wherein the machine learning model is a deep neural network or one or more Gaussian processes. [10] Computer-implemented method for executing a machine learning network for machine-learned interatomic potentials, comprising: Receiving data specifying a request from a customer to determine the total energy of an atomic system, wherein the request includes atomic positions and atomic species of respective atoms in the atomic system; Embedding the data specifying the atomic positions and atomic species into atomic descriptors; Mapping the embedded atomic descriptors through a latent space with limited dimension of an autoencoder to learn a discretized set of auxiliary states of the atomic system; Generating a Hamiltonian auxiliary description based on the learned auxiliary states; Executing a machine learning model to learn local energies of the atomic system based on the embedded atomic descriptors; Output of the total energy of the atomic system based on Hamilton's auxiliary description and on the learned local energies; and Providing the total energy for the customer. [11] Computer-implemented method according to claim 10, wherein: The atomic descriptors are inputs for interatomic potentials that describe the atomic system; and the atomic descriptors are invariant or covariant with respect to a symmetry group of the atomic system. [12] Computer-implemented method according to claim 10, wherein: Providing the learned auxiliary states as an input to the machine learning model; and Executing the machine learning model to learn the local energies of the atomic system based on the embedded atomic descriptors and the learned auxiliary states. [13] Computer-implemented method according to claim 10, wherein the discretized set of auxiliary states is one or more of: charge states; oxidation states; or magnetic states. [14] Computer-implemented method according to claim 10, wherein: The requirement further includes a specification of a type of discretized set of auxiliary states to be learned; and The procedure further includes determining a dimension for the latent space with bounded dimension, which is to be applied, at least partially based on the complexity of the atomic system or the type of auxiliary states. [15] Computer-implemented method according to claim 14, further comprising: Determine, based on the requirement to enforce charge neutrality of the atomic system during the mapping of the embedded atomic descriptors by the latent space with limited dimension of the autoencoder; and Providing a specification of the required load neutrality for the autoencoder. [16] Computer-implemented method according to claim 14, further comprising: Determine, based on the requirement not to enforce charge neutrality of the atomic system during mapping of the embedded atomic descriptors by the latent space with limited dimension of the autoencoder; and Providing a statement regarding the non-restriction of load neutrality for the autoencoder. [17] Non-volatile, computer-readable medium that stores program instructions which, when executed on or through one or more processors, cause the one or more processors to: Receiving data that provides an atomic description of an atomic system, wherein the atomic description includes atomic positions and atomic species of respective atoms in the atomic system; Embedding the data specifying the atomic positions and atomic species into atomic descriptors; Mapping the embedded atomic descriptors through a latent space with limited dimension of an autoencoder to learn a discretized set of auxiliary states of the atomic system; Generating a Hamiltonian auxiliary description based on the learned auxiliary states; Executing a machine learning model to learn a local scalar vector or a local tensor property of the atomic system based on the embedded atomic descriptors; and Outputting an overall property of the atomic system based on Hamilton's auxiliary description and on the learned local scalar vector or the learned local tensor property. [18] Non-volatile, computer-readable medium according to claim 17, wherein the program instructions further cause the one or more processors to: Providing the learned auxiliary states as an input to the machine learning model; and Executing the machine learning model to learn the local scalar vector or local tensor property of the atomic system based on the embedded atomic descriptors and the learned auxiliary states. [19] Non-volatile, computer-readable medium according to claim 17, wherein the discretized set of auxiliary states is one or more of: charge states; oxidation states; or magnetic states. [20] Non-volatile, computer-readable medium according to claim 17, wherein: the local scalar vector or the local tensor property is local energies and the overall property is total energy; or the local scalar vector or local tensor property is related forces and the overall property is total force.