Prediction of the equilibrium distribution of molecular systems
The computing system uses a diffusion model with graph neural networks to predict molecular system equilibrium distributions, addressing inefficiencies in current methods by providing accurate and cost-effective predictions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2023-03-16
- Publication Date
- 2026-04-14
AI Technical Summary
Current computational methods struggle to efficiently predict the equilibrium distribution of molecular systems, which is crucial for understanding the properties and functions of molecules, as they are computationally complex and require expensive data collection.
A computing system utilizing a diffusion model based on graph neural networks and stochastic processes to predict the equilibrium distribution of molecular systems, enabling efficient sampling and accurate estimation of free energy, even with limited data.
Enables efficient and accurate prediction of molecular system equilibrium distributions, facilitating the understanding of molecular properties and functions, and reducing the computational complexity and cost associated with data collection.
Smart Images

Figure 2026511390000001_ABST
Abstract
Description
[Background technology]
[0001] background
[0001] In the field of computational chemistry, computer-based techniques have been developed to predict molecular properties through computer simulations. These molecular properties can have a wide range of effects on the appearance and function of molecules or materials, and are therefore of great interest in a variety of fields. For example, in the field of drug design, changes in molecular properties can affect the efficacy of drugs. In the field of drug discovery, molecular properties can affect the possibility of naturally occurring materials being used for therapeutic purposes. In the field of quantum chemistry, the quantum mechanical calculation of the electron contribution to the physical and chemical properties of molecules and materials is a fundamental area of exploration. As will be discussed below, there are still opportunities to improve computational methods for predicting molecular properties, which would have useful applications well beyond the field of computational chemistry. [Overview of the Initiative] [Means for solving the problem]
[0002] overview
[0002] To address the problems discussed herein, a computing system and method for predicting the equilibrium distribution of a molecular system are provided. In one embodiment, the computing system includes a processor that executes instructions using a portion of the associated memory to implement an equilibrium distribution prediction model. In the inference stage, the processor is configured to receive input data representing the molecular system and to create a graph representation of the molecular system, including positional information of each atom in the molecular system. The processor is further configured to input the graph representation of the molecular system into a graph neural network and to receive a plurality of predicted conformations of the molecular system as output from the graph neural network. The processor is further configured to input each of the plurality of predicted conformations into a diffusion model, to predict the equilibrium distribution of the conformations of the molecular system, and to output the equilibrium distribution.
[0003]
[0003] This summary is provided in a simplified form to introduce certain concepts that will be further described in the following detailed description. This summary is not intended to identify any important or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to any implementation that solves any defects described in any part of this disclosure. [Brief explanation of the drawing]
[0004] Brief explanation of the drawing [Figure 1]
[0004] A schematic diagram of a computing system for predicting the equilibrium distribution of a molecular system according to one embodiment of the present disclosure is shown. [Figure 2A]
[0005] Figure 1 shows the training stages of the equilibrium distribution prediction model for the computing system. [Figure 2B]
[0005] Figure 1 shows the training stages of the equilibrium distribution prediction model of the computing system. [Figure 3]
[0006] Figure 1 shows an exemplary free energy landscape of the equilibrium distribution predicted by the computing system. [Figure 4]
[0007] Figure 1 shows the inference stage of the trained equilibrium distribution prediction model of the computing system. [Figure 5]
[0008] A flowchart illustrating a method for predicting the equilibrium distribution of a molecular system according to an exemplary embodiment of this disclosure is shown. [Figure 6]
[0009] This document illustrates exemplary computing environments in which several embodiments of this disclosure may be implemented. [Modes for carrying out the invention]
[0005] Details of the invention
[0010] The rise of deep learning models has led to rapid advancements in predicting the molecular properties of molecular systems, culminating in significant milestones such as accurate protein structure prediction from amino acid sequences. However, comparing the conformation of a molecular system at its lowest energy to the distribution of its possible configurations—that is, the energy landscape within its conformational space—remains more relevant to its macroscopic properties. Therefore, developing computational tools to predict the microscopic states of molecules within an equilibrium population will provide significant progress in understanding the properties and functions of molecular systems, and will enable the evaluation of macroscopic properties using statistical mechanics methods.
[0006]
[0011] However, developing efficient sampling methods to predict the equilibrium distribution of molecular systems remains a computational challenge. Rapid advances in deep learning techniques have enabled data-driven methods to be used without issue to predict molecular structures corresponding to the lowest energy of molecular systems, resulting in breakthroughs in this area. For example, artificial intelligence (AI) structure prediction models have improved the accuracy of protein structure prediction at the atomic level, leading to successful applications in structural biology. The development of high-speed computational docking methods based on deep neural networks for ligand-binding structure prediction enables rapid virtual screening and drug engineering. Deep learning models have been designed to predict the relaxation structure of adsorbates on catalyst surfaces. While each of these developments demonstrates that deep learning methods offer promising solutions for understanding microscopic structures and states, accurate prediction of the lowest energy structure of molecular systems reveals only a fraction of the information needed to understand molecular systems in equilibrium. In reality, molecules can take on different structures depending on the probability governed by the energy function when the system reaches equilibrium. This equilibrium distribution is important for studying statistical mechanical properties. For example, the properties and functions of biomolecules can be inferred from diverse structures with variable energies. In addition, determining the equilibrium distribution is necessary to calculate the free energy associated with many important macroscopic properties.
[0007]
[0012] Computational methods for efficiently predicting the equilibrium distribution of the conformational states of a molecular system directly from basic descriptors (e.g., the amino acid sequence of a protein) would greatly improve the understanding of the functions and properties of the molecular system. However, for such models to be practically useful, for example, it is necessary to support independent sampling that enables efficient exploration in microstates separated by energy barriers in conformational space with a fluctuating energy landscape, and it would be necessary to accurately predict the density states for efficient calculation of the free energy of the state of interest. Independent sampling overcomes the inherent problems in sequential sampling methods, such as Markov Chain Monte Carlo (MCMC) or Molecular Dynamics (MD), which have relatively low sampling efficiency due to independent and non-identical distributions and due to sample correlations. Predicting the equilibrium distribution over all conformations is much more technically difficult than predicting a single conformation corresponding to the most stable, e.g., lowest energy, state, especially when the distribution is highly complex, due to anisotropy and multiple local optima.
[0008]
[0013] Design Principles
[0014] In view of the problems discussed above, a computing system is provided that utilizes an equilibrium distribution prediction model. The computing system has applicability to predicting the equilibrium distribution of a molecular system and the distribution of other types of systems that can be represented as graphs. The following discussion provides an overview of the theoretical basis and design principles behind which the equilibrium distribution prediction model was conceived. Following this discussion, specific exemplary embodiments of the equilibrium distribution prediction model are described in detail.
[0009]
[0015] Framework
[0016] Deep neural networks have been demonstrated to predict accurate molecular structures from molecular descriptors D. The equilibrium distribution prediction model described herein predicts the equilibrium distribution on the state space. To achieve this, the equilibrium distribution prediction model is based on the following diffusion model. When noise to the system is gradually downscaled and injected, the equilibrium distribution finally becomes a standard Gaussian distribution, making prediction easier and thus enabling efficient sampling. A forward diffusion process is implemented to reach the Gaussian distribution, and then a subsequent reverse diffusion process is used to gradually recover the original equilibrium distribution.
[0010]
[0017] Previous approaches for predicting the equilibrium distribution from a single distribution, such as simulation annealing using the Monte Carlo method and simulation annealing using Langevin dynamics, have the drawbacks of being computationally complex, slow, and expensive to implement. As described herein, these issues in predicting the equilibrium distribution of a molecular system can be reduced by formulating the equilibrium distribution of the molecular system into a probabilistic process based on the forward diffusion process and the reverse diffusion process.
[0011]
[0018] The forward diffusion process and the reverse diffusion process can be regarded as a pair of mutual probabilistic processes that simulate the conversion between the equilibrium distribution and the system-independent single distribution p , D,0 ,
[0019] , . The equilibrium distribution of the molecular descriptor D is taken as the initial distribution q D,0 , and a probabilistic process of the molecular descriptor D is constructed that gradually converts q simple in the direction of p D,0 through a period τ in the forward process. Next, the corresponding reverse process converts p simple to the equilibrium distribution q D,0 . This defines the generation process of the equilibrium distribution prediction model.
[0012]
[0019] For actual simulations, the deep neural network is trained from the forward process to predict the reverse process from a given system descriptor D. Compared to directly predicting the desired distribution from the system descriptor D, the mutual forward and reverse diffusion methods described herein significantly reduce the difficulty of this problem. p simple is selected such that its samples can be drawn independently and has a closed-form density function, so the equilibrium distribution prediction model is p simple By simulating the reverse process from the samples, it is possible to independently sample the equilibrium distribution and also provide the density function of the distribution by tracking the simulation. Specifically, P simple := N(0,1) was selected as the standard Gaussian distribution. The forward process is constructed as Langevin dynamics (Ornstein-Uhlenbeck process) targeting p simple (where the time expansion scheme β t increases with t) and is described as a stochastic differential equation (SDE) in Equation (1).
Number
Number
[0013]
[0020] According to the theory of stochastic process, the reverse process is also a stochastic process that follows the steps of a reverse-time Markov chain described as the following SDE in Equation (2).
Number
[0014]
[0021] Here,
number
number
number
number
number
[0015]
[0022] Since the model is defined in conformational space, a graph neural network trained to predict the conformation of a molecular system is implemented as the backbone of the neural network architecture of the equilibrium distribution prediction model, enabling the modeling of molecular structures and their generalization across systems.
number
number
number
[0016]
[0023] Here, the index i is defined as time t = ih and β i :=hβ t=ih This corresponds to the direct Euler-Maruyama discretization of equation (2) has o(h) discretization error.
number
number
number
number
[0017]
[0024] Fine-tuning the equilibrium distribution prediction model using data.
[0025] The scoring model is
number
number
number
[0018]
[0026] In other words, it is Fisher's identity. The score matching loss is as follows:
number
[0019]
[0027] Here, the second term in the final representation is a constant of θ. Optimizing the score matching loss is equivalent to minimizing the first term in the final representation. This term is shown below as the support equation (S1).
number
[0020]
[0028] Noise reduction score matching loss.
[0029] The optimization of the denoising score matching loss equation (S1) is fortunately available in the closed form of equation (1) for the forward process, which is a conditional distribution.
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0021]
[0030] Diffusion pre-training based on physical information
[0031] Equilibrium distribution prediction models can be trained with conformational data samples across a series of molecular systems. However, collecting sufficient experimental and / or simulation data to well characterize the equilibrium distributions of various systems can be prohibitively expensive. To address the data scarcity problem, a pre-training algorithm called Physically Information-Based Diffusion Pre-training (PIDP) is implemented to effectively optimize equilibrium distribution prediction models in the absence of data. The training information comes from the energy function of each molecular system descriptor D, which is already well defined in the equilibrium distribution. The true score function ∇logq is derived from equation (1) of the forward process. D,T It is understood that this follows a partial differential equation known as the Fokker-Planck equation. Therefore, the score model
number
number
[0022]
[0032] Here, λ1 weights the initial conditions, and
number
number
number
number
number
number
[0023]
[0033] Calculation of free energy
[0034] In addition to generating IID samples, the equilibrium distribution prediction model can also evaluate the density function of the distribution defined as shown in equation (6).
number
[0024]
[0035] Here, D is the dimension of the state space, and
number
number
[0025]
[0036] The initial condition R0, which can be solved using the ODE solver, provides another aspect of the distribution and allows for more calculations related to the distribution. For example, the free energy of molecular system descriptor D can be estimated using the following:
number
[0026]
[0377] Here, the expected value can be estimated by independent samples that can be easily derived using an equilibrium distribution prediction model.
number
number
number
number
number
number
number
[0027]
[0038] To evaluate the density of the equilibrium distribution prediction model, the model uses a reverse diffusion process to p simple Note that the distribution is defined by a sequence of distributional transformations. Under continuous time constraints, the general diffusion process
number
number
[0028]
[0039] Here,
number
number
number
number
[0029]
[0040] Here,
number
number
[0030]
[0041] Here,
number
number
number
number
number
number
number
[0031]
[0042] Compared to the FPE type (S2), this indicator
number
number
[0032]
[0043] Compared with equation (S5), this gives the following equation:
number
[0033]
[0044] Integrating with respect to t gives the following equation:
number
[0034]
[0045] This matches equation (6).
[0035]
[0046] Score model for each process
number
number
number
[0036]
[0047] By taking the slope of the above equation, we obtain the following equation.
number
[0037]
[0048] This is the score function
number
number
number
number
[0038]
[0049]
number
number
number
[0039]
[0050] To combine loss terms at various time steps, the term at t is expressed in the same way as for database loss, i.e.,
number
[0040]
[0051] Alternatively, for weighting schemes with different t values, the weighting of the time step t is inversely proportional to the following equation:
number
[0041]
[0052] FPE is for each q D, As long as t is normalized, there are no or no boundary conditions. With respect to the initial conditions, the target equilibrium distribution
number
number
[0042]
[0053] Inverse design using equilibrium distribution prediction models
[0054] Another benefit of the equilibrium distribution prediction model is that it allows for the direct prediction of a conditioned distribution with respect to a given microscopic characteristic c, without the need to generate data for the conditional distribution and without the need to retrain the model. This is necessary for more demanding inverse design tasks where a state satisfying specific characteristics, such as a base gap and conductivity, is desired. From the data generation process in equation (2), what needs to be fitted is:
number
number
number
number
[0043]
[0055] Interpolation between states
[0056] Equilibrium distribution prediction models can also reveal reaction pathways between two given states, which can be used to discover reaction coordinates or collective variables, and can discover intermediate metastable states in transitions. This is because equation (1), the distribution transformation process behind the equilibrium distribution prediction model, is deterministic and reversible, and is therefore equivalent to the process in equation (7) that establishes the correspondence between real space and potential space.
number
[0044]
[0057] Exemplary Embodiments
[0058] In accordance with the principles discussed above, specific exemplary embodiments of the equilibrium distribution prediction model according to this disclosure are described here with reference to Figures 1-6.
[0045]
[0059] Referring first to Figure 1, the computing system 10 includes at least one computing device. The computing system 10 is shown having a first computing device 14 including a processor 18 and memory 22, and a second computing device 16 including a processor 20 and memory 24. The embodiments shown are illustrative in nature, and other configurations are possible. In the following description, the first computing device is described as a server 14, and the second computing device is described as a client computing device 16, and the respective functions performed by each device are described. In other configurations, it is recognized that the computing system 12 includes a single computing device that performs the essential functions of both the server 14 and the client computing device 16, and the first computing device may be a computing device other than a server. In other alternative configurations, the functions described as being performed by the server 14 may be performed by the client computing device 16 instead, and vice versa.
[0046]
[0060] Continuing with Figure 1, the processor 18 is configured to execute instructions using a portion of the associated memory 22 to implement an equilibrium distribution prediction model 26 hosted on the server 14. At a high level, the equilibrium distribution prediction model 26 processes data 28 representing molecular systems to predict the equilibrium distribution of the molecular system's conformation based on a score function of each intermediate distribution over a set of predetermined time steps relating to the diffusion process, i.e., the slope of the logarithm of the density function. As will be described in detail below with respect to Figure 2A, the data 28 representing a molecular system, which includes at least one of the amino acid sequence D1 of the molecular system, the two-dimensional structure D2 of the molecular system, and / or the chemical formula D3 of the molecular system, is a descriptor D. Additionally or alternatively, the descriptor D may be the three-dimensional structure of the molecular system, the protein name, the Simple Molecular Input Line Input System (SMILES), etc. The data 28 may be stored in molecular system databases 30 such as UniProt, Swiss-Prot, or the Protein Research Foundation (PRF). When the client computing device 16 receives user input via the user interface 32, the data 28 can be sent to the graph representation generator 34 included in the equilibrium distribution prediction model 26.
[0047]
[0061] The graph representation generator 34 is configured to receive input data 28. Based on the data 24 representing the molecular system, the graph representation generator 34 creates a graph representation 36 of the molecular system, which includes the positional information of each atom in the molecular system. A preprocessing algorithm may be implemented to create the graph representation, as described below with reference to Figure 2A. The graph representation 36 is then input to the graph neural network 38. The graph neural network 38 is configured to generate multiple predicted conformations 40 of the molecular system. The predicted conformations 40 may be stored as predicted conformation data 44 in the predicted conformation database 42.
[0048]
[0062] Each of the predicted conformations 40 output from the graph neural network 38 is input to the diffusion model 46. The equilibrium distribution 48 of the molecular system is predicted, as will be explained in detail above and below with respect to Figures 2A, 2B, and 3. Data 50 representing the predicted equilibrium distribution 48 may be stored in the equilibrium distribution database 52.
[0049]
[0063] The balanced distribution 48 is output and can be observed, for example, on the user interface 54 of the display 56 of the client computing device 16. In the embodiments described herein, it is understood that the server 14 communicates with the client computing device 16 via the network 58.
[0050]
[0064] Moving on to Figures 2A and 2B, the training phase of the equilibrium distribution prediction model 26 is shown. As discussed and shown above in Figure 2A, the graph representation 36 of the molecular system is generated based on the descriptor D of the molecular system via the preprocessing algorithm 60. The graph representation 36 includes multiple normal nodes connected by edges, as indicated by the keys in Figure 2A. Each normal node represents an atom in the molecular system. The graph representation 36 further includes one virtual node, which is fully connected to all normal nodes of the graph representation 36 by virtual edges. The difference between virtual nodes and normal nodes is understood to be that normal nodes represent atoms, while virtual nodes are provided solely for computational purposes and therefore do not represent any physical components of the molecular system.
[0051]
[0065] Based on the graph representation 36, the processor 12 of the computing system 10 is configured to provide, for example, a training dataset 62 for a graph neural network 38, which includes multiple training data pairs, by computationally generating or reading it from a storage location in memory. Each training data pair includes the graph representation 36 of the molecular system and energy parameter values 64 that represent energy changes within the molecular system that may be due to energy transformations resulting from, for example, molecular relaxation of the molecular system.
[0052]
[0066] As described above, the graph neural network 38 is configured to generate multiple predicted conformations 40 of molecular systems, each of which is output from the graph neural network 38 and input to the diffusion model 46 (shown in Figure 2B). The diffusion model 46 is a score-based generative diffusion model configured to introduce random noise at predetermined time steps based on a Gaussian distribution in the forward process, and to reduce the noise in the reverse process. As discussed in detail in the "Framework" section, the forward process is defined as a Markov chain of steps, for example, in equation (1), p simple It is constructed as Langevin dynamics targeting the (Ornstein-Uhlenbeck process). The reverse process is also a stochastic process that follows a reverse-time Markov chain of several steps. Equation (2) is an example of a reverse process. The diffusion model additionally implements a score model defined by equation (3) to predict the true score function of the distribution of multiple instances from a molecular system descriptor D. The score model can be trained by equation (4) to minimize the loss function described above, see, for example, the "Physically Informed Diffusion" section.
[0053]
[0067] Equilibrium distribution prediction models can be trained using conformational data samples across a series of molecular systems; however, collecting sufficient experimental and / or simulation data to adequately characterize the equilibrium distributions of various systems can be prohibitively expensive. Therefore, PIDP is implemented to optimize equilibrium distribution prediction models when data is unavailable. As described above, the diffusion model follows a Boltzmann distribution and is trained using molecular dynamics simulation data, with the energy function of each molecular system descriptor D serving as training information.
[0054]
[0068] The diffusion model 46 is configured to generate an equilibrium distribution by varying the noise applied to the predicted conformation over a series of time steps shown in Figure 2B as T1, T2, and T3 instances of the model. This noise increases through the simulation over the time steps. The increased noise between time steps indicates that there is a greater potential change in the conformation predicted by the model during the reverse diffusion process. Advancing the series of time steps draws out the IID sample to generate the equilibrium distribution 48, driving the conformation of the molecular system to follow a single distribution (e.g., a standard Gaussian distribution) so that its density function is evaluated in the reverse diffusion process. Shaded regions within the equilibrium distribution 48 represent various levels of free energy associated with each conformation of the molecular system, with the lowest level of free energy having the darkest shading to represent a region of stability. The two stability regions shown in Figure 2B may represent two different stable states of the molecular system, e.g., the bound / unbound conformation of an enzyme or the open / closed state of a membrane channel.
[0055]
[0069] Figure 3 shows an exemplary free energy landscape of the equilibrium distribution 48. As discussed above, the shaded regions of the equilibrium distribution 48 represent the free energies associated with various conformations of the molecular system. The darkest regions represent the regions with the lowest free energies corresponding to the metastable conformations of the molecular system, as shown in Figure 3, by conformations C1 and C3. Conformation C2 is the semistable transition conformation of the molecular system when transitioning between the metastable states of conformations C1 and C3.
[0056]
[0070] The inference stage of the equilibrium distribution prediction model 26 is shown in Figure 4. As described in detail above with reference to Figures 1, 2A and 2B, the equilibrium distribution prediction model 26 receives a molecular system descriptor D, creates a graph representation 36 of the molecular system based on the molecular system descriptor D via a preprocessing algorithm 60, and inputs the graph representation 36 into the diffusion model 46.
[0057]
[0071] The equilibrium distribution prediction model 26 additionally includes an independent identical-distribution (IID) sampler 66 that performs conformational sampling of the equilibrium distribution 48 to generate statistically independent samples. One conventional method to such statistical sampling is to use a simulation annealing algorithm, an adaptation of the Metropolis-Hastings algorithm, to explore the solution space and arrive at the complex conformational distribution of the target. The simulation annealing algorithm can be used to make multiple predictions from multiple initial states of a molecular system and arrive at the distribution of the predicted conformations of the molecular system. Temperature is a parameter within the simulation annealing algorithm that governs how far the algorithm is allowed to explore in the direction away from the target (minimum or maximum) in a given time step. Typically, temperature decreases over the simulation to allow the solution to converge. While useful, this conventional class of simulation annealing algorithms is computationally complex, slow, and expensive to implement. This is due to the fact that the simulation of the distribution change is an approximation (when using Markov chain Monte Carlo sampling) or experiences particle degeneracy (when using importance sampling). In contrast to such simulation annealing methods, to avoid such computational complexity, this disclosure utilizes a diffusion model that takes the conformation of a molecular system from a Gaussian distribution to a complex target distribution via an explicit diffusion process that can be simulated more accurately.
[0058]
[0072] The equilibrium distribution prediction model 26 further includes a density module 68 configured to determine the density function of the equilibrium distribution 48 based on computational statistics. The diffusion process formulation allows for the estimation of the density function of the target distribution by solving an ordinary differential equation, as described in detail above in the "Calculation of Free Energy" section.
[0059]
[0073] Figure 5 shows a flowchart of method 500 for predicting the equilibrium distribution of a molecular system. Method 500 can be carried out by the hardware and software of the computing system 10 described above or other suitable hardware and software. In step 502, method 500 may include pre-training a diffusion model via a Physically Information-Based Diffusion Pre-Training (PIDP) algorithm during the training phase. As described above, an equilibrium distribution prediction model can be trained with conformational data samples across a series of molecular systems, but collecting sufficient experimental and / or simulation data to well characterize the equilibrium distributions of various systems can be prohibitively expensive. Therefore, PIDP is implemented to optimize the equilibrium distribution prediction model when data is unavailable.
[0060]
[0074] Proceeding from step 502 to step 504, method 500 may further include training an energy function following a Boltzmann distribution as training information. The training information arises from the energy function of each molecular system descriptor D, which is already well defined in the equilibrium distribution. When the energy of a sample descriptor is unknown, calculating the free energy difference is made possible using a data sample that follows a Boltzmann distribution. Compared to conventional free energy calculation methods, the equilibrium distribution predictive model does not rely on harmonic approximations for the energy function near the metastable state, and IID sample generation allows for faster sample coverage across the relevant state space than MD.
[0061]
[0075] Moving from step 504 to step 506, method 500 may further include training a diffusion model using molecular dynamics (MD) simulation data as training information. While using dynamic processes to characterize the distribution is inefficient, MD simulations are used to generate sufficient data that can be applied as training functions in equilibrium predictive distribution models.
[0062]
[0076] In step 508, method 500 may further include receiving input data representing a molecular system during the inference stage. As described in detail above, the data representing a molecular system may be a descriptor including at least one of the amino acid sequence, two-dimensional structure, and chemical formula of the molecular system. Additionally or alternatively, the descriptor may be the three-dimensional structure, protein name, or Simple Molecular Input Line Input System (SMILES). The data may be stored in molecular system databases such as UniProt, Swiss-Prot, or the Protein Research Foundation (PRF).
[0063]
[0077] Continuing from step 508 to step 510, method 500 may further include creating a graph representation of the molecular system, which includes the positional information of each atom in the molecular system. A preprocessing algorithm may be implemented to create the graph representation, as described above with reference to Figure 2A. The graph representation may include a plurality of normal nodes connected by edges, each normal node representing an atom in the molecular system. The graph representation may further include virtual nodes, which are sufficiently connected to all normal nodes in the graph representation by virtual edges.
[0064]
[0078] Moving from step 510 to step 512, method 500 may further include inputting a graph representation of the molecular system into a graph neural network. The graph neural network may be configured to generate multiple predicted conformations of the molecular system, which can be stored in a predicted conformation database as predicted conformation data.
[0065]
[0079] Proceeding from step 512 to step 514, method 500 may further include receiving multiple predicted conformations of the molecular system as output from a graph neural network. Continuing from step 514 to step 516, method 500 may further include inputting each of the multiple predicted conformations into a diffusion model. The diffusion model may be a score-based generative diffusion model configured to introduce random noise at predetermined time steps based on a Gaussian distribution in the forward process and to reduce the noise in the reverse process. As described in detail above, the forward process is defined as a Markov chain of steps, p simple The Langevin dynamics (Ornstein-Uhlenbeck process) targeting p can be constructed, and the reverse process can also be a stochastic process following an inverse time Markov chain of steps. The diffusion model can additionally implement a scoring model that is trained to predict the true score function of multiple instances of the distribution from a molecular system descriptor by minimizing the loss function. Furthermore, the diffusion model can be configured to vary the noise scale of each predicted conformation among multiple predicted conformations over a set of predetermined time steps. By advancing the time steps to inject increased levels of noise into the molecular system, the conformation of the system becomes p simple It is driven to follow a specific pattern and utilizes a score model to simulate a reverse diffusion process, allowing the score model to generate an equilibrium distribution.
[0066]
[0080] Proceeding from step 516 to step 518, method 500 may further include receiving a predicted equilibrium distribution of the molecular system. Data representing the predicted equilibrium distribution may be stored in an equilibrium distribution database. The equilibrium distribution may be output to a user interface on the display of a client computing device and observed.
[0067]
[0081] The equilibrium distribution predictive models described herein enable sampling without simulation or model retraining and may be more efficient and feasible for large or complex molecular systems. Equilibrium distribution predictive models can sample from multiple long-term states that may be difficult to access using conventional methods. In addition, equilibrium distribution predictive models have the potential to accurately calculate free energies, which is important in many areas of chemistry and physics. Finally, equilibrium distribution predictive models can inform inverse design and facilitate the development of new materials or molecules with specific properties.
[0068]
[0082] The use of diffusion pre-training based on physical information allows models to satisfy physical constraints and perform well in the case of insufficient data, while diffusion-based generative modeling facilitates the reduction of problem complexity by breaking down the complexity of the problem into smaller pieces. The advanced neural network architectures of our models, such as the use of graph neural networks, enable more effective processing and understanding of descriptors of molecular systems. While the development of deep learning-based structure prediction techniques has advanced rapidly in recent years, elucidating the equilibrium distribution of the microscopic states of molecular systems remains a challenge. By developing accurate equilibrium distribution prediction methods and leveraging large simulation databases assembled by the simulation community and well-developed force fields, the equilibrium distribution prediction models described herein have the potential to accelerate the development of computational science and make a significant contribution to the study of molecular systems and beyond.
[0069]
[0083] In some embodiments, the methods and processes described herein may relate to computing systems of one or more computing devices. In particular, such methods and processes may be implemented as computer application programs or services, application programming interfaces (APIs), libraries and / or other computer program products.
[0070]
[0084] Figure 6 schematically illustrates a non-limiting embodiment of a computing system 600 that can carry out one or more of the methods and processes described above. The computing system 600 is shown in a simplified form. The computing system 600 can embody the computer device 10 described above and shown in Figure 1. The computing system 600 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones) and / or other computing devices, as well as wearable computing devices such as smartwatches and head-mounted augmented reality devices.
[0071]
[0085] The computing system 600 includes a logical processor 602, a volatile memory 604, and a non-volatile storage device 606. The computing system 600 may optionally include a display subsystem 608, an input subsystem 610, a communication subsystem 612, and / or other components not shown in Figure 1.
[0072]
[0086] The logical processor 602 includes one or more physical devices configured to execute instructions. For example, the logical processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform a task, implement a data type, transform the state of one or more components, achieve a technical effect, or otherwise reach a desired result.
[0073]
[0087] A logic processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, a logic processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logic processor 602 may be single-core or multi-core, and the instructions executed thereon may be configured for serial, parallel, and / or distributed processing. Individual components of the logic processor may optionally be distributed across two or more separate devices that may be remotely located and / or configured for cooperative processing. Embodiments of the logic processor may be virtualized and executed by remotely accessible network computing devices configured in a cloud computing configuration. In such cases, it is understood that these virtualized embodiments run on various physical logic processors of diverse machines.
[0074]
[0088] The non-volatile storage device 606 includes one or more physical devices configured to hold instructions executable by a logical processor for carrying out the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile storage device 606 can be changed, for example, to hold various types of data.
[0075]
[0089] The non-volatile storage device 606 may include removable and / or embedded physical devices. The non-volatile storage device 606 may include optical memory (e.g., CD, DVD, HD DVD, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, FLASH memory, etc.), and / or magnetic memory (e.g., hard disk drive, floppy disk drive, tape drive, MRAM, etc.), or other mass storage device technologies. The non-volatile storage device 606 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location addressable, file addressable, and / or content accessable devices. The non-volatile storage device 606 is recognized as being configured to retain instructions even when power to the non-volatile storage device 606 is cut off.
[0076]
[0090] The volatile memory 604 may include a physical device containing random access memory. The volatile memory 604 is typically used by the logical processor 602 to temporarily store information during the processing of software instructions. The volatile memory 604 is typically recognized as not continuing to store instructions when power to the volatile memory 604 is cut off.
[0077]
[0091] Embodiments of the logic processor 602, volatile memory 604, and non-volatile storage device 606 may be integrated together within one or more hardware logic components. Such hardware logic components may include, for example, field-programmable gate arrays (FPGAs), program and application-specific integrated circuits (PASICs / ASICs), program and application-specific standard products (PSSPs / ASSPs), systems-on-a-chip (SOCs), and complex-programmable logic devices (CPLDs).
[0078]
[0092] The terms “module,” “program,” and “engine” may be used to describe a form of computing system 600 that is typically implemented in software by a processor to perform a specific function using a portion of volatile memory. This function involves a translation process that specifically configures the processor to perform this function. Thus, a module, program, or engine may be instantiated via a logical processor 602 that executes instructions held by a non-volatile storage device 606 using a portion of volatile memory 604. It is understood that various modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine may be instantiated by various applications, services, code block, object, routine, API, function, etc. The terms “module,” “program,” and “engine” may encompass individual executable files, data files, libraries, drivers, scripts, database records, etc., or groups thereof.
[0079]
[0093] If included, the display subsystem 608 may be used to present a visual representation of the data held by the non-volatile storage device 606. The visual representation may take the form of a graphical user interface (GUI). If the methods and processes described herein modify the data held by the non-volatile storage device and thereby transform the state of the non-volatile storage device, the state of the display subsystem 608 may also be transformed to visually represent the change in the underlying data. The display subsystem 608 may include one or more display devices utilizing substantially any type of technology. Such display devices may be combined with the logical processor 602, volatile memory 604 and / or non-volatile storage device 606 in a shared enclosure, or such display devices may be peripheral display devices.
[0080]
[0094] If included, the input subsystem 610 may include or interface with one or more user input devices such as a keyboard, mouse, touchscreen, or game controller. In some embodiments, the input subsystem may include or interface with selected natural user input (NUI) components. Such components may be integrated or peripheral components, and the conversion and / or processing of input actions may be handled on or off the board. Exemplary NUI components may include microphones for speech and / or speech recognition, infrared cameras, color cameras, stereo cameras, and / or depth cameras for machine vision and / or gesture recognition, head trackers, eye-trackers, accelerometers, and / or gyroscopes for motion detection and / or intent recognition, and electric field sensing components and / or any other suitable sensors for evaluating brain activity.
[0081]
[0095] If included, the communication subsystem 612 may be configured to communicate with each other and with other devices, depending on the type of computing device described herein. The communication subsystem 612 may include wireless and / or wireless communication devices compatible with one or more different communication protocols. In non-limiting examples, the communication subsystem may be configured for communication over a wireless telephone network, or over a wired or wireless local or wide area network (such as HDMI over a Wi-Fi connection). In some embodiments, the communication subsystem may enable the computing system 600 to send and / or receive messages to and from other devices over a network such as the Internet.
[0082]
[0096] The following paragraphs provide further explanation of aspects of the present disclosure. One aspect provides a computing system for predicting the equilibrium distribution of a molecular system. The computing system may include a processor that executes instructions using a portion of the associated memory to implement an equilibrium distribution prediction model. The processor may be configured to, in the inference stage, receive input data representing the molecular system; create a graph representation of the molecular system including positional information of each atom in the molecular system; input the graph representation of the molecular system into a graph neural network; receive a plurality of predicted conformations of the molecular system as output from the graph neural network; input each of the plurality of predicted conformations into a diffusion model; predict the equilibrium distribution of the molecular system; and output the predicted equilibrium distribution.
[0083]
[0097] In this embodiment, additionally or alternatively, the diffusion model may vary the noise applied to the predicted conformation over a series of time steps, thereby generating an equilibrium distribution.
[0084]
[0098] In this embodiment, the equilibrium distribution prediction model may additionally or alternatively include an independent and identically distributed (IID) sampler module that generates statistically independent samples from the equilibrium distribution.
[0085]
[0099] In this embodiment, the equilibrium distribution prediction model may additionally or alternatively include a density module that determines the density function of the equilibrium distribution.
[0086]
[0100] In this embodiment, the diffusion model may be additionally or alternatively pre-trained via a diffusion pre-training algorithm based on physical information.
[0087]
[0101] In this embodiment, the diffusion model may be additionally or alternatively trained to minimize the loss function.
[0088]
[0102] In this embodiment, the diffusion model may be trained using an energy function following a Boltzmann distribution as training information, either additionally or alternatively.
[0089]
[0103] In this embodiment, the diffusion model may be additionally or alternatively trained using molecular dynamics simulation data as training information.
[0090]
[0104] In this embodiment, additionally or alternatively, the diffusion model may be a score-based generative diffusion model configured to introduce random noise at predetermined time steps based on a Gaussian distribution in the forward process.
[0091]
[0105] In this embodiment, additionally or alternatively, the data representing the molecular system may include at least one of the following: the amino acid sequence of the molecular system, the two-dimensional structure of the molecular system, the chemical formula, the three-dimensional structure of the molecular system, the protein name, and / or the Simple Molecular Input Line Input System (SMILES) of the molecular system.
[0092]
[0106] Another embodiment provides a method for predicting the equilibrium distribution of a molecular system. This method may include receiving input data representing the molecular system, creating a graph representation of the molecular system including positional information of each atom in the molecular system, inputting the graph representation of the molecular system into a graph neural network, receiving multiple predicted conformations of the molecular system as output from the graph neural network, inputting each of the multiple predicted conformations into a diffusion model, and receiving the predicted equilibrium distribution of the molecular system.
[0093]
[0107] In this embodiment, the method may further include, either additionally or alternatively, varying the noise applied to the predicted conformation over a series of time steps to thereby generate an equilibrium distribution.
[0094]
[0108] In this embodiment, the method may further include, additionally or alternatively, generating statistically independent samples from an equilibrium distribution via an independent and identically distributed (IID) sampler module.
[0095]
[0109] In this embodiment, the method may additionally or alternatively further include determining the density function of the equilibrium distribution via a density module.
[0096]
[0110] In this embodiment, the method may further include, either additionally or alternatively, pre-training a diffusion model via a diffusion pre-training algorithm based on physical information.
[0097]
[0111] In this embodiment, the method may additionally or alternatively further include training a diffusion model to minimize a loss function.
[0098]
[0112] In this embodiment, the method may further include, either additionally or alternatively, training a diffusion model using an energy function following a Boltzmann distribution as training information.
[0099]
[0113] This method may further include training a diffusion model using molecular dynamics simulation data as training data.
[0100]
[0114] In this embodiment, additionally or alternatively, the diffusion model may be a score-based generative diffusion model configured to introduce random noise at predetermined time steps based on a Gaussian distribution in the forward process.
[0101]
[0115] Another embodiment provides a computing system for predicting the equilibrium distribution of a molecular system. The computing system may include a processor that uses a portion of the associated memory to execute instructions in order to implement an equilibrium distribution prediction model. The processor may be configured to, in the training phase, receive training phase input data representing the molecular system; create a graph representation of the molecular system including the positional information of each atom in the molecular system via a preprocessing algorithm; receive energy parameter values of the graph representation; input the graph representation of the molecular system and a training dataset including the energy parameter values of the graph representation into a graph neural network; receive multiple predicted conformations of the molecular system as output from the graph neural network; input each of the multiple predicted conformations into a diffusion model; train the diffusion model using an energy function following a Boltzmann distribution as training data; train the diffusion model using molecular dynamics simulation data as training data; and receive the predicted equilibrium distribution of the molecular system from the diffusion model.
[0102]
[0116] The configurations and / or methods described herein are, in their essence, illustrative, and it should be understood that these particular embodiments or examples should not be considered restrictive, as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Accordingly, the various actions shown and / or described may be performed in parallel or omitted in one shown and / or described sequence compared to another. Similarly, the order of the processes described above may be altered.
[0103]
[0117] The subject matter of this disclosure includes all novel and non-obvious combinations and partial combinations of the various processes, systems and configurations and other features, functions, actions and / or characteristics disclosed herein and all their equivalents.
Claims
1. A computing system for predicting the equilibrium distribution of molecular systems, To implement the equilibrium distribution prediction model, it includes a processor that uses a portion of the associated memory to execute instructions. The aforementioned processor, in the inference stage, Receiving input data representing a molecular system, To create a graph representation of the molecular system, including the positional information of each atom within the molecular system, Inputting the graph representation of the molecular system into a graph neural network, The output from the graph neural network is to receive multiple predicted conformations of the molecular system, Inputting each of the multiple predicted conformations into the diffusion model, To predict the equilibrium distribution of the aforementioned molecular system, Outputting the predicted equilibrium distribution A computing system configured to perform the following actions.
2. The computing system according to claim 1, wherein the diffusion model varies the noise applied to the predicted conformation over a series of time steps to generate the equilibrium distribution.
3. The computing system according to claim 1, wherein the equilibrium distribution prediction model includes an independent identically distributed (IID) sampler module that generates statistically independent samples from the equilibrium distribution.
4. The computing system according to claim 1, wherein the equilibrium distribution prediction model includes a density module for determining the density function of the equilibrium distribution.
5. The computing system according to claim 1, wherein the diffusion model is pre-trained via a diffusion pre-training algorithm based on physical information.
6. The computing system according to claim 1, wherein the diffusion model is trained to minimize a loss function.
7. The computing system according to claim 1, wherein the diffusion model is trained using an energy function following a Boltzmann distribution as training information.
8. The computing system according to claim 1, wherein the diffusion model is trained using molecular dynamics simulation data as training information.
9. The computing system according to claim 1, wherein the diffusion model is a score-based generative diffusion model configured to introduce random noise at predetermined time steps based on a Gaussian distribution in a forward process.
10. The computing system according to claim 1, wherein the data representing the molecular system includes at least one of the amino acid sequence of the molecular system, the two-dimensional structure of the molecular system, the chemical formula, the three-dimensional structure of the molecular system, the protein name, and / or a simplified molecular input line input system (SMILES) for the molecular system.
11. A method for predicting the equilibrium distribution of a molecular system, Receiving input data representing a molecular system, To create a graph representation of the molecular system, including the positional information of each atom within the molecular system, Inputting the graph representation of the molecular system into a graph neural network, The output from the graph neural network is to receive multiple predicted conformations of the molecular system, Inputting each of the multiple predicted conformations into the diffusion model, To receive the predicted equilibrium distribution of the aforementioned molecular system and A method that includes this.
12. The method according to claim 11, further comprising varying the noise applied to the predicted conformation over a series of time steps to generate the equilibrium distribution.
13. The method according to claim 11, further comprising generating statistically independent samples from the equilibrium distribution via an independent and identically distributed (IID) sampler module.
14. The method according to claim 11, further comprising determining the density function of the equilibrium distribution via a density module.
15. The method according to claim 11, further comprising pre-training the diffusion model via a diffusion pre-training algorithm based on physical information.
16. The method according to claim 11, further comprising training the diffusion model to minimize the loss function.
17. The method according to claim 11, further comprising training the diffusion model using an energy function following a Boltzmann distribution as training information.
18. The method according to claim 11, further comprising training the diffusion model using molecular dynamics simulation data as training data.
19. The method according to claim 11, wherein the diffusion model is a score-based generative diffusion model configured to introduce random noise at predetermined time steps based on a Gaussian distribution in a forward process.
20. A computing system for predicting the equilibrium distribution of molecular systems, A processor that uses a portion of the relevant memory to execute instructions in order to implement an equilibrium distribution prediction model. The processor includes, during the training phase, Receiving training stage input data representing molecular systems, A graph representation of the molecular system is created, including the positional information of each atom within the molecular system, via a preprocessing algorithm. Receiving the energy parameter values of the aforementioned graph representation, Inputting the graph representation of the molecular system and a training dataset including the energy parameter values of the graph representation into a graph neural network, The output from the graph neural network is to receive multiple predicted conformations of the molecular system, Inputting each of the multiple predicted conformations into the diffusion model, Training the diffusion model using an energy function following a Boltzmann distribution as training information, Training the diffusion model using molecular dynamics simulation data as training information, The predicted equilibrium distribution of the molecular system is received from the aforementioned diffusion model. A computing system configured to perform the following actions.