Methods and systems for generating a boltzmann distribution of functionally important metastable states of macromolecules
The method addresses the limitations of current approaches by using molecular dynamics simulations and gated attention units within a Boltzmann generator to efficiently generate a Boltzmann distribution of metastable states for macromolecules, achieving low-energy conformations and diverse metastable states.
Patent Information
- Application Number
- PCT/US2024/057103
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-12
AI Technical Summary
Current methods for generating a Boltzmann distribution of functionally important metastable states of macromolecules, such as proteins, are limited by their inability to penetrate high energy barriers and are computationally expensive, especially for larger systems.
A novel method and system that utilize molecular dynamics simulations, internal coordinate representations, and gated attention units within a Boltzmann generator to determine long-range interactions and generate a Boltzmann distribution of metastable states, constrained by a Frechet Inception Distance matrix to ensure native conformational ensemble.
The method efficiently generates Boltzmann distributions and important experimental structures for proteins, capturing low-energy conformations and diverse metastable states that are not accessible by conventional methods, while maintaining computational feasibility.
Smart Images

Figure US2024057103_12062025_PF_FP_ABST
Abstract
Description
Docket No: 361213.00201 METHODS AND SYSTEMS FOR GENERATING A BOLTZMANN DISTRIBUTION OF FUNCTIONALLY IMPORTANT METASTABLE STATES OF MACROMOLECULES CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 607,622, filed December 8, 2023. The foregoing application is incorporated by reference herein in its entirety. FIELD This invention relates generally to methods and systems for generating a Boltzmann distribution of functionally important metastable states of macromolecules. BACKGROUND The structural ensemble of a protein determines its functions. The probabilities of the ground and metastable states of a protein at equilibrium for a given temperature determine the interactions of the protein with other proteins, effectors, and drugs, which are keys for pharmaceutical development. However, enumeration of the equilibrium conformations and their probabilities are infeasible. Since complete knowledge is inaccessible, a sampling approach must be adopted. Conventional approaches toward sampling the equilibrium ensemble rely on Markov- chain Monte Carlo or molecular dynamics (MD). These approaches explore the local energy landscape adjacent to a starting point; however, they are limited by their inability to penetrate high energy barriers. In addition, MD simulations are expensive and scale poorly with system size. This results in incomplete exploration of the equilibrium conformational ensemble. Noe et al. proposed a normalizing flow model that is trained on the energy function of a many-body system, termed Boltzmann generators (BGs) (F. Noe, et al. Science, 365(6457), 2019). The model learns an invertible transformation from a system’s configurations to a latent space representation, in which the low-energy configurations of different states can be easily sampled. However, existing BGs have often struggled with even moderate-sized proteins, due to the complexity of conformation dynamics and scarcity of available data. Most works have focused on small systems such as alanine dipeptide with only 22 atoms. The limited scope of flow model BG applications is in part due to the high computational expense of their training process. TheirDocket No: 361213.00201 invertibility requirement limits expressivity when modeling targets whose supports have complicated topologies, necessitating the use of many transformation layers. Another hurdle in scaling BGs is that proteins often involve long-range interactions; atoms far apart in sequence can interact with each other. Therefore, there exists a need for methods and systems capable of generating a Boltzmann distribution of functionally important metastable states of macromolecules. SUMMARY This disclosure addresses the need mentioned above in a number of aspects. In one aspect, this disclosure provides a method of generating a Boltzmann distribution of functionally important metastable states of a macromolecule. In some embodiments, the method comprises: (a) generating representative conformations of a macromolecular structure of the macromolecule by molecular dynamics simulations; (b) generating an internal coordinate representation with fixed bond length and side chain angles for each of the representative conformations of the macromolecular structure; (c) dividing internal coordinate representations into a backbone channel and a side chain channel; (d) inputting the internal coordinate representations of the backbone channel and the internal coordinate representations of the side-chain channel into gated attention units of a Boltzmann generator, wherein the Boltzmann generator determines long range interactions within the macromolecule based on a combination of rotary positional embeddings with global absolute positional embeddings; and (e) determining by the Boltzmann generator a Boltzmann distribution of functionally important metastable states of the macromolecule using a Frechet Inception Distance matrix as a loss function to constrain global backbone structures to a space of native conformational ensemble. In some embodiments, the method comprises training the Boltzmann generator based on negative log-likelihood. In some embodiments, the method comprises training the Boltzmann generator based on a combination of negative log-likelihood and 2-Wasserstein loss. In some embodiments, the method comprises training the Boltzmann generator based on a combination of negative log-likelihood, 2-Wasserstein loss, and Kullback-Leibler divergence. In some embodiments, the macromolecule is a protein. In some embodiments, the internal coordinate representations are translation and rotation invariant. In some embodiments, theDocket No: 361213.00201 Boltzmann distribution of functionally important metastable states comprises low-energy conformations of the macromolecule. In some embodiments, the gated attention units comprise gated attention rational quadratic spline coupling blocks. In some embodiments, the method comprises implementing relative positional embeddings on a global level. In some embodiments, the method further comprises concatenating backbone latent embeddings with side chain features and passing the backbone latent embeddings through the gated attention units. In some embodiments, the Boltzmann generator comprises a machine-learning model. In some embodiments, the Boltzmann generator comprises a deep neural network model that has been trained to learn the transformation between a mathematic Gaussian distribution and a physical Boltzmann distribution of macromolecular structures. In some embodiments, the method comprises generating by the Boltzmann generator metastable conformational states of the macromolecule by drawing samples from a von Mises distribution. In another aspect, this disclosure provides a system of generating a Boltzmann distribution of functionally important metastable states of a macromolecule. In some embodiments, the system comprises one or more processors configured to: (i) generate representative conformations of a macromolecular structure of the macromolecule by molecular dynamics simulations; (ii) generate an internal coordinate representation with fixed bond length and side chain angles for each of the representative conformations of the macromolecular structure; (iii) divide internal coordinate representations into a backbone channel and a side chain channel; (iv) input the internal coordinate representations of the backbone channel and the internal coordinate representations of the side- chain channel into gated attention units of a Boltzmann generator, wherein the Boltzmann generator determines long range interactions within the macromolecule based on a combination of rotary positional embeddings with global absolute positional embeddings; and (v)determine by the Boltzmann generator a Boltzmann distribution of functionally important metastable states of the macromolecule using a Frechet Inception Distance matrix as a loss function to constrain global backbone structures to a space of native conformational ensemble. In some embodiments, the one or more processors are configured to train the Boltzmann generator based on negative log-likelihood. In some embodiments, the one or more processors areDocket No: 361213.00201 configured to train the Boltzmann generator based on a combination of negative log-likelihood and 2-Wasserstein loss. In some embodiments, the one or more processors are configured to train the Boltzmann generator based on a combination of negative log-likelihood, 2-Wasserstein loss, and Kullback-Leibler divergence. In some embodiments, the macromolecule is a protein. In some embodiments, the internal coordinate representations are translation and rotation invariant. In some embodiments, the Boltzmann distribution of functionally important metastable states comprises low-energy conformations of the macromolecule. In some embodiments, the gated attention units comprise gated attention rational quadratic spline coupling blocks. In some embodiments, the one or more processors are configured to implement relative positional embeddings on a global level. In some embodiments, the one or more processors are configured to concatenate backbone latent embeddings with side chain features and passing the backbone latent embeddings through the gated attention units. In some embodiments, the Boltzmann generator comprises a machine-learning model. In some embodiments, the Boltzmann generator comprises a deep neural network model that has been trained to learn the transformation between a mathematic Gaussian distribution and a physical Boltzmann distribution of macromolecular structures. In some embodiments, the one or more processors are configured to generate by the Boltzmann generator metastable conformational states of the macromolecule by drawing samples from a von Mises distribution. The foregoing summary is not intended to define every aspect of the disclosure, and additional aspects are described in other sections, such as the following detailed description. The entire document is intended to be related as a unified disclosure, and it should be understood that all combinations of features described herein are contemplated, even if the combinations of features are not found together in the same sentence, or paragraph, or section of this document. Other features and advantages of the invention will become apparent from the following detailed description. It should be understood, however, that the detailed description and the specific examples, while indicating specific embodiments of the disclosure, are given by way of illustration only, because various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description.Docket No: 361213.00201 BRIEF DESCRIPTION OF THE DRAWINGS Fig.1 shows an example flowchart of a Boltzmann generator for proteins. Figs.2a, 2b, and 2c show an example scheme for generating a Boltzmann distribution of functionally important metastable states of a protein. Fig.2a shows a split flow architecture. Fig. 2b shows each transformation block consisting of a gated attention rational quadratic spline (RQS) coupling layer. Fig. 2c shows example structures of protein G from the flow qθ(left) and from molecular dynamics simulation p (right). Sample distance matrices D�x^^^^θ� and D�x^^^^�. Fig.3 shows a two-residue chain. Backbone atoms areAn example of a bond length d, a bond angle θ, and a dihedral / torsion angle Ø. Figs.4a, 4b, 4c, and 4d show sample conformations generated by a Boltzmann generator via different training strategies. Fig. 4a shows Root mean square fluctuation (RMSF) computed for each residue (Ca atoms) in HP35 and protein G. Matching the training dataset’s plot is desirable, Fig. 4b shows examples of HP35 from ground truth training data, generated samples from the model, and generated samples from the baseline model. Fig. 4c shows an example of two metastable states from protein G training data. Fig.5d shows low-energy conformations of protein G generated by the model superimposed on each other. Some examples of pathological structures generated after training with different training paradigms are also shown: negative log-likelihood (NLL) (maximum likelihood), both NLL and Kullback-Leibler (KL) divergence, and NLL and the 2-Wasserstein loss. Atom clashes are highlighted with red circles. Figs. 5a, 5b, 5c, and 5d show that Boltzmann generators can generate novel sample conformations. Fig.5a shows Protein G two-dimensional (2D) Uniform Manifold Approximation and Projection (UMAP) embeddings for the training data, test data, and 2 x 105generated samples. Fig.5b shows a representative example of generated structures by the Boltzmann generator model which was not found in training data and the closest structure in the training dataset by Root Mean Square Deviation (RMSD). Fig. 5c shows Protein G energy distribution of training dataset and samples generated by the model. The second energy peak of the sampled conformations covers the novel structure shown in Fig.5b. Fig.5d shows an overlay of high-resolution, lowest-energy all- atom structures of protein G generated by the Boltzmann generator model. This demonstrates that the model is capable of sampling low-energy conformations at atomic resolution.Docket No: 361213.00201 Figure 6 shows an example computing system for implementing the disclosed methods. DETAILED DESCRIPTION A Boltzmann distribution of a protein provides a roadmap to all functional states of the protein. Normalizing flows are a promising tool for modeling this distribution, but current methods are intractable for typical pharmacological targets. They are computationally intractable due to the size of the system, heterogeneity of intra-molecular potential energy, and long-range interactions. To address these issues, this disclosure presents a novel flow architecture that utilizes split channels and gated attention to efficiently learn the conformational distribution of proteins defined by internal coordinates. It was demonstrated that while standard architectures and training strategies (such as maximum likelihood alone) fail, the disclosed architecture and multi-stage training strategy can model the conformational distributions of proteins. In one aspect, this disclosure provides a method of generating a Boltzmann distribution of functionally important metastable states of a macromolecule. In some embodiments, the method comprises: (a) generating representative conformations of a macromolecular structure of the macromolecule by molecular dynamics simulations; (b) generating an internal coordinate representation with fixed bond length and side chain angles for each of the representative conformations of the macromolecular structure; (c) dividing internal coordinate representations into a backbone channel and a side chain channel; (d) inputting the internal coordinate representations of the backbone channel and the internal coordinate representations of the side- chain channel into gated attention units of a Boltzmann generator, wherein the Boltzmann generator determines long range interactions within the macromolecule based on a combination of rotary positional embeddings with global absolute positional embeddings; and (e) determining by the Boltzmann generator a Boltzmann distribution of functionally important metastable states of the macromolecule using a Frechet Inception Distance matrix as a loss function to constrain global backbone structures to a space of native conformational ensemble. In another aspect, this disclosure provides a system of generating a Boltzmann distribution of functionally important metastable states of a macromolecule. In some embodiments, the system comprises one or more processors configured to: (i) generate representative conformations of a macromolecular structure of the macromolecule by molecular dynamics simulations; (ii) generate an internal coordinate representation with fixed bond length and side chain angles for each of theDocket No: 361213.00201 representative conformations of the macromolecular structure; (iii) divide internal coordinate representations into a backbone channel and a side chain channel; (iv) input the internal coordinate representations of the backbone channel and the internal coordinate representations of the side- chain channel into gated attention units of a Boltzmann generator, wherein the Boltzmann generator determines long range interactions within the macromolecule based on a combination of rotary positional embeddings with global absolute positional embeddings; and (v)determine by the Boltzmann generator a Boltzmann distribution of functionally important metastable states of the macromolecule using a Frechet Inception Distance matrix as a loss function to constrain global backbone structures to a space of native conformational ensemble. The Boltzmann distribution is a probability distribution that describes the likelihood of a system being in a specific state. It is also known as the Gibbs distribution or probability distribution. A probability distribution is the mathematical function that gives the probabilities of occurrence of different possible outcomes for an experiment. As used herein, the term “functionally important metastable states” refers to metastable states (or conformations) that are important for one or more functions of macromolecules, such as protein-ligand interactions, protein-protein interactions, enzymeatic activities, protein folding, transportation, storage, structural functions, etc. The Fréchet inception distance (FID) is a metric that is often used for generative adversarial networks (GANs). The FID calculates the distance between the feature vectors of real images and the feature vectors of fake images. A lower FID score indicates that the generated images are more similar to real images in terms of their visual quality, diversity, overall appearance, and statistical properties. In some embodiments, the method comprises training the Boltzmann generator based on negative log-likelihood. In some embodiments, the method comprises training the Boltzmann generator based on a combination of negative log-likelihood and 2-Wasserstein loss. In some embodiments, the method comprises training the Boltzmann generator based on a combination of negative log-likelihood, 2-Wasserstein loss, and Kullback-Leibler divergence. Negative log-likelihood (NLL) is a cost function used in machine learning models to measure their performance. It is a loss function used in multi-class classification. The lower the NLL, the better the performance. NLL is calculated as: −log(y). y: is a prediction correspondingDocket No: 361213.00201 to the true label. The loss for a mini-batch is calculated by taking the mean or sum of all items in the batch. NLL can be interpreted as the “unhappiness” of the network with respect to its parameters. The gradient of the negative log likelihood function is ∇θ(−logL)=−XT(y−eXθ). X∈RM×N is the data matrix with M the number of samples and N the number of features in each input vector xi. y∈IM×1 is the scores vector θ∈RN×1 is the parameters vector. The Wasserstein loss function is a metric that measures the distance between two measures on a target label space. It is also known as earth mover’s distance. The Wasserstein loss function increases the gap between the scores for real and produced imagery. It is defined as: Critic loss, average critic score on real images, and average critic score on fake images. Kullback-Leibler divergence (KL-divergence) is a statistical measurement that compares the difference between two probability distributions. It is also known as relative entropy. KL- divergence is a concept of information theory to quantify the difference between one probability distribution from a reference probability distribution. The KL-divergence between two distributions p(X) and q(X) of the same variable X is: DKL(p(X)||q(X)) = ∫xp(x)logp(x)q(x)dx. The KL-divergence score can be written as: KL(P||Q). In some embodiments, the macromolecule includes a polypeptide, such as a protein. As used herein, the term “protein” as used throughout this specification generally encompasses macromolecules comprising one or more polypeptide chains, i.e., polymeric chains of amino acid residues linked by peptide bonds. The term may encompass naturally, recombinantly, semi- synthetically or synthetically produced proteins. The term also encompasses proteins that carry one or more co- or post-expression-type modifications of the polypeptide chain(s), such as, without limitation, glycosylation, acetylation, phosphorylation, sulfonation, methylation, ubiquitination, signal peptide removal, N-terminal Met removal, conversion of pro-enzymes or pre-hormones into active forms, etc. The term further also includes protein variants or mutants which carry amino acid sequence variations vis-à-vis a corresponding native protein, such as, e.g., amino acid deletions, additions and / or substitutions. The term contemplates both full-length proteins and protein parts or fragments, e.g., naturally occurring protein parts that ensue from processing of such full-length proteins.Docket No: 361213.00201 In some embodiments, the internal coordinate representations are translation and rotation invariant. In some embodiments, the Boltzmann distribution of functionally important metastable states comprises low-energy conformations of the macromolecule. In some embodiments, the gated attention units comprise gated attention rational quadratic spline coupling blocks. In some embodiments, the method comprises implementing relative positional embeddings on a global level. In some embodiments, the method further comprises concatenating backbone latent embeddings with side chain features and passing the backbone latent embeddings through the gated attention units. In some embodiments, the Boltzmann generator comprises a machine-learning model. In some embodiments, the Boltzmann generator comprises a deep neural network model that has been trained to learn the transformation between a mathematic Gaussian distribution and a physical Boltzmann distribution of macromolecular structures. In some embodiments, the method comprises generating by the Boltzmann generator metastable conformational states of the macromolecule by drawing samples from a von Mises distribution. The von Mises distribution is a continuous probability distribution, also known as the circular normal or Tikhonov distribution. The von Mises distribution is a maximum entropy distribution for circular data. It is the equivalent of the normal distribution for data defined with directional coordinates. The distribution's coordinates exist on a circular plane. As used herein, a “machine learning model,” a “model,” or a “classifier” refers to a set of algorithmic routines and parameters that can predict an output(s) for a process input based on a set of input features, with or without being explicitly programmed. A structure of the software routines (e.g., number of subroutines and relation between them) and / or the values of the parameters can be determined in a training process, which can use actual results of the process that is being modeled. Such systems or models are understood to be necessarily rooted in computer technology, and in fact, cannot be implemented or even exist in the absence of computing technology. While machine learning systems utilize various types of statistical analyses, machine learning systems are distinguished from statistical analyses by virtue of the ability to learn without explicit programming and being rooted in computer technology. A neural network or an artificial neural network is one set of algorithms used in machine learning for modeling the data using graphs ofDocket No: 361213.00201 neurons. Any network structure may be used. Any number of layers, nodes within layers, types of nodes (activations), types of layers, interconnections, learnable parameters, and / or other network architectures may be used. Machine training uses the defined architecture, training data, and optimization to learn values of the learnable parameters of the architecture based on the samples and ground truth of training data. A typical machine learning pipeline may include building a machine learning model from a sample dataset (referred to as a “training set”), evaluating the model against one or more additional sample datasets (referred to as a “validation set” and / or a “test set”) to decide whether to keep the model and to benchmark how good the model is, and using the model in “production” to make predictions or decisions against live input data captured by an application service. For training the model to be applied as a machine-learned model, training data is acquired and stored in a database or memory. The training data is acquired by aggregation, mining, loading from a publicly or privately formed collection, transfer, and / or access. Ten, hundreds, or thousands of samples of training data are acquired. The samples are from scans of different patients and / or phantoms. Simulation may be used to form the training data. The training data includes the desired output (ground truth), such as segmentation, and the input, such as protocol data and imaging data. In some embodiments, the training set will be used to create a single classifier using any now or hereafter-known methods. In other embodiments, a plurality of training sets will be created to generate a plurality of corresponding classifiers. Each of the plurality of classifiers can be generated based on the same or different learning algorithm that utilizes the same or different features in the corresponding one of the pluralities of training sets. Once trained, the machine-learned or trained classifier is stored for later application. The training determines the values of the learnable parameters of the network. The network architecture, values of non-learnable parameters, and values of the learnable parameters are stored as the machine-learned network. Once stored, the machine-learned network may be fixed. The same machine-learned network may be applied to different patients, different scanners, and / or with different imaging protocols for the scanning. The machine-learned network may be updated. As additional training data is acquired, such as through application of the network for patients and corrections by experts to that output, the additional training data may be used to re-train or update the training.Docket No: 361213.00201 For the machine learning model, input data structures of subreads can be used for the training. The input data structure may correspond to a window of in a sequence. Each training sample can include one of the first plurality of first data structures and a pocket prediction score likelihood of a residue being a ligand binding residue. The training is performed by optimizing parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels and optionally the second labels when the first plurality of first data structures and optionally the second plurality of second data structures are input to the model. An output of the model specifies whether a residue in sequence is likely a ligand binding residue. In some embodiments, the output of the model may include a probability of being in each of a plurality of states. The state with the highest probability can be taken as the state. In some embodiments, the machine learning model may further include a supervised learning model. Supervised learning models may include different approaches and algorithms including analytical learning, artificial neural network, backpropagation, boosting (meta- algorithm), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, group method of data handling, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), multilinear subspace learning, naive Bayes classifier, maximum entropy classifier, conditional random field, Nearest Neighbor Algorithm, probably approximately correct learning (PAC) learning, ripple down rules, a knowledge acquisition methodology, symbolic machine learning algorithms, subsymbolic machine learning algorithms, support vector machines, Minimum Complexity Machines (MCM) , random forests, ensembles of classifiers, ordinal classification, data pre-processing, handling imbalanced datasets, statistical relational learning, or Proaftn, a multicriteria classification algorithm, linear regression, logistic regression, deep recurrent neural network (e.g., long short term memory, LSTM), Bayes classifier, hidden Markov model (HMM) , linear discriminant analysis (LDA) , k-means clustering, density- based spatial clustering of applications with noise (DBSCAN) , random forest algorithm, support vector machine (SVM), or any model described herein. In some embodiments, the classifier may include a supervised or unsupervised Machine Learning or Deep Learning algorithm, Logistic Regression, Naive Bayes, Support Vector Machine, Decision Tree, Random Forest, Gradient Boosting, Regularizing Gradient Boosting, K-Nearest Neighbors, a continuous regression approach, Ridge Regression, Kernel Ridge Regression,Docket No: 361213.00201 Support Vector Regression, deep learning approach, Neural Networks, Convolutional Neural Network (CNNs), Recurrent Neural Networks (RNNs), Gated Recurrent Units (GRUs), Long Short Term Memory Networks (LSTMs), Generative Models, Generative Adversarial Networks (GANs), Deep Belief Networks (DBNs), Feedforward Neural Networks, Autoencoders, Variational Autoencoders, Normalizing Flow Models, Deniosing Diffusion Probabilistic Models (DDPMs), Score Based Generative Models (SGMs), Radial Basis Function Networks (RBFNs), Multilayer Perceptrons (MLPs), Stochastic Neural Networks, or any combination thereof. In some embodiments, the model may include a convolutional neural network (CNN). The CNN may include a set of convolutional filters configured to filter the first plurality of data structures and, optionally, the second plurality of data structures. The filter may be any filter described herein. The number of filters for each layer may be from 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, 90 to 100, 100 to 150, 150 to 200, or more. The kernel size for the filters can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, from 15 to 20, from 20 to 30, from 30 to 40, or more. The CNN may include an input layer configured to receive the filtered first plurality of data structures and, optionally, the filtered second plurality of data structures. The CNN may also include a plurality of hidden layers, including a plurality of nodes. The first layer of the plurality of hidden layers is coupled to the input layer. The CNN may further include an output layer coupled to a last layer of the plurality of hidden layers and configured to output an output data structure. The output data structure may include the properties. The terms or acronyms like “convolutional neural network,” “CNN,” “neural network,” “NN,” “deep neural network,” “DNN,” “recurrent neural network,” “RNN,” and / or the like may be interchangeably referenced throughout this document. Figure 6 is a functional diagram illustrating a programmed computer system in accordance with some embodiments. As will be apparent, other computer system architectures and configurations can be used to perform the described methods. Computer system 600, which includes various subsystems as described below, includes at least one microprocessor subsystem (also referred to as a processor or a central processing unit (CPU) 606). For example, processor 606 can be implemented by a single-chip processor or by multiple processors. In some embodiments, processor 606 is a general-purpose digital processor that controls the operation of the computer system 600. In some embodiments, processor 606 also includes one or moreDocket No: 361213.00201 coprocessors or special purpose processors (e.g., a graphics processor, a network processor, etc.). Using instructions retrieved from memory 607, processor 606 controls the reception and manipulation of input data received on an input device (e.g., image processing device 603, I / O device interface 602), and the output and display of data on output devices (e.g., display 601). Processor 606 is coupled bi-directionally with memory 607, which can include, for example, one or more random access memories (RAM) and / or one or more read-only memories (ROM). As is well known in the art, memory 607 can be used as a general storage area, a temporary (e.g., scratchpad) memory, and / or a cache memory. Memory 607 can also be used to store input data and processed data, as well as to store programming instructions and data, in the form of data objects and text objects, in addition to other data and instructions for processes operating on processor 606. Also, as is well known in the art, memory 607 typically includes basic operating instructions, program code, data, and objects used by the processor 606 to perform its functions (e.g., programmed instructions). For example, memory 607 can include any suitable computer-readable storage media described below, depending on whether, for example, data access needs to be bi-directional or uni-directional. For example, processor 606 can also directly and very rapidly retrieve and store frequently needed data in a cache memory included in memory 607. A removable mass storage device 608 provides additional data storage capacity for the computer system 600 and is optionally coupled either bi-directionally (read / write) or uni- directionally (read-only) to processor 606. A fixed mass storage 609 can also, for example, provide additional data storage capacity. For example, storage devices 608 and / or 609 can include computer-readable media such as magnetic tape, flash memory, PC-CARDS, portable mass storage devices such as hard drives (e.g., magnetic, optical, or solid-state drives), holographic storage devices, and other storage devices. Mass storages 608 and / or 609 generally store additional programming instructions, data, and the like that typically are not in active use by the processor 606. It will be appreciated that the information retained within mass storages 608 and 609 can be incorporated, if needed, in a standard fashion as part of memory 607 (e.g., RAM) as virtual memory. In addition to providing processor 606 access to storage subsystems, bus 610 can be used to provide access to other subsystems and devices as well. As shown, these can include a displayDocket No: 361213.00201 601, a network interface 604, an input / output (I / O) device interface 602, an image processing device 603, as well as other subsystems and devices. For example, image processing device 603 can include a camera, a scanner, etc.; I / O device interface 602 can include a device interface for interacting with a touchscreen (e.g., a capacitive touch sensitive screen that supports gesture interpretation), a microphone, a sound card, a speaker, a keyboard, a pointing device (e.g., a mouse, a stylus, a human finger), a global positioning system (GPS) receiver, a differential global positioning system (DGPS) receiver, an accelerometer, and / or any other appropriate device interface for interacting with system 600. Multiple I / O device interfaces can be used in conjunction with computer system 600. The I / O device interface can include general and customized interfaces that allow the processor 606 to send and, more typically, receive data from other devices such as keyboards, pointing devices, microphones, touchscreens, transducer card readers, tape readers, voice or handwriting recognizers, biometrics readers, cameras, portable mass storage devices, and other computers. The network interface 604 allows processor 606 to be coupled to another computer, computer network, or telecommunications network using a network connection as shown. For example, through the network interface 604, the processor 606 can receive information (e.g., data objects or program instructions) from another network, or output information to another network in the course of performing method / process steps. Information, often represented as a sequence of instructions to be executed on a processor, can be received from and outputted to another network. An interface card or similar device and appropriate software implemented by (e.g., executed / performed on) processor 606 can be used to connect the computer system 600 to an external network and transfer data according to standard protocols. For example, various process embodiments disclosed herein can be executed on processor 606 or can be performed across a network such as the Internet, intranet networks, or local area networks, in conjunction with a remote processor that shares a portion of the processing. Additional mass storage devices (not shown) can also be connected to processor 606 through network interface 604. In addition, various embodiments disclosed herein further relate to computer storage products with a computer-readable medium that includes program code for performing various computer-implemented operations. The computer-readable medium includes any data storage device that can store data that can thereafter be read by a computer system. Examples of computer- readable media include, but are not limited to: magnetic media such as disks and magnetic tape;Docket No: 361213.00201 optical media such as CD-ROM disks; magneto-optical media such as optical disks; and specially configured hardware devices such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), and ROM and RAM devices. Examples of program code include both machine code as produced, for example, by a compiler, or files containing higher level code (e.g., script) that can be executed using an interpreter. The computer system as shown in Figure 6 is an example of a computer system suitable for use with the various embodiments disclosed herein. Other computer systems suitable for such use can include additional or fewer subsystems. In some computer systems, subsystems can share components (e.g., for touchscreen-based devices such as smartphones, tablets, etc., I / O device interface 602 and display 601 share the touch-sensitive screen component, which both detects user inputs and displays outputs to the user). In addition, bus 610 is illustrative of any interconnection scheme serving to link the subsystems. Other computer architectures having different configurations of subsystems can also be utilized. Additional Definitions To aid in understanding the detailed description of the compositions and methods according to the disclosure, a few express definitions are provided to facilitate an unambiguous disclosure of the various aspects of the disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless specified otherwise. Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. In some embodiments, the flowchart andDocket No: 361213.00201 block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, a segment, or a portion of instructions, which may include one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. These computer readable program instructions may be provided to a processor of a general- purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks. The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks. It will be understood that, although the terms “first,” “second,” etc., may be used herein to describe various elements, components, regions, layers and / or sections, these elements,Docket No: 361213.00201 components, regions, layers and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Thus, a first element, component, region, layer or section discussed below could be termed a second element, component, region, layer or section without departing from the teachings of example embodiments. Unless specifically stated otherwise, as apparent from the above discussion, it is appreciated that throughout the description, discussions utilizing terms such as “processing,” “performing,” “receiving,” “computing,” “calculating,” “determining,” “identifying,” “displaying,” “providing,” “merging,” “combining,” “running,” “transmitting,” “obtaining,” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (or electronic) quantities within the computer system memories or registers or other such information storage, transmission or display devices. As used herein, the term “if may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event],” depending on the context. An “electronic device” or a “computing device” refers to a device that includes a processor and memory. Each device may have its own processor and / or memory, or the processor and / or memory may be shared with other devices as in a virtual machine or container arrangement. The memory will contain or receive programming instructions that, when executed by the processor, cause the electronic device to perform one or more operations according to the programming instructions. The terms “memory,” “memory device,” “computer-readable medium,” “data store,” “data storage facility” and the like each refer to a non-transitory device on which computer-readable data, programming instructions or both are stored. Except where specifically stated otherwise, the terms “memory,” “memory device,” “computer-readable medium,” “data store,” “data storage facility” and the like are intended to include single device embodiments, embodiments in whichDocket No: 361213.00201 multiple memory devices together or collectively store a set of data or instructions, as well as individual sectors within such devices. The terms “processor” and “processing device” refer to a hardware component of an electronic device that is configured to execute programming instructions, such as a microprocessor or other logical circuit. A processor and memory may be elements of a microcontroller, custom configurable integrated circuit, programmable system-on-a-chip, or other electronic device that can be programmed to perform various functions. Except where specifically stated otherwise, the singular term “processor” or “processing device” is intended to include both single-processing device embodiments and embodiments in which multiple processing devices together or collectively perform a process. In this document, the terms “communication link” and “communication path” mean a wired or wireless path via which a first device sends communication signals to and / or receives communication signals from one or more other devices. Devices are “communicatively connected” if the devices are able to send and / or receive data via a communication link. “Electronic communication” refers to the transmission of data via one or more signals between two or more electronic devices, whether through a wired or wireless network, and whether directly or indirectly via one or more intermediary devices. As used herein, “in vitro” refers to events that occur in an artificial environment, e.g., in a test tube or reaction vessel, in cell culture, etc., rather than within a multi-cellular organism. As used herein, “in vivo” refers to events that occur within a multi-cellular organism, such as a non-human animal. It is noted here that, as used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise. As used herein, “plurality” means two or more. As used herein, a “set” of items may include one or more of such items. As used herein, “including,” “comprising,” “containing,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional subject matter unless otherwise noted.Docket No: 361213.00201 As used herein, the phrases “in one embodiment,” “in various embodiments,” “in some embodiments,” and the like do not necessarily refer to the same embodiment, but may unless the context dictates otherwise. As used herein, the terms “and / or” or “ / ” means any one of the items, any combination of the items, or all of the items with which this term is associated.As used herein, the term “substantially” does not exclude “completely,” e.g., a composition which is “substantially free” from Y may be completely free from Y. Where necessary, the word “substantially” may be omitted from the definition of the present disclosure. As used herein, the term “approximately” or “about,” as applied to one or more values of interest, refers to a value that is similar to a stated reference value. In some embodiments, the term “approximately” or “about” refers to a range of values that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the stated reference value unless otherwise stated or otherwise evident from the context (except where such number would exceed 100% of a possible value). Unless indicated otherwise herein, the term “about” is intended to include values, e.g., weight percents, proximate to the recited range that are equivalent in terms of the functionality of the individual ingredient, the composition, or the embodiment. As used herein, the term “each,” when used in reference to a collection of items, is intended to identify an individual item in the collection but does not necessarily refer to every item in the collection. Exceptions can occur if explicit disclosure or context clearly dictates otherwise. As disclosed herein, a number of ranges of values are provided. It is understood that each intervening value, to the tenth of the unit of the lower limit, unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the present disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither, or both limits are included in the smaller ranges is also encompassed within the present disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the present disclosure.Docket No: 361213.00201 The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the present disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the present disclosure. All methods described herein are performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. In regard to any of the methods provided, the steps of the method may occur simultaneously or sequentially. When the steps of the method occur sequentially, the steps may occur in any order, unless noted otherwise. In cases in which a method comprises a combination of steps, each and every combination or sub- combination of the steps is encompassed within the scope of the disclosure, unless otherwise noted herein. Each publication, patent application, patent, and other reference cited herein is incorporated by reference in its entirety to the extent that it is not inconsistent with the present disclosure. Publications disclosed herein are provided solely for their disclosure prior to the filing date of the present disclosure. Nothing herein is to be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed. It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. Examples In this example, a new Boltzmann Generator (“BG”) method for general proteins with one or more of the following characteristics: (a) A global internal coordinate representation with fixed bond-lengths and side-chain angles is used. From a global structure and energetics point-of-view, little information is lost by allowing side-chain bonds to only rotate. Such a representation not only reduces the number ofDocket No: 361213.00201 variables but also samples conformations more efficiently than Cartesian coordinates (F. Noe, et al. Science, 365(6457), 2019; A. H. Mahmoud, et al. J. Chem. Inf. Modeling, 62(7), 2022). (b) The global internal coordinate representation is initially split into a backbone channel and a side-chain channel. This allows the model to efficiently capture the distribution of backbone internal coordinates, which most controls the overall global conformation. (c) A new neural network (NN) architecture for learning the transformation parameters of the coupling layers of the flow model which makes use of gated attention units (GAUs) (Weizhe Hua, et al. Transformer quality in linear time, 2022) and a combination of rotary positional embeddings (Jianlin Su, et al. arXiv preprint arXiv:2104.09864, 2021) with global, absolute positional embeddings for learning long range interactions. (d) To handle global conformational changes, a new loss-function, similar in spirit to the Fréchet Inception Distance (FID) (Martin Heusel, et al., Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017), is introduced to constrain the global backbone structures to the space of native conformational ensemble. As demonstrated in this example, the disclosed methods can efficiently generate Boltzmann distributions and important experimental structures in two different protein systems. While the traditional maximum likelihood training for training flow models is insufficient for proteins, the disclosed multi-stage training strategy can generate samples with high Boltzmann probability. Normalizing Flows Normalizing flow models learn an invertible map f: Rd ↦ Rd to transform a randomvariable ^^^^ ∼ ^^^^^^^^ to the random variable x = f(z) with distribution:−1 ^^^^^^^^(^^^^) = ^^^^^^^^(^^^^)�det�^^^^^^^^(^^^^)(1)where Jf(z) = ∂f / ∂z is the. targetdistribution ^^^^(^^^^). To simplify notation, the flow distribution is referred to as ^^^^θ, where θ are the parameters of the flow. If samples from the target distribution are available, the flow can be trained via maximum likelihood. If the unnormalized target density p(x) is known, the flow can be trainedDocket No: 361213.00201by minimizing the KL divergence between ^^^^^^^^ and ^^^^ , i.e., KL(qθ, p) = qθ(x) log�qθ(x) / p(x)�dx. Distance Matrix A protein distance matrix is a square matrix of Euclidean distances from each atom to all other atoms. Practitioners typically use ^^^^αatoms or backbone atoms only. Protein distance matrices have many applications including structural alignment, protein classification, and finding homologous proteins. They have also been used as representations for protein structure prediction algorithms, including the first iteration of AlphaFold. 2-Wasserstein Distance The 2-Wasserstein Distance is a measure of the distance between two probabilitydistributions. Let ^^^^ = ^^^^(^^^^^^^^,^^^^^^^^) and ^^^^ = ^^^^�^^^^ ^^^^^^^^,^^^^^^^^� be two normal distributions in ^^^^ . Then,with respect to 2-Wasserstein distance between ^^^^ and ^^^^ isdefined as W(P, Q)2 = ||μ 21 / 22P − μQ||2 + trace�ΣP + ΣQ − 2�ΣPΣQ��. (2) In computer vision, the Fréchet Inceptionin Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017) computes the 2-Wasserstein distance and is often used as an evaluation metric to measure generated image quality. Scalable Boltzmann Generators Problem Setup As shown in Fig. 1, BGs are generative models that are trained to sample from theBoltzmann distribution for physical systems, i.e., p(x) ∝ e−u(x) / (kT), where u(x) is the potentialenergy of the conformation x, ^^^^ is the Boltzmannis the temperature. A protein conformation is defined as the arrangement in space of its constituent atoms (Fig.3), specifically, by the set of 3D Cartesian coordinates of its atoms. Enumeration of metastable conformations for a protein at equilibrium is quite challenging with standard sampling techniques. This problem was addressed with generative modeling. Throughout this example, ^^^^ is referred to as the ground truthDocket No: 361213.00201 conformation distribution and qθas the distribution parameterized by the normalizing flow model ^^^θ^ . Reduced Internal Coordinates Energetically favored conformational changes take place via rotations around single chemical bonds while bond vibrations and angle bending at physiologic temperature result in relatively small spatial perturbations (Vaidehi & Jain, 2015). The focus on near ground and meta-stable states therefore motivates the use of internal coordinates: ^^^^ − 1 bond lengths ^^^^, ^^^^ − 2bond angles θ, and ^^^^ − 3 torsion angles ϕ, where ^^^^ is the number of atoms of the system (Fig. 3)In addition, internal coordinate representation is translation and rotation invariant. A protein can be described as a branching graph structure with a set of backbone atoms and non-backbone atoms (colloquially referred to as side-chain atoms). Previous works have noted the difficulty in working with internal coordinate representations for the backbone atoms. This is due to the fact that protein conformations are sensitive to small changes in the backbone torsion angles. Noe et al. introduced a coordinate transformation whereby the side-chain atom coordinates are mapped to internal coordinates while the backbone atom coordinates are linearly transformed via principal component analysis (PCA) and the six coordinates with the lowest variance are eliminated (F. Noe, et al. Science, 365(6457), 2019). However, the mapping of vectors onto a fixed set of principal components is generally not invariant to translations and rotations. In addition, PCA suffers from distribution shift. A full internal coordinate system requires 3N-6 dimensions where N is the number of atoms. Bond lengths hardly vary in equilibrium distributions while torsion angles can vary immensely. Non-backbone bond angles are treated as constant, again replaced by their mean. Heterocycles in the sidechains of Trp, Phe, Tyr and His residues are treated as rigid bodies. The final representation is: x= [θ^^^^^^^^,ϕ^^^^^^^^,ϕ^^^^^^^^],where the subscripts bb and sc indicate backbone and sidechain, respectively. This dramatically reduces the input dimension and keeps the most important features for learning global conformation changes in the equilibrium distribution. Training and EvaluationDocket No: 361213.00201 BGs was trained with MD simulation data at equilibrium, i.e., the distribution of conformations is constant and not changing as with, for example, folding. The simulation was seeded with energetically stable native conformations. BG training aims to learn to sample from the Boltzmann distribution of protein conformations. The energy of samples generated by the model under the AMBER14 forcefield (D. A. Case, et al. Amber 14, 2014) was computed and their mean is reported. In addition, to evaluate how well the flow model generates the proper backbone distribution, a new measure is defined as follows: Definition (Distance Distortion). Let Abbdenote the indices of backbone atoms. DefineD(x) as the pairwise distance matrix for the backbone atoms of x . Define ^^^^ = {(^^^^, ^^^^)|^^^^, ^^^^ ∈^^^^^^^^^^^^ and ^^^^ < ^^^^}. The distance distortion is defined asΔD ≔ E1 x^^^^^^^^∼qθ, x^^^^∼p�|^^^^|∑(i,j)∈^^^^�D�x^^^^^^^^�ij − D�x^^^^�ij, (3)Split FlowNeural Spline Flows (NSF) with rational quadratic splines (Conor Durkan, et al. Neural spline flows. Advances in Neural Information Processing Systems, 2019) having 8 bins each was used. The conditioning was performed via coupling. Torsion angles ϕ can freely rotate and were therefore treated as periodic coordinates (Danilo Jimenez Rezende, et al. Proceedings of the 37th International Conference on Machine Learning, 119, 8083–8092, 2020). The full architectural details are highlighted in Fig.2. the input was first split into backboneand sidechain channels. The backbone inputs were then passed through Lbb = 48 gated attentionrational quadratic spline (GA-RQS) coupling blocks. As all the features are angles in[−π,π], the features were augmented with their mapping on the unit circle. In order to utilize an efficient attention mechanism, gated attention units (GAUs) were employed. In addition, relative positional embeddings (Peter Shaw, et al. arXiv preprint arXiv:1803.02155, 2018) on a global level were implemented so as to allow each coupling block to utilize the correct embeddings. The backbonelatent embeddings were then concatenated with the side chain features and passed through L = 10more GA-RQS coupling blocks. Multi-stage training strategy Normalizing flows are most often trained by maximum likelihood, i.e., minimizing the negative log likelihood (NLL)Docket No: 361213.00201 LNLL(θ) ≔ −Ex∼p[log qθ (x)], (4)or by minimizing the reverse “KLdivergence” or “KL loss”, as is : LKL(θ) ≔ Ex∼qθ�log�qθ(x) / p(x), (5)In the BG literature, minimizing the KL divergence is often referred to as “training-by- energy,” as the expression can be rewritten in terms of the energy of the system. The reverse KL divergence suffers from mode-seeking behavior, which is problematic when learning multimodal target distributions. While minimizing the NLL is mass-covering, samples generated from flows trained in this manner suffer from high variance. In addition, for larger systems, maximum likelihood training often results in high-energy generated samples. Noe et al. (F. Noe, S. Olsson, J. Kohler, and H. Wu. Science, 365(6457), 2019.) used a convex combination of the two loss terms, in the context of BGs, in order to both avoid mode- collapse and generate low-energy samples. However, for larger molecules, target evaluation is computationally expensive and dramatically slows iterative training with the reverse KL divergence objective. In addition, during the early stages of training, the KL divergence explodes and leads to unstable training. One way to circumvent these issues is to train with the NLL loss, followed by a combination of both loss terms. Unfortunately, for larger systems, the KL term tends to dominate, and training often get stuck at non-optimal local minima. To remedy these issues, a sequential training scheme (shown below) was used, whereby the transition from maximum likelihood training to the combination of maximum likelihood and reverse KL divergence minimization was smoothed. (1) As mentioned previously, training with maximum likelihood to convergence was first performed. (2) Afterward, training with a combination of the NLL and the 2-Wasserstein loss with respect to distance matrices of the backbone atoms was performed: L(θ) ≔ ||μ − μ ||2 + trace�Σ + Σ −1 / 2W^^^^^^^^ ^^^^ 2 ^^^^^^^^ ^^^^ 2�Σ^^^^^^^^Σ^^^^��, (6) whereDocket No: 361213.00201 μ^^^^ ≔ ^^^^^^^^∼^^^^[^^^^^^^^^^^^], Σ^^^^ ≔ Ex∼px^^^^^^^^ − μ^^^^x^^^^^^^^ − μ^^^^�⊤�, (7) (8) are(3) As a third stage of training, training with a combination of the NLL, the 2-Wasserstein loss, and the KL divergence was performed. (4) In the final stage of training, the 2-Wasserstein loss term was dropped, and training was carried out to minimize a combination of the NLL and the KL divergence. Results Protein Systems Alanine dipeptide (ADP) is a two residue (22-atoms) common benchmark system for evaluating BGs. The MD simulation datasets provided by Midgley et al. (Midgley et al., arXiv preprint arXiv:2208.01893 (2022)) were used for training and validation. HP35 (nle-nle), a 35-residue double-mutant of the villin headpiece subdomain, is a well- studied structure whose folding dynamics have been observed and documented. For training, the MD simulation dataset was used (Beauchamp et al. PNAS, 2012) and faulty trajectories and unfolded structures were removed as described by Ichinomiya (Ichinomiya, Scientific Reports volume 12, Article number: 2719 (2022)). Protein G is a 56-residue cell surface-associated protein from Streptococcus that binds to IgG with high affinity. In order to train the model, samples were generated by running a MD simulation. The crystal structure of protein G, 1PGA, was used as the seed structure. The conformational space of Protein G was first explored by simulations with ClustENMD (Burak T. Kaynak, et al. Bioinformatics, 37(21):3956–3958, 2021). From 3 rounds of ClustENMD iteration and approximately 300 generated conformations, 5 distinctly diverse structures were selected as the starting point for equilibrium MD simulation by Amber. On each starting structure, 5 replica simulations were carried out in parallel with different random seeds for 400 ns at 300 K. The total simulation time of all the replicas was accumulated to 1 microsecond. Thus, 106structures of protein G were saved over all the MD trajectories.Docket No: 361213.00201 As a baseline model for comparison, Neural Spline Flows (NSF) with 58 rational quadratic spline coupling layers was used. NSFs have been used in many recent works on BGs utilized the NSF model (with fewer coupling layers) in their experiments with alanine dipeptide, a two-residue system. In the experiments with ADP and HP35, 48 GA-RQS coupling layers for the backbone followed by 10 GA-RQS coupling layers for the full latent size were utilized. All models have a similar number of trainable parameters. A Gaussian base distribution was use for non-dihedral coordinates. For ADP and HP35, a uniform distribution was use for dihedral coordinates. For protein G, a von Mises base distribution was use for dihedral coordinates. It was found that using a von Mises base distribution improved training for the protein G system as compared to a uniform or truncated normal distribution. Table 1: Training BGs with different strategiesΔ^^^^, energy u(⋅), and mean NLL of 106generated structures were computed after training with different training strategies with ADP, protein G, and Villin HP35. Δ^^^^ is computed for batches of 103samples. Means and standard deviations are reported. Statistics for ^^^^(⋅) are reported for structures with energy below the median sample energy. Best results are bold-faced. For reference,the energy for training data structures is −317.5 ± 125.5 kcal / mol for protein G and −1215.5 ±Docket No: 361213.00201 222.2 kcal / mol for villin HP35. The results were compared against a Neural Spline Flows (NSF) baseline model. As shown in Table 1, the model has marginal improvements over the baseline model for ADP. This is not surprising as the system is extremely small, and maximum likelihood training sufficiently models the conformational landscape. For both proteins, the model closely captures the individual residue flexibility as analyzed by the root mean square fluctuations (RMSF) of the generated samples from the various training schemes in Fig.4(a). This is a common metric for MD analysis, where larger per-residue values indicate larger movements of that residue relative to the rest of the protein. Fig.4(a) indicates that the model generates samples that present with similar levels of per-residue flexibility as the training data. Table 1 also shows the distance distortion ΔD, the mean energy, and the mean NLL of structures generated from flow models trained with different strategies. For each model(architecture and training strategy), 3 × 106 conformations (106 structures over 3 random seeds)were generated after training with either protein G or Villin HP35. Due to the cost of computing Δ^^^^, it was computed for batches of 103samples and report statistics (mean and standard deviation)over the 3 × 103 batches. Before sample statistics for the energy u was computed, the sampleswith energy higher than the median value were first filtered out, to remove high energy outliers that are not of interest and would noise the reported mean and standard deviations. The mean and standard deviation (across 3 seeds) for the average NLL were also provided. The model is capable of generating low-energy, stable conformations for these two systems while the baseline method and ablated training strategies produce samples with energies that are positive and five or more orders of magnitude greater. Table 1 highlights a key difference in the results for protein G and villin HP35. For villin, models trained by reverse KL and without the 2-Wasserstein loss do not result in completely unraveled structures. This is consistent with the notion that long-range interactions become much more important in larger structures. From Fig.4(b), Villin HP35 is not densely packed, and local interactions, e.g., as seen in α-helices, are more prevalent than long-range interactions / contacts. In addition, the model generates diverse alternative modes of the folded villin HP35 protein that are energetically stable compared to the structures obtained from the baseline model.Docket No: 361213.00201 Fig. 4(d) visualizes pathological structures of protein G generated via different training schemes. In Fig.4(d, left), minimizing the NLL generally captures local structural motifs, such as the α-helix. However, structures generated by training only with the NLL loss tend to have clashes in the backbone, as highlighted with red circles in Fig 3(d), and / or long-range distortions. This results in large van der Waals repulsion as evidenced by the high average energy values in Table 1. As shown in Fig. 4(d, middle), structures generated by minimizing a combination of the NLL loss and the reverse KL divergence unravel and present with large global distortions. This results from large, perpetually unstable gradients during training. In Fig.4(d, right), training with a combination of the NLL loss and the 2-Wasserstein loss properly captures the backbone structural distribution but tends to have clashes in the sidechains. Table 1 demonstrates that only a model with the disclosed multistage training strategy is able to achieve both low energy samples and proper global structures. The 2-Wasserstein loss prevents large backbone distortions, and thus, simultaneously minimizing the reverse KL divergence accelerates learning for the side chain marginals with respect to the backbone and other side chain atoms. BGs can generate novel samples One of the primary goals for development of BG models is to sample important metastable states that are unseen or difficult to sample by conventional MD simulations. Protein G is a medium-size protein with diverse metastable states that provide a system for us to evaluate the capability of the BG model. First, 2D UMAP embeddings for the training data set, test dataset, andfor 2 × 105 generated samples of protein G in Fig. 5(a) were visualized. the test dataset (Fig. 5(a,middle)), an independent MD dataset, covers far less conformational space than an equivalent number of BG samples (Fig.5(a, right)). Secondly, the energy distributions of the training set from MD simulations and sample set from the BG model as shown by Fig. 5(c) were computed, respectively. Unlike the training set, the BG sample energy distribution is bimodal. Analysis of structures in the second peak revealed that a set of conformations are not present in the training set. These new structures are characterized by a large bent conformation in the hair-pin loop which links beta-strand 3 and 4 of protein G. Fig. 5(b) compares representative structures of the discovered new structure with the closest structure (by RMSD) in the training dataset. Vastly different sidechain conformations alongDocket No: 361213.00201 the bent loops between two structures were observed. Energy minimization on the discovered new structures demonstrated that these new structures are local-minimum metastable conformations. Thirdly, the lowest-energy conformations generated by the BG model were carefully examined. Fig.5(d) showcases a group of lowest-energy structures generated by the BG model, overlaid by backbone and all-atom side chains shown explicitly. All of these structures are very similar to the crystal structure of protein G, demonstrating that the trained BG model is capable of generating protein structures with high quality at the atomic level. The present disclosure is not to be limited in scope by the specific embodiments described herein. Indeed, various modifications of the invention, in addition to those described herein, will become apparent to those skilled in the art from the foregoing description and the accompanying figures. Such modifications are intended to fall within the scope of the appended claims.
Claims
Docket No: 361213.00201 CLAIMS What is claimed is:
1. A method of generating a Boltzmann distribution of functionally important metastable states of a macromolecule, comprising: generating representative conformations of a macromolecular structure of the macromolecule by molecular dynamics simulations; generating an internal coordinate representation with fixed bond length and side chain angles for each of the representative conformations of the macromolecular structure; dividing internal coordinate representations into a backbone channel and a side chain channel; inputting the internal coordinate representations of the backbone channel and the internal coordinate representations of the side-chain channel into gated attention units of a Boltzmann generator, wherein the Boltzmann generator determines long range interactions within the macromolecule based on a combination of rotary positional embeddings with global absolute positional embeddings; and determining by the Boltzmann generator a Boltzmann distribution of functionally important metastable states of the macromolecule using a Frechet Inception Distance matrix as a loss function to constrain global backbone structures to a space of native conformational ensemble.
2. The method of claim 1, comprising training the Boltzmann generator based on negative log-likelihood.
3. The method of claim 1, comprising training the Boltzmann generator based on a combination of negative log-likelihood and 2-Wasserstein loss.
4. The method of claim 1, comprising training the Boltzmann generator based on a combination of negative log-likelihood, 2-Wasserstein loss, and Kullback-Leibler divergence.
5. The method of claim 1, wherein the macromolecule is a protein.Docket No: 361213.00201 6. The method of claim 1, wherein the Boltzmann distribution of functionally important metastable states comprises low-energy conformations of the macromolecule.
7. The method of claim 1, wherein the internal coordinate representations are translation and rotation invariant.
8. The method of claim 1, wherein the gated attention units comprise gated attention rational quadratic spline coupling blocks.
9. The method of claim 1, comprising implementing relative positional embeddings on a global level.
10. The method of claim 7, further comprising concatenating backbone latent embeddings with side chain features and passing the backbone latent embeddings through the gated attention units.
11. The method of claim 1, wherein the Boltzmann generator comprises a machine-learning model.
12. The method of claim 11, wherein the Boltzmann generator comprises a deep neural network model that has been trained to learn the transformation between a mathematic Gaussian distribution and a physical Boltzmann distribution of macromolecular structures.
13. The method of claim 1, comprising generating by the Boltzmann generator metastable conformational states of the macromolecule by drawing samples from a von Mises distribution.
14. A system of generating a Boltzmann distribution of functionally important metastable states of a macromolecule, comprising one or more processors configured to: generate representative conformations of a macromolecular structure of the macromolecule by molecular dynamics simulations; generate an internal coordinate representation with fixed bond length and side chain angles for each of the representative conformations of the macromolecular structure; divide internal coordinate representations into a backbone channel and a side chain channel;Docket No: 361213.00201 input the internal coordinate representations of the backbone channel and the internal coordinate representations of the side-chain channel into gated attention units of a Boltzmann generator, wherein the Boltzmann generator determines long range interactions within the macromolecule based on a combination of rotary positional embeddings with global absolute positional embeddings; and determine by the Boltzmann generator a Boltzmann distribution of functionally important metastable states of the macromolecule using a Frechet Inception Distance matrix as a loss function to constrain global backbone structures to a space of native conformational ensemble.
15. The system of claim 14, wherein the one or more processors are configured to train the Boltzmann generator based on negative log-likelihood.
16. The system of claim 14, wherein the one or more processors are configured to train the Boltzmann generator based on a combination of negative log-likelihood and 2-Wasserstein loss.
17. The system of claim 14, the one or more processors are configured to train the Boltzmann generator based on a combination of negative log-likelihood, 2-Wasserstein loss, and Kullback- Leibler divergence.
18. The system of claim 14, wherein the macromolecule is a protein.
19. The system of claim 14, wherein the Boltzmann distribution of functionally important metastable states comprises low-energy conformations of the macromolecule.
20. The system of claim 14, wherein the internal coordinate representations are translation and rotation invariant.
21. The system of claim 14, wherein the gated attention units comprise gated attention rational quadratic spline coupling blocks.
22. The system of claim 14, wherein the one or more processors are configured to implement relative positional embeddings on a global level.Docket No: 361213.00201 23. The system of claim 20, wherein the one or more processors are configured to concatenate backbone latent embeddings with side chain features and passing the backbone latent embeddings through the gated attention units.
24. The system of claim 14, wherein the Boltzmann generator comprises a machine-learning model.
25. The system of claim 24, wherein the Boltzmann generator comprises a deep neural network model that has been trained to learn the transformation between a mathematic Gaussian distribution and a physical Boltzmann distribution of macromolecular structures.
26. The system of claim 14, wherein the one or more processors are configured to generate by the Boltzmann generator metastable conformational states of the macromolecule by drawing samples from a von Mises distribution.
Citation Information
Patent Citations
System and methods for electrostatic analysis with machine learning model
US20220277804A1
Protein kinetic ensemble platform for discovery of binding pockets and novel ligands
US20240185963A1
Constrained langevin dynamics method for simulating molecular conformations
US5553004A
Rapidly convergent method for boltzmann-weighted ensemble generation in free energy simulations
US5740072A
Trained machine learning model for forecasting molecular conformations
WO2024158510A1