Protein conformation ensemble modeling method based on diffusion model
By combining a diffusion model with a large language model and physical constraints, protein conformation sampling is optimized, which solves the shortcomings of existing methods in predicting protein dynamic conformation and achieves more accurate protein conformation modeling.
Patent Information
- Application Number
- CN202510878961.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-11
AI Technical Summary
Existing methods for predicting protein dynamic conformation cannot accurately capture conformational changes of proteins in different states, and the generated models lack physical constraints, making it difficult to comprehensively characterize the global conformational polymorphism of complex proteins.
By combining large language models and graph neural networks, protein sequence features are learned through diffusion models, physical constraints and Modeller algorithms are introduced to correct conformations, and a protein conformation ensemble that conforms to physical laws is generated, thus optimizing the sampling process.
It improves the accuracy and diversity of protein conformation sampling, and provides the ability to accurately model the dynamic behavior and function of proteins.
Smart Images

Figure CN120932723A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and computational intelligence, and in particular to a method for protein conformation ensemble modeling based on a diffusion model. Background Technology
[0002] Protein structure is central to understanding the molecular mechanisms of life. It not only determines protein function but also profoundly influences key biological processes such as cellular metabolism, signal transduction, and immune responses. Therefore, accurate resolution of protein structures is crucial for life science research, especially in disease research and drug development, where protein structural information provides the theoretical foundation for a deeper understanding of pathological mechanisms and drug targets. In recent years, with the rapid development of artificial intelligence technology, breakthroughs have been achieved in the field of protein structure prediction. Cutting-edge methods such as Google's AlphaFold2 and David Baker's RoseTTAfold have achieved prediction accuracy comparable to experimentally resolved structures. In 2024, the Nobel Prize in Chemistry was awarded to three scientists for their pioneering contributions to protein design and structure prediction. Behind this prestigious award lies the profound impact and significance of protein structure prediction technology on biological science. However, the realization of protein function does not solely depend on its static structure; many functions are accomplished through the switching between different conformations of the protein.
[0003] Despite significant breakthroughs in protein structure prediction using deep learning, capturing dynamic conformational changes and comprehensively sampling the conformational space of proteins remains a core challenge in protein dynamics research. With increased computational power, molecular dynamics (MD) simulations can now run on timescales from microseconds to milliseconds, accurately depicting the dynamic properties of proteins and enabling effective exploration of their structural evolution and dynamic behavior. However, MD simulations are computationally expensive, making it difficult to efficiently sample the global conformational space of complex protein systems, especially when studying large-scale conformational changes. Traditional MD methods may be limited by computation time and the complexity of potential energy surfaces. Furthermore, Monte Carlo (MC) sampling methods exhibit unique advantages in studying the free energy landscape, transition states, and the stability of individual amino acids at different temperatures. By rapidly sampling different conformations, MC sampling methods can effectively explore the diversity of protein conformations. This is of great value in revealing protein conformational changes and function-related conformational states. However, MC sampling methods rely on appropriate conformational sampling strategies and may get trapped in local minima in high-dimensional free energy surfaces, making it difficult to comprehensively characterize the global conformational polymorphism of complex proteins. In recent years, AlphaFold-based protein dynamic conformation prediction methods have gradually emerged. Their core idea is to obtain the conformation of proteins in different states after changing the model input (such as templates or multiple sequence alignments), thereby revealing, to some extent, function-related conformational changes. However, these methods are still mainly based on static structure prediction frameworks, unable to provide accurate time-dependent information, and limited by the quality of the input multiple sequence alignments (MSA) or templates, making it difficult to directly analyze highly dynamic protein systems. Besides AlphaFold-driven methods, generative models are becoming important tools for studying protein dynamic conformation. These models learn patterns from large-scale training data, extract low-dimensional representations in high-dimensional space, and directly generate diverse sets of protein conformations without relying on traditional large-scale sampling operations, providing a more efficient solution for protein dynamics research. However, existing generative models still face the problem of insufficient constraints on the physical rationality of conformations. Therefore, with the rapid development of deep learning, generative modeling, and augmented sampling techniques, how to combine existing large language models with deep learning methods and incorporate physical constraints to more accurately capture the dynamics of protein evolution over time and model the conformational ensemble of proteins remains a significant challenge to be addressed in this field. Summary of the Invention
[0004] To overcome the limitations of existing protein dynamic ensemble modeling methods, this invention proposes a diffusion-based protein ensemble modeling method. This method combines a large language model to deeply mine the dynamic information contained in protein sequences and incorporates physical constraints to guide the model in generating protein conformational ensembles that conform to physical laws, thereby improving the accuracy of modeling protein structural dynamics. This method can effectively improve the accuracy and diversity of protein conformational sampling, providing new insights for a deeper understanding of the dynamic characteristics and functional mechanisms of proteins.
[0005] The technical solution adopted by this invention to solve its technical problem is:
[0006] A protein ensemble modeling method based on a diffusion model is proposed. First, key features of the protein at the atomic, residue, and overall topological structures are extracted and input into a diffusion model based on a graph neural network (GNN) for learning. During the denoising process, the Modeller algorithm is introduced to physically correct the generated conformations to ensure their rationality and biological credibility. Subsequently, through iterative optimization, a protein conformation ensemble that conforms to physical constraints is continuously generated, thereby accurately capturing the dynamic behavior and functional conformational changes of proteins.
[0007] Furthermore, the method includes the following steps:
[0008] 1) Dataset Construction: First, the amino acid sequences of all protein chains in the PDB database are clustered according to 80% sequence similarity. Then, all structures within each cluster are compared and the structures are pruned to ensure consistency. Based on this, classes with sequence lengths between 50 and 500 are selected, and the two conformations with the greatest structural differences in each class are selected as the final structures, ensuring that their TM-scores are less than 0.98. Finally, the training set and validation set are divided with a time limit.
[0009] 2) Creating labeled data: For two protein conformations within the same class, firstly, calculate their respective solvent accessible area (SASA) values, and designate the conformation with the larger solvent accessible area as state 1 and the conformation with the smaller solvent accessible area as state 2. Using state 1 as a baseline, calculate the conformational difference between state 1 and state 2, mainly considering the rotation, translation, and changes in the dihedral angle of the side chains between amino acid pairs. Then, convert this difference into Gaussian noise and add this noise to state 1 to generate a new conformation. Here, the noise is not only added to state 1 as a perturbation but also used as a label for training the model.
[0010] 3) Feature extraction: Extract the sequence and structural features of each protein, and describe its features at the atomic level, residue level and overall topological structure level respectively;
[0011] 4) The ESM features and amino acid physicochemical properties obtained in step 3) are spliced together as node features, and the radius of gyration, side chain dihedral angle and backbone orientation features are spliced together as edge features. The features are then input into the network for further processing.
[0012] 5) Building the network model: In the four-layer diffusion model based on graph neural network, each layer includes a linear layer, a ReLU activation function and a dropout layer. The model is trained using the Adam optimizer with an appropriate learning rate and the L1 loss function is used for optimization training.
[0013] 6) Constructing an energy field: During the model denoising process, a physical potential energy function is introduced. As a guide to optimize the sampling process, It contains two energy terms: vdw and hb_srbb. The vdw term represents only steric repulsion, avoiding unreasonable conformations during atomic collisions. hb_srbb is the short-range main-chain hydrogen bond energy term, which allows...
[0014] Spirals or adjacent β hairpins form rapidly in the initial state. The definition is as follows:
[0015]
[0016] 7) Training model parameters: Based on the network architecture built in step 5), firstly, by setting training and validation sets, the Adam optimizer is used to optimize the weights and bias parameters of the network model. During training, gradient updates are performed according to the set learning rate, and the L1 loss function is used as the objective function. The gradient is calculated through the backpropagation algorithm, and the parameters of each layer are updated to optimize the performance of the model. The training process will adjust the model parameters through several rounds to minimize the loss function.
[0017] 8) Select the ATLAS test set of protein MD simulation data, input the simulated initial structure into the trained network for noise reduction. For the output structure, first use the Modeller tool for physical correction, and then input the corrected structure back into the trained model. Repeat this process until N conformations are obtained, thus obtaining the final protein ensemble.
[0018] Furthermore, the process described in 3) is as follows:
[0019] 3.1) Extracting ESM features: The amino acid sequence is input into the protein big language model ESM3. The model encodes each amino acid position to generate the corresponding high-dimensional representation vector.
[0020] 3.2) Extraction of physicochemical properties of amino acids: This feature set contains seven main features of amino acids, namely: the physical volume of the molecule, the size of the side chain and the radius of the sphere, the variability of the electron cloud, the probability of forming α-helical and β-sheet structures, the pH value when the number of positive and negative charges of zwitterions is equal, this feature helps to understand the charge characteristics of amino acids and their stability under specific pH conditions, and the degree of hydrophobicity.
[0021] 3.3) Extracting the radius of gyration feature: Radius of gyration R g R is an important parameter for measuring the compactness of protein structure, representing the average squared distance between the centers of mass and geometric centers of all atoms in a protein. g The smaller the value, the more compact the protein structure; conversely, the smaller the R value, the more compact the protein structure. g The larger the value, the looser the protein structure. The equation for the radius of gyration is given by the following formula:
[0022]
[0023] Where N is the total number of atoms in the protein, r i It is the position vector of the i-th atom, r cm It is the mass center position vector of all atoms;
[0024] 3.4) Extraction of side chain dihedral angles and framework atom orientations: Calculations based on C α The sine and cosine values corresponding to the dihedral angles of the outward-extending side chains of atoms are used. Each amino acid corresponds to a maximum of five rotatable χ angles [χ1,χ2,χ3,χ4,χ5] and two symmetrical χ angles [altχ1,altχ2]. To ensure the uniqueness of the side chain angles for a given structure, the dihedral angle of a particular amino acid's side chain is always defined as [max(χ1,altχ1),min(χ1,altχ1),max(χ2,altχ2),min(χ2,altχ2),χ3,χ4,χ5]. In the design of the skeletal atom orientation features, α-carbon (C) is selected. α ), main chain nitrogen (N), side chain first carbon atom (C) β The oxygen atom (O) and carboxyl carbon (C) are used as the skeleton atoms, and the relative position and orientation of each amino acid in three-dimensional space are extracted by using x, y, and z values in three-dimensional space.
[0025] The beneficial effects of this invention are as follows: By combining data from the PDB database with the powerful capabilities of large language models, it deeply mines the dynamic conformational information hidden in protein amino acid sequences and static databases, providing a more accurate modeling and analysis method for the dynamic behavior and function of proteins. In this invention, while using a diffusion model to generate diverse conformations, the Modeller tool is used to gradually guide conformational optimization, ensuring the rationality and scientific validity of the generated conformations. Attached Figure Description
[0026] Figure 1 This is a flowchart of a protein conformation ensemble modeling method based on a diffusion model.
[0027] Figure 2 The ensemble structure of protein 4AKE_B is predicted by a protein conformational ensemble modeling method based on a diffusion model, where A is the ensemble generated by molecular dynamics simulation and B is the conformational ensemble generated in this invention. Detailed Implementation
[0028] The present invention will now be further described with reference to the accompanying drawings.
[0029] Reference Figure 1 and Figure 2 A method for protein conformational ensemble modeling based on a diffusion model, the method comprising the following steps:
[0030] 1) Dataset Construction: First, the amino acid sequences of all protein chains in the PDB database (as of January 1, 2025) were clustered according to 80% sequence similarity. Then, all structures within each cluster were compared and pruned to ensure consistency. Based on this, classes with sequence lengths between 50 and 500 were selected, and the two conformations with the greatest structural differences from each class were selected as the final structures, ensuring that their TM-scores were less than 0.98. Finally, the training and validation sets were divided with January 1, 2024 as the cutoff date.
[0031] 2) Creating labeled data: For two protein conformations within the same class, first calculate their respective solvent accessible area (SASA) values, and set the conformation with the larger solvent accessible area as state 1 and the conformation with the smaller solvent accessible area as state 2. Using state 1 as the baseline, calculate the conformational difference between state 1 and state 2, mainly considering the rotation, translation between amino acid pairs and the change of dihedral angle of the side chain. Then, convert this difference into Gaussian distributed noise and add this noise to state 1 to generate a new conformation. Here, the noise is not only added to state 1 as a perturbation, but also used as a label for training the model.
[0032] 3) Feature Extraction: For a protein with a sequence length of L, its sequence and structural features are extracted to describe its characteristics at the atomic, residue, and overall topological levels, respectively. The process is as follows:
[0033] 3.1) Extracting ESM features: The amino acid sequence is input into the protein big language model ESM3. The model encodes each amino acid position to generate a corresponding high-dimensional representation vector with a feature dimension of L×1280.
[0034] 3.2) Extraction of physicochemical properties of amino acids: This feature set contains seven main features of amino acids, namely: the physical spatial volume of the molecule, the size of the side chain and the radius of the sphere, the variability of the electron cloud, the probability of forming α-helical and β-sheet structures, the pH value when the number of positive and negative charges of zwitterions is equal, this feature helps to understand the charge characteristics of amino acids and their stability under specific pH conditions, and the degree of hydrophobicity. The corresponding feature dimension extracted is L×7.
[0035] 3.3) Extracting the radius of gyration feature: Radius of gyration R g It is an important parameter for measuring the compactness of protein structure, representing the average squared distance between the center of mass of all atoms in the protein and their geometric center. The extracted feature dimension is L×1, R g The smaller the value, the more compact the protein structure; conversely, the smaller the R value, the more compact the protein structure. g The larger the value, the looser the protein structure. The equation for the radius of gyration is given by the following formula:
[0036]
[0037] Where N is the total number of atoms in the protein, r i It is the position vector of the i-th atom, r cm It is the mass center position vector of all atoms;
[0038] 3.4) Extraction of side chain dihedral angles and framework atom orientations: Calculations based on C α The sine and cosine values corresponding to the dihedral angles of the outward-extending side chains of atoms are given. Each amino acid corresponds to a maximum of five rotatable χ angles [χ1,χ2,χ3,χ4,χ5] and two symmetrical χ angles [altχ1,altχ2]. To ensure the uniqueness of the side chain angles for a given structure, the dihedral angle of a particular amino acid is always defined as [max(χ1,altχ1),min(χ1,altχ1),max(χ2,altχ2),min(χ2,altχ2),χ3,χ4,χ5]. Therefore, the characteristic dimension of the side chain dihedral angle is L×7×2. In the design of the skeletal atom orientation features, α-carbon (C) is selected. α ), main chain nitrogen (N), side chain first carbon atom (C) β Oxygen (O) and carboxyl carbon (C) are used as skeletal atoms. They are represented by x, y, and z values in three-dimensional space. The relative position and orientation of each amino acid in three-dimensional space are extracted, and the corresponding feature dimension is L×5×3.
[0039] 4) The ESM features and amino acid physicochemical properties obtained in step 3) are spliced together as node features, and the radius of gyration, side chain dihedral angle and backbone orientation features are spliced together as edge features. The features are then input into the network for further processing.
[0040] 5) Building the network model: In the four-layer diffusion model based on graph neural network, each layer includes a linear layer, a ReLU activation function and a dropout layer. The model is trained using the Adam optimizer, with weights selected as β1 = 0.9 and β2 = 0.999, and an appropriate learning rate lr of 1e-4 is set. At the same time, the L1 loss function is used for optimization training.
[0041] 6) Constructing an energy field: During the model denoising process, a physical potential energy function is introduced. As a guide to optimize the sampling process, It contains two energy terms: vdw and hb_srbb. The vdw term represents only steric repulsion, avoiding unreasonable conformations during atomic collisions. hb_srbb is the short-range main-chain hydrogen bond energy term, which allows...
[0042] Spirals or adjacent β hairpins form rapidly in the initial state. The definition is as follows:
[0043]
[0044] 7) Training model parameters: Based on the network architecture built in step 5), firstly, by setting training and validation sets, the Adam optimizer is used to optimize the weights and bias parameters of the network model. During training, gradient updates are performed according to the set learning rate, and the L1 loss function is used as the objective function. The gradient is calculated through the backpropagation algorithm, and the parameters of each layer are updated to optimize the performance of the model. The training process will adjust the model parameters through several rounds to minimize the loss function.
[0045] 8) Select the ATLAS test set of protein MD simulation data, input the simulated initial structure into the trained network for noise reduction. For the output structure, first use the Modeller tool for physical correction, and then input the corrected structure back into the trained model. Repeat this process until N (N=250) conformations are obtained, thus obtaining the final protein ensemble.
[0046] This embodiment uses the protein 6d7y_A with a sequence length of 96 as an example, and successfully predicted the conformational ensemble of this protein using the above method. The predicted protein conformational ensemble is comparable to the result of MD simulation, with a corresponding root mean square Wasserstein distance of 2.5. The results are shown in the figure below. Figure 2 As shown.
[0047] The above describes the results of an example provided by the present invention. Obviously, the present invention is not only suitable for the above embodiments, but can also be implemented in various ways without departing from the basic spirit of the present invention and without exceeding the content involved in the substantive content of the present invention.
Claims
1. A protein conformation ensemble modeling method based on a diffusion model, characterized in that, First, key features of the protein at the atomic, residue, and overall topological structure are extracted and then fed into a diffusion model based on graph neural networks (GNNs) for learning. During the denoising process, the Modeller algorithm is introduced to physically correct the generated conformation to ensure its rationality and biological credibility. Subsequently, through iterative optimization, a protein conformation ensemble that conforms to physical constraints is continuously generated, thereby accurately capturing the dynamic behavior and functional conformational changes of proteins.
2. The protein conformation ensemble modeling method based on a diffusion model as described in claim 1, characterized in that, The method includes the following steps: 1) Dataset Construction: First, the amino acid sequences of all protein chains in the PDB database are clustered according to 80% sequence similarity. Then, all structures within each cluster are compared and the structures are pruned to ensure consistency. Based on this, classes with sequence lengths between 50 and 500 are selected, and the two conformations with the greatest structural differences in each class are selected as the final structures, ensuring that their TM-scores are less than 0.
98. Finally, the training set and validation set are divided with a time limit. 2) Creating tag data: For two protein conformations within the same class, first calculate their respective solvent accessible area (SASA) values, and designate the conformation with the larger solvent accessible area as state 1 and the conformation with the smaller solvent accessible area as state 2; using state 1 as the baseline, calculate the conformational difference between state 1 and state 2, mainly considering the rotation, translation between amino acid pairs and the change in the dihedral angle of the side chain; then, convert this difference into Gaussian distributed noise and add this noise to state 1 to generate a new conformation; The noise here is not only added to state 1 as a perturbation, but also used as a label for training the model; 3) Feature extraction: Extract the sequence and structural features of each protein, and describe its features at the atomic level, residue level and overall topological structure level respectively; 4) The ESM features and amino acid physicochemical properties obtained in step 3) are spliced together as node features, and the radius of gyration, side chain dihedral angle and backbone orientation features are spliced together as edge features. The features are then input into the network for further processing. 5) Building the network model: In the four-layer diffusion model based on graph neural network, each layer includes a linear layer, a ReLU activation function and a dropout layer. The model is trained using the Adam optimizer with an appropriate learning rate and the L1 loss function is used for optimization training. 6) Constructing an energy field: During the model denoising process, a physical potential energy function is introduced. As a guide to optimize the sampling process, It contains two energy terms: vdw and hb_srbb. The vdw term only represents steric repulsion, avoiding unreasonable conformations during atomic collisions. The hb_srbb term is the short-range main-chain hydrogen bond energy term, which allows helices or adjacent β hairpins to form rapidly in the initial state. The definition is as follows: 7) Training model parameters: Based on the network architecture built in step 5), firstly, by setting training and validation sets, the Adam optimizer is used to optimize the weights and bias parameters of the network model. During training, gradient updates are performed according to the set learning rate, and the L1 loss function is used as the objective function. The gradient is calculated through the backpropagation algorithm, and the parameters of each layer are updated to optimize the performance of the model. The training process will adjust the model parameters through several rounds to minimize the loss function. 8) Select the ATLAS test set of protein MD simulation data, input the simulated initial structure into the trained network for noise reduction. For the output structure, first use the Modeller tool for physical correction, and then input the corrected structure back into the trained model. Repeat this process until N conformations are obtained, thus obtaining the final protein ensemble.
3. The protein conformation ensemble modeling method based on a diffusion model as described in claim 2, characterized in that, The process described in 3) is as follows: 3.1) Extracting ESM features: The amino acid sequence is input into the protein big language model ESM3. The model encodes each amino acid position to generate the corresponding high-dimensional representation vector. 3.2) Extraction of physicochemical properties of amino acids: This feature set contains seven main features of amino acids, namely: the physical volume of the molecule, the size of the side chain and the radius of the sphere, the variability of the electron cloud, the probability of forming α-helical and β-sheet structures, the pH value when the number of positive and negative charges of zwitterions is equal, this feature helps to understand the charge characteristics of amino acids and their stability under specific pH conditions, and the degree of hydrophobicity. 3.3) Extracting the radius of gyration feature: Radius of gyration R g R is an important parameter for measuring the compactness of protein structure, representing the average squared distance between the centers of mass and geometric centers of all atoms in a protein. g The smaller the value, the more compact the protein structure; conversely, the smaller the R value, the more compact the protein structure. g The larger the value, the looser the protein structure. The equation for the radius of gyration is given by the following formula: Where N is the total number of atoms in the protein, r i It is the position vector of the i-th atom, r cm It is the mass center position vector of all atoms; 3.4) Extraction of side chain dihedral angles and framework atom orientations: Calculations based on C α The sine and cosine values corresponding to the dihedral angles of the side chains extending outward from the atoms are given. Each amino acid corresponds to a maximum of five rotatable χ angles [χ1,χ2,χ3,χ4,χ5] and two symmetrical χ angles [altχ1,altχ2]. To ensure the uniqueness of the side chain angles of a given structure, the dihedral angle of the side chain of a certain amino acid is always defined as [max(χ1,altχ1),min(χ1,altχ1),max(χ2,altχ2),min(χ2,altχ2),χ3,χ4,χ5]. In the design of the skeletal atomic orientation features, α-carbon (C) was selected. α ), main chain nitrogen (N), side chain first carbon atom (C) β The oxygen atom (O) and carboxyl carbon (C) are used as the skeleton atoms, and the relative position and orientation of each amino acid in three-dimensional space are extracted by using x, y, and z values in three-dimensional space.