RNA dynamic conformation ensemble generation method based on generative artificial intelligence model

By using a denoised diffusion probability model based on a generative artificial intelligence model and an isovariant graph neural network, an RNA dynamic conformation ensemble can be directly generated. This solves the problems of high computational cost and multiple sequence alignment in existing technologies, and achieves rapid and accurate RNA dynamic conformation generation, which is suitable for RNA structural biology and synthetic biology research.

CN121661231APending Publication Date: 2026-03-13SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and cost-effectively generate ensembles of dynamic RNA conformations that conform to physical laws. Furthermore, traditional methods rely on multiple sequence alignments, resulting in high computational resource consumption and making it difficult to accurately characterize the dynamic conformational diversity and excited-state conformations of RNA.

Method used

A generative artificial intelligence model-based approach is adopted, which utilizes a denoised diffusion probability model and an isovariant graph neural network. By adding noise through the forward diffusion process and using the isovariant graph neural network for denoising modeling, a dynamic conformation ensemble of RNA is directly generated, avoiding the need for multiple sequence alignment.

Benefits of technology

It enables the rapid and efficient generation of dynamic RNA conformation ensembles that conform to physical laws, balancing structural fidelity and conformational diversity. It can accurately capture the excited and high-energy conformations of RNA, reduce computational costs, and is suitable for RNA structural biology and synthetic biology research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661231A_ABST
    Figure CN121661231A_ABST
Patent Text Reader

Abstract

The invention relates to an RNA dynamic conformation ensemble generation method based on a generative artificial intelligence model. The method comprises the following steps: S1, obtaining RNA three-dimensional structure training data; s2, converting the RNA three-dimensional structure data into an RNA three-dimensional space map structure, and extracting RNA three-dimensional space map structure features; s3, inputting the RNA three-dimensional space graph structural features into a diffusion probability model, adding noise through a forward diffusion process to obtain noise-added data, and performing denoising modeling on the noise-added data based on an isotropic graph neural network to obtain reconstructed data; s4, training a diffusion probability model based on the reconstructed data; and S5, outputting a conformation ensemble for the diffusion probability model after single conformation training is completed. Compared with the prior art, the method has the advantages that the RNA dynamic conformation ensemble conforming to the physical law can be quickly and efficiently generated in an end-to-end manner, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic conformational ensemble analysis of RNA, and in particular to a method for generating dynamic conformational ensembles of RNA based on a generative artificial intelligence model. Background Technology

[0002] RNA exhibits far greater structural diversity and functional complexity in cells than previously thought. Besides serving as messenger RNA (mRNA) to transmit genetic information, it also participates extensively in important physiological processes such as regulation, catalysis, and transcriptional regulation as non-coding RNA (ncRNA). RNA's unique three-dimensional structure and dynamic conformation are crucial to its various biological functions. While experimental structural analysis methods such as X-ray crystallography, nuclear magnetic resonance (NMR), and cryo-electron microscopy (Cryo-EM) can obtain high-resolution static structures, they struggle to resolve the dynamic conformational ensemble of RNA. Traditional molecular dynamics (MD) simulations, as a supplementary theoretical approach, suffer from the following bottlenecks: (1) extremely low efficiency in sampling RNA conformational space, typically requiring significant computational resources; (2) the RNA molecular force field is still immature, leading to simulation results that deviate from experimental data, such as excessive bias towards incorrect nested conformations.

[0003] In recent years, deep learning, especially diffusion probabilistic models, has shown great potential in protein structure modeling and graph generation tasks. However, current mainstream structure prediction tools such as AlphaFold2 / 3 mainly focus on proteins and tend to generate static structures, failing to solve the problem of generating dynamic conformational ensembles of RNA. Summary of the Invention

[0004] The purpose of this invention is to provide a method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model, in order to achieve rapid and efficient end-to-end generation of RNA dynamic conformation ensembles that conform to physical laws.

[0005] The objective of this invention can be achieved through the following technical solutions: A method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model, comprising the following steps: S1. Obtain RNA three-dimensional structure training data; S2. Convert the RNA three-dimensional structure data into an RNA three-dimensional spatial map structure and extract the features of the RNA three-dimensional spatial map structure. S3, RNA three-dimensional spatial graph structural features are input into the diffusion probability model. Noise is added through the forward diffusion process to obtain noisy data. The noisy data is then modeled for denoising based on the equivariant graph neural network to obtain reconstructed data. S4. Train a diffusion probability model based on reconstructed data; S5. Output the conformational ensemble from the diffusion probability model trained on a single conformation.

[0006] Furthermore, the specific steps for adding noise through a forward diffusion process to obtain noisy data are as follows: The structural features of the RNA three-dimensional spatial map at the previous time step are obtained, and Gaussian noise is added to them to obtain the structural features of the RNA three-dimensional spatial map at the next time step as the noisy data.

[0007] Furthermore, the structural features of the RNA three-dimensional spatial graph include C4' node features and edge features.

[0008] Furthermore, the structural features of the RNA three-dimensional spatial map at the next time step are as follows: in, and The three-dimensional spatial structure features of RNA represent the next time step and the preceding time step, respectively. This represents the noise level at the preceding time step. Describe the Gaussian noise of the preceding time step.

[0009] Furthermore, the noise level at the preceding time step is: in, and These represent the noise levels at the preceding time step, time step 0, and the total number of time steps, respectively.

[0010] Furthermore, the noisy data is modeled using an equivariant graph neural network to obtain the reconstructed data. The specific steps are as follows: The noise vector at each time step is predicted based on the isotropic graph neural network, and the noise-reducing model is used to model the noise-added data based on the predicted noise vector to obtain the reconstructed data.

[0011] Furthermore, The reconstructed data for each moment is as follows: in, express Reconstruction data at any given moment Represents signal retention rate. , representing the cumulative signal retention rate, This represents the noise vector predicted by the isomorphic graph neural network. This represents the Gaussian noise newly introduced during the sampling phase.

[0012] Furthermore, the isovariant graph neural network is trained using mean squared error.

[0013] Furthermore, the C4' node features are the three-dimensional coordinates and time step embedding information of the C4' atom.

[0014] Furthermore, edge features represent the spatial relationships and topological connections between adjacent nucleotides.

[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention provides a method for generating RNA dynamic conformation ensembles based on a diffusion model, specifically a method for generating RNA dynamic conformations using a denoised diffusion probability model and an isovariant graph neural network. The method of this invention does not rely on multiple sequence alignment (MSA) information, which avoids the time-consuming and costly problem of traditional structure prediction methods that require large-scale complex multiple sequence alignments. It can quickly and efficiently generate RNA dynamic conformation ensembles that conform to physical laws end-to-end.

[0016] (2) This invention takes into account both structural fidelity and conformational diversity. It takes the real structure as the learning target and fits the energy distribution of the real RNA conformation by learning the noise perturbation pattern of the real structure. Therefore, it can accurately generate an RNA conformational ensemble that conforms to the Boltzmann distribution. At the same time, it can overcome the bottleneck of difficult sampling of excited state and high-energy state conformations, take into account physical rationality, and can efficiently sample the diverse structures of RNA. It can also be extended to other systems such as DNA and protein-nucleic acid complexes. Attached Figure Description

[0017] Figure 1 This is a schematic flowchart of a method for generating RNA dynamic conformation ensembles based on a deep learning diffusion model, provided by an embodiment of the present invention. Figure 1 A represents the framework of DynaRNA. Figure 1 B is a comparison of the C4' neighbor distance distribution of the conformation set generated by DynaRNA with the experimental structure distribution in PDB. Figure 1 C is an isovariate graphical neural network (EGNN) used to predict noise and the denoising process. Figure 1 D is a comparison of the C4' adjacent angle distribution of the conformation set generated by DynaRNA with the experimental structure distribution of PDB; Figure 2 This is a graph showing the test results of the tetranucleotide system in an embodiment of the present invention, wherein... Figure 2 A represents the intercalation ratio of the set of conformations generated by molecular dynamics simulations starting from the type A conformation to the set of conformations generated by DynaRNA. Figure 2B represents the intercalation ratio of the set of conformations generated by molecular dynamics simulations starting from the intercalated conformations to the set of conformations generated by DynaRNA. Figure 3 This is a diagram showing the test results of the HIV inverse activation response (TAR) element system in an embodiment of the present invention. Figure 3 A represents the secondary structure of HIV-1 TAR in its ground and excited states. Figure 3 B represents the tertiary structure of the ground and excited states generated by DynaRNA. Figure 3 C represents the principal component analysis results of the conformational assemblage generated from the ground-state initialized DynaRNA. Figure 3 D represents the principal component analysis results of the conformation set generated by the excited-state DynaRNA. Detailed Implementation

[0018] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0019] This invention relates to a generative artificial intelligence model-based method for generating RNA dynamic conformation ensembles (DynaRNA). Addressing the limitations and high computational costs of existing traditional experimental techniques and molecular dynamics (MD) simulations in characterizing RNA dynamic conformation ensembles, this invention provides a method for directly modeling the three-dimensional spatial coordinates of RNA using a denoised diffusion probability model (DDPM) and an equal variation graphical neural network (EGNN). This method generates RNA dynamic conformation ensembles end-to-end without requiring multiple sequence alignment (MSA) information. DynaRNA can efficiently and accurately explore RNA conformation space, effectively reduce the base intercalation error rate in tetranucleotide simulations, successfully capture rare excited-state conformations of key RNAs such as HIV transactivation response (TAR) elements, and de novo folding of four-loop structures. This invention provides a universal and efficient computational platform for RNA structural dynamics research, with broad application prospects in RNA structural biology, synthetic biology, and RNA therapy development.

[0020] The present invention Figure 1 A represents the framework of DynaRNA, which includes two processes: the forward denoising process, represented by solid lines, in which Gaussian noise is gradually added to the input structure; and the reverse denoising process, represented by dashed lines. Figure 1 B is a comparison of the C4' neighbor distance distribution of the conformation set generated by DynaRNA with the structure distribution in the PDB experiment. Figure 1 C is an isovariate graphical neural network (EGNN) used to predict noise and the denoising process. Figure 1 D is a comparison of the C4' adjacent angle distribution of the conformation set generated by DynaRNA with the experimental structure distribution of PDB.

[0021] Figure 2 This is a graph showing the test results of the tetranucleotide system in an embodiment of the present invention, wherein... Figure 2 A represents the intercalation ratio of the set of conformations generated by molecular dynamics simulations starting from the A-type conformation to the set of conformations generated by DynaRNA. Figure 2 B represents the intercalation ratio of the set of conformations generated by molecular dynamics simulations starting from the intercalated conformations to the set of conformations generated by DynaRNA.

[0022] Figure 3 This is a diagram showing the test results of the HIV inverse activation response (TAR) element system in an embodiment of the present invention. Figure 3 A represents the secondary structure of HIV-1 TAR in its ground and excited states. Figure 3 B represents the tertiary structure of the ground and excited states generated by DynaRNA. Figure 3 C represents the principal component analysis results of the conformational assemblage generated from the ground-state initialized DynaRNA. Figure 3 D represents the principal component analysis results of the conformation set generated by the excited-state DynaRNA.

[0023] This invention proposes a method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model, the method comprising the following steps: S1. Obtain RNA three-dimensional structure training data; S2. Convert the RNA three-dimensional structure data into an RNA three-dimensional spatial map structure and extract the features of the RNA three-dimensional spatial map structure. S3, RNA three-dimensional spatial graph structural features are input into the diffusion probability model. Noise is added through the forward diffusion process to obtain noisy data. The noisy data is then modeled for denoising based on the equivariant graph neural network to obtain reconstructed data. S4. Train a diffusion probability model based on reconstructed data; S5. Output the conformational ensemble from the diffusion probability model trained on a single conformation.

[0024] This invention directly models the three-dimensional coordinates of the C4' atom in RNA molecules based on a diffusion generation model, enabling efficient generation of dynamic RNA conformation sets, reconstructing experimental conformation features, capturing rare excited states, and predicting RNA folding paths.

[0025] The specific steps of this invention include: (1) Input data modeling: The three-dimensional structure of the input RNA is simplified into a coarse-grained model in which each nucleotide is represented by a C4' atom. A spatial graph structure is constructed, in which nodes are C4' atoms and edges contain features such as distance between nucleotides and connectivity.

[0026] (2) Forward Diffusion: The innovative partial diffusion strategy only diffuses to the intermediate noise step (e.g., step 800), gradually adding Gaussian noise to the original RNA structure to form perturbation data; this partial diffusion strategy effectively balances the effectiveness and rationality of RNA structure generation, greatly improving the generation effect.

[0027] (3) Reverse Denoising: The added noise is predicted using an Equivariant Graph Neural Network (EGNN) (E(3)) to gradually perform reverse denoising and then reconstruct the original structure. (4) Model Training: The training dataset consisted of 6820 high-quality RNA 3D structures selected from the PDB (Protein Data Bank) database. The model was implemented using the PyTorch framework and trained for 14 days on an NVIDIA 4090D GPU. The training optimizer was Adam with a learning rate of 0.0001. The network training employed the L2 norm loss function to minimize the difference between the actual noise and the model's predicted noise. (5) Dynamic conformation generation: The trained model can be used to directly model the coordinates of the single-frame RNA structure input by the user and generate a dynamic conformation ensemble.

[0028] This invention utilizes the Denoising Diffusion Probability Model (DDPM) to perturb the coordinates of RNA atoms, thereby generating multi-conformation coordinates.

[0029] This invention uses an Equivariant Graph Neural Network (EGNN) (E(3)) to denoise the data and then reconstruct the RNA dynamic conformation ensemble.

[0030] This invention provides a novel method for generating end-to-end dynamic conformational ensembles of RNA without multiple sequence alignment (MSA) and based on C4' atom modeling. The core of this method is to progressively perturb and reconstruct the RNA atom coordinates using a diffusion model and an isovariant graphical neural network, thereby efficiently and physically reliably generating a series of three-dimensional structures that accurately reflect the dynamic conformational ensemble characteristics of RNA. The technical problem solved by this invention is achieved through the following technical solution: (1) Training set collection and data processing: 6820 high-quality RNA three-dimensional structures were collected based on the most commonly used PDB database, while non-RNA components such as proteins and small molecules were removed. (2) Input data modeling: The coordinates of the input RNA atomic structure are extracted and coarsely represented. With the C4' atom as the center, each nucleotide is represented as a node to construct a three-dimensional spatial graph structure as the model input; (3) Data noise perturbation: Gaussian noise is added to the input data using a diffusion probability model. The model employs a partial diffusion strategy, not completely adding noise to a pure Gaussian noise distribution, but only adding noise to a moderate number of steps to retain some information of the initial input structure while introducing perturbation. The forward diffusion process involves adding Gaussian noise to gradually perturb the original RNA conformation to an approximate Gaussian noise state, constructing a trajectory from the data distribution to the prior distribution (usually a standard normal distribution). The detailed principle and process of this procedure are as follows: Given the preceding time Data distribution Add Gaussian noise to it Thus, the next moment is obtained. Data distribution This process is defined by the following stochastic differential equation: In the above formula, and These represent the data distribution at the corresponding time points. This represents the noise level at a given moment, and its intensity varies with the number of time steps, as defined by the following formula: In the above formula, and These represent the noise levels at the corresponding times, respectively. It is 0.0001. It is 0.02.

[0031] By combining the above formulas, Gaussian noise can be gradually added to the initial distribution of the input data, thereby constructing the forward diffusion noise addition process of the model.

[0032] (4) Data Denoising and Reconstruction: Denoising modeling is performed based on the E(3) equivariant graph neural network to reconstruct RNA structural data, thereby ensuring that the generated results satisfy rotation and translation symmetry; the denoising process is guided by the denoising function learned by the above equivariant graph neural network, and the noise predicted by the denoising function is used to perform the inverse process of approximating the noise addition process on the data distribution, and new samples can be generated iteratively through back diffusion. Specifically, it is implemented through the following formula: in, express Distribution of predicted data at time points Represents signal retention rate. Represents the cumulative signal retention rate. The noise represented by the EGNN network model prediction. This represents the Gaussian noise newly introduced during the sampling phase.

[0033] (5) Model training: Based on PyTorch, it supports multi-GPU training and breakpoint recovery. It uses L2 noise regression loss as the training target to avoid the instability caused by direct regression of structural coordinates. It uses features such as time step embedding, graph structure embedding, and spatial distance encoding to enhance the representation capability. (6) Dynamic conformation generation: After the user inputs a single conformation, the trained model can be used to add noise and remove noise, thereby quickly outputting a set of structures for conformation analysis, experimental data fitting, etc.

[0034] The beneficial effects of this invention are: (1) This invention provides a method for generating RNA dynamic conformation ensembles based on a diffusion model, specifically a method for generating RNA dynamic conformations using a denoised diffusion probability model and an isovariant graph neural network. The method of this invention does not rely on multiple sequence alignment (MSA) information, which avoids the time-consuming and costly problem of traditional structure prediction methods that require large-scale complex multiple sequence alignments. It can quickly and efficiently generate RNA dynamic conformation ensembles that conform to physical laws end-to-end.

[0035] (2) This invention takes into account both structural fidelity and conformational diversity, and can accurately generate RNA conformational ensembles that conform to the Boltzmann distribution. At the same time, it can overcome the bottleneck of difficult sampling of excited state and high-energy state conformations, take into account physical rationality, and can efficiently sample RNA diverse structures. It can also be extended to other systems such as DNA and protein-nucleic acid complexes.

[0036] (3) The present invention, DynaRNA, is a method for generating dynamic RNA conformations based on a diffusion model. It belongs to a new application field of artificial intelligence in life sciences. It breaks through the traditional method’s trade-off between sampling efficiency and physical rationality, and provides a new tool and platform for RNA structural biology and computational drug design. It has extremely high scientific research and industrial application value.

[0037] In this invention, Figure 1This is a schematic flowchart of a method for generating an RNA dynamic conformation ensemble based on a deep learning diffusion model, provided by an embodiment of the present invention. Figure 2 This is a graph showing the test results of the tetranucleotide system in an embodiment of the present invention; Figure 3 This is a diagram showing the test results of the HIV inverse activation response (TAR) element system in an embodiment of the present invention.

[0038] This invention addresses the problem of difficulty in resolving current RNA dynamic conformation ensembles by providing an RNA structure generation method called DynaRNA. It employs a denoising diffusion probabilistic model (DDPM) combined with an equivariant graph neural network (EGNN) to directly model the three-dimensional spatial structure of RNA. This method can efficiently generate RNA dynamic conformation ensembles that conform to experimental statistical laws, while also capturing high-energy excited-state conformations, thus fully validating the effectiveness of the invention.

[0039] The development process of this invention can be summarized as follows: by constructing a high-quality RNA structure training dataset, and using coarse-grained processing as model input, the model combines a diffusion model and an isovariant graph neural network. Through "forward diffusion" and "reverse denoising", the two constitute a symmetrical, probability-driven generation loop, ultimately realizing the generation of RNA dynamic conformation ensemble.

[0040] The method described above for generating RNA dynamic conformation ensembles using a deep learning diffusion model specifically includes the following steps: (1) Training set construction: High-quality RNA crystal structures ranging from 5 to 200 nucleotides in length were extracted from the largest and most commonly used biomolecular structure database, PDB (Protein Data Bank). Sections with low resolution, incomplete structures, and non-standard nucleotides containing modifications were removed. Sections containing non-RNA components such as proteins, DNA, water, ions, and small molecules were also screened out, resulting in 6,820 high-quality structures as the training dataset for the model.

[0041] (2) Data coarsening and input To reduce computational costs during model training and thus improve training efficiency, we performed coarse-grained processing of the data. Considering the geometric features and physicochemical properties of RNA molecules, the C4' atom is centrally located within the RNA ribose, phosphate, and base groups. Furthermore, the C4' atom plays a crucial role in maintaining the overall structure and function of the RNA molecule. Therefore, we selected the C4' atom coordinates as the coarse-grained coordinate input for RNA nucleotides. Simultaneously, all structures were standardized to ensure they reside in a unified coordinate system, removing overall translational and rotational degrees of freedom. Each nucleotide was represented as a particle, with its three-dimensional spatial coordinates being the C4' atom coordinates. The RNA sequence was treated as a graph structure: nodes represent particles (C4'), and edges represent the spatial relationships and topological connections between adjacent nucleotides. This graph structure, along with the particle coordinates, was fed into the diffusion model. This coarse-grained processing of the input data minimizes the number of model parameters and training costs while preserving the characteristics of the input data as much as possible, maximizing the model's training efficiency.

[0042] (3) Forward diffusion noise addition process of the model The forward diffusion process involves adding Gaussian noise to gradually perturb the original RNA conformation to an approximately Gaussian noise state, constructing a trajectory from the data distribution to the prior distribution (usually a standard normal distribution). The detailed principle and process of this procedure are as follows: Given the preceding time Data distribution Add Gaussian noise to it Thus, the next moment is obtained. Data distribution This process is defined by the following stochastic differential equation: In the above formula, and These represent the data distribution at the corresponding time points. This represents the noise level at a given moment, and its intensity varies with the number of time steps, as defined by the following formula: In the above formula, and These represent the noise levels at the corresponding times, respectively. It is 0.0001. It is 0.02.

[0043] By combining the above formulas, Gaussian noise can be gradually added to the initial distribution of the input data, thereby constructing the forward diffusion noise addition process of the model.

[0044] (4) Backward Process To achieve the reverse denoising process, this invention trains a translationally and rotationally equivalent graph neural network (EGNN). This EGNN network architecture is specifically designed to follow geometric symmetry (such as translational and rotational equivalence), making it highly suitable for molecular-level or geometric data modeling. Its core function is noise prediction. Details are as follows: Each layer of the EGNN includes node and edge update operations to capture the complex geometric relationships between nucleotides. Node features contain the three-dimensional coordinates and temporal step embedding information of the C4' atom; edge features simultaneously encode molecular connectivity and spatial distance. The hidden dimension of each layer is set to 128, and LayerNorm regularization and the SiLU activation function are used to enhance training stability and nonlinear modeling capabilities. Temporal information is encoded using sine function embedding, commonly used in standard diffusion models. To prevent overfitting, a 0.1 Dropout is added after the output of each layer.

[0045] The network's final output is the predicted noise vector at each time step, dependent on the geometric and graph topological structures. The denoising process is guided by the denoising function learned by the aforementioned equivariant graph neural network. It involves approximating the data distribution with noise predicted by this function through the inverse process of adding noise, and then iteratively generating new samples through back-diffusion. Specifically, this is achieved through the following formula: in, express Distribution of predicted data at time points Represents signal retention rate. Represents the cumulative signal retention rate. The noise represented by the EGNN network model prediction. This represents the Gaussian noise newly introduced during the sampling phase.

[0046] (5) Model training process The goal of model training is to minimize the mean squared error (MSE) between the model's predicted noise and the actual noise. The loss function is shown in Equation (4): in, For loss function, Represents Gaussian noise. The loss function represents the model's predicted noise under the corresponding data distribution. Minimizing this loss function can improve the accuracy of the model's noise prediction as much as possible. During training, all gradients are backpropagated through the Adam optimizer, and a learning rate scheduling strategy is used to improve convergence speed and performance. The initial learning rate is set to 0.0001. To avoid gradient explosion and ensure numerical stability, a gradient pruning mechanism is introduced, and the maximum norm (Max Norm) is limited to 1.0. The DynaRNA model described in this invention is implemented based on PyTorch and the PyTorch-Lightning framework, possessing a good modular structure and distributed training compatibility, facilitating model iteration and expansion. The entire training process is completed on a single NVIDIA 4090D GPU, with a total training time of approximately 14 days.

[0047] (6) Model usage Based on the deep learning generative model DynaRNA trained using the above steps, users can obtain an ensemble of dynamic RNA conformations.

[0048] See Figure 1 As shown in the embodiment of the present invention, a method for generating an RNA dynamic conformation ensemble based on a deep learning diffusion model is provided. The method consists of two steps: first, the RNA structure input by the user is perturbed to make it diverse; then, a denoising process is used to obtain an RNA dynamic conformation ensemble that conforms to the Boltzmann distribution.

[0049] The accuracy of DynaRNA was assessed by dynamic conformational ensemble generation of tetranucleotides, while the efficiency of DynaRNA was assessed by dynamic conformational ensemble generation of the human immunodeficiency virus type 1 transcriptional activation response element system.

[0050] 4.1 Tetranucleotide Tetranucleotides, composed of four nucleotides, are small in size and computationally inexpensive, while also providing abundant experimental data for easy comparison, making them a common case study in computational RNA structure research. Previously, computational structure studies of tetranucleotide RNA primarily relied on molecular dynamics simulations; however, existing simulation methods are prone to generating erroneous intercalation conformations in tetranucleotides, resulting in significant discrepancies with experimental results. In this invention, we used DynaRNA to generate dynamic conformational ensembles for five tetranucleotide systems: AAAA, CCCC, UUUU, CAAU, and GACC, and compared the results with the most commonly used nucleic acid force fields. ff99bsc0χOL3 The conformational ensembles generated by BSFF1 and BSFF2 simulations were compared, and the results are as follows: Figure 2 As shown, where Figure 2A and 2B represent the proportions of erroneous intercalation conformations generated when the initial conformation is the standard experimental conformation and the erroneous intercalation conformation, respectively. The four columns from left to right for each system show the dynamic conformation ensemble results generated using the force fields of OL3, BSFF1, and BSFF2 and DynaRNA. Conventional molecular dynamics simulations can generate a large number of erroneous intercalation structures, and this problem is particularly serious in the initial intercalation conformation. DynaRNA successfully reduced the frequency of erroneous intercalation conformations and its effect is very robust when the initial conformation is an erroneous intercalation structure, while increasing the frequency of reasonable conformations.

[0051] Figure 2 Based on the test results of the tetranucleotide system, by analyzing the base stacking order of each cluster structure, it can be found that DynaRNA is not only less computationally expensive than traditional molecular dynamics simulations, but also more accurate.

[0052] 4.2 Human Immunodeficiency Virus Type 1 Transcription Activation Response Element System Human immunodeficiency virus type 1 (HIV-1) is the pathogenic virus for HIV / AIDS, severely impacting human health. Its transcriptional activation response element (TAR) has been widely recognized as a highly promising therapeutic target in recent years. This element consists of two helical regions connected by a bulge, forming a typical hairpin loop conformation at the apex. Its unique secondary structure has attracted considerable research attention. TARs are non-coding RNA elements located on messenger RNA (mRNA). They can bind small ligands, allowing them to sample several low-abundance excited states (ES) in addition to the dominant ground state (GS). Although these excited states are scarce and short-lived, they play crucial roles in biochemical reaction mechanisms, disease progression, and therapeutic strategy development. However, ES conformations are typically rich in atypical base mispairing and have unfavorable energy, resulting in poor stability and transient nature, significantly limiting their availability for structural characterization. Traditional experimental techniques struggle to capture the excited-state conformation of RNA, while conventional molecular dynamics simulations face challenges such as high energy barriers and low sampling efficiency, making it difficult to effectively obtain the complete conformational space. The DynaRNA method described in this invention provides a novel technical approach to overcome energy barriers and directly explore the set of RNA conformational spaces. Through a bidirectional diffusion mechanism, DynaRNA can achieve efficient conversion and sampling between the excited and ground states of RNA under different initial conformational conditions. Results are as follows... Figure 3As shown, DynaRNA can sample another conformation starting from both the dominant conformation state and the excited conformation state, which fully demonstrates that this method has the ability to overcome the energy barrier of complex conformations.

[0053] Figure 3 Figures 3A and 3B show the test results of the human immunodeficiency virus type 1 transcriptional activation response element system. As shown in Figures 3A and 3B, starting from GS, DynaRNA can effectively generate a conformation set containing ES2; conversely, starting from ES2, the GS conformation can also be sampled. Principal component analysis (PCA) was performed on the conformation sets generated from the two initial states, and the results are as follows: Figure 3 C and Figure 3 As shown in D, two distinct clustering regions are presented, corresponding to GS and ES2 respectively.

[0054] The beneficial effects of this invention: The present invention proposes a method for generating dynamic RNA conformational ensembles based on a deep learning diffusion model. By combining a diffusion model and an isovariant graphical neural network, it can accurately and efficiently model the dynamic structural ensembles of nucleic acid molecules. This method is of great help in studying the relationship between nucleic acid structure and function, and is beneficial to the development of related pharmaceutical industries such as RNA therapy.

[0055] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model, characterized in that, The method includes the following steps: S1. Obtain RNA three-dimensional structure training data; S2. Convert the RNA three-dimensional structure data into an RNA three-dimensional spatial map structure and extract the features of the RNA three-dimensional spatial map structure. S3, RNA three-dimensional spatial graph structural features are input into the diffusion probability model. Noise is added through the forward diffusion process to obtain noisy data. The noisy data is then modeled for denoising based on the equivariant graph neural network to obtain reconstructed data. S4. Train a diffusion probability model based on reconstructed data; S5. Output the conformational ensemble from the diffusion probability model trained on a single conformation.

2. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 1, characterized in that, The specific steps for adding noise through a forward diffusion process to obtain noisy data are as follows: The structural features of the RNA three-dimensional spatial map at the previous time step are obtained, and Gaussian noise is added to them to obtain the structural features of the RNA three-dimensional spatial map at the next time step as the noisy data.

3. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 2, characterized in that, The structural features of the RNA three-dimensional spatial graph include C4' node features and edge features.

4. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 3, characterized in that, The structural features of the RNA three-dimensional spatial map at the next moment are as follows: Where, x t-1 and x t The three-dimensional spatial structure features of RNA at the next and preceding time steps, respectively, β t Represents the noise level of the preceding time step, ∈ t Describe the Gaussian noise of the preceding time step.

5. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 4, characterized in that, The noise level at the preceding time step is: Where, β t ,β0 and β T These represent the noise levels at the preceding time step, time step 0, and the total number of time steps, respectively.

6. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 1, characterized in that, The specific steps for denoising the noisy data using an isotropic graph neural network to obtain the reconstructed data are as follows: The noise vector at each time step is predicted based on the isotropic graph neural network, and the noise-reducing model is used to model the noise-added data based on the predicted noise vector to obtain the reconstructed data.

7. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 6, characterized in that, The reconstructed data at time t-1 is as follows: in, Represents the reconstructed data at time t-1, α t Represents signal retention rate. Represents the cumulative signal retention rate, ∈ θ (x t ,t) represents the noise vector predicted by the equivariant graph neural network, and z represents the Gaussian noise newly introduced during the sampling phase.

8. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 1, characterized in that, The equivariant graph neural network is trained using mean squared error.

9. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 3, characterized in that, The C4' node features are the three-dimensional coordinates and time step embedding information of the C4' atom.

10. The method for generating RNA dynamic conformation ensembles based on a generative artificial intelligence model according to claim 9, characterized in that, Edge features represent the spatial relationships and topological connections between adjacent nucleotides.