Protein-rna complex molecular dynamics trajectory generation method, device, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI INSTITUTE OF SCIENCE & INTELLIGENCE
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]而上述得到替代方案中或者缺乏蛋白质-RNA复合物的动力学生成能力,无方法能生成其连续动力学轨迹;或者缺乏跨模态等变去噪能力,无方法能在统一等变框架下同时处理蛋白质残基和RNA核苷酸;或者缺乏蛋白质-RNA界面的动态建模能力,静态预测无法捕捉界面时序演化,单体方法无法建模跨模态动态耦合
[0015]可见,本申请公开了获取目标蛋白质-RNA复合物的氨基酸序列、核苷酸序列以及静态参考构象;利用跨模态编码器对所述氨基酸序列和所述核苷酸序列进行统一编码,以生成包含蛋白质残基与RNA核苷酸的单表示特征以及包含蛋白质-蛋白质、RNA-RNA和蛋白质-RNA跨模态界面特征的对表示特征;将随机初始化的噪声结构、所述单表示特征、所述对表示特征以及所述静态参考构象输入预设轨迹生成网络,其中,所述噪声结构包含M个轨迹时间帧的蛋白质残基刚体表示和RNA核苷酸刚体表示;在每一个扩散时间步上,对所述噪声结构的所有M个轨迹时间帧并行执行根据当前扩散时间步的噪声结构与所述蛋白质-RNA跨模态界面特征计算跨模态几何偏置信息,并通过不变点注意力机制并利用所述跨模态几何偏置信息、所述对表示特征更新所述单表示特征,以得到第一中间特征;将所述静态参考构象的蛋白质-RNA界面拓扑约束与所述第一中间特征融合,输出第二中间特征;将运动参考特征与所述第二中间特征沿时间维度进行拼接处理,并利用正弦位置编码对拼接后的第二中间特征进行顺序标记,然后沿时间轴对标记顺序的第二中间特征进行自注意力运算,并裁剪出与所述第二中间特征对应的时间帧部分作为第三中间特征;根据所述第三中间特征,分别预测M个轨迹时间帧上的蛋白质残基刚体更新量、RNA核苷酸刚体更新量以及RNA碱基原子坐标更新量,并更新当前扩散时间步的噪声结构的刚体表示与碱基原子坐标,以得到更新后的噪声结构;将满足预设迭代次数条件的更新后的表征M个轨迹时间帧的蛋白质-RNA复合物全原子动力学轨迹的去噪结构进行输出。由此可见,利用跨模态编码器同时编码复合物的蛋白质和RNA序列,生成包含蛋白质-RNA跨模态界面特征的对表示,为等变去噪提供先验;在扩散模型的每一个扩散时间步上,对所有M个轨迹时间帧并行执行不变点注意力,结合从当前噪声结构计算的跨模态几何偏置信息与对表示特征,在各自局部刚体坐标系下更新节点特征,实现蛋白质残基和RNA核苷酸的统一等变去噪,保证了长时程轨迹的旋转平移不变性;将静态参考构象的蛋白质-RNA界面拓扑约束与第一中间特征融合,解决了跨模态界面空间合理性约束缺失的问题;进一步引入运动参考特征,与第二中间特征沿时间维度拼接并施加正弦位置编码,通过时间自注意力聚合历史运动趋势,实现了蛋白质-RNA界面相互作用的时序动态耦合;最终并行预测所有时间帧的蛋白质及RNA刚体更新量和RNA碱基原子坐标更新量,迭代去噪后输出全原子轨迹。
Smart Images

Figure CN122531480A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for generating molecular dynamic trajectories of protein-RNA complexes. Background Technology
[0002] Protein-RNA interactions play a central role in life processes such as gene regulation, translation, and viral replication. Understanding the dynamic conformational changes of protein-RNA complexes is crucial for elucidating their functional mechanisms and designing drugs that target RNA. Currently, the main method for obtaining the dynamic trajectory of protein-RNA complexes is all-atom molecular dynamics (MD) simulation, but its computational cost is extremely high and it is difficult to cover long-term conformational changes.
[0003] To address computational cost issues, generative models based on deep learning have emerged in recent years as an alternative to molecular dynamics (MD) simulations. AlphaFold3 performs only static structure prediction, abandons the Invariant Point Attention (IPA) mechanism, uses non-isotropic atomic diffusion, and cannot generate dynamic trajectories. ATMOS uses an autoregressive SSM, is applicable only to protein-ligand pairs, does not involve RNA, and suffers from long-term error accumulation. AlphaFolding is limited to single-stranded proteins; neither the encoder nor the rigid body representation supports RNA. RNAGenesis is limited to RNA monomers and does not support protein and complex interfaces.
[0004] The aforementioned alternative solutions either lack the ability to generate the dynamics of protein-RNA complexes, and there is no method to generate their continuous dynamic trajectories; or they lack the ability to perform cross-modal isovariant denoising, and there is no method to simultaneously process protein residues and RNA nucleotides within a unified isovariant framework; or they lack the ability to dynamically model the protein-RNA interface, and static predictions cannot capture the temporal evolution of the interface, while single-methods cannot model cross-modal dynamic coupling. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for generating molecular dynamic trajectories of protein-RNA complexes, capable of generating continuous, physically consistent all-atom dynamic trajectories of complexes containing protein and RNA chains through a unified isotropic diffusion architecture. The specific scheme is as follows: In a first aspect, this application discloses a method for generating molecular dynamics trajectories of protein-RNA complexes, comprising: Obtain the amino acid sequence, nucleotide sequence, and static reference conformation of the target protein-RNA complex; The amino acid sequence and the nucleotide sequence are uniformly encoded using a cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features. The randomly initialized noise structure, the single representation feature, the pair representation feature, and the static reference conformation are input into a preset trajectory generation network, wherein the noise structure includes rigid body representations of protein residues and rigid body representations of RNA nucleotides for M trajectory time frames. At each diffusion time step, for all M trajectory time frames of the noise structure, the cross-modal geometric bias information is calculated in parallel based on the noise structure and the protein-RNA cross-modal interface features at the current diffusion time step. The single representation feature is updated using the cross-modal geometric bias information and the pair representation features through an invariant point attention mechanism to obtain the first intermediate feature. The protein-RNA interface topological constraints of the static reference conformation are fused with the first intermediate feature to output the second intermediate feature; The motion reference feature and the second intermediate feature are concatenated along the time dimension, and the concatenated second intermediate feature is sequentially marked using sinusoidal position coding. Then, self-attention operation is performed on the second intermediate features with the marked order along the time axis, and the time frame portion corresponding to the second intermediate feature is cropped out as the third intermediate feature. Based on the third intermediate feature, the rigid body update amount of protein residues, the rigid body update amount of RNA nucleotides, and the update amount of RNA base atom coordinates on M trajectory time frames are predicted respectively, and the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step are updated to obtain the updated noise structure. The denoised structure of the protein-RNA complex all-atom dynamic trajectory of M trajectory time frames, which meets the preset iteration number condition, is output.
[0006] Optionally, the cross-modal encoder is a PairFormer; Accordingly, the method of uniformly encoding the amino acid sequence and the nucleotide sequence using a cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features, includes: The protein residues of the amino acid sequence are encoded using amino acid type tokens to obtain the encoded amino acid sequence. The RNA nucleotides of the nucleotide sequence are encoded using base-type tokens to obtain the encoded nucleotide sequence; The encoded amino acid sequence and the encoded nucleotide sequence are input into the PairFormer so that the protein evolution information of the encoded amino acid sequence can be extracted by multiple sequence alignment, and the RNA structure information of the encoded nucleotide sequence can be extracted by secondary structure prediction. The single representation feature and the pair representation feature are generated through a triangular multiplication update mechanism based on the protein evolution information and the RNA structure information.
[0007] Optionally, calculating the cross-modal geometric bias information based on the noise structure at the current diffusion time step and the protein-RNA cross-modal interface features includes: The relative rigid body transformation between the rigid body of protein residues and the rigid body of RNA nucleotides is calculated based on the noise structure at the current diffusion time step, and the relative distance and relative angle are extracted from the relative rigid body transformation. The relative distance and relative angle are combined with the protein-RNA cross-modal interface features to generate cross-modal geometric bias information.
[0008] Optionally, the step of obtaining the first intermediate feature by using the invariant point attention mechanism and the cross-modal geometric bias information, and updating the single representation feature with the representation feature, includes: The single representation feature of the current diffusion time step is used as the query vector and key vector, the pair representation feature is used as the edge bias, and the cross-modal geometric bias information is used as the geometric bias. In the invariant point attention module, attention weights are calculated based on the query vector, the key vector, the edge bias, and the geometric bias. Based on these attention weights, protein residues and RNA nucleotides that have adjacent or interacting relationships in spatial structure are aggregated to update the single representation features and obtain the first intermediate feature.
[0009] Optionally, fusing the protein-RNA interface topological constraints of the static reference conformation with the first intermediate feature to output a second intermediate feature includes: The reference distances and reference angles between protein residues and RNA nucleotides are extracted from the static reference conformation to construct protein-RNA interface topological constraints. The first intermediate feature is spliced and fused with the protein-RNA interface topological constraints to obtain the second intermediate feature.
[0010] Optionally, before concatenating the motion reference feature and the second intermediate feature along the time dimension, the method further includes: The rigid body difference information between adjacent frames is calculated from the rigid body representation of the noise structure in the historical time frame; wherein, the rigid body difference information includes translation difference and rotation difference; The rigid body difference information is geometrically aggregated to obtain the motion reference features.
[0011] Optionally, the step of predicting the rigid body update amounts of protein residues, RNA nucleotides, and RNA base atom coordinates on M trajectory time frames based on the third intermediate feature, and updating the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step, includes: Based on the third intermediate feature, the rotation update amount of protein residues, the translation update amount of protein residues, the rotation update amount of RNA nucleotides, the translation update amount of RNA nucleotides, and the coordinate update amount of RNA base atoms in the corresponding local rigid body coordinate system are predicted and output on M trajectory time frames respectively. The rigid body representation of the protein residues at the current diffusion time step is updated based on the rotation matrix of the protein residues at the current diffusion time step, the translation vector of the protein residues at the current diffusion time step, the rotation update amount of the protein residues, the translation update amount of the protein residues, and the first rigid body update equation. The rigid body representation of the RNA nucleotide at the current diffusion time step is updated based on the rotation matrix of the RNA nucleotide at the current diffusion time step, the translation vector of the RNA nucleotide at the current diffusion time step, the rotation update amount of the RNA nucleotide, the translation update amount of the RNA nucleotide, and the second rigid body update equation. The global coordinates of the RNA base atoms are updated based on the rigid body representation of the protein residues at the current diffusion time step, the rigid body representation of the RNA nucleotides at the current diffusion time step, the coordinate update amount of the RNA base atoms in the corresponding local rigid body coordinate system, and the atomic coordinate update equation.
[0012] Secondly, this application discloses a protein-RNA complex molecular dynamics trajectory generation device, comprising: The information acquisition module is used to acquire the amino acid sequence, nucleotide sequence, and static reference conformation of the target protein-RNA complex; The encoding module is used to uniformly encode the amino acid sequence and the nucleotide sequence using a cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features. An information input module is used to input a randomly initialized noise structure, the single representation feature, the pair representation feature, and the static reference conformation into a preset trajectory generation network, wherein the noise structure includes rigid body representations of protein residues and rigid body representations of RNA nucleotides for M time frames; The first feature generation module is used to perform parallel execution on all M trajectory time frames of the noise structure at each diffusion time step to calculate cross-modal geometric bias information based on the noise structure and the protein-RNA cross-modal interface features at the current diffusion time step, and update the single representation feature using the cross-modal geometric bias information and the pair representation features through an invariant point attention mechanism to obtain the first intermediate feature; The second feature generation module is used to fuse the protein-RNA interface topological constraints of the static reference conformation with the first intermediate feature and output the second intermediate feature. The third feature generation module is used to concatenate the motion reference feature and the second intermediate feature along the time dimension, and use sinusoidal position coding to sequentially mark the concatenated second intermediate feature. Then, it performs self-attention operation on the second intermediate feature with the marked order along the time axis, and crops out the time frame part corresponding to the second intermediate feature as the third intermediate feature. The update module is used to predict the rigid body update amount of protein residues, the rigid body update amount of RNA nucleotides, and the update amount of RNA base atom coordinates on M trajectory time frames according to the third intermediate feature, and update the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step to obtain the updated noise structure. The trajectory generation module is used to output the denoised structure of the updated protein-RNA complex all-atom dynamic trajectory representing M trajectory time frames that meets the preset iteration number condition.
[0013] Thirdly, this application discloses an electronic device, including: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the aforementioned disclosed method for generating molecular dynamic trajectories of protein-RNA complexes.
[0014] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned method for generating molecular dynamic trajectories of protein-RNA complexes.
[0015] As can be seen, this application discloses the acquisition of the amino acid sequence, nucleotide sequence, and static reference conformation of the target protein-RNA complex; the use of a cross-modal encoder to uniformly encode the amino acid sequence and the nucleotide sequence to generate single representation features containing protein residues and RNA nucleotides, and pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features; inputting a randomly initialized noise structure, the single representation features, the pair representation features, and the static reference conformation into a preset trajectory generation network, wherein the noise structure contains rigid body representations of protein residues and rigid body representations of RNA nucleotides for M trajectory time frames; at each diffusion time step, performing parallel calculations of cross-modal geometric bias information based on the noise structure and the protein-RNA cross-modal interface features of the current diffusion time step for all M trajectory time frames of the noise structure, and using an invariant point attention mechanism and the cross-modal geometric bias information and the pair representations The single representation feature is updated to obtain a first intermediate feature; the protein-RNA interface topological constraint of the static reference conformation is fused with the first intermediate feature to output a second intermediate feature; the motion reference feature and the second intermediate feature are spliced along the time dimension, and the spliced second intermediate feature is sequentially marked using sinusoidal position coding. Then, self-attention operation is performed on the second intermediate features with the marked order along the time axis, and the time frame portion corresponding to the second intermediate feature is cropped as a third intermediate feature; based on the third intermediate feature, the rigid body update amount of protein residues, the rigid body update amount of RNA nucleotides, and the update amount of RNA base atom coordinates on M trajectory time frames are predicted respectively, and the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step are updated to obtain the updated noise structure; the denoised structure of the protein-RNA complex all-atom dynamic trajectory representing M trajectory time frames that meets the preset iteration number condition is output.Therefore, by using a cross-modal encoder to simultaneously encode the protein and RNA sequences of the complex, a pair representation containing protein-RNA cross-modal interface features is generated, providing a priori information for equivariant denoising. At each diffusion time step of the diffusion model, invariant point attention is performed in parallel on all M trajectory time frames. Combining the cross-modal geometric bias information calculated from the current noise structure with the pair representation features, node features are updated in their respective local rigid body coordinate systems, achieving unified equivariant denoising of protein residues and RNA nucleotides, ensuring the rotation and translation invariance of long-term trajectories. The topological constraints of the protein-RNA interface in the static reference conformation are fused with the first intermediate feature, solving the problem of missing spatial rationality constraints of the cross-modal interface. Furthermore, a motion reference feature is introduced, which is concatenated with the second intermediate feature along the time dimension and sinusoidal position encoding is applied. By aggregating historical motion trends through temporal self-attention, the temporal dynamic coupling of protein-RNA interface interactions is achieved. Finally, the rigid body update amounts of protein and RNA and the atom coordinate update amounts of RNA bases are predicted in parallel for all time frames, and the full atom trajectory is output after iterative denoising. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0017] Figure 1 This is a flowchart of a method for generating molecular dynamics trajectories of a protein-RNA complex disclosed in this application; Figure 2 This is a structural diagram of an isotropic diffusion model disclosed in this application; Figure 3 This is a flowchart of a specific method for generating molecular dynamics trajectories of a protein-RNA complex disclosed in this application; Figure 4 This is a schematic diagram of the structure of a protein-RNA complex molecular dynamics trajectory generation device disclosed in this application; Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] Protein-RNA interactions play a central role in life processes such as gene regulation, translation, and viral replication. Understanding the dynamic conformational changes of protein-RNA complexes is crucial for elucidating their functional mechanisms and designing drugs that target RNA. Currently, the main method for obtaining the dynamic trajectory of protein-RNA complexes is all-atom molecular dynamics (MD) simulation, but its computational cost is extremely high and it is difficult to cover long-term conformational changes.
[0020] To address the computational cost issue, generative models based on deep learning have emerged in recent years as an alternative to MD simulations, mainly including: AlphaFold3 uses PairFormer to uniformly encode biomolecules such as proteins, RNA, and ligands, and generates static three-dimensional structures through a diffusion module. However, AlphaFold3 has the following fundamental limitations: (1) Its structural modules explicitly abandon the Invariant Point Attention (IPA) mechanism in AlphaFold2 and instead use non-equivariant atomic coordinates for direct diffusion, resulting in a lack of strict rotation and translation invariance constraints when generating long-term trajectories, and severe energy drift; (2) Its diffusion module is only used for static structure prediction and does not have the ability to model in the time dimension, and cannot generate continuous dynamic trajectories; (3) Its output is a single static conformation, rather than a time-continuous sequence of conformational evolution.
[0021] ATMOS proposes using a state-space model (SSM) and autoregressive decoding to generate biomolecular dynamic trajectories. However, ATMOS has the following limitations: (1) Its technical approach is SSM autoregressive generation, which is completely different from the mathematical framework of diffusion model; (2) Its statement applies only to protein monomers and protein-ligand complexes, and does not involve RNA molecules, let alone protein-RNA complexes; (3) Its generation process adopts frame-by-frame autoregressive decoding instead of one-time denoising generation of the diffusion model. The time consistency depends on the autoregressive condition, and the error accumulates significantly when extrapolating over a long time period.
[0022] AlphaFolding proposes a protein dynamics simulator based on a ternary attention architecture, using an IPA isovariant denoising network to generate continuous dynamic trajectories of single protein chains. However, AlphaFolding has the following limitations: (1) Its technical solution is only designed for single-chain proteins, does not involve RNA molecules, and does not involve cross-modal dynamics modeling of protein-RNA complexes; (2) Its encoder is EvoFormer, which does not have the ability to process RNA nucleotide sequences; (3) Its data representation and denoising network is only adapted to the rigid framework of protein residues and not to the sugar-phosphate backbone framework of RNA nucleotides.
[0023] RNAGenesis proposes a method for RNA monomer conformation generation based on SE(3) flow matching, using the IPA mechanism for isovariant denoising under rigid nucleotide representation. However, RNAGenesis has the following limitations: (1) Its technical solution is only designed for single-stranded RNA, without involving protein molecules, let alone cross-modal interface modeling of protein-RNA complexes; (2) Its encoder only processes RNA sequences and does not have the ability to process protein amino acid sequences; (3) Its target of generation is the equilibrium conformational ensemble of RNA monomers, rather than the continuous dynamic trajectory of protein-RNA complexes.
[0024] However, among the aforementioned alternative methods, AlphaFold3 only performs static structure prediction; ATMOS only covers protein-ligand pairs and not RNA; AlphaFolding only covers single-stranded proteins and not RNA; and RNAGenesis only covers RNA monomers and not proteins. No existing technology can generate continuous dynamic trajectories of protein-RNA complexes. AlphaFold3 uses non-isotropic atomic coordinate diffusion, which cannot guarantee the physical consistency of long-term trajectories. ATMOS uses SSM autoregression and lacks SE(3) isotropic properties. AlphaFolding has IPA isotropic properties but only applies to protein residues; and RNAGenesis has IPA isotropic properties but only applies to RNA nucleotides. No existing technology can simultaneously handle cross-modal denoising of protein residues and RNA nucleotides within a unified isotropic framework.
[0025] Furthermore, the core function of the protein-RNA complex lies in the interfacial interactions (hydrogen bonds, electrostatic interactions, base stacking, etc.) between the protein and RNA. Existing static prediction methods (such as AlphaFold3) cannot capture the temporal evolution of the interface; existing monomer dynamics methods (such as AlphaFolding and RNAGenesis) cannot model the dynamic coupling of cross-modal interfaces.
[0026] To this end, the present invention provides a scheme for generating molecular dynamics trajectories of protein-RNA complexes, which can generate continuous and physically consistent all-atom dynamic trajectories of complexes containing protein chains and RNA chains through a unified isovariant diffusion architecture.
[0027] like Figure 1As shown, the present invention provides a method for generating molecular dynamics trajectories of protein-RNA complexes, comprising: Step S11: Obtain the amino acid sequence, nucleotide sequence, and static reference conformation of the target protein-RNA complex.
[0028] In this embodiment, the amino acid sequence and nucleotide sequence of the target protein-RNA complex are obtained from public bioinformatics databases (such as UniProt, NCBI GenBank, Rfam). Alternatively, the amino acid sequence and nucleotide sequence of the target protein-RNA complex can be obtained by direct input by the user or by gene sequencing technology. In addition, for complexes with unknown sequences, they can be obtained by assembling sequences after high-throughput sequencing. No specific limitation is made in this regard.
[0029] In this embodiment, the static reference conformation is a three-dimensional structure that serves as the spatial anchor and source of geometric constraints throughout the entire dynamic trajectory generation process. The static reference conformation can be determined experimentally, and these structures are usually obtained from protein data banks or nucleic acid databases (NDBs). In the absence of experimental structures, the three-dimensional structure model of the complex predicted by computational methods (such as complex structure prediction tools like AlphaFold3 and RoseTTAFoldNA) can still provide the overall folding topology of the complex and the approximate layout of the protein-RNA interface, even though the prediction may have some errors.
[0030] Step S12: Use a cross-modal encoder to uniformly encode the amino acid sequence and the nucleotide sequence to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features.
[0031] In this embodiment, the cross-modal encoder is a PairFormer. Correspondingly, the step of uniformly encoding the amino acid sequence and the nucleotide sequence using the cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, and pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features, includes: encoding the protein residues of the amino acid sequence with amino acid-type tokens to obtain an encoded amino acid sequence; encoding the RNA nucleotides of the nucleotide sequence with base-type tokens to obtain an encoded nucleotide sequence; inputting the encoded amino acid sequence and the encoded nucleotide sequence into the PairFormer to extract protein evolution information from the encoded amino acid sequence through multiple sequence alignment and extract RNA structure information from the encoded nucleotide sequence through secondary structure prediction; and generating single and pair representation features based on the protein evolution information and the RNA structure information using a triangular multiplication update mechanism.
[0032] Understandably, for the input target protein-RNA complex, each residue in its protein amino acid sequence is mapped to an amino acid-type token (a universal embedding of 20 standard amino acids plus rare amino acids), resulting in the encoded amino acid sequence. Simultaneously, each base in the RNA nucleotide sequence is mapped to a base-type token (a universal embedding of A / U / G / C plus modified bases), resulting in the encoded nucleotide sequence. Both the encoded amino acid and nucleotide sequences are input into PairFormer. In PairFormer, evolutionary information of protein residues is extracted through multiple sequence alignment (MSA), and secondary structure information of RNA nucleotides (e.g., helices, loops, etc.) is extracted through an external secondary structure prediction tool. PairFormer utilizes this information, combined with reference conformational geometric features (e.g., residue-nucleotide Cα-C1' distance and interface angle), and iteratively updates the single and paired representations using a triangular multiplication update mechanism. The output dimension of PairFormer is... Single representation feature ,in, and dimension are The pair represents the features The features represented include three categories: protein-protein, RNA-RNA, and protein-RNA. It's important to note that for the protein-RNA cross-modal interface feature (i,j), where i is the protein residue index and j is the RNA nucleotide index, it is updated as follows: ; in, The single representation feature of residue i The single representation feature of nucleotide j For the reference conformation, the Cα-C1' distance between residue i and nucleotide j From the perspective of the interface.
[0033] This ensures that the encoder simultaneously captures the independent properties of proteins and RNA, as well as the prior knowledge of their interactions, providing the isovariant denoising network with single and pair representations containing biological information.
[0034] Step S13: Input the randomly initialized noise structure, the single representation feature, the pair representation feature, and the static reference conformation into the preset trajectory generation network, wherein the noise structure includes rigid body representations of protein residues and rigid body representations of RNA nucleotides for M trajectory time frames.
[0035] In this embodiment, the preset trajectory generation network is an isovariant diffusion model, such as... Figure 2 As shown, its input information includes a randomly initialized noise structure, single representation features and pair representation features output by PairFormer, and a static reference conformation. Based on the set trajectory length M (i.e., the number of time frames), a noise rigid body is independently sampled for each time frame to form the initial noise structure. For the i-th protein residue in the s-th time frame (1≤s≤M), the protein residue rigid body representation... ,in This is a rotation matrix of protein residues. Let be the translation vector of the protein residues. For the i-th RNA nucleotide in the s-th time frame, the rigid body representation of the RNA nucleotide is: ,in, This is a rotation matrix of RNA nucleotides. This is the translation vector of the RNA nucleotide, with the origin at the C4' atom of the sugar ring, the x-axis along the C4'-C3' direction, the y-axis in the C4'-C3'-O4' plane, and the z-axis determined by the right-hand rule. The coordinates of the base atoms are... Represented in a local rigid body coordinate system, through local coordinate transformation Mapped to global coordinates.
[0036] Furthermore, for ternary complexes containing small molecule ligands, the ligand atoms are represented by a generic token without a predefined chemical type, and their chemical identity is inferred from the spatial context through the pairwise attention mechanism of PairFormer. The ligand atom coordinates are updated with coordinate-level denoising in the global coordinate system.
[0037] The isovariant diffusion model is constructed through forward diffusion and backward generation. For the forward diffusion process, noise is simultaneously applied to the rigid structure of all M time frames in the protein-RNA complex trajectory, and the translation component X adopts a variance-preserving stochastic differential equation (VP-SDE). ; in, , , .
[0038] The rotation component R is injected with noise using the variance explosion stochastic differential equation (VE-SDE), specifically expressed as follows: ; Where g(t) can be obtained through Obtained, and others , , .
[0039] The protein side chain torsion angle is predicted for each residue based on the clean structure of the protein after noise reduction, while the RNA base atom coordinates are represented as VP-SDE (in the local coordinate system).
[0040] Then, each sub-module is constructed for updating the rigid body and coordinates. The model training strategy and updates are as follows: The total loss function includes a denoising score matching loss to supervise the model's denoising ability for rotational and translational components, ensuring the generated structure conforms to the data distribution; a denoising angle loss to supervise the model's denoising ability for protein residues and RNA nucleotide rigid bodies; a torsion angle loss to calculate the difference between the predicted torsion angle and the true value, supervising the accuracy of side chain conformations; a base coordinate loss to supervise the accuracy of RNA base atom coordinates; and an auxiliary geometry loss. This auxiliary geometry loss is introduced to further improve the stereochemical rationality of the structure, penalizing non-physical atomic spacing (such as peptide bond breakage) and atomic collisions. Specifically, this loss is only enabled in stages with small diffusion time steps (e.g., t < 0.25) to fine-tune the structure quality in the final stage of generation, avoiding interference with early global structure generation and penalizing non-physical atomic spacing and collisions.
[0041] Step S14: At each diffusion time step, perform parallel calculation of cross-modal geometric bias information based on the noise structure and the protein-RNA cross-modal interface features at the current diffusion time step for all M trajectory time frames of the noise structure, and update the single representation feature using the cross-modal geometric bias information and the pair representation features through an invariant point attention mechanism to obtain the first intermediate feature.
[0042] In this embodiment, based on a trained scoring network, the dynamic trajectory of a protein, conforming to physical laws and possessing spatiotemporal coherence, is gradually recovered from a disordered noise distribution by solving inverse stochastic differential equations. First, the relative rigid body transformation between the rigid bodies of protein residues and RNA nucleotides is calculated based on the noise structure at the current diffusion time step, and the relative distance and relative angle are extracted from this transformation. The relative distance and relative angle are then combined with the protein-RNA cross-modal interface features to generate cross-modal geometric bias information. It can be understood that for protein residue i and RNA nucleotide j, the relative rigid body transformation in the local coordinate system is calculated as follows: ; in, for The inverse rigid body transformation represents the transformation of a point in the global coordinate system to the local coordinate system of protein residue i. Let be the rigid body transformation of the i-th protein residue in the current noisy structure, belonging to the SE(3) group, including the rotation matrix and translation vector, representing the position and orientation of the residue in the global coordinate system. The j-th RNA nucleotide represents the rigid body transformation in the current noise structure. It also belongs to the SE(3) group and consists of a rotation matrix and a translation vector, representing the position and orientation of the nucleotide in the global coordinate system. The compound operator for rigid body transformation indicates that the transformation on the right-hand side is applied first, followed by the transformation on the left-hand side. Therefore, the specific meaning of this formula is to apply the transformation on the right-hand side first. Map the local coordinates of RNA nucleotides to global coordinates, and then apply... Mapping global coordinates to the local coordinate system of protein residue i, the final result This represents a relative rigid body transformation, describing the position and orientation of RNA nucleotide j in the local coordinate system of protein residue i (i.e., the relative transformation from the local coordinate system of residue i to the local coordinate system of nucleotide j).
[0043] Then, from relative rigid body transformation The relative translation distance and relative rotation angle are extracted. These geometric quantities are the cross-modal geometric bias information, which are used as geometric bias terms in the attention calculation in the subsequent invariant point attention module to enhance the model's perception of the protein-RNA spatial relationship.
[0044] Furthermore, the single representation feature of the current diffusion time step is used as the query vector and the key vector, the pair representation feature is used as the edge bias, and the cross-modal geometric bias information is used as the geometric bias. In the invariant point attention module, attention weights are calculated based on the query vector, the key vector, the edge bias, and the geometric bias. Based on the attention weights, protein residues and RNA nucleotides that have adjacent or interacting relationships in spatial structure are aggregated to update the single representation feature and obtain the first intermediate feature.
[0045] It is understandable that in the reverse generation process of the isovariant diffusion model, the scoring function is learned by training a neural network. By training a neural network to learn a scoring function, during the inference phase, the model starts from randomly sampled Gaussian noise tracks and uses the learned scoring function to solve inverse stochastic differential equations, gradually denoising and recovering the dynamic trajectory of the protein-RNA complex that conforms to physical laws. Specifically, in the inverse generation process, the computational unit used is an isovariant denoising network, which includes L layers of isovariant denoising modules (in this embodiment, L=4, hidden dimension C=256, and the total number of parameters is approximately 23M). Each layer contains three sequential sub-modules. The first sub-module is a local geometric isovariant denoising sub-module, whose input is the rigid body of the currently noisy protein residues. RNA nucleotide rigid body The single representation feature is used as node feature V, and the pair representation feature is used as edge feature Z. After the above input information is input to the local geometric equivariance denoising submodule, the submodule calculates the cross-modal distance bias weight based on the 3D spatial information provided by the rigid body structure in the local coordinate system of protein residues and the local coordinate system of RNA nucleotides. At the same time, it calculates the feature weight based on the node feature, and then updates the node feature based on the distance bias weight and the feature weight to obtain the first intermediate feature. This process ensures that the feature update satisfies the SE(3) rotation and translation invariance.
[0046] Step S15: Fuse the protein-RNA interface topological constraints of the static reference conformation with the first intermediate feature to output the second intermediate feature.
[0047] In this embodiment, reference distances and angles between protein residues and RNA nucleotides are extracted from the static reference conformation to construct protein-RNA interface topological constraints. The first intermediate feature is then spliced and fused with the protein-RNA interface topological constraints to obtain a second intermediate feature. It is understood that the features of the initial reference conformation... (shape is) Noise structure characteristics at the current time step (shape is () The input is concatenated with the local geometric equivariance denoising submodule, followed by the next-level global topological constraint injection submodule, which employs a self-attention mechanism. The reference features are copied M times to obtain... Then and By concatenating along the feature dimension C, a joint feature representation containing reference information is constructed. This process injects global topological constraints (such as interface distance and relative orientation) of the protein-RNA interface in the reference conformation into the first intermediate feature to obtain the second intermediate feature. The calculation formula is as follows: .
[0048] Step S16: The motion reference feature and the second intermediate feature are concatenated along the time dimension, and the concatenated second intermediate feature is sequentially marked using sinusoidal position coding. Then, self-attention operation is performed on the second intermediate features with the marked order along the time axis, and the time frame portion corresponding to the second intermediate feature is cropped out as the third intermediate feature.
[0049] In this embodiment, before concatenating the motion reference feature and the second intermediate feature along the time dimension, the method further includes: calculating rigid body difference information between adjacent frames from the rigid body representation of the noise structure in historical time frames; wherein the rigid body difference information includes translation difference and rotation difference; and aggregating the rigid body difference information using geometric features to obtain the motion reference feature. It can be understood that, for the current diffusion time step, the rigid body representation of the noise structure of the generated historical time frames (usually several consecutive frames prior to the current time frame) is obtained. Suppose there are Mmot historical frames, each containing a rigid body of protein residues and a rigid body of RNA nucleotides. Then, the rigid body difference information between adjacent frames is calculated, i.e., the translation difference (displacement of the centroid between adjacent frames) and the rotation difference (relative rotation matrix between adjacent frames) are calculated. All translation and rotation differences between adjacent frames are arranged in chronological order to form the original difference sequence. This original difference sequence is input into an invariant point attention module with the same structure as the local geometric equivariance denoising submodule for geometric feature aggregation. This IPA module takes the rigid body differences in the difference sequence as input (considered as moving rigid bodies) and combines them with the corresponding node features (single representation features of historical frames), outputting a geometrically consistent feature vector, which is the motion reference feature. The shape is ( ).
[0050] Then, the second intermediate feature and motion reference features Input the concatenated temporal correlation modeling submodule and construct joint features along the time axis. Its shape is By injecting temporal location information using sinusoidal positional encoding, and then calculating self-attention along the time axis to obtain updated features. In this invention, the self-attention output is pruned, retaining only the features corresponding to the input, resulting in a third intermediate feature output. The shape is (M×N×C). This module aggregates motion information from adjacent time frames to ensure that the trajectory is smooth and continuous in the time dimension.
[0051] Step S17: Based on the third intermediate feature, predict the rigid body update amount of protein residues, the rigid body update amount of RNA nucleotides, and the update amount of RNA base atom coordinates on M trajectory time frames respectively, and update the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step to obtain the updated noise structure.
[0052] In this embodiment, based on the third intermediate feature, the rotation update amount of protein residues, the translation update amount of protein residues, the rotation update amount of RNA nucleotides, the translation update amount of RNA nucleotides, and the coordinate update amount of RNA base atoms in the corresponding local rigid body coordinate system are predicted and output for M trajectory time frames respectively. It can be understood that the third intermediate feature after processing by the above sub-modules... The rigid body update of protein residues is predicted and output using a multilayer perceptron (MLP). and RNA nucleotide rigid body turnover and Protein side chain twist angle update α, RNA base atom coordinate update (In the local coordinate system).
[0053] Furthermore, the rigid body representation of the protein residues at the current diffusion time step is updated based on the rotation matrix, translation vector, rotation update amount, translation update amount, and first rigid body update equation of the protein residues at the current diffusion time step; the rigid body representation of the RNA nucleotides at the current diffusion time step is updated based on the rotation matrix, translation vector, rotation update amount, translation update amount, and second rigid body update equation of the RNA nucleotides at the current diffusion time step; and the global coordinates of the RNA base atoms are updated based on the rigid body representation of the protein residues, the rigid body representation of the RNA nucleotides at the current diffusion time step, the coordinate update amount of the RNA base atoms in the corresponding local rigid body coordinate system, and the atomic coordinate update equation. It can be understood that the first rigid body update equation... The rigid body representation of protein residues at the current diffusion time step is updated with the corresponding update amount; this is achieved through the second rigid body update equation. The RNA nucleotide rigid body representation at the current diffusion time step is updated with the corresponding update amount. Based on the RNA base atom coordinate update amount and the RNA base atom coordinates in the local rigid body coordinate system, the local base coordinates x are updated using the atomic coordinate update equation. base Multiply by the rotation matrix R, and add This allows us to obtain the global three-dimensional coordinates of the base atom. After all the above information has been updated, the noise structure at the current diffusion time step can be obtained.
[0054] Step S18: Output the denoised structure of the updated protein-RNA complex all-atom dynamic trajectory representing M trajectory time frames that meets the preset iteration number condition.
[0055] In this embodiment, the diffusion model requires a fixed number of steps (e.g., 1000 steps) of reverse process. When the number of steps reaches a preset value (t=0), the iteration stops. At this point, the noise structure has been sufficiently denoised and is theoretically consistent with the distribution of real data, becoming the desired full-atom structure of the protein-RNA complex for the final output. This output structure contains all M time frames, each of which is a complete three-dimensional conformation of the protein-RNA complex (containing the full-atom coordinates of all protein residues and RNA nucleotides). These frames are arranged in chronological order, forming a dynamic trajectory.
[0056] like Figure 3 The diagram shows the end-to-end flow of the protein-RNA complex dynamic trajectory generation system. First, based on the set trajectory length M, an initial noise trajectory is sampled in the SE(3) space. Specifically, based on the input reference complex, features are extracted using PairFormer, and based on the set trajectory length (time frame number M), an initial noise trajectory is sampled in the SE(3) space. Wherein, the translation component (X): from the standard Gaussian distribution Sampling, rotation component (R): Sampling from a uniform distribution on SO(3).
[0057] Reverse SDE solution: Perform reverse iteration from 1 to 0 at discrete time step t.
[0058] Score prediction: At each denoising step t, the current noisy trajectory is predicted. Sequence features and reference structural features are input into the diffusion model.
[0059] Structural correction: The network uses a spatial attention module to inject the geometric constraints of the reference structure into the current step, and a temporal attention module to aggregate the motion trends of adjacent frames, calculating the estimated value of the scoring function at the current time (i.e., the denoised gradient). ).
[0060] State update: For translation components, updates are performed according to the inverse equation of VP-SDE (variance preservation), gradually shrinking the atomic coordinate distribution. For rotation components, updates are performed according to the inverse equation of VE-SDE (variance explosion), gradually orienting the rigid body direction.
[0061] When the iteration reaches t≈0, the output is a continuous three-dimensional structure of S frames after denoising, containing the complete full-atom coordinates of proteins and RNA.
[0062] Therefore, by using a cross-modal encoder to simultaneously encode the protein and RNA sequences of the complex, a pair representation containing protein-RNA cross-modal interface features is generated, providing a priori information for equivariant denoising. At each diffusion time step of the diffusion model, invariant point attention is performed in parallel on all M trajectory time frames. Combining the cross-modal geometric bias information calculated from the current noise structure with the pair representation features, node features are updated in their respective local rigid body coordinate systems, achieving unified equivariant denoising of protein residues and RNA nucleotides, ensuring the rotation and translation invariance of long-term trajectories. The topological constraints of the protein-RNA interface in the static reference conformation are fused with the first intermediate feature, solving the problem of missing spatial rationality constraints of the cross-modal interface. Furthermore, a motion reference feature is introduced, which is concatenated with the second intermediate feature along the time dimension and sinusoidal position encoding is applied. By aggregating historical motion trends through temporal self-attention, the temporal dynamic coupling of protein-RNA interface interactions is achieved. Finally, the rigid body update amounts of protein and RNA and the atom coordinate update amounts of RNA bases are predicted in parallel for all time frames, and the full atom trajectory is output after iterative denoising.
[0063] like Figure 4 As shown, the present invention also discloses a protein-RNA complex molecular dynamics trajectory generation device, comprising: Information acquisition module 11 is used to acquire the amino acid sequence, nucleotide sequence and static reference conformation of the target protein-RNA complex; Encoding module 12 is used to uniformly encode the amino acid sequence and the nucleotide sequence using a cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA and protein-RNA cross-modal interface features; Information input module 13 is used to input the randomly initialized noise structure, the single representation feature, the pair representation feature and the static reference conformation into the preset trajectory generation network, wherein the noise structure includes rigid body representations of protein residues and rigid body representations of RNA nucleotides for M time frames; The first feature generation module 14 is used to perform parallel execution of cross-modal geometric bias information calculation based on the noise structure and the protein-RNA cross-modal interface features of the noise structure at the current diffusion time step for all M trajectory time frames of the noise structure, and to update the single representation feature by using the cross-modal geometric bias information and the pair representation features through an invariant point attention mechanism to obtain the first intermediate feature; The second feature generation module 15 is used to fuse the protein-RNA interface topological constraints of the static reference conformation with the first intermediate feature and output the second intermediate feature. The third feature generation module 16 is used to concatenate the motion reference feature and the second intermediate feature along the time dimension, and use sinusoidal position coding to sequentially mark the concatenated second intermediate feature. Then, it performs self-attention operation on the second intermediate feature with the marked order along the time axis, and cuts out the time frame part corresponding to the second intermediate feature as the third intermediate feature. The update module 17 is used to predict the rigid body update amount of protein residues, the rigid body update amount of RNA nucleotides, and the update amount of RNA base atom coordinates on M trajectory time frames according to the third intermediate feature, and update the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step to obtain the updated noise structure. The trajectory generation module 18 is used to output the denoised structure of the updated protein-RNA complex all-atom dynamic trajectory representing M trajectory time frames that meets the preset iteration number condition.
[0064] This demonstrates that by uniformly processing protein and RNA sequences through a cross-modal coding module, a pair representation containing interface features is generated, providing prior information for subsequent denoising. The first feature generation module updates features across all time frames in parallel within a local coordinate system, ensuring rotation and translation invariance. The second feature generation module introduces interface constraints from a static reference conformation to avoid structural drift during long-term generation. The third feature generation module utilizes motion reference features and temporal self-attention to maintain continuous and smooth motion between frames. The entire iterative denoising process generates microsecond-level all-atom trajectories in just a few seconds to tens of seconds, three to four orders of magnitude faster than traditional molecular dynamics simulations, while maintaining energy drift below 1.0 kcal / mol, collision rate below 1%, and interface hydrogen bond retention rate highly correlated with the real simulation. In short, this device achieves, for the first time, long-term dynamic simulation of protein-RNA complexes, combining physical accuracy and computational efficiency.
[0065] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0066] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the protein-RNA complex molecular dynamics trajectory generation method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0067] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0068] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0069] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0070] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. It can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the protein-RNA complex molecular dynamics trajectory generation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.
[0071] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned method for generating molecular dynamic trajectories of protein-RNA complexes. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0072] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0073] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs (Compact Disc-Read Only Memory), or any other form of storage medium known in the art.
[0074] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0075] The solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for generating molecular dynamics trajectories of protein-RNA complexes, characterized in that, include: Obtain the amino acid sequence, nucleotide sequence, and static reference conformation of the target protein-RNA complex; The amino acid sequence and the nucleotide sequence are uniformly encoded using a cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features. The randomly initialized noise structure, the single representation feature, the pair representation feature, and the static reference conformation are input into a preset trajectory generation network, wherein the noise structure includes rigid body representations of protein residues and rigid body representations of RNA nucleotides for M trajectory time frames. At each diffusion time step, for all M trajectory time frames of the noise structure, the cross-modal geometric bias information is calculated in parallel based on the noise structure and the protein-RNA cross-modal interface features at the current diffusion time step. The single representation feature is updated using the cross-modal geometric bias information and the pair representation features through an invariant point attention mechanism to obtain the first intermediate feature. The protein-RNA interface topological constraints of the static reference conformation are fused with the first intermediate feature to output the second intermediate feature; The motion reference feature and the second intermediate feature are concatenated along the time dimension, and the concatenated second intermediate feature is sequentially marked using sinusoidal position coding. Then, self-attention operation is performed on the second intermediate features with the marked order along the time axis, and the time frame portion corresponding to the second intermediate feature is cropped out as the third intermediate feature. Based on the third intermediate feature, the rigid body update amount of protein residues, the rigid body update amount of RNA nucleotides, and the update amount of RNA base atom coordinates on M trajectory time frames are predicted respectively, and the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step are updated to obtain the updated noise structure. The denoised structure of the protein-RNA complex all-atom dynamic trajectory of M trajectory time frames, which meets the preset iteration number condition, is output.
2. The method for generating molecular dynamics trajectories of protein-RNA complexes according to claim 1, characterized in that, The cross-modal encoder is a PairFormer; Accordingly, the method of uniformly encoding the amino acid sequence and the nucleotide sequence using a cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features, includes: The protein residues of the amino acid sequence are encoded using amino acid type tokens to obtain the encoded amino acid sequence. The RNA nucleotides of the nucleotide sequence are encoded using base-type tokens to obtain the encoded nucleotide sequence; The encoded amino acid sequence and the encoded nucleotide sequence are input into the PairFormer so that the protein evolution information of the encoded amino acid sequence can be extracted by multiple sequence alignment, and the RNA structure information of the encoded nucleotide sequence can be extracted by secondary structure prediction. The single representation feature and the pair representation feature are generated through a triangular multiplication update mechanism based on the protein evolution information and the RNA structure information.
3. The method for generating molecular dynamics trajectories of protein-RNA complexes according to claim 1, characterized in that, The calculation of cross-modal geometric bias information based on the noise structure at the current diffusion time step and the protein-RNA cross-modal interface features includes: The relative rigid body transformation between the rigid body of protein residues and the rigid body of RNA nucleotides is calculated based on the noise structure at the current diffusion time step, and the relative distance and relative angle are extracted from the relative rigid body transformation. The relative distance and relative angle are combined with the protein-RNA cross-modal interface features to generate cross-modal geometric bias information.
4. The method for generating molecular dynamics trajectories of protein-RNA complexes according to claim 1, characterized in that, The step of updating the single representation feature using the invariant point attention mechanism and the cross-modal geometric bias information to obtain the first intermediate feature includes: The single representation feature of the current diffusion time step is used as the query vector and key vector, the pair representation feature is used as the edge bias, and the cross-modal geometric bias information is used as the geometric bias. In the invariant point attention module, attention weights are calculated based on the query vector, the key vector, the edge bias, and the geometric bias. Based on these attention weights, protein residues and RNA nucleotides that have adjacent or interacting relationships in spatial structure are aggregated to update the single representation features and obtain the first intermediate feature.
5. The method for generating molecular dynamics trajectories of protein-RNA complexes according to claim 1, characterized in that, The step of fusing the protein-RNA interface topological constraints of the static reference conformation with the first intermediate feature to output a second intermediate feature includes: The reference distances and reference angles between protein residues and RNA nucleotides are extracted from the static reference conformation to construct protein-RNA interface topological constraints. The first intermediate feature is spliced and fused with the protein-RNA interface topological constraints to obtain the second intermediate feature.
6. The method for generating molecular dynamics trajectories of protein-RNA complexes according to claim 1, characterized in that, Before concatenating the motion reference feature and the second intermediate feature along the time dimension, the process further includes: The rigid body difference information between adjacent frames is calculated from the rigid body representation of the noise structure in the historical time frame; wherein, the rigid body difference information includes translation difference and rotation difference; The rigid body difference information is geometrically aggregated to obtain the motion reference features.
7. The method for generating molecular dynamics trajectories of protein-RNA complexes according to claim 1, characterized in that, The step of predicting the rigid body update amounts of protein residues, RNA nucleotides, and RNA base atom coordinates on M trajectory time frames based on the third intermediate feature, and updating the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step, includes: Based on the third intermediate feature, the rotation update amount of protein residues, the translation update amount of protein residues, the rotation update amount of RNA nucleotides, the translation update amount of RNA nucleotides, and the coordinate update amount of RNA base atoms in the corresponding local rigid body coordinate system are predicted and output on M trajectory time frames respectively. The rigid body representation of the protein residues at the current diffusion time step is updated based on the rotation matrix of the protein residues at the current diffusion time step, the translation vector of the protein residues at the current diffusion time step, the rotation update amount of the protein residues, the translation update amount of the protein residues, and the first rigid body update equation. The rigid body representation of the RNA nucleotide at the current diffusion time step is updated based on the rotation matrix of the RNA nucleotide at the current diffusion time step, the translation vector of the RNA nucleotide at the current diffusion time step, the rotation update amount of the RNA nucleotide, the translation update amount of the RNA nucleotide, and the second rigid body update equation. The global coordinates of the RNA base atoms are updated based on the rigid body representation of the protein residues at the current diffusion time step, the rigid body representation of the RNA nucleotides at the current diffusion time step, the coordinate update amount of the RNA base atoms in the corresponding local rigid body coordinate system, and the atomic coordinate update equation.
8. A device for generating molecular dynamics trajectories of protein-RNA complexes, characterized in that, include: The information acquisition module is used to acquire the amino acid sequence, nucleotide sequence, and static reference conformation of the target protein-RNA complex; The encoding module is used to uniformly encode the amino acid sequence and the nucleotide sequence using a cross-modal encoder to generate single representation features containing protein residues and RNA nucleotides, as well as pair representation features containing protein-protein, RNA-RNA, and protein-RNA cross-modal interface features. An information input module is used to input a randomly initialized noise structure, the single representation feature, the pair representation feature, and the static reference conformation into a preset trajectory generation network, wherein the noise structure includes rigid body representations of protein residues and rigid body representations of RNA nucleotides for M time frames; The first feature generation module is used to perform parallel execution on all M trajectory time frames of the noise structure at each diffusion time step to calculate cross-modal geometric bias information based on the noise structure and the protein-RNA cross-modal interface features at the current diffusion time step, and update the single representation feature using the cross-modal geometric bias information and the pair representation features through an invariant point attention mechanism to obtain the first intermediate feature; The second feature generation module is used to fuse the protein-RNA interface topological constraints of the static reference conformation with the first intermediate feature and output the second intermediate feature. The third feature generation module is used to concatenate the motion reference feature and the second intermediate feature along the time dimension, and use sinusoidal position coding to sequentially mark the concatenated second intermediate feature. Then, it performs self-attention operation on the second intermediate feature with the marked order along the time axis, and crops out the time frame part corresponding to the second intermediate feature as the third intermediate feature. The update module is used to predict the rigid body update amount of protein residues, the rigid body update amount of RNA nucleotides, and the update amount of RNA base atom coordinates on M trajectory time frames according to the third intermediate feature, and update the rigid body representation and base atom coordinates of the noise structure at the current diffusion time step to obtain the updated noise structure. The trajectory generation module is used to output the denoised structure of the updated protein-RNA complex all-atom dynamic trajectory representing M trajectory time frames that meets the preset iteration number condition.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method for generating molecular dynamic trajectories of protein-RNA complexes as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when executed by a processor, the computer program implements the steps of the method for generating molecular dynamic trajectories of protein-RNA complexes as described in any one of claims 1 to 7.