Quaternion rectification method for protein skeleton generation task
Through the quaternary rectification method, Gaussian distribution and exponential spherical linear interpolation are used to solve the shortcomings of the existing protein generation methods in rotational characterization, and efficient and high-quality protein backbone generation, especially the generation of long-chain proteins.
Patent Information
- Application Number
- CN202510374576.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-08
AI Technical Summary
Existing protein generation methods do not fully utilize the advantages of manifold learning and numerical stability of quaternions in rotational characterization, resulting in low generation quality and high computational complexity, especially poor production effect of long-chain proteins.
Using the quaternary rectification method, by modeling the residues of the protein skeleton structure into frames containing translation and rotation, using Gaussian distribution and isotropic Gaussian distribution sampling noise, combining exponential spherical linear interpolation and Hamiltonian multiplication, loss function is constructed for model training, and efficient and high-quality protein skeleton generation is achieved.
Significantly reduce the number of inference steps and time, generate high-quality protein backbones, and especially show excellent designability in long-chain protein production.
Smart Images

Figure CN120279977A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and more particularly, to a quaternion rectification method for protein backbone generation tasks. Background Art
[0002] In the field of biomedicine, de novo protein design aims to design proteins with specific properties or functions from scratch, which has important applications in fields such as biocatalysis and drug development. For example, designing novel enzymes for biocatalytic reactions or developing innovative drugs for diseases. However, the protein design space is extremely large, making this task extremely challenging. Therefore, the mainstream de novo protein design strategy takes protein backbone generation (i.e., generating a three-dimensional protein structure without side chains) as the core step, and its generation quality directly determines the rationality and basic properties of the designed protein. For the problem of protein backbone generation, researchers have proposed a variety of deep generative models, especially methods based on diffusion models and flow matching models. However, the designability (a key indicator to measure the generation quality) of the protein backbones generated by existing models is generally low, especially for the generation of long-chain proteins. In addition, they usually require a large number of sampling steps to generate results, resulting in high computational complexity and long inference time. The above defects in generation quality and computational efficiency limit the potential of these models in actual large-scale applications.
[0003] Methods based on diffusion models: FrameDiff generates local translation and rotation through two independent diffusion processes respectively; RFDiffusion uses a large pre-trained model as the backbone network to improve the generation quality; Genie represents the protein structure in an asymmetric manner in the forward and reverse processes; Genie2 extends Genie through architectural innovation and large-scale data augmentation, enabling it to capture a larger and more diverse protein structure space.
[0004] Methods based on flow matching models: FrameFlow and FoldFlow construct flows in the SO(3) space using geodesic interpolation of rotation matrices; GAFL further uses the Clifford frame attention module to improve the accuracy of the model in generating protein secondary structures.
[0005] In terms of three-dimensional rotation representation, quaternion algebra is widely used in the fields of computer graphics and robotics due to its compactness, computational efficiency, and the property of avoiding gimbal lock. The axis-angle rotation can be converted into a unit quaternion through exponential mapping, and spherical linear interpolation (SLERP) is used to construct a smooth trajectory on the manifold. In recent years, quaternion technology has been extended to scientific tasks such as molecular conformation modeling and molecular generation, but its application in protein generation is still blank.
[0006] Existing protein generation methods mostly use rotation matrices for rotation representation, which does not fully utilize the advantages of quaternions in manifold learning and numerical stability. In addition, the traditional diffusion model has many sampling steps and long inference time, while the flow matching model has crossover sampling paths, requiring a large number of steps to ensure generation quality. Although rectification technology can achieve non-crossover path generation by retraining the model through noise samples, its application on the three-dimensional rotation manifold SO(3) has not yet been explored. Therefore, how to combine quaternion algebra and rectification technology to achieve efficient and high-quality protein skeleton generation is a key issue that needs to be solved urgently. Summary of the invention
[0007] The purpose of the embodiments of the present disclosure is to provide a quaternion rectification method for protein skeleton generation tasks.
[0008] 2. In a general aspect, a quaternion rectification method for protein skeleton generation task is provided, comprising five steps:
[0009] The first step is to initialize the model based on the known protein dataset D Modeling the residues of the protein backbone structure in the dataset D as a frame including translation and rotation;
[0010] The second step is to sample the local transformation from the noise distribution, that is in is the translation noise sampled from a Gaussian distribution, is the rotation noise sampled from the isotropic Gaussian distribution on SO(3). The sampled local transformation is randomly paired with the local transformation corresponding to the protein skeleton in the dataset D and the interpolation trajectory between the local transformations is calculated, where the translation trajectory is obtained by linear interpolation and the rotation trajectory is calculated by exponential spherical linear interpolation:
[0011] x t =(1-t)x0+tx1
[0012]
[0013] The translation transformation is represented by x, and the rotation transformation is represented by the unit quaternion q; x0 and q0 are the translation and rotation transformations at time 0, corresponding to the local transformation of the sampled noise; x1 and q1 are the translation and rotation transformations at time 1, corresponding to the local transformation of the paired protein backbone in the dataset D; represents Hamiltonian multiplication; exp(·) represents the conversion of the axis-angle representation of the rotation into a unit quaternion, and log(·) represents its inverse transformation; the formula in the first row calculates the linear interpolation of the translation transformation, and the formula in the second row calculates the exponential spherical linear interpolation of the rotation transformation, respectively obtaining the translation x at time t t and rotation q t。
[0014] Step 3: During the training process, use the interpolation trajectory and the model to construct a loss function to update the weight matrix of the model . The specific final loss function used is
[0015]
[0016] where is the loss function term for translational transformation; is the loss function term for rotational transformation; is the loss function term for reducing physical conflicts in the generated skeleton structure, α is 's weight, ∈ is the threshold of the indicator function 1, and t ∈ [0, 1] is the sampled time step.
[0017] Step 4: Use the trained model to gradually generate protein backbone samples starting from the local noise transformation. The specific implementation of the gradual update is as follows: Given the local transformation T t , including the translational transformation x t and the rotational transformation q t , the model predicts the translation and rotation at the end time. This process is expressed as:
[0018]
[0019] where T t is the local transformation at time t; T θ,1 is the local transformation at the end time predicted by the model , including the translational transformation x θ,1 and the rotational transformation q θ,1 . Then calculate the translational velocity υ θ,t and the rotational angular velocity ω θ,t from time t to the end time:
[0020]
[0021] For the sampling process with a time step interval of Δt, use the Euler method to achieve the gradual update of the translational transformation and the rotational transformation:
[0022] x t+△t = x t + υ θ,t ·△t
[0023]
[0024] x t and q tis the translation transformation and rotation transformation at time t; x t+Δt and q t+Δt are the translation transformation and rotation transformation after a time interval Δt; υ θ,t and ω θ,t are to calculate the translation speed and rotational angular velocity from time t to the end time; denotes Hamilton multiplication; exp(·) denotes converting the axis - angle representation of rotation into a unit quaternion; γ is a hyperparameter controlling sampling acceleration.
[0025] Step 5, using the said model and the generated noise - sample paired data, re - apply the method of the second step to calculate the interpolation trajectory, and further apply the method of the third step to update the weight matrix of the model M again to obtain the updated model. Finally, sample translational noise from a Gaussian distribution and sample rotational noise from an isotropic Gaussian distribution, use the noise as the input of the model to generate a high - quality protein backbone structure according to the method of the fourth step.
[0026] The specific method of modeling the residues of the protein backbone structure as frames containing translation and rotation is as follows: Starting from the idealized central coordinates, through the translation and rotation transformations contained in the frame, the coordinates of N, C α , C, and O in each residue can be obtained:
[0027]
[0028] where N represents a nitrogen atom, C α represents an α - carbon atom, C represents a carbon atom, O represents an oxygen atom; the superscript i represents the i - th residue, * represents the central coordinates; the operator ○ represents group action; the local transformation T contains a translation transformation and a rotation transformation: the translation transformation is denoted as x, the rotation axis of the rotation transformation is u, the rotation angle is φ, and through the exponential mapping of quaternions, it is denoted as the unit quaternion q:
[0029] where ω represents the angular velocity; the inverse logarithmic mapping is expressed as ω = 2log(q); the Hamilton multiplication of quaternions is denoted as the inverse of the unit quaternion is expressed as the operation of taking the imaginary part of the quaternion is denoted as Im, and for any coordinate v, the translation and rotation transformations contained in the frame are applied to v:
[0030]
[0031] The use of the trajectory and the said model during training to construct a loss function pair The specific method for updating the weight matrix is as follows: Given the local transformation T at time t t , including the translation transformation x t and the rotation transformation q t , the model predicts the translation and rotation at the end time, and this process is expressed as:
[0032]
[0033] where T t is the local transformation at time t; T θ,1 is the local transformation at the end time predicted by the model , including the translation transformation x θ,1 and the rotation transformation q θ,1 . Then calculate the translation velocity υ θ,t and the rotational angular velocity ω θ,t from time t to the end time:
[0034]
[0035] where represents the Hamilton multiplication. The loss functions corresponding to the translation and rotation are specifically:
[0036]
[0037] where is the loss function term for the translation transformation; is the loss function term for the rotation transformation.
[0038] The innovation of the embodiments of the present invention lies in:
[0039] Compared with the existing methods, the present invention can achieve efficient and high-quality protein backbone generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Through the following description with reference to the drawings, the above and other objects and features of the present disclosure will become clearer.
[0041] Figure 1 is a structural diagram showing a quaternion rectification method for a protein backbone generation task according to an embodiment of the present disclosure;
[0042] Figure 2 is a corresponding graph showing the training parameters, inference time, and designability of different protein backbone generation models according to an embodiment of the present disclosure;
[0043] Figure 3 is a comparison showing the error after the quaternion and rotation matrix are converted back and forth between the sum axis angle representation according to an embodiment of the present disclosure;
[0044] Figure 4 Shows the statistics of the number and corresponding frequencies of large angles in the PDB dataset and the SCOPe dataset according to an embodiment of the present disclosure;
[0045] Figure 5 Shows the statistics of the average number of small angles in protein backbones of different lengths generated according to an embodiment of the present disclosure;
[0046] Figure 6 Shows the comparison of rotation matrix geodesic interpolation, additive spherical linear interpolation, and exponential spherical linear interpolation in terms of interpolation form, Euler solver, numerical stability, and application scenarios according to an embodiment of the present disclosure;
[0047] Figure 7 Shows the comparison of the results of generating protein backbones by different protein backbone generation models on the PDB dataset according to an embodiment of the present disclosure;
[0048] Figure 8 Shows the visualization effect of the secondary structure distribution of protein backbones generated by different protein backbone generation models according to an embodiment of the present disclosure;
[0049] Figure 9 Shows the comparison of the scores of the designability of long-chain protein backbones generated by different protein backbone generation models according to an embodiment of the present disclosure;
[0050] Figure 10 Shows the comparison of the sRMSD of the designability of long-chain protein backbones generated by different protein backbone generation models according to an embodiment of the present disclosure;
[0051] Figure 11 Shows the visualization effect of generating protein backbones of different lengths using a quaternion rectification model according to an embodiment of the present disclosure;
[0052] Figure 12 Shows the comparison of the results of whether the quaternion rectification model uses an exponential scheduler, rectification, and data filtering according to an embodiment of the present disclosure;
[0053] Figure 13 Shows the comparison of the results of generating protein backbones by different protein backbone generation models on the SCOPe dataset according to an embodiment of the present disclosure;
[0054] Figure 14 Shows the comparison of the scores of the designability of protein backbones generated by different protein backbone generation models on the SCOPe dataset according to an embodiment of the present disclosure. Detailed implementation manners
[0055] The following specific embodiments are provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a particular order. Additionally, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0056] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, devices, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.
[0057] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more of them.
[0058] Although terms such as "first", "second", and "third" may be used herein to describe various components, elements, regions, layers, or parts, these components, elements, regions, layers, or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, element, region, layer, or part from another. Thus, the first component, first element, first region, first layer, or first part referred to in the examples described herein may also be referred to as the second component, second element, second region, second layer, or second part without departing from the teachings of the examples.
[0059] In the specification, when an element (such as a layer, region, or substrate) is described as "on", "connected to", or "coupled to" another element, the element may be directly "on", directly "connected to", or "coupled to" the other element, or there may be one or more other elements therebetween. Conversely, when an element is described as "directly on", "directly connected to", or "directly coupled to" another element, there may be no other elements therebetween.
[0060] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0061] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding this disclosure. Unless explicitly defined herein, terms (such as those defined in a general dictionary) shall be construed to have a meaning consistent with their meaning in the context of the relevant art and this disclosure, and shall not be construed in an idealized or overly formal manner.
[0062] In addition, in the description of the examples, when it is considered that a detailed description of a related structure or function that is well-known will cause an ambiguous interpretation of this disclosure, such a detailed description will be omitted.
[0063] Figure 1 is a schematic diagram showing a quaternion rectification method for protein backbone generation tasks according to an embodiment of this disclosure.
[0064] It includes the following steps:
[0065] The first step: Given a protein dataset D and an initialized model is an SE(3)-equivariant neural network, which consists of a node feature module, an edge feature module, and an invariant point attention module. In the sequence of a protein, the amino group and carboxyl group between amino acids dehydrate to form a bond. Since part of its groups participate in the formation of the peptide bond, the remaining structural part is called an amino acid residue. The residues of the protein backbone structure in the dataset D are modeled as frames containing translation and rotation according to the existing method. Specifically, starting from an idealized central coordinate, through the translation and rotation transformations contained in the frame, the coordinates of N, C α in each residue can be obtained:
[0066]
[0067] where, N represents a nitrogen atom, C α represents an α-carbon atom, C represents a carbon atom, and O represents an oxygen atom; the superscript i represents the i-th residue, * represents the central coordinate; the operator ○ represents a group action; the local transformation T includes a translation transformation and a rotation transformation;
[0068] The second step: Sample local transformations from a noise distribution, that is where is translation noise sampled from a Gaussian distribution, is rotation noise sampled from an isotropic Gaussian distribution on SO(3). Randomly pair the sampled local transformations with the local transformations corresponding to the protein backbone in the dataset D and calculate the interpolation trajectory between the local transformations, where the translation trajectory is calculated by linear interpolation and the rotation trajectory is calculated by exponential spherical linear interpolation:
[0069] x t = (1 - t)x0 + tx1
[0070]
[0071] The translation transformation is represented by x, and the rotation transformation is represented by the unit quaternion q; x0 and q0 are the translation and rotation transformations at time 0, corresponding to the local transformations of the sampled noise; x1 and q1 are the translation and rotation transformations at time 1, corresponding to the local transformations of the paired protein backbones in the dataset D. denotes the Hamilton multiplication; exp(·) represents converting the axis-angle representation of rotation into a unit quaternion, and log(·) represents its inverse transformation; the first row of the formula calculates the linear interpolation of the translation transformation, and the second row of the formula calculates the exponential spherical linear interpolation of the rotation transformation, respectively obtaining the translation x t and rotation q t .
[0072] Step 3, during the training process, use the interpolation trajectory and the model to construct a loss function to update the weight matrix of the model . The specific final loss function adopted is
[0073]
[0074] where is the loss function term for the translation transformation; is the loss function term for the rotation transformation; is the loss function term for reducing the physical conflicts of the generated backbone structure, α is 's weight, ∈ is the threshold of the indicator function 1, and t ∈ [0, 1] is the sampled time step.
[0075] Step 4, use the trained model to gradually generate protein backbone samples starting from the local transformation of the noise. The specific implementation of the gradual update is: given the local transformation T t , including the translation transformation x t and the rotation transformation q t , the model predicts the translation and rotation at the end time. This process is expressed as:
[0076]
[0077] where, T t is the local transformation at time t; T θ,1 is the local transformation at the end time predicted by the model , including the translation transformation x θ,1 and the rotation transformation qθ,1 Then calculate the translational velocity υ from time t to the end time θ,t and the rotational angular velocity ω θ,t :
[0078]
[0079] For the sampling process with a time step interval of Δt, use the Euler method to achieve the step-by-step update of translational transformation and rotational transformation:
[0080] x t+△t = x t + υ θ,t ·Δt
[0081]
[0082] x t and q t are the translational transformation and rotational transformation at time t; x t+Δt and q t+Δt are the translational transformation and rotational transformation after the time interval Δt; υ θ,t and ω θ,t are the calculated translational velocity and rotational angular velocity from time t to the end time; represents Hamilton multiplication; exp(·) represents converting the axis-angle representation of rotation into a unit quaternion; γ is a hyperparameter that controls sampling acceleration.
[0083] Step 5, use the model and the generated noise-sample paired data, reapply the method of the second step to calculate the interpolation trajectory, and further apply the method of the third step to update the weight matrix of the model M again to obtain an updated model. Finally, sample translational noise from a Gaussian distribution and sample rotational noise from an isotropic Gaussian distribution, and use the noise as the input of the model to generate a high-quality protein backbone structure according to the method of the fourth step.
[0084] Specifically:
[0085] First, give the task definition: Given a dataset D and an initialized model Then use this dataset to train the model when needed. During the training process, the loss function is:
[0086]
[0087] After training, generate samples starting from noise. And use the noise-sample pairs to train the model again to achieve the rectification operation.
[0088] Then the task to be solved by this method: In the present invention, two protein datasets are used: the Protein Data Bank (PDB) dataset and the SCOPe dataset. By training the model and then sampling noise to obtain samples, and training again with the noise-sample pairs to achieve rectification. The finally obtained model can quickly generate high-quality protein backbone structures with fewer iteration steps.
[0089] Given a dataset D and an initialized model when using the dataset to train the model, first, the residues of the protein are modeled as frames: starting from the idealized central coordinates, through the translation and rotation transformations contained in the frame, the coordinates of N, C α , C, and O in each residue can be obtained:
[0090]
[0091] where N represents a nitrogen atom, C α represents the α-carbon atom, C represents a carbon atom, and O represents an oxygen atom; the superscript i represents the i-th residue, and * represents the central coordinates; the operator ○ represents the group action; the local transformation T includes a translation transformation and a rotation transformation: the translation transformation is denoted as x, the rotation axis of the rotation transformation is u, the rotation angle is φ, and through the exponential mapping of quaternions, it is denoted as the unit quaternion q:
[0092]
[0093] where ω represents the angular velocity; the inverse logarithmic mapping is represented as ω = 2log(q); the Hamilton multiplication of quaternions is denoted as the inverse of the unit quaternion is represented as the operation of taking the imaginary part of the quaternion is denoted as Im, and for any coordinate v, the translation and rotation transformations contained in the frame are applied to v:
[0094]
[0095] Then sample noise and randomly pair it with the proteins in the dataset D. For each pair, calculate the interpolation trajectory between the corresponding frames: the translation trajectory is calculated by linear interpolation, and the rotation trajectory is calculated using exponential spherical linear interpolation:
[0096] x t =(1 - t)x0 + tx1
[0097]
[0098] The translational transformation is represented by x, and the rotational transformation is represented by the unit quaternion q; x0 and q0 are the translational and rotational transformations at time 0, corresponding to the locally sampled noise transformations; x1 and q1 are the translational and rotational transformations at time 1, corresponding to the local transformations of the paired protein backbones in the dataset D. denotes the Hamilton multiplication; exp(·) represents the conversion of the axis-angle representation of rotation into a unit quaternion, and log(·) represents its inverse transformation; the first row of the formula calculates the linear interpolation of the translational transformation, and the second row of the formula calculates the exponential spherical linear interpolation of the rotational transformation, respectively obtaining the translation x t and rotation q t at time t.
[0099] During the training process, the trajectory and the model are used to construct a loss function to update the weight matrix of t . Given the local transformation T t at time t, including the translational transformation x t and the rotational transformation q , the model
[0100]
[0101] predicts the translation and rotation at the end time, and the calculation method is: t where T θ,1 is the local transformation at time t; T is the local transformation at the end time predicted by the model θ,1 , including the translational transformation x θ,1 and the rotational transformation q θ,t . Then, the translational velocity υ θ,t and the rotational angular velocity ω
[0102]
[0103] from time t to the end time are calculated as: where
[0104]
[0105] denotes the Hamilton multiplication. The loss functions corresponding to the translation and rotation are specifically: is the loss function term for the translational transformation; is the loss function term for the rotational transformation.
[0106] Considering the auxiliary loss function , the final loss function is:
[0107]
[0108] It is a loss function term for reducing the physical conflicts in generating the skeletal structure, and α is its weight, ∈ is the threshold of the indicator function 1, and t ∈ [0, 1] is the sampled time step.
[0109] After the training is completed, inference needs to be performed. Starting from the noise sample, the protein backbone is sampled:
[0110] x t+△t = x t + υ θ,t ·△t
[0111]
[0112] x t and q t are the translational transformation and rotational transformation at time t; x t+Δt and q t+Δt are the translational transformation and rotational transformation after the time interval Δt; υ θ,t and ω θ,t are the translational velocity and rotational angular velocity calculated from time t to the end time; denotes Hamilton multiplication; exp(·) denotes converting the axis-angle representation of rotation into a unit quaternion; γ is a hyperparameter controlling the sampling acceleration.
[0113] Furthermore, the noise-sample pair is further used to retrain the model to achieve the rectification effect. The finally obtained model can quickly generate a high-quality protein backbone structure with fewer iteration steps.
[0114] The technical effects to be achieved by the present invention are as follows:
[0115] Significantly reduce the number of inference steps: The present invention adopts a flow matching technical route and is solved based on ordinary differential equations (ODE). Compared with the diffusion model based on stochastic differential equations (SDE), it naturally has fewer inference steps. In addition, by introducing the rectification operation, the intersection between translational trajectories and rotational trajectories is effectively reduced, further reducing the transmission cost, thereby realizing a further reduction in the number of inference steps.
[0116] Significantly shorten the inference time: As mentioned above, the present invention reduces the number of inference steps. At the same time, due to the high efficiency of the exponential quaternion interpolation calculation, the calculation efficiency of a single-step inference is further improved compared with the rotation matrix representation method. Combining these two points, the present invention significantly shortens the time required for inference.
[0117] Generate high-quality protein backbones: The exponential spherical quaternion interpolation of the present invention shows good stability both in dealing with large angles and small angles. In addition, the model reduces the transmission cost after rectification. These two points together improve the generation quality of the protein backbone.
[0118] Support for the generation of long-chain proteins: The designability score of the present invention in the generation of long-chain proteins is significantly better than that of existing models, providing an effective technical path for the generation of long-chain proteins.
[0119] Experimental verification
[0120] Figure 2 It is a corresponding graph of the training parameters, inference time, and designability of different protein backbone generation models. The horizontal axis is the inference time, the vertical axis is the designability score of the generated protein backbone, and the area of each circle represents the number of parameters of the model.
[0121] Figure 7 It is a comparison of the results of generating protein backbones by different protein backbone generation models on the PDB dataset. Among the rows with a designability score greater than 0.8, the top three best results are marked, and the best result is shown in bold. It includes the designability score, sRMSD of designability, diversity, and novelty.
[0122] Figure 9 It is a comparison of the designability scores of different protein backbone generation models for generating long-chain protein backbones according to embodiments of the present disclosure. The horizontal axis is the length of the protein chain, and the vertical axis is the designability score.
[0123] Figure 13 It is a comparison of the results of generating protein backbones by different protein backbone generation models on the SCOPe dataset according to embodiments of the present disclosure. Among the rows with a designability score greater than 0.8, the top three best results are marked, and the best result is shown in bold. It includes the designability score, sRMSD of designability, diversity, and novelty.
[0124] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure as defined by the claims and their equivalents.
Claims
1. A quaternion rectification method for protein backbone generation tasks, characterized in that, It includes five steps: Step 1: According to the dataset D of known proteins and the initialized model Model the residues of the protein backbone structure in the dataset D as local transformations including translation and rotation; Step 2: Sample local transformations from the noise distribution, i.e., where is the translational noise sampled from the Gaussian distribution, is the rotational noise sampled from the isotropic Gaussian distribution on SO(3), where SO(3) refers to the special orthogonal group in three-dimensional Euclidean space; randomly pair the sampled local transformations with the local transformations corresponding to the protein backbone in the dataset D and calculate the interpolation trajectory between the local transformations, where the translational trajectory is obtained by linear interpolation and the rotational trajectory is obtained by exponential spherical linear interpolation: x t = (1 - t)x0 + tx1 The translation transformation is represented by x, and the rotation transformation is represented by the unit quaternion q; x0 and q0 are the translation and rotation transformations at time 0, corresponding to the locally transformed noise sampled; x1 and q1 are the translation and rotation transformations at time 1, corresponding to the local transformation of the paired protein backbones in the dataset D. denotes the Hamilton multiplication; exp(·) represents the conversion of the axis-angle representation of rotation into a unit quaternion, and log(·) represents its inverse transformation; the first row of the formula calculates the linear interpolation of the translation transformation, and the second row of the formula calculates the exponential spherical linear interpolation of the rotation transformation, respectively obtaining the translation x t and rotation q t ; Step 3: During the training process, use the interpolation trajectory and the model to construct a loss function for the model to update the weight matrix. The specific final loss function used is as follows: Among them is the loss function term for translational transformation; is the loss function term for rotational transformation; is the loss function term for reducing physical conflicts in the generated skeleton structure, α is the weight of, ε is the threshold of the indicator function 1, and t ∈ [0, 1] is the sampled time step; Step 4: Utilize the trained model Start from the local transformation of noise and gradually generate protein backbone samples. The specific implementation of the gradual update is as follows: Given the local transformation T t , including the translation transformation x t and the rotation transformation q t , the model predicts the translation and rotation at the end time. This process is expressed as: Among them, T t is the local transformation at time t; T θ,1 is the local transformation at the end time predicted by the model , including the translation transformation x θ,1 and the rotation transformation q θ,1 . Then calculate the translation velocity v θ,t and the rotational angular velocity ω θ,t from time t to the end time: For the sampling process with a time step interval of Δt, the Euler method is used to gradually update the translation transformation and the rotation transformation: x t+△t = x t + υ θ,t · Δ t x t and q t are the translation transformation and rotation transformation at time t; x t+Δt and q t+Δt are the translation transformation and rotation transformation after a time interval Δ t ; v θ,t and ω θ,t are to calculate the translational velocity and rotational angular velocity from time t to the end time; denotes Hamilton multiplication; exp(·) denotes converting the axis - angle representation of rotation into a unit quaternion; γ is a hyper - parameter for controlling sampling acceleration; Step 5: Using the said model and the generated noise-sample paired data, reapply the method of the second step to calculate the interpolation trajectory, and further apply the method of the third step to update the weight matrix of the model M again to obtain an updated model; Finally, sample translational noise from a Gaussian distribution and rotational noise from an isotropic Gaussian distribution, and use the noise as the input to the model to generate high-quality protein backbone structures according to the method in the fourth step.
2. A quaternion rectification method for protein backbone generation tasks according to claim 1, wherein The specific method of modeling the residues of the protein backbone structure as frames containing translation and rotation is as follows: Starting from the idealized central coordinates, through the translation and rotation transformations contained in the frames, the coordinates of N, C α , C, and O in each residue can be obtained: wherein, N represents a nitrogen atom, C α represents an α-carbon atom, C represents a carbon atom, and O represents an oxygen atom; the superscript i represents the i-th residue, and * represents the central coordinate; the operator ○ represents a group action; the local transformation T includes a translation transformation and a rotation transformation: the translation transformation is denoted as x, the rotation axis of the rotation transformation is u, the rotation angle is φ, and through the exponential mapping of quaternions, it is denoted as the unit quaternion q: where ω represents the angular velocity; the inverse logarithmic mapping is expressed as ω = 2 log(q); the Hamilton multiplication of quaternions is denoted as the inverse of the unit quaternion is expressed as The operation of taking the imaginary part of a quaternion is denoted as Im, and for any coordinate v, the translation and rotation transformations contained in the frame are applied to v:
3. A quaternion rectification method for protein backbone generation tasks according to claim 1, characterized in that, During the training process, the trajectory and the model are used to construct a loss function for updating the weight matrix of. The specific method is as follows: Given the local transformation T at time t t , including the translation transformation x t and the rotation transformation q t , the model predicts the translation and rotation at the end time, and its calculation method is: Among them, T t is the local transformation at time t; T θ,1 is the local transformation at the end time predicted by the model, including the translation transformation x and the rotation transformation q θ,1 , and then calculate the translation velocity υ θ,1 from time t to the end time and the rotational angular velocity ω θ,t : θ,t Among them represents Hamilton multiplication. The specific loss functions corresponding to the translation and rotation are as follows: Among them is the loss function term for translational transformation; is the loss function term for rotational transformation.