Protein structure generation method based on LexFrameDiff diffusion and electronic equipment thereof
Through the protein structure generation method based on LexFrameDiff diffusion, the problem of difficult to generate protein multi-conformational states in the prior art is solved, and the ability to sample a reasonable protein structure from noise is realized, and the problem of local minimum value is avoided.
Patent Information
- Application Number
- CN202510282250.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art lacks a general deep learning framework, which makes it difficult to effectively solve the problem of multi-conformational state generation of proteins, especially when designing such as binder proteins and higher-order symmetric structures.
A protein structure generation method based on LexFrameDiff diffusion is proposed. This method uses the diffusion model to gradually denoise by inputting monomer representation, paired representation and noise information to generate a three-dimensional structure of the protein.
This method can sample reasonable protein structures from noise, solve the problem that traditional optimization methods are prone to falling into local minimum values, and has the ability to generate multiple reasonable protein structures.
Smart Images

Figure SMS_1 
Figure SMS_11 
Figure SMS_13
Abstract
Description
Technical Field
[0001] The present invention relates to the field of protein design, and more specifically, to a method for generating protein structures based on LexFrameDiff diffusion and an electronic device thereof. Background Art
[0002] Proteins are the main executors of life activities, and the correct execution of their functions depends on specific three-dimensional structures. However, proteins are not static entities but dynamically transition between multiple conformational states under physiological conditions to perform specific biological functions such as signal transduction, enzyme catalysis, and molecular recognition. Therefore, accurately predicting and designing the multi-conformational states of proteins is of great significance for deeply understanding their functional mechanisms and regulatory principles, as well as for directed evolution and drug design.
[0003] In recent years, significant progress has been made in protein design using deep learning methods, and various algorithms have emerged in an endless stream. Although the progress is gratifying, the drawbacks are still visible. Researchers may need to specifically design a framework for a particular problem. Currently, there is still a lack of a general deep learning framework for solving more extensive problems such as binder protein design and higher-order symmetric structure design. These problems can all be regarded as generation problems, and many generation models have emerged in the development of deep learning, such as VAE, GAN, Diffusion, etc. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, one of the purposes of the present invention is to propose a new method, aiming to provide a new protein structure generation model that can sample protein structures from noise.
[0005] To achieve the above object, the present invention provides the following technical solution: A method for generating protein structures based on LexFrameDiff diffusion, comprising the following steps: 1) Input of information: Input monomer representation, pairwise representation, and noise information; the monomer representation reflects the information of individual amino acids in the protein; the pairwise representation reflects the mutual relationship between amino acids in pairs; the noise information consists of a rotation matrix and a translation vector, and the initial noise information is pure noise; The monomer representation includes time encoding information, and the time encoding information reflects the denoising time nodes of L dimensions, representing the denoising degree of L dimensions; The pairwise representation includes moment encoding information, and the moment encoding information reflects the denoising time nodes of L*L dimensions, representing the denoising degree of L*L dimensions; 2) Denoising of information: Input the monomer representation into the IPA module for denoising to obtain a new monomer representation. Input the new monomer representation into the Transition layer to obtain the first transitional monomer information. Use the paired representation as a bias to input into the Row-attention module to obtain new monomer information. Input the new monomer information into the affine_update linear layer for denoising to obtain the updated noise information, and output the noise information; Input the noise information into the affine2pair module to obtain a new paired representation. After merging the old paired representation and the new paired representation, input them into the Transition layer to generate the paired representation updated again, and output the paired representation updated again; Input the first transitional monomer information and the paired representation updated again into the Transition layer to obtain the second transitional monomer information. Input the second transitional monomer information into the Row-attention module to obtain the monomer representation updated again, and output the monomer representation updated again; 3) Repeated denoising: Repeating the calculation in step 2) n times completes one denoising. The next denoising still starts from step 1). The time encoding information and the moment encoding information are advanced forward. The number of denoising times corresponds to the time encoding information and the moment encoding information. After denoising T times, the denoised noise information, monomer representation, and paired representation are obtained. The finally denoised noise information is the structural information; both n and T are integers greater than 1; 4) Generation of the structure: Generate the three-dimensional structure of the protein according to the finally obtained structural information, monomer representation, and paired representation, and obtain the corresponding amino acid sequence according to the three-dimensional structure of the protein.
[0006] Preferably, in step 1), the monomer representation further includes amino acid type information, amino acid index encoding information, and dihedral angle information. The amino acid encoding information reflects the type of amino acid. The amino acid index encoding information reflects the index encoding of the amino acid in the L dimension. The dihedral angle information reflects the dihedral angle of the amide plane and the C α atom of the Concatenate the amino acid type information, amino acid index encoding information, and time encoding information together on the feature dimension, and then project them through a linear layer to obtain the initialized monomer representation. Concatenate the time encoding information and the dihedral angle information together on the feature dimension, and then project them through a multi-layer perceptron to obtain the dihedral angle representation. Concatenate the initialized monomer representation and the dihedral angle representation together on the feature dimension, and then project them through a multi-layer perceptron to obtain the new monomer representation, and output the new monomer representation; The paired representation further includes relative position index coding information and amino acid distance binning information. The relative position index coding information reflects the relative distance coding of the indexes of amino acids in the L*L dimension, and the amino acid distance binning information reflects the distance between pairwise amino acids in three-dimensional space. The relative position index coding information is projected through a linear layer to obtain relative position index coding. The time coding information and the amino acid distance binning information are concatenated in the feature dimension and then projected through a linear layer to obtain an initialized paired representation. The initialized paired representation is input into the TemplatePairStack module to obtain a new paired representation. The new paired representation and the relative position index coding are superimposed to obtain a further updated paired representation, and the further updated paired representation is output. In step 3), after one denoising, the monomer representation and the paired representation are output. The dihedral angle information after denoising is obtained from the monomer representation and output. The amino acid distance binning information after denoising is obtained from the paired representation and output.
[0007] Preferably, the protein structure generation method based on LexFrameDiff diffusion includes model parameters, which are obtained through training. The training method is as follows: perform the processes of step 1), step 2), and step 3), calculate the loss value between the structure information or / and monomer representation or / and paired representation output in step 3) and the true structure information or / and monomer representation or / and paired representation, and obtain the model parameters according to the loss value.
[0008] Preferably, the loss value includes a frame loss value, which reflects the difference between the output structure information and the true structure information. The calculation method of the frame loss value is as follows:
[0009] where, frame represents the frame loss value, N res represents the number of amino acids; w trans represents the loss weight of the translation distance, w rot represents the loss weight of the rotation matrix, z represents the translation distance; r represents the rotation matrix, d clamp represents the fixed maximum translation distance error.
[0010] Preferably, the loss value includes a dihedral angle loss value, which reflects the difference between the output dihedral angle information and the true dihedral angle information. The calculation method of the dihedral angle loss value is as follows: where, torsionRepresents the dihedral angle loss value, , , represents the dihedral angle information, where i and f respectively represent the f-th dihedral angle information on the i-th amino acid, represents the dihedral angle length, represents the normalized dihedral angle, represents the true dihedral angle information.
[0011] Preferably, the loss value includes a distance loss value, which reflects the difference between the output amino acid distance binning information and the true amino acid distance binning information. The calculation method of the distance loss value is:
[0012] where, dist represents the distance loss value, N res represents the number of amino acids, p b ij represents the probability distribution of the amino acid distance binning information, where i and j represent pairs of amino acids, b represents the bin index, and y b ij represents the distance bin of the true structure.
[0013] Aiming at the deficiencies of the prior art, the second object of the present invention is a device capable of running the above algorithm.
[0014] To achieve the above object, the present invention provides the following technical solution: An electronic device, including: a processor and a memory, the memory stores executable code, and when the executable code is executed by the processor, the processor executes the protein structure generation method based on LexFrameDiff diffusion as described above.
[0015] Compared with the prior art, the advantages of the present invention are as follows: This is a novel protein structure generation model that can sample protein structures from noise. LexFrameDiff uses a diffusion model to gradually denoise and maps random noise to the structure space of proteins. This process is similar to the biological behavior of proteins from the unfolded state to stable secondary and tertiary structures, ensuring that the generated structures are reasonable in terms of physical properties such as chemical bond lengths and angle constraints. By controlling the template information and secondary structure information, proteins can generate local structures in specific domains. Due to the diversity and randomness of noise, multiple reasonable protein structures can be finally generated. Protein structure generation involves high-dimensional space optimization problems. LexFrameDiff gradually samples through the diffusion model and adjusts the protein backbone conformation from rough to fine, effectively solving the local minimum problem of traditional optimization methods. Detailed implementation manners
[0016] The present invention will be further described in detail below in conjunction with embodiments.
[0017] The linear layer is as follows: The linear layer described in the present invention is a common fully connected layer in a neural network, which is a process of linearly transforming input data and mapping it to output data. All the linear layers mentioned in the present invention are this operation, and the difference lies in the different feature dimensions of the input features and output features.
[0018] The multi-layer perceptron is as follows: The multi-layer perceptron described in the present invention is a module that alternately stacks multiple linear layers and activation functions (non-linear functions). Compared with the linear layer, the multi-layer perceptron can complete a more complex mapping.
[0019] The IPA module is as follows: The IPA described in the present invention is the same as the IPA used in AlphaFold2. The purpose is to fuse monomer representation information, pairwise representation information, and amino acid rotation and translation information, and output richer monomer representation information with positional and relative relationships.
[0020] The TemplatePairStack is as follows: The TemplatePairStack described in the present invention is basically the same as the architecture used in AlphaFold2, and the difference lies in adding information about time encoding when stacking template features.
[0021] The Transition layer is as follows: The transition layer used in the present invention is similar in function to the transition layer used in AlphaFold2, connecting between two modules and mapping the output features of the previous module to the input features of the next module.
[0022] The Row-attention module is as follows: The Row-attention module described in the present invention is the same as that in AlphaFold2. However, due to the unknownness of the sequence in the diffusion model, the input features change from multi-sequence representation to single-sequence representation. The role of this module in the present invention is to extract the attention relationship of sequence representations between different index positions.
[0023] The Affine-update module is as follows: Map the monomer representation to the form of regularized quartiles and translation vectors through a linear layer, and use it to update the original quartiles and translation vectors, and then convert the quartiles into rotation matrices (mathematical calculations).
[0024] The Affine2pair module is as follows: Set the default amino acid framework (N, CA, C, CB) as the basis, and use the rotation matrix and translation vector in the updated Affine to calculate the coordinates of the amino acid backbone atoms after rotation and translation. Calculate the distances between amino acids using the CB atom coordinates, bin the distances according to the proximity, and project them into the pairwise representation of Affine through a linear layer projection. Concatenate the previous pairwise representation and the pairwise representation features of Affine, and perform feature fusion through the Transition layer to generate a new pairwise representation. Embodiment
[0025] A protein structure generation method based on LexFrameDiff diffusion, comprising the following steps: 1) Input of information: Input monomer representation, pairwise representation, and noise information; the monomer representation reflects the information of individual amino acids in the protein; the pairwise representation reflects the mutual relationship between pairwise amino acids; the noise information consists of a rotation matrix and a translation vector, and the initial noise information is pure noise; The monomer representation includes time encoding information, which reflects the denoising time nodes of the L dimension and represents the degree of denoising of the L dimension; The pairwise representation includes moment encoding information, which reflects the denoising time nodes of the L*L dimension and represents the degree of denoising of the L*L dimension; 2) Denoising of information: Input the monomer representation into the IPA module for denoising to obtain a new monomer representation, input the new monomer representation into the Transition layer to obtain the first transitional monomer information, input the pairwise representation as a bias into the Row-attention module to obtain new monomer information, input the new monomer information into the affine_update linear layer for denoising to obtain the updated noise information, and output the noise information; Input the noise information into the affine2pair module to obtain a new pairwise representation, merge the old pairwise representation and the new pairwise representation and input them into the Transition layer to generate a pairwise representation updated again, and output the pairwise representation updated again; Input the first transitional monomer information and the pairwise representation updated again into the Transition layer to obtain the second transitional monomer information, input the second transitional monomer information into the Row-attention module to obtain the monomer representation updated again, and output the monomer representation updated again; 3) Repeated denoising: Repeating the calculation in step 2) n times completes one denoising. The next denoising still starts from step 1). The time encoding information and the moment encoding information are advanced forward. The number of denoising times corresponds to the time encoding information and the moment encoding information. After T times of denoising, the denoised noise information, monomer representation, and pairwise representation are obtained. The finally denoised noise information is the structural information; both n and T are integers greater than 1. 4) Structure generation: Generate the three-dimensional protein structure based on the finally obtained structural information, monomer representation, and pairwise representation, and obtain the corresponding amino acid sequence according to the three-dimensional protein structure.
[0026] This is a novel protein structure generation model that can sample protein structures from noise. LexFrameDiff uses a diffusion model to gradually denoise, mapping random noise to the structural space of proteins. This process is similar to the biological behavior of proteins from the unfolded state to stable secondary and tertiary structures, ensuring that the generated structures are reasonable in terms of physical properties such as chemical bond lengths and angle constraints. Through the control of template information and secondary structure information, local structures of proteins can be generated under specific domains. Due to the diversity and randomness of noise, multiple reasonable protein structures can ultimately be produced. Protein structure generation involves high-dimensional space optimization problems. LexFrameDiff gradually samples through a diffusion model, adjusting the conformation of the protein backbone and side chains from rough to fine, effectively solving the local minimum problem of traditional optimization methods. Example
[0027] The difference from Example 1 is as follows: In step 1), the monomer representation further includes amino acid type information, amino acid index encoding information, and dihedral angle information. The amino acid encoding information reflects the type of amino acid, the amino acid index encoding information reflects the index encoding of amino acids in the L dimension, and the dihedral angle information reflects the dihedral angle between the amide plane and the C α atom.
[0028] Concatenate the amino acid type information, amino acid index encoding information, and time encoding information in the feature dimension, and then project them through a linear layer to obtain the initialized monomer representation. Concatenate the time encoding information and the dihedral angle information in the feature dimension, and then project them through a multi-layer perceptron to obtain the dihedral angle representation. Concatenate the initialized monomer representation and the dihedral angle representation in the feature dimension, and then project them through a multi-layer perceptron to obtain a new monomer representation, and output the new monomer representation. The paired representation further includes relative position index encoding information and amino acid distance binning information. The relative position index encoding information reflects the relative distance encoding of the indices of amino acids in the L*L dimension, and the amino acid distance binning information reflects the distances between pairwise amino acids in three-dimensional space. The relative position index encoding information is projected through a linear layer to obtain relative position index encoding. The time encoding information and the amino acid distance binning information are concatenated in the feature dimension and then projected through a linear layer to obtain an initialized paired representation. The initialized paired representation is input into the TemplatePairStack module to obtain a new paired representation. The new paired representation and the relative position index encoding are superimposed to obtain a further updated paired representation, and the further updated paired representation is output. In step 3), after one denoising, the monomer representation and the paired representation are output. The denoised dihedral angle information is obtained from the monomer representation and output, and the denoised amino acid distance binning information is obtained from the paired representation and output. Embodiment
[0029] The difference from Embodiment 1 is that the protein structure generation method based on LexFrameDiff diffusion includes model parameters, which are obtained through training. The training method is as follows: perform the processing of steps 1), 2), and 3), calculate the loss value between the structure information and / or monomer representation and / or paired representation output in step 3) and the true structure information and / or monomer representation and / or paired representation, and obtain the model parameters according to the loss value. Embodiment
[0030] The difference from Embodiment 1 is as follows: The loss value includes a frame loss value, which reflects the difference between the output structure information and the true structure information. The calculation method of the frame loss value is as follows: , where frame represents the frame loss value, N res represents the number of amino acids; w trans represents the loss weight of the translation distance, w rot represents the loss weight of the rotation matrix, z represents the translation distance; r represents the rotation matrix, d clamp represents the fixed maximum translation distance error.
[0031] The backbone loss function describes the backbone information of amino acids. The rotation matrix depicts the orientation of the overall amino acid at the origin position, and the translation vector depicts the direction and distance of the amino acid's translation from the origin. Among common amino acids, the relative position information of the backbone atoms is basically fixed, such as the bond length. Through the position of the CA atom and the orientation of the amino acid, the coordinates of other backbone atoms can be calculated. Compared with the direct description of the coordinates of all backbone atoms of amino acids, the description method using the rotation matrix and translation vector is more physical and implies the relative distance relationship between the internal atoms of the amino acid. Example
[0032] The difference from Example 1 is as follows: The loss value includes the dihedral angle loss value, which reflects the difference between the output dihedral angle information and the true dihedral angle information. The calculation method of the dihedral angle loss value is as follows: Where, torsion represents the dihedral angle loss value, , , represents the dihedral angle information. i and f respectively represent the f-th dihedral angle information on the i-th amino acid. represents the dihedral angle length, represents the normalized dihedral angle, represents the true dihedral angle information.
[0033] The dihedral angle loss is another description of the amino acid structure. Through the dihedral angle, the structure information of the protein can be roughly obtained. This function depicts the dihedral angle with a unit circle, that is, represented by the sine and cosine values of the rotation angle. Therefore, the length of the dihedral angle is 1, and the latter part of this loss is to constrain its length. Secondly, the dihedral angle describes the direction of the vector on the unit circle. Therefore, the former part of the loss describes the component error of the sine and cosine values in the vector. Example
[0034] The difference from Example 1 is as follows: The loss value includes the distance loss value, which reflects the difference between the output amino acid distance binning information and the true amino acid distance binning information. The calculation method of the distance loss value is as follows:
[0035] Where, dist represents the distance loss value, N res represents the number of amino acids, p b ij represents the probability distribution of the amino acid distance binning information. i and j represent pairs of amino acids, b represents the bin index, y bij Distance binning representing the true structure.
[0036] The distance loss function describes the relative distance relationship between individual amino acids and is the projection of the three-dimensional protein structure onto a two-dimensional space. With the aid of the relative distance relationship, the three-dimensional structure can be constrained. By binning the distances, the classification of the proximity between individual amino acids can be obtained, and thus the problem of distance error can be transformed into a classification problem of distance proximity. When classified into the same proximity level, the classification is correct. Therefore, cross-entropy can be used to characterize the loss of classification. Embodiment
[0037] In addition to the method of calculating the loss value between the structural information or / and monomer representation or / and pairwise representation output in step 3) listed in Embodiments 4 to 6 and the true structural information or / and monomer representation or / and pairwise representation, there are other calculation methods in the prior art.
[0038] For example, the secondary structure loss value can be calculated. The secondary structure information, which is also one-dimensional information with the same length as the protein sequence length, is used to describe the protein structure information. If the site is a helix, it is represented by the letter H; if the site is a β-sheet, it is represented by E; if the site is a loop, it is represented by L. This information is also put into the monomer representation through a linear projection.
[0039] The secondary structure data is similar to the following: LLLLHHHHHHHHHLLLEEEEEEEELLLEEEEEEEELLLLHHHHHHHHLLLLL The addition method of the secondary structure information is the same as that of the dihedral angle. It is also through a linear projection and then spliced into the monomer representation. At the end of the model, a head for predicting the secondary structure information is added, and the secondary structure information is projected through the monomer representation. During training, it is supervised with the secondary structure loss.
[0040] Adding the secondary structure information helps to adjust the structural domain of the generated protein by adjusting the input secondary structure information during the sampling time.
[0041] By performing softmax on the projection of the linear layer, the distribution p of the secondary structure is obtained b i , and the calculation method of the secondary structure loss value is:
[0042] Where ss represents the secondary structure loss value; N res represents the number of amino acids, b represents the secondary structure type index, and there are three in total (H, E, L); y b iIt is the secondary structure information of the true label in one-hot form.
[0043] The types of loss values calculated by different loss value calculation methods are different, and the emphasis of the corrected information is also different. The loss value calculation method disclosed in the present invention is the research result of the inventor. In particular, it describes the protein information after step-by-step denoising from the framework dimension, sequence dimension, and pairwise dimension, forms the mutual update and fusion of information inside the model, improves the self-consistency and richness between information, and thus generates a more reasonable protein structure. Embodiment
[0044] An electronic device, comprising: a processor and a memory, the memory stores executable code, and when the executable code is executed by the processor, the processor is caused to execute the algorithm corresponding to the protein structure generation method based on LexFrameDiff diffusion disclosed in Embodiments 1-6.
[0045] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary researchers in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A protein structure generation method based on LexFrameDiff diffusion, characterized in that The following steps are involved: 1) Information input: input monomer representation, pair representation and noise information; the monomer representation reflects the information of individual amino acids in the protein; the pair representation reflects the relationship between amino acids; the noise information consists of rotation matrix and translation vector, and the initial noise information is pure noise; The monomer representation includes time coding information, and the time coding information reflects the time node of the denoising of the L dimension and represents the degree of the denoising of the L dimension; The paired representation includes moment encoding information, which reflects the time node of the denoising in the L*L dimension and represents the degree of denoising in the L*L dimension; 2) Information denoising: The monomer representation is input into the IPA module for denoising to obtain a new monomer representation, the new monomer representation is input into the Transition layer to obtain the first transition monomer information, the paired representation is used as a bias input into the Row-attention module to obtain new monomer information, the new monomer information is input into the affine_update linear layer for denoising to obtain updated noise information, and the noise information is output; The noise information is input into the affine2pair module to obtain a new paired representation. The old paired representation and the new paired representation are combined and input into the Transition layer to generate an updated paired representation. The updated paired representation is output. Input the first transition monomer information and the updated paired representation into the Transition layer to obtain the second transition monomer information, input the second transition monomer information into the Row-attention module to obtain the updated monomer representation, and output the updated monomer representation; 3) Repeat denoising: Repeat the calculation of step 2) n times to complete one denoising. The next denoising still starts from step 1). The time coding information and the moment coding information are advanced. The number of denoising times corresponds to the time coding information and the moment coding information. After denoising T times, the denoised noise information, monomer representation and paired representation are obtained. The final denoised noise information is the structural information. n and T are both integers greater than 1. 4) Structure generation: Generate a three-dimensional protein structure based on the final structural information, monomer representation and pair representation, and obtain the corresponding amino acid sequence based on the three-dimensional protein structure.
2. The protein structure generation method based on LexFrameDiff diffusion according to claim 1, characterized in that: In step 1), the monomer representation also includes amino acid type information, amino acid index coding information and dihedral angle information, wherein the amino acid coding information reflects the type of amino acid, the amino acid index coding information reflects the index coding of amino acid in L dimension, and the dihedral angle information reflects the relationship between amide plane and C α The dihedral angles of atoms, The amino acid type information, amino acid index encoding information and time encoding information are concatenated together in the feature dimension, and then projected through a linear layer to obtain an initialized monomer representation, the time encoding information and dihedral angle information are concatenated together in the feature dimension, and then projected through a multi-layer perceptron to obtain a dihedral angle representation, the initialized monomer representation and the dihedral angle representation are concatenated together in the feature dimension, and then projected through a multi-layer perceptron to obtain a new monomer representation, and the new monomer representation is output; The paired representation also includes relative position index encoding information and amino acid distance binning information, wherein the relative position index encoding information reflects the relative distance encoding of the amino acid index in L*L dimension, and the amino acid distance binning information reflects the distance between two amino acids in three-dimensional space. The relative position index coding information is projected through a linear layer to obtain the relative position index coding, the moment coding information and the amino acid distance binning information are spliced together in the feature dimension, and then projected through a linear layer to obtain an initialized paired representation, the initialized paired representation is input into the TemplatePairStack module to obtain a new paired representation, the new paired representation and the relative position index coding are superimposed to obtain a further updated paired representation, and the further updated paired representation is output; In step 3), after one denoising is completed, a monomer representation and a paired representation are output, denoised dihedral angle information is obtained according to the monomer representation, and the denoised dihedral angle information is output, and denoised amino acid distance binning information is obtained according to the paired representation, and the denoised amino acid distance binning information is output.
3. The protein structure generation method based on LexFrameDiff diffusion according to claim 2, characterized in that: The protein structure generation method based on LexFrameDiff diffusion includes model parameters, and the model parameters are obtained through training. The training method is: performing the processing of step 1), step 2) and step 3), calculating the loss value between the structural information or / and monomer representation or / and paired representation output in step 3) and the real structural information or / and monomer representation or / and paired representation, and obtaining the model parameters according to the loss value.
4. The protein structure generation method based on LexFrameDiff diffusion according to claim 3, characterized in that: The loss value includes a framework loss value, which reflects the difference between the output structure information and the real structure information. The calculation method of the framework loss value is: ,in, frame Represents the framework loss value, N res represents the number of amino acids; w trans represents the loss weight of the translation distance, w rot represents the loss weight of the rotation matrix, z represents the translation distance; r represents the rotation matrix, d clamp Indicates the fixed maximum translation distance error.
5. The protein structure generation method based on LexFrameDiff diffusion according to claim 3, characterized in that: The loss value includes a dihedral angle loss value, which reflects the difference between the output dihedral angle information and the real dihedral angle information. The calculation method of the dihedral angle loss value is: ,in, torsion represents the dihedral angle loss value, , , represents the dihedral angle information, i and f represent the fth dihedral angle information on the i-th amino acid, represents the dihedral length, represents the unitized dihedral angle, Represents the true dihedral angle information.
6. The protein structure generation method based on LexFrameDiff diffusion according to claim 3, characterized in that: The loss value includes a distance loss value, which reflects the difference between the output amino acid distance binning information and the actual amino acid distance binning information. The distance loss value is calculated as follows: ,in, dist represents the distance loss value, N res represents the number of amino acids, p b ij represents the probability distribution of amino acid distance bin information, i, j represent two amino acid pairs, b represents the bin index, y b ij Distance bins representing true structure.
7. An electronic device, characterized in that: include: Processor and A memory storing executable codes, which, when executed by the processor, enable the processor to execute a protein structure generation method based on LexFrameDiff diffusion as claimed in any one of claims 1 to 6.