Tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism

By constructing a protein radius graph and introducing a graph attention mechanism of the diffusion model, the shortcomings of tumor necrosis factor receptor sequence design in the existing technology are addressed, the quality of the generated sequences and the focus on the structural domain characteristics of specific functional proteins are improved, and more efficient tumor necrosis factor receptor sequence design is achieved.

CN120673834APending Publication Date: 2025-09-19NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510803374.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When designing tumor necrosis factor receptor protein sequences, existing technologies cannot accurately reflect the density characteristics of atomic distribution in proteins, the main chain structure information is insufficient, and the training data set is small, resulting in low quality of the generated sequences.

Method used

The tumor necrosis factor receptor sequence design model AttenGNNDiff based on the graph attention mechanism is adopted. By constructing a protein radius graph, introducing a diffusion model and a graph attention mechanism, a denoising network is designed that can update node features and edge features separately. The model is optimized with the self-attention mechanism and trained and tuned using the CATH4.2 dataset and the tumor necrosis factor receptor dataset.

Benefits of technology

It has improved the sequence design capability of tumor necrosis factor receptor protein, significantly improved the quality of generated sequences and the focus on the structural domain characteristics of specific functional proteins, and enhanced the adaptability and generation capability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673834A_ABST
    Figure CN120673834A_ABST
Patent Text Reader

Abstract

The invention provides a tumor necrosis factor receptor sequence design algorithm AttenGNNDiff based on a graph attention mechanism, and the algorithm employs a radius graph to replace a K neighbor graph as a main chain modeling mode, so as to better reflect the density characteristics of atom distribution in a tumor necrosis factor receptor main chain. A neighborhood attention mechanism is introduced to improve the learning efficiency of local structure features of the tumor necrosis factor receptor, an edge feature updating module based on the attention mechanism is combined to improve the weight of edge feature updating in feature updating of a graph, and a diffusion model framework is adopted to enable the method to learn part of sequence information in a training stage. And a prompt tuning method is designed to tune the model, so that the characteristics of the tumor necrosis factor receptor homologous domain can be better extracted, and stronger adaptability and generation capability are shown in a tumor necrosis factor receptor protein sequence design task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of immune system research, and in particular relates to a tumor necrosis factor receptor sequence design model AttenGNNDiff based on a graph attention mechanism. Background Art

[0002] Tumor necrosis factor (TNF) is a cytokine secreted by the immune system. By binding to TNF receptors, it triggers a series of intracellular signaling pathways, leading to inflammatory responses and tissue damage, thereby causing the occurrence of various pathological conditions, including autoimmune diseases, chronic inflammation, exacerbated infections, and the occurrence and development of tumors. For immunodeficiency diseases caused by TNF, designing antagonists with stronger affinity for TNF has become one of the important approaches to treat these diseases. That is, by designing receptors with stronger binding ability to TNF, so that they replace the original receptors to bind to TNF, the pathogenic effects of TNF can be effectively inhibited, providing more precise treatment methods.

[0003] As a protein, the function of the TNF receptor is determined by its unique three-dimensional structure. Therefore, designing a novel TNF receptor with enhanced TNF binding requires a more refined structural design using the principles of molecular docking, regenerating an amino acid sequence that precisely folds into this structure. This design results in a protein with both structural stability and superior performance.

[0004] Focusing on the specific structural features of the tumor necrosis factor receptor, precisely generating receptor protein sequences that conform to fixed backbone constraints is known as fixed-backbone tumor necrosis factor receptor protein sequence design. In recent years, with the rapid development of computational protein design technology, numerous algorithms have emerged for designing protein sequences based on protein backbone structural information. However, these methods still have several challenges: First, most methods model proteins as K-nearest-neighbor graphs, which cannot accurately reflect the density distribution of atoms within proteins; second, the model training relies solely on backbone structural information, resulting in insufficient representation of features in sequence-related regions; further, the relatively small size of protein structure datasets available for training prevents the model from fully learning the protein's sequence evolution; and finally, the output residue probability matrix often contains low confidence levels for some residue probability vectors, further impacting the quality of the generated sequences. Summary of the Invention

[0005] In view of the above-mentioned problems in the prior art, the main purpose of this invention is to provide a TNF receptor sequence design model, AttenGNNDiff, based on a graph attention mechanism. This TNF receptor sequence design model, AttenGNNDiff, based on a graph attention mechanism, has stronger adaptability and generative capabilities in TNF receptor protein sequence design tasks.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] The present invention provides a tumor necrosis factor receptor sequence design model AttenGNNDiff based on a graph attention mechanism, comprising the following steps:

[0008] The graph neural network model is trained by constructing a sample protein radius graph as a training set;

[0009] Based on the diffusion model, protein sequence data is introduced into the training of the graph neural network model;

[0010] A denoising network is designed using the graph attention mechanism to update the node features and edge features of the protein radius graph separately, thereby denoising the graph neural network model that introduces protein sequence data.

[0011] Pre-training the denoised graph neural network model on the CATH4.2 dataset;

[0012] Tuning the pre-trained graph neural network model on a tumor necrosis factor receptor dataset to obtain a target graph neural network model, wherein the target graph neural network model is used for sequence design of a tumor necrosis factor receptor protein with a known main chain structure;

[0013] Wherein, the protein radius map is obtained by the following method:

[0014] The residues of the protein are taken as nodes, and the C, N, C α , atomic coordinates of O;

[0015] The obtained C α The atomic coordinates of are set as the node coordinates of the protein graph structure, and the nodes whose distance is less than the preset radius threshold are considered to be connected to establish a radius graph;

[0016] According to the acquisition of C, N, C α , the atomic coordinates of O construct the vectors u and v between atoms, and then construct the local coordinate system Q of node i i ;

[0017] According to the local coordinate system Q of the node i obtained iThe edge features and node features of the radius graph are extracted, and the extracted edge features and node features of the radius graph are stored in the form of a matrix.

[0018] As a further description of the above technical solution, the step of obtaining the preset radius threshold is specifically as follows: for each protein, the average distance between each node and its neighboring nodes in its K nearest neighbor graph is obtained, and then the average value d of the average distance of each node is calculated, and finally the obtained average value d is multiplied by the preselected radius coefficient μ1 to obtain the radius threshold; the calculation method of the average value d is as follows:

[0019]

[0020] Where L represents the protein main chain structure, K represents the number of neighbors of each node in the K-nearest neighbor graph; B represents the set of protein main chain residues, j represents any neighbor node of node i; N i represents the neighbor set of node i; d ij Represents the distance between node i and node j.

[0021] As a further description of the above technical solution, according to the acquisition of C, N, C α , the atomic coordinates of O construct the vectors u and v between atoms, and then construct the local coordinate system Q of node i i , specifically including:

[0022]

[0023] Q i =[b,n,b×n];

[0024] Among them, u means N points to C α vector; v represents C α The vector pointing to C; P x represents the x-coordinate of the atom; b represents the unit vector coplanar with u and v; n represents the unit normal vector perpendicular to the plane where u and v are located.

[0025] As a further description of the above technical solution, in step "according to the local coordinate system Q of the node i obtained i Extracting edge features and node features of the radius graph, wherein the extracted edge features and node features of the radius graph are stored in a matrix form, wherein the edge features of the radius graph include distance features, direction features, and torsion features between nodes;

[0026] The node features of the radius graph include distance features, direction features, and angle features of a single node.

[0027] As a further description of the above technical solution, the distance feature between nodes is constructed by calculating the radial basis distance between atoms, specifically including:

[0028] The Euclidean distance matrix and the number of radial basis function kernels are used to obtain the center value μ2 of the radial basis function kernel RBF and the standard deviation σ of each radial basis function kernel RBF:

[0029] μ2=linspace(D min ,D max ,numrbf);

[0030]

[0031] Among them, D min represents the minimum distance of the Euclidean distance matrix; D max represents the maximum distance of the Euclidean distance matrix; numrbf is the number of RBF kernels; the linspace function stratifies the distance interval into multiple uniform parts, the number of which is equal to numrbf;

[0032] The center value μ2 of the radial basis function kernel and the standard deviation σ of each radial basis function kernel are obtained and used to calculate the radial basis function kernel RBF distance matrix using the exp function:

[0033]

[0034] Introducing virtual atom C β , to utilize the virtual atom C β The coordinates of the set of atoms between two nodes [C, N, C α ,O,C β ] the radial basis distance between every two atoms; the virtual atom C β The coordinates in the local coordinate system are (x k ,y k ,z k ),satisfy:

[0035]

[0036] Among them, x k 、y k 、z k All are learnable parameters;

[0037] The calculated radial basis distances are concatenated to obtain the distance feature between the two nodes:

[0038]

[0039] Among them, ‖M i -N j ‖ represents the radial basis distance between the M atoms of node i and the N atoms of its neighbor node j, and Cat represents the concatenation operation of the matrix.

[0040] As a further description of the above technical solution, the direction feature between the nodes adopts the C of the central node. α Atoms point to neighbor nodes C, N, C α , O four atoms between the vector set construction:

[0041]

[0042] in, Represents the direction vector from atom N in node i to atom M in its neighbor node j.

[0043] As a further description of the above technical solution, the torsion feature between the nodes is captured by constructing a rotation matrix R between the nodes and converting it into a quaternion, specifically including:

[0044] Compute the 3D part of the quaternion:

[0045]

[0046] Among them, R xx 、R yy 、R zz are the diagonal elements of the rotation matrix R;

[0047] Calculate the sign of the antisymmetric component of the rotation matrix and use it to adjust the signs of x, y, and z to ensure the correct orientation of the four elements:

[0048]

[0049] Update x, y, z to: [x, y, z] = [x, y, z] × [S x , S y , S z ];

[0050] Compute the scalar part of the quaternion:

[0051]

[0052] Combine x, y, z, and w into a quaternion and normalize it to get the torsion characteristics between nodes:

[0053]

[0054] As a further description of the above technical solution, according to the obtained C, N, C α , the atomic coordinates of O and the virtual atom C β The atomic coordinates of the atoms are used to calculate the radial basis distances between atoms, and then the calculated radial basis distances are spliced ​​to obtain the distance characteristics of a single node

[0055]

[0056] The directional characteristics of a single node are represented by its C α Atoms pointing to C, N, C α , vector set construction of O atoms:

[0057]

[0058] Extract the sine and cosine values ​​of the six angles in the node as the angle features of a single node:

[0059]

[0060] Among them, α i Chemical bond The angle between i for The angle between i for The angle between Chemical bond The torsion angle; For chemical bonds The torsion angle; ω i For chemical bond C i —N i+1 The torsion angle.

[0061] As a further description of the above technical solution, in the step "using the graph attention mechanism to design a denoising network that can separately update the node features and edge features of the protein radius graph to denoise the graph neural network model that introduces protein sequence data", the neighborhood attention mechanism is used to update the node features:

[0062] In the lth layer of the graph neural network model, the attention weights of the central node and its neighboring nodes are:

[0063]

[0064] Among them, softmax represents the activation function; AttMLP represents the attention-seeking multi-layer perceptron, which consists of three fully connected layers and two ReLU activation layers; is the node feature of neighbor node j; is the feature of the edge connecting the central node i and its neighbor node j; represents the node feature of the central node i; || represents the splicing operation; represents the scaling factor; d k Represents the quotient of the node feature dimension and the number of attention heads;

[0065] After obtaining the attention weights of the central node i and its neighbor node j, the node feature of the central node i is updated to the weighted sum of the v values ​​of the neighbor node j:

[0066]

[0067] Among them, the v value of neighbor node j is obtained by concatenating the neighbor node's own features with the edge features and then mapping them through the multi-layer perceptron NodeMLP; NodeMLP consists of three fully connected layers and two GELU activation layers.

[0068] As a further description of the above technical solution, in the step of "using the graph attention mechanism to design a denoising network that can update the node features and edge features of the protein radius graph respectively to denoise the graph neural network model introduced with protein sequence data", the node features of the connected nodes are concatenated with the original edge features between the connected nodes to obtain the new edge features e' of the connected nodes. ij , thereby obtaining an edge feature matrix E' containing the preliminary updated features of all edges:

[0069]

[0070] in, Represents the node features of the central node i in the l+1th layer of the graph neural network model; Represents the edge features of the lth layer of the graph neural network model; Represents the node features of neighbor node j in the l+1th layer of the graph neural network model;

[0071] And use the self-attention mechanism to calculate the attention weight between each pair of edge features in the set of edges connected to the same node:

[0072] First, the new feature e' of the edge connecting the central node i and the neighbor node j ij Mapped to the query, key and value space through linear transformation, the corresponding representation Q is obtained ij , K ij and V ij :

[0073] Q ij =W q ·e' ij ;K ij =W k ·e' ij ; V ij =W v ·e' ij ;

[0074] Among them, W q , W k , W vLinear transformation matrices representing queries, keys, and values;

[0075] Then calculate the attention score between each pair of edge features to obtain the attention weight between each pair of edge features;

[0076] Secondly, the weighted sum method is used to calculate the edge e ij The l+1th feature of the network is used, and the layer normalization operation is performed on the feature matrix after edge update.

[0077] In summary, the outstanding effects of the present invention are:

[0078] 1. The graph-attention-based tumor necrosis factor receptor sequence design model, AttenGNNDiff, provided by this invention, replaces traditional protein K-nearest neighbor graph modeling with geometric distance-based radius graph modeling. This model demonstrates greater biological plausibility in capturing local protein topological relationships and offers significant advantages for sequence design of tumor necrosis factor receptor proteins with specialized structural domains.

[0079] 2. The graph-attention-based tumor necrosis factor receptor sequence design model AttenGNNDiff provided by the present invention introduces a neighborhood attention mechanism, which combines the feature information of the central node with the feature information of the edge. The information of the central node's neighbors is aggregated to the central node through the edge. This enhances the model's perception of the local geometric structure of the protein and improves the model's attention to the special structural domain features of proteins with specific functions, significantly improving the effectiveness of sequence design.

[0080] 3. In the graph-attention-based tumor necrosis factor receptor sequence design model AttenGNNDiff, the present invention introduces a self-attention-based edge feature update method that further exploits the high-dimensional nature of complex interactions between protein residues. This increases the weight of edge feature updates in updating protein structure graph features, improves their effectiveness, and provides important support for the model's adaptability to complex protein structures. The application of a diffusion model optimizes the integration of structural and sequence features through an implicit denoising process, significantly improving the quality of generated sequences.

[0081] 4. In the tumor necrosis factor receptor sequence design model AttenGNNDiff based on the graph attention mechanism provided by the present invention, a function-based prompt tuning algorithm is designed to fine-tune the model. By effectively embedding functional information, the model's design performance for tumor necrosis factor receptor sequence design is optimized, thereby improving the model's ability to learn the special structural domain characteristics of the tumor necrosis factor family. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0083] Figure 1 Schematic diagram of a radius graph modeling method for proteins in an embodiment of the present invention;

[0084] Figure 2 Schematic diagram of the local coordinate system of the central residue in an embodiment of the present invention;

[0085] Figure 3 Schematic diagram of the chemical bond torsion angles and the angles between chemical bonds of the protein main chain according to an embodiment of the present invention;

[0086] Figure 4 is an algorithm flow chart of the diffusion model used in the embodiment of the present invention;

[0087] Figure 5 This is a framework diagram of the AttenGNN model in an embodiment of the present invention;

[0088] Figure 6 This is a flowchart of prompting and tuning using tumor necrosis factor data in an embodiment of the present invention;

[0089] Figure 7 This is a boxplot comparing the effects of the AttenGNNDiff model on the CATH4.2 dataset with other models in an embodiment of the present invention;

[0090] Figure 8 This is a box plot of the recovery rate of proteins with different secondary structures using the AttenGNNDiff model in an embodiment of the present invention;

[0091] Figure 9 This is a box plot of the recovery rate of protein hydrophilic and hydrophobic residues by the AttenGNNDiff model in an embodiment of the present invention;

[0092] Figure 10 A residue confusion matrix heat map showing the overall differences between the sequence generated by the AttenGNNDiff model and the natural sequence residues in an embodiment of the present invention;

[0093] Figure 11 A protein backbone sequence generated by the AttenGNNDiff model in an embodiment of the present invention and a residue type distribution diagram of a natural protein backbone sequence;

[0094] Figure 12This is a structural comparison diagram of the structure folded into the A chain generated sequence of protein 1qwj in the embodiment of the present invention and the native structure;

[0095] Figure 13 This is a structural comparison diagram of the structure folded into by the A chain generation sequence of protein 2dkw in an embodiment of the present invention and the native structure;

[0096] Figure 14 This is a structural comparison diagram of the structure folded into the A chain generated sequence of protein 2lwy in an embodiment of the present invention and the native structure. DETAILED DESCRIPTION

[0097] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0098] In the description of the present invention, it should be noted that the terms "upper," "middle," "lower," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Those skilled in the art will understand the specific meanings of the above terms in the present invention in specific contexts.

[0099] The embodiment of the present invention provides a tumor necrosis factor receptor sequence design model AttenGNNDiff based on a graph attention mechanism, comprising the following steps:

[0100] The following steps are involved:

[0101] The graph neural network model is trained by constructing a sample protein radius graph as a training set;

[0102] Based on the diffusion model, protein sequence data is introduced into the training of the graph neural network model;

[0103] A denoising network is designed using the graph attention mechanism to update the node features and edge features of the protein radius graph separately, thereby denoising the graph neural network model that introduces protein sequence data.

[0104] Pre-training the denoised graph neural network model on the CATH4.2 dataset;

[0105] Tuning the pre-trained graph neural network model on a tumor necrosis factor receptor dataset to obtain a target graph neural network model, wherein the target graph neural network model is used for sequence design of a tumor necrosis factor receptor protein with a known main chain structure;

[0106] Wherein, the protein radius map is obtained by the following method:

[0107] First, the protein residues are taken as nodes, and the C, N, C α , atomic coordinates of O;

[0108] Then the obtained C α The atomic coordinates of C are set as the node coordinates of the protein graph structure, and the nodes whose distance is less than the preset radius threshold are considered to be connected to establish a radius graph. α The reason why the atomic coordinates of are set as the node coordinates of the protein graph structure is that each residue in the protein main chain is a collection of atoms, and the central carbon atom C α It is the core component of each residue and determines the basic skeleton of the protein main chain. α The change of atoms has a great influence on the overall structure of protein, and C α The spatial arrangement of atoms is closely related to the secondary structure of proteins. All the sequence design methods based on the main chain structure of proteins model proteins as directed K-nearest neighbor graphs, that is, the K residues closest to a certain residue are regarded as the neighbors of this residue, and this residue is called the central residue. The edges between the central residue and its neighbors are directed from the central residue to the neighbors. This modeling can ensure that each node in the protein graph has the same number of neighbors, so that the feature extraction and update of the nodes and edges of the graph can use matrix operations, thereby improving the learning and computational efficiency of the neural network model. However, proteins are a type of biological molecules and usually have irregular three-dimensional structures, which leads to different residue densities in the environment around each residue on the main chain. This density is an important condition for maintaining the stability of the protein structure and an important guarantee for the protein to perform its function. The modeling of the K-nearest neighbor graph makes the number of neighbors of each residue the same, because it cannot well reflect this structural feature of the protein molecule. In this embodiment, the protein main chain is modeled as a radius graph, that is, the residues within a certain radius threshold from the central residue are regarded as neighbors of the central residue. The modeling method is as follows. Figure 1 As shown in the figure, the red origin represents the central residue, the green origin represents the neighbors of the central residue, and the blue origin represents other residues. The radius of the red dotted circle is the radius threshold. The nodes within the radius threshold are the neighbors of the central node. The black solid line vector represents the real edge. The black dotted line in the figure represents the protein main chain, which reflects the order between residues and is irrelevant to the construction of the radius graph.

[0109] In order to continue to use matrix processing for subsequent graph processing, this example chooses to model the radius graph based on the K-nearest neighbor graph when K=30. For each protein, the average distance between each center node i and its neighbor nodes in its K-nearest neighbor graph is calculated, and then the average distance of each center node i is averaged, that is:

[0110]

[0111] Where L represents the protein main chain structure; K represents the number of neighbors of each node in the K-nearest neighbor graph; B represents the set of protein main chain residues, j represents any neighbor node of the central node i; N i represents the neighbor set of the central node i; d ij Represents the distance between node i and node j.

[0112] After obtaining the average value of the average distance of each central node of the protein main chain, the radius threshold is obtained by multiplying the pre-selected radius coefficient μ1 by it. In order to enable the model to accurately learn the density of protein residues, μ1 = 0.8 is taken in this implementation. This is the optimal radius coefficient obtained by sensitivity experimental analysis. For experimental results, please see the relevant explanation of the ablation of model performance and parameter sensitivity analysis below. For a protein main chain with a length of L, the distance between the residues in the main chain can be calculated, thereby obtaining the distance matrix D∈R B,L,L , where B is the batch size. Sorting the third dimension of the distance matrix and slicing the first K values ​​can obtain the distance matrix D between residues in the K nearest neighbor graph k ∈R B,L,k Adjacency list E of protein graph id ∈R B,L,k , use the distance matrix and radius threshold to obtain the mask matrix M∈R B,L,K ,Right now:

[0113] M=int(D k <0.8d);

[0114] Among them, the int function can convert Boolean values ​​into integers 0 and 1, and then use M to convert the adjacency list E id The adjacency matrix of the radius graph can be obtained by performing a mask operation, namely:

[0115] E' id =E id ⊙M;

[0116] Among them, E' id ∈R B,L,k is the adjacency matrix of the radius graph, which represents the index of the neighbors connected to each residue. The number of neighbors of each residue is K'≤K, E' idSimilar to the neighbor index vector matrix obtained after filling and splicing, it can represent the interaction relationship between nodes in the protein radius graph.

[0117] The modeled graph is represented as G = (V, E), where V = {v1,…,v N} represents each residue, E={e ij} i≠j Represents the edges connecting the nodes, that is, the relationships between residues.

[0118] Then, according to the obtained C, N, C α , the atomic coordinates of O construct the vectors u and v between atoms, and then construct the local coordinate system Q of node i i , so that the deep learning model can focus on the interaction between the microscopic atoms of each residue, as well as the interaction between the atoms of the residues and the more complex interaction between the residues as a whole as a set of atoms. More specifically, since hydrogen atoms are light, their role in determining the conformation of the residue can be ignored. In the composition, the role of hydrogen atoms is ignored, and only the coordinates of the four heavy atoms in the residue are considered, namely the carboxyl carbon atom C, the amino nitrogen atom N, and the central carbon atom C α and oxygen atom O. After obtaining the atomic coordinates of the four heavy atoms, we can construct the vectors u and v between atoms, where Indicates that N points to C α vector of Indicates C α The vector pointing to C; P x Represents the coordinate of atom x; then construct the unit vector b coplanar with u and v, and the unit normal vector n perpendicular to the plane where u and v are located, that is:

[0119]

[0120] Among them, b and n are perpendicular to each other, and the vector b×n is perpendicular to the plane where b and n are located. Therefore, b, n, and b×n are three perpendicular vectors, which meet the conditions for the formation of the coordinate system in the vector space. So, Figure 2 As shown, for node i, Q can be constructed i =[b,n,b×n] is the local coordinate system.

[0121] Afterwards, the local coordinate system Q of the node i can be obtained i Extract the edge features and node features of the radius graph, and store the extracted edge features and node features of the radius graph in the form of a matrix

[0122] Specifically, in this embodiment, in step "according to the local coordinate system Q of the node i obtained iExtracting edge features and node features of the radius graph, and storing the extracted edge features and node features of the radius graph in the form of a matrix. The edge features of the radius graph include distance features, direction features, and torsion features between nodes; the node features of the radius graph include distance features, direction features, and angle features of a single node.

[0123] When constructing the distance feature between two residues in a protein residue graph, directly using the Euclidean distance cannot effectively capture the complex nonlinear relationship in the protein structure. The Euclidean distance is a simple measure of the straight-line distance between two points, but in the protein structure, the similarity between residues does not only depend on their positional distance in three-dimensional space. The folding structure of the protein, the interaction between residues, and the influence of the local environment on the similarity all need to be taken into account. In contrast, the radial basis function can better measure the local similarity between residues by introducing a smooth Gaussian kernel function. The RBF kernel converts the similarity between two residues into a local measure by exponentially decaying the distance, which makes the similarity between the residues higher when the distance between them is closer, and the similarity decreases sharply when the distance is farther. This characteristic is more in line with the characteristics of local interactions and spatial organization in protein structure. Therefore, in this embodiment, the distance feature between the nodes is constructed by calculating the radial basis distance between atoms.

[0124] Assuming that the Euclidean distance matrix is ​​D, according to the spatial relationship between atoms in the protein molecule, this embodiment first sets the minimum distance D min =2 angstroms, maximum distance D max = 22 angstroms, where angstroms are the commonly used units for calculating interatomic distances, and 1 angstrom = 0.1 nanometers. min 、D max With the number of radial basis function kernels RBF, the center value μ2 of the radial basis function kernel and the standard deviation σ of each radial basis function kernel RBF can be obtained:

[0125]

[0126] Among them, D min represents the minimum distance of the Euclidean distance matrix; D max represents the maximum distance of the Euclidean distance matrix; numrbf is the number of RBF kernels; the linspace function stratifies the distance interval into multiple uniform parts, the number of which is equal to numrbf;

[0127] Then use the exp function to calculate the radial basis function kernel RBF distance matrix:

[0128]

[0129] The distance between two nodes in the protein backbone is the distance between two sets of atoms, and an atom in one set of atoms interacts with atoms in the other set of atoms. In order to more comprehensively represent the interaction between atoms, when calculating the distance between nodes, the C, N, and C α , O The RBF distances between the four atoms should be calculated. In order to express the implicit position information between the two nodes in addition to the position information of the four heavy atoms, this embodiment also introduces a virtual atom C β , its coordinates in the local coordinate system are (x k ,y k ,z k ), satisfying the following relationship:

[0130]

[0131] Among them, x k 、y k 、z k are all learnable parameters.

[0132] Get virtual atom C β After the coordinates of the two nodes are obtained, the atomic set [C, N, C α ,O,C β ]; then the calculated radial basis distances are concatenated to obtain the distance feature between two nodes i and j:

[0133]

[0134] Among them, ‖M i -N j ‖ represents the radial basis distance between the M atoms of node i and the N atoms of its neighbor node j, and Cat represents the concatenation operation of the matrix.

[0135] The directional characteristics between protein main chain residues are used to describe the relative orientation between residues. In the protein structure, the residues are not only close to each other, but their directionality (i.e., relative orientation) is crucial to the overall folding and functionality of the protein. In the protein residue graph, the directional characteristics of the central residue and its neighbors can be represented by a vector pointing from the central node to the neighboring node. Similarly, due to the complexity of the residue structure, the relative orientation between the residues cannot be well reflected by only the vector from the Cα atom of the central residue to the Cα atom of the neighboring residue. Therefore, specifically, in this embodiment, the directional characteristics between the nodes (residues) are represented by the Cα atom of the central node. α Atoms point to neighbor nodes C, N, C α , O four atoms between the vector set construction:

[0136]

[0137] Among them, Q i is the local coordinate system of node i, Represents the direction vector from atom N in node i to atom M in its neighbor node j.

[0138] Specifically, in this embodiment, the torsion characteristics between nodes are captured by constructing a rotation matrix R between the nodes and converting it into a quaternion. The quaternion encoding function is used to convert the rotation matrix into a quaternion, which can avoid the gimbal lock problem while improving computational efficiency. Specifically, it includes:

[0139] First calculate the three-dimensional part of the quaternion: the three-dimensional part x, y, z of the quaternion is determined by the diagonal elements R of the rotation matrix xx , R yy , R zz To calculate, the formula is as follows:

[0140]

[0141] Among them, R xx 、R yy 、R zz are the diagonal elements of the rotation matrix R.

[0142] Next, calculate the sign of the antisymmetric component of the rotation matrix and use it to adjust the signs of x, y, and z to ensure the correct direction of the four elements:

[0143]

[0144] After that, update x, y, z to: [x, y, z] = [x, y, z] × [S x , S y , S z ];

[0145] Then calculate the scalar part of the quaternion, the scalar part w is calculated by the diagonal elements of the rotation matrix:

[0146]

[0147] Finally, x, y, z, and w are combined into quaternions and normalized to obtain the torsion characteristics between nodes. The resulting torsional characteristics between nodes For subsequent protein sequence design and structure optimization:

[0148]

[0149] Specifically, in this embodiment, according to the obtained C, N, C α , the atomic coordinates of O and the virtual atom Cβ The atomic coordinates of the atoms are used to calculate the radial basis distance between atoms (the calculation method is the same as above and will not be repeated here), and then the calculated radial basis distances are spliced ​​to obtain the distance characteristics of a single node

[0150]

[0151] For a single node (residue), the central carbon atom plays an important role in the construction of its conformation. Therefore, the directional characteristics of a single node are represented by its C α Atoms pointing to C, N, C α , vector set construction of O atoms:

[0152]

[0153] C in residue α Atoms are connected to N atoms and C atoms through chemical bonds, and N atoms are connected to C atoms of adjacent residues through chemical peptide bonds. These chemical bonds form angles with each other, and the complex conformation of the protein backbone causes some chemical bonds to have torsion angles. The chemical bonds between residue atoms and their angles and torsion angles are as follows: Figure 3 As shown. The torsion angle and angle of the chemical bond jointly determine the overall three-dimensional structure of the protein and its dynamic properties. In the folding process of the protein, not only a single torsion angle or angle plays a role, but the interaction of these angles in space jointly affects the secondary structure (such as α-helix, β-folding, etc.) and tertiary structure (such as enzyme active site, binding site, etc.) of the protein and the ultimate biological function. Therefore, accurately describing these angular features is of great significance for improving the accurate description of the protein main chain structure and improving the accuracy of sequence generation. Therefore, in this embodiment, the sine and cosine values ​​of the 6 angles in the node are extracted as the angular features of a single node (residue):

[0154]

[0155] Among them, α i For chemical bonds The angle between i for The angle between i for The angle between For chemical bonds The torsion angle; For chemical bonds The torsion angle; ω i For chemical bond C i —N i+1 The torsion angle.

[0156] Specifically, this embodiment, taking into account the unique training mechanism of the diffusion model, feeds both noisy sequences and protein structures into the model for training. During the generation process, the model inputs consist solely of random noise sampled from a uniform distribution and the protein structure, without inputting protein sequence information. This approach effectively avoids data leakage while still incorporating sequence data into the model training process. This approach not only enables the model to generate protein sequences that match known structures, but also offers greater design flexibility and precision while ensuring sequence plausibility.

[0157] The core idea of ​​the diffusion model is to transform the data from a disordered noisy state into target data through a step-by-step denoising process. In the training of the diffusion model, the data in the training set are first denoised for t steps, where t is a value randomly sampled in the range [0, T1] and T1 is the total number of denoising steps. The data x after denoising is t It is then input into the model, and the model's optimization goal is to generate the corresponding noise-free data x0. After the model training is completed, the generation process will gradually restore the target data through a specific calculation formula and a denoising process from steps t to t-1.

[0158] Using the diffusion model for protein sequence generation has many advantages. First, through the process of gradual noise reduction, the diffusion model can effectively handle the uncertainty and complexity in the protein sequence and gradually recover high-quality sequences from the noise. Secondly, the training process of the diffusion model uses a step-by-step generation and denoising method, so that each step of generation can be optimized based on the existing sequence information, thereby improving the quality and stability of the generated sequence. In addition, the diffusion model can gradually learn the correlation characteristics of protein sequence and structure during the training process, which helps to generate protein sequences with higher functionality and structural rationality. The algorithm flow of the diffusion model is as follows Figure 4 As shown in Figure 2. During the diffusion process, after T1 steps of noise addition, the protein sequence can be considered sampled from a uniform distribution. During the generation process, the protein sequence probability matrix and protein structure sampled from the uniform distribution are input into the network for stepwise noise reduction. After a specified number of noise reduction steps, a noise-free protein sequence capable of folding into the desired structure is generated.

[0159] Since the protein residue sequence X aaFor discrete data, the diffusion process is completed by introducing a state transition matrix to perform state transformation. The state transition matrix promotes the transition of amino acid sequence states by providing the probability of transitioning from the current time step to the next time step. It reflects the possibility of transitioning from one residue type to another, and plays a key role in both the diffusion and generation processes. In the diffusion stage, the transition matrix is ​​iteratively multiplied by the protein sequence matrix to transfer its state, thereby adding noise to the protein sequence. As the diffusion time increases, the probability of the original residue type gradually decays, and eventually converges to a uniform distribution among all residue types. In the generation stage, the conditional probability p(x t-1 |x t ) is determined by the model prediction and the characteristics of the transfer matrix Q.

[0160] Given the biological specificity of amino acid substitutions, the transition probabilities between amino acids in natural proteins are not uniformly distributed. In order to make the state changes of proteins in the noise-adding process conform to the laws of amino acid state transitions in nature, the state transitions during the diffusion process cannot be defined as random directions. As an alternative, the diffusion process can reflect evolutionary pressure by utilizing alternative scoring matrices that can retain the functional structural characteristics or stability of proteins in wild-type protein families. Formally speaking, the amino acid substitution scoring matrix quantifies the rate at which various amino acids in a protein are replaced by other amino acids over time. Therefore, this embodiment uses a block permutation matrix (BLOSUM) as the state transition matrix of amino acids, incorporating the diffusion and generation processes. The matrix can identify conserved regions in proteins that are considered to have greater functional relevance. BLOSUM provides an estimate of the possibility of substitution between different amino acids based on empirical observations of protein evolution. By using this matrix to refine the transition probability, the generation space to be sampled is effectively reduced, so that the model's predictions converge to a meaningful subspace.

[0161] Use BLOSUM to set the state transfer matrix Q t After that, the protein residue sequence can be denoised. The protein residue sequence can be encoded into a matrix X = [x1, x2, ..., x n ],X∈R L×d , L is the sequence length, d = 20 represents the type of amino acids. Let X and Q t By multiplying them step by step, the vector of each residue can gradually obey the uniform distribution, thereby achieving the noise addition of the protein sequence. The noise addition process can be expressed as follows:

[0162]

[0163] in, is the transition probability matrix from the beginning to step t. When t approaches infinity, Obey the uniform distribution, so we can choose a sufficiently large T1 so that It approximately obeys a uniform distribution, that is, after adding noise for T1 steps, the distribution of the natural sequence obeys a uniform distribution. In this embodiment, T1=500 is selected.

[0164] Specifically, in this embodiment, during the model training phase, the data input to the denoising network is the sequence probability matrix after t steps of noise addition. and the protein backbone structure, where t is a random integer sampled in the range [0, T1]. To enable the denoising model to identify high-quality components of a noisy sequence, the sequence is masked based on the entropy of each residue, masking residues with high entropy and low quality. Since sequences are represented as probability matrices, the entropy of each residue in the sequence can be calculated, and high-entropy residues can be masked.

[0165] For each residue i in the protein backbone sequence, its entropy is calculated as follows:

[0166]

[0167] Among them, C = 20 is the number of amino acid categories, p ic is the probability that residue i corresponds to category C.

[0168] After obtaining the entropy H of the sequence, the conditional mask mask can be obtained and the mask operation can be performed on the sequence, namely:

[0169] Q th =quantile(H,θ);

[0170]

[0171] Among them, Q th is the quantile of entropy, which is determined by the parameter θ (e.g., 0.5 corresponds to the median). In this embodiment, θ=0.1.

[0172] After the noisy sequence is masked, it can be input into the denoising network together with the protein main chain structure. The goal of the denoising network is to transform the noisy sequence into Denoise to original sequence Therefore the loss of the model is defined as and The sequence generation problem is essentially a classification problem, that is, assigning one of the 20 amino acids to each position on the main chain. Therefore, the loss function is defined as the relationship between the residue type A of each node and the predicted probability p(X aa ) is the cross entropy loss between:

[0173]

[0174] Where L is the length of the protein backbone, i.e. the number of residues, and the final loss is the average of the losses of all samples.

[0175] During model training, the Adam method is used to update the gradient, and the learning rate is set to 5x10 -4 , and trained for 40 rounds. To prevent overfitting during training and enhance the robustness of the model, the model was regularized using dropout. The dropout value was set to 0.2, meaning that during each training step, 20% of the neurons in the neural network would be randomly disabled (set to 0).

[0176] Specifically, during the generation process, the input of the model is the residue sequence sampled from the uniform distribution and the natural structure of the protein. The sampled residue sequence can be regarded as the one obtained by performing T1-step denoising on the natural sequence. Therefore, by using the denoising network to perform T1-step denoising, the noise-free sequence can be restored, thereby realizing the generation of the protein sequence. In the denoising process from step t to step t-1, the probability distribution pθ(x t-1 |x t ) is predicted by the neural network based on the probability Make an estimate. Then we get the probability distribution of the sequence output by the network After that, the generation distribution of each iteration can be calculated:

[0177]

[0178] Among them, the posterior distribution q(x t-1 |x t ,x aa ) is calculated as follows:

[0179]

[0180] Among them, x aa It is sampled from the prediction results of the denoising network.

[0181] A significant drawback of the diffusion model is the speed of the generation process, which is relatively slow because the generation process usually consists of multiple incremental steps. To address this problem, this embodiment adopts a deterministic denoising implicit model (DDIM) to accelerate the generation process of the continuous variable diffusion generation model. DDIM operates based on a non-Markov forward diffusion process, where the generation of each step is conditioned on the current input rather than the output of the previous step. By setting the noise variance of each step to zero, the inverse generation process becomes completely deterministic under the condition of a given initial sample, thereby significantly improving the generation efficiency.

[0182] Similarly, according to the predicted x aa and the posterior distribution q(xt-1 |x t ,x aa ) can generate the probability pθ(x t-1 |x t ) can also be calculated by controlling p(x aa |x t ) to make the generation model deterministic. Therefore, this embodiment defines a multi-step generation process in the following way:

[0183]

[0184] Among them, the temperature T2 controls the randomness of the generation process, and the multi-step posterior distribution is:

[0185]

[0186] In this embodiment, k is set to 50, that is, 50 steps of denoising are performed on the sequence each time. When , sequence denoising is completed and the model outputs high-quality sequences.

[0187] Specifically, in order to enable the diffusion model to generate high-quality sequences, this embodiment uses the graph attention mechanism structure to design an efficient denoising model with reference to the overall architecture and self-attention mechanism of the Transformer structure. This model uses the graph attention mechanism to update the point features and edge features in the protein radius graph structure respectively, fully learning the protein structure information. Since the input of the denoising network is the noisy protein sequence and its structure, the denoising process can be regarded as a partial protein sequence design task. This task aims to use the known protein spatial structure and partial sequence information to generate a complete protein sequence through a deep learning model, that is, to learn a function:

[0188] F θ :X, Spartial→Sentire;

[0189] Among them, X represents the spatial structure of the protein main chain, and Spartial represents the known partial sequence.

[0190] Specifically, in this embodiment, a protein structure encoder module AttenGNN is used, which uses a neighborhood attention mechanism to update node features and a specific edge attention mechanism to update edge features; its structure is as follows Figure 5 shown.

[0191] The protein graph encoder consists of ten layers of AttenGNN modules. Given the relatively simple structure of the generated sequence compared to the three-dimensional structure of a protein, the decoder consists of only one fully connected layer. After encoding and decoding, the output is projected onto the [0, 1] interval using a softmax activation function, generating a sequence probability matrix. The value at each position represents the probability of belonging to a different amino acid type. The amino acid with the highest probability at each position is then selected as the final generated sequence.

[0192] More specifically, in this embodiment, a neighborhood attention mechanism is used to update node features, and by limiting the scope of attention calculation, only the interactions between each node and its neighbors are focused on. The node features of each central node and its neighbors, as well as the edge features connecting them, are taken as input and passed to a simple attention module (AttnMLP). This module calculates the attention weights between each pair of adjacent nodes based on the edge features and the features of the connected nodes. In addition, in order to better capture the diverse feature relationships between residues and improve the expressiveness of the model, this embodiment adopts a multi-head attention mechanism with four attention heads. In the lth layer of the graph neural network model, for the central node i, its attention weight with the neighbor node j is:

[0193]

[0194] Among them, softmax represents the activation function; AttnMLP represents the attention-seeking multi-layer perceptron, which consists of three fully connected layers and two ReLU activation layers; is the node feature of neighbor node j; is the feature of the edge connecting the central node i and its neighbor node j; represents the node feature of the central node i; || represents the splicing operation; represents the scaling factor; d k Represents the quotient of the node feature dimension and the number of attention heads. k When it is too large, the attention weight will grow exponentially as the number of network layers goes deeper, and the final value will fall into the area where the gradient of the softmax function is extremely small, resulting in the gradient disappearance problem. The introduction of the scaling factor can avoid this problem.

[0195] After obtaining the attention weights of the central node i and its neighbor node j, the node feature of the central node i is updated to the weighted sum of the v values ​​of the neighbor node j:

[0196]

[0197] Among them, the v value of neighbor node j is obtained by concatenating the neighbor node's own features with the edge features and then mapping them through the multi-layer perceptron NodeMLP; NodeMLP consists of three fully connected layers and two GELU activation layers.

[0198] In this way, the model can efficiently capture the relationships between local residues without having to calculate interactions between global nodes, thereby reducing computational complexity. Furthermore, this approach is more consistent with the characteristics of the local structure of the protein backbone, helping to better simulate the protein's spatial structure and function. Through this local attention calculation, the model of this embodiment can better capture the local structural features of the protein backbone, improving the accuracy of protein sequence design.

[0199] Specifically, in the graph neural network model of this embodiment, the edge feature update process first performs a preliminary update by combining the features of the connected nodes with the original edge features. That is, for each edge in the graph, the updated node features of the two nodes connecting the edge are first spliced ​​together with the original features of the edge to form a new edge feature representation. Suppose the node features of nodes i and j in the l+1 layer of the graph neural network model are and The edge features of the lth layer of the graph neural network model are Then the new feature of the edge is:

[0200]

[0201] In this way, we can obtain an edge feature matrix E, which contains the initially updated features of all edges. Given the dependencies and associations between edges of the same node in a protein graph structure, fully and effectively extracting these associations can enable more efficient dissemination of feature information. Therefore, this embodiment utilizes a self-attention mechanism to further update edge features to enhance their expressiveness.

[0202] In the edge update process, the classic self-attention mechanism is used to calculate the attention weights between each pair of edge features in the set of edges connected to the same node. For the new feature e' of the edge between node i and neighbor node j ij First, it is mapped to the space of query, key and value through linear transformation to obtain the corresponding representation Q ij , K ij and V ij :

[0203] Q ij =W q ·e' ij ;K ij =W k ·e' ij ; V ij=W v ·e' ij ;

[0204] Among them, W q , W k , W v Represents the linear transformation matrix of query, key and value. Then, the attention score between each pair of edge features is calculated. Usually, the similarity between query and key is calculated by dot product, and the attention scores of all edge features are normalized using softmax function to obtain the attention weight between each pair of edge features. For example, the edge feature e' ij With e' ik The attention weight is:

[0205]

[0206] α ik =softmax.Attn(Q ij ,K ik ) / ;

[0207] After obtaining the attention weights of other edges connected to node i, edge e ij The features of the l+1th layer of the graph neural network model are calculated using the weighted average sum method, namely:

[0208]

[0209] Among them, N i is the set of edges connected to node i.

[0210] In order to improve the optimization efficiency of the graph neural network model and increase the convergence speed, the feature matrix after edge update is also normalized. The calculation process is:

[0211]

[0212] Among them, E(x) contains the mean of the input x in the specified dimension, Var(x) represents the standard deviation of x in the specified dimension, ∈ is a very small integer introduced to prevent the denominator from being zero, and γ and β represent two learnable hyperparameters.

[0213] Compared to multi-layer perceptron (MLP) updating methods, using a self-attention mechanism for edge feature updates offers significant advantages. Traditional MLP methods typically process each edge feature independently, only considering the direct relationship between the current edge feature and the input feature during updates. This approach lacks modeling of the relationships between edges, failing to capture the dependencies between edge features within the graph and effectively integrating information between different edges. However, using a self-attention mechanism for edge feature updates allows for modeling the relationships between each pair of edge features, focusing on the contextual information of each edge feature within other edge features within the graph, thereby capturing more complex local and global feature dependencies. Therefore, this approach not only improves the model's expressive power but also enhances its ability to handle complex graph-structured data, making edge feature updates within protein graphs more consistent with actual biological relationships.

[0214] Experimental setup:

[0215] Specifically, the experimental part of this embodiment tests the common protein structure dataset and the tumor necrosis factor receptor protein structure dataset respectively. Since the number of tumor necrosis factor receptor datasets is relatively small, the model needs to be pre-trained on the CATH4.2 dataset with a larger data volume, and then fine-tuned using the tumor necrosis factor receptor dataset, so as to achieve the model's high-precision design capability for tumor necrosis factor receptor sequences. Since the sequence of natural proteins can spontaneously fold into natural structures, the higher the similarity between the sequence generated by the model and the natural sequence, the greater the probability that the generated sequence will fold into the target structure. Therefore, the sequence recovery rate indicator is introduced to reflect the similarity between the generated sequence and the natural sequence. At the same time, this embodiment performs a number of downstream analyses on the CATH4.2 dataset to conduct a more comprehensive evaluation of the generated sequence. A detailed description is given below.

[0216] (1) Protein structure dataset

[0217] 1. CATH Dataset

[0218] The CATH dataset is a resource focused on protein structural classification, aiming to efficiently classify and annotate proteins based on their three-dimensional structures. The core concept of the CATH dataset is to categorize proteins hierarchically based on their spatial structure and evolutionary history, thereby helping scientists understand protein functions and evolutionary relationships. Its main feature is its four-tier structural classification system, with each tier providing progressively more detailed information, from initial folding types to more specific homologous superfamilies.

[0219] The dataset used in this example is version 4.2 of the CATH dataset. After removing protein data containing unknown residues, the dataset contains 19,751 protein structures, including 18,023 protein structures in the training set, 608 protein structures in the validation set, and 1,120 protein structures in the test set.

[0220] 2. TS50 dataset

[0221] The TS50 dataset was constructed by Li et al., the authors of the SPIN method. It contains 50 proteins with sequence identity less than 30% and a resolution greater than 3.0 Å. It is a commonly used test set for inverse folding models.

[0222] 3. Tumor Necrosis Factor Receptor Dataset

[0223] To verify the performance of the proposed model in the design of various tumor necrosis factor receptor sequences, this example selected different tumor necrosis factor receptor-related proteins as a test set, and independently tested each type of receptor protein dataset. These include TNF Receptor 1 (TNFR1), TNF Receptor 2 (TNFR2), TNF Receptor Superfamily Member 1lb (RANK), TNF Receptor Superfamily Member 6 (Fas, CD95), Death Receptor 3 (DR3), TNF Receptor Superfamily Member 13 (BAFF-R), and Glucocorticoid-Induced TNF Receptor (GITR). By searching for these tumor necrosis factor receptor family members on the PDB website, the relevant protein data can be obtained. The relevant protein data is downloaded, the protein backbone is split, and the data is cleaned and deduplicated to obtain the relevant datasets for each type of receptor protein family member. The detailed information of the dataset is shown in Table 1 below.

[0224] Table 1 Tumor necrosis factor receptor dataset

[0225]

[0226] 4. Tumor Necrosis Factor Superfamily Dataset

[0227] The homologous domains in the tumor necrosis factor receptor family have a high degree of similarity to the cysteine-rich domains in tumor necrosis factor, and other proteins in the tumor necrosis factor superfamily also exhibit similar structural features. In order to enhance the stability and robustness of the fine-tuned model and improve the effect of prompt tuning, this embodiment adds data on the tumor necrosis factor superfamily and its related proteins to the fine-tuning training set as a supplement. These proteins share similar structural features, which helps the model better understand the sequence design rules of the tumor necrosis factor receptor family. In the data preprocessing stage, the structure of the relevant protein is first downloaded from the PDB database and split into single-chain forms. Next, a pairwise sequence alignment is performed to remove proteins with a sequence similarity of more than 50%, and proteins for which GO terms cannot be queried are excluded. Finally, a data set of 158 protein main chains is obtained. By combining it with 70% of the data in the receptor data set, a training set for model fine-tuning can be obtained.

[0228] (2) Evaluation indicators for protein sequence design based on backbone structure

[0229] 1. Sequence-level evaluation metrics

[0230] (1) Sequence recovery rate: In the task of generating protein sequences based on the main chain structure, the sequence recovery rate is an important evaluation indicator. It measures the degree of match between the sequence generated by the model and the target sequence (usually the natural sequence) at a specific position, reflecting the model's ability to capture real sequence information. A high sequence recovery rate indicates that the model can fully understand and utilize the structural information of the protein main chain to generate a reasonable sequence, which is crucial for verifying the effectiveness and design capabilities of the model. At the same time, the sequence recovery rate can also directly reflect the accuracy and reliability of the sequence generated by the model, providing a strong feedback basis for further optimization of the model. In practical applications, this indicator is of great significance for evaluating the potential of the generated sequence in terms of functionality and structural fidelity, especially when designing new proteins or modifying existing proteins. It can serve as one of the core references for evaluating the quality of the generated model. The sequence recovery rate is calculated as follows:

[0231]

[0232] Among them, s gen is the generated protein sequence, s nat The native protein sequence.

[0233] (2) Perplexity: Perplexity is a commonly used evaluation metric in natural language processing, especially for language models. It measures the degree of "confusion" of the model's prediction of a given sequence, or in other words, the difficulty of the model in predicting the next element (the next amino acid in a protein sequence). Perplexity is the "inverse" measure of the probability distribution predicted by the model. Perplexity reflects the average prediction error of the model for the sequence. The lower the perplexity, the stronger the model's prediction ability and the higher the quality of sequence generation. It is calculated as follows:

[0234]

[0235] Where L is the length of the protein main chain sequence, p(x i ) is the model's correct label x at position i i The predicted probability of .

[0236] 2. Structural-level evaluation indicators

[0237] (1) Recovery rate of protein sequences containing different types of secondary structures: The recovery rate of protein secondary structures is an important indicator for evaluating the performance of protein reverse folding models. This indicator measures whether the model can accurately predict the corresponding amino acid sequence when the input is a protein structure of all α helix, all β fold, or a mixture of the two. A higher secondary structure recovery rate indicates that the model can effectively capture different secondary structure features and generate amino acid sequences that match the original sequence, thereby reflecting its generalization ability on proteins of different structural types. According to the two common secondary structures of proteins (α helix and β fold), the protein test set is divided into three categories: proteins containing only α helix, proteins containing only β fold, and proteins containing both structures, and their recovery rates are calculated separately.

[0238] (2) Recovery rate of hydrophilic and hydrophobic residues in proteins: The recovery rate of hydrophilic residues refers to the degree of recovery of hydrophilic amino acids (such as D, E, K, R, H, N, Q, S, T) in the recovered protein sequence. It indicates the proportion of residues that are hydrophilic amino acids in the natural sequence that are successfully recovered in the generated sequence. The recovery rate of hydrophobic residues refers to the degree of recovery of hydrophobic amino acids (such as A, C, F, I, L, M, P, V, W, Y) in the recovered protein sequence. It indicates the proportion of residues that are hydrophobic amino acids in the natural sequence that are successfully recovered in the generated sequence. Hydrophilic amino acids are usually located on the surface or exposed areas of proteins and participate in the interaction with water molecules and other polar molecules. A high recovery rate of hydrophilic residues indicates that the model performs well in recovering the surface and functional areas of proteins. Hydrophobic amino acids are usually located inside proteins to form a hydrophobic core, which is crucial for maintaining the three-dimensional structure and stability of proteins. A high hydrophobic residue recovery rate indicates that the model is able to accurately restore the protein's intrinsic structure and maintain the formation of the hydrophobic core. The recovery rate of hydrophilic and hydrophobic residues reflects the structural and functional preservation of the generated protein sequence and the recovery of the hydrophobic core, and is an important indicator of protein sequence recovery quality.

[0239] (3) The difference between the folding structure of the generated sequence and the natural sequence: This metric focuses on the accuracy and rationality of the three-dimensional structure of the protein sequence generated by the model, especially the difference in folding structure from the real, natural protein sequence. Since the function of a protein is usually closely related to its three-dimensional structure, evaluating the difference between the folding structure of the generated sequence and the natural sequence is an important criterion for the effectiveness of the reverse folding model. This metric is mainly quantified by the RMSD value and TM-score value of the folding structure of the generated sequence and the natural sequence.

[0240] RMSD calculates the root mean square difference in distance between corresponding atoms in two structures. The formula is:

[0241]

[0242] Where N is the number of corresponding atom pairs, x i and y i are the corresponding atomic coordinates.

[0243] TM-score (Template Modeling Score) is a commonly used metric for evaluating the structural similarity of two proteins. The calculation takes into account the distance differences between atomic pairs between the two structures and is defined by the following formula:

[0244]

[0245] Where L is the number of amino acids in the reference structure (usually the length of the target protein), N is the number of overlapping amino acid pairs, i.e., the amino acid pairs that can be effectively aligned between the two structures, and d i is the C of the i-th pair of amino acids α The distance between atoms (usually C α The distance between atoms). d0 is a constant, usually taken as d0 = 1.24 × (L-1) 1 / 3 -1.8, which represents the average distance between amino acids within the same protein.

[0246] (III) Model fine-tuning using the tumor necrosis factor dataset

[0247] In order to improve the performance of the model in the tumor necrosis factor receptor sequence design task, this example uses the prompt tuning method to fine-tune the model on the tumor necrosis factor receptor dataset. The fine-tuning process is as follows: Figure 6 shown.

[0248] Tumor necrosis factor receptors belong to a family of proteins with specific biological functions. Each member possesses unique functional properties, which can be described using Gene Ontology (GO) terms. GO terms are a standardized set of terms widely used in biology to describe protein functions, cellular localization, and biological processes. By incorporating these GO terms as part of the model input, the protein sequence generation process can be further guided based on structural features, thereby improving the model's effectiveness in the TNF receptor sequence design task.

[0249] Taking into account the functional differences between different proteins, the number and type of GO terms may be different in each protein. In order to unify the data format, this embodiment counts the number of GO terms for all proteins in the training data set, and uses the maximum number of GO terms as the length of the prompt vector. For proteins with an insufficient number of GO terms, their prompt vectors will be filled with 0 to ensure that all proteins have prompt vectors of the same length when input. Before being input into the model, the GO terms will first be embedded through an embedding layer with learnable parameters to generate embedded GO term features, and then the embedded features will be converted into soft prompts through multi-layer perceptron MLP encoding. The MLP structure used in this embodiment consists of two fully connected layers and one activation layer. Then, the generated soft prompts are spliced ​​with the node feature matrix of the protein structure diagram to obtain an input matrix containing GO term prompts. Using this matrix as the input of the baseline model can introduce the functional information of the protein into the node features, thereby guiding sequence generation.

[0250] The fine-tuning process consists of two main stages. First, the model is pre-trained on the CATH4.2 dataset to learn the basic structural and sequence features of the protein. Then, the model is fine-tuned on the TNF receptor dataset. During fine-tuning, all parameters of the baseline model are frozen, and only the parameters of the hint embedding layer are optimized to enhance the model's ability to learn the functional characteristics of the TNF receptor. Specifically, the hint embedding vector is combined with the protein's structural features and processed through the model's encoder to generate the optimized TNF receptor sequence.

[0251] Considering the small size of the dataset, this example sets the fine-tuning rounds to 5 and the learning rate to 1×10 -5 , to avoid overfitting and ensure a stable training process.

[0252] Experimental results

[0253] (I) Generating sequence quality assessments on common protein datasets

[0254] Table 2 below shows the test results of AttenGNNDiff and other mainstream classic sequence design models on the CATH4.2 dataset. To make the results more clear and intuitive, Figure 7 The corresponding boxplots are shown. The methods compared in the chart are all existing research methods, both domestically and internationally, and therefore will not be detailed here. Compared to the baseline model, the model achieved the highest sequence recovery rate, outperforming existing mainstream classical methods, demonstrating its ability to more accurately capture the mapping relationship between protein backbone structure and sequence. Furthermore, the model achieved excellent perplexity, reaching the lowest value. This reflects the high consistency between the sequence distribution generated by the model and the true distribution, demonstrating its superior performance in protein sequence design tasks. The boxplot results show that the median of the boxes for the method in this example is significantly higher than that of other models, and the box length is relatively small, indicating that the recovery rate is concentrated at a high level and the data distribution is relatively stable. Specifically, the shorter box length indicates a smaller fluctuation range in the recovery rate and the model's consistent performance across different protein samples. Furthermore, the higher median reflects the superior performance of this method in terms of recovery rate, enabling it to generate relatively reasonable protein sequences in the majority of test samples. This further demonstrates its effectiveness and robustness in protein sequence recovery tasks. This result not only verifies the efficiency of the method but also demonstrates its broad applicability in complex protein sequence design.

[0255] Table 2 Recovery rate and median perplexity of AttenGNNDiff and other classic models on the CATH4.2 test set

[0256]

[0257] Table 3 below shows the test results of AttenGNNDiff and other mainstream classic sequence design models on the TS50 dataset. In experiments on the TS50 dataset, the model also achieved excellent results, achieving the highest recovery rate and the lowest perplexity, consistent with its performance on the CATH4.2 dataset. This result validates the model's robustness and generalization ability, demonstrating that it can consistently generate high-quality sequences for both common and complex protein structures. This excellent performance also demonstrates the model's broad applicability in protein sequence design tasks.

[0258] Table 3 Recovery rate and median perplexity of AttenGNNDiff and other classic models in the TS50 test set

[0259]

[0260] In the protein sequence design task, this example further evaluated the sequence recovery rates for short-chain and single-chain proteins, with the results shown in Table 4. Short-chain proteins are proteins with a length of less than 100 residues, while single-chain proteins are proteins composed of a single peptide chain. Due to their shorter length and simpler structure, the sequence recovery rates of short-chain proteins can reflect the model's ability to generate simple targets. On the other hand, since single-chain proteins are not affected by interchain interactions, their recovery rates more directly reflect the model's ability to capture the relationship between main-chain conformation and sequence.

[0261] Through detailed analysis of these two types of proteins, we can see that the model's recovery rate for short-chain proteins is higher than the baseline model, demonstrating its excellent performance under simple structures; the recovery rate for single-chain proteins is also significantly improved, indicating the model's strong ability to model independent chain conformations. This analysis of different categories verifies the model's broad applicability and demonstrates its excellent performance in designing specific proteins, such as short-chain active peptides or single-chain functional proteins.

[0262] Table 4 Recovery rate and perplexity of short chains and single chains of AttenGNNDiff and other classic models in the CATH4.2 test set

[0263]

[0264] In order to evaluate the model's ability to design sequences for proteins with different structures, this example divides the CATH4.2 test set into three categories: full α-helix, full β-sheet, and α / β mixed proteins, and calculates the model's sequence recovery rate for each category. To make the experimental results more intuitive, the results are presented in the form of box plots, as shown in the figure below. Figure 8The experimental results reflect the model's adaptability across diverse structural categories. The low recovery rate for α-helical structures, due to their strong local dependencies and high sequence redundancy, allows the model to achieve diverse generation that balances functionality and stability through flexible design. The high recovery rate for β-pleated structures demonstrates the model's ability to capture long-range sequence dependencies and excel in modeling structure-sequence relationships. The highest recovery rate for α / β hybrid structures further demonstrates the model's superior performance in complex topologies, enabling simultaneous optimization of both short- and long-range sequence dependencies. This refined classification analysis fully validates the model's broad applicability and superior performance in complex structures.

[0265] The experiment on the restoration effect of the model on hydrophobic and hydrophilic residues is as follows Figure 9 As shown, it reflects the unique strategy of the model in generating sequence residues.

[0266] The high recovery rate of hydrophobic residues indicates that the model has a strong ability to capture the formation patterns of the hydrophobic core of proteins. The hydrophobic core is a key driving force for protein folding. By preferentially generating hydrophobic residues, the model effectively enhances the stability of the target conformation. The low recovery rate of hydrophilic residues may be due to the model's targeted adjustments to surface residues during the generation process. Hydrophilic residues are distributed on the protein surface, and their location and functional requirements are highly dependent on the specific characteristics of the target structure. When generating sequences, the model may have broken away from the constraints of the native sequence and optimized the distribution of hydrophilic residues, making the generated sequences more functional and diverse.

[0267] This distribution pattern demonstrates that the model is able to prioritize the regions most important for the target conformation during generation while adapting to the actual needs of sequence distribution. This design strategy not only improves the structural stability of the generated sequences but also provides more possibilities for optimizing sequence function.

[0268] To show the overall difference between the residues in the sequence generated by the reaction model and the natural sequence, this example compares the contents of 20 residues in the generated sequence and the natural sequence, and presents them in the form of a residue confusion matrix heat map and distribution map, as shown in the figure. Figure 10-11 shown.

[0269] Figure 10The following is a residue confusion matrix heatmap, with native protein residue types on the horizontal axis and generated protein residue types on the vertical axis. Darker cells in the heatmap indicate a higher proportion of residues that have been decoded into the corresponding native side chain types. Inspection of the heatmap clearly shows a high degree of consistency between the generated residue types and those found in the native protein sequence. In particular, the darkest colors are found at positions corresponding to glycine (G) and proline (P), indicating that the model accurately generated these residues. This demonstrates that the model effectively captures the specific roles of glycine and proline in protein structure, particularly in fixed-backbone sequence design tasks, where the specificity and structural requirements of these two amino acids are particularly prominent. Furthermore, the shaded area in the heatmap (the diagonal line running from the upper left corner to the lower right corner) also appears significantly darker, indicating a high degree of match between the native and generated residues. This further demonstrates that the model effectively maintains the target protein's amino acid distribution and types. In particular, at positions for some common hydrophobic and polar amino acids, the generated residues have good overlap with the native sequence.

[0270] Figure 11 The distribution bar graph of residue types between the model-generated sequence and the natural sequence shows the relative difference in the content of each residue between the generated sequence and the natural sequence.

[0271] The amino acid distribution of the generated sequences showed a significantly higher proportion of glutamic acid (E) and leucine (L) residues compared to the native sequences, while the proportion of methionine (M) and glutamine (Q) residues was relatively low. This distribution reflects the model's precise adaptability to the target structure and the rationality of the sequence design.

[0272] Glutamic acid is a negatively charged hydrophilic residue that is usually distributed on the surface of proteins and forms interactions with solvents (such as water molecules). The increase in the proportion of glutamate in the generated sequence indicates that the model tends to design more hydrophilic residues in the surface area of ​​the protein to improve the solubility and surface stability of the target protein. This optimization can enhance the stability and functional performance of the target protein under physiological conditions. Leucine is a typical hydrophobic residue that is usually located in the hydrophobic core of the protein and participates in the formation of a stable folded structure. The increase in the proportion of leucine in the generated sequence indicates that the model has an accurate ability to capture the formation rules of the hydrophobic core and gives priority to the residue types that contribute most to folding and stability. This design strategy further illustrates the excellent performance of the model in optimizing folding stability.

[0273] In contrast, the reduction in methionine and glutamine reflects the model's streamlined optimization of complex residue distributions. While methionine is a primed amino acid, it is generally not essential for the core structure of the target protein. Glutamine, a residue with a polar side chain, may rely on specific geometric conformations for its function and distribution. By reducing the proportion of these complex residues, the model avoids redundant design found in natural sequences, resulting in more concise sequences with greater adaptability in terms of function and stability.

[0274] Furthermore, this distributional property demonstrates that the model can transcend the constraints of natural sequence evolution and focus on the practical requirements of the target structure. For example, in specific application scenarios, the amino acid distribution of the generated sequence is more consistent with the functional optimization goal, making it more feasible and efficient in experimental verification and practical application.

[0275] (II) Evaluation on multiple tumor necrosis factor receptor datasets

[0276] This example tested the fine-tuned AttenGNNDiff model against other classic models on a sequence generation task on the Tumor Necrosis Factor Receptor dataset. The performance of the sequences generated by the different models was compared in terms of recovery rate and perplexity. The experimental results are shown in Tables 5 and 6. The experimental results show that the AttenGNNDiff model based on the hint-based fine-tuning method exhibits significant advantages on all datasets.

[0277] Table 5 Recovery rates of AttenGNNDiff and other classic models after fine-tuning on various tumor necrosis factor receptor datasets

[0278]

[0279] Table 6 The perplexity of AttenGNNDiff and other classic models after fine-tuning on various tumor necrosis factor receptor datasets

[0280]

[0281]

[0282] All models fine-tuned with hint tuning achieved sequence recovery rates exceeding 50%. In particular, the model combining AttenGNNDiff with hint tuning achieved the highest sequence recovery rate and the lowest perplexity across all test datasets. Sequence recovery rates exceeded 60% for all datasets, demonstrating that hint tuning can significantly improve the model's sequence generation quality and better preserve the structural features of protein sequences. This result demonstrates the effectiveness of hint tuning in the tumor necrosis factor receptor sequence generation task, helping the model extract more accurate functional features from existing structural information, thereby generating more reasonable protein sequences.

[0283] Experimental results demonstrate that hinted tuning significantly improves performance in protein sequence generation tasks, particularly in the tumor necrosis factor receptor (TNFR) sequence design task. Compared to traditional general tuning methods, hinted tuning effectively improves sequence recovery and leads to higher accuracy and reliability in generation tasks. These results demonstrate that hinted tuning significantly enhances the model's sequence generation quality and function extraction capabilities.

[0284] (III) Comparison of the differences between the folded result and the original structure of the generated sequence

[0285] In the protein structure comparison experiment, this example compares the predicted structure of the generated sequence with the target reference structure. Figure 12-14 As shown in the figure, the structure prediction tool used to generate the sequence was SWISS-MODEL, and the structure visualization software was PyMOL. Further analysis of three representative structures: an α / β hybrid, an all-α helical structure, and an all-β sheet structure. Visualization showed a high degree of overlap between the generated structures and the target structures. Both numerical indicators and visualization results verified the accuracy of the model.

[0286] The α / β hybrid structure combines the local dependence of the α helix with the long-range interaction of the β sheet, forming a complex secondary structure. The generation of this structure requires not only capturing the relationship between the local and global aspects but also accurately reflecting the synergistic effects between different structural units. Figure 3-12This image shows a structural alignment of the generated sequence folding into the native structure of chain A of the α / β hybrid protein 1qwj. The green structure represents the native structure, while the yellow structure represents the generated sequence. The RMSD for both structures is 0.931, and the TM-score is 0.9757, indicating a high degree of overlap between the generated sequence and the native structure. Visualization reveals that the generated hybrid structure maintains a smooth transition between the α-helix and β-sheet regions, and key spatial arrangement features, such as the hydrophobic core, are fully consistent with the target structure. Furthermore, the generated sequence captures the geometric properties of both the helical and β-sheet regions, demonstrating the model's broad adaptability for diverse structural design applications. This high consistency further demonstrates the model's robust global optimization capabilities, enabling it to generate highly plausible sequences for complex targets.

[0287] As a regular secondary structure, the stability of the α-helix structure depends on the short-range interactions of local hydrogen bonds and the regular distribution of amino acid side chains. Figure 13 This image shows a structural alignment of the generated sequence folded into the native structure of protein 2dkw's A chain. The pink structure represents the native structure, while the white structure represents the generated sequence folded into the native structure. The RMSD (RMSD) and TM-score (TM-score) of the two structures are 0.870 and 0.948, respectively, demonstrating a high degree of similarity between the generated sequence and the native structure. The visualization results show that the α-helix of the generated sequence is highly consistent with the coiling and directionality of the reference structure, accurately reproducing the helical properties of the target structure. This demonstrates the model's strong ability to model short-range dependent features (such as bond angles and dihedral angles) and local geometric distributions. This further validates the model's exceptional performance in sequence optimization for a single type of regular structure, enabling reliable generation of physically consistent sequences.

[0288] The formation of β-sheets depends on long-range interactions, especially the hydrogen bond network between parallel or antiparallel β-strands and the matching distribution of hydrophobic residues. Figure 14 This is a structural alignment of the generated sequence folded into the native structure of protein 2lwy's A chain. The green structure represents the native structure, while the blue structure represents the generated sequence. The RMSD between the two structures is 0.096, and the TM-score is 0.9926. The high consistency of the generated structures demonstrates that the model accurately captures long-range dependencies and the unique topological properties of β-sheets. Visual analysis shows that the orientation, spacing, and interactions of the β-strands in the generated structures fully conform to the target conformation. This result not only validates the model's reliability in long-range structural design but also demonstrates its ability to model complex interactions.

[0289] The low RMSD and high TM-score results not only demonstrate the reliability and adaptability of the model-generated sequences under the target structure, but also reflect the universality of the model in a wide range of protein structures and its potential biological application value.

[0290] (IV) Ablation of model performance and parameter sensitivity analysis

[0291] To explore the impact of the various modules of AttenGNNDiff on the final performance of the model, this example conducted ablation experiments on the CATH4.2 dataset. Ablation studies were performed on the radius graph modeling approach, the neighborhood attention mechanism for node updates, and the edge attention mechanism for edge updates. For the ablation of radius graph modeling, this example replaced the radius graph modeling with a conventional K-nearest neighbor graph; for the neighborhood attention mechanism, it was replaced with the classic Transformer that uses Q, K, and V to solve attention; and for the ablation of the edge attention mechanism, the edge feature update layer was replaced with three fully connected layers. The experimental results are shown in Table 7.

[0292] Table 7 Ablation experiment results

[0293]

[0294] When the radius graph was replaced by the K-nearest neighbor graph, the sequence recovery rate dropped from 52.19% to 51.35%. This shows that radius graph modeling has greater advantages in the model. The K-nearest neighbor graph defines node relationships by selecting a fixed number of neighbors, but in complex geometric structures such as proteins, a fixed number of neighbors may lead to the loss of key structural information, especially in areas with uneven local density distribution. In contrast, the radius graph selects neighbors based on geometric distance and can more accurately depict the local residue relationships of proteins, especially in sparse or dense areas. Therefore, the use of the radius graph not only improves the accuracy of local feature expression, but also enhances the network's overall perception of protein geometry, which is an important factor in improving model performance.

[0295] After replacing the improved neighborhood attention mechanism module with the classic Transformer module, the sequence recovery rate dropped from 52.19% to 50.34%, the largest decrease in recovery rate in the experiment. This demonstrates that the neighborhood attention mechanism module is a core module of the denoising network and outperforms the classic Transformer module. The classic Transformer's attention calculation method only calculates weights by the dot product of the node's query value and key value, ignoring the importance of edge features in the protein graph structure. The improved module, however, introduces edge features, concatenating the target node, neighbor nodes, and edge information, and then encoding them using a multilayer perceptron, enabling the network to explicitly perceive the local topological structure of the protein. This approach not only enhances the ability to model local interactions but also effectively suppresses interference from noise information. Ablation experiments clearly demonstrate the key role of this module in capturing the local geometric properties of proteins and improving sequence design performance.

[0296] When the edge feature update module was replaced with a multilayer perceptron (MLP) instead of an edge attention mechanism, the sequence recovery rate dropped from 52.19% to 51.01%, demonstrating that the self-attention mechanism is more effective at updating edge features than the MLP. The MLP-based edge feature update method only concatenates the original edge features with updated node features, ignoring the interactions between edges. The edge attention mechanism, on the other hand, achieves deep modeling of edge features by calculating attention between edges, enabling the model to capture higher-dimensional protein structural features. The introduction of this module significantly enhances the model's ability to represent complex protein interactions, providing important support for improving model performance. The performance degradation after removing this module further demonstrates the importance of edge feature updates for enhancing protein sequence design.

[0297] Ablation experiments clearly demonstrate the contribution of each module to model performance, validating the key roles of radius graph modeling, neighborhood attention, and edge attention modules in improving sequence recovery rate and accuracy. These results provide valuable guidance for model optimization and further demonstrate the potential of graph neural network-based protein sequence generation methods in capturing local structure and global interactions.

[0298] During data processing, the radius threshold setting plays a crucial role in the effectiveness of protein structure modeling. To explore the impact of different radius coefficients on model performance, this paper conducted a sensitivity analysis, adjusting the radius coefficient from 0.4 to 1.3 in steps of 0.1. Model training and validation were performed at each coefficient value. The experimental results are shown in Table 8. By analyzing the model performance under different radius coefficients, it was concluded that the model performance was optimal when the radius coefficient was 0.8.

[0299] Table 8 Effect of radius coefficient on recovery rate

[0300]

[0301] When the radius coefficient is set to 0.4, the connection information between the central residue and its neighbors is relatively sparse due to the small radius threshold, which makes the model unable to fully extract the topological structure information of the protein main chain. At this time, the structural information provided to the model is too brief, which affects the model's ability to capture local structural features. When the radius coefficient is too large (such as set to 1.3), the radius graph is almost the same as the K-nearest neighbor graph, and the adjacency relationship is completely determined by the K-nearest neighbor graph. As a result, the model lacks a detailed description of the density of the protein main chain, which in turn affects the effective learning of the overall structure. Therefore, it is crucial to choose a suitable radius coefficient. Sensitivity analysis experiments confirmed that when the radius coefficient is 0.8, the model can ensure sufficient local structural information while not oversimplifying or over-complicating the topological structure of the protein. Therefore, this paper selects a radius coefficient of 0.8, so that the model can effectively extract protein main chain structural information.

[0302] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any changes, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism, characterized by: The following steps are involved: The graph neural network model is trained by constructing a sample protein radius graph as a training set; Based on the diffusion model, protein sequence data is introduced into the training of the graph neural network model; A graph attention mechanism is used to design a denoising network that can update the node features and edge features of the protein radius graph separately, so as to denoise the graph neural network model that introduces protein sequence data. Pre-training the denoised graph neural network model on the CATH4.2 dataset; Tuning the pre-trained graph neural network model on a tumor necrosis factor receptor dataset to obtain a target graph neural network model, wherein the target graph neural network model is used for sequence design of a tumor necrosis factor receptor protein with a known main chain structure; Wherein, the protein radius map is obtained by the following method: The residues of the protein are taken as nodes, and the C, N, C α , atomic coordinates of O; The obtained C α The atomic coordinates of are set as the node coordinates of the protein graph structure, and the nodes whose distance is less than the preset radius threshold are considered to be connected to establish a radius graph; According to the acquisition of C, N, C α , the atomic coordinates of O construct the vectors u and v between atoms, and then construct the local coordinate system Q of node i i ; According to the local coordinate system Q of the node i obtained i The edge features and node features of the radius graph are extracted, and the extracted edge features and node features of the radius graph are stored in the form of a matrix.

2. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 1, characterized in that The steps for obtaining the preset radius threshold are as follows: for each protein, the average distance between each node and its neighboring nodes in its K-nearest neighbor graph is obtained, and then the average value d of the average distance of each node is calculated, and finally the radius threshold is obtained by multiplying the obtained average value d by the preselected radius coefficient μ1; the calculation method of the average value d is as follows: Where L represents the protein main chain structure, K represents the number of neighbors of each node in the K-nearest neighbor graph; B represents the set of protein main chain residues, j represents any neighbor node of node i; N i represents the neighbor set of node i; d ij Represents the distance between node i and node j.

3. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 1, characterized in that According to the acquisition of C, N, C α , the atomic coordinates of O construct the vectors u and v between atoms, and then construct the local coordinate system Q of node i i , specifically including: u=P Cα -P N ; v=P C -P Cα ; Q i =[b,n,b×n]; Among them, u means N points to C α vector; v represents C α The vector pointing to C; P x represents the x-coordinate of the atom; b represents the unit vector coplanar with u and v; n represents the unit normal vector perpendicular to the plane where u and v are located.

4. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 1, characterized in that In step ", the local coordinate system Q of the node i is obtained i Extracting edge features and node features of the radius graph, wherein the extracted edge features and node features of the radius graph are stored in a matrix form, wherein the edge features of the radius graph include distance features, direction features, and torsion features between nodes; The node features of the radius graph include distance features, direction features, and angle features of a single node.

5. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 4, characterized in that The distance feature between nodes is constructed by calculating the radial basis distance between atoms, specifically including: The Euclidean distance moment and the number of radial basis function kernels are used to obtain the center value μ2 of the radial basis function kernel RBF and the standard deviation σ of each radial basis function kernel RBF: μ2=linspace(D min ,D max ,numberbf); Among them, D min represents the minimum distance of the Euclidean distance matrix; D max represents the maximum distance of the Euclidean distance matrix; numrbf is the number of RBF kernels; the linspace function stratifies the distance interval into multiple uniform parts, the number of which is equal to numrbf; Use the exp function to calculate the radial basis function kernel RBF distance matrix: Introducing virtual atom C β , to utilize the virtual atom C β The coordinates of the set of atoms between two nodes [C, N, C α ,O,C β ] the radial basis distance between every two atoms; the virtual atom C β The coordinates in the local coordinate system are (x k ,y k ,z k ),satisfy: Among them, x k 、y k 、z k All are learnable parameters; The calculated radial basis distances are concatenated to obtain the distance feature between the two nodes: Among them, ‖M i -N j ‖ represents the radial basis distance between the M atoms of node i and the N atoms of its neighbor node j, and Cat represents the concatenation operation of the matrix.

6. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 5, characterized in that The directional feature between the nodes adopts the C α Atoms point to neighbor nodes C, N, C α , O four atoms between the vector set construction: in, Represents the direction vector from atom N in node i to atom M in its neighbor node j.

7. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 5, characterized in that The torsion feature between nodes is captured by constructing a rotation matrix R between nodes and converting it into a quaternion, specifically including: Compute the 3D part of the quaternion: Among them, R xx 、R yy 、R zz are the diagonal elements of the rotation matrix R; Calculate the sign of the antisymmetric component of the rotation matrix and use it to adjust the signs of x, y, and z: Update x, y, z to: [x, y, z] = [x, y, z] × [S x , S y , S z ; Compute the scalar part of the quaternion: Combine x, y, z, and w into a quaternion and normalize it to get the torsion characteristics between nodes:

8. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 5, characterized in that According to the obtained C, N, C α , the atomic coordinates of O and the virtual atom C β The atomic coordinates of the atoms are used to calculate the radial basis distances between atoms, and then the calculated radial basis distances are spliced ​​to obtain the distance characteristics of a single node The directional characteristics of a single node are represented by its C α Atoms pointing to C, N, C α , vector set construction of O atoms: Extract the sine and cosine values ​​of the six angles in the node as the angle features of a single node: Among them, α i Chemical bond The angle between i for The angle between i for Angle; Chemical bond The torsion angle; Chemical bond The torsion angle; ω i For chemical bond C i —N i+1 The torsion angle.

9. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 1, characterized in that In the step "Using the graph attention mechanism to design a denoising network that can update the node features and edge features of the protein radius graph separately to denoise the graph neural network model that introduces protein sequence data", the neighborhood attention mechanism is used to update the node features: In the lth layer of the graph neural network model, the attention weights of the central node and its neighboring nodes are: Among them, softmax represents the activation function; AttnMLP represents the attention-seeking multi-layer perceptron, which consists of three fully connected layers and two ReLU activation layers; is the node feature of neighbor node j; is the feature of the edge connecting the central node i and its neighbor node j; represents the node feature of the central node i; || represents the splicing operation; represents the scaling factor; d k Represents the quotient of the node feature dimension and the number of attention heads; After obtaining the attention weights of the central node i and its neighbor node j, the node feature of the central node i is updated to the weighted sum of the v values ​​of the neighbor node j: Among them, the v value of neighbor node j is obtained by concatenating the neighbor node's own features with the edge features and then mapping them through the multi-layer perceptron NodeMLP; NodeMLP consists of three fully connected layers and two GELU activation layers.

10. The tumor necrosis factor receptor sequence design model AttenGNNDiff based on graph attention mechanism according to claim 1, characterized in that In the step of "using the graph attention mechanism to design a denoising network that can update the node features and edge features of the protein radius graph respectively to denoise the graph neural network model that introduces protein sequence data", the node features of the connected nodes are concatenated with the original edge features between the connected nodes to obtain the new edge features of the connected nodes e' ij , thereby obtaining an edge feature matrix E' containing the preliminary updated features of all edges: in, Represents the node features of the central node i in the l+1th layer of the graph neural network model; Represents the edge features of the lth layer of the graph neural network model; Represents the node features of neighbor node j in the l+1th layer of the graph neural network model; And use the self-attention mechanism to calculate the attention weight between each pair of edge features in the set of edges connected to the same node: First, the new feature e' of the edge connecting the central node i and the neighbor node j ij Mapped to the query, key and value space through linear transformation, the corresponding representation Q is obtained ij , K ij and V ij : Q ij =W q ·e’ ij ;K ij =W k ·e’ ij ;V ij =W v ·e’ ij ; Among them, W q , W k , W v Linear transformation matrices representing queries, keys, and values; Then calculate the attention score between each pair of edge features to obtain the attention weight between each pair of edge features; Secondly, we use the weighted sum method to calculate the edge e ij The l+1th feature of the network is used, and the layer normalization operation is performed on the feature matrix after edge update.