Training method of amino acid sequence generation model and application thereof
By using graph neural network models and multi-dimensional geometric coding systems, combined with training data from real and simulated data, the problem of insufficient generalization ability of amino acid sequence generation models was solved, achieving more efficient protein structure prediction.
Patent Information
- Application Number
- CN202511002929.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-10
AI Technical Summary
Existing amino acid sequence generation models have insufficient generalization capabilities, limited by the scarcity of training data and model limitations, resulting in poor performance when processing unknown data.
A graph neural network model is used to model protein structure through a multi-dimensional geometric coding system and hierarchical perception method. The fusion training data of real and simulated data is combined, and decoding is performed using a random mask matrix and a hierarchical mask mechanism. The loss function is optimized to improve the generalization and robustness of the model.
The robustness and generalization ability of the prediction results of the amino acid sequence generation model are improved, ensuring the accuracy and adaptability in protein structure prediction.
Smart Images

Figure CN120766770A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of sequence prediction, and in particular to a training method for an amino acid sequence generation model and its application. Background Art
[0002] Protein reverse folding is the computational process of inferring the amino acid sequence that stably folds into a given three-dimensional protein backbone structure. Understanding the sequence-structure relationship of proteins is crucial for protein sequence prediction and design. Precise amino acid sequence design is central to creating new functional proteins.
[0003] In related technologies, due to the scarcity of training data and the limitations of existing models, existing models generally have the problem of insufficient generalization ability. How to design a highly generalizable amino acid sequence generation model is a problem that needs to be solved at present. Summary of the Invention
[0004] In order to solve or partially solve the problems existing in the related art, the present application provides a training method for an amino acid sequence generation model and its application, which can improve the generalization and robustness of the model.
[0005] The first aspect of the present application provides a method for training an amino acid sequence generation model, comprising:
[0006] Acquiring training data, wherein the training data includes sample structure information of a sample protein and corresponding sample sequence information;
[0007] Performing feature encoding on the sample structure information to obtain node features and edge features corresponding to the sample protein; wherein the node features are used to represent residue information, and the edge features are used to represent the spatial geometric relationship between adjacent residues;
[0008] Alternately updating node features and edge features through multiple message passing layers of the graph neural network model to be trained to obtain updated target node features and target edge features; wherein the number of neighbor nodes aggregated by each message passing layer increases sequentially;
[0009] According to the corresponding target node features and target edge features, the graph neural network model to be trained is used to decode and predict each node in a random order, and output prediction sequence information;
[0010] The network parameters of the graph neural network model are iterated according to the loss value calculated based on the predicted sequence information and the sample sequence information to obtain a trained amino acid sequence generation model.
[0011] The second aspect of the present application provides a method for generating an amino acid sequence, comprising:
[0012] Obtain the main chain structure information of the target protein;
[0013] The main chain structure information is input into the amino acid sequence generation model described in the first aspect, and the corresponding target amino acid sequence is output.
[0014] A third aspect of the present application provides a training device for an amino acid sequence generation model, comprising:
[0015] A sample acquisition module is used to acquire training data, wherein the training data includes sample structure information of the sample protein and corresponding sample sequence information;
[0016] A feature embedding module is used to perform feature encoding on the sample structure information to obtain node features and edge features corresponding to the sample protein; wherein the node features are used to represent residue information, and the edge features are used to represent the spatial geometric relationship between adjacent residues;
[0017] A message passing module is used to alternately update node features and edge features through multiple message passing layers of the graph neural network model to be trained to obtain updated target node features and target edge features; wherein the number of neighbor nodes aggregated by each message passing layer increases sequentially;
[0018] A node decoding module is used to decode and predict each node in a random order through the graph neural network model to be trained according to the corresponding target node features and target edge features, and output prediction sequence information;
[0019] A back propagation module is used to iterate the network parameters of the graph neural network model according to the loss value calculated based on the predicted sequence information and the sample sequence information to obtain a trained amino acid sequence generation model.
[0020] A fourth aspect of the present application provides a device for generating an amino acid sequence, comprising:
[0021] Information acquisition module, used to obtain the main chain structure information of the target protein;
[0022] A sequence generation module is used to input the main chain structure information into the amino acid sequence generation model trained by the training method described in the first aspect, and output the corresponding target amino acid sequence.
[0023] A fifth aspect of the present application provides an electronic device, including:
[0024] processor; and
[0025] A memory having executable code stored thereon, which, when executed by the processor, causes the processor to execute the method described in the first aspect or the second aspect above.
[0026] In a sixth aspect, the present application provides a computer-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described in the first or second aspect above.
[0027] In a seventh aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, they implement the method described in the first or second aspect above.
[0028] The technical solution provided by this application may include the following beneficial results:
[0029] The training method of the amino acid sequence generation model of the present application, in the process of modeling protein structure, is based on the framework of graph neural network, and the nodes of the graph are used to represent amino acid residues, and the edges are used to represent the spatial geometric relationship between residues. By introducing a multi-dimensional geometric coding system, it breaks through the limitations of traditional distance threshold coding and realizes the modeling of complex geometric systems in the protein folding process. In the encoding process of the graph neural network, the encoder takes the defined graph node and edge features as input, and alternately updates the node and edge features through multiple message passing layers; during this period, the perception range of the node is gradually expanded through each layer by increasing the neighbor range layer by layer. This hierarchical perception method effectively captures the local and global information in the graph, enabling the graph neural network to perceive the structural characteristics of proteins at multiple levels, thereby improving the model's learning and prediction capabilities for protein sequences. In the decoding process of the graph neural network, a random mask matrix is used to combine structural confidence with random noise to achieve order-independent decoding. Finally, by introducing structural confidence as a loss weight adjustment factor in the loss function, the contribution of high-confidence regions to model training is strengthened. At the same time, a hierarchical masking mechanism is used to filter low-confidence noise data, so that the model can dynamically adjust the loss weight according to the confidence of each predicted amino acid residue, ensuring that the model pays more attention to high-confidence regions during training, avoiding the impact of low-confidence regions on loss calculation, and effectively balancing the model's prediction accuracy and generalization performance.
[0030] In addition, in terms of constructing training data, an enhanced strategy of data fusion of real structure and predicted structure is adopted to expand the amount, diversity and representativeness of data required for training, and ensure the quality of training data.
[0031] The amino acid sequence generation method of the present application achieves the prediction of amino acid sequences based on the three-dimensional protein backbone by feature encoding the structure of the target protein and updating the hierarchical perception of the amino acid sequence generation model, thereby ensuring the robustness of the prediction results and the generalizability of the application.
[0032] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The above and other objects, features and advantages of the present application will become more apparent by describing in more detail exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.
[0034] Figure 1 1 is a flow chart of the training method for the amino acid sequence generation model shown in the embodiment of the present application;
[0035] Figure 2 1 is another flow chart of the training method for the amino acid sequence generation model shown in the embodiment of the present application;
[0036] Figure 3 1 is a flow chart of a method for generating an amino acid sequence generation model according to an embodiment of the present application;
[0037] Figure 4 Schematic diagram of the structure of the training device for the amino acid sequence generation model shown in the embodiment of the present application;
[0038] Figure 5 Schematic diagram of the structure of the device for generating an amino acid sequence generation model shown in an embodiment of the present application;
[0039] Figure 6 It is a structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0041] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0042] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, initial information may also be referred to as second information, and similarly, second information may also be referred to as initial information. Thus, features defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of this application, "plurality" means two or more, unless otherwise clearly and specifically defined.
[0043] In the related art, the Graph Neural Network (GNN model) is a deep learning model based on graph structure data, which can effectively process graph structure data. In protein structure modeling, graph neural networks can be used to simulate the amino acid residues of proteins and their interaction relationships. However, the lack of model training data is one of the key bottlenecks in the current field of protein sequence design. Although a large number of protein structures have been resolved, the protein structures obtained experimentally only account for 0.1% of the known protein sequence space. This lack of data volume and diversity significantly limits the performance of deep learning models. In addition, current protein sequence design methods oversimplify structural features in modeling. Due to the scarcity of training data and the limitations of existing models, current protein sequence generation methods generally have the problem of insufficient generalization ability, which makes the model perform poorly when processing unknown data.
[0044] To address the above problems, the embodiments of the present application provide a training method for an amino acid sequence generation model and its application, which can improve the generalization and robustness of the model prediction results.
[0045] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0046] See also Figure 1 and Figure 2 , a training method for an amino acid sequence generation model shown in one embodiment of the present application includes:
[0047] S110, obtaining training data, where the training data includes sample structure information of sample proteins and corresponding sample sequence information.
[0048] In this application, the training data comes from real experimental data and simulated prediction data. It is understandable that since the real experimental data is limited, for example, there are more than 100,000 experimental data, this application has enriched the data through millions of high-confidence prediction data. Among them, the sample structure information of the sample protein includes the three-dimensional coordinate information corresponding to the protein skeleton, and the sample sequence information includes the amino acid sequence corresponding to the protein. Among them, each residue in the amino acid sequence is selected from 20 natural amino acids (i.e., standard amino acids); optionally, some residues in the amino acid sequence of the sample protein can also be selected from non-standard amino acids to enrich the quantity and diversity of the sample protein as much as possible.
[0049] In order to obtain a large amount of high-quality training data, in some implementations, the training data can be obtained according to the following method:
[0050] S111 , clustering the real proteins according to a first preset percentage of sequence identity to obtain a plurality of first protein clusters.
[0051] In this step, real protein structures can be screened from various known protein databases or public literature, for example, protein structures determined by X-ray crystallography or cryo-electron microscopy can be screened from the PDB database (Protein Data Bank). In order to improve the quality of protein data, protein structures with a resolution less than a preset value and / or a sequence length (i.e., the number of residues) less than a preset length threshold can be selected. The preset value of the resolution can be, for example, Or other values, the preset length threshold can be selected from 800 to 1000. For example, select a resolution less than and protein structures with fewer than 1,000 residues. It is understood that the finer the protein resolution, the higher the structural accuracy, and thus the higher the quality of the training data generated; an appropriate number of residues can improve computational efficiency and reduce memory usage.
[0052] Furthermore, relevant tools can be used to cluster the selected real proteins according to the preset sequence identity to obtain multiple first protein clusters. Among them, sequence identity refers to the percentage of sites with exactly the same amino acids in any two compared amino acid sequences. Among them, the first preset percentage can be selected from 25% to 35%. For example, when the first preset percentage is 30%, real proteins containing at least 30% of the same sites are selected and clustered into a cluster to form a corresponding first protein cluster. By analogy, more than 100,000 first protein clusters can be obtained, and each first protein cluster contains multiple real proteins.
[0053] S112, clustering the simulated proteins according to a dual clustering rule to obtain a plurality of second protein clusters, wherein the dual clustering rule includes clustering according to a second preset percentage of sequence overlap and / or a third preset percentage of sequence identity, and clustering according to a fourth preset percentage of structural overlap.
[0054] In this step, simulated proteins can be selected from known protein structure prediction platforms (such as AlphaFold, Swiss-Model, I-TASSER, etc.) as the source of training data. Among them, the structure of simulated proteins is generally predicted by artificial intelligence models. For example, there are hundreds of millions of predicted simulated proteins in the AlphaFold Database. In this application, these predicted data are aggregated and filtered using dual clustering rules to obtain high-quality training data.
[0055] The dual clustering rules of the present application include clustering according to sequence overlap and / or consistency, and clustering according to structural overlap. When multiple simulated proteins meet the dual clustering rules of sequence clustering and structural clustering at the same time, they can be clustered into one cluster. In some specific embodiments, the dual clustering rules include: clustering simulated proteins with a sequence overlap of not less than a second preset percentage and / or a sequence identity of not less than a third preset percentage, and a protein structural overlap of not less than a fourth preset percentage. Sequence overlap refers to the percentage of the total length of the shorter sequence that can be compared with the longer sequence during the alignment of any two amino acid sequences. Optionally, the second preset percentage can be selected from 80% to 90%, the third preset percentage can be selected from 40% to 60%, and the fourth preset percentage can be selected from 80% to 90%. Exemplarily, multiple simulated proteins with a sequence overlap of not less than 90% and a sequence identity of not less than 50%, and a protein structural overlap of not less than 90% are clustered to obtain corresponding second protein clusters, that is, each second protein cluster contains multiple simulated proteins. Through clustering, more than 200 million simulated proteins can be clustered into 18 million second protein clusters.
[0056] S113 , determining sample proteins from each of the first protein clusters and the second protein clusters.
[0057] The sample proteins in the training data of the present application are composed of a mixture of real proteins and simulated proteins. In some embodiments, a preset number of real proteins are selected from each of the first protein clusters as sample proteins, and simulated proteins with a confidence score greater than a preset confidence threshold are selected from each of the second protein clusters as sample proteins.
[0058] Specifically, the same preset number can be set, and a preset number of real proteins can be selected from each first protein cluster. Alternatively, a random preset number can be set, and at least one real protein can be randomly selected from each first protein cluster.
[0059] After obtaining multiple second protein clusters, in order to improve the data quality of sample proteins, simulated proteins with low confidence can be filtered out for data enhancement, that is, simulated proteins with higher confidence are selected as sample proteins. In order to obtain simulated proteins whose confidence meets the preset confidence threshold, in some embodiments, the confidence of each simulated protein is obtained separately; the simulated proteins with confidence greater than the preset confidence threshold in each second protein cluster are retained as sample proteins. Optionally, the confidence of the present application can be characterized by the pLDDT average value of the protein. Of course, in other embodiments, other indicators can also be selected to represent the confidence, which is not limited here. Among them, pLDDT (Predicted Local Distance Difference Test) is a confidence indicator for evaluating protein structure prediction, and its value range is 0 to 100. The higher the value, the higher the accuracy of the prediction. Its value is automatically evaluated by the AlphaFold Database. The pLDDT average is the average of the pLDDT scores of each residue in a single protein.
[0060] In this embodiment, the pLDDT average value of each protein can be obtained as an evaluation index of the corresponding confidence, and then the proteins in each second protein cluster whose pLDDT average value is greater than a preset confidence threshold are retained. The preset confidence threshold can be selected from 80 to 90, for example, the simulated proteins with a pLDDT average value greater than 80 are retained as sample proteins. It can be understood that the number of sample proteins determined in each second protein cluster may be different. In order to obtain more balanced structural data, in some embodiments, the maximum and / or minimum number of sample proteins selected from each second protein cluster can be pre-set. That is, the number of qualified sample proteins selected from each second protein cluster does not exceed the maximum value, and / or the number of qualified sample proteins selected from each second protein cluster is not less than the minimum value. Furthermore, the simulated proteins in the second protein cluster can be sorted according to the size of the confidence, and simulated proteins with higher confidence can be screened as sample proteins.
[0061] By screening according to the aforementioned confidence levels, we can obtain 3.2 million pieces of quality-enhanced simulated protein data, effectively expanding the quality and diversity of model training data.
[0062] In this step, the first protein clusters obtained in S111 and the second protein clusters obtained in S112 may be deduplicated in advance, that is, redundant sequences and proteins with the same structure may be deleted before selecting sample proteins therefrom.
[0063] As can be understood, the sample proteins are selected from real proteins and simulated proteins, that is, the training data contains both real proteins and simulated proteins, making the training data both reliable and rich. As the model training continues, after multiple rounds (e.g., 8-10 epochs), new proteins can be selected to form new training data.
[0064] Furthermore, the training data is pre-divided into training set, validation set and test set. The sample proteins in different sets are different, thus avoiding data leakage and overfitting, and ensuring the generalization ability and evaluation accuracy of the trained model in practical applications.
[0065] S120, feature encoding is performed on the sample structure information to obtain node features and edge features corresponding to the sample protein; wherein the node features are used to represent residue information, and the edge features are used to represent the spatial geometric relationship between adjacent residues.
[0066] For each sample protein in the training data, its sample structure information can be feature encoded according to the preset feature dimensions. The amino acid sequence generation model of the present application is a model based on a graph neural network architecture, in which the nodes of the graph are used to represent each amino acid residue in the protein, and the edges are used to represent the spatial geometric relationship between adjacent residues.
[0067] In some embodiments, each node feature includes the periodic trigonometric function values corresponding to the torsion angles φ and ψ of the residue, the secondary structure, and the structural confidence of the residue. Each node feature includes 25 preset feature dimensions, and each node feature is a feature vector composed of 25 features, specifically including:
[0068] The periodic trigonometric function values corresponding to the torsion angles φ and ψ of the residue include the 16-dimensional periodic features obtained by fourth-order Fourier expansion of the torsion angles φ and ψ, namely sin(nφ), cos(nφ), sin(nψ), cos(nψ), where n = 1 to 4. When n = 1, the values are sin(φ), cos(φ), sin(φ), and cos(φ) of the residue, respectively; when n = 2, the values are sin(2φ), cos(2φ), sin(2φ), and cos(2φ), respectively, of the residue, and so on, for a total of 16-dimensional features.
[0069] The secondary structure of the residues, including: α-helix, 3 10There are 8 secondary structures, including -helix, π-helix, extended β-strand, isolated β-bridge, turn, bend, and coiled / random structure, corresponding to 8-dimensional features. The corresponding features of each secondary structure can be represented by one-hot encoding.
[0070] The structural confidence of a residue is expressed as a pLDDT score or 100. When the sample protein is a simulated protein, the structural confidence can be expressed using the pLDDT score of each residue automatically evaluated using the AlphaFold Database. When the sample protein is a real protein, the structural confidence can default to 100. The structural confidence of a residue can be expressed after normalization.
[0071] In some embodiments, a single edge feature includes the distance information between the preset atoms of two residues and the confidence of the relative positions of the two residues. α atoms, C atoms and O atoms, including a virtual C to simulate the side chain position β Atoms, any two atoms of the two residues are paired, and 25 pairs of atomic combinations can be obtained, and then the distance information corresponding to each pair of atomic combinations can be obtained. Furthermore, the distance information of each pair of atomic combinations can be calculated according to multiple RBFs (Radial Basis Functions) to obtain the distance features calculated by each RBFs. Taking the number of RBFs as an example, a total of 25×16=400-dimensional distance features can be obtained. Among them, the virtual C β The atoms are not real atoms, but artificially defined virtual atoms. The virtual C β The coordinates of atoms can be determined according to pre-set formulas. β The atomic coordinates of are calculated as follows:
[0072] C β =-0.58273431×a+0.56802827×b-0.54067466×c+C α
[0073] Where, vector b = C α -N, vector c = CC α , normal vector a=cross(b,c).
[0074] In some embodiments, the confidence of the relative position of two residues can be expressed using PAE. PAE (Predicted Aligned Error) is an indicator used in protein structure prediction tools such as AlphaFold to evaluate the accuracy of protein structure prediction. The larger the PAE value, the greater the deviation between the predicted structure and the actual structure, and vice versa. When the sample protein is a simulated protein, PAE is automatically evaluated by the AlphaFold Database, for example. When the sample protein is a real protein, the PAE value defaults to 0.01.
[0075] In summary, a single node feature can be generated from a feature vector consisting of 25-dimensional features, and a single edge feature can be generated from a feature vector consisting of 401-dimensional features. In this step, for each sample protein, the corresponding 25-dimensional node features and 401-dimensional edge features are encoded based on the corresponding sample structure information. These serve as the initial node features and initial edge features input to the encoder in subsequent steps.
[0076] This application captures the periodic characteristics of the torsion angle of the residue by fourth-order Fourier expansion, using virtual C β Atoms construct a spatial reference system for virtual side chains, combining multiple RBFs to finely encode the distances between atoms and integrate secondary structure and confidence features. This design overcomes the limitations of traditional distance threshold encoding through a multi-scale geometric feature fusion mechanism, enabling the modeling of complex geometric systems during protein folding. This in turn improves the ability to model complex geometric constraints on proteins and enhances the model's understanding of protein structure.
[0077] S130, alternately updating the node features and edge features through multiple message passing layers of the graph neural network model to be trained to obtain updated target node features and target edge features; wherein the number of neighbor nodes aggregated by each message passing layer increases sequentially.
[0078] The graph neural network model of the present application includes an encoder and a decoder. In this step, the encoder includes a plurality of message passing layers, which take each node feature and edge feature of the graph as input, and alternately update the node features and edge features through a plurality of message passing layers to obtain updated target node features and target edge features. Among them, each message passing layer is used to aggregate the node features of a preset number of neighbor nodes and the edge features connected to each neighbor node. The number of neighbor nodes aggregated by the next message passing layer is greater than the number of neighbor nodes aggregated by the previous message passing layer. By gradually increasing the number of aggregated neighbor nodes, the perception range of the node is expanded, thereby helping to capture structural information of different scales. This hierarchical perception method effectively captures local and global information in the graph, enabling the graph neural network to perceive the structural characteristics of proteins at multiple levels, thereby improving the model's learning and prediction capabilities for protein sequences.
[0079] The present application includes at least two message passing layers. In some embodiments, the multiple message passing layers of the present application include at least a first message passing layer and a second message passing layer. The first message passing layer is used to aggregate the current node's X neighbor nodes and corresponding connecting edges, and the second message passing layer is used to aggregate the current node's Y neighbor nodes and corresponding connecting edges. Preferably, the message passing layer also includes a third message passing layer, which is used to aggregate the current node's Z neighbor nodes and corresponding connecting edges. 0 < X < Y < Z, where X, Y, and Z are all natural numbers. For example, X can be selected from 3 to 12, Y can be selected from 13 to 24, and Z can be selected from 25 to 48. In some specific embodiments, X can be 12, Y can be 24, and Z can be 48. In other words, the first message passing layer is used to aggregate the current node's 12 closest neighbor nodes in space, the second message passing layer is used to aggregate the current node's 24 closest neighbor nodes in space, and the third message passing layer is used to aggregate the current node's 48 closest neighbor nodes in space. The number of neighbor nodes in each message transmission layer may be increased randomly or in a certain regular pattern, such as an arithmetic change or a geometric change, without any restriction.
[0080] In some specific implementations, the initial node feature of the current node is updated by aggregating X neighbor nodes and corresponding connecting edges through the first message passing layer to obtain the first node feature corresponding to the current node; the initial edge feature of the connecting edge is updated based on the first node feature of the current node and the current node features of the neighbor nodes to obtain the first edge feature corresponding to the current connecting edge; the first node feature of the current node is updated by aggregating Y neighbor nodes and corresponding connecting edges through the second message passing layer to obtain the second node feature corresponding to the current node; the first edge feature of the connecting edge is updated based on the second node feature of the current node and the current node features of the neighbor nodes to obtain the second edge feature corresponding to the current connecting edge. When a third message passing layer exists, the second node feature of the current node is updated by aggregating Z neighbor nodes and corresponding connecting edges through the third message passing layer to obtain the target node feature corresponding to the current node; the second edge feature of the connecting edge is updated based on the target node feature of the current node and the current node features of the neighbor nodes to obtain the target edge feature corresponding to the current connecting edge.
[0081] In other words, each message passing layer differs only in the number of aggregated neighbor nodes. Multiple message passing layers update node features sequentially, with the node features updated by the first message passing layer serving as input to the second message passing layer. The node features updated by the second message passing layer then serve as input to the third message passing layer, and the node features updated by the third message passing layer serve as the target node features. Simultaneously, for each message passing layer, the node features are updated first, followed by the corresponding edge features. This cycle of alternating node and edge feature updates continues until the last message passing layer outputs the final target node features.
[0082] For ease of understanding, take any node i as the current node to be updated as an example, and combine the following node feature update formula (1) to update its initial node feature V through the first message passing layer. i Update to obtain the updated first node feature V' i , and then combine the following edge feature update formula (2) to update its initial edge feature E ij Update to obtain the updated first edge feature E' ij The node feature update formula is as follows:
[0083]
[0084] Among them, i represents the current node, j represents the neighbor node, V represents the node feature, and E represents the edge feature.
[0085] It can be understood that the neighbor node refers to the node closest to the current node in the current space. Take the first message passing layer to aggregate 12 neighbor nodes as an example. First, the node feature V is transformed into i , node features V of neighboring nodes j , and the edge features E corresponding to the current node and the neighboring nodes ij Processing is performed to generate 12 corresponding side information respectively; then the 12 side information of the current node i are aggregated to obtain aggregate information, and then the aggregate information is processed by the multi-layer perceptron MLP2. Preferably, Dropout (random dropout) can be applied to the processed aggregate information to prevent overfitting. The Dropout layer can randomly return 10% to 20% of the aggregate information to zero, and then compare the information after Dropout with the initial node feature V i Add and normalize the layer to get the updated node feature V' of the current message passing layer i Optionally, the hidden dimension of each multilayer perceptron in the present application is 512 dimensions.
[0086] In order to facilitate distinction, the node feature updated by the first message passing layer is defined as the first node feature. After obtaining the first node feature corresponding to the current node, the edge feature E is updated in combination with the following edge feature update formula (2). ij Update and obtain the first edge feature E' after the current first message passing layer is updated ij .
[0087] E′ ij =LayerNorm(E ij +Dropout(MLP3([V′ i , V′ j , E ij ]))) (2)
[0088] Among them, the updated first node feature V' is updated through the multi-layer perceptron MLP3 i , the current node feature V' corresponding to the neighbor node j j and the initial edge feature E ij Perform splicing aggregation. Preferably, Dropout (random discard) can be applied to the processed aggregate information to prevent overfitting. Then the information after Dropout is combined with the initial edge feature E ij Add and normalize the layer to get the first edge feature E' after the current message passing layer is updated ij .
[0089] It should be noted that the current node feature V' corresponding to the neighbor node j jIt refers to the latest feature of the node j at present. The current feature of the neighbor node j may be the initial node feature that has not been updated yet, or it may be the first node feature that has been updated, which is determined according to the actual situation.
[0090] After sequentially updating the first node features corresponding to the current node i and the 12 first edge features corresponding to the current node i and its 12 neighboring nodes j through the first message passing layer, the updated second node features and 24 second edge features are obtained through the second message passing layer, using the latest feature information as input, according to the aforementioned principles. If the third message passing layer does not exist, feature updating ceases, and the second node features and 24 second edge features are used as the target node features and target edge features corresponding to the current node i, which the encoder needs to output. If the third message passing layer exists, the second node features and 24 second edge features are used as input to the third message passing layer, and the updated third node features and 48 third edge features are obtained, according to the aforementioned principles. This process continues until the last message passing layer outputs the target node features and target edge features corresponding to the current node i. This multi-level perception-based update method updates the features of each node and edge, providing more accurate input for the subsequent decoding process.
[0091] It can be understood that the number of residues in a single protein does not exceed 1000, and the corresponding number of nodes does not exceed 1000. After completing the update of the node features and edge features of the first node, the node features and edge features of the second node are then updated according to the multi-layer message passing layer update method described above until the feature updates of all nodes are completed, which serves as the input data for the subsequent decoder.
[0092] In order to further reduce the negative impact of the quality of the training data itself on the reliability of the training results, in some embodiments, during the information aggregation process of each message passing layer, a mask matrix is used to mask the abnormal residues in each sample protein. Among them, for the simulated proteins in the training data, abnormal residues refer to residues with structural confidence, such as pLDDT scores, less than 90; for the real proteins in the training data, abnormal residues refer to non-standard amino acids and residues with missing coordinates. It can be understood that these masked abnormal residues do not need to participate in the update of the corresponding node features. For example, in the process of processing the aggregated information according to the multi-layer perceptron MLP2 mentioned above, these abnormal residues are masked, and then message passing will not be performed. These masked abnormal nodes do not need to participate in the node residue type prediction and loss function calculation in the subsequent decoding steps.
[0093] In order to distinguish between abnormal residues and normal residues in a sample protein, in some embodiments, corresponding preset codes are used in a mask matrix to mark abnormal residues and normal residues, respectively. The label corresponding to abnormal residues is 0, and the label corresponding to normal residues is 1.
[0094] S140: According to the corresponding target node features and target edge features, the graph neural network model to be trained is used to decode and predict each node in a random order, and output prediction sequence information.
[0095] In this step, all target node and edge features updated in the previous step are input into the decoder of the graph neural network model for decoding. The residue type corresponding to each node is predicted one by one based on the autoregressive masking mechanism. The residue types include 20 natural amino acids.
[0096] Furthermore, the decoder of the present application has a symmetrical design with the encoder, i.e., the decoder has the same number of message passing layers as the encoder, and each layer aggregates a corresponding number of neighbor nodes. The mechanism of operation of each message passing layer of the decoder is the same as that of the encoder. For example, if the encoder has three message passing layers and aggregates 12, 24, and 48 neighbor nodes in ascending order, then the decoder has a first message passing layer, a second message passing layer, and a third message passing layer, and each layer aggregates 12, 24, and 48 neighbor nodes in ascending order.
[0097] In some embodiments, the current node to be predicted is determined based on a random mask matrix; the target node features of the node to be predicted are updated in a cyclical manner according to multiple message passing layers in the decoder to obtain corresponding node update features; based on the residue type probability distribution converted from the node update features, the corresponding residue type is predicted and the corresponding prediction sequence information is generated.
[0098] It's important to note that traditional autoregressive decoding mechanisms typically rely on a fixed predicted order of residues to generate amino acid sequences. This approach fails to fully consider the structural characteristics of proteins and lacks a dynamic adjustment mechanism. Given the high complexity of protein folding, fixed-order decoding strategies fail to adapt to the diverse folding pathways, limiting the model's flexibility and adaptability.
[0099] Specifically, unlike the related art, which uses a fixed prediction order from the beginning to the end of the amino acid sequence, this application uses a random mask matrix to randomly determine the nodes to be predicted at each step. The node prediction order of different sample proteins is randomly determined. In the node prediction of the current step, the target node features input by the encoder are updated through multiple message passing layers of the decoder. The working mechanism of each message passing layer of the decoder is the same as that of the encoder. The only difference is the number of aggregated neighbor nodes. The input data of the next message passing layer is the output data of the previous message passing layer. The last message passing layer only needs to output the node update features of the last update, and no longer needs to output the updated edge features.
[0100] For example, using the mechanism of the first message passing layer of the decoder as an example, the target node feature corresponding to the predicted node, the target edge features of the 12 connecting edges, and the sequence relative position feature corresponding to the current node are concatenated to obtain the corresponding concatenated feature. The concatenated feature is then aggregated with the target node features corresponding to the 12 neighboring nodes to obtain the corresponding node aggregated feature. The node aggregated feature is then further updated through layer normalization, dropout, and a feedforward network to obtain the updated node feature after the first message passing layer. After the node feature update is completed in the first message passing layer, the target edge features are alternately updated using the same update method as the encoder to obtain the corresponding 12 updated edge features. The updated node feature and the 12 updated edge features are input to the second message passing layer for further updating, obtaining a new node update feature and 24 new updated edge features. The updated node feature and the 24 updated edge features are then input to the third message passing layer for further updating, obtaining the final node update feature and 48 new updated edge features. Based on the final node update feature, a linear layer and a softmax function are used to output the probability distribution of the 20 natural amino acids corresponding to the current node. According to these probability distributions, the amino acid with the highest probability is selected as the prediction result of the current node. Based on the predicted results of the residue types of each node, the complete amino acid sequence of the sample protein is composed, and the predicted sequence information is obtained.
[0101] Furthermore, in order to overcome the limitation of autoregressive decoding that relies on a fixed generation order (from left to right), the random mask matrix T' of the present application is obtained by randomly permuting the original triangular mask T. In some embodiments, the random mask matrix T' is obtained according to the following formula (3), specifically:
[0102] T'=PTP T (3)
[0103] P is a random masking matrix of the same size as T. T is an L×L square matrix, where L represents the length of the protein's amino acid sequence. Matrix T indicates which nodes at each step in the decoding process need to be masked, preventing the decoder from directly accessing the masked portion of the input sequence. The elements of the masking matrix are either 0 or 1, with 1 indicating that the node at that position is not masked and 0 indicating that the node at that position is masked.
[0104] From the above, we can see that by using the random mask matrix T', the rows and columns of the mask are disrupted at the same time, so that the positions of nodes that should be adjacent in sequence are not in a fixed generation order, so as to achieve order-independent decoding.
[0105] S150, iterating the network parameters of the graph neural network model according to the loss value calculated based on the predicted sequence information and the sample sequence information to obtain a trained amino acid sequence generation model.
[0106] In this step, the loss value (loss) is calculated based on the generated predicted sequence information and the actual sample sequence information, where the actual sample sequence information is the training label. In order to improve the accuracy of the prediction results, the loss function used in the graph neural network model of this application is a hybrid loss function that is weighted and adjusted based on the structural confidence of each node in the sample protein. In some specific embodiments, the loss value of the predicted sequence information and the sample sequence information is calculated according to the following formula (4), as shown below:
[0107]
[0108] Among them, -log(probs(S i )) is the negative log-likelihood loss, which is used to measure the logarithmic difference between the predicted sequence information and the sample sequence information, probs(S i ) is the probability value of the residue category predicted by node i, pLDDT i Indicates the structural confidence corresponding to each residue, mask i It is 0 or 1 to ensure that only the loss of normal nodes that are not obscured is calculated.
[0109] By adopting the above-mentioned hybrid loss function that includes structural confidence weighting, the structural confidence represented by pLDDT is used as the loss weight adjustment factor to enhance the contribution of high-confidence regions to model training. At the same time, a hierarchical masking mechanism is used to filter low-confidence noise data. This dynamic optimization strategy effectively balances prediction accuracy and generalization performance.
[0110] After calculating the loss value, backpropagation can be performed based on the loss value to adjust the network parameters of each hidden layer of the encoder and decoder. This iteration is repeated until the calculated loss value converges, and the trained graph neural network model is obtained as an amino acid sequence generation model for application in the generation prediction of amino acid sequences.
[0111] It can be understood that the trained graph neural network model is verified and tested using a validation set and a test set that are completely non-repetitive with the training set, and sequence and structure deduplication is strictly implemented to ensure that the training set, validation set, and test set are independent of each other at the sequence and structure levels, thereby effectively avoiding data leakage and overfitting problems.
[0112] In summary, the training method of the amino acid sequence generation model of the present application adopts an enhanced strategy of data fusion of real structure and predicted structure in the construction of training data, which expands the amount, diversity and representativeness of data required for training and ensures the quality of training data. In the process of modeling protein structure, based on the framework of graph neural network, the nodes of the graph are used to represent amino acid residues, and the edges are used to represent the spatial geometric relationship between residues. By introducing a multi-dimensional geometric coding system, the limitations of traditional distance threshold coding are broken through, and the modeling of complex geometric systems in the protein folding process is realized. In the encoding process of the graph neural network, the encoder takes the defined graph node and edge features as input, and alternately updates the features of the nodes and edges through multiple message passing layers; during this period, the perception range of the nodes is gradually expanded through three layers by increasing the neighbor range layer by layer. This hierarchical perception method effectively captures the local and global information in the graph, enabling the graph neural network to perceive the structural characteristics of proteins at multiple levels, thereby improving the model's learning and prediction capabilities for protein sequences. In the decoding process of the graph neural network, a random mask matrix is used to combine structural confidence with random noise to achieve order-independent decoding. Finally, by introducing structural confidence as a loss weight adjustment factor in the loss function, the contribution of high-confidence regions to model training is strengthened. At the same time, a hierarchical masking mechanism is used to filter low-confidence noise data, so that the model can dynamically adjust the loss weight according to the confidence of each predicted amino acid residue, ensuring that the model pays more attention to high-confidence regions during training, avoiding the impact of low-confidence regions on loss calculation, and effectively balancing the model's prediction accuracy and generalization performance.
[0113] See also Figure 3 One embodiment of the present application further provides a method for generating an amino acid sequence, which comprises:
[0114] S210, obtaining the main chain structure information of the target protein.
[0115] The amino acid sequence generation model of this application is applied to the task of protein reverse folding. Based on the three-dimensional structure of the target protein, the main chain structure information is obtained, and then the amino acid sequence that can stably fold into this three-dimensional structure is predicted. The main chain structure information includes the three-dimensional coordinates of each heavy atom on the backbone. Optionally, the target protein can be a known real protein or a virtual protein, without limitation.
[0116] Based on the main chain structure information, the node features corresponding to each residue in the target protein can be generated based on the embedded module encoding, and the edge features corresponding to each node and its 12 to 48 neighboring nodes can be generated. In some embodiments, each node feature includes:
[0117] The torsion angles φ and ψ of the residues were expanded by the fourth-order Fourier transform to obtain the 16-dimensional periodic features: sin(nφ), cos(nφ), and sin(nψ), cos(nψ), respectively, where n = 1 to 4;
[0118] 8-dimensional residue secondary structure: α-helix, 3 10 - Helix, π-helix, extended β-strand, isolated β-bridge, turn, bend, and coiled / random structures;
[0119] Structural confidence of 1-D residues: pLDDT score.
[0120] That is, each node feature contains at most 25-dimensional features.
[0121] In some implementations, each edge feature includes:
[0122] The nitrogen atoms, carbon atoms, and α Atoms, C atoms, O atoms and virtual C β The distance information between any two atoms;
[0123] The confidence of the predicted relative position of residue i to residue j.
[0124] Preferably, the distance information between any two atoms can be distance information encoded according to multiple radial basis functions (RBFs). When the number of radial basis functions is 16, each edge feature contains at most 401-dimensional features. By calculating according to 16 RBFs, a total of 25×16=400-dimensional geometric distance features can be obtained.
[0125] S220, inputting the main chain structure information into the trained amino acid sequence generation model, and outputting the corresponding target amino acid sequence.
[0126] The node features and the related edge features corresponding to each residue of the target protein are input into the amino acid sequence generation model, and the amino acid type corresponding to each residue is predicted in a random decoding order, and the target amino acid sequence is sequentially composed. The amino acid sequence generation model updates the node features in a hierarchical manner according to the plurality of message passing layers in the encoder and the decoder, and details are not repeated here.
[0127] As can be seen from the example, the amino acid sequence generation method of the present application encodes the structure of the target protein, and realizes the prediction of the amino acid sequence based on the hierarchical perception update mode of the amino acid sequence generation model based on the three-dimensional protein main chain, to ensure the robustness and generalization of the prediction result.
[0128] Corresponding to the foregoing application function implementation method embodiment, the present application also provides a training device, a generation device, an electronic device and corresponding embodiments of an amino acid sequence generation model.
[0129] Figure 4 FIG. 1 is a structural schematic diagram of a training device of an amino acid sequence generation model according to an embodiment of the present application.
[0130] Referring to Figure 4 , the training device of the amino acid sequence generation model according to an embodiment of the present application includes a sample acquisition module 310, a feature embedding module 320, a message passing module 330, a node decoding module 340 and a back propagation module 350. Among them:
[0131] The sample acquisition module 310 is configured to acquire training data, and the training data includes sample structure information and corresponding sample sequence information of a sample protein.
[0132] The feature embedding module 320 is configured to encode the sample structure information to obtain node features and edge features corresponding to the sample protein; wherein the node features are used to represent residue information, and the edge features are used to represent the spatial geometric relationship between adjacent residues.
[0133] The message passing module 330 is configured to alternately update the node features and the edge features through a plurality of message passing layers of a graph neural network model to be trained to obtain updated target node features and target edge features; wherein the number of neighbor nodes aggregated by each message passing layer increases in order.
[0134] The node decoding module 340 is configured to decode and predict each node in a random order according to the corresponding target node features and target edge features through the graph neural network model to be trained, and output the predicted sequence information.
[0135] The back propagation module 350 is used to iterate the network parameters of the graph neural network model according to the loss value calculated based on the predicted sequence information and the sample sequence information to obtain a trained amino acid sequence generation model.
[0136] In some specific embodiments, the sample acquisition module 310 is used to cluster each real protein according to a first preset percentage of sequence consistency to obtain multiple first protein clusters; cluster the simulated proteins according to a dual clustering rule to obtain multiple second protein clusters, wherein the dual clustering rule includes clustering according to a second preset percentage of sequence overlap and / or a third preset percentage of sequence consistency, and clustering according to a fourth preset percentage of structural overlap; and determine the sample protein from each first protein cluster and the second protein cluster.
[0137] In some specific embodiments, the sample acquisition module 310 determines sample proteins from each of the first protein clusters and the second protein clusters, including: selecting a preset number of real proteins from each of the first protein clusters as sample proteins, and selecting simulated proteins with a confidence level greater than a preset confidence threshold from each of the second protein clusters as sample proteins.
[0138] In some specific embodiments, the message passing module 330 is used to aggregate X neighbor nodes and corresponding connection edges through the first message passing layer, update the initial node feature of the current node, and obtain the first node feature corresponding to the current node; update the initial edge feature of the connection edge based on the first node feature of the current node and the current node feature of the neighbor node, and obtain the first edge feature corresponding to the current connection edge; aggregate Y neighbor nodes and corresponding connection edges through the second message passing layer, update the first node feature of the current node, and obtain the second node feature corresponding to the current node; update the first edge feature of the connection edge based on the second node feature of the current node and the current node feature of the neighbor node, and obtain the second edge feature corresponding to the current connection edge; when the third message passing layer exists, aggregate Z neighbor nodes and corresponding connection edges through the third message passing layer, update the second node feature of the current node, and obtain the target node feature corresponding to the current node; update the second edge feature of the connection edge based on the target node feature of the current node and the current node feature of the neighbor node, and obtain the target edge feature corresponding to the current connection edge.
[0139] In some specific implementations, the node decoding module 340 is used to determine the node to be predicted at the current step based on a random mask matrix; based on the corresponding target node features and target edge features, the target node features of the node to be predicted are updated in a cyclical manner according to multiple message passing layers in the decoder to obtain corresponding node update features; based on the probability distribution of residue types converted from the node update features, the corresponding residue types are predicted and corresponding prediction sequence information is generated.
[0140] In some specific embodiments, the back propagation module 350 is used to calculate the loss value of the predicted sequence information and the sample sequence information according to the mixed loss function, and back propagate the network parameters of the iterative graph neural network model according to the loss value to obtain a trained amino acid sequence generation model.
[0141] As can be seen from this example, the training device of the amino acid sequence generation model of the present application obtains large-scale and high-quality training data through a data enhancement strategy, and extracts data features through a multi-dimensional geometric feature encoding system, thereby improving the modeling ability of complex geometric constraints of proteins and enhancing the model's understanding of protein structure. By combining structural confidence with random noise using a random permutation matrix, order-independent decoding is achieved. A hybrid loss function containing confidence weighting is constructed to overcome the generalization capability defects caused by data scarcity and model limitations. Such a design, combining the expansion of the training data set, the update method of hierarchical perception and the construction of a weighted loss function, effectively enhances the robustness and generalization ability of the amino acid sequence generation model on unknown data.
[0142] Figure 5 It is a structural schematic diagram of the generation device of the amino acid sequence generation model shown in the embodiment of the present application.
[0143] See also Figure 5 The embodiment of the present application shows an apparatus for generating an amino acid sequence, which includes: an information acquisition module 410 and a sequence generation module 420.
[0144] The information acquisition module 410 is used to obtain the main chain structure information of the target protein;
[0145] The sequence generation module 420 is used to input the main chain structure information into the trained amino acid sequence generation model and output the corresponding target amino acid sequence.
[0146] It can be seen from this example that the amino acid sequence generation device of the present application can be widely used in the prediction of amino acid sequences of unknown proteins, ensuring the generalization and robustness of the prediction results.
[0147] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.
[0148] Figure 6 It is a structural diagram of an electronic device shown in an embodiment of the present application.
[0149] See also Figure 6 , the electronic device 1000 includes a memory 1010 and a processor 1020.
[0150] The processor 1020 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0151] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. ROM may store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage may be a readable and writable storage device. The permanent storage may be a non-volatile storage device that retains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (e.g., a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all instructions and data required by the processor during operation. In addition, the memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0152] The memory 1010 stores executable codes. When the executable codes are processed by the processor 1020 , the processor 1020 may execute part or all of the above-mentioned methods.
[0153] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0154] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium), which stores executable code (or computer program or computer instruction code) and, when executed by a processor of an electronic device (or server, etc.), enables the processor to execute part or all of the steps of the above-mentioned method according to the present application.
[0155] The present application also provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.
[0156] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A training method for an amino acid sequence generation model, characterized in that: include: Acquiring training data, wherein the training data includes sample structure information of a sample protein and corresponding sample sequence information; Performing feature encoding on the sample structure information to obtain node features and edge features corresponding to the sample protein; wherein the node features are used to represent residue information, and the edge features are used to represent the spatial geometric relationship between adjacent residues; Alternately updating node features and edge features through multiple message passing layers of the graph neural network model to be trained to obtain updated target node features and target edge features; wherein the number of neighbor nodes aggregated by each message passing layer increases sequentially; According to the corresponding target node features and target edge features, the graph neural network model to be trained is used to decode and predict each node in a random order, and output prediction sequence information; The network parameters of the graph neural network model are iterated according to the loss value calculated based on the predicted sequence information and the sample sequence information to obtain a trained amino acid sequence generation model.
2. The training method according to claim 1, wherein: The single node features include the periodic trigonometric function values corresponding to the torsion angles φ and ψ of the residue, the secondary structure and the structural confidence of the residue; The single edge feature includes the distance information between the preset atoms of the two residues and the confidence of the relative positions of the two residues.
3. The training method according to claim 1, characterized in that In the encoder of the graph neural network model, the plurality of message passing layers include at least a first message passing layer and a second message passing layer; The first message transmission layer is used to aggregate X neighbor nodes of the current node and the corresponding connection edges, and the second message transmission layer is used to aggregate Y neighbor nodes of the current node and the corresponding connection edges; Preferably, the message transmission layer further includes a third message transmission layer, and the third message transmission layer is used to aggregate Z neighbor nodes of the current node and corresponding connection edges; 0<X<Y<Z, X, Y, Z are all natural numbers.
4. The training method according to claim 3, characterized in that The node features and edge features are alternately updated through multiple message passing layers of the graph neural network model to be trained to obtain updated target node features and target edge features, including: Aggregate X neighboring nodes and their corresponding connection edges through the first message passing layer, update the initial node feature of the current node, and obtain the first node feature corresponding to the current node; based on the first node feature of the current node and the current node features of the neighboring nodes, update the initial edge feature of the connection edge and obtain the first edge feature corresponding to the current connection edge; Aggregate Y neighbor nodes and their corresponding connection edges through the second message passing layer, update the first node feature of the current node, and obtain the second node feature corresponding to the current node; based on the second node feature of the current node and the current node feature of the neighbor node, update the first edge feature of the connection edge to obtain the second edge feature corresponding to the current connection edge; When the third message passing layer exists, Z neighbor nodes and corresponding connection edges are aggregated through the third message passing layer, the second node feature of the current node is updated, and the target node feature corresponding to the current node is obtained; according to the target node feature of the current node and the current node feature of the neighbor node, the second edge feature of the connection edge is updated to obtain the target edge feature corresponding to the current connection edge.
5. The training method according to claim 3, characterized in that: The number of message passing layers in the decoder of the graph neural network model is the same as the number of message passing layers in the encoder; wherein: The graph neural network model to be trained decodes and predicts each node in a random order and outputs prediction sequence information, including: Determine the current node to be predicted according to the random mask matrix; updating the target node feature of the node to be predicted in a loop in sequence according to the multiple message passing layers in the decoder to obtain the corresponding node update feature; According to the residue type probability distribution converted from the node update feature, the corresponding residue type is predicted and the corresponding predicted sequence information is generated.
6. The training method according to any one of claims 1 to 5, characterized in that: The loss value is calculated according to the following hybrid loss function, including: Among them, -log(probs(S i )) is the negative log-likelihood loss, which is used to measure the logarithmic difference between the predicted sequence information and the sample sequence information, probs(S i ) is the probability value of the residue category predicted by node i, pLDDT i Indicates the confidence of the prediction of residue i, mask i 0 or 1.
7. The training method according to any one of claims 1 to 6, characterized in that: The obtaining of training data includes: According to the first preset percentage of sequence identity, each real protein is clustered to obtain a plurality of first protein clusters; preferably, the resolution of the real protein is less than and / or the number of residues is less than 1000; Clustering the simulated proteins according to a dual clustering rule to obtain a plurality of second protein clusters, wherein the dual clustering rule includes clustering according to a second preset percentage of sequence overlap and / or a third preset percentage of sequence identity, and clustering according to a fourth preset percentage of structural overlap; Sample proteins are determined from each of the first protein cluster and the second protein cluster.
8. The training method according to claim 7, characterized in that: The determining of sample proteins from each of the first protein clusters and the second protein clusters comprises: A preset number of real proteins are selected from each of the first protein clusters as sample proteins, and simulated proteins with a confidence score greater than a preset confidence threshold are selected from each of the second protein clusters as sample proteins.
9. A method for generating an amino acid sequence, characterized in that: include: Obtain the main chain structure information of the target protein; The main chain structure information is input into the amino acid sequence generation model trained by the training method according to any one of claims 1 to 8, and the corresponding target amino acid sequence is output.
10. The generation method according to claim 9, characterized in that The obtaining of the main chain structure information of the target protein includes: encoding and obtaining the node features and edge features corresponding to the target protein according to the main chain structure information; wherein: The node features include some or all of the following features: The torsion angles φ and ψ of the residues were respectively expanded by the fourth-order Fourier transform to obtain the 16-dimensional periodic features, sin(nφ), cos(nφ), and sin(nψ), cos(nψ), where n = 1 to 4; The secondary structure of the residues, including: α-helix, 3 10 - Helix, π-helix, extended β-strand, isolated β-bridge, turn, bend, and coiled / random structures; Structural confidence of the residues; The edge features include some or all of the following features: The nitrogen atoms, carbon atoms, and α atoms, C atoms, O atoms, and virtual C β The distance information between any two atoms; The confidence of the predicted relative position of residue i to residue j.
11. A training device for an amino acid sequence generation model, characterized in that: include: A sample acquisition module is used to acquire training data, wherein the training data includes sample structure information of the sample protein and corresponding sample sequence information; A feature embedding module is used to perform feature encoding on the sample structure information to obtain node features and edge features corresponding to the sample protein; wherein the node features are used to represent residue information, and the edge features are used to represent the spatial geometric relationship between adjacent residues; A message passing module is used to alternately update node features and edge features through multiple message passing layers of the graph neural network model to be trained to obtain updated target node features and target edge features; wherein the number of neighbor nodes aggregated by each message passing layer increases sequentially; A node decoding module is used to decode and predict each node in a random order through the graph neural network model to be trained according to the corresponding target node features and target edge features, and output prediction sequence information; A back propagation module is used to iterate the network parameters of the graph neural network model according to the loss value calculated based on the predicted sequence information and the sample sequence information to obtain a trained amino acid sequence generation model.
12. A device for generating an amino acid sequence, characterized in that: include: Information acquisition module, used to obtain the main chain structure information of the target protein; A sequence generation module is used to input the main chain structure information into the amino acid sequence generation model trained by the training method according to any one of claims 1 to 8, and output the corresponding target amino acid sequence.
13. An electronic device, characterized in that: include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to execute the training method according to any one of claims 1 to 8 or the generation method according to any one of claims 9 to 10.
14. A computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the training method according to any one of claims 1 to 8 or the generation method according to any one of claims 9 to 10.
15. A computer program product, characterized in that The computer program product comprises computer instructions, which, when executed by a processor, implement the training method according to any one of claims 1 to 8 or the generating method according to any one of claims 9 to 10.
Citation Information
Patent Citations
Drug-target interaction prediction method and apparatus, device, and storage medium
WO2022222231A1
Predicting stability of protein on the basis of graph neural network
WO2025021024A1
Cited By
Phenotype prediction and phenotype prediction model training method
CN121459933A
Protein inverse folding method based on conditional information
CN121922190A