Complex structure prediction method and device, computer equipment and storage medium

CN116959577BActive Publication Date: 2026-09-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310631412.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2026-09-11
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

[0004]然而,基于搜索的通用蛋白质对接方法在没有蛋白质结合口袋相关的先验信息的情况下,盲对接的搜索空间巨大且崎岖,效率较低且准确性较差

Benefits of technology

[0016] On the other hand, embodiments of this application provide a computer program product including computer instructions stored in a computer-readable storage medium. A terminal's processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal to perform the methods described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116959577B_ABST
    Figure CN116959577B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of complex structure prediction method, device, computer equipment and storage medium, involve artificial intelligence field.The application includes: obtaining antigen peptide sequence, major histocompatibility complex MHC structure and T cell receptor TCR structure;Antigen peptide sequence and MHC structure are input into first graph neural network, and the first structure prediction result of first complex is obtained, and first complex is pMHC complex, and first graph neural network is used to flexible docking of antigen peptide sequence and MHC structure;TCR structure and first structure prediction result are input into second graph neural network, and the second structure prediction result of second complex is obtained, and second complex is TCR-pMHC complex, and second graph neural network is used to rigid docking of TCR structure and first structure prediction result.The method of the embodiment of the application can realize the efficient prediction of TCR-pMHC complex structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus, computer device, and storage medium for predicting the structure of a complex. Background Technology

[0002] T cells (thymus-dependent lymphocytes) are an important component of the adaptive immune system, capable of responding specifically to certain pathogens. In the adaptive immune response, the MHC (Major Histocompatibility Complex) of infected cells presents peptides (antigen peptides) to the cell surface, where they are specifically recognized by the T cell receptor (TCR), forming the TCR-pMHC complex (T Cell Receptor-peptide-Major Histocompatibility Complex). Understanding the structure of this complex at the molecular level is crucial for designing TCR-based therapies and vaccines.

[0003] Currently, there is no method for predicting the structure of the TCR-pMHC complex in related technologies. In general protein docking structure prediction methods, search-based general protein docking methods (such as ClusPro and PatchDock) are typically used. These methods sample pre-assumed binding modes using a search algorithm and evaluate them using an energy scoring function; based on the evaluation scores, all assumed binding modes are ranked, and the binding mode with the highest score is finally selected as the predicted structure.

[0004] However, search-based general protein docking methods suffer from a huge and rugged search space for blind docking without prior information about protein binding pockets, resulting in low efficiency and poor accuracy. Summary of the Invention

[0005] This application provides a method, apparatus, computer device, and storage medium for predicting the structure of a complex. The technical solution is as follows:

[0006] On one hand, embodiments of this application provide a method for predicting the structure of a complex, the method comprising:

[0007] Obtain the antigen peptide sequence, the major histocompatibility complex (MHC) structure, and the T cell receptor (TCR) structure;

[0008] The antigen peptide sequence and the MHC structure are input into a first graph neural network to obtain the first structure prediction result of the first complex, which is a pMHC complex. The first graph neural network is used to perform flexible docking of the antigen peptide sequence and the MHC structure.

[0009] The TCR structure and the prediction result of the first structure are input into the second graph neural network to obtain the prediction result of the second structure of the second complex, which is a TCR-pMHC complex. The second graph neural network is used to rigidly connect the prediction results of the TCR structure and the first structure.

[0010] On the other hand, embodiments of this application provide a complex structure prediction device, the device comprising:

[0011] The acquisition module is used to acquire antigen peptide sequences, major histocompatibility complex (MHC) structures, and T cell receptor (TCR) structures.

[0012] The first prediction module is used to input the antigen peptide sequence and the MHC structure into the first graph neural network to obtain the first structure prediction result of the first complex, wherein the first complex is a pMHC complex, and the first graph neural network is used to perform flexible docking of the antigen peptide sequence and the MHC structure.

[0013] The second prediction module is used to input the TCR structure and the prediction result of the first structure into the second graph neural network to obtain the prediction result of the second structure of the second complex, wherein the second complex is a TCR-pMHC complex, and the second graph neural network is used to rigidly connect the prediction results of the TCR structure and the first structure.

[0014] On the other hand, embodiments of this application provide a computer device including a processor and a memory; the memory stores at least one instruction, which is executed by the processor to implement the method described above.

[0015] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction that is loaded and executed by a processor to implement the method described above.

[0016] On the other hand, embodiments of this application provide a computer program product including computer instructions stored in a computer-readable storage medium. A terminal's processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal to perform the methods described above.

[0017] In this embodiment, by inputting the antigen peptide sequence and MHC structure into a first graph neural network, the antigen peptide sequence and MHC structure are flexibly docked to obtain the first structure prediction result of the pMHC complex; the TCR structure and the first structure prediction result of the pMHC complex are input into a second graph neural network for rigid docking to obtain the second structure prediction result of the TCR-pMHC complex, which can achieve efficient prediction of the structure of the TCR-pMHC complex. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a complex structure prediction method provided in an exemplary embodiment of this application;

[0020] Figure 2 This is a schematic diagram of a complex structure prediction method provided in an exemplary embodiment of this application;

[0021] Figure 3 This is a partial structural schematic diagram of a hierarchical diagram provided in an exemplary embodiment of this application;

[0022] Figure 4 This is a schematic diagram illustrating the global and local adjustments made to the initial structure prediction results of the first complex through iteration, provided by an exemplary embodiment of this application.

[0023] Figure 5 This is a schematic diagram illustrating the prediction result of determining the second structure of the second complex via a messaging network, provided in an exemplary embodiment of this application.

[0024] Figure 6 This is a schematic diagram of the internal structure of a messaging network provided in an exemplary embodiment of this application;

[0025] Figure 7 This is a schematic diagram of a message passing mechanism provided in an exemplary embodiment of this application;

[0026] Figure 8 This is a schematic diagram illustrating the training process of a first graph neural network and a second graph neural network provided in an exemplary embodiment of this application;

[0027] Figure 9 This is a bar chart comparing the TCR-pMHC prediction success rates of different models provided in an exemplary embodiment of this application;

[0028] Figure 10 This is a structural block diagram of a complex structure prediction device provided in an exemplary embodiment of this application;

[0029] Figure 11 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0031] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0032] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0033] The human immune system consists of innate immunity and adaptive immunity. Adaptive immunity is an immune response that is generated upon contact with a specific pathogen (antigen) and is initiated against that pathogen.

[0034] T cells are an important component of the adaptive immune system, capable of responding specifically to certain pathogens, thereby activating the immune system to eliminate cancer cells, invading viruses, and other pathogens. Therefore, T cells play a vital role in maintaining human health.

[0035] T cell receptors are specific receptors on the surface of T cells that are responsible for recognizing antigens presented by the MHC and activating the immune response.

[0036] MHC is the major histocompatibility complex of infected cells, used to bind to antigenic peptides and present them on the cell surface as antigens. The pMHC complex formed by MHC and antigenic peptides can bind to T cell receptors on the cell surface, activating T cells and thus inducing a T cell-mediated immune response.

[0037] Antigenic peptides are compounds formed by three or more amino acids linked together by peptide bonds within infected cells.

[0038] TCR-T therapy (T Cell Receptor-gene engineered T cells) is a T-cell therapy that utilizes genetically engineered cell receptors. This method enhances the affinity of TCRs that specifically recognize tumor-associated antigens and strengthens the immune cells by transducing chimeric antigen receptors or TCRα / β heterodimers into ordinary T cells. This allows T lymphocytes to re-establish efficient recognition of target cells and exert a stronger anti-tumor immune effect in vivo. Clinical trials have demonstrated that TCR-T therapy is a viable strategy for treating cancer and also shows promise in treating other diseases, such as tuberculosis and HIV / AIDS.

[0039] In the adaptive immune response, the MHC of infected cells presents antigenic peptides to the cell surface, where they are specifically recognized by T cell receptors. This process plays a crucial role in TCR-T therapy. Therefore, studying the TCR-pMHC complex formed by the binding of MHC, antigenic peptides, and TCR is of great significance for the research of TCR-T therapy and the development of related drugs.

[0040] Experimental crystallization and structural determination of the TCR-pMHC complex is an expensive and time-consuming process. Therefore, predictive methods for the structure of the TCR-pMHC complex can greatly reduce costs and make a significant contribution to the research of TCR-T therapy.

[0041] Currently, there is no method for predicting the structure of the TCR-pMHC complex in related technologies. In general protein docking structure prediction methods, search-based general protein docking methods (such as ClusPro and PatchDock) are typically used. These methods sample pre-assumed binding modes using a search algorithm and evaluate them using an energy scoring function; based on the evaluation scores, all assumed binding modes are ranked, and the binding mode with the highest score is finally selected as the predicted structure.

[0042] However, search-based general protein docking methods suffer from a huge and rugged search space for blind docking without prior information about protein binding pockets, resulting in low efficiency and poor accuracy.

[0043] This application proposes a method for predicting the structure of a complex, which can predict the structure of the TCR-pMHC complex with high efficiency.

[0044] See Figure 1 , Figure 1 This is a flowchart of a complex structure prediction method provided in an exemplary embodiment of this application. The method includes the following steps:

[0045] Step 101: Obtain the antigen peptide sequence, the major histocompatibility complex (MHC) structure, and the T cell receptor (TCR) structure.

[0046] Optionally, the antigenic peptide sequence is a compound sequence formed by three or more amino acids linked by peptide bonds. The antigenic peptide sequence includes the number of each amino acid in the antigenic peptide and their linkage relationship, but does not have a molecular structure.

[0047] Optionally, the MHC structure is the protein structure of the major histocompatibility complex, including the coordinates of each amino acid and atom in the MHC and their interconnections, and has a molecular structure.

[0048] Optionally, the TCR structure is the protein structure of the T cell receptor, including the coordinates of each amino acid in the TCR and their linkage relationships, and has a molecular structure.

[0049] In some embodiments, the antigen peptide sequence, MHC structure, and TCR structure can be obtained based on user input.

[0050] Step 102: Input the antigen peptide sequence and MHC structure into the first graph neural network to obtain the first structure prediction result of the first complex. The first complex is a pMHC complex. The first graph neural network is used to perform flexible docking of the antigen peptide sequence and MHC structure.

[0051] The first complex, the pMHC complex, is a complex formed after the antigen peptide binds to MHC.

[0052] Flexible docking is a molecular docking method in which the antigen peptide sequence and MHC structure undergo free conformational changes during the docking process. Flexible docking is commonly used to precisely examine the docking and folding of molecules. In the pMHC complex formed through flexible docking, both the antigen peptide and the MHC possess molecular structures.

[0053] Optionally, the first graph neural network is a graph neural network with rotation and translation equivariance, such as a hierarchical graph neural network with rotation and translation equivariance.

[0054] When the first graph neural network has rotational and translational isovariability, for the same antigen peptide sequence and MHC structure, arbitrarily rotating and translating them before inputting them into the first graph neural network will not change the prediction result of the pMHC complex. Therefore, the first graph neural network with rotational and translational isovariability can improve the robustness and stability of the prediction results.

[0055] Optionally, the first graph neural network can be a hierarchical graph neural network, which includes different hierarchical encoders for encoding antigenic peptides and MHC at different levels, such as atomic-level encoding and amino acid residue-level encoding. The amino acid residue-level encoding reflects the overall structure of the antigenic peptide and MHC, while the atomic-level encoding reflects their local structure.

[0056] Optionally, the first structure prediction result of the first complex can be an atomic-level prediction result, where the coordinates of a node in the first structure prediction result represent the coordinates of an atom.

[0057] Optionally, the first structure prediction result of the first complex can be the prediction result at the amino acid residue level, where the coordinates of a node in the first structure prediction result represent the coordinates of the α carbon atom of an amino acid residue.

[0058] Optionally, the first structure prediction result of the first complex can be a mixture of atomic level and amino acid residue level prediction results, where the coordinates of a node in the first structure prediction result represent the coordinates of an atom or the coordinates of the α carbon atom of an amino acid residue.

[0059] Step 103: Input the TCR structure and the first structure prediction results into the second graph neural network to obtain the second structure prediction result of the second complex. The second complex is a TCR-pMHC complex. The second graph neural network is used to rigidly connect the TCR structure and the first structure prediction results.

[0060] The second complex, the TCR-pMHC complex, is a complex formed by the specific recognition of pMHC by TCR.

[0061] Rigid docking is a molecular docking method in which the conformation of pMHC and TCR remains unchanged during the rigid docking process. Rigid docking is typically used to examine the docking processes of large molecules such as proteins. Since both pMHC and TCR are proteins with large molecular weights, treating the docking process of pMHC and TCR as rigid docking, as an idealized yet realistic assumption, can improve the efficiency of prediction while ensuring the accuracy of the complex prediction results.

[0062] Optionally, the second graph neural network can also be a graph neural network with rotational and translational equivariance, such as SE(3) equivariant graph neural networks (Special Euclidean group equivariant graph neural networks).

[0063] When the second graph neural network has rotation and translation equivariance, for the same TCR and pMHC structures, arbitrarily rotating and translating them before inputting them into the second graph neural network will not change the prediction result of the TCR-pMHC complex. Therefore, the second graph neural network with rotation and translation equivariance can also improve the robustness and stability of the prediction result.

[0064] Optionally, the second graph neural network can be a message-passing network with a message-passing mechanism. For example, the second graph neural network can be IEGMN (Independent E(3)-equivariant Graph Matching Network). IEGMN is a graph neural network obtained by improving upon the SE(3) equivariant graph neural network and the graph matching network. The message-passing network can update the node feature vectors and node coordinates of the nodes in the graph based on the message-passing mechanism.

[0065] Optionally, the second structure prediction result of the second complex can be the prediction result at the amino acid residue level, where the coordinates of a node in the second structure prediction result represent the coordinates of the α carbon atom of an amino acid residue.

[0066] In summary, by inputting the antigen peptide sequence and MHC structure into a first graph neural network and performing flexible docking between the antigen peptide sequence and MHC structure, the first structure prediction result of the pMHC complex is obtained; by inputting the TCR structure and the first structure prediction result of the pMHC complex into a second graph neural network and performing rigid docking, the second structure prediction result of the TCR-pMHC complex is obtained, thus achieving efficient prediction of the TCR-pMHC complex structure.

[0067] See Figure 2 , Figure 2 This is a schematic diagram of a complex structure prediction method provided in an exemplary embodiment of this application.

[0068] like Figure 2 As shown, the antigen peptide sequence 201 and the MHC structure 202-2 are input into the first graph neural network 210 to obtain the first structure prediction result of the first complex, namely the first structure prediction result of the p-MHC complex structure 204.

[0069] In some embodiments, the obtained MHC data may not have a molecular structure, but exist in the form of MHC sequence 202-1. Therefore, MHC sequence 202-1 can be input into the first protein structure prediction model 230 to predict MHC structure 202-2; then MHC structure 202-2 and antigen peptide sequence 201 can be input into the first graph neural network 210 to obtain the first structure prediction result of p-MHC complex structure 204.

[0070] Optionally, the first protein structure prediction model 230 can be a protein structure prediction model such as AlphaFold2, RoseTTAFold, Tfold, or other protein structure prediction models.

[0071] The p-MHC complex structure 204 and the TCR structure 203-2 are input into the second graph neural network 220 to obtain the second structure prediction result of the second complex, namely the second structure prediction result of the TCR-pMHC complex structure 205.

[0072] In some embodiments, the obtained TCR data may not have a molecular structure, but exist in the form of TCR sequence 203-1. Therefore, TCR sequence 203-1 can be input into the second protein structure prediction model 240 to predict TCR structure 203-2; then TCR structure 203-2 and p-MHC complex structure 204 can be input into the second graph neural network 220 to obtain the second structure prediction result of TCR-pMHC complex structure 205.

[0073] Optionally, the second protein structure prediction model 240 is a protein structure prediction model such as AlphaFold2, RoseTTAFold, Tfold, or other protein structure prediction models.

[0074] In this embodiment, the obtained MHC and TCR sequences can be predicted using a first protein structure prediction model and a second protein structure prediction model. The predicted MHC and TCR structures are then used in a first and a second graph neural network to obtain the second structure prediction result of the second complex. This method achieves end-to-end prediction results from monomeric sequences to complex structures, exhibiting high prediction efficiency.

[0075] In the process of encoding antigen peptide sequences and MHC structures, an atom in the antigen peptide sequence and MHC structure can be encoded as a node, or an amino acid residue can be encoded as a node.

[0076] To balance prediction effectiveness and efficiency, this embodiment employs an amino acid residue-level encoding of the entire MHC structure and an atomic-level encoding of the docking interface between the antigen peptide sequence and the MHC structure. This approach ensures prediction accuracy near the docking interface while improving overall prediction efficiency.

[0077] The docking interface refers to the interface at which the antigen peptide sequence binds to the MHC structure after docking. The docking interface can be estimated in advance.

[0078] In one possible implementation, the first graph neural network is a hierarchical graph neural network, which includes a hierarchical encoder. In some embodiments, the hierarchical encoder is divided into three types, with the following functions:

[0079] (1) First-level encoder.

[0080] The first-level encoder is used to encode the amino acid residues in the MHC structure, obtaining the first node features of the first type of nodes. The first type of nodes characterizes the α-carbon atom of the amino acid residues in the MHC structure.

[0081] (2) Second-level encoder.

[0082] The second-level encoder is used to encode the atoms in the docking interface, obtaining the second-node features of the second type of nodes. The second-type nodes represent atoms in the docking interface; that is, the second-type nodes may be atoms on the MHC structure in the docking interface, or atoms on the antigen peptide sequence in the docking interface.

[0083] (3) Third-level encoder.

[0084] The third-level encoder is used to encode the amino acid residues in the docking interface, obtaining the third-node features of the third type of nodes. The third type of nodes characterizes the α-carbon atom of the amino acid residues in the docking interface; that is, the third type of node may be the α-carbon atom of an amino acid residue on the MHC structure in the docking interface, or it may be the α-carbon atom of an amino acid residue on the antigenic peptide sequence in the docking interface.

[0085] See Figure 3 , Figure 3 This is a partial structural schematic diagram of a hierarchical diagram provided in an exemplary embodiment of this application.

[0086] like Figure 3As shown, the first-level encoder encodes amino acid residues in the MHC structure, obtaining the first node feature of the first type of node 311. The first type of node 311 is shown as a white circle with a thick border, representing an amino acid residue in the MHC. The first type of node 311 may be an α-carbon atom of an amino acid residue at the docking interface in the MHC structure, or it may be an α-carbon atom of an amino acid residue outside the docking interface in the MHC structure (hereinafter, it is assumed that the first type of node 311 represents an α-carbon atom of an amino acid residue at the docking interface in the MHC structure, and that this α-carbon atom corresponds to the same α-carbon atom as the third type of node 332).

[0087] The second-level encoder encodes atoms on the MHC structure and atoms on the antigen peptide sequence at the docking interface, resulting in second-class nodes. These second-class nodes may include second-class node 321 and second-class node 322.

[0088] The second type of node 321 is shown as a white circle with a thin border, representing atoms at the docking interface in the MHC.

[0089] The second type of node 322 is shown as a gray circle with a thin border, representing atoms at the docking interface of the antigenic peptide.

[0090] The third-level encoder encodes the amino acid residues on the MHC structure and the amino acid residues on the antigenic peptide sequence in the docking interface, thus obtaining the third node features of the third type of node.

[0091] The third type of node can include third type nodes 331, 332 and 333. These three third type nodes are all shown as white circles with thick borders, representing amino acid residues in the MHC. All three third type nodes are α carbon atoms of amino acid residues at the docking interface in the MHC.

[0092] The third type of node may also include third type nodes 334 and 335. Both of these third type nodes are shown as gray circles with thick borders, representing amino acid residues in the antigenic peptide. Both of these third type nodes are α carbon atoms of amino acid residues at the docking interface of the antigenic peptide.

[0093] Regarding the method for determining the third node features of the third type of node, in some embodiments, it can be determined by fusing the first node features of the first type of node and the second node features of the second type of node.

[0094] In some embodiments, a third node feature of the third type of node can be obtained by fusing the first node feature corresponding to the α carbon atom represented by the third type of node and the second node feature corresponding to the atom connected to the α carbon atom represented by the third type of node through a third-level encoder.

[0095] by Figure 3Taking the third type node 332 shown as an example, in some embodiments, the first node features of the first type node 311 and the second node features of the second type nodes 323 and 324 can be fused (e.g., directly added) using a third-level encoder to obtain the third node features of the third type node 332. Here, the first type node 311 and the third type node 332 represent the α-carbon atom of the same amino acid residue at the docking interface in MHC, the difference being that the first type node 311 is obtained through a first-level encoder, while the third type node 332 is obtained through a third-level encoder. The second type nodes 323 and 324 represent atoms connected to the α-carbon atom of the first type node 311, that is, MHC atoms with edges existing between them and the α-carbon atom of the first type node 311.

[0096] In one possible implementation, nodes in a hierarchical graph neural network can have edges between them. For example, there may be an edge between third-class nodes 331 and 332, an edge between third-class node 332 and second-class node 323, or an edge between third-class node 331 and first-class node 311.

[0097] In some embodiments, the edge between three type nodes indicates that two amino acid residues are adjacent amino acid residues. For example, three type nodes 331 and 332 characterize the α carbon atoms of two adjacent amino acid residues.

[0098] In some embodiments, the edge between the third type of node and the second type of node represents the dependency relationship between an atom and an amino acid residue. For example, if there is an edge between the third type of node 332 and the second type of node 323, it means that the atom represented by the second type of node 323 belongs to the amino acid residue represented by the third type of node 332.

[0099] In some embodiments, the edge between the third type of node and the first type of node indicates that the two nodes correspond to the same α-carbon atom. For example, the third type of node 331 and the first type of node 311 correspond to the α-carbon atom of an amino acid residue on the same MHC docking interface.

[0100] In one possible implementation, the existence of edges between nodes can be determined based on distance. For example, the K-nearest neighbor algorithm can be used to determine the edges between nodes based on their distance.

[0101] In some embodiments, h(a_i, k) can represent the node features (including second-class and third-class node features) of a node in an antigenic peptide, and h(b_j, l) can represent the node features (including first-class, second-class, and third-class node features) of a node in the MHC. For example, h(a_i, k) represents the node feature of the k-th atom of the i-th amino acid residue in the antigenic peptide, and h(b_j, l) represents the node feature of the l-th atom of the j-th amino acid residue in the MHC.

[0102] In this embodiment, the MHC structure and antigenic peptide sequence are encoded by three hierarchical encoders, which can map the MHC structure as a whole at the amino acid residue level and map the MHC structure and antigenic peptide sequence at the docking interface at the atomic level. This can not only obtain more accurate prediction results of the first complex, but also improve the prediction efficiency.

[0103] After mapping the MHC structure and antigen peptide sequence using a hierarchical encoder, the first structure prediction result of the first complex can be determined based on the first node features, the second node features, and the third node features.

[0104] During the docking process between MHC and antigenic peptide, the antigenic peptide sequence folds into an antigenic peptide structure with molecular structure. Influenced by intermolecular and interatomic forces within the molecular structure, the molecular conformation of both the MHC structure and the antigenic peptide structure changes.

[0105] For the reasons mentioned above, the initial structure prediction of the first complex is not accurate enough, and further global and local adjustments are needed based on the inter-residue and inter-atomic forces.

[0106] In some embodiments, the initial structure prediction results of the first complex can be globally and locally adjusted through multiple iterations.

[0107] See Figure 4 , Figure 4 This is a schematic diagram illustrating the global and local adjustments made to the initial structure prediction results of the first complex through iteration, as provided in an exemplary embodiment of this application.

[0108] like Figure 4 As shown, the computer device inputs the antigen peptide sequence 411 and MHC structure 412 into the hierarchical encoder 420 for encoding, resulting in an encoded hierarchical graph 430. The encoded hierarchical graph 430 includes first node features, second node features, and third node features. In some embodiments, the encoded hierarchical graph 430 can be input into a fully connected layer to obtain the initial structure prediction result of the first complex.

[0109] In some embodiments, multiple iterations can be performed to make global and local adjustments to the initial structure prediction results of the first complex.

[0110] Taking the i-th iteration as an example (i is a positive integer), in the i-th iteration process, the initial structure prediction result of the first complex is determined firstly based on the encoded hierarchical graph 430 (containing the first node feature, the second node feature and the third node feature) through a fully connected layer.

[0111] The computer device determines the inter-residue forces between amino acid residues based on the first node features and the third node features in the encoded hierarchical graph 430.

[0112] In some embodiments, the first node features and the third node features in the encoded hierarchical graph 430 are input into the global adjustment network 440 to output the inter-residue interaction forces between amino acid residues.

[0113] Among them, the inter-residue forces include the forces between amino acid residues of the first type of node and amino acid residues of the first type of node, the forces between amino acid residues of the first type of node and amino acid residues of the third type of node, and the forces between amino acid residues of the third type of node and amino acid residues of the third type of node.

[0114] The computer device determines the interatomic forces between atoms based on the second node features in the encoded hierarchical graph 430.

[0115] In some embodiments, the second node features in the encoded hierarchical graph 430 are input into the local adjustment network 450, and the interatomic forces between atoms are output.

[0116] Among them, the interatomic forces between atoms include the interatomic forces between MHC atoms and antigen peptide atoms at the docking interface, the interatomic forces between MHC atoms at the docking interface, and the interatomic forces between antigen peptide atoms at the docking interface.

[0117] After obtaining the inter-residue forces between amino acid residues and the inter-atomic forces between atoms, the coordinates of the amino acid residues in the initial structure prediction results of the first complex can be adjusted based on the inter-residue forces, and the coordinates of the atoms in the amino acid residues in the initial structure prediction results of the first complex can be adjusted based on the inter-atomic forces to obtain the i-th round of structure prediction results.

[0118] If the iteration stopping condition is met in the i-th iteration, the structure prediction result of the i-th iteration is determined as the first structure prediction result 460 of the first complex.

[0119] Optionally, the iteration stopping condition is that the number of iterations reaches a threshold (e.g., 8 times), or the degree of change in the coordinates of the amino acid residues and the degree of change in the coordinates of the atoms in the amino acid residues are less than a preset value.

[0120] If the iteration stopping condition is not met in the i-th iteration, the features of the first node, the second node, and the third node are updated based on the structural prediction results of the i-th iteration.

[0121] In some embodiments, the structure prediction result of the i-th round can be input again into the hierarchical encoder 420 for re-encoding to obtain an updated hierarchical graph 430. The updated hierarchical graph 430 contains updated first node features, second node features, and third node features. The initial structure prediction result of the (i+1)-th round is then predicted based on the updated hierarchical graph 430, and the initial structure prediction result is globally and locally adjusted based on the global adjustment network 440 and the local adjustment network 450 to obtain the structure prediction result of the (i+1)-th round... This process continues until the iteration stopping condition is met.

[0122] In some embodiments, the multi-round structure prediction results of the first complex obtained from multiple iterations can be sampled and evaluated using an energy scoring function; the multi-round structure prediction results can be sorted according to the evaluation scores, and the structure prediction result with the highest score can be selected as the first structure prediction result.

[0123] In this embodiment, by iteratively adjusting the initial structure prediction results of the first complex globally and locally, the coordinates of amino acid residues and atoms can be adjusted based on inter-residue forces and inter-atomic forces, thereby improving the accuracy of the first structure prediction results of the first complex.

[0124] To assemble the TCR structure and the pMHC complex structure, it is necessary to determine the location of protein-binding pockets on both the TCR and pMHC complex structures. Protein-binding pockets are cavities on the surface or inside a protein that are suitable for ligand binding.

[0125] In the embodiments of this application, the protein binding pockets on the TCR structure and the pMHC complex structure are also referred to as docking key points.

[0126] In some embodiments, the docking critical points on the TCR structure and the pMHC complex structure can be determined based on a messaging network.

[0127] See Figure 5 , Figure 5 This is a schematic diagram illustrating the prediction result of determining the second structure of the second complex via a messaging network, provided in an exemplary embodiment of this application.

[0128] like Figure 5 As shown, the first structure prediction result 511 of the first complex is first mapped to obtain the structure diagram 521 of the first complex; and the TCR structure 512 is mapped to obtain the structure diagram 522 of the TCR.

[0129] Optionally, the mapping can employ mapping methods with rotational and translational properties, such as mapping based on distance or amino acid types.

[0130] Optionally, the nodes in the first complex structure diagram 521 and the TCR structure diagram 522 represent the α-carbon atoms on the amino acid residues.

[0131] Optionally, the first complex structure diagram 521 and the TCR structure diagram 522 include node coordinates, node features, and edge information.

[0132] Next, the computer device inputs the first complex structure diagram 521 and the TCR structure diagram 522 into the message passing network 530 to predict the key point coordinates 540 of the docking key points on the TCR structure and the pMHC complex structure.

[0133] Optionally, the message passing network 530 is a neural network model with a message passing mechanism, such as IEGMN.

[0134] The docking keypoint is the protein binding pocket. In some embodiments, a docking keypoint on the TCR structure corresponds to a docking keypoint on the pMHC complex structure, indicating that the two docking keypoints theoretically overlap.

[0135] Optionally, the key point coordinates 540 of the docking key point are three-dimensional coordinates in three-dimensional space.

[0136] After obtaining the key point coordinates 540, the computer device overlays the corresponding key point coordinates on the TCR structure and the pMHC complex structure to obtain the second structure prediction result 550 of the second complex, that is, the second structure prediction result of the TCR-pMHC complex.

[0137] Optionally, the computer device may use the Kabsch point cloud overlay algorithm or other point cloud overlay algorithms to overlay the coordinates of corresponding key points on the TCR structure and the pMHC complex structure. This application embodiment does not limit this.

[0138] In this embodiment, the key point coordinates of the docking key points on the TCR structure and the pMHC complex structure are predicted by the message passing network. The corresponding key point coordinates can be superimposed by the point cloud overlay algorithm to achieve rigid docking of the TCR structure and the pMHC complex structure.

[0139] See Figure 6 , Figure 6 This is a schematic diagram of the internal structure of a messaging network provided in an exemplary embodiment of this application.

[0140] like Figure 6 As shown, the input of message passing network 61 is the structure diagram of the first complex and the structure diagram of TCR, and the output of message passing network 61 is the key point coordinates of the docking key points on the TCR structure and the pMHC complex structure.

[0141] The internal structure of the message passing network 61 includes a message generation layer 610, a message propagation layer 620, and an attention layer 630.

[0142] The message generation layer 610 is used to generate intra-graph interaction features (shown as m1) and inter-graph interaction features (shown as m2) for each node in the first complex structure diagram and the TCR structure diagram. Among them, the intra-graph interaction feature m1 represents the interaction information between nodes within the TCR structure diagram or the first complex structure diagram, and the inter-graph interaction feature m2 represents the interaction information between nodes between the TCR structure diagram and the first complex structure diagram.

[0143] In some embodiments, the node coordinates and node features of the first complex structure graph and the TCR structure graph can be updated through multiple iterations, passing through the message generation layer 610 and the message propagation layer 620. The updated node coordinates are denoted as x′, and the updated node features are denoted as h.

[0144] The message generation layer 610 includes an MLP (Multilayer Perceptron) 611 and a shallow neural network 612.

[0145] Specifically, in the first iteration, the updated node features equal the initial node features, i.e., h = f; the updated node coordinates equal the initial node coordinates, i.e., x′ = x.

[0146] Taking the j-th iteration in a multi-round iteration as an example, MLP611 is used to process the updated node features (shown as h), updated node coordinates (shown as x′), and edge information (shown as e) in the first complex structure graph or TCR structure graph obtained in the (j-1)-th round to obtain the intra-graph interaction features (shown as m1) of each node in the first complex structure graph or TCR structure graph corresponding to the j-th round.

[0147] The shallow neural network 612 is used to process the updated node features (indicated by h) in the first complex structure graph and TCR structure graph obtained in the (j-1)th round to obtain the inter-graph interaction features (indicated by m2) of each node in the first complex structure graph and TCR structure graph in the jth round.

[0148] At this point, the message generation layer 610 has obtained the intra-graph interaction features and inter-graph interaction features of the j-th round through MLP 611 and shallow neural network 612, respectively. Next, the message propagation layer 620 will propagate the intra-graph interaction features and inter-graph interaction features of the j-th round to update the node coordinates and node features of each node.

[0149] The message propagation layer 620 includes MLP621 and MLP622.

[0150] MLP621 is used to pass the intra-graph interaction features corresponding to each node in the first complex structure diagram or TCR structure diagram to the neighboring nodes, so that the neighboring nodes update their node coordinates based on the initial node coordinates (shown as x), the node coordinates after the (j-1)th round update (shown as x′), and the intra-graph interaction features in the jth round (shown as m1), to obtain the node coordinates after the jth round update (shown as x′).

[0151] See Figure 7 , Figure 7 This is a schematic diagram of a message passing mechanism provided in an exemplary embodiment of this application.

[0152] like Figure 7 As shown, nodes x1, x2, x3, and x4 in the diagram represent nodes in the first complex structure diagram or the TCR structure diagram. Node x1 passes its corresponding intra-graph interaction feature v1 to its neighboring nodes, such as nodes x2 and x3; node x2 passes its corresponding intra-graph interaction feature v2 to its neighboring nodes, such as nodes x1 and x3; node x3 passes its corresponding intra-graph interaction feature v3 to its neighboring nodes, such as nodes x1, x2, and x4; node x4 passes its corresponding intra-graph interaction feature v4 to its neighboring nodes, such as node x3… and so on. Each node receives the intra-graph interaction feature passed from its neighboring nodes.

[0153] The MLP622 is used to transfer the inter-graph interaction features corresponding to each node in the first complex structure diagram and the TCR structure diagram to the nodes in the other diagram, so that the nodes in the other diagram can update their node features based on the intra-graph interaction features in the j-th round (shown as m1), the inter-graph interaction features in the j-th round (shown as m2), the node features updated in the (j-1)-th round (shown as h), and the initial node features (shown as f), to obtain the node features updated in the j-th round (shown as h).

[0154] At this point, message propagation layer 620 has updated the node coordinates and node features of each node in the first complex structure diagram and TCR structure diagram in the j-th round through MLP621 and MLP622, respectively.

[0155] If the iteration stopping condition is met, the node coordinates and node features of each node in the first complex structure diagram and TCR structure diagram after the j-th round update are input into the attention layer 630 to predict the key point coordinates.

[0156] Optionally, attention layer 630 is a machine learning model with an attention mechanism.

[0157] Optionally, the iteration stopping condition is when the number of iterations reaches a threshold (e.g., 8 times).

[0158] If the iteration stopping condition is not met, the node coordinates and node features updated in the j-th round are input again into the message generation layer 610 and the message propagation layer 620 to update the node coordinates and node features in the (j+1)-th round until the iteration stopping condition is met.

[0159] In this embodiment, the message passing network can calculate the intra-graph interaction features and inter-graph interaction features corresponding to each node in the first complex structure graph and the TCR structure graph through the message generation layer; through the message propagation layer, the intra-graph interaction features and inter-graph interaction features of the nodes can be passed to adjacent nodes or nodes in another graph to iteratively update the node coordinates and node features, so that the node coordinates and node features of each node are integrated with the intra-graph and inter-graph interaction information, thereby improving the prediction effect of key point coordinates after inputting into the attention layer, making the docking of TCR and pMHC more accurate.

[0160] During the training of the first and second graph neural networks, since the available sample size of the TCP-pMHC complex is relatively small, the first and second graph neural networks can be pre-trained on other datasets first, and then fine-tuned on the dataset corresponding to the TCP-pMHC complex.

[0161] In some embodiments, a first graph neural network is pre-trained based on a first dataset, a second graph neural network is pre-trained based on a second dataset, and the pre-trained first and second graph neural networks are fine-tuned based on a third dataset.

[0162] The third dataset contains samples of TCR-pMHC complexes, and the first and second datasets have a larger data volume than the third dataset.

[0163] See Figure 8 , Figure 8 This is a schematic diagram illustrating the training process of a first graph neural network and a second graph neural network provided in an exemplary embodiment of this application.

[0164] like Figure 8As shown, the initial first graph neural network is first pre-trained on the first dataset 821 to obtain the pre-trained first graph neural network; the pre-trained first graph neural network is then fine-tuned on the third dataset 823 to obtain the fine-tuned first graph neural network.

[0165] The initial second graph neural network is first pre-trained on the second dataset 822 to obtain the pre-trained second graph neural network; the pre-trained second graph neural network is then fine-tuned on the third dataset 823 to obtain the fine-tuned second graph neural network.

[0166] Optionally, the first dataset 821 can be collected from SAbDab (the Structural Antibody Database). For example, an antigen sample dataset and an antibody sample dataset from SAbDab can be collected as the first dataset, and a first graph neural network can be pre-trained on the first dataset.

[0167] Optionally, the second dataset 822 can be acquired from DIPS (Database of Interacting Protein Structures). For example, a protein docking sample dataset from DIPS can be acquired as the second dataset, and a second graph neural network can be pre-trained on the second dataset.

[0168] Optionally, the third dataset 823 can be collected via STCRDab (the Structural T-Cell Receptor Database).

[0169] In some embodiments, TCR-pMHC complex samples can be screened from a third dataset based on screening criteria.

[0170] Optionally, the screening criteria are that the TCR structure has two chains, the antigen chain contains a short peptide, and the MHC structure has at least one chain. These screening criteria ensure that the resulting third dataset meets the biological characteristics required for training samples in both the first and second graph neural networks.

[0171] In some embodiments, the sample TCR-pMHC complex in the third dataset can be split to obtain sample TCR, sample antigen peptide, sample MHC and sample pMHC.

[0172] In some embodiments, the first graph neural network can be fine-tuned based on the sample antigen peptide, sample MHC, and sample pMHC.

[0173] For example, the sample antigen peptide and sample MHC can be input into a pre-trained first graph neural network to obtain the predicted coordinates of atoms in the predicted pMHC; based on the true coordinates and predicted coordinates of atoms in the sample pMHC, a first loss is calculated; and the first graph neural network is fine-tuned based on the first loss.

[0174] Optionally, the first loss is the Huber loss.

[0175] In some embodiments, the second graph neural network can be fine-tuned based on sample pMHC, sample TCR, and sample TCR-pMHC complex.

[0176] For example, sample pMHC and sample TCR can be input into a second graph neural network to obtain predicted coordinates of nodes in the TCR-pMHC complex and predicted coordinates of key points; the key point coordinates are used to characterize the coordinates of the protein binding pocket; a second loss is calculated based on the true and predicted coordinates of nodes in the sample TCR-pMHC complex; a third loss is calculated based on the true and predicted coordinates of key points in the sample TCR-pMHC complex; and the second graph neural network is fine-tuned based on the sum of the second and third losses.

[0177] Optionally, the second loss is the mean squared error loss, and the third loss is the OT (Optimal Transport) loss.

[0178] By fine-tuning the second graph neural network based on the sum of the second and third losses, the prediction performance of keypoint coordinates and TCR-pMHC complex can be evaluated simultaneously, resulting in a more accurate prediction capability for the trained second graph neural network.

[0179] In one possible implementation, the split sample TCR, sample antigenic peptide, sample MHC, and sample pMHC can be randomly translated and rotated. Then, the first and second graph neural networks are fine-tuned based on the randomly translated and rotated data. For input data with the same structure, random translation and rotation do not affect the model's output. Therefore, fine-tuning the model with randomly translated and rotated data can improve the model's stability and robustness.

[0180] In this embodiment, the first graph neural network and the second graph neural network are pre-trained on the first dataset and the second dataset, which have a large amount of data, respectively. Then, the first graph neural network and the second graph neural network are fine-tuned on the third dataset. This can solve the problem of insufficient sample data and enhance the prediction effect of the model.

[0181] See Figure 9 , Figure 9This is a bar chart comparing the TCR-pMHC prediction success rates of different models provided in an exemplary embodiment of this application.

[0182] like Figure 9 As shown in the bar chart, the TCR-pMHC prediction success rates of the five models are illustrated. The first four models are general protein docking models based on search algorithms, namely the ClusPro model, HADDOCK model, LightDOCK model, and ZDOCK model, with TCR-pMHC prediction success rates of approximately 28%, 34%, 8%, and 17% respectively on the test set.

[0183] The fifth model, T-DOCK, is a model that uses the method described in the embodiments of this application. Its TCR-pMHC prediction success rate on the test set is approximately 32%.

[0184] according to Figure 9 As can be seen, the prediction success rates of the T-DOCK model and the HADDOCK model proposed in this application are quite similar. However, the HADDOCK model is a general protein docking method based on a search algorithm. Without prior information about the protein binding pocket, the search space for blind docking is huge and rugged, resulting in low efficiency. In contrast, using the T-DOCK model in this application, the computer device can quickly derive the prediction result of the second structure of the second complex based on the input data, and can predict the protein binding pocket, thus exhibiting higher prediction efficiency.

[0185] From the perspective of TCR-pMHC prediction success rate, the method of this application embodiment is comparable to or better than existing general protein docking models based on search algorithms. From the perspective of prediction efficiency, the method of this application embodiment has higher prediction efficiency. Therefore, whether in terms of TCR-pMHC prediction success rate or prediction efficiency, the method described in this application embodiment has greater advantages compared with general protein docking methods based on search algorithms.

[0186] See Figure 10 , Figure 10 This is a structural block diagram of a complex structure prediction device provided in an exemplary embodiment of this application. The device includes:

[0187] The acquisition module 1001 is used to acquire the antigen peptide sequence, the major histocompatibility complex (MHC) structure, and the T cell receptor (TCR) structure.

[0188] The first prediction module 1002 is used to input the antigen peptide sequence and the MHC structure into a first graph neural network to obtain the first structure prediction result of the first complex, wherein the first complex is a pMHC complex, and the first graph neural network is used to perform flexible docking of the antigen peptide sequence and the MHC structure.

[0189] The second prediction module 1003 is used to input the TCR structure and the prediction result of the first structure into the second graph neural network to obtain the second structure prediction result of the second complex, wherein the second complex is a TCR-pMHC complex, and the second graph neural network is used to rigidly connect the TCR structure and the prediction result of the first structure.

[0190] Optionally, the first graph neural network is a hierarchical graph neural network, which includes a hierarchical encoder; the first prediction module 1002 is used for:

[0191] The amino acid residues in the MHC structure are encoded by the first-level encoder to obtain the first node features of the first type of node, and the first type of node represents the α carbon atom of the amino acid residue in the MHC structure.

[0192] The atoms in the docking interface are encoded by the second-level encoder to obtain the second node features of the second type of node. The docking interface is the binding interface between the antigen peptide sequence and the MHC structure. The second type of node characterizes the atoms in the docking interface.

[0193] The amino acid residues in the docking interface are encoded by a third-level encoder to obtain the third node features of the third type of node, wherein the third type of node represents the α carbon atom of the amino acid residue in the docking interface.

[0194] Based on the first node features, the second node features, and the third node features, the first structure prediction result of the first complex is determined.

[0195] Optionally, the first prediction module 1002 is used for:

[0196] The third node feature of the third type of node is obtained by fusing the first node feature corresponding to the α carbon atom and the second node feature corresponding to the atom connected to the α carbon atom through the third-level encoder.

[0197] Optionally, the first prediction module 1002 is used for:

[0198] During the i-th iteration, the initial structure prediction result of the first complex is determined based on the first node feature, the second node feature and the third node feature, where i is a positive integer;

[0199] The initial structure prediction result is adjusted to obtain the i-th round structure prediction result;

[0200] If the iteration stopping condition is met in the i-th iteration, the i-th iteration structure prediction result is determined as the first structure prediction result;

[0201] If the iteration stopping condition is not met in the i-th iteration, the first node feature, the second node feature, and the third node feature are updated based on the i-th iteration structure prediction result.

[0202] Optionally, the first prediction module 1002 is used for:

[0203] Based on the first node features and the third node features, the inter-residue forces between the amino acid residues are determined;

[0204] Based on the characteristics of the second node, the interatomic forces between atoms are determined;

[0205] The coordinates of the amino acid residues in the initial structure prediction result are adjusted based on the inter-residue forces, and the coordinates of the atoms in the amino acid residues are adjusted based on the inter-atomic forces to obtain the i-th round of structure prediction result.

[0206] Optionally, the second graph neural network is a message passing network, and the second prediction module 1003 is used for:

[0207] The step of inputting the TCR structure and the prediction result of the first structure into a second graph neural network to obtain the prediction result of the second structure of the second complex includes:

[0208] Mapping is performed on the TCR structure and the predicted results of the first structure to obtain a TCR structure map and a first complex structure map;

[0209] Input the TCR structure diagram and the first complex structure diagram into the message passing network to obtain the key point coordinates of the docking key points in the TCR structure diagram and the first complex structure diagram;

[0210] Based on the key point coordinates, the TCR structure diagram and the first complex structure diagram are connected to obtain the second structure prediction result of the second complex.

[0211] Optionally, the messaging network includes a message generation layer, a message propagation layer, and an attention layer; the second prediction module 1003 is used for:

[0212] The TCR structure diagram and the first complex structure diagram are input into the message generation layer to obtain the intra-graph interaction features and inter-graph interaction features corresponding to the nodes. The intra-graph interaction features characterize the interaction information between nodes inside the TCR structure diagram or the first complex structure diagram, and the inter-graph interaction features characterize the interaction information between nodes between the TCR structure diagram and the first complex structure diagram.

[0213] The intra-graph interaction features and inter-graph interaction features are input into the message propagation layer to obtain the updated node coordinates and node features corresponding to the node.

[0214] The updated node coordinates and node features are input into the attention layer to obtain the key point coordinates of the docking key points.

[0215] Optionally, the device further includes a training module for:

[0216] The first graph neural network is pre-trained based on the first dataset;

[0217] The second graph neural network is pre-trained based on the second dataset;

[0218] The first and second graph neural networks were fine-tuned based on the pre-trained third dataset;

[0219] The third dataset contains samples of TCR-pMHC complexes, and the data volume of the first and second datasets is greater than that of the third dataset.

[0220] Optional, training module, used for:

[0221] The sample TCR-pMHC complex in the third dataset was split to obtain sample TCR, sample antigen peptide, sample MHC and sample pMHC;

[0222] The first graph neural network is fine-tuned based on the sample antigen peptide, the sample MHC, and the sample pMHC.

[0223] The second graph neural network is fine-tuned based on the sample pMHC, the sample TCR, and the sample TCR-pMHC complex.

[0224] Optionally, before splitting the sample TCR-pMHC complex in the third dataset to obtain sample TCR, sample antigen peptide, sample MHC, and sample pMHC, the training module is used for:

[0225] Based on the screening criteria, the sample TCR-pMHC complex was selected from the third dataset; the screening criteria were that the TCR structure has two chains, the antigen chain contains a short peptide, and the MHC structure has at least one chain.

[0226] Optionally, after splitting the sample TCR-pMHC complex in the third dataset to obtain sample TCR, sample antigen peptide, sample MHC, and sample pMHC, the training module is used for:

[0227] The separated sample TCR, sample antigen peptide, sample MHC and sample pMHC are randomly translated and randomly rotated.

[0228] Optional, training module, used for:

[0229] The sample antigen peptide and the sample MHC are input into the first graph neural network to obtain the predicted coordinates of atoms in the predicted pMHC.

[0230] The first loss is calculated based on the true coordinates and predicted coordinates of atoms in the sample pMHC.

[0231] The first graph neural network is fine-tuned based on the first loss.

[0232] Optional, training module, used for:

[0233] The sample pMHC and the sample TCR are input into the second graph neural network to obtain the predicted coordinates of the nodes in the predicted TCR-pMHC, as well as the predicted coordinates of the key points; the key point coordinates are used to characterize the coordinates of the protein binding pocket.

[0234] The second loss is calculated based on the true coordinates and predicted coordinates of the nodes in the sample TCR-pMHC complex.

[0235] The third loss is calculated based on the true values ​​of the key point coordinates in the sample TCR-pMHC complex and the predicted values ​​of the key point coordinates;

[0236] The second graph neural network is fine-tuned based on the second loss and the third loss.

[0237] See Figure 11 , Figure 11 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application.

[0238] The computer device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including random access memory 1102 and read-only memory 1103, and a system bus 1105 connecting the system memory 1104 and the CPU 1101. The computer device 1100 also includes a basic input / output system (I / O system) 1106 to facilitate information transfer between various components within the computer, and a mass storage device 1107 for storing the operating system 1113, application programs 1114, and other program modules 1115.

[0239] The basic input / output system 1106 includes a display 1108 for displaying information and an input device 1109 for user input, such as a mouse or keyboard. Both the display 1108 and the input device 1109 are connected to the central processing unit 1101 via an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include the input / output controller 1110 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1110 also provides output to a display screen, printer, or other types of output devices.

[0240] The mass storage device 1107 is connected to the central processing unit 1101 via a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable media provide non-volatile storage for the computer device 1100. That is, the mass storage device 1107 may include computer-readable media (not shown) such as a hard disk or drive.

[0241] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage media are not limited to the above-mentioned types. The system memory 1104 and mass storage device 1107 described above can be collectively referred to as memory.

[0242] The memory stores one or more programs, which are configured to be executed by one or more central processing units 1101. The one or more programs contain instructions for implementing the methods described above, and the central processing unit 1101 executes the one or more programs to implement the methods provided in the various method embodiments described above.

[0243] According to various embodiments of this application, the computer device 1100 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1100 can be connected to the network 1112 via the network interface unit 1111 connected to the system bus 1105, or the network interface unit 1111 can be used to connect to other types of networks or remote computer systems (not shown).

[0244] The memory further includes one or more programs stored in the memory, and the one or more programs include steps performed by a computer device in the methods provided in the embodiments of this application.

[0245] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method described in any of the above embodiments.

[0246] Optionally, the computer-readable storage medium may include ROM, RAM, solid-state drives (SSDs), or optical discs, etc. The RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0247] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods described in the above embodiments.

[0248] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for predicting the structure of a complex, characterized in that, The method includes: Obtain the antigen peptide sequence, the major histocompatibility complex (MHC) structure, and the T cell receptor (TCR) structure; The antigen peptide sequence and the MHC structure are input into a first graph neural network to obtain a first structure prediction result of the first complex, which is a pMHC complex. The first graph neural network is used to perform flexible docking of the antigen peptide sequence and the MHC structure. The first graph neural network is a hierarchical graph neural network. Different hierarchical encoders in the hierarchical graph neural network are used to encode the MHC structure at the amino acid residue level, encode the docking interface at the atomic level, and encode the docking interface at the amino acid residue level. The docking interface is the binding interface between the antigen peptide sequence and the MHC structure. The TCR structure and the prediction results of the first structure are input into a second graph neural network to obtain the prediction results of the second structure of the second complex, which is a TCR-pMHC complex. The second graph neural network is used to rigidly connect the TCR structure and the prediction results of the first structure. The second graph neural network is a message passing network, which is used to predict the key point coordinates of the docking key points on the TCR structure and the pMHC complex structure. The second structure prediction results are obtained by overlapping the corresponding key point coordinates on the TCR structure and the pMHC complex structure.

2. The method according to claim 1, characterized in that, The step of inputting the antigen peptide sequence and the MHC structure into a first graph neural network to obtain the first structure prediction result of the first complex includes: The amino acid residues in the MHC structure are encoded by the first-level encoder to obtain the first node features of the first type of node, and the first type of node represents the α carbon atom of the amino acid residue in the MHC structure. The atoms in the docking interface are encoded by the second-level encoder to obtain the second node features of the second type of nodes, which represent the atoms in the docking interface. The amino acid residues in the docking interface are encoded by a third-level encoder to obtain the third node features of the third type of node, wherein the third type of node represents the α carbon atom of the amino acid residue in the docking interface. Based on the first node features, the second node features, and the third node features, the first structure prediction result of the first complex is determined.

3. The method according to claim 2, characterized in that, The third-level encoder encodes the amino acid residues in the docking interface to obtain the third-node features of the third type of node, including: The third node feature of the third type of node is obtained by fusing the first node feature corresponding to the α carbon atom and the second node feature corresponding to the atom connected to the α carbon atom through the third-level encoder.

4. The method according to claim 2, characterized in that, The step of determining the first structure prediction result of the first complex based on the first node features, the second node features, and the third node features includes: During the i-th iteration, the initial structure prediction result of the first complex is determined based on the first node feature, the second node feature and the third node feature, where i is a positive integer; The initial structure prediction result is adjusted to obtain the i-th round structure prediction result; If the iteration stopping condition is met in the i-th iteration, the i-th iteration structure prediction result is determined as the first structure prediction result; If the iteration stopping condition is not met in the i-th iteration, the first node feature, the second node feature, and the third node feature are updated based on the i-th iteration structure prediction result.

5. The method according to claim 4, characterized in that, The adjustment of the initial structure prediction result to obtain the i-th round structure prediction result includes: Based on the first node features and the third node features, the inter-residue forces between the amino acid residues are determined; Based on the characteristics of the second node, the interatomic forces between atoms are determined; The coordinates of the amino acid residues in the initial structure prediction result are adjusted based on the inter-residue forces, and the coordinates of the atoms in the amino acid residues are adjusted based on the inter-atomic forces to obtain the i-th round of structure prediction result.

6. The method according to claim 1, characterized in that, The step of inputting the TCR structure and the prediction result of the first structure into a second graph neural network to obtain the prediction result of the second structure of the second complex includes: Mapping is performed on the TCR structure and the predicted results of the first structure to obtain a TCR structure map and a first complex structure map; Input the TCR structure diagram and the first complex structure diagram into the message passing network to obtain the key point coordinates of the docking key points in the TCR structure diagram and the first complex structure diagram; Based on the key point coordinates, the TCR structure diagram and the first complex structure diagram are connected to obtain the second structure prediction result of the second complex.

7. The method according to claim 6, characterized in that, The messaging network includes a message generation layer, a message propagation layer, and an attention layer; The step of inputting the TCR structure diagram and the first complex structure diagram into the messaging network to obtain the key point coordinates of the docking key points in the TCR structure diagram and the first complex structure diagram includes: The TCR structure diagram and the first complex structure diagram are input into the message generation layer to obtain the intra-graph interaction features and inter-graph interaction features corresponding to the nodes. The intra-graph interaction features characterize the interaction information between nodes inside the TCR structure diagram or the first complex structure diagram, and the inter-graph interaction features characterize the interaction information between nodes between the TCR structure diagram and the first complex structure diagram. The intra-graph interaction features and inter-graph interaction features are input into the message propagation layer to obtain the updated node coordinates and node features corresponding to the node. The updated node coordinates and node features are input into the attention layer to obtain the key point coordinates of the docking key points.

8. The method according to claim 1, characterized in that, The method further includes: The first graph neural network is pre-trained based on the first dataset; The second graph neural network is pre-trained based on the second dataset; The first and second graph neural networks were fine-tuned based on the pre-trained third dataset; The third dataset contains samples of TCR-pMHC complexes, and the data volume of the first and second datasets is greater than that of the third dataset.

9. The method according to claim 8, characterized in that, The first and second graph neural networks, obtained by fine-tuning pre-training based on the third dataset, include: The sample TCR-pMHC complex in the third dataset was split to obtain sample TCR, sample antigen peptide, sample MHC and sample pMHC; The first graph neural network is fine-tuned based on the sample antigen peptide, the sample MHC, and the sample pMHC. The second graph neural network is fine-tuned based on the sample pMHC, the sample TCR, and the sample TCR-pMHC complex.

10. The method according to claim 9, characterized in that, Before splitting the sample TCR-pMHC complex in the third dataset to obtain sample TCR, sample antigen peptide, sample MHC, and sample pMHC, the method includes: Based on the screening criteria, the sample TCR-pMHC complex was selected from the third dataset; the screening criteria were that the TCR structure has two chains, the antigen chain contains a short peptide, and the MHC structure has at least one chain.

11. The method according to claim 9, characterized in that, After splitting the sample TCR-pMHC complex in the third dataset to obtain sample TCR, sample antigen peptide, sample MHC, and sample pMHC, the method further includes: The separated sample TCR, sample antigen peptide, sample MHC and sample pMHC are randomly translated and randomly rotated.

12. The method according to claim 9, characterized in that, The fine-tuning of the first graph neural network based on the sample antigen peptide, the sample MHC, and the sample pMHC includes: The sample antigen peptide and the sample MHC are input into the first graph neural network to obtain the predicted coordinates of atoms in the predicted pMHC. The first loss is calculated based on the true coordinates and predicted coordinates of atoms in the sample pMHC. The first graph neural network is fine-tuned based on the first loss.

13. The method according to claim 9, characterized in that, The fine-tuning of the second graph neural network based on the sample pMHC, the sample TCR, and the sample TCR-pMHC complex includes: The sample pMHC and the sample TCR are input into the second graph neural network to obtain the predicted coordinates of the nodes in the predicted TCR-pMHC, as well as the predicted coordinates of the key points; the key point coordinates are used to characterize the coordinates of the protein binding pocket. The second loss is calculated based on the true coordinates and predicted coordinates of the nodes in the sample TCR-pMHC complex. The third loss is calculated based on the true values ​​of the key point coordinates in the sample TCR-pMHC complex and the predicted values ​​of the key point coordinates; The second graph neural network is fine-tuned based on the second loss and the third loss.

14. A complex structure prediction device, characterized in that, The device includes: The acquisition module is used to acquire antigen peptide sequences, major histocompatibility complex (MHC) structures, and T cell receptor (TCR) structures. A first prediction module is used to input the antigen peptide sequence and the MHC structure into a first graph neural network to obtain a first structure prediction result of a first complex, wherein the first complex is a pMHC complex. The first graph neural network is used to perform flexible docking of the antigen peptide sequence and the MHC structure. The first graph neural network is a hierarchical graph neural network. Different hierarchical encoders in the hierarchical graph neural network are used to encode the MHC structure at the amino acid residue level, encode the docking interface at the atomic level, and encode the docking interface at the amino acid residue level. The docking interface is the binding interface between the antigen peptide sequence and the MHC structure. The second prediction module is used to input the prediction results of the TCR structure and the first structure into a second graph neural network to obtain the second structure prediction result of the second complex, wherein the second complex is a TCR-pMHC complex. The second graph neural network is used to rigidly connect the prediction results of the TCR structure and the first structure. The second graph neural network is a message passing network, which is used to predict the key point coordinates of the docking key points on the TCR structure and the pMHC complex structure. The second structure prediction result is obtained by overlapping the corresponding key point coordinates on the TCR structure and the pMHC complex structure.

15. A computer device, characterized in that, The computer device includes a processor and a memory; the memory stores at least one instruction, which is executed by the processor to implement the complex structure prediction method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement the complex structure prediction method as described in any one of claims 1 to 13.

17. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; the processor of the terminal reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal to perform the complex structure prediction method as described in any one of claims 1 to 13.