Method and apparatus for processing mhc class molecule and polypeptide affinity prediction data
By combining a protein pre-training model and a Uni-Fold fine-tuning model, a multimodal feature tensor is generated. An affinity prediction model is then used to predict the binding state of MHC molecules and peptides and optimize their structures. This solves the problem of high research complexity in existing technologies and achieves efficient affinity prediction and structure optimization.
Patent Information
- Application Number
- CN202310473498.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2043-04-27
AI Technical Summary
Existing technologies are not efficient and convenient for studying the affinity between MHC molecules and peptides, and conventional experimental methods are complex and cumbersome to operate.
A protein pre-trained model and a Uni-Fold fine-tuning model were used to encode amino acid sequence features and predict three-dimensional structures. The output tensors of the Uni-Fold fine-tuning model were extracted by a specified module to generate multimodal feature tensors. The affinity prediction model was then used to predict binding states and optimize structures.
This reduces the operational complexity of studying the affinity of MHC molecules for peptides, provides an efficient prediction scheme beyond conventional experimental methods, and can optimize three-dimensional structures.
Smart Images

Figure CN116386761B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a method and device for processing MHC class molecule and polypeptide affinity prediction data. BACKGROUND
[0002] Major histocompatibility complex (MHC) refers to a group of tightly linked genes existing on a certain chromosome of a vertebrate, and human leukocyte antigen (HLA) is an expression product of human MHC, which is closely related to specific immune response.
[0003] MHC class I molecule is a heterodimeric molecule formed by a heavy chain (α chain) encoded by MHC class I gene and a light chain (β2 microglobulin) encoded by a non-MHC class I gene through disulfide bond. MHC class II molecule is composed of two polypeptide chains, which are called α chain and β chain respectively; unlike class I molecule, both chains of class II molecule are encoded by HLA gene. In T cell-mediated immune response, the binding of MHC class I and MHC class II molecules to polypeptide is an essential step: MHC class I molecule mainly presents endogenous antigen to CD8+T cell, and MHC class II molecule mainly presents exogenous antigen to CD4+T cell, and the affinity of MHC class I and MHC class II molecules to polypeptide can affect the effect and intensity of T cell immune response. Therefore, it is of great significance to study the affinity of MHC class I and MHC class II molecules to polypeptide.
[0004] The study of the affinity of MHC class I and MHC class II molecules to polypeptide includes two parts: 1) analysis of whether MHC class I and MHC class II molecules correctly bind to polypeptide; 2) analysis of the binding strength of MHC class I and MHC class II molecules to polypeptide. At present, the commonly used experimental methods include crystallography, nuclear magnetic resonance, surface plasmon resonance, etc., although these methods can provide high-resolution structural information, but the complex sample preparation and structure analysis process makes it difficult for these conventional experimental methods to better complete the above analysis tasks. SUMMARY
[0005] The present application aims at the defects of the prior art, and provides a method and device for processing MHC class molecule and polypeptide affinity prediction data, an electronic equipment and a computer readable storage medium. A protein pre-training model is used to perform amino acid sequence feature coding on a molecular sequence including an MHC class molecule and a polypeptide to obtain a corresponding first feature tensor. A Uni-Fold fine-tuning model is used to perform three-dimensional structure prediction on the molecular sequence to generate a corresponding three-dimensional structure. In the prediction processing, a specified module output tensor of the Uni-Fold fine-tuning model is extracted to form a corresponding second feature tensor. The hydrogen bond, salt bridge, van der Waals force, charge feature, relative solvent accessible surface area and secondary structure feature of the MHC class molecule and the polypeptide in the three-dimensional structure are extracted to generate a corresponding third feature tensor. The first, second and third feature tensors are fused based on a tensor splicing manner. An affinity prediction model is used to predict the binding state of the MHC class molecule and the polypeptide in the molecular sequence according to the fused features. The three-dimensional structure is optimized according to the prediction result. Through the prediction method based on the artificial intelligence algorithm model, on the one hand, a technical solution is added to the conventional experimental method to study the affinity of MHC I class and MHC II class molecules and polypeptides, thereby reducing the complexity of the operation. On the other hand, the three-dimensional structure can be optimized based on the prediction result.
[0006] To achieve the above object, the first aspect of the embodiment of the present application provides a method for processing MHC class molecule and polypeptide affinity prediction data, which comprises the following steps.
[0007] receiving a first molecular sequence; the first molecular sequence includes a molecular sequence of an MHC class molecule and a polypeptide; the MHC class molecule includes an MHC I class molecule and an MHC II class molecule;
[0008] performing amino acid sequence feature coding processing on the first molecular sequence by using a preset protein pre-training model to generate a corresponding first feature tensor; performing three-dimensional structure prediction processing on the first molecular sequence by using a preset Uni-Fold fine-tuning model to generate a corresponding first three-dimensional structure; extracting a specified module output tensor of the Uni-Fold fine-tuning model to form a corresponding second feature tensor in the three-dimensional structure prediction processing; extracting the hydrogen bond, salt bridge, van der Waals force, charge feature, relative solvent accessible surface area and secondary structure feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding third feature tensor; and performing feature tensor splicing on the first, second and third feature tensors to generate a corresponding first fusion tensor;
[0009] The binding state of the MHC class molecule and the polypeptide is predicted and processed according to the first fusion tensor to generate a corresponding first prediction vector using a preset affinity prediction model; and the first three-dimensional structure is optimized and processed according to the first prediction vector; the first prediction vector includes a first binding probability and a first non-binding probability; the first binding probability is the predicted probability of the MHC class molecule binding to the polypeptide; and the first non-binding probability is the predicted probability of the MHC class molecule not binding to the polypeptide.
[0010] Preferably, the protein pre-training model includes an ESM2 model, an ESM-1b model, an ESM-1v model, an ESM-IF1 model, and a protBERT model.
[0011] The Uni-Fold fine-tuning model is a fine-tuning model based on a Uni-Fold model; the Uni-Fold fine-tuning model includes an Evoformer network and a Structure module, and an output tensor of the Evoformer network is an input tensor of the Structure module.
[0012] The affinity prediction model includes a Transformer model and a prediction neural network.
[0013] Preferably, in the three-dimensional structure prediction processing, a specified module output tensor of the Uni-Fold fine-tuning model is extracted to form a corresponding second feature tensor, specifically including:
[0014] In the three-dimensional structure prediction processing, output tensors of the Evoformer network and the Structure module of the Uni-Fold fine-tuning model are extracted as corresponding first and second extraction tensors; and the first and second extraction tensors form a corresponding second feature tensor.
[0015] Preferably, the hydrogen bond, salt bridge, van der Waals force, charge feature, relative solvent accessible surface area, and secondary structure feature of the MHC class molecule and the polypeptide in the first three-dimensional structure are extracted to generate a corresponding third feature tensor, specifically including:
[0016] calculating the hydrogen bond force between the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding first sub-feature tensor; calculating the salt bridge force between the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding second sub-feature tensor; calculating the van der Waals force between the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding third sub-feature tensor; extracting the charge feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to obtain a corresponding fourth sub-feature tensor; extracting the relative solvent accessible surface area feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to obtain a corresponding fifth sub-feature tensor; extracting the secondary structure feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to obtain a corresponding sixth sub-feature tensor; and generating the third feature tensor by the first, second, third, fourth, fifth, and sixth sub-feature tensors.
[0017] Preferably, the using a preset affinity prediction model to predict the binding state of the MHC class molecule and the polypeptide according to the first fusion tensor to generate a corresponding first prediction vector, specifically comprising:
[0018] inputting the first fusion tensor into the affinity prediction model as a corresponding current model input tensor; and performing feature encoding processing on the affinity feature of the MHC class molecule and the polypeptide by the Transformer model according to the current model input tensor to generate a corresponding first encoding tensor; and predicting the binding probability and the unbinding probability of the MHC class molecule and the polypeptide by the prediction neural network according to the first encoding tensor to obtain the first binding probability and the first unbinding probability, which constitute the first prediction vector.
[0019] Preferably, the optimizing the first three-dimensional structure according to the first prediction vector, specifically comprising:
[0020] Step 61, taking the first three-dimensional structure as a corresponding current three-dimensional structure; and performing credibility evaluation according to the first binding probability and the first unbinding probability of the first prediction vector to generate a corresponding current credibility;
[0021] Step 62, identifying whether the current credibility exceeds a preset credibility threshold; if not, turning to step 63; if yes, turning to step 64;
[0022] Step 63, molecular dynamics simulation is performed on the current three-dimensional structure to obtain a corresponding simulated three-dimensional structure, and the simulated three-dimensional structure is taken as a new current three-dimensional structure; a binding state of the MHC class molecule and the polypeptide on the current three-dimensional structure is predicted to obtain a corresponding second prediction vector; and a credibility evaluation is performed according to a second binding probability and a second non-binding probability of the second prediction vector to generate a new current credibility, and the step 62 is turned to; the second prediction vector includes the second binding probability and the second non-binding probability;
[0023] Step 64, the current three-dimensional structure is taken as an optimization processing output.
[0024] The second aspect of the embodiment of the application provides a device for implementing the processing method of the MHC class molecule and polypeptide affinity prediction data in the first aspect, and the device comprises a receiving module, a first data processing module and a second data processing module.
[0025] The receiving module is used for receiving a first molecular sequence; the first molecular sequence comprises a molecular sequence of an MHC class molecule and a polypeptide; the MHC class molecule comprises an MHC I class molecule and an MHC II class molecule;
[0026] The first data processing module is used for performing amino acid sequence feature coding processing on the first molecular sequence by using a preset protein pre-training model to generate a corresponding first feature tensor; performing three-dimensional structure prediction processing on the first molecular sequence by using a preset Uni-Fold fine-tuning model to generate a corresponding first three-dimensional structure; extracting a tensor output by a specified module of the Uni-Fold fine-tuning model in the three-dimensional structure prediction processing process to form a corresponding second feature tensor; extracting hydrogen bond, salt bridge, van der Waals force, charge feature, relative solvent accessible surface area and secondary structure feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding third feature tensor; and performing feature tensor splicing on the first, second and third feature tensors to generate a corresponding first fusion tensor;
[0027] The second data processing module is used for performing prediction processing on a binding state of the MHC class molecule and the polypeptide according to the first fusion tensor by using a preset affinity prediction model to generate a corresponding first prediction vector; and performing optimization processing on the first three-dimensional structure according to the first prediction vector; the first prediction vector comprises a first binding probability and a first non-binding probability; the first binding probability is a prediction probability of the MHC class molecule and the polypeptide binding; and the first non-binding probability is a prediction probability of the MHC class molecule and the polypeptide not binding.
[0028] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0029] The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above;
[0030] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0031] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.
[0032] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for processing MHC-like molecules and peptide affinity prediction data. The method uses a protein pre-trained model to encode the amino acid sequence features of molecular sequences containing MHC-like molecules and peptides to obtain corresponding first feature tensors. A Uni-Fold fine-tuning model is used to predict the three-dimensional structure of the molecular sequences to generate corresponding three-dimensional structures. During the prediction process, the output tensors of designated modules of the Uni-Fold fine-tuning model are extracted to form corresponding second feature tensors. Hydrogen bonds, salt bridges, van der Waals forces, charge characteristics, relative solvent contact surface area, and secondary structure characteristics of MHC-like molecules and peptides in the three-dimensional structure are extracted to generate corresponding third feature tensors. Multimodal feature fusion is performed on the first, second, and third feature tensors based on tensor splicing. An affinity prediction model is used to predict the binding state of MHC-like molecules and peptides in the molecular sequences based on the fused features. The three-dimensional structure is then optimized based on the prediction results. The prediction method based on artificial intelligence algorithm model of this invention provides an additional technical solution to study the affinity of MHCII and MHCII molecules for peptides, in addition to conventional experimental methods, thus reducing the operational complexity of such studies. Furthermore, it can also optimize the three-dimensional structure based on the prediction results. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of a method for processing MHC molecules and peptide affinity prediction data provided in Embodiment 1 of the present invention;
[0034] Figure 2 This is a module structure diagram of a data processing device for predicting the affinity of MHC molecules to peptides, provided in Embodiment 2 of the present invention.
[0035] Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0037] The embodiment one of the present application provides a processing method for MHC class molecule and polypeptide affinity prediction data, which comprises the following steps: Figure 1 The processing method for MHC class molecule and polypeptide affinity prediction data provided by the embodiment one of the present application is shown in a schematic diagram, and the method mainly comprises the following steps:
[0038] Step 1, receiving a first molecular sequence;
[0039] The first molecular sequence comprises the molecular sequence of the MHC class molecule and the polypeptide; the MHC class molecule comprises the MHC I class molecule and the MHC II class molecule;
[0040] Step 2, using a preset protein pre-training model to perform amino acid sequence feature coding processing on the first molecular sequence to generate a corresponding first feature tensor; using a preset Uni-Fold fine-tuning model to perform three-dimensional structure prediction processing on the first molecular sequence to generate a corresponding first three-dimensional structure; extracting the output tensor of the specified module of the Uni-Fold fine-tuning model in the three-dimensional structure prediction processing process to form a corresponding second feature tensor; extracting the hydrogen bond, salt bridge, van der Waals force, charge feature, relative solvent accessible surface area and secondary structure feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding third feature tensor; and performing feature tensor splicing on the first, second and third feature tensors to generate a corresponding first fusion tensor;
[0041] Specifically, step 21 comprises: using a preset protein pre-training model to perform amino acid sequence feature coding processing on the first molecular sequence to generate a corresponding first feature tensor;
[0042] The protein pre-training model comprises an ESM2 model, an ESM-1b model, an ESM-1v model, an ESM-IF1 model and a protBERT model;
[0043] Here, the protein pre-training model used by the embodiment of the application is an artificial intelligence model that has completed model training in advance, and at least includes ESM2 model, ESM-1b model, ESM-1v model, ESM-IF1 model and protBERT model and other model structures; according to the technical content disclosed by the above models, the ESM2 model, the ESM-1b model, the ESM-1v model, the ESM-IF1 model and the protBERT model and other models can perform two-dimensional structure feature coding on the input one-dimensional molecular sequence to obtain the corresponding two-dimensional structure coding, that is, the first feature tensor;
[0044] Step 22, using a preset Uni-Fold fine-tuning model to perform three-dimensional structure prediction processing on the first molecular sequence to generate a corresponding first three-dimensional structure;
[0045] Among them, the Uni-Fold fine-tuning model is a fine-tuning model based on the Uni-Fold model; the Uni-Fold fine-tuning model includes an Evoformer network and a Structure module, and the output tensor of the Evoformer network is the input tensor of the Structure module;
[0046] Here, the Uni-Fold fine-tuning model used by the embodiment of the application is a fine-tuning model based on the Uni-Fold model, and the Uni-Fold model is a full-size protein three-dimensional structure prediction model that reproduces the precision of the AlphaFold2 model. The technical implementation of the Uni-Fold model can be understood by referring to the paper “Uni-Fold: An Open-Source Platform for Developing Protein Folding Models beyond AlphaFold”. After training the Uni-Fold model in a supervised model training manner, the Uni-Fold fine-tuning model can be obtained. The training manner is to combine the known one-dimensional MHC+ polypeptide analysis sequence and MHC+ polypeptide complex three-dimensional structure as training-label data combination, and use the training-label data combination to train the Uni-Fold fine-tuning model in a supervised manner. The Uni-Fold fine-tuning model obtained after training can predict the protein three-dimensional structure based on the input one-dimensional molecular sequence and output the corresponding three-dimensional structure, that is, the first three-dimensional structure;
[0047] Step 23, in the three-dimensional structure prediction processing process, the specified module output tensor of the Uni-Fold fine-tuning model is extracted to form a corresponding second feature tensor;
[0048] Specifically comprising: in the three-dimensional structure prediction processing, the output tensors of the Evoformer network and the Structure module of the Uni-Fold fine-tuning model are extracted as corresponding first and second extracted tensors; and the second feature tensor is composed of the first and second extracted tensors;
[0049] Here, because the Uni-Fold model is a replication of the AlphaFold2 model, the Uni-Fold fine-tuning model is trained from the Uni-Fold model, so the Uni-Fold fine-tuning model is also a replication of the AlphaFold2 model, and the internal logic modules are basically the same as the AlphaFold2 model; The AlphaFold2 model has two important modules: the Evoformer network and the Structure module. In the processing of the Uni-Fold fine-tuning model for the three-dimensional structure prediction processing of the first molecular sequence, the output results of the Evoformer network and the Structure module are retained and the corresponding feature tensor, i.e., the second feature tensor, is composed of the output results of the two modules;
[0050] Step 24, extracting the hydrogen bond, salt bridge, van der Waals force, charge feature, relative solvent accessible surface area and secondary structure feature between the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding third feature tensor;
[0051] Specifically comprising: calculating the hydrogen bond force between the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding first sub-feature tensor; and calculating the salt bridge force between the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding second sub-feature tensor; and calculating the van der Waals force between the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding third sub-feature tensor; and extracting the charge feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to obtain a corresponding fourth sub-feature tensor; and extracting the relative solvent accessible surface area feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to obtain a corresponding fifth sub-feature tensor; and extracting the secondary structure feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to obtain a corresponding sixth sub-feature tensor; and the first, second, third, fourth, fifth and sixth sub-feature tensors are composed of the corresponding third feature tensor;
[0052] Here, the three-dimensional structure output by the Uni-Fold fine-tuning model in the first three-dimensional structure will describe the characteristics of each atom type, bond, bond angle, charge, coordinates, etc. Under the premise of known MHC class molecules and polypeptides, the hydrogen bond force, salt bridge force and van der Waals force between the MHC class molecules and the polypeptides, the charge amount on the MHC class molecules and the polypeptides, the accessible surface area of the MHC class molecules and the polypeptides relative to the solvent, and the secondary structure of the MHC class molecules and the polypeptides can be estimated and analyzed to obtain the corresponding six types of sub-feature tensors: the first to sixth sub-feature tensors; The six types of sub-feature tensors are spliced to obtain the third feature tensor;
[0053] Step 25, the first, second and third feature tensors are spliced to generate a corresponding first fusion tensor.
[0054] Here, the first fusion tensor obtained by splicing is actually the result of multi-modal feature fusion of the first, second and third feature tensors.
[0055] Step 3, using a preset affinity prediction model, the binding state of the MHC class molecules and the polypeptides is predicted and processed according to the first fusion tensor to generate a corresponding first prediction vector; and the first three-dimensional structure is optimized according to the first prediction vector;
[0056] The first prediction vector includes a first binding probability and a first unbinding probability; the first binding probability is the predicted probability of the MHC class molecules and the polypeptides binding; and the first unbinding probability is the predicted probability of the MHC class molecules and the polypeptides not binding;
[0057] Specifically, it includes: step 31, using a preset affinity prediction model, the binding state of the MHC class molecules and the polypeptides is predicted and processed according to the first fusion tensor to generate a corresponding first prediction vector;
[0058] The affinity prediction model includes a Transformer model and a prediction neural network.
[0059] Specifically, the first fusion tensor is input into the affinity prediction model as a corresponding current model input tensor; and the Transformer model is used to perform feature encoding processing on the affinity features of the MHC class molecules and the polypeptides according to the current model input tensor to generate a corresponding first encoding tensor; and the prediction neural network is used to predict the binding probability and the unbinding probability of the MHC class molecules and the polypeptides according to the first encoding tensor to obtain the corresponding first binding probability and the first unbinding probability, which constitute the corresponding first prediction vector;
[0060] Here, the Transformer model used by the embodiments of the present application is a Transformer model based on multi-modal fusion features for feature encoding; before using the affinity prediction model, it needs to be trained in a supervised model training manner, and the training manner is to train the affinity prediction model using a first original data set composed of data of MHC class I molecules and polypeptide affinity and a second original data set composed of data of MHC class II molecules and polypeptide affinity respectively: 1) obtain the data of MHC class I molecules and polypeptide affinity from the public data set NetMHCpan to constitute the first original data set; then convert the data format and set of the first original data set according to the literature "Reliable prediction of T-cell epitopes using neural networks with novel sequence representations" and "Improved methods for predicting peptide binding affinity to MHC class II molecules"; then divide the first original data set into a first training set and a first test set in a ratio of 8:2; then perform data deduplication processing on the first training set and the first test set; then train the affinity prediction model based on the obtained first training set and first test set; 2) obtain the data of MHC class II molecules and polypeptide affinity from the public data set NetMHCIIPan to constitute the second original data set; then convert the data format and set of the second original data set according to the literature "Reliable prediction of T-cell epitopes using neural networks with novel sequence representations" and "Improved methods for predicting peptide binding affinity to MHC class II molecules"; then divide the second original data set into a second training set and a second test set in a ratio of 8:2; then perform data deduplication processing on the second training set and the second test set; then train the affinity prediction model based on the obtained second training set and second test set;
[0061] In addition, the prediction neural network is a binary classification prediction neural network, which is used for generating two corresponding prediction probabilities (a first binding probability and a first non-binding probability) according to the first encoded tensor; the first binding probability is a prediction probability of the MHC class molecule binding with the polypeptide in the first molecular sequence, and the first non-binding probability is a prediction probability of the MHC class molecule failing to bind with the polypeptide in the first molecular sequence; it can be known from the foregoing that the correct binding positions of the MHC I and MHC II class molecules with the polypeptide are different in the T cell-mediated immune response: the MHC I class molecule mainly presents endogenous antigens to CD8+ T cells, and the MHC II class molecule mainly presents exogenous antigens to CD4+ T cells, so the binding states of different MHC class molecules with the polypeptide can be binary classified and predicted based on such prior knowledge; in addition, the prediction neural network can be implemented in various ways, and is usually implemented based on a multilayer perception neural network, can also be implemented based on a random forest network, and can also be implemented based on any other neural network capable of realizing binary classification prediction and outputting corresponding prediction probabilities;
[0062] Step 32, optimizing the first three-dimensional structure according to the first prediction vector;
[0063] Specifically, step 321 comprises: taking the first three-dimensional structure as a corresponding current three-dimensional structure; and performing credibility evaluation according to the first binding probability and the first non-binding probability of the first prediction vector to generate a corresponding current credibility;
[0064] Here, the embodiment of the present application adopts a conventional classifier credibility evaluation method to realize the credibility evaluation based on the two classification probabilities, and no further elaboration is made on the implementation steps; it should be noted that the conventional classifier credibility evaluation method can obtain N credibilities corresponding to N classification probabilities based on the total number N (N is an integer greater than 1) of classifications of the classifier, and the current credibility obtained by the embodiment of the present application is only the credibility corresponding to the first binding probability;
[0065] Step 322, identifying whether the current credibility exceeds a preset credibility threshold; if not, turning to step 323; if yes, turning to step 324;
[0066] Here, the credibility threshold is a preset threshold parameter;
[0067] Step 323, performing molecular dynamics simulation on the current three-dimensional structure to obtain a corresponding simulated three-dimensional structure, taking the simulated three-dimensional structure as a new current three-dimensional structure; and performing prediction on the binding state of the MHC class molecule with the polypeptide on the current three-dimensional structure to obtain a corresponding second prediction vector; and performing credibility evaluation according to the second binding probability and the second non-binding probability of the second prediction vector to generate a new current credibility, and turning to step 322; the second prediction vector comprises the second binding probability and the second non-binding probability;
[0068] Here, the embodiment of the application uses conventional means to perform simulation (such as enhanced sampling, enhanced dynamics and extension dynamics) when performing molecular dynamics simulation, and each simulation needs to be started with a corresponding configuration (such as a minimum energy function configuration, a simulation number configuration, etc.), and one or more simulation three-dimensional structures can be obtained based on the configuration during the simulation process; after each simulation three-dimensional structure is obtained, a prediction model similar to the affinity prediction model of the embodiment of the application can be used to input the simulation three-dimensional structure to encode the affinity characteristics of the MHC class molecule and the polypeptide and perform binary classification prediction on the binding state of the MHC class molecule and the polypeptide based on the feature encoding to obtain a prediction vector composed of two prediction probabilities (a second binding probability and a second non-binding probability), i.e. a second prediction vector; in addition, a pre-set linear regression calculation model for calculating the binding strength of the MHC class molecule and the polypeptide can be used to regressively predict the binding strength of the MHC class molecule and the polypeptide on each simulation three-dimensional structure, and after the predicted strength reaches the set strength range, a prediction model similar to the affinity prediction model of the embodiment of the application is used to encode the affinity characteristics of the MHC class molecule and the polypeptide on the input simulation three-dimensional structure and perform binary classification prediction on the binding state of the MHC class molecule and the polypeptide based on the feature encoding to obtain a second prediction vector;
[0069] Step 324, taking the current three-dimensional structure as the output of the optimization process.
[0070] Figure 2 A module structure diagram of an MHC class molecule and polypeptide affinity prediction data processing device provided for the second embodiment of the application is shown in FIG. 8. The device is a terminal device or a server for implementing the method embodiments, or a device capable of enabling the terminal device or the server to implement the method embodiments, such as a device or a chip system of the terminal device or the server. As shown in the figure, the device includes a receiving module 201, a first data processing module 202 and a second data processing module 203. Figure 2
[0071] The receiving module 201 is configured to receive a first molecular sequence; the first molecular sequence includes a molecular sequence of an MHC class molecule and a polypeptide; the MHC class molecule includes an MHC I class molecule and an MHC II class molecule.
[0072] The first data processing module 202 is configured to perform amino acid sequence feature coding processing on the first molecular sequence using a preset protein pre-training model to generate a corresponding first feature tensor; perform three-dimensional structure prediction processing on the first molecular sequence using a preset Uni-Fold fine-tuning model to generate a corresponding first three-dimensional structure; extract a tensor output by a specified module of the Uni-Fold fine-tuning model during the three-dimensional structure prediction processing to form a corresponding second feature tensor; extract hydrogen bond, salt bridge, van der Waals force, charge feature, relative solvent accessible surface area and secondary structure feature of the MHC class molecule and the polypeptide in the first three-dimensional structure to generate a corresponding third feature tensor; and perform feature tensor splicing on the first, second and third feature tensors to generate a corresponding first fusion tensor.
[0073] The second data processing module 203 is configured to perform binding state prediction processing on the MHC class molecule and the polypeptide according to the first fusion tensor using a preset affinity prediction model to generate a corresponding first prediction vector; and perform optimization processing on the first three-dimensional structure according to the first prediction vector; the first prediction vector includes a first binding probability and a first non-binding probability; the first binding probability is a prediction probability of the MHC class molecule and the polypeptide binding; and the first non-binding probability is a prediction probability of the MHC class molecule and the polypeptide not binding.
[0074] The MHC class molecule and polypeptide affinity prediction data processing device provided in the embodiment of the application can perform the method steps in the method embodiment, and has similar implementation principles and technical effects, which will not be described here.
[0075] It should be noted that the division of each module of the above device is only a logical function division, and all or part of the modules can be integrated into one physical entity, or can be physically separated. The modules can all be implemented in the form of software called by a processing element; or all be implemented in the form of hardware; or some modules are implemented in the form of software called by a processing element, and some modules are implemented in the form of hardware. For example, the receiving module can be a separate processing element, or can be integrated in a chip of the above device, and in addition, the receiving module can be stored in the memory of the above device in the form of program code, and the function of the receiving module can be called and executed by a processing element of the above device. The implementation of other modules is similar. In addition, all or part of the modules can be integrated together, or can be independently implemented. The processing element described herein can be an integrated circuit having a signal processing capability. In the implementation process, each step of the above method or each module can be completed by an integrated logic circuit of hardware or an instruction in the form of software in the processing element.
[0076] For example, the above modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of a processing element scheduling code, the processing element can be a general purpose processor, such as a Central Processing Unit (CPU) or other processor that can invoke code. For another example, the modules can be integrated together to implement in the form of a System-on-a-chip (SOC).
[0077] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed by a computer, the computer instructions generate all or part of the processes or functions described in the above method embodiments. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0078] Figure 3 A structural schematic diagram of an electronic device is provided for Embodiment Three of the present application. The electronic device can be a terminal device or a server implementing the method of the above embodiments, or a terminal device or a server connected to the terminal device or the server implementing the method of the above embodiments. As shown in FIG. 3, the electronic device includes a processor 301, a memory 302, a transceiver 303 and an antenna 304. The processor 301, the memory 302, the transceiver 303 and the antenna 304 can be connected to each other through a bus or other suitable connection means. The processor 301 can be configured to implement the method of the above embodiments. The memory 302 can be configured to store the computer instructions to be executed by the processor 301. The transceiver 303 can be configured to transmit and receive signals. The antenna 304 can be configured to transmit and receive signals. Figure 3As shown, the electronic device can include a processor 301 (for example, a CPU), a memory 302, a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiving action of the transceiver 303. The memory 302 can store various instructions for completing various processing functions and implementing the processing steps described in the foregoing method embodiments. Preferably, the electronic device related to the embodiments of the present application further includes a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize the communication connection between elements. The communication port 306 described above is used for the connection communication between the electronic device and other peripherals.
[0079] In Figure 3 The system bus 305 mentioned in the foregoing can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus. The communication interface is used to realize the communication between the database access device and other devices (for example, a client, a read-write library and a read-only library). The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory, for example, at least one disk memory.
[0080] The processor described above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; can also be a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0081] It should be noted that the embodiments of the present application also provide a computer readable storage medium, which stores instructions, when running on a computer, causes the computer to execute the method and processing procedure provided in the above embodiments.
[0082] The embodiments of the present application also provide a chip for running instructions, which is used to execute the processing steps described in the foregoing method embodiments.
[0083] The embodiment of the present application provides a kind of MHC class molecule and the processing method, device, electronic equipment and computer readable storage medium of polypeptide affinity prediction data, using protein pre-training model to the molecular sequence including MHC class molecule and polypeptide is carried out amino acid sequence feature coding to obtain corresponding first feature tensor, and using Uni-Fold fine-tuning model to molecular sequence is carried out three-dimensional structure prediction to obtain corresponding three-dimensional structure, and in the prediction processing process, the specified module output tensor of Uni-Fold fine-tuning model is extracted to form corresponding second feature tensor, and the hydrogen bond, salt bridge, van der waals force, charge feature, relative solvent accessible surface area and secondary structure feature of MHC class molecule and polypeptide in three-dimensional structure are extracted to generate corresponding third feature tensor, and based on tensor splicing mode, first, second and third feature tensors are multi-modal feature fusion;And using affinity prediction model, the binding state of MHC class molecule and polypeptide in molecular sequence is predicted according to fusion feature;And according to the prediction result, the structure optimization of three-dimensional structure is carried out.The prediction mode based on artificial intelligence algorithm model of the present application, on the one hand, a technical solution is added to the conventional experimental method for studying the affinity of MHC I class and MHC II class molecules and polypeptides, which reduces the operation complexity of this kind of research;On the other hand, three-dimensional structure can also be optimized based on the prediction result.
[0084] The skilled person should further realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0085] The steps of the method or algorithm described in connection with the embodiments disclosed herein can be implemented in hardware, software executed by a processor, or a combination of both. The software module can be placed in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0086] The above detailed description of the specific embodiments of the present application has been given to understand the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for processing MHC molecules' affinity prediction data for peptides, characterized in that, The method includes: Receive a first molecular sequence; the first molecular sequence includes molecular sequences of MHC class molecules and peptides; the MHC class molecules include MHC class I molecules and MHC class II molecules; The first molecular sequence is processed using a pre-trained protein model to encode amino acid sequence features, generating a corresponding first feature tensor. A pre-trained Uni-Fold fine-tuning model is used to predict the three-dimensional structure of the first molecular sequence, generating a corresponding first three-dimensional structure. During the three-dimensional structure prediction process, the output tensors of designated modules of the Uni-Fold fine-tuning model are extracted to form a corresponding second feature tensor. Hydrogen bonds, salt bridges, van der Waals forces, charge characteristics, relative solvent-accessible surface area, and secondary structure characteristics of the MHC molecules and peptides in the first three-dimensional structure are extracted to generate a corresponding third feature tensor. The first, second, and third feature tensors are then concatenated to generate a corresponding first fusion tensor. A preset affinity prediction model is used to predict the binding state of the MHC molecule and the peptide based on the first fusion tensor, generating a corresponding first prediction vector; and the first three-dimensional structure is optimized based on the first prediction vector; the first prediction vector includes a first binding probability and a first non-binding probability; the first binding probability is the predicted probability that the MHC molecule and the peptide bind; the first non-binding probability is the predicted probability that the MHC molecule and the peptide do not bind. The protein pre-training models include ESM2, ESM-1b, ESM-1v, ESM-IF1 and protBERT models. The Uni-Fold fine-tuning model is a fine-tuning model based on the Uni-Fold model; the Uni-Fold fine-tuning model includes an Evoformer network and a Structure module, and the output tensor of the Evoformer network is the input tensor of the Structure module; The affinity prediction model includes a Transformer model and a predictive neural network; The step of extracting the output tensor of a specified module of the Uni-Fold fine-tuning model to form the corresponding second feature tensor during the 3D structure prediction process specifically includes: In the process of 3D structure prediction, the output tensors of the Evoformer network and the Structure module of the Uni-Fold fine-tuning model are extracted as the corresponding first and second extracted tensors; and the first and second extracted tensors are used to form the corresponding second feature tensor.
2. The method for processing MHC molecule-peptide affinity prediction data according to claim 1, characterized in that, The step of extracting the hydrogen bonds, salt bridges, van der Waals forces, charge characteristics, relative solvent-accessible surface area, and secondary structure characteristics of the MHC molecules and peptides in the first three-dimensional structure to generate the corresponding third feature tensor specifically includes: The hydrogen bonding forces between the MHC molecules and the polypeptide in the first three-dimensional structure are calculated to generate a first sub-feature tensor; the salt bridge forces between the MHC molecules and the polypeptide in the first three-dimensional structure are calculated to generate a second sub-feature tensor; the van der Waals forces between the MHC molecules and the polypeptide in the first three-dimensional structure are calculated to generate a third sub-feature tensor; the charge characteristics of the MHC molecules and the polypeptide in the first three-dimensional structure are extracted to obtain a fourth sub-feature tensor; the relative solvent-accessible surface area characteristics of the MHC molecules and the polypeptide in the first three-dimensional structure are extracted to obtain a fifth sub-feature tensor; and the secondary structure characteristics of the MHC molecules and the polypeptide in the first three-dimensional structure are extracted to obtain a sixth sub-feature tensor; and the obtained first, second, third, fourth, fifth, and sixth sub-feature tensors constitute the corresponding third feature tensor.
3. The method for processing MHC molecule-peptide affinity prediction data according to claim 1, characterized in that, The step of using a preset affinity prediction model to predict the binding state of the MHC molecule and the peptide based on the first fusion tensor to generate a corresponding first prediction vector specifically includes: The first fusion tensor is input into the affinity prediction model as the corresponding current model input tensor; the Transformer model performs feature encoding processing on the affinity features between the MHC molecule and the peptide based on the current model input tensor to generate the corresponding first encoding tensor; and the prediction neural network predicts the binding probability and non-binding probability of the MHC molecule and the peptide based on the first encoding tensor to obtain the corresponding first binding probability and first non-binding probability, which together form the corresponding first prediction vector.
4. The method for processing MHC molecule-peptide affinity prediction data according to claim 1, characterized in that, The optimization process of the first three-dimensional structure based on the first prediction vector specifically includes: Step 61: Take the first three-dimensional structure as the corresponding current three-dimensional structure; and generate the corresponding current confidence level by performing a confidence assessment based on the first combination probability and the first non-combination probability of the first prediction vector. Step 62: Identify whether the current credibility exceeds a preset credibility threshold; if not, proceed to step 63; if yes, proceed to step 64. Step 63: Perform molecular dynamics simulation on the current three-dimensional structure to obtain the corresponding simulated three-dimensional structure, and use the simulated three-dimensional structure as the new current three-dimensional structure; predict the binding state of the MHC molecule and the polypeptide on the current three-dimensional structure to obtain the corresponding second prediction vector; and perform a confidence evaluation based on the second binding probability and the second non-binding probability of the second prediction vector to generate a new current confidence level, and then proceed to step 62; the second prediction vector includes the second binding probability and the second non-binding probability; Step 64: Output the current three-dimensional structure as an optimization process.
5. An apparatus for processing MHC molecule-peptide affinity prediction data according to any one of claims 1-4, characterized in that, The device includes: a receiving module, a first data processing module, and a second data processing module; The receiving module is used to receive a first molecular sequence; the first molecular sequence includes molecular sequences of MHC class molecules and peptides; the MHC class molecules include MHC class I molecules and MHC class II molecules; The first data processing module is used to perform amino acid sequence feature encoding on the first molecular sequence using a preset protein pre-training model to generate a corresponding first feature tensor; and to perform three-dimensional structure prediction on the first molecular sequence using a preset Uni-Fold fine-tuning model to generate a corresponding first three-dimensional structure; and to extract the output tensor of the specified module of the Uni-Fold fine-tuning model to form a corresponding second feature tensor during the three-dimensional structure prediction process; and to extract the hydrogen bonds, salt bridges, van der Waals forces, charge characteristics, relative solvent contact surface area, and secondary structure characteristics of the MHC molecules and the polypeptide in the first three-dimensional structure to generate a corresponding third feature tensor; and to perform feature tensor splicing on the first, second, and third feature tensors to generate a corresponding first fusion tensor; The second data processing module is used to use a preset affinity prediction model to predict the binding state of the MHC molecule and the polypeptide based on the first fusion tensor to generate a corresponding first prediction vector; and to optimize the first three-dimensional structure based on the first prediction vector; the first prediction vector includes a first binding probability and a first non-binding probability; the first binding probability is the predicted probability that the MHC molecule and the polypeptide bind; the first non-binding probability is the predicted probability that the MHC molecule and the polypeptide do not bind.
6. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-4; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1-4.
Citation Information
Patent Citations
MHC-I epitope affinity prediction method based on deep learning
CN112002374A
Method and device for acquiring affinity prediction model of HLA (human leukocyte antigen) II type molecule and polypeptide
CN114446385A