A deep learning-based prediction method and system for peptide-TCR binding
Through deep learning methods, the correlation and three-dimensional structural characteristics of polypeptides and TCR were extracted through deep learning methods, which solved the problem of failure to effectively capture the binding characteristics of polypeptides and TCR in the prior art, and improved the prediction accuracy.
Patent Information
- Application Number
- CN202510184676.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-02-19
AI Technical Summary
When the prior art predicts that the binding of polypeptides to TCR, it is difficult to effectively capture the binding characteristics of unoccupied sequences, and the spatial structural characteristics of the sequence are not fully considered, resulting in a decrease in prediction accuracy.
A deep learning-based method is adopted, combining the amino acid number dictionary and HelixFoldSingle protein structure prediction model, correlation characteristics and three-dimensional structural characteristics of polypeptides and TCRs are extracted, and predictions are made through linear weighting. Multimodal features are captured using the cross-attention Transformer model and convolutional neural network.
The accuracy of the prediction of binding of peptides to TCR is improved, and the predictive ability of unseen sequences is improved by binding sequence correlation and spatial structural characteristics.
Smart Images

Figure CN120108510B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of neoantigen immunotherapy, and in particular relates to a method and system for predicting peptide-TCR binding based on deep learning. Background Art
[0002] In neoantigen immunotherapy, the interaction between the T cell receptor (TCR) and peptides presented by the major histocompatibility complex (MHC) is the foundation of the immune response. This interaction activates T cells, enabling them to effectively fight pathogens in the body, such as bacteria or tumor cells. Therefore, accurately predicting TCR-peptide binding is crucial for improving the prevention and treatment of various diseases.
[0003] With the remarkable performance of deep learning in fields such as natural language processing and computer vision, various deep learning frameworks have been applied to predict peptide-TCR binding, including convolutional neural networks (CNNs), long short-term memory neural networks (LSTMs), Transformers, and prediction methods that use pre-training as embeddings. While these models demonstrate promising results, their performance degrades significantly when applied to novel sequences not seen in the training data. Therefore, accurately predicting peptide-TCR binding for previously unseen sequences remains challenging.
[0004] So far, most methods for predicting peptide-TCR binding are based on sequence features, and rarely deal with the correlation of the interaction between peptides and TCRs. How to obtain the interaction characteristics between peptides and TCR sequences has become a research difficulty.
[0005] Cross-attention has made significant progress in areas such as text-image matching, video understanding, and machine translation. This cross-attention mechanism, primarily used to process information interactions between different sequences or modalities, enhances the ability to perceive and fuse information across different sequences, and has great potential in multimodal tasks and complex scenarios. Google's Transformer model has achieved outstanding results in natural language processing and machine translation. Its encoding layer, through self-attention, is able to effectively understand the context of the input, capture long-range dependencies, and aggregate information across the entire input sequence.
[0006] However, so far, the prediction of TCR and peptide sequence binding is mainly based on sequence information, and the spatial structural characteristics of the sequence are rarely considered. The amino acid sequence alone cannot determine all the information about TCR and peptide binding. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a technical solution for a method for predicting the binding of peptides to TCRs based on deep learning to solve the above technical problems.
[0008] The first aspect of the present invention discloses a method for predicting peptide-TCR binding based on deep learning, the method comprising:
[0009] Step S1: digitize the sequences of the peptide and TCR according to the digital dictionary of amino acids to extract the coding features of the peptide and TCR sequences;
[0010] Step S2: Obtain the three-dimensional structure of the peptide and TCR sequence according to the HelixFoldSingle protein structure prediction model, and extract the three-dimensional structural features of the peptide and TCR;
[0011] Step S3: linearly weight the correlation characteristics between the polypeptide and the TCR sequence and the three-dimensional structural characteristics of the polypeptide and the TCR sequence to obtain a prediction result of the binding between the polypeptide and the TCR.
[0012] According to the method of the first aspect of the present invention, in step S1, the steps of digitizing the sequences of the polypeptide and TCR according to the digital dictionary of amino acids and extracting correlation features between the polypeptide and TCR sequences include:
[0013] The sequences of the polypeptide and TCR are digitized according to the digital dictionary of amino acids to obtain word units and position codes of the polypeptide sequence of a predefined length and word units and position codes of the TCR sequence; the word units and position codes of the polypeptide sequence and the TCR sequence are combined to obtain a first coding combination and a second coding combination; the first coding combination is input into a first cross-attention Transformer model to obtain a first correlation feature; the second coding combination is input into a second cross-attention Transformer model to obtain a second correlation feature; the first correlation feature and the second correlation feature are concatenated and input into a first feedforward fully connected neural network to obtain the correlation feature of the polypeptide and TCR sequences; the word unit and position code is the sum of the word unit code and the position code.
[0014] According to the method of the first aspect of the present invention, in step S1, the digitizing of the polypeptide and TCR sequences according to the digital dictionary of amino acids to obtain word units and position codes of the polypeptide sequence of a predefined length and word units and position codes of the TCR sequence includes:
[0015] The amino acid sequence is tokenized according to a digital dictionary; in the digital dictionary, "*" represents an ambiguous amino acid in the sequence; the amino acid sequence converted by the digital dictionary is used to form a tokenized sequence of the polypeptide and TCR; the length of the sequence of the polypeptide and TCR is aligned to a predefined length, and if the length of the sequence of the polypeptide and TCR is less than the predefined length, it is padded with 0; the amino acids in the sequence of the polypeptide and TCR aligned to the predefined length are token-encoded and position-encoded to obtain the token and position coding of the polypeptide sequence of the predefined length and the token and position coding of the TCR sequence.
[0016] According to the method of the first aspect of the present invention, in step S1, the first coding combination includes: a first input of the first coding combination is a word unit and position code of a polypeptide sequence, a second input of the first coding combination is a word unit and position code of a TCR sequence, and a third input of the first coding combination is a word unit and position code of a TCR sequence;
[0017] The second coding combination includes: the first input of the second coding combination is the word element and position code of the TCR sequence, the second input of the second coding combination is the word element and position code of the polypeptide sequence, and the third input of the second coding combination is the word element and position code of the polypeptide sequence.
[0018] According to the method of the first aspect of the present invention, in step S2, obtaining the three-dimensional structure of the polypeptide and TCR sequence according to the HelixFoldSingle protein structure prediction model, and extracting the three-dimensional structural features of the polypeptide and TCR sequence include:
[0019] The predicted structures of the polypeptide and TCR are obtained according to the HelixFoldSingle protein structure prediction model; and the three-dimensional structural characteristics of the polypeptide and TCR sequences are obtained according to the three-dimensional coordinates of the predicted structures.
[0020] According to the method of the first aspect of the present invention, in step S2, obtaining the three-dimensional coordinates of the predicted structure includes:
[0021] The coordinates of the 'CA' atom of each residue in the predicted structure were taken as the three-dimensional coordinates of each amino acid.
[0022] According to the method of the first aspect of the present invention, in step S2, obtaining the three-dimensional structural features of the polypeptide and TCR sequence based on the three-dimensional coordinates of the predicted structure includes:
[0023] Obtaining the three-dimensional coordinates of the polypeptide and TCR sequence according to the three-dimensional coordinates of the predicted structure;
[0024] The three-dimensional coordinates of the peptide sequence are input into the first convolutional neural network to obtain the three-dimensional structural features of the peptide;
[0025] The three-dimensional coordinates of the TCR sequence are input into the second convolutional neural network to obtain the three-dimensional structural features of the TCR;
[0026] After splicing the three-dimensional structural features of the polypeptide and the three-dimensional structural features of TCR, the three-dimensional structural features are input into a second feedforward fully connected neural network to obtain the three-dimensional structural features of the polypeptide and TCR sequences.
[0027] A second aspect of the present invention discloses a deep learning-based prediction system for peptide-TCR binding, comprising:
[0028] A first processing module is configured to digitally process the sequences of the peptide and TCR according to a digital dictionary of amino acids to extract coding features of the peptide and TCR sequences;
[0029] The second processing module is configured to obtain the three-dimensional structure of the polypeptide and TCR sequence according to the HelixFoldSingle protein structure prediction model and extract the three-dimensional structural features of the polypeptide and TCR sequence;
[0030] The third processing module is configured to linearly weight the correlation characteristics between the polypeptide and the TCR sequence and the three-dimensional structural characteristics of the polypeptide and the TCR sequence to obtain a prediction result of the binding between the polypeptide and the TCR.
[0031] A third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of any one of the deep learning-based methods for predicting peptide-TCR binding according to the first aspect of the present disclosure.
[0032] A fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods for predicting peptide-TCR binding based on deep learning according to the first aspect of the present disclosure.
[0033] In summary, the proposed approach considers both the correlation characteristics of peptide-TCR interactions and the spatial structural features of sequences. By combining these inter-sequence correlation characteristics with sequence structural features through a multimodal deep learning model, a more comprehensive prediction approach is captured, further improving the accuracy of peptide-TCR sequence binding prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0035] Figure 1 Flowchart of a method for predicting peptide-TCR binding based on deep learning according to an embodiment of the present invention;
[0036] Figure 2 2. A structural diagram of a cross-attention Transformer model according to an embodiment of the present invention;
[0037] Figure 3 A diagram showing the structure of a convolutional neural network according to an embodiment of the present invention;
[0038] Figure 4 2. A structural diagram of a system for predicting peptide-TCR binding based on deep learning according to an embodiment of the present invention;
[0039] Figure 5 FIG. 4 is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0041] The first aspect of the present invention discloses a method for predicting peptide-TCR binding based on deep learning. Figure 1 Flowchart of a method for predicting peptide-TCR binding based on deep learning according to an embodiment of the present invention, as shown in FIG. Figure 1 As shown, the method includes:
[0042] Step S1: digitally process the peptide and TCR sequences according to the digital dictionary of amino acids to extract correlation features between the peptide and TCR sequences, wherein the correlation features are coding features;
[0043] Step S2: obtaining the three-dimensional structures of the polypeptide and TCR sequence based on the HelixFoldSingle protein structure prediction model, and extracting the three-dimensional structural features of the polypeptide and TCR sequence;
[0044] Step S3: linearly weight the correlation characteristics between the polypeptide and the TCR sequence and the three-dimensional structural characteristics of the polypeptide and the TCR sequence to obtain a prediction result of the binding between the polypeptide and the TCR.
[0045] In step S1, the sequences of the peptide and TCR are digitized according to the digital dictionary of amino acids to extract correlation features of the peptide and TCR sequences.
[0046] In some embodiments, in step S1, the digitizing of the sequences of the polypeptide and the TCR according to the digital dictionary of amino acids to extract correlation features of the sequences of the polypeptide and the TCR comprises:
[0047] The sequences of the polypeptide and TCR are digitized according to the digital dictionary of amino acids to obtain word units and position codes of the polypeptide sequence of a predefined length and word units and position codes of the TCR sequence; the word units and position codes of the polypeptide sequence and the TCR sequence are combined to obtain a first coding combination and a second coding combination; the first coding combination is input into a first cross-attention Transformer model to obtain a first correlation feature; the second coding combination is input into a second cross-attention Transformer model to obtain a second correlation feature; the first correlation feature and the second correlation feature are concatenated and input into a first feedforward fully connected neural network to obtain the correlation feature of the polypeptide and TCR sequences; the word unit and position code is the sum of the word unit code and the position code.
[0048] The method of digitizing the sequences of the polypeptide and TCR according to the digital dictionary of amino acids to obtain the word element and position code of the polypeptide sequence of the predefined length and the word element and position code of the TCR sequence includes:
[0049] The amino acid sequence is tokenized according to a digital dictionary; in the digital dictionary, "*" represents an ambiguous amino acid in the sequence; the amino acid sequence converted by the digital dictionary is used to form a tokenized sequence of the polypeptide and TCR; the length of the sequence of the polypeptide and TCR is aligned to a predefined length, and if the length of the sequence of the polypeptide and TCR is less than the predefined length, it is padded with 0; the amino acids in the sequence of the polypeptide and TCR aligned to the predefined length are token-encoded and position-encoded to obtain the token and position coding of the polypeptide sequence of the predefined length and the token and position coding of the TCR sequence.
[0050] In some embodiments, the predefined length is 20.
[0051] The first coding combination includes: a first input of the first coding combination is a word unit and position code of a polypeptide sequence, a second input of the first coding combination is a word unit and position code of a TCR sequence, and a third input of the first coding combination is a word unit and position code of a TCR sequence;
[0052] The second coding combination includes: the first input of the second coding combination is the word element and position code of the TCR sequence, the second input of the second coding combination is the word element and position code of the polypeptide sequence, and the third input of the second coding combination is the word element and position code of the polypeptide sequence.
[0053] Specifically, each cross-attention Transformer model consists of two parts: the cross-attention Transformer encoder and the self-attention Transformer encoder. The cross-attention encoder layer mainly consists of multi-head cross-attention, normalization layer, fully connected layer, and residual layer; the self-attention encoder layer consists of multi-head self-attention, normalization layer, fully connected layer, and residual layer.
[0054] In some embodiments, as Figure 2 As shown in the figure, the Cross-Attention Transformer model consists of a layer of Cross-Attention Transformer encoders and multiple layers of Self-Attention Transformer encoders. The network is mainly composed of multiple layers of encoders stacked together. Each encoder layer consists of two parts: the first sub-layer includes multi-head attention, a fully connected layer, a residual layer, and a normalization layer; the second sub-layer includes two fully connected layers, a residual layer, and a normalization layer.
[0055] The first feed-forward fully connected neural network has a fully connected layer, a Dropout layer, a normalization layer, and uses the LeakyReLU activation function.
[0056] Amino acid word dictionary:
[0057] {'A': 1, 'R': 2, 'N': 3, 'D': 4, 'C': 5, 'Q': 6, 'E': 7, 'G': 8, 'H': 9, 'I': 10, 'L': 11, 'K': 12, 'M': 13, 'F': 14, 'p': 15, 'S': 16, 'T': 17, 'W': 18, 'Y': 19, 'V': 20, '*': 21}
[0058] In step S2, the three-dimensional structural features of the polypeptide and TCR sequence are obtained based on the three-dimensional structure of the HelixFoldSingle protein structure prediction model.
[0059] In some embodiments, in step S2, obtaining the three-dimensional structure of the polypeptide and TCR sequence according to the HelixFoldSingle protein structure prediction model, and extracting the three-dimensional structural features of the polypeptide and TCR sequence include:
[0060] The predicted structures of the polypeptide and TCR are obtained according to the HelixFoldSingle protein structure prediction model; and the three-dimensional structural characteristics of the polypeptide and TCR sequences are obtained according to the three-dimensional coordinates of the predicted structures.
[0061] Obtaining the three-dimensional coordinates of the predicted structure includes:
[0062] The coordinates of the 'CA' atom of each residue in the predicted structure were taken as the three-dimensional coordinates of each amino acid.
[0063] Obtaining the three-dimensional structural features of the polypeptide and TCR sequence based on the three-dimensional coordinates of the predicted structure includes:
[0064] Obtaining the three-dimensional coordinates of the polypeptide and TCR sequence according to the three-dimensional coordinates of the predicted structure;
[0065] The three-dimensional coordinates of the peptide sequence are input into the first convolutional neural network to obtain the three-dimensional structural features of the peptide;
[0066] The three-dimensional coordinates of the TCR sequence are input into the second convolutional neural network to obtain the three-dimensional structural features of the TCR;
[0067] After splicing the three-dimensional structural features of the polypeptide and the three-dimensional structural features of TCR, the three-dimensional structural features are input into a second feedforward fully connected neural network to obtain the three-dimensional structural features of the polypeptide and TCR sequences.
[0068] Specifically, each convolutional neural network consists of multiple convolutional layers, a normalization layer, a dropout layer, and a maximum pooling layer.
[0069] In some embodiments, as Figure 3 As shown, the convolutional neural network consists of multiple layers of 2D convolutional layers, normalization layers, Dropout layers, and maximum pooling layers. The fully connected network consists of Dropout layers, normalization layers, the activation function of the middle layer uses LeakyReLU, and the activation function of the last layer uses Sigmoid.
[0070] The second feedforward fully connected neural network has a Dropout layer and a normalization layer. The activation function of the middle layer is LeakyReLU, and the activation function of the last layer is Sigmoid function.
[0071] The entire network model uses the Adam optimizer and adds a weight decay rate of 1e-5 to obtain the final prediction results of the binding between peptide and TCR sequence.
[0072] The prediction structure is as follows:
[0073]
[0074] In some embodiments, two data sets are used, the first data set is divided into a training set and a validation set according to an 8:2 ratio, and the second data set, in which the peptides and TCR sequences that do not appear in the training set and the validation set are used as the test set.
[0075] Example 1
[0076] Experimental Data: In this experimental evaluation, the TEP-merge dataset from the article TEPCAM: Prediction of T-cell receptor–epitope binding specificity via interpretable deep learning was used as the training and validation sets. This dataset primarily comes from three public databases, including 129,654 pairs of TCR-peptide data screened from VDJdb, McPAS, and IEDB data, with a positive-to-negative ratio of 1:1. It contains 1,523 unique epitopes and 60,342 unique TCRs. The ImmuneCODE dataset was used as the test set. This dataset contains 56,606 pairs of TCR-peptide data, including 28,162 TCR sequences and 131 peptide sequences. Neither the peptides nor the TCR sequences in this dataset appear in the training or validation sets.
[0077] Data processing: For the cross-attention Transformer network, the sequence is processed according to the digital dictionary and padded with 0 to the set length as the network input; for the convolutional neural network model with three-dimensional spatial structure, the three-dimensional coordinates of amino acid atoms generated by HelixFoldSingle are used as the network input.
[0078] It should be understood that the present invention disclosed is not limited only to the specific methods, schemes and materials described, because these are all variable. It should also be understood that the terms used herein are only for the purpose of illustrating specific embodiment schemes, rather than being intended to limit the scope of the present invention, which is limited only by the appended claims.
[0079] Model experiment: In the experiment, for the multimodal peptide and TCR binding prediction network model of the present invention, the sequence pair, the three-dimensional coordinate information of the sequence pair, and the label constitute the input data of the network. The sequence pair input is the value corresponding to the amino acid dictionary, the input dimension is 20, and it is encoded including word unit encoding and position encoding, and the encoded data is floating point type, and the encoding dimension is set to 64; the three-dimensional coordinates of the sequence pair are as follows Figure 5As shown, the position coordinates corresponding to the 'CA' position are extracted as the three-dimensional coordinates of the amino acid, which are floating-point data with a dimension of 3. The label data is 0, 1, where 0 indicates that the peptide does not bind to the TCR and 1 indicates that the peptide binds to the TCR. The multimodal peptide-TCR binding prediction network is a binary classification model. The network output is softmax processed. If the prediction value is greater than 0.5, it indicates that the peptide binds to the TCR; otherwise, it does not bind to the TCR. During training, the network parameters are initialized using nn.init.xavier_uniform_, and the Adam optimizer is used with a weight decay of 1e-5.
[0080] In summary, the proposed approach considers both the correlation characteristics of peptide-TCR interactions and the spatial structural features of sequences. By combining these inter-sequence correlation characteristics with sequence structural features through a multimodal deep learning model, a more comprehensive prediction approach is captured, further improving the accuracy of peptide-TCR sequence binding prediction.
[0081] The second aspect of the present invention discloses a prediction system for peptide-TCR binding based on deep learning. Figure 4 is a structural diagram of a prediction system for peptide and TCR binding based on deep learning according to an embodiment of the present invention; Figure 4 As shown, the system 100 includes:
[0082] The first processing module 101 is configured to digitally process the sequences of the peptide and TCR according to the digital dictionary of amino acids and extract correlation features between the peptide and TCR sequences;
[0083] The second processing module 102 is configured to obtain the three-dimensional structure of the polypeptide and TCR sequence according to the HelixFoldSingle protein structure prediction model, and extract the three-dimensional structural features of the polypeptide and TCR sequence;
[0084] The third processing module 103 is configured to linearly weight the correlation characteristics between the polypeptide and the TCR sequence and the three-dimensional structural characteristics of the polypeptide and the TCR sequence to obtain a prediction result of the binding between the polypeptide and the TCR.
[0085] According to the system of the second aspect of the present invention, the first processing module 101 is specifically configured to perform digital processing on the sequences of the polypeptide and TCR according to the digital dictionary of amino acids to extract correlation features between the polypeptide and TCR sequences, including:
[0086] The sequences of the polypeptide and TCR are digitized according to the digital dictionary of amino acids to obtain word units and position codes of the polypeptide sequence of a predefined length and word units and position codes of the TCR sequence; the word units and position codes of the polypeptide sequence and the TCR sequence are combined to obtain a first coding combination and a second coding combination; the first coding combination is input into a first cross-attention Transformer model to obtain a first correlation feature; the second coding combination is input into a second cross-attention Transformer model to obtain a second correlation feature; the first correlation feature and the second correlation feature are concatenated and input into a first feedforward fully connected neural network to obtain the correlation feature of the polypeptide and TCR sequences; the word unit and position code is the sum of the word unit code and the position code.
[0087] The method of digitizing the sequences of the polypeptide and TCR according to the digital dictionary of amino acids to obtain the word element and position code of the polypeptide sequence of the predefined length and the word element and position code of the TCR sequence includes:
[0088] The amino acid sequence is tokenized according to a digital dictionary; in the digital dictionary, "*" represents an ambiguous amino acid in the sequence; the amino acid sequence converted by the digital dictionary is used to form a tokenized sequence of the polypeptide and TCR; the length of the sequence of the polypeptide and TCR is aligned to a predefined length, and if the length of the sequence of the polypeptide and TCR is less than the predefined length, it is padded with 0; the amino acids in the sequence of the polypeptide and TCR aligned to the predefined length are token-encoded and position-encoded to obtain the token and position coding of the polypeptide sequence of the predefined length and the token and position coding of the TCR sequence.
[0089] In some embodiments, the predefined length is 20.
[0090] The first coding combination includes: a first input of the first coding combination is a word unit and position code of a polypeptide sequence, a second input of the first coding combination is a word unit and position code of a TCR sequence, and a third input of the first coding combination is a word unit and position code of a TCR sequence;
[0091] The second coding combination includes: the first input of the second coding combination is the word element and position code of the TCR sequence, the second input of the second coding combination is the word element and position code of the polypeptide sequence, and the third input of the second coding combination is the word element and position code of the polypeptide sequence.
[0092] Specifically, each cross-attention Transformer model consists of two parts: the cross-attention Transformer encoder and the self-attention Transformer encoder. The cross-attention encoder layer mainly consists of multi-head cross-attention, normalization layer, fully connected layer, and residual layer; the self-attention encoder layer consists of multi-head self-attention, normalization layer, fully connected layer, and residual layer.
[0093] In some embodiments, as Figure 2 As shown in the figure, the Cross-Attention Transformer model consists of a layer of Cross-Attention Transformer encoders and multiple layers of Self-Attention Transformer encoders. The network is mainly composed of multiple layers of encoders stacked together. Each encoder layer consists of two parts: the first sub-layer includes multi-head attention, a fully connected layer, a residual layer, and a normalization layer; the second sub-layer includes two fully connected layers, a residual layer, and a normalization layer.
[0094] The first feed-forward fully connected neural network has a fully connected layer, a Dropout layer, a normalization layer, and uses the LeakyReLU activation function.
[0095] According to the system of the second aspect of the present invention, the second processing module 102 is specifically configured to obtain the three-dimensional structure of the polypeptide and TCR sequence according to the HelixFoldSingle protein structure prediction model, and extract the three-dimensional structural features of the polypeptide and TCR sequence, including:
[0096] The predicted structures of the polypeptide and TCR are obtained according to the HelixFoldSingle protein structure prediction model; and the three-dimensional structural characteristics of the polypeptide and TCR sequences are obtained according to the three-dimensional coordinates of the predicted structures.
[0097] Obtaining the three-dimensional coordinates of the predicted structure includes:
[0098] The coordinates of the 'CA' atom of each residue in the predicted structure were taken as the three-dimensional coordinates of each amino acid.
[0099] Obtaining the three-dimensional structural features of the polypeptide and TCR sequence based on the three-dimensional coordinates of the predicted structure includes:
[0100] Obtaining the three-dimensional coordinates of the polypeptide and TCR sequence according to the three-dimensional coordinates of the predicted structure;
[0101] The three-dimensional coordinates of the peptide sequence are input into the first convolutional neural network to obtain the three-dimensional structural features of the peptide;
[0102] The three-dimensional coordinates of the TCR sequence are input into the second convolutional neural network to obtain the three-dimensional structural features of the TCR;
[0103] After splicing the three-dimensional structural features of the polypeptide and the three-dimensional structural features of TCR, the three-dimensional structural features are input into a second feedforward fully connected neural network to obtain the three-dimensional structural features of the polypeptide and TCR sequences.
[0104] Specifically, each convolutional neural network consists of multiple convolutional layers, a normalization layer, a dropout layer, and a maximum pooling layer.
[0105] In some embodiments, as Figure 3 As shown, the convolutional neural network consists of multiple layers of 2D convolutional layers, normalization layers, Dropout layers, and maximum pooling layers. The fully connected network consists of Dropout layers, normalization layers, the activation function of the middle layer uses LeakyReLU, and the activation function of the last layer uses Sigmoid.
[0106] The second feedforward fully connected neural network has a Dropout layer and a normalization layer. The activation function of the middle layer is LeakyReLU, and the activation function of the last layer is Sigmoid function.
[0107] The entire network model uses the Adam optimizer and adds a weight decay rate of 1e-5 to obtain the final prediction results of the binding between peptide and TCR sequence.
[0108] In some embodiments, two data sets are used, the first data set is divided into a training set and a validation set according to an 8:2 ratio, and the second data set, in which the peptides and TCR sequences that do not appear in the training set and the validation set are used as the test set.
[0109] A third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of any one of the deep learning-based methods for predicting peptide-TCR binding according to the first aspect of the present invention.
[0110] Figure 5 FIG. 1 is a structural diagram of an electronic device according to an embodiment of the present invention. Figure 5As shown, the electronic device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, near field communication (NFC) or other technologies. The display screen of the electronic device can be a liquid crystal display or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the electronic device housing, or an external keyboard, touchpad or mouse.
[0111] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a structural diagram of the part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0112] A fourth aspect of the present invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of any one of the methods for predicting peptide-TCR binding based on deep learning disclosed in the first aspect of the present invention.
[0113] Please note that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. The above embodiments only express several implementation methods of the present application. The description is relatively specific and detailed, but it cannot be understood as a limitation on the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, without departing from the concept of this application, several variations and improvements can be made, which all fall within the scope of protection of this application. Therefore, the scope of protection of the patent in this application shall be based on the attached claims.
Claims
1. A method for predicting peptide-TCR binding based on deep learning, characterized in that: The method comprises: Step S1: digitize the sequences of the peptide and TCR according to the digital dictionary of amino acids to extract correlation features between the peptide and TCR sequences; Step S2: obtaining the three-dimensional structure of the polypeptide and TCR sequence according to the HelixFoldSingle protein structure prediction model, and extracting the three-dimensional structural features of the polypeptide and TCR sequence; Step S3: linearly weighting the correlation features of the polypeptide and TCR sequence and the three-dimensional structural features of the polypeptide and TCR sequence to obtain a prediction result of the binding of the polypeptide to the TCR; In step S1, the steps of digitizing the sequences of the peptide and TCR according to the digital dictionary of amino acids and extracting correlation features of the peptide and TCR sequences include: The sequences of the polypeptide and TCR are digitized according to a digital dictionary of amino acids to obtain word units and position codes of the polypeptide sequence of a predefined length and word units and position codes of the TCR sequence; the word units and position codes of the polypeptide sequence and the TCR sequence are combined to obtain a first coding combination and a second coding combination; the first coding combination is input into a first cross-attention Transformer model to obtain a first correlation feature; the second coding combination is input into a second cross-attention Transformer model to obtain a second correlation feature; the first correlation feature and the second correlation feature are concatenated and input into a first feedforward fully connected neural network to obtain a correlation feature of the polypeptide and TCR sequences; the word unit and position code is the sum of the word unit code and the position code; In step S1, the digitization of the polypeptide and TCR sequences according to the digital dictionary of amino acids to obtain the word unit and position code of the polypeptide sequence of the predefined length and the word unit and position code of the TCR sequence includes: The amino acid sequence is tokenized according to a digital dictionary; in the digital dictionary, "*" represents an ambiguous amino acid in the sequence; the amino acid sequence converted by the digital dictionary is used to form a tokenized sequence of the polypeptide and TCR; the length of the sequence of the polypeptide and TCR is aligned to a predefined length, and if the length of the sequence of the polypeptide and TCR is less than the predefined length, it is padded with 0; the amino acids in the sequence of the polypeptide and TCR aligned to the predefined length are token-encoded and position-encoded to obtain the token and position coding of the polypeptide sequence of the predefined length and the token and position coding of the TCR sequence.
2. The method for predicting peptide-TCR binding based on deep learning according to claim 1, characterized in that: In step S1, the first coding combination includes: a first input of the first coding combination is a word unit and position code of a polypeptide sequence, a second input of the first coding combination is a word unit and position code of a TCR sequence, and a third input of the first coding combination is a word unit and position code of a TCR sequence; The second coding combination includes: the first input of the second coding combination is the word element and position code of the TCR sequence, the second input of the second coding combination is the word element and position code of the polypeptide sequence, and the third input of the second coding combination is the word element and position code of the polypeptide sequence.
3. The method for predicting peptide-TCR binding based on deep learning according to claim 1, characterized in that: In step S2, obtaining the three-dimensional coordinates of the polypeptide and the TCR sequence according to the HelixFoldSingle protein structure prediction model and extracting the three-dimensional structural features of the polypeptide and the TCR sequence includes: The predicted structures of the polypeptide and TCR are obtained according to the HelixFoldSingle protein structure prediction model; and the three-dimensional structural characteristics of the polypeptide and TCR sequences are obtained according to the three-dimensional coordinates of the predicted structures.
4. The method for predicting peptide-TCR binding based on deep learning according to claim 3, characterized in that: In step S2, obtaining the three-dimensional coordinates of the predicted structure includes: The coordinates of the 'CA' atom of each residue in the predicted structure were taken as the three-dimensional coordinates of each amino acid.
5. The method for predicting peptide-TCR binding based on deep learning according to claim 3, characterized in that: In step S2, obtaining the three-dimensional structural features of the polypeptide and TCR sequence based on the three-dimensional coordinates of the predicted structure includes: Obtaining the three-dimensional coordinates of the polypeptide and TCR sequence according to the three-dimensional coordinates of the predicted structure; The three-dimensional coordinates of the peptide sequence are input into the first convolutional neural network to obtain the three-dimensional structural features of the peptide; The three-dimensional coordinates of the TCR sequence are input into the second convolutional neural network to obtain the three-dimensional structural features of the TCR; After splicing the three-dimensional structural features of the polypeptide and the three-dimensional structural features of TCR, the three-dimensional structural features are input into a second feedforward fully connected neural network to obtain the three-dimensional structural features of the polypeptide and TCR sequences.
6. A deep learning-based prediction system for peptide-TCR binding, characterized in that: The system comprises: The first processing module is configured to digitally process the sequences of the peptide and TCR according to the digital dictionary of amino acids and extract correlation features between the peptide and TCR sequences; The step of digitally processing the sequences of the peptide and TCR according to the digital dictionary of amino acids and extracting correlation features between the peptide and TCR sequences includes: The sequences of the polypeptide and TCR are digitized according to a digital dictionary of amino acids to obtain word units and position codes of the polypeptide sequence of a predefined length and word units and position codes of the TCR sequence; the word units and position codes of the polypeptide sequence and the TCR sequence are combined to obtain a first coding combination and a second coding combination; the first coding combination is input into a first cross-attention Transformer model to obtain a first correlation feature; the second coding combination is input into a second cross-attention Transformer model to obtain a second correlation feature; the first correlation feature and the second correlation feature are concatenated and input into a first feedforward fully connected neural network to obtain a correlation feature of the polypeptide and TCR sequences; the word unit and position code is the sum of the word unit code and the position code; The method of digitizing the sequences of the polypeptide and TCR according to the digital dictionary of amino acids to obtain the word element and position code of the polypeptide sequence of the predefined length and the word element and position code of the TCR sequence includes: The amino acid sequence is tokenized according to a digital dictionary; in the digital dictionary, "*" represents an ambiguous amino acid in the sequence; the amino acid sequence after the digital dictionary conversion is used to form a tokenized sequence of the polypeptide and TCR; the length of the sequence of the polypeptide and TCR is aligned to a predefined length, and if the length of the sequence of the polypeptide and TCR is less than the predefined length, it is padded with 0; the amino acids in the sequence of the polypeptide and TCR aligned to the predefined length are token-encoded and position-encoded to obtain the token and position coding of the polypeptide sequence of the predefined length and the token and position coding of the TCR sequence; The second processing module is configured to obtain the three-dimensional coordinates of the polypeptide and TCR sequence according to the HelixFoldSingle protein structure prediction model and extract the three-dimensional structural features of the polypeptide and TCR sequence; The third processing module is configured to linearly weight the correlation characteristics between the polypeptide and the TCR sequence and the three-dimensional structural characteristics of the polypeptide and the TCR sequence to obtain a prediction result of the binding between the polypeptide and the TCR.
7. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the steps in the method for predicting polypeptide and TCR binding based on deep learning according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps in the method for predicting polypeptide and TCR binding based on deep learning according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system for predicting interaction between T cell receptor and peptide
CN116665769A
Polypeptide toxicity determination method, apparatus and device, and storage medium
CN118471346A
Cited By
Polypeptide target spot prediction method based on sequence and structural characteristics
CN121528292A