A method for establishing a human alpha beta TCR and pMHC binding probability prediction model
By constructing a convolutional neural network prediction model, the problems of low throughput and high cost in the detection of TCR and pMHC in the existing technology are solved, and the specific identification and binding of TCR and pMHC are realized through rapid large-scale screening and efficient prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN UNIV
- Filing Date
- 2022-07-27
- Publication Date
- 2026-06-02
AI Technical Summary
Existing experimental methods have low throughput and high cost when detecting the specific recognition and binding of human TCR and pMHC, making them difficult to adapt to large-scale screening, and lack prediction methods based on neural networks.
A convolutional neural network was used to establish a probability prediction model for the binding of human αβTCR and pMHC. By acquiring specific binding data of TCR and pMHC, data preprocessing and encoding were performed to generate positive and negative samples. The convolutional neural network was then used for training and cross-validation to construct the prediction model.
This study enables rapid, large-scale screening for the specific recognition and binding of TCR and pMHC. The model exhibits good generalization and predictive performance, and assists experimental methods in verifying the specific recognition and binding of TCR and pMHC.
Smart Images

Figure CN115188410B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedicine, specifically relating to a method for establishing a probability prediction model of human αβTCR and pMHC binding. Background Technology
[0002] The human adaptive immune system is divided into two main categories: humoral immunity and cellular immunity. Humoral immunity is mediated by B cells and primarily targets extracellular antigens. Cellular immunity is mediated by T cells and primarily targets intracellular antigens presented by antigen-presenting cells (APCs). Cellular immunity can be divided into three stages: the T cell-specific antigen recognition stage; the T cell activation, proliferation, and differentiation stage; and the effector T cell generation and effector stage. The key to T cell-specific antigen recognition is the specific binding of the T cell receptor (TCR) on the cell surface to the peptide-major histocompatibility complex (pMHC) presented by APCs; this process is also known as antigen recognition.
[0003] The genes encoding the α and β chains of the TCR are located at the TRA and TRB loci, respectively, and exist as numerous, separated gene segments, including the V, D, J, and C genes. During maturation, the TCR undergoes gene rearrangement to form a V-(D)-J link, which then joins with the C gene segment to ultimately form the gene encoding the TCR. Due to the randomness of gene rearrangement, with random insertions, deletions, or substitutions of nucleotides and random pairing of the α and β chains, the diversity of TCRs in the human body reaches 10-1. 6 ~10 7 The variety of species makes it extremely difficult to detect the specific recognition between TCR and pMHC through experimental methods.
[0004] In 2018, Zhang et al. developed a single-cell sequencing technology, TetTCR-seq, which enables paired sequencing of the T cell TCR α and β chains, while simultaneously obtaining information on the pMHC that the cells specifically recognize. This method utilizes in vitro transcription and translation during peptide synthesis, and uses the oligonucleotide encoding the peptide as a DNA tag for the pMHC tetramer. After staining T cells with the pMHC tetramer, the T cells are sorted into single cells, followed by sequencing of the DNA tag and the TCR α and β chains. TCR information can be linked to the antigens that the DNA tag specifically binds to. Another similar single-cell sequencing method was developed by 10x Genomics, a company specializing in single-cell sequencing solutions. Compared to TetTCR-seq, this single-cell sequencing protocol uses 10x Genomics' water-in-oil emulsion droplet method, but also employs the MHC multimer method to detect the specific recognition and binding of TCRs to pMHC.
[0005] Although current experimental methods have high specificity and sensitivity, their throughput for pMHC detection remains low. TetTCR-seq can detect 315 pMHCs in a single run, while the 10x Genomics protocol can only detect 44. Furthermore, these methods are time-consuming, labor-intensive, and costly, making them unsuitable for large-scale antigen screening.
[0006] Currently, deep neural networks have been widely used in fields such as image recognition and text processing. Their characteristic is that they extract features through training with big data for solving classification problems. However, no method based on neural networks that can effectively predict whether TCR and pMHC specifically recognize and bind has been reported. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a method for establishing a probability prediction model of the combination of human αβTCR and pMHC based on a convolutional neural network.
[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0009] A method for establishing a probability prediction model for the binding of human αβTCR and pMHC includes the following steps:
[0010] (1) Obtain raw data on TCR and pMHC specific binding, including CDR3β amino acid sequence, polypeptide sequence and / or TRAV gene type, TRBV gene type, CDR3α amino acid sequence, and MHC typing information;
[0011] (2) Data preprocessing: The amino acid sequences of CDR1α, CDR2α, CDR1β and CDR2β were obtained from the annotation information of TRAV and TRBV, and the amino acid sequence of MHC was obtained from the annotation information of MHC; the positions of amino acid residues with upward side chains in the two α helices of the HLA-A*02:01 subtype were determined according to the protein structure, and the amino acid residue sequences at the same position on the two α helices of all other MHC subtypes were designated as MHCα1contact and MHCα2contact, respectively.
[0012] (3) Generate negative samples: Define the data that specifically binds αβTCR and pMHC as positive samples. Randomly combine TCR and pMHC in the positive samples, remove data that are repeated with the positive samples, and randomly sample to generate negative samples according to a certain proportion to ensure that the frequency of all pMHC types in the negative samples is the same as that in the positive samples.
[0013] (4) Data encoding: The amino acid sequences of TCR and pMHC are arranged into two-dimensional matrices respectively. The two amino acid two-dimensional matrices are encoded using the five Atchley factors of amino acids. Missing amino acid sequences are filled with 0.
[0014] (5) Training and testing of convolutional neural networks: Using the two-dimensional matrices of TCR and pMHC as input, positive samples are labeled as "combined" and negative samples are labeled as "non-combined". The positive and negative samples are used as training data by convolutional neural networks respectively, and the convolutional neural networks are trained according to the corresponding labels. The output results of both are then used as input to further train the convolutional neural network and perform cross-validation to obtain a prediction model of the combination probability of TCR and pMHC.
[0015] Furthermore, the raw data obtained in step (1) is used as positive sample data.
[0016] Furthermore, the data preprocessing method in step (2) includes the following sub-steps:
[0017] 2.1 The amino acid sequences of CDR1α, CDR2α, CDR1β and CDR2β were obtained based on the annotation information of human TRAV and TRBV genes on the IMGT website;
[0018] 2.2 Based on the annotation information of human MHC (i.e., HLA, hereinafter MHC and HLA have the same meaning, both referring to human MHC) on the IMGT website, the amino acid sequences of all MHC isoforms were obtained; based on the protein structure of the HLA-A*02:01 isoform (PDB database search number 6RSY), the positions of the amino acid residues with upward-facing side chains in its two α-helices α1 and α2 were determined. The selected residue positions in the α1 helix were 59, 62, 63, 66, 67, 69, 70, 73, 74, 76, 77, 80, and 81. HLA-A*02:01α1 The corresponding amino acid residues are E, D, G, R, K, K, A, Q, T, R, V, G, T. The selected residue positions on the α2 helix are 147, 150, 151, 152, 155, 156, 159, 163, 164, 167, 168. The corresponding amino acid residues for HLA-A*02:01α1 are K, A, A, H, E, Q, A, G, T, E, W. The amino acid residue sequences at the same positions on the two α helices of all other MHC subtypes are designated as MHCα1contact and MHCα2contact, respectively.
[0019] Furthermore, in step (3), the ratio of positive samples to negative samples is 1:0.5 to 1. Preferably, the ratio of positive samples to negative samples is 1:1.
[0020] Furthermore, step (4) data encoding includes the following sub-steps:
[0021] 4.1 Arrange the CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, and CDR3β amino acid sequences of the TCR into a two-dimensional matrix. The first row of the matrix contains CDR1β+CDR2β amino acid residues, the second row contains CDR3β amino acid residues, the third row contains CDR3α amino acid residues, and the fourth row contains CDR1α+CDR2α amino acid residues. Each row has a length of 24, and empty parts are filled with 0. The matrix size is (4,24).
[0022] 4.2 Arrange the amino acid sequences of the pMHC peptide, MHCα1contact, and MHCα2contact into a two-dimensional matrix. The first row of the matrix contains MHCα1contact amino acid residues, the second row contains peptide amino acid residues, and the third row contains MHCα2contact amino acid residues. Each row is 13 in length, and empty parts are filled with 0. The matrix size is (3,13).
[0023] 4.3 The two amino acid two-dimensional matrices above were encoded using five Atchley factors, with missing amino acid sequences filled with 0. Atchley factors were used to transform the letters in the subsequence listing into a set of five data points representing their physicochemical properties. Atchley factors are a set of highly interpretable numerical patterns of amino acid variation. These high-dimensional attribute data are summarized by multidimensional patterns reflecting the covariation of five attributes: polarity, secondary structure, molecular volume, codon diversity, and electrostatic charge. This provides a natural and fundamental metric for comparing these letter data representing amino acids.
[0024] Furthermore, in step (5), the convolutional neural network used for TCR training includes three convolutional layers TC1, TC2, and TC3, a fully connected layer, and an output layer. The data of the fully connected layer is batch normalized and then output through the ReLU function. The convolutional neural network used for pMHC training includes three convolutional layers PC1, PC2, and PC3, a fully connected layer, and an output layer. The data of the fully connected layer is batch normalized and then output through the ReLU function.
[0025] Furthermore, the convolution kernels of TC1, TC2, and TC3 are (3,3) with a stride of 1, and are initialized using a Kaiming normal distribution.
[0026] Furthermore, the convolution kernels of PC1, PC2, and PC3 are (3,3) with a stride of 1, and are initialized using a Kaiming normal distribution.
[0027] Furthermore, the output values of the fully connected layers of the convolutional neural network used for TCR training and the convolutional neural network used for pMHC training are merged and input into the fully connected layer, then output after batch normalization and ReLU function, and then output after passing through the connected layer and softmax function.
[0028] Furthermore, cross-validation is used to test the model, and one or more of the following restrictions are imposed:
[0029] (1) CDR3α in the test set has not appeared in the training set;
[0030] (2) CDR3β in the test set has not appeared in the training set;
[0031] (3) The data in the test set contains only one or more of CDR1β, CDR2β and CDR3β;
[0032] (4) The data in the test set contains only one or more of CDR1α, CDR2α and CDR3α;
[0033] (5) The data in the test set must include CDR1α, CDR2α, CDR3α, CDR1β, CDR2β and CDR3β.
[0034] The beneficial effects of this invention are as follows:
[0035] (1) Based on the structural features of TCR-pMHC recognition, this invention extracts important sequences from the full-length amino acid sequence, encodes the amino acid sequence using a special encoding method, trains the convolutional neural network using the encoded data, and tests the model under strict conditions. The final model has good performance in the specific recognition and binding prediction of TCR and pMHC.
[0036] (2) This invention provides a prediction method for the specific recognition and binding of TCR and pMHC, including a model construction method and an effective prediction model. The model has good generalization ability and can quickly screen TCR and pMHC on a large scale. It assists experimental methods in verifying the specific recognition and binding of TCR and pMHC and has important practical value for the study of T cell immunity. Attached Figure Description
[0037] Figure 1 A schematic diagram of TCR and pMHC encoding.
[0038] Figure 2 This is a schematic diagram of a convolutional neural network structure.
[0039] Figure 3 The results of 5-fold cross-validation are shown for models trained using different TCR and pMHC input information.
[0040] Figure 4 This represents the 5-fold cross-validation results of the prediction model on the test set under different conditions.
[0041] Figure 5 The ROC curves are for the two models.
[0042] Figure 6 This is a prediction of CMV virus antigen-specific TCR. Detailed Implementation
[0043] To better understand the present invention, the following embodiments further illustrate the content of the present invention, but the content of the present invention is not limited to the following embodiments.
[0044] Example 1
[0045] A method for establishing a probabilistic prediction model for the binding of human αβTCR and pMHC, comprising the following steps:
[0046] Three thousand αβTCR-pMHC specific recognition data were collected as positive samples. Each data point included the CDR3β amino acid sequence, CDR3α amino acid sequence, TRBV gene type, TRAV gene type, polypeptide sequence, and MHC type. For each pMHC (peptide and MHC) in the positive samples, TCRs other than those paired with the pMHC in the positive samples were randomly combined to generate negative samples at a positive sample:negative sample ratio of 1:1, resulting in a total of 3000 negative sample data points.
[0047] The amino acid sequences of CDR1α, CDR2α, CDR1β, and CDR2β were obtained based on the annotation information of the human TRAV and TRBV genes on the IMGT website. The amino acid sequences of all MHC isotypes were obtained based on the annotation information of the human MHC on the IMGT website. Amino acid residues at positions 59, 62, 63, 66, 67, 69, 70, 73, 74, 76, 77, 80, and 81 of the MHC amino acid sequence were selected as the MHCα1 contact sequence, and amino acid residues at positions 147, 150, 151, 152, 155, 156, 159, 163, 164, 167, and 168 of the MHC amino acid sequence were selected as the MHCα2 contact sequence.
[0048] Twelve different information combinations were designed by combining all the information from TCR and pMHC. The information combinations and the information used were: TRBcdr3-pep (CDR3β and peptide), TRBcdr3-pep-MHC (CDR3β, peptide, MHCα1 contact sequence and MHCα2 contact sequence), TRBcdr123-pep (CDR1β, CDR2β, CDR3β and peptide), TRBcdr123-pep-MHC (CDR1β, CDR2β, CDR3β, peptide, MHCα1 contact sequence and MHCα2 contact sequence), TRAcdr3-pep (CDR3α and peptide), and TRAcdr3-pep-MHC (CDR3α, peptide, MHCα1 contact sequence and MHCα2 contact sequence). TRAcdr123-pep (CDR1α, CDR2α, CDR3α and peptide), TRAcdr123-pep-MHC (CDR1α, CDR2α, CDR3α, peptide, MHCα1 contact sequence and MHCα2 contact sequence), TRAcdr3-TRBcdr3-pep (CDR3α, CDR3β and peptide), TRAcdr3-TRBcdr3-pep-MHC (CDR3α, CDR3β, peptide, MHCα1 contact sequence and MHCα2 contact sequence ... (contact sequences), TRAcdr123-TRBcdr123-pep (CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, CDR3β and peptides) and TRAcdr123-TRBcdr123-pep-MHC (CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, CDR3β, peptides, MHCα1 contact sequence and MHCα2 contact sequence).
[0049] Next, the TCR-pMHC amino acid sequences with different information combinations were arranged into a two-dimensional matrix. Taking one data point as an example, its CDR1α is TSINN, CDR2α is IRSNERE, CDR3α is CATGLTGGGNKLTF, CDR1β is MNHEY, CDR2β is SMNVEV, CDR3β is CASSLGGQNYGYTF, the polypeptide is KLGGALQAK, the MHCα1 contact is EDQRNKAQTRVGT, and the MHCα2 contact is KAAHEQAGTEW. The specific arrangements of the twelve information combinations are as follows:
[0050] TRBcdr3-pep
[0051] TCR
[0052] C A S S L G G Q N Y G Y T F * * * * * * * * * *
[0053] pMHC
[0054] K L G G A L Q A K * * * *
[0055] TRBcdr3-pep-MHC
[0056] TCR
[0057] C A S S L G G Q N Y G Y T F * * * * * * * * * *
[0058] pMHC
[0059] E D Q R N K A Q T R V G T K L G G A L Q A K * * * * K A A H E Q A G T E W * *
[0060] TRBcdr123-pep
[0061] TCR
[0062] M N E Y * * * * * * * * * * * * * * S M N V E V C A S S L G G Q N Y G Y T F * * * * * * * * * *
[0063] pMHC
[0064] K L G G A L Q A K * * * *
[0065] TRBcdr123-pep-MHC
[0066] TCR
[0067] M N H E Y * * * * * * * * * * * * * S M N V E V C A S S L G G Q N Y G Y T F * * * * * * * * * *
[0068] pMHC
[0069] E D Q R N K A Q T R V G T K L G G A L Q A K * * * * K A A H E Q A G T E W * *
[0070] TRAcdr3-pep
[0071] TCR
[0072] C A T G L T G G G N K L T F * * * * * * * * * *
[0073] pMHC
[0074] K L G G A L Q A K * * * *
[0075] TRAcdr3-pep-MHC
[0076] TCR
[0077] C A T G L T G G G N K L T F * * * * * * * * * *
[0078] pMHC
[0079] E D Q R N K A Q T R V G T K L G G A L Q A K * * * * K A A H E Q A G T E W * *
[0080] TRAcdr123-pep
[0081] TCR
[0082] C A T G L T S G G N K L T F * * * * * * * * * * T S I N N * * * * * * * * * * * * I R S N E R E
[0083] pMHC
[0084] K L G G A L Q A K * * * *
[0085] TRAcdr123-pep-MHC
[0086] TCR
[0087] G A T G L T G G G N K L T F * * * * * * * * * * T S I N N * * * * * * * * * * * * I R S N E R E
[0088] pMHC
[0089] E D Q R N K A Q T R V G T K L G G A L Q A K * * * * K A A H E Q A G T E W * *
[0090] TRAcdr3-TRBcdr3-pep
[0091] TCR
[0092] C A S S L G G Q N Y G Y T F * * * * * * * * * * C A T G L T G G G N K L T F * * * * * * * * * *
[0093] pMHC
[0094] K L G G A L Q A K * * * *
[0095] TRAcdr3-TRBcdr3-pep-MHC
[0096] TCR
[0097] C A S S L G G Q N Y G Y T F * * * * * * * * * * C A T G L T G G G N K L T * * * * * * * * * * *
[0098] E D Q R N K A Q T R V G T K L G G A L Q A K * * * * K A A H E Q A G T E W * *
[0099] TRAcdr123-TRBcdr123-pep
[0100] M N H E Y * * * * * * * * * * * * * S M N V E V C A S S L G G Q N Y G Y T F * * * * * * * * * * C A T G L T G G G N K L T F * * * * * * * * * * T S I N N * * * * * * * * * * * * I R S N E R E
[0101] K L G G A L Q A K * * * *
[0102] TRAcdr123-TRBcdr123-pep-MHC
[0103] M N H E Y * * * * * * * * * * * * * S M N V E V C A S S L G G Q N Y G Y T F * * * * * * * * * * C A T G L T G G G N K L T F * * * * * * * * * * T S I N N * * * * * * * * * * * * I R S N E R E
[0104] E D Q R N K A Q T R V G T K L G G A L Q A K * * * * K A A H E Q A G T E W * *
[0105] The TRAcdr123-TRBcdr123-pep-MHC combination uses Atchley factors for encoding. The specific values of the Atchley factors for the 20 amino acid residues are shown in the table below:
[0106]
[0107]
[0108] After encoding, two three-dimensional matrices are obtained, as follows:
[0109] TCR, size (5, 4, 24)
[0110] [[[-0.663,0.945,0.336,1.357,0.26,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,-0.228,-0.663,0.945,-1.337,1.357,-1.337],
[0111] [-1.343,-0.591,-0.228,-0.228,-1.019,-0.384,-0.384,0.931,0.945,0.26,-0.384,0.26,-0.032,-1.006,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0112] [-1.343,-0.591,-0.032,-0.384,-1.019,-0.032,-0.384,-0.384,-0.384,0.945,1.831,-1.019,-0.032,-1.006,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0113] [-0.032,-0.228,-1.239,0.945,0.945,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,-1.239,1.538,-0.228,0.945,1.357,1.538,1.357]],
[0114] [[-1.524,0.828,-0.417,-1.453,0.83,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,1.399,-1.524,0.828,-0.279,-1.453,-0.279],
[0115] [0.465,-1.302,1.399,1.399,-0.987,1.652,1.652,-0.179,0.828,0.83,1.652,0.83,0.326,-0.59,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0116] [0.465,-1.302,0.326,1.652,-0.987,0.326,1.652,1.652,1.652,0.828,-0.561,-0.987,0.326,-0.59,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0117] [0.326,1.399,-0.547,0.828,0.828,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,-0.547,-0.055,1.399,0.828,-1.453,-0.055,-1.453]],
[0118] [[2.219,1.299,-1.673,1.477,3.097,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,-4.76,2.219,1.299,-0.544,1.477,-0.544],
[0119] [-0.862,-0.733,-4.76,-4.76,-1.505,1.33,1.33,-3.005,1.299,3.097,1.33,3.097,2.213,1.891,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0120] [-0.862,-0.733,2.213,1.33,-1.505,2.213,1.33,1.33,1.33,1.299,0.533,-1.505,2.213,1.891,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0121] [2.213,-4.76,2.131,1.299,1.299,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,2.131,1.502,-4.76,1.299,1.477,1.502,1.477]],
[0122] [[-1.005,-0.169,-1.474,0.113,-0.838,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.67,-1.005,-0.169,1.242,0.113,1.242],
[0123] [-1.02,1.57,0.67,0.67,1.266,1.045,1.045,-0.503,-0.169,-0.838,1.045,-0.838,0.908,-0.397,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0124] [-1.02,1.57,0.908,1.045,1.266,0.908,1.045,1.045,1.045,-0.169,-0.277,1.266,0.908,-0.397,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0125] [0.908,0.67,0.393,-0.169,-0.169,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.393,0.44,0.67,-0.169,0.113,0.44,0.113]],
[0126] [[1.212,0.933,-0.078,-0.837,1.512,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,-2.647,1.212,0.933,-1.262,-0.837,-1.262],
[0127] [-0.255,-0.146,-2.647,-2.647,-0.912,2.064,2.064,-1.853,0.933,1.512,2.064,1.512,1.313,0.412,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0128] [-0.255,-0.146,1.313,2.064,-0.912,1.313,2.064,2.064,2.064,0.933,1.648,-0.912,1.313,0.412,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0],
[0129] [1.313,-2.647,0.816,0.933,0.933,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.816,2.897,-2.647,0.933,-0.837,2.897,-0.837]]]
[0130] pMHC, size (5, 3, 14)
[0131] [[[1.357,1.05,0.931,1.538,0.945,1.831,-0.591,0.931,-0.032,1.538,-1.337,-0.384,-0.032],
[0132] [1.831,-1.019,-0.384,-0.384,-0.591,-1.019,0.931,-0.591,1.831,0.0,0.0,0.0,0.0],
[0133] [1.831,-0.591,-0.591,0.336,1.357,0.931,-0.591,-0.384,-0.032,1.357,-0.595,0.0,0.0]],
[0134] [[-1.453,0.302,-0.179,-0.055,0.828,-0.561,-1.302,-0.179,0.326,-0.055,-0.279,1.652,0.326],
[0135] [-0.561,-0.987,1.652,1.652,-1.302,-0.987,-0.179,-1.302,-0.561,0.0,0.0,0.0,0.0],
[0136] [-0.561,-1.302,-1.302,-0.417,-1.453,-0.179,-1.302,1.652,0.326,-1.453,0.009,0.0,0.0]],
[0137] [[1.477,-3.656,-3.005,1.502,1.299,0.533,-0.733,-3.005,2.213,1.502,-0.544,1.33,2.213],
[0138] [0.533,-1.505,1.33,1.33,-0.733,-1.505,-3.005,-0.733,0.533,0.0,0.0,0.0,0.0],
[0139] [0.533,-0.733,-0.733,-1.673,1.477,-3.005,-0.733,1.33,2.213,1.477,0.672,0.0,0.0]],
[0140] [[0.113,-0.259,-0.503,0.44,-0.169,-0.277,1.57,-0.503,0.908,0.44,1.242,1.045,0.908],
[0141] [-0.277,1.266,1.045,1.045,1.57,1.266,-0.503,1.57,-0.277,0.0,0.0,0.0,0.0],
[0142] [-0.277,1.57,1.57,-1.474,0.113,-0.503,1.57,1.045,0.908,0.113,-2.128,0.0,0.0]],
[0143] [[-0.837,-3.242,-1.853,2.897,0.933,1.648,-0.146,-1.853,1.313,2.897,-1.262,2.064,1.313],
[0144] [1.648,-0.912,2.064,2.064,-0.146,-0.912,-1.853,-0.146,1.648,0.0,0.0,0.0,0.0],
[0145] [1.648,-0.146,-0.146,-0.078,-0.837,-1.853,-0.146,2.064,1.313,-0.837,-0.184,0.0,0.0]]]
[0146] Other combinations of information can yield the corresponding three-dimensional matrix.
[0147] The encoded TCR and pMHC 3D matrices are input into a convolutional neural network (CNN) for training. The CNN used for TCR training consists of three convolutional layers (TC1, TC2, TC3), a fully connected layer, and an output layer. The fully connected layer is batch normalized and then output using the ReLU function. The CNN used for pMHC training consists of three convolutional layers (PC1, PC2, PC3), a fully connected layer, and an output layer. The fully connected layer is batch normalized and then output using the ReLU function. The output values of the fully connected layers from the TCR-trained CNN and the pMHC-trained CNN are combined and then input into the fully connected layer. After batch normalization and ReLU function, the output is passed through the fully connected layer and the softmax function. All convolutional kernels used in the model are (3,3) with a stride of 1, initialized using a Kaiming normal distribution. The CNN structure is as follows: Figure 2 As shown.
[0148] The model structure code implemented using pyTorch is as follows:
[0149] class CNNnet_dual(torch.nn.Module):
[0150] def__init__(self,nrow_pmhc,nrow_tcr):
[0151] super().__init__()
[0152] self.nrow_pmhc = nrow_pmhc
[0153] self.nrow_tcr = nrow_tcr
[0154] self.pmhc_conv1Nn.Sequential(nn.Conv2d(5,32,3,1,1),nn.BatchNorm2d(32),nn.ReLU());
[0155] self.pmhc_conv2Nn.Sequential(nn.Conv2d(32,64,3,1,1),nn.BatchNorm2d(64),nn.ReLU());
[0156] self.pmhc_conv3(nn.Sequential(nn.Conv2d(64,128,3,1,1),nn.BatchNorm2d(128),nn.ReLU());
[0157] self.pmhc_fc1Nn.Sequential(nn.Linear(128*self.nrow_pmhc*13,512),nn.BatchNorm1d(512),nn.ReLU());
[0158] self.tcr_conv1Nn.Sequential(nn.Conv2d(5,32,3,1,1),nn.BatchNorm2d(32),nn.ReLU());
[0159] self.tcr_conv2Nn.Sequential(nn.Conv2d(32,64,3,1,1),nn.BatchNorm2d(64),nn.ReLU());
[0160] self.tcr_conv3Nn.Sequential(nn.Conv2d(64,128,3,1,1),nn.BatchNorm2d(128),nn.ReLU())
[0161] self.tcr_fc1Nn.Sequential(nn.Linear(128*self.nrow_tcr*24,512),nn.BatchNorm1d(512),nn.ReLU());
[0162] self.combine_fc1Nn.Sequential(nn.Linear(1024,256),nn.BatchNorm1d(256),nn.ReLU());
[0163] self.combine_fc2=nn.Sequential(nn.Linear(256,128),nn.BatchNorm1d(128),nn.ReLU())
[0164] self.combine_fc3=nn.Sequential(nn.Linear(128,32),nn.BatchNorm1d(32),nn.ReLU())
[0165] self.combine_fc4=nn.Sequential(nn.Linear(32,2),nn.Softmax(dim=1))
[0166] def forward(self,x):
[0167] pmhc=torch.tensor([i[0]for i in x])
[0168] tcr=torch.tensor([i[1]for i in x])
[0169] pmhc=self.pmhc_conv1(pmhc)
[0170] pmhc=self.pmhc_conv2(pmhc)
[0171] pmhc=self.pmhc_conv3(pmhc)
[0172] pmhc=pmhc.view(pmhc.size()[0],-1)
[0173] pmhc=self.pmhc_fc1(pmhc)
[0174] tcr=self.tcr_conv1(tcr)
[0175] tcr=self.tcr_conv2(tcr)
[0176] tcr=self.tcr_conv3(tcr)
[0177] tcr=tcr.view(tcr.size()[0],-1)
[0178] tcr=self.tcr_fc1(tcr)
[0179] x=torch.cat([pmhc,tcr],dim=1)
[0180] x = self.combine_fc1(x)
[0181] x = self.combine_fc2(x)
[0182] x = self.combine_fc3(x)
[0183] x = self.combine_fc4(x)
[0184] return x
[0185] The results of the 5-fold cross-validation on a total of 6000 positive and negative samples are as follows: Figure 3 As shown. Figure 3 The horizontal axis represents the information combination used during training with paired TCR and pMHC data, as shown in the table below:
[0186]
[0187] The results show that the model has the best prediction performance when it uses information on CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, CDR3β, peptides, MHCα1contact, and MHCα2contact.
[0188] Example 2
[0189] Predictive performance testing:
[0190] 16,263 TCR-pMHC specific identification data were collected, including αβTCR and βTCR. The CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, and CDR3β of αβTCR were known, and the CDR1β, CDR2β, and CDR3β of βTCR were known, while information on CDR1α, CDR2α, and CDR3α was missing. Negative samples were generated 1:1 using a random combination method. We used CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, CDR3β, peptide, MHCα1contact, and MHCα2contact as input information to encode all 32,526 data points, using the same encoding method as in Example 1. The data was input into a convolutional neural network for 5-fold cross-validation, with the model structure the same as in Example 1, while the following restrictions were imposed on the test set:
[0191]
[0192] The results are as follows Figure 4As shown, the model has good predictive performance regardless of whether βTCR or αβTCR is unknown, and the model does not show obvious bias on βTCR, αβTCR, or mixed TCR data.
[0193] Example 3
[0194] On a mixed dataset containing 28,331 βTCR and αβTCR data points, two input information combinations were trained: the first, TRABcdr123-pep-MHC, used CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, CDR3β, peptide, MHCα1contact, and MHCα2contact information; the second, TRBcdr3-pep-MHC, used CDR3β, peptide, MHCα1contact, and MHCα2contact information. The data encoding and model structure were the same as in Example 1. Both models were tested on a test set consisting of 1,909 mixed data points. The ROC curves of the models on the test set are shown below. Figure 5 As shown. From Figure 5 It can be seen that the model trained using the TRABcdr123-pep-MHC combination achieved an AUC of 0.827 on the test set, while the model trained using the TRBcdr3-pep-MHC combination had an AUC of 0.697, indicating that the former has better prediction performance.
[0195] Example 4
[0196] Using CDR1α, CDR2α, CDR3α, CDR1β, CDR2β, CDR3β, peptides, MHCα1contact, and MHCα2contact as input information, a convolutional neural network model was trained on a mixed dataset containing 32,526 βTCR and αβTCR data points. The data encoding and convolutional neural network structure were the same as in Example 1. Subsequently, peripheral blood TCR data from 9 CMV seropositive and 10 CMV seronegative cases were collected, including TRBV gene and CDR3β amino acid sequences. CDR1β and CDR2β corresponding to the genotypes were obtained based on the annotation of the human TRBV gene on the IMGT website. For four CMV antigenic peptides (CRVLCCYVL, FRCPRRFCF, RPHERNGFTVL, and TPRVTGGGAM), the trained model was used to predict the antigen-specificity of the top 100 TCR clones in the TCR database for each sample. The MHC subtypes corresponding to the four antigenic peptides are: CRVLCCYVL, HLA-C*07:02; FRCPRRFCF, HLA-C*07:02; RPHERNGFTVL, HLA-B*07:02; and TPRVTGGGAM, HLA-B*07:02. Based on the MHC annotations on the IMGT website, amino acid residues at positions 59, 62, 63, 66, 67, 69, 70, 73, 74, 76, 77, 80, and 81 were identified as MHCα1 contact, and amino acids at positions 147, 150, 151, 152, 155, 156, 159, 163, 164, 167, and 168 were identified as MHCα2 contact. The percentage of TCR clones with a binding probability greater than 0.5 was statistically analyzed for each sample. The results are as follows: Figure 6 As shown, Figure 6 (A) represents the percentage of TCR clones with a predicted binding probability greater than 0.5 against the antigenic peptide CRVLCCYVL in both the CMV seropositive and negative groups. Figure 6 (B) shows the prediction results for the antigenic peptide FRCPRRFCF. Figure 6 (C) shows the prediction results for the antigenic peptide RPHERNGFTVL. Figure 6 (D) shows the prediction results for the antigenic peptide TPRVTGGGAM. The results indicate that among the top 100 clones, the proportion of CMV antigen-specific clones in CMV seropositive samples was significantly higher than that in CMV seronegative samples, indicating that the model's prediction results can distinguish whether a sample is infected with CMV virus.
[0197] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the invention.
Claims
1. A method for establishing a probability prediction model of human αβTCR and pMHC binding, characterized in that, Includes the following steps: (1) Obtain raw data on the specific binding of human αβTCR and pMHC, including CDR3β amino acid sequence, polypeptide sequence and / or TRAV gene type, TRBV gene type, CDR3α amino acid sequence, and MHC typing information; (2) Data preprocessing: The amino acid sequences of CDR1α, CDR2α, CDR1β, and CDR2β were obtained from the annotation information of TRAV and TRBV; the MHC amino acid sequences were obtained from the MHC annotation information; and the sequences were processed according to HLA-A... The protein structure of the 02:01 subtype was used to determine the positions of the amino acid residues with the side chains facing upward in its two α helices, and the amino acid residue sequences at the same positions on the two α helices of all other MHC subtypes were designated as MHCα1 contact and MHCα2 contact, respectively. (3) Generate negative samples: Define the data that specifically binds αβTCR and pMHC as positive samples, randomly combine TCR and pMHC in the positive samples, remove the data that are repeated with the positive samples, and randomly sample to generate negative samples according to a certain proportion to ensure that the frequency of all pMHC types in the negative samples is the same as that in the positive samples. (4) Data encoding: Arrange the amino acid sequences of TCR and pMHC into two-dimensional matrices respectively. Encode the two amino acid two-dimensional matrices using the five Atchley factors of the amino acids. Fill missing amino acid sequences with 0 to obtain a three-dimensional matrix. (5) Training and testing of convolutional neural networks: The three-dimensional matrices of TCR and pMHC are input into a convolutional neural network for training. Positive samples are labeled "bound" and negative samples are labeled "non-bound," and both positive and negative samples are used as training data. The convolutional neural network includes a convolutional neural network for TCR training and a convolutional neural network for pMHC training. The convolutional neural network for TCR training includes three convolutional layers, a fully connected layer, and an output layer. The fully connected layer is batch normalized and then output through the ReLU function. The convolutional neural network for pMHC training includes three convolutional layers, a fully connected layer, and an output layer. The fully connected layer is batch normalized and then output through the ReLU function. The output values of the fully connected layers of the convolutional neural networks for TCR training and pMHC training are merged and then input into the fully connected layer. After batch normalization and ReLU function, the output is passed through the fully connected layer and softmax function.
2. The method according to claim 1, characterized in that: The data preprocessing method in step (2) includes the following sub-steps: 2.1 The amino acid sequences of CDR1α, CDR2α, CDR1β and CDR2β were obtained based on the annotation information of human TRAV and TRBV genes on the IMGT website; 2.2 All MHC isotype amino acid sequences were obtained based on the annotation information of human MHC from the IMGT website; based on HLA-A... The protein structure of the 02:01 subtype determines the positions of the amino acid residues with upward-facing side chains in its two α-helices α1 and α2, and identifies the amino acid residue sequences at the same positions on the two α-helices of all other MHC subtypes as MHCα1 contact and MHCα2 contact, respectively.
3. The method according to claim 1, characterized in that: In step (3), the ratio of positive samples to negative samples is 1:0.5~1.
4. The method according to claim 1, characterized in that: The data encoding in step (4) includes the following sub-steps: 4.1 Arrange the CDR1α, CDR2α, CDR3α, CDR1β, CDR2β and CDR3β amino acid sequences of TCR into a two-dimensional matrix. The first row of the matrix is CDR1β+CDR2β amino acid residues, the second row is CDR3β amino acid residues, the third row is CDR3α amino acid residues, and the fourth row is CDR1α+CDR2α amino acid residues. Each row is 24 in length, and empty parts are filled with 0. 4.2 Arrange the amino acid sequences of the pMHC peptide, MHCα1 contact, and MHCα2 contact into a two-dimensional matrix. The first row of the matrix contains amino acid residues of the MHCα1 contact, the second row contains amino acid residues of the peptide, and the third row contains amino acid residues of the MHCα2 contact. Each row is 13 in length, and empty parts are filled with 0. 4.3 The two amino acid two-dimensional matrices above were encoded using five Atchley factors of amino acids, and missing amino acid sequences were filled with 0 to obtain a three-dimensional matrix.
5. The method for establishing a probability prediction model for the binding of human αβTCR and pMHC according to claim 1, characterized in that, The model is tested using cross-validation, and one or more of the following restrictions are applied: (1) CDR3α in the test set has not appeared in the training set; (2) CDR3β in the test set did not appear in the training set; (3) The data in the test set contains only one or more of CDR1β, CDR2β and CDR3β; (4) The data in the test set contains only one or more of CDR1α, CDR2α and CDR3α; (5) The data in the test set must include CDR1α, CDR2α, CDR3α, CDR1β, CDR2β and CDR3β.