TCR-pMHC binding affinity prediction method based on reinforcement learning

By combining reinforcement learning with sequence, structural, and evolutionary information, the TCR-pMHC affinity prediction method solves the problems of insufficient structural information, missing evolutionary information, and poor data adaptability in existing technologies, achieving efficient and accurate prediction results and supporting applications such as cancer immunotherapy and vaccine design.

CN120895088APending Publication Date: 2025-11-04ZHONGYUAN ARTIFICIAL INTELLIGENCE IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510963614.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing TCR-pMHC affinity prediction methods suffer from insufficient utilization of structural information, lack of evolutionary information, poor data adaptability, and insufficient interpretability, resulting in insufficient prediction accuracy and difficulty in adapting to practical application scenarios.

Method used

We employ a reinforcement learning-based approach that combines sequence, structural, and evolutionary information. We generate structural embeddings using the ProFOLD model, train affinity ranking using Direct Preference Optimization, and interpret key identification residues using attention heatmaps to support sequence completion for missing data.

Benefits of technology

It significantly improves the accuracy and flexibility of affinity prediction, can adapt to complex data situations, reduces computational costs, provides highly interpretable prediction results, and supports applications in fields such as cancer immunotherapy and vaccine design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895088A_ABST
    Figure CN120895088A_ABST
Patent Text Reader

Abstract

The invention relates to the crossing field of bioinformatics and artificial intelligence, and provides a combination recognition technology based on reinforcement learning in order to overcome the defects of an existing TCR-pMHC affinity prediction method. According to the method, the sequence, the structure and evolutionary information are fused, a ProFOLD model is used for extracting structural features, and a Direct Prediction Optimization method is used for training, so that TCR-pMHC affinity prediction is realized. The technology supports partial data input, can output a visual result, and has accuracy, flexibility and interpretability. Compared with a traditional method, the prediction precision is greatly improved, and the calculation cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of bioinformatics and artificial intelligence, and specifically relates to a T cell receptor (TCR) and antigen peptide-MHC complex (pMHC) binding recognition technology based on reinforcement learning, including a calculation method and system for predicting the binding affinity of TCR and pMHC. BACKGROUND

[0002] T cells recognize and bind to antigen peptide pMHC presented by MHC molecules through their receptors TCR, triggering downstream immune response reactions. The TCR-pMHC recognition mechanism is a key link in vaccine development, cancer immunotherapy, allergy and autoimmune disease research. Due to the high specificity and individual variability of TCR and pMHC binding, accurate prediction of their binding relationship is crucial for understanding immune mechanisms and developing personalized treatment plans.

[0003] Existing TCR-pMHC affinity prediction methods can be mainly divided into the following three categories:

[0004] Clustering methods based on sequence similarity: such as TCRdist, GLIPH and iSMART, which are mainly used to find similar TCR patterns, but only achieve pattern clustering and cannot provide quantitative affinity evaluation, making it difficult to meet the demand for accurate prediction.

[0005] Deep learning-based sequence modeling methods: such as DeepTCR, NetTCR, PanPep, etc., which have improved prediction accuracy to some extent, but most methods only focus on the CDR3 region of TCR, ignoring structural information and the context of the entire sequence. Since TCR and pMHC binding not only depends on the sequence, but also relies on the structure, this limitation leads to a deviation between the prediction results and the actual situation.

[0006] Methods based on structural modeling: such as using AlphaFold2 combined with deep learning, which can utilize structural information to some extent, but has the problems of high computational cost and poor scalability. The complex calculation process and large computational resource requirements make it difficult to handle large-scale TCR sequence prediction tasks, limiting its use in actual large-scale application scenarios.

[0007] The above methods have the following main problems:

[0008] Insufficient use of structural information: most existing methods ignore the three-dimensional structure information of TCR and pMHC, and cannot fully consider the spatial interaction between molecules, resulting in insufficient accuracy of affinity prediction.

[0009] Evolution information loss: Residue co-evolution relationship is not considered, and the model lacks evolutionary background support. When facing new antigens, the model is difficult to generalize, and the reliability of the prediction result is low.

[0010] Data adaptability is poor: Most methods have strict requirements for input, such as providing complete double-chain TCR. In actual application, data loss is common, and these methods are difficult to adapt, limiting their application range.

[0011] Lack of interpretability: Unable to effectively identify key recognition residues or potential binding sites, which is not conducive to experimental verification and further mechanism research, hindering the practical application and development of the technology.

[0012] Therefore, there is an urgent need for a new modeling framework that can integrate sequence, structure and evolutionary information, improve prediction performance, and consider model flexibility and interpretability to meet the demand for accurate prediction of TCR-pMHC affinity in the fields of immunotherapy, vaccine development, etc. SUMMARY

[0013] The present application aims to provide a TCR-pMHC binding affinity prediction method based on reinforcement learning, which can be widely used in cancer immunotherapy, virus infection monitoring, vaccine design, TCR drug screening, individualized immune evaluation, etc. It provides core technical support for immune mechanism research and personalized treatment plan development, including the following steps:

[0014] Obtain the amino acid sequences of TCR alpha and beta chains, peptide segment sequences and MHC sequences, and perform one-hot encoding or amino acid index conversion;

[0015] Use BLAST or jackhmmer tools to retrieve homologous structure sequences similar to the current TCR-pMHC complex sequence from the PDB structure database, and input them into the ProFOLD model to generate single residue representation and structure embedding of residue representation;

[0016] The structure embedding and input sequence are sent to the sequence reconstruction module for mask language modeling, and the reconstructed amino acid distribution is output through Softmax, and the consistency score with the database sequence is calculated;

[0017] Using the Direct Preference Optimization method, construct positive and negative TCR-pMHC pairs, minimize the ranking loss function P(positive>negative)=sigmoid(w T ·(δh)+b) for affinity ranking training, where δh represents the difference between the two pairs in the embedding space, and w and b are learnable parameters;

[0018] Input TCR-pMHC sequence combination, output its affinity score, and explain key recognition residues and possible binding sites through attention heat map and residue sensitivity analysis.

[0019] Further, when part of the TCR alpha chain and beta chain is missing, the missing part is completed by a sequence completion module based on sequence similarity and evolutionary information.

[0020] Further, the ProFOLD model uses a 64-layer Evoformer module, with 3 iterations of structure recovery, and an output alignment resolution of 1 residue.

[0021] Further, the mask language modeling process includes masking part of the input sequence, learning sequence context information to mine potential evolutionary conserved sites, and reconstructing high-affinity candidate sequences.

[0022] Further, the other aspect provides a TCR-pMHC binding affinity prediction system based on reinforcement learning, comprising:

[0023] The data acquisition and preprocessing module is used to acquire the amino acid sequences of TCR alpha chain and beta chain, peptide sequence and MHC sequence, and perform encoding and completion processing;

[0024] The structure information extraction module is used to retrieve homologous structure sequences using BLAST or jackhmmer tools, and generate structure embeddings through the ProFOLD model;

[0025] The sequence and evolution analysis module is used to mask language model the structure embedding and the input sequence, and extract evolutionary information;

[0026] The reinforcement learning training module is used to perform affinity ranking training using the Direct Preference Optimization method;

[0027] The prediction and output module is used to input TCR-pMHC sequence combination, output affinity score, and generate visual results of key recognition residues and possible binding sites.

[0028] Further, the data acquisition and preprocessing module, when part of the TCR chain is missing, completes the missing part based on sequence similarity and evolutionary information.

[0029] Further, the sequence and evolution analysis module masks part of the input sequence, learns sequence context information to mine potential evolutionary conserved sites, and reconstructs high-affinity candidate sequences.

[0030] A computer device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the TCR-pMHC binding affinity prediction method based on reinforcement learning when executing the computer program.

[0031] A computer readable storage medium is also provided, which stores a computer program, wherein the computer program implements the steps of the TCR-pMHC binding affinity prediction method based on reinforcement learning when executed by a processor.

[0032] Advantages:

[0033] Improved prediction accuracy: By integrating structural and evolutionary information, the method fully considers various factors of molecular binding, making the affinity prediction more consistent with biological reality. On multiple test datasets, the prediction accuracy is significantly improved compared to traditional methods, providing a more reliable theoretical basis for immunotherapy and vaccine design.

[0034] Enhanced task adaptability: The training method based on reinforcement learning enables the model to have ranking and generation capabilities, which can flexibly handle different task requirements such as TCR candidate screening, affinity ranking, and new pairing design, improving the practicality and application value of the technology.

[0035] Expanded application scope: The method supports partial input and data missing, making it adaptable to complex and variable data situations in actual applications, and can be widely applied to large-scale TCR dataset analysis, playing an important role in virus infection monitoring and individualized immune evaluation.

[0036] Cost and efficiency advantages: High computational performance reduces the demand for computing resources, significantly improves training speed, and significantly reduces computational cost while ensuring prediction accuracy, which is conducive to large-scale promotion and industrial application of the technology. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 : System overall architecture schematic diagram. DETAILED DESCRIPTION

[0038] Embodiment one, embodiment specific process

[0039] (I) Input preparation and data encoding

[0040] In this embodiment, TCR-pMHC binding affinity prediction for antigens is selected as the application scenario. Python scripts are written through UniProt API to obtain the amino acid sequences of TCR alpha and beta chains, corresponding peptide sequences, and MHC sequences. During actual data acquisition, it was found that some TCR beta chain sequences were missing.

[0041] At this time, the partial sequence of the missing chain is aligned with the known homologous TCR beta chain sequence database using the Needleman-Wunsch algorithm. Set the matching score matrix, such as setting the matching score of the same amino acid to +2, the matching score of different amino acids to -1, and the gap penalty to -3. Calculate the optimal alignment path by dynamic programming to filter out the highest similarity sequence fragments from the homologous sequence database, and reasonably complete the missing TCR beta chain part.

[0042] After completion, the sequence is one-hot encoded using the numpy library of Python. Taking 20 common amino acids as an example, each amino acid is encoded as a 20-dimensional vector, where the corresponding amino acid position is set to 1 and the remaining positions are set to 0. For example, for the amino acid "alanine (A)", the encoding is [1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]. In this way, all the obtained sequences are converted into a unified input format, providing a basis for subsequent processing.

[0043] (II) Homologous sequence retrieval and structure embedding extraction

[0044] The BLAST tool is called using Python, and the e-value threshold is set to 1e-5 to retrieve similar homologous structure sequences in the PDB structure database to the current TCR-pMHC complex sequence. After retrieval, 10 similar homologous structure sequences with high similarity are obtained.

[0045] The 10 sequences are input into the pre-trained ProFOLD model. During the running process of the ProFOLD model, its 64-layer Evoformer module performs deep processing on the input sequence. In each structure recovery iteration, the module continuously optimizes the extraction of sequence structure information. After 3 structure recovery iterations, single residue representation and structure embedding information for residue representation are generated with 1residue output alignment resolution.

[0046] Single residue representation accurately reflects the preferred characteristics of each amino acid site in the structure, such as the tendency of certain amino acids to form specific secondary structures at specific positions; and residue representation reveals the spatial coevolution relationship between residue pairs in detail, such as the key role of the interaction between certain residue pairs in space in maintaining the stability of the complex structure. These structure embedding information provides accurate and rich structure feature data for subsequent analysis.

[0047] (III) Sequence reconstruction and evolution information extraction

[0048] The sequence reconstruction module is built using the TensorFlow framework. The previously obtained structure embedding information is input into the module along with the input sequence. In the sequence reconstruction module, a masked language modeling approach is used, and 30% of the amino acid positions in the input sequence are randomly selected for masking.

[0049] The context information around the masked positions is learned through a multi-layer neural network, and the reconstructed amino acid distribution is output using a Softmax function. For example, for a certain masked position, the model calculates the probability of the occurrence of 20 amino acids at that position based on the context information, and the amino acid with the highest probability is the predicted reconstructed amino acid.

[0050] At the same time, the reconstructed sequence is compared with the known high-affinity TCR-pMHC sequences in the database to calculate the consistency score. This score is used for subsequent preference learning guidance to ensure that the model can capture sequence features that conform to evolutionary laws during the learning process, and to mine potential evolutionary conserved sites, thereby reconstructing high-affinity candidate sequences.

[0051] (IV) Affinity ranking training

[0052] The Direct Preference Optimization (DPO) method is implemented based on the stable-baselines3 library in Python. First, construct a positive and negative TCR-pMHC pairing dataset. From a large number of existing TCR-pMHC pairing data, select 1000 pairs of data related to the current research antigen, including positive pairs and negative pairs (pairs with known low or no affinity).

[0053] During training, the model parameters w and b are continuously adjusted by minimizing the ranking loss function

[0054] P(positive>negative)=sigmoid(wT·(δh)+b)

[0055] The learning rate is set to 1e-4, and the number of training iterations is set to 10000. In each iteration, the difference δh between the positive and negative pairs in the embedding space is calculated, the loss value is calculated according to the loss function, and the parameters w and b are updated through the backpropagation algorithm to gradually enhance the model's accuracy in ranking TCR-pMHC affinity during training, and to converge towards the direction that conforms to the evolutionary constraints.

[0056] (V) Affinity prediction and visualization output

[0057] After the training is completed, a new set of TCR-pMHC sequences to be predicted is input into the model. The model calculates and outputs its affinity score, for example, the obtained TCR-pMHC binding score is 0.75, indicating that the TCR has a high binding possibility with the pMHC.

[0058] The attention heat map is generated using the matplotlib and seaborn libraries, which visually displays the residues that play a key role in binding. In the heat map, the darker the color, the higher the importance of the residue in the binding process. At the same time, through the residue sensitivity analysis algorithm, the influence degree of each residue on the affinity is calculated, and a visual structure contact map is generated. From the map, the specific contact sites and interaction modes of TCR and pMHC in structure can be clearly seen, providing intuitive evidence for in-depth study of the binding mechanism of the two.

[0059] II. Explanation of the innovation and efficiency

[0060] (1) Multi-information fusion innovation and efficiency

[0061] In this embodiment, by fusing the sequence information, three-dimensional structure information and residue co-evolution information of TCR-pMHC, compared with the traditional prediction method based on sequence only, the accuracy of affinity prediction has been significantly improved. In the test data set of this cancer antigen, the average prediction accuracy of the traditional method is 65%, and the accuracy of the present technology is improved by 37% through multi-information fusion.

[0062] This is because the introduction of structure information enables the model to consider the spatial interaction between molecules, and evolution information provides the model with more extensive background knowledge, so that the model can better generalize when facing new antigens and more accurately predict the binding affinity of TCR-pMHC, providing a more reliable theoretical basis for screening effective TCR in cancer immunotherapy, which helps to improve the targeting and effectiveness of treatment programs.

[0063] (2) Reinforcement learning application innovation and efficiency

[0064] The DPO method based on reinforcement learning is used for affinity ranking training, which gives the model strong ranking and generation capabilities. In the TCR candidate screening task, the traditional method needs to manually set complex screening rules, and the screening efficiency is low and the accuracy is poor. The present technology can automatically rank a large number of TCR candidates according to the affinity score, and quickly screen out TCRs with high affinity.

[0065] The screening efficiency of the present technology is significantly improved compared to traditional methods on the same candidate dataset. At the same time, the model also has the ability to generate new pairs, providing more potential effective solutions for vaccine design and TCR drug research, greatly improving the efficiency and success rate of related research and development work.

[0066] (III) Flexible data adaptability innovation and efficiency

[0067] In the input preparation stage, the present technology supports partial input and data missing conditions, such as successfully completing the missing TCR beta chain and completing subsequent prediction in this embodiment. In actual large-scale TCR dataset analysis, data missing conditions are common, and traditional methods often cannot effectively handle such data, resulting in a waste of a large amount of data.

[0068] (IV) Strong interpretability innovation and efficiency

[0069] Through attention visualization and mutation perturbation analysis, the present technology can accurately locate key recognition sites. In the attention heat map and structure contact map generated in this embodiment, researchers can clearly see which residues play a key role in the TCR-pMHC binding process.

[0070] This provides clear theoretical guidance for experimental verification, and in subsequent experimental design, researchers can target these key residues for research and verification, greatly accelerating the immune mechanism research process and promoting the development of related technologies.

[0071] (V) High-efficiency computing innovation and efficiency

[0072] Compared to structure prediction methods such as AlphaFold2, the present technology only introduces structure modules in the structure reference stage, greatly improving the training speed and reducing the computing cost. At the same time, the present technology significantly reduces the demand for computing resources and can run on ordinary workstations while ensuring prediction accuracy, while AlphaFold2 requires high-performance computing clusters. This makes the present technology more conducive to large-scale promotion and industrial application, and can meet the needs of industrial-level large-scale training and inference deployment, providing a strong guarantee for the commercial application of the technology.

Claims

1. A reinforcement learning-based method for predicting TCR-pMHC binding affinity, characterized in that, Includes the following steps: Obtain the amino acid sequences, peptide sequences, and MHC sequences of the TCR α and β chains, and perform one-hot encoding or amino acid indexing conversion; Using BLAST or jackhmmer tools, homologous structural sequences similar to the current TCR-pMHC complex sequence are retrieved from the PDB structural database and input into the ProFOLD model to generate single-residue representations and structural embeddings of residue representations. The structure is embedded and fed into the sequence reconstruction module for masked language modeling. The amino acid distribution is reconstructed by outputting Softmax and the consistency score with the database sequence is calculated. The Direct Preference Optimization method is used to construct positive and negative TCR-pMHC pairings and minimize the ranking loss function. Affinity ranking training is performed, where δh represents the difference between two pairs in the embedding space, and w and b are learnable parameters; Input TCR-pMHC sequence combinations, output their affinity scores, and interpret key identification residues and possible binding sites through attention heatmaps and residue sensitivity analysis.

2. The method according to claim 1, characterized in that, When some strands of the TCRα and β chains are missing, the missing parts are filled in by a sequence completion module based on sequence similarity and evolutionary information.

3. The method according to claim 1, characterized in that, The ProFOLD model uses a 64-layer Evoformer module, with a structure recycling iteration count of 3 and an output alignment resolution of 1 residue.

4. The method according to claim 1, characterized in that, The masked language modeling process includes masking a portion of the input sequence, using the model to learn sequence context information to mine potential evolutionarily conserved sites, and reconstructing high-affinity candidate sequences.

5. A TCR-pMHC binding affinity prediction system based on reinforcement learning, characterized in that, include: The data acquisition and preprocessing module is used to acquire the amino acid sequences, peptide sequences, and MHC sequences of the TCR α and β chains, and to perform encoding and completion processing. The structural information extraction module is used to retrieve homologous structural sequences using BLAST or Jackhmmer tools and generate structural embeddings using the ProFOLD model. The sequence and evolutionary analysis module is used to perform masked language modeling on the structural embedding and input sequence to extract evolutionary information; The reinforcement learning training module is used for affinity ranking training using the Direct Preference Optimization method. The prediction and output module is used to take TCR-pMHC sequence combinations as input, output affinity scores, and generate visualizations of key identification residues and possible binding sites.

6. The system according to claim 5, characterized in that, When a portion of the TCR chain is missing, the data acquisition and preprocessing module completes the missing portion based on sequence similarity and evolutionary information.

7. The system according to claim 5, characterized in that, The sequence and evolution analysis module performs masking on a portion of the input sequence, learns sequence context information to discover potential evolutionarily conserved sites, and reconstructs high-affinity candidate sequences.