A method for intelligent recognition of T-cell antigens at the atomic level

By combining graph convolutional networks and multi-head cross-attention mechanisms with atomic-level features of antigens, HLA, and TCR, an artificial intelligence model was constructed. This model solved the problem of insufficient accuracy in T-cell antigen recognition, enabling efficient screening of highly immunogenic T-cell antigens and promoting the development of tumor mRNA vaccines.

CN119864088BActive Publication Date: 2025-10-31HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178950.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-10-31
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Current technologies cannot accurately predict the binding and interaction between T cell antigens and HLA or TCR at the atomic level, resulting in insufficient T cell antigen recognition efficiency and accuracy, which limits the development of tumor mRNA vaccines.

Method used

By employing graph convolutional networks and multi-head cross-attention mechanisms, and combining atomic-level features of antigens, HLA, and TCRs, an artificial intelligence model is constructed through graph representation and digital encoding to predict the binding probability and contact sites of antigens with HLA or TCRs. The model is then fine-tuned using crystal structure data to screen for highly immunogenic T-cell antigens.

Benefits of technology

This technology enables efficient and accurate identification of T-cell antigens at the atomic level, improving the efficiency and accuracy of T-cell antigen screening and promoting the development of tumor mRNA vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864088B_ABST
    Figure CN119864088B_ABST
Patent Text Reader

Abstract

This invention discloses a method for intelligently recognizing T-cell antigens at the atomic level, comprising the following steps: representing antigens, HLA, and TCR as a graph with atoms as nodes and chemical bonds as edges; constructing an atomic-level intelligent prediction model for antigen-HLA binding to screen antigens that can be presented by HLA; constructing an atomic-level intelligent prediction model for antigen-TCR interaction to screen antigens that can be recognized by TCR; calculating the immunogenicity of antigens based on T-cell clone frequencies and identifying highly immunogenic antigens, thus forming an intelligent T-cell antigen recognition method. This invention overcomes the problem of existing technologies that only recognize T-cell antigens at the sequence or residue level, resulting in most candidate antigens failing to elicit an immune response, and can accurately screen for immunogenic neotumor antigens at the atomic level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, specifically a method for intelligently identifying T-cell antigens at the atomic level. Background Technology

[0002] T-cell antigens are short peptide epitopes that can be presented by major histocompatibility complex (MHC) molecules (also known as human leukocyte antigens, HLA in humans) and recognized by T-cell receptors (TCRs) on the surface of T cells, thus activating an immune response. T-cell antigen recognition is crucial for infection, autoimmune diseases, and tumor immunotherapy, especially in the development of tumor mRNA vaccines. A vast array of antigens exist in nature, and the TCR diversity generated by VDJ recombination is as high as 10. 16 -10 18 Furthermore, the binding or interaction between antigens and HLA or TCR is multispecific. Therefore, biological experimental techniques such as high-throughput sequencing, mass flow cytometry, and microfluidics cannot individually verify the immunogenicity of a vast number of potential antigens, thus limiting the recognition of T-cell antigens. Therefore, utilizing artificial intelligence to identify T-cell antigens is crucial for accelerating the development of mRNA vaccines.

[0003] Two key elements in using artificial intelligence to identify T-cell antigens are the prediction of antigen-HLA binding and the prediction of antigen-TCR interaction. Existing methods only use sequence or residue-level binding information to predict whether an antigen can be presented by HLA or recognized by the TCR. Most identified T-cell antigens fail to elicit a T-cell immune response. Recent research indicates that the atomic-level contact between antigen and TCR, especially the non-covalent interaction, is crucial for determining the activation of the T-cell immune response. During antigen-TCR dissociation, newly formed hydrogen bonds or salt bridges can prolong the contact time between immunogenic antigens and the TCR (reverse locking), while non-immunogenic antigens do not form new atomic-level non-covalent interactions with the TCR. Reverse locking has been used to precisely regulate TCR activity and has shown potential in cancer immunotherapy. However, the development of computational methods capable of predicting antigen-HLA binding and antigen-TCR interaction at the atomic level, and further inferring antigen immunogenicity, has not yet been explored.

[0004] Therefore, by fully utilizing information at the atomic level of antigens, HLA, and TCR, as well as the crystal structure of the complex, to develop computational methods to reveal the binding or interaction mechanisms between antigens and HLA or TCR at the atomic level, and by comprehensively considering both antigen presentation and T cell recognition processes to infer the immunogenicity of antigens, the screening efficiency and accuracy of T cell antigens will be greatly improved, providing a new breakthrough for the development of tumor mRNA vaccines. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing T-cell antigen recognition technologies by utilizing the structural information of antigen-HLA or TCR complexes, combined with artificial intelligence algorithms, to provide a method for efficiently and accurately predicting the binding or interaction of antigens with HLA or TCR at the atomic level. This integrates biotechnology and information technology to achieve efficient and precise recognition of massive amounts of T-cell antigens.

[0006] To achieve the above objectives, the technical solution adopted by this invention is: a method for intelligently recognizing T-cell antigens at the atomic level, comprising the following steps:

[0007] 1) Representation and feature encoding of antigen, HLA, and TCR data: The amino acid sequences of antigen, HLA, and TCR are represented as graphs with atoms as nodes and chemical bonds as edges using the structural formulas of amino acids. The graphs are then digitally encoded in combination with the physicochemical properties of atoms and chemical bonds.

[0008] 2) Construction of intelligent prediction model for antigen-HLA binding: Based on the graphical representation of antigen and HLA, an artificial intelligence model is established to predict the binding probability of antigen with HLA class I molecules or HLA class II molecules and the atomic-level contact sites, and to screen out antigens that can be presented by HLA.

[0009] 3) Construction of intelligent prediction model for antigen-TCR interaction: Based on the graph representation of antigen and TCR, an artificial intelligence model is established to predict the binding probability of antigen and TCR and the atomic-level contact site, and to screen out antigens that can be recognized by TCR.

[0010] 4) Construction of intelligent T-cell antigen recognition method: Based on the probability of antigen interacting with each TCR and the frequency of T-cell clones, calculate the comprehensive immunogenicity score of antigens that can be presented by HLA, screen for antigens with high immunogenicity, and obtain the method for intelligent recognition of T-cell antigens at the atomic level.

[0011] The representation and feature encoding of the above antigen, HLA, and TCR data include the following steps:

[0012] 1.1) Based on HLA polymorphic sites, amino acids at corresponding positions were extracted from the full-length sequences of HLA α and β chains and arranged in sequence to form HLA pseudo sequences. Based on the alignment results from the IMGT database, the CDR3 sequence was extracted from the full-length sequence of the TCR β chain.

[0013] 1.2) Analyze each amino acid in the input antigen sequence, HLA pseudo-sequence, and CDR3 sequence, and convert each amino acid into a molecular structure according to the general structural formula of 20 common amino acids;

[0014] 1.3) Amino acids are linked sequentially, with the carboxyl group of the previous amino acid and the amino group of the next amino acid undergoing dehydration condensation to form a peptide bond, and hydrogen atoms are added to the structure to meet the requirements of valence.

[0015] 1.4) Convert the molecular structure into a graph with atoms as nodes and chemical bonds as edges. The nodes are digitally encoded by atomic attributes (atom type, degree, hybridization state, formal charge, aromaticity, chirality), and the edges are digitally encoded by chemical bond attributes (bond type, conjugation, ring, stereoconformation) to obtain the graphical representation of antigens, HLA, and TCR.

[0016] The construction of the above-mentioned intelligent prediction model for antigen-HLA binding includes the following steps:

[0017] 2.1) Graph convolutional networks are used to extract atomic-level features of antigens and HLA, and atoms that play key roles in the interaction are identified from antigens and HLA respectively. Multi-head cross-attention mechanism is used to predict the binding reactivity of antigens and HLA, and the parameters of the prediction model are adjusted by training set and validation set data.

[0018] 2.2) Calculate the Euclidean distance between paired atoms of antigen and HLA using the resolved 3D structure of the antigen-HLA complex. Use this as prior information to fine-tune the model parameters, guide the graph convolutional network to identify atoms that play a key role in the binding process of antigen and HLA, and predict the contact probability between key atoms.

[0019] The construction of the above-mentioned intelligent prediction model for antigen-TCR interaction includes the following steps:

[0020] 3.1) Graph convolutional networks are used to extract atomic-level features of antigens and TCRs, and atoms that play a key role in the interaction are screened from antigens and TCRs respectively. Multi-head cross-attention mechanism is used to predict the binding reactivity of antigens to HLA, and the parameters of the prediction model are adjusted by training set and validation set data.

[0021] 3.2) The Euclidean distance between paired atoms of the antigen and TCR is calculated using the 3D structure of the antigen-TCR complex. This distance is used to fine-tune the model parameters with the resolved prior information, guide the graph convolutional network to identify atoms that play a key role in the interaction between the antigen and TCR, and predict the contact probability between key atoms.

[0022] Steps 2.1) and 3.1) above, which utilize graph convolutional networks to extract atomic-level features of antigens, HLA, and TCR, include the following steps:

[0023] S1) A graphical representation of each antigen, HLA, or TCR sequence is shown below. , Let i represent the initial characteristics of the i-th node. The features representing the edge between the i-th node and the j-th node are transformed by a single-layer neural network. The space is as shown in formula (1):

[0024] (1)

[0025] It is a learnable weight matrix. It is a non-linear activation function;

[0026] S2) The atomic features are processed by an L-layer graph convolutional network. In each iteration, message aggregation, information update and gated recurrent unit (GRU) are used in sequence, as shown in formulas (2), (3) and (4):

[0027] (2)

[0028] (3)

[0029] (4)

[0030] in This indicates that the i-th node is in the... l Layer aggregation characteristics, This indicates that the i-th node is in the... l Layer update characteristics This indicates that the i-th node after GRU reset and update is at position i. l Features of the layer This represents all neighboring nodes of the i-th node. This indicates a serial operation. , ;

[0031] S3) Each atom is scored using a neural network, as shown in formula (5):

[0032] (5)

[0033] , The hyperbolic tangent activation function is used to select the k highest-scoring atoms to represent the atomic-level features of antigen, HLA, and TCR using the Top-k pooling method.

[0034] S4) Calculate the pairwise interaction characteristic F between the Top-k atom of the antigen and the Top-k atom of HLA or TCR, as shown in formulas (6) and (7):

[0035] (6)

[0036] (7)

[0037] This represents the antigen atom features updated after L-layer graph convolution. This represents the HLA or TCR atomic features updated after L layers of graph convolution, where M represents the number of attention heads. This represents the average value of the multi-head attention weights;

[0038] S5) Based on the interaction characteristic F, the probability P (range 0-1) of the antigen binding to HLA or interacting with TCR is calculated using a multilayer perceptron, as shown in formula (8):

[0039] (8)

[0040] Based on the interaction probability, a binary discrimination can be made, namely whether the antigen binds to HLA or interacts with TCR, and the model parameters can be optimized using the cross-entropy loss function.

[0041] The fine-tuning of the model using the 3D structural information of the complex in steps 2.2) and 3.2) above includes the following steps:

[0042] S6) The positional information is added to the atomic feature through positional encoding, as shown in formulas (9), (10), and (11):

[0043] (9)

[0044] (10)

[0045] (11)

[0046] S7) Calculate the joint score of the antigen with paired atoms in HLA or TCR. The loss function based on the Pearson correlation coefficient makes the joint score negatively correlated with the actual distance, as shown in formulas (12), (13), and (14):

[0047] (12)

[0048] (13)

[0049] (14)

[0050] It is covariance. This is the standard deviation, where S and D represent the joint score and actual distance between the antigen and paired atoms in HLA or TCR, respectively. These are the coordinates of the atom in space;

[0051] S8) While ensuring that the model can identify key atoms, the model is further fine-tuned using the cross-entropy loss function to predict the contact probability of paired atoms.

[0052] The above-mentioned intelligent T-cell antigen recognition method is constructed by the following steps:

[0053] 4.1) The binding strength between antigen and HLA allele was predicted using an intelligent prediction model for antigen-HLA binding, and antigens with a binding strength greater than 0.5 were screened as candidates.

[0054] 4.2) The antigen-HLA binding intelligent prediction model is used to predict the probability of clonal interaction between the candidate antigen and each TCR in the TCR library, and the interaction probability is weighted according to the TCR clonal frequency to calculate the immunogenicity score of the i-th candidate antigen, as shown in formula (15):

[0055] (15)

[0056] Where p ij freq is the interaction strength between the i-th candidate antigen and the j-th TCR in the TCR library. j It is the cloning frequency of the j-th TCR;

[0057] 4.3) Candidate antigens are sorted according to their immunogenicity scores, and antigens with high immunogenicity scores are selected as T cell antigens. The contact sites between the antigen and HLA or TCR atoms are also output.

[0058] The beneficial effects of this invention are as follows: This invention utilizes atomic-level information on antigens, HLA, and TCRs, employing graph neural networks and multi-head attention mechanisms to accurately capture the atomic-level differences between T-cell antigens and non-T-cell antigens. By combining rare atomic-level complex crystal structure data with abundant sequence-level interaction data, it effectively integrates potential atomic-level contact information from sequence-level data, thereby accurately revealing the atomic-level contact between antigens and HLA or TCRs, solving the problems of limited crystal structure data and insufficient prediction accuracy of antigen-HLA binding or TCR interactions. The specific beneficial effects of this invention include the following:

[0059] I. Based on the general structural formula of amino acids and the dehydration condensation between amino acids, this invention constructs a graph representation with atoms as nodes and chemical bonds as edges, refining the representation of antigens, HLA, and TCR from the sequence level or residue level to the atomic level, laying the foundation for distinguishing the differences between T cell antigens and non-T cell antigens at the atomic level and the contacts between atoms.

[0060] Second, this invention encodes the graph through the biochemical properties of atoms and chemical bonds, extracts high-dimensional latent features through convolutional neural networks, and scores each atom using the neural network. This can identify a few atoms that play a key role in the interaction process, while reducing the amount of computation. It realizes an atomic-level antigen-HLA binding prediction model and an antigen-TCR interaction prediction model, thus improving the prediction accuracy.

[0061] Third, this invention first pre-trains the model on abundant sequence-level binding data to capture potential atomic-level contact information, and then fine-tunes the model using a small amount of atomic-level complex structure data. This solves the problem of limited atomic-level data and enables accurate prediction of the contact between antigens and HLA or TCR atoms through small-sample learning, providing a more comprehensive and clear perspective for meaningful biological discoveries such as key contact site detection and mutation effect prediction.

[0062] Fourth, this invention constructs an intelligent T-cell antigen recognition method, which makes full use of artificial intelligence methods and omics data analysis technology, comprehensively considers the two processes of antigen presentation by HLA and recognition by TCR, improves the accuracy of T-cell antigen recognition, will greatly accelerate the screening of immunogenic tumor neoantigens, and promote the development of clinical tumor mRNA vaccines. Attached Figure Description

[0063] Figure 1 This is a flowchart of a method for intelligently recognizing T-cell antigens at the atomic level according to the present invention;

[0064] Figure 2 This is a flowchart of deepAntigen, a T-cell antigen recognition framework based on graph convolutional networks;

[0065] Figure 3 This is a graph showing the results of deepAntigen's prediction of the binding performance of antigens to HLA molecules;

[0066] Figure 4 This is a graph showing the results of deepAntigen's prediction of the interaction performance between the antigen and the TCR;

[0067] Figure 5 This is a diagram showing the results of deepAntigen's identification of atomic-level critical contact sites;

[0068] Figure 6 This is a graph showing the results of an enzyme-linked spot assay to verify the immunogenicity of neoantigens in clinical cancer patients identified by deepAntigen. Detailed Implementation

[0069] The following detailed description of a specific embodiment of the present invention is provided in conjunction with the accompanying drawings. However, it should be understood that the scope of protection of the present invention is not limited to the specific embodiment.

[0070] See Figures 1-6 This invention discloses a method for intelligently recognizing T-cell antigens at the atomic level. For example... Figure 1 As shown, the specific implementation steps are as follows:

[0071] 1) Data Preprocessing: Sequence-level data of antigen-HLA binding, determined by experimental methods such as affinity measurement and mass spectrometry identification of naturally eluted peptides, were obtained from public resource databases such as IEDB as positive samples. This included antigen-HLA class I pairing data and antigen-HLA class II pairing data. Each entry in the data set included two characteristics: the antigen sequence and the HLA allele genotype. The corresponding protein sequences for the HLA alleles were found in the IMGT database. For each antigen, an antigen sequence of equal length was randomly cut from its corresponding gene as a negative sample. Furthermore, atomic-level antigen-HLA complex crystal structure data were obtained from protein data banks, and the atomic pair distance matrix between the antigen and HLA was calculated. Atomic pairs with a distance less than 5 Å were defined as contact pairs, and the remaining atomic pairs were non-contact pairs. Sequence-level data of antigen-TCR interactions validated through biological experiments were collected from databases such as VDJdb, PIRD, McPas-TCR, and ImmuneAccess as positive samples. Each entry in this data set included both the antigen sequence and the TCR sequence. For each antigen, non-interacting TCRs were randomly sampled from approximately 600 million inactive TCRs from healthy individuals to balance positive samples as negative samples. Furthermore, atomic-level antigen-TCR complex crystal structure data were obtained from a protein data bank, and the atomic pair distance matrix between the antigen and TCR was calculated. Atomic pairs with a distance less than 5 Å were defined as contact pairs, and the remaining pairs were defined as non-contact pairs.

[0072] 2) Representation and Feature Encoding of Antigen, HLA, and TCR Data: The amino acid sequences of antigens, HLA, and TCR are represented as graphs with atoms as nodes and chemical bonds as edges using the structural formulas of amino acids. The graphs are then digitally encoded based on the physicochemical properties of the atoms and chemical bonds, laying the foundation for identifying T-cell antigens using atomic-level features. Figure 2 As shown in a, b.

[0073] 2.1) Based on HLA polymorphic sites, amino acids at corresponding positions were extracted from the full-length sequences of HLA α and β chains and arranged in sequence to form HLA pseudo sequences. Based on the alignment results of the IMGT database, the complementarity determination region 3 (CDR3) sequence was extracted from the full-length sequence of TCR β chain.

[0074] 2.2) Analyze each amino acid in the input antigen sequence, HLA pseudo-sequence, and CDR3 sequence, and convert each amino acid into a molecular structure according to the general structural formula of 20 common amino acids;

[0075] 2.3) The amino acids are linked in sequence. The carboxyl group of the previous amino acid and the amino group of the next amino acid undergo dehydration condensation to form a peptide bond, and hydrogen atoms are added to the structure to meet the requirements of valence.

[0076] 2.4) The molecular structure is converted into a graph with atoms as nodes and chemical bonds as edges. The nodes are represented by attribute vectors of length 25 by splicing together the one-hot encoding of atomic attributes (atom type, degree, hybridization state, formal charge, aromaticity, chirality). The edges are represented by attribute vectors of length 11 by one-hot encoding of chemical bond attributes (bond type, conjugation, ring, stereoconformation). The graph representation of antigen, HLA, and TCR is obtained.

[0077] 2.5) During the model training phase, a portion of the nodes and their associated edges in the graph are randomly deleted with a 5% probability to augment the graph, expand the dataset, and prevent overfitting. This process is ignored during the model testing phase.

[0078] 3) Construction of intelligent prediction model for antigen-HLA binding: Based on the graphical representation of antigen and HLA, an artificial intelligence model is established to predict the binding probability of antigen with HLA class I molecules or HLA class II molecules and the atomic-level contact sites, and to screen out antigens that can be presented by HLA.

[0079] 3.1) Graph convolutional networks are used to extract atomic-level features of antigens and HLA, identifying atoms that play key roles in the interaction between antigens and HLA. Multi-head cross-attention mechanisms are used to predict the binding reactivity of antigens and HLA. The parameters of the prediction model are adjusted using training and validation set data, such as... Figure 2 As shown in a.

[0080] 3.1.1) A graphical representation of each antigen or HLA sequence is shown below. , This represents the 25-dimensional initial features of the i-th node. The 11-dimensional features representing the edge between the i-th node and the j-th node are transformed by a single-layer neural network. Space, h=128, as shown in formula (1):

[0081] (1)

[0082] It is a learnable weight matrix. It is a non-linear activation function;

[0083] 3.1.2) The atomic features are processed by an L=5 layer graph convolutional network. In each iteration, message aggregation, information update and gated recurrent unit (GRU) are used in sequence, as shown in formulas (2), (3) and (4):

[0084] (2)

[0085] (3)

[0086] (4)

[0087] in This indicates that the i-th node is in the... l Layer aggregation characteristics, This indicates that the i-th node is in the... l Layer update characteristics This indicates that the i-th node after GRU reset and update is at position i. l Features of the layer This represents all neighboring nodes of the i-th node. This indicates a serial operation. , ;

[0088] 3.1.3) Each atom is scored using a neural network, as shown in formula (5):

[0089] (5)

[0090] , This represents the hyperbolic tangent activation function, which uses Top-k pooling to select the k atoms with the highest scores to represent antigen and HLA atomic-level features. For antigen-HLA class I molecule binding prediction, the k value is set to 10; for antigen-HLA class II molecule binding prediction, the k value is set to 15.

[0091] 3.1.4) Calculate the pairwise interaction characteristic F between the Top-k atom of the antigen and the Top-k atom of HLA, as shown in formulas (6) and (7):

[0092] (6)

[0093] (7)

[0094] This represents the antigen atom features updated after L=5 layers of graph convolution. This represents the HLA or TCR atomic features updated after L layers of graph convolution, where M indicates that the number of attention heads is 4. This represents the average value of the multi-head attention weights;

[0095] 3.1.5) Based on the interaction characteristic F, the probability P of antigen binding to HLA (ranging from 0 to 1) is calculated using a multilayer perceptron, as shown in formula (8):

[0096] (8)

[0097] Based on the interaction probability, a binary discrimination can be made, namely whether the antigen binds to HLA, and the model parameters can be optimized using the cross-entropy loss function;

[0098] 3.1.6) The sequence-level antigen-HLA binding data were split into training and testing sets in an 8:2 ratio, and the model parameters were optimized on the training set using 10-fold cross-validation.

[0099] 3.2) Calculate the Euclidean distance between paired atoms of the antigen and HLA complex using the resolved 3D structure of the antigen-HLA complex. This distance serves as prior information to fine-tune model parameters, guiding the graph convolutional network to identify atoms playing a key role in the antigen-HLA binding process and predicting the contact probability between key atoms, such as... Figure 2 As shown in c.

[0100] 3.2.1) Add positional information to the atomic features through positional encoding, as shown in formulas (9), (10), and (11):

[0101] (9)

[0102] (10)

[0103] (11)

[0104] 3.2.2) Calculate the joint score of the antigen with paired atoms in HLA or TCR. The loss function based on the Pearson correlation coefficient makes the joint score negatively correlated with the actual distance, as shown in formulas (12), (13), and (14):

[0105] (12)

[0106] (13)

[0107] (14)

[0108] It is covariance. This is the standard deviation, where S and D represent the joint score and actual distance between the antigen and paired atoms in HLA or TCR, respectively. These are the coordinates of the atom in space;

[0109] 3.2.3) While ensuring that the model can identify key atoms, the model is further fine-tuned using the cross-entropy loss function to predict the contact probability of paired atoms;

[0110] 3.2.4) The atomic-level antigen-HLA crystal structure data were divided into training and validation sets using the leave-one-out method, and the model parameters were optimized through cross-validation.

[0111] 4) Construction of an intelligent prediction model for antigen-TCR interaction: Based on the graphical representation of antigen and TCR, an artificial intelligence model is established to predict the binding probability of antigen and TCR and the atomic-level contact sites, screening out antigens that can be recognized by TCR, such as... Figure 2 As shown in a and c. Antigen-TCR interaction prediction models were constructed using antigen-TCR sequence-level binding data according to specific implementation steps 3.1.1)-3.1.6); antigen-TCR atomic-level contact site prediction models were constructed using antigen-TCR atomic-level complex crystal structure data according to specific implementation steps 3.2.1)-3.2.4. The difference lies in the screening of k=20 key atoms for antigen and TCR.

[0112] 5) Construction of intelligent T-cell antigen recognition method: Based on the probability of antigen interacting with each TCR and the frequency of T-cell clones, calculate the comprehensive immunogenicity score of antigens that can be presented by HLA, screen for antigens with high immunogenicity, and obtain the method for intelligent recognition of T-cell antigens at the atomic level.

[0113] 5.1) The binding strength between antigen and HLA allele was predicted using an intelligent prediction model for antigen-HLA binding, and antigens with a binding strength greater than 0.5 were screened as candidates.

[0114] 5.2) The antigen-HLA binding intelligent prediction model is used to predict the probability of clonal interaction between candidate antigens and each TCR in the TCR library, and the interaction probability is weighted according to the TCR clonal frequency to calculate the immunogenicity score of the i-th candidate antigen, as shown in formula (15):

[0115] (15)

[0116] Where p ij freq is the interaction strength between the i-th candidate antigen and the j-th TCR in the TCR library. j It is the cloning frequency of the j-th TCR;

[0117] 5.3) Candidate antigens are sorted according to their immunogenicity scores, and antigens with high immunogenicity scores are selected as T cell antigens. The contact sites between the antigen and HLA or TCR atoms are output.

[0118] 6) Model performance evaluation: The model performance is evaluated using the test set. In this invention, both sequence-level binding prediction and atom-level contact prediction are binary classification tasks. A confusion matrix is ​​generated using the binding probability or contact probability, including true positives (TP), true negatives (TN), false positives (FP) and false negatives (FN). Then, the sensitivity and specificity can be calculated as shown in formulas (16) and (17).

[0119] (16)

[0120] (17)

[0121] Based on specificity and sensitivity, receiver operating characteristic (ROC) curves can be plotted, and the area under the ROC curve (AUROC) can be calculated to quantitatively evaluate the overall performance of the model.

[0122] In addition, precision can be calculated according to formula (18).

[0123] (18)

[0124] Recall can be calculated according to formula (16). Based on recall and precision, a precision-recall (PR) curve can be plotted, and the area under the PR curve (AUPR) can be calculated to quantitatively evaluate the performance in identifying positive samples.

[0125] Figure 3 a and b demonstrate the performance of predicting the binding reactivity of antigen-HLA class I and antigen-HLA class II molecules, respectively. Figure 4 The performance of antigen-TCR interaction prediction is demonstrated. Figure 5 This indicates that the predicted atomic-level contact sites reflect the actual binding conformation. Figure 6 This demonstrates that the neoantigens of breast and lung cancer patients identified by deepAntigen are immunogenic, as verified by an enzyme-linked spot (ELISPOT) assay.

[0126] In summary, this invention extends beyond sequences or residues to the atomic level, designing an AI-based computational method to identify T-cell antigens based on biological mechanisms. It comprehensively considers antigen presentation and T-cell recognition processes, providing a one-stop solution for identifying whether an antigen can activate an immune response. Furthermore, it offers high-precision atomic-level contact sites between antigens and HLA or TCR, overcoming the coarse-grained and low-precision problems of existing technologies in identifying T-cell antigens. This allows for efficient and accurate screening of T-cell antigens and extends to the atomic-precision prediction of interactions between small molecule drugs, nucleic acids, and proteins.

[0127] The above-disclosed embodiments are merely a few specific examples of the present invention. However, the embodiments of the present invention are not limited thereto, and any variations that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. A method for intelligently recognizing T-cell antigens at the atomic level, characterized in that, Includes the following steps: 1) Representation and feature encoding of antigen, HLA, and TCR data: The amino acid sequences of antigen, HLA, and TCR are represented as graphs with atoms as nodes and chemical bonds as edges using the structural formulas of amino acids. The graphs are then digitally encoded in combination with the physicochemical properties of atoms and chemical bonds. 2) Construction of intelligent prediction model for antigen-HLA binding: Based on the graphical representation of antigen and HLA, an artificial intelligence model is established to predict the binding probability of antigen with HLA class I or HLA class II molecules and the atomic-level contact sites, and to screen out antigens that can be presented by HLA; 3) Construction of intelligent prediction model for antigen-TCR interaction: Based on the graphical representation of antigen and TCR, an artificial intelligence model is established to predict the binding probability of antigen with TCR and the atomic-level contact sites, and to screen out antigens that can be recognized by TCR; 4) Construction of intelligent T cell antigen recognition method: Based on the probability of antigen interacting with each TCR and the frequency of T cell clones, calculate the comprehensive immunogenicity score of antigens that can be presented by HLA, screen for antigens with high immunogenicity, and obtain the method for intelligent recognition of T cell antigens at the atomic level. The construction of the intelligent prediction model for antigen-HLA binding includes the following steps: 1) Extract atomic-level features of antigen and HLA using graph convolutional networks, identify the atoms that play key roles in the interaction between antigen and HLA, use multi-head cross-attention mechanism to predict the binding reactivity of antigen and HLA, and adjust the parameters of the prediction model using training set and validation set data. 2) Calculate the Euclidean distance between paired atoms of antigen and HLA using the resolved 3D structure of the antigen-HLA complex. Use this as prior information to fine-tune the model parameters, guide the graph convolutional network to identify atoms that play a key role in the binding process of antigen and HLA, and predict the contact probability between key atoms. The construction of the intelligent prediction model for antigen-TCR interaction includes the following steps: 1) Extract atomic-level features of antigen and TCR using graph convolutional networks, screen out atoms that play key roles in the interaction from antigen and TCR respectively, use multi-head cross-attention mechanism to predict the binding reactivity of antigen and HLA, and adjust the parameters of the prediction model using training set and validation set data. 2) Calculate the Euclidean distance between paired atoms of the antigen and TCR using the 3D structure of the antigen-TCR complex. This distance is used to fine-tune the model parameters with the resolved prior information, guiding the graph convolutional network to identify atoms that play a key role in the antigen-TCR interaction and predict the contact probability between key atoms. The extraction of antigen, HLA, and TCR atomic-level features using the graph convolutional network includes the following steps: 1) Represent each antigen, HLA, or TCR sequence graphically and convert it into atomic features using a single-layer neural network; 2) The atomic features are processed through a multi-layer graph convolutional network. In each iteration, message aggregation, information update and gated recurrent unit (GRU) are used sequentially to obtain the updated atomic features. 3) Each atom is scored using a neural network, and the top-k pooling method is used to select the k atoms with the highest scores to represent the atomic-level features of antigens, HLA, and TCR; The method of predicting the interaction between antigens and HLA or TCR using a multi-head cross-attention mechanism includes the following steps: 1) Based on the atomic features of antigen, HLA or TCR after graph convolution update, calculate the pairwise interaction features between the Top-k atoms of antigen and the Top-k atoms of HLA or TCR using the average value of multi-head attention weights; 2) Based on this interaction characteristic, the probability of the antigen binding to HLA or interacting with TCR is calculated using a multilayer perceptron, and a binary discrimination is made accordingly, i.e. whether the antigen binds to HLA or interacts with TCR. At the same time, the model parameters are optimized using the cross-entropy loss function.

2. The method for intelligently recognizing T-cell antigens at the atomic level according to claim 1, characterized in that, The representation and feature encoding of antigen, HLA, and TCR data include the following steps: 1) Extract amino acids at corresponding positions from the full-length sequences of HLA α and β chains based on HLA polymorphic sites and assemble them into HLA pseudo sequences in chronological order. Extract the CDR3 sequence from the full-length sequence of TCR β chain based on the alignment results from the IMGT database. 2) Analyze each amino acid in the input antigen sequence, HLA pseudo-sequence, and CDR3 sequence, and convert each amino acid into a molecular structure according to the general structural formula of common amino acids; 3) Connect amino acids in sequence. The carboxyl group of the previous amino acid and the amino group of the next amino acid undergo dehydration condensation to form a peptide bond, and hydrogen atoms are added to the structure to meet the requirements of valence. 4) Convert the molecular structure into a graph with atoms as nodes and chemical bonds as edges. Nodes are digitally encoded through atomic properties, and edges are digitally encoded through chemical bond properties to obtain a graph representation of antigens, HLA, and TCR.

3. The method for intelligently recognizing T-cell antigens at the atomic level according to claim 1, characterized in that, The extraction of antigen, HLA, and TCR atomic-level features using graph convolutional networks includes the following steps: 1) A graphical representation of each antigen, HLA, or TCR sequence is shown below. , Let i represent the initial characteristics of the i-th node. The features representing the edge between the i-th node and the j-th node are transformed by a single-layer neural network. The space is as shown in formula (1): (1) It is a learnable weight matrix. It is a non-linear activation function; 2) The atomic features are processed by an L-layer graph convolutional network. In each iteration, message aggregation, information update, and gated recurrent unit (GRU) are used sequentially, as shown in formulas (2), (3), and (4): (2) (3) (4) in This indicates that the i-th node is in the... l Layer aggregation characteristics, This indicates that the i-th node is in the... l Layer update characteristics This indicates that the i-th node after GRU reset and update is at position i. l Features of the layer This represents all neighboring nodes of the i-th node. This indicates a serial operation. , ; 3) Each atom is scored using a neural network, as shown in formula (5): (5) , The hyperbolic tangent activation function is used to select the k highest-scoring atoms to represent the atomic-level features of antigen, HLA, and TCR using the Top-k pooling method.

4. The method for intelligently recognizing T-cell antigens at the atomic level according to claim 1, characterized in that, The method of predicting antigen-HLA or TCR interaction using a multi-head cross-attention mechanism includes the following steps: 1) Calculate the pairwise interaction characteristic F between the Top-k atom of the antigen and the Top-k atom of HLA or TCR, as shown in formulas (6) and (7): (6) (7) This represents the antigen atom features updated after L-layer graph convolution. This represents the HLA or TCR atomic features updated after L layers of graph convolution, where M represents the number of attention heads. This represents the average value of the multi-head attention weights; 2) Based on the interaction characteristic F, the probability P (ranging from 0 to 1) of the antigen binding to HLA or interacting with TCR is calculated using a multilayer perceptron, as shown in formula (8): (8) Based on the interaction probability, a binary discrimination can be made, namely whether the antigen binds to HLA or interacts with TCR, and the model parameters can be optimized using the cross-entropy loss function.

5. The method for intelligently recognizing T-cell antigens at the atomic level according to claim 1, characterized in that, The process of fine-tuning the model using the 3D structural information of the complex includes the following steps: 1) Add positional information to the atomic features through positional encoding, as shown in formulas (9), (10), and (11): (9) (10) (11) 2) Calculate the joint score of the antigen with paired atoms in HLA or TCR. The loss function based on the Pearson correlation coefficient makes the joint score negatively correlated with the actual distance, as shown in formulas (12), (13), and (14): (12) (13) (14) It is covariance. This is the standard deviation, where S and D represent the joint score and actual distance between the antigen and paired atoms in HLA or TCR, respectively. These are the coordinates of the atom in space; 3) While ensuring that the model can identify key atoms, the model is further fine-tuned using the cross-entropy loss function to predict the contact probability of paired atoms.

6. The method for intelligently recognizing T-cell antigens at the atomic level according to claim 1, characterized in that, The construction of the intelligent T-cell antigen recognition method includes the following steps: 1) Utilize an intelligent prediction model for antigen-HLA binding to predict the binding strength between antigen and HLA alleles, and screen antigens with high binding strength as candidates. 2) The antigen-HLA binding intelligent prediction model is used to predict the probability of clonal interaction between the candidate antigen and each TCR in the TCR library. The interaction probability is weighted according to the TCR clonal frequency, and the immunogenicity score of the i-th candidate antigen is calculated as shown in formula (15): (15) Where p ij freq is the interaction strength between the i-th candidate antigen and the j-th TCR in the TCR library. j It is the cloning frequency of the j-th TCR; 3) Sort candidate antigens according to their immunogenicity scores, select antigens with high immunogenicity scores as T cell antigens, and output the contact sites between the antigen and HLA or TCR atoms.