Methods for identifying high avidity t cell receptors or t cells
By using a logistic regression model to analyze solvent accessibility and other characteristics of TCR amino acids, high avidity TCRs are identified, addressing the need for effective tumor infiltration and improving cancer immunotherapy.
Patent Information
- Application Number
- PCT/EP2024/088417
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
Current methods for identifying T cell receptors (TCRs) with high avidity against antigens are inadequate, particularly in the context of cancer immunotherapy, where the success of treatments like immune checkpoint blockade and adoptive cell therapy depends on the avidity of TCRs, which is not well understood, and there is a need for methods to identify high avidity TCRs for effective tumor infiltration.
A method involving selecting specific amino acids in the CDR3β of the TCR, determining their solvent accessibility, and using a logistic regression model to predict high avidity status, with additional characteristics like hydrophilicity and secondary structure, to identify TCRs with high avidity against antigens.
This approach allows for the identification of TCRs with high avidity, which are associated with better tumor infiltration and functional efficacy, enabling personalized cancer immunotherapy by enhancing the effectiveness of T cell-based treatments.
Smart Images

Figure EP2024088417_03072025_PF_FP_ABST
Abstract
Description
[0001] Docket No.: FR 084276.00403 METHODS FOR IDENTIFYING HIGH AVIDITY T CELL RECEPTORS OR T CELLS FIELD OF THE INVENTION This invention relates to methods for identifying a T cell receptor (TCR) with high avidityagainst an antigen or a T cell comprising the TCR. BACKGROUND OF THE INVENTION Tumor neoantigen recognition is a major factor in the success of clinical immunotherapiesand patients with high tumor mutational burden (TMB), and a significant number of neoantigensbenefit more from immune checkpoint blockade as well as adoptive cell therapy (ACT) using natural tumor-infiltrating lymphocytes (TILs). TIL therapy also showed success in low TMBcancers. Unlike T cell clones that recognize tumor-associated antigens (TAAs), which are self-antigens, those recognizing neoantigens are presumably not subject to negative thymic selection. Consequently, neoantigen-specific T cells may be of higher functionality and their superior antigensensitivity was recently demonstrated. T cell functionality is partly determined by the avidity ofTCRs for their cognate pMHC. Because this is dictated by structure, it is referred to as structuralavidity, and it is determined through the dissociation kinetics of monomeric pMHCs and TCRs.Conversely, sensitivity to antigen reflects the properties of the TCR but also the functional state of the cells and is thus referred to as functional avidity. Studies in mice and humans indicate that structural and functional avidities of CD8 T cells correlate and determine T cells performance. In TIL-ACT, clinical efficacy has been correlated with the persistence of adoptively transferred TIL clones in vivo, linked to unique gene expression patterns. However, how avidity affects tumor engraftment of tumor-specific T cells is presently not well understood. Yet, this is a key parameter affecting the success of T cell-based immunotherapy. Accordingly, there remains a need for methods and systems for identifying TCRs with highavidity with important applications in cancer treatment.SUMMARY OF THE INVENTION This disclosure addresses the need mentioned above in a number of aspects. In one aspect, this disclosure provides a method of identifying a T cell receptor (TCR) with high avidity against an antigen or an immune cell (e.g., T cell) comprising the TCR. In some embodiments, the method Docket No.: FR 084276.00403comprises: (a) selecting a set of amino acids in an amino acid sequence of a CDR3β of the TCR;(b) determining solvent accessibility of each amino acid of the set of amino acids; (c) inputting thesolvent accessibility of each amino acid of the set of amino acids into a logistic regression model, wherein the logistic regression model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to eachamino acid of the set of amino acids; (d) determining a probability value of a high avidity statusof the TCR as an output of the logistic regression model; and (e) identifying the TCR or the T cellas having high avidity against the antigen if the probability value is greater than or equal to a threshold value. The bias term b0 and the weights Wn were determined using a set of 48 with known Koff, maximizing the likelihood that each avidity prediction for these TCRs is correct. In some embodiments, the set of amino acids comprises one or more of Arg, Asn, Asp, Gly,Ile, Leu, and Phe. In some embodiments, the method comprises determining the probability value by: wherein p is the probability value, b0 represents a bias term, and R, N, D, G, I, L, and F areamino acids Arg, Asn, Asp, Gly, Ile, Leu, and Phe, respectively. In some embodiments, the threshold value is about 0.5. In some embodiments, the method comprises selecting a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 30%. In some embodiments, the method comprises inputting the solvent accessibility of thesubset of amino acids into the logistic regression model. In some embodiments, the methodcomprises determining the solvent accessibility of each amino acid of the set of amino acids as arelative solvent excluded surface area (SESA). In some embodiments, the SESA is determined bynormalizing surface area of an amino acid in the TCR against surface area of the amino acid in areference state. The reference state for a given amino-acid X is the SESA of X in the tripeptideGly-X-Gly. In some embodiments, the method comprises determining solvent accessibility of an amino acid based on a three-dimensional model of the TCR. Docket No.: FR 084276.00403 In some embodiments, the TCR with high avidity has a pMHC-TCR half-life of less than about 60 seconds. In some embodiments, the method comprises inputting into the logistic regression modelone or more additional characteristics of each amino acid of the set of amino acids. In someembodiments, the additional characteristics comprise hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of β-turn, molecular weight, relative mutability,normalized frequency of coil, average volume of buried residue, conformational parameter of β-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in β-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, α- helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in α-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed α-helix, heat capacity, free energy in α-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of β-sheet, unweighted, normalized frequency of β-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof. In some embodiments, the additional characteristics comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof. In some embodiments, the method comprises determining a nucleotide sequence of the TCR by deep sequencing. In some embodiments, the antigen comprises a neoantigen or a tumor-associated antigen. In some embodiments, the T cell comprises a CD8+ T cell or a CD4+ T cell. In another aspect, this disclosure provides a TCR or a T cell, which is identified accordingto the method described herein. Docket No.: FR 084276.00403 In another aspect, this disclosure provides a method of producing an engineered T cell withhigh avidity against an antigen. In some embodiments, the method comprises transfecting ortransducing a T cell with a nucleic acid molecule encoding a TCR identified according to the method described herein. In another aspect, this disclosure provides a method of treating cancer in a subject. In some embodiments, the method comprises administering to the subject a T cell identified according tothe method described herein or a T cell produced according to the method of claim described herein.In some embodiments, the cancer is selected from adrenal gland tumors, biliary cancer, bladder cancer, brain cancer, breast cancer, carcinoma, central or peripheral nervous system tissue cancer, cervical cancer, colon cancer, endocrine or neuroendocrine cancer or hematopoietic cancer, esophageal cancer, fibroma, gastrointestinal cancer, glioma, head and neck cancer, Li-Fraumeni tumors, liver cancer, lung cancer, lymphoma, melanoma, meningioma, multiple neuroendocrine type I and type II tumors, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteogenic sarcoma tumors, ovarian cancer, pancreatic cancer, pancreatic islet cell cancer, parathyroid cancer, pheochromocytoma, pituitary tumors, prostate cancer, rectal cancer, renal cancer, respiratory cancer, sarcoma, skin cancer, stomach cancer, testicular cancer, thyroid cancer, tracheal cancer, urogenital cancer, and uterine cancer. In yet another aspect, this disclosure provides a system for identifying a TCR with highavidity against an antigen or an immune cell (e.g., T cell) comprising the TCR. In someembodiments, the system comprises one or more processors configured to: (i) select a set of aminoacids in an amino acid sequence of a CDR3β of the TCR; (ii) determine solvent accessibility ofeach amino acid of the set of amino acids; (iii) input the solvent accessibility of each amino acidof the set of amino acids into a logistic regression model, wherein the logistic regression model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set ofamino acids and a weight value assigned to each amino acid of the set of amino acids; (iv)determine a probability value of a high avidity status of the TCR as an output of the logisticregression model; and (v) identify the TCR or the T cell as having high avidity against the antigenif the probability value is greater than or equal to a threshold value. Docket No.: FR 084276.00403 In some embodiments, the set of amino acids comprises one or more of Arg, Asn, Asp, Gly,Ile, Leu, and Phe. In some embodiments, the one or more processors are further configured todetermine the probability value by: wherein p is the probability value, b0 represents a bias term, and R, N, D, G, I, L, and F areamino acids Arg, Asn, Asp, Gly, Ile, Leu, and Phe, respectively. In some embodiments, the one or more processors are further configured to select a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 30%. In some embodiments, the one or more processors are further configured to input the solvent accessibility of the subset of amino acids into the logistic regression model. In some embodiments, the one or more processors are further configured to determine the solvent accessibility of each amino acid of the set of amino acids as a relative solvent excludedsurface area (SESA). In some embodiments, the SESA is determined by normalizing surface areaof an amino acid in the TCR against surface area of the amino acid in a reference state. In someembodiments, the one or more processors are further configured to determine solvent accessibility of an amino acid based on a three-dimensional model of the TCR. In some embodiments, the TCR with high avidity has a pMHC-TCR half-life of less than about 60 seconds. In some embodiments, the one or more processors are further configured to input into the logistic regression model one or more additional characteristics of each amino acid of the set of amino acids. In some embodiments, the additional characteristics comprise hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of β-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of β-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in β-strand region, side chain volume, amino acid composition of total proteins, average relative probability Docket No.: FR 084276.00403 of helix, α-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in α-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed α-helix, heat capacity, free energy in α-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of β-sheet, unweighted, normalized frequency of β-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof. In some embodiments, the additional characteristics comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof. The foregoing summary is not intended to define every aspect of the disclosure, andadditional aspects are described in other sections, such as the following detailed description. Theentire document is intended to be related as a unified disclosure, and it should be understood thatall combinations of features described herein are contemplated, even if the combination of featuresare not found together in the same sentence, or paragraph, or section of this document. Otherfeatures and advantages of the invention will become apparent from the following detaileddescription. It should be understood, however, that the detailed description and the specificexamples, while indicating specific embodiments of the disclosure, are given by way of illustration only, because various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description. BRIEF DESCRIPTION OF THE DRAWINGS FIGs. 1A, 1B, 1C, and 1D show structural avidity of neoantigen-, TAA-, and virus-specific CD8 T cells. FIG. 1A shows an example process for predicting structural avidity of TCR.The sequences in panel 2 are assigned SEQ ID Nos 330-389 and are also shown in Table 7. Thesequence in panel 3 is assigned with SEQ ID NO: 361, and the sequence in panel 4 is assignedwith SEQ ID NO: 361. FIG. 1B shows that neoantigen- and TAA-specific CD8 T cells werepurified from in vitro expanded CD8 T cells of melanoma, ovarian, lung or colorectal cancerpatients and virus-specific CD8 were isolated from healthy donors and cancer patients. After Docket No.: FR 084276.00403 single-cell cloning and expansion, individual clones were subjected to antigen sensitivity andstructural avidity measurements as well as TCR sequencing. Molecular modeling of pMHC-TCRinteractions were also performed. FIG. 1C shows cumulative structural avidities per classes ofantigen- (virus, TAAs, and neoantigen)-specific CD8 T cells. The number of clones is indicated inbrackets. Mann-Whitney tests were used to calculate P-values and are indicated when significant.FIG. 1D shows the coefficient of determination R2 of the regression analyses between pMHCbinding / stability and immunogenicity predictors values and the medians of T1 / 2 (s) or EC50 (M) of antigen-specific CD8 T cells. Pearson coefficients were calculated and mentioned when significant. Patients and clones are described in Tables 1-2. FIGs 2A, 2B, and 2C show association between structural avidity and tumor tropism. FIG.2A shows correlation between the structural profile of antigen-specific CD8 T cells and tropismin a steady state. FIG. 2B shows structural avidity of TAA- and neoantigen-specific PBLs andTILs. The number of clones is indicated in brackets. Mann-Whitney tests were used to calculateP-values and are indicated when significant. FIG. 2C shows relative frequency of clones 1, 3, and5 among UTP20-specific CD8 TILs (left) and PBLs (right). Structural avidity for each clone is also plotted. FIGs. 3A, 3B, 3C, 3D, 3E, and 3F show that tumor infiltration of high avidity clones isassociated to CXCR3 expression. FIG. 3A shows preferential tumor infiltration by high avidityclones and CXCR3-mediated tumor homing was validated by in vivo ACT in mice and in fourmelanoma patients receiving T cell therapy. FIG. 3B shows reactivity of V49I, WT and DMβ-transduced T cell was measured through IFN-γ secretion upon coculture with Me275 tumor cells(left), and monomeric pMHC-TCR dissociation kinetics of the three mutants were determinedusing reversible pMHC multimers (right). FIG.3C shows Me275 tumor growth in IL-2 NOG miceadoptively-transferred at day 7 post-tumor engraftment with 5x106 primary CD8 T cells transducedwith V49I, WT or DMβ NY-ESO-I157–165-specific TCRs. Log rank test was used to determine P-values. FIG. 3D shows P-values for tumor infiltration by CD8 T cells (cells / mm2) calculated byMann-Whitney test. Analyses were performed using Inform (Kramer, A. S. et al. Sci. Rep.8, 3418(2018)). FIG.3E shows Me275 tumor growth in IL-2 NOG mice adoptively transferred with 2x106DMβ-transduced primary CD8 T cells at day 5 and co-injected or not with anti-CXCR3 blockingantibody (100 μg at day 5 and day 10). Log rank test was used to determine P-value. FIG. 3Fshows quantitative measurement of tumor infiltration by CD8 T cells (cells / mm2) 10 days post- Docket No.: FR 084276.00403 ACT of 2x106DMβ-transduced primary CD8 T cells co-injected or not with anti-CXCR3 blocking antibody. Mann-Whitney test was used to calculate the P-value. FIGs.4A, 4B, 4C, and 4D show that tumor infiltration after ACT correlates with predictedstructural avidity inferred from TCR clustering analyses. FIG. 4A shows computational analysisof TCR features led to the establishment of a predictor of TCR avidity and its application onpatients’ TIL-ACT products allowed tracking of predicted low and high avidity TCRs in post-ACTtumor samples. FIG. 4B shows cumulative analysis for four melanoma patients of the percentageof predicted high avidity CD8 T cells in blood and tumor samples. Values for individual patients are plotted (grey and black) as well as the cumulative analysis for which the P-value was calculated as described in the method section. The number of clones is indicated below for each patientindividually. FIG. 4C shows monomeric pMHC-TCR dissociation kinetics of Jurkat cellstransfected with neoantigen KIF1BS918F-specific TCR#1 and TCR#2, respectively, predicted ashigh and low avidity TCRs. FIG. 4D shows autologous (Mel8) tumor growth in IL-2 NOG miceadoptively transferred at day 22 with 5x106 primary T cells transduced with KIF1BS918F-specificTCRs. Log rank test was used to determine the P-value. FIG. 5 shows amino acid frequencies in high and low structural avidity TCRs. Amino acidfrequencies of the CDR3β (excluding the first four and last three residues that may not contact thepeptide). High avidity is considered as T1 / 2 > 60s.DETAILED DESCRIPTION OF THE INVENTION This disclosure describes methods for identifying a T cell receptor (TCR) with high avidity against an antigen or an immune cell comprising the TCR. The success of cancer immunotherapy depends in part on the strength of antigen recognition by T cells. Relative to tumor-associated antigen (TAA)-specific T cells, neoantigen-specific T cells were of higher structural avidity and, consistently, were preferentially detected in tumors. Effective tumor infiltration in mice models was associated with high structural avidity and CXCR3 expression. Based on TCRsbiophysicochemical properties, an in silico model was developed and applied for predicting TCRstructural avidity and validated the enrichment in high avidity T cells in patients’ tumors. Theresults demonstrate a direct relationship between neoantigen recognition, T cell functionality, andtumor infiltration. Thus, the disclosed methods represent a rational approach for identifying potentT cells for personalized cancer immunotherapy. Docket No.: FR 084276.00403 Method of Identifying High Avidity TCRs and T Cells In one aspect, this disclosure provides a method of identifying a TCR with high avidity against an antigen or an immune cell comprising the TCR. In some embodiments, the method comprises: (a) selecting a set of amino acids in an amino acid sequence of the TCR; (b) determining solvent accessibility of each amino acid of the set of amino acids; (c) inputting the solvent accessibility of each amino acid of the set of amino acids into a supervised machine learning model, wherein the supervised machine learning model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids; (d) determining a probability value of a high avidity status of the TCR as an output of the supervised machine learning model; and (e) identifying theTCR or the T cell as having high avidity against the antigen if the probability value is greater thanor equal to a threshold value. The bias term b0and the weights Wn were determined using a set of 48 with known Koff, maximizing the likelihood that each avidity prediction for these TCRs is correct. In yet another aspect, this disclosure provides a system for identifying a TCR with high avidity against an antigen or an immune cell comprising the TCR. In some embodiments, the system comprises one or more processors configured to: (i) select a set of amino acids in an amino acid sequence of the TCR; (ii) determine solvent accessibility of each amino acid of the set of amino acids; (iii) input the solvent accessibility of each amino acid of the set of amino acids into a supervised machine learning model, wherein the supervised machine learning model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids; (iv) determine a probability value of a high avidity status of the TCR as an output of the supervised machine learning model; and (v) identify the TCR or the T cell as having high avidity against the antigen if the probability value is greater than or equal to a threshold value. As used herein, the supervised machine learning model may include, without limitation,analytical learning, artificial neural network, backpropagation, boosting (meta-algorithm), bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, gaussian process regression, group method of data handling, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), Docket No.: FR 084276.00403 multilinear subspace learning, naive bayes classifiers, maximum entropy classifiers, conditional random fields, nearest neighbor algorithms, probably approximately correct learning (PAC)learning, ripple down rules, support vector machines, minimum complexity machines (MCM),random forests, ensembles of classifiers, ordinal classification, data pre-processing, and statistical relational learning. In some embodiments, the supervised machine learning model may include classification- type supervised machine learning techniques or regression-type supervised machine learning techniques. Examples of classification-type supervised machine learning techniques include support vector machines (SVM), neural networks, naive bayes classifiers, decision trees, Adaptive Boosting (AdaBoost), Extreme Gradient Boosting (XGBoost), discriminant analysis, and nearest neighbors (kNN). Examples of regression-type supervised machine learning techniques includelinear regression, lasso regression, ridge regression, elasticnet regression, partial least squaresregression, polynomial regression, random forests, SVM, XGBoost, Adaboost, nonlinear regression, generalized linear models, decision trees, and neural networks. In some embodiments, the supervised machine learning model comprises a logistic regression model. As used herein, the term “logistic regression model” refers to the widely used statistical model that, in its basic form, may use a logistic function to model a binary dependent variable; many more complex extensions exist. In regression analysis, logistic regression (or logit regression) may estimate the parameters of a logistic model; it is a form of binomial regression. Mathematically, a binary logistic model has a dependent variable with two possible values, such as pass / fail, win / lose, alive / dead or healthy / sick; these are represented by an indicator variable,where the two values are labeled “0” and “1.” In the logistic model, the log-odds (the logarithm ofthe odds) for the value labeled “1” is a linear combination of one or more independent variables(“predictors”); the independent variables can each be a binary variable (two classes, coded by anindicator variable) or a continuous variable (any real value). The corresponding probability of thevalue labeled “1” can vary between 0 (certainly the value “0”) and 1 (certainly the value “1”),hence the labeling; the function that converts log-odds to probability is the logistic function, hence the name. Docket No.: FR 084276.00403 As used herein, the term “avidity” refers to an informative measure of the overall stabilityor strength of the pMHC-TCR complex. It is controlled by three major factors, including TCRepitope affinity, the valency of both the antigen and TCR, and the structural arrangement of theinteracting parts. These factors also define the specificity of TCR, that is, the likelihood that the particular antibody is binding to a precise antigen epitope. As used herein, the term “affinity” refers to the strength of interaction between TCR andantigen at single antigenic sites. Within each antigenic site, the variable region of the TCR “arm” interacts through weak non-covalent forces with antigen at numerous sites; the more interactions, the stronger the affinity. In some embodiments, the avidity of a TCR may be defined by either pMHC-TCRassociation (Kassoc or Ka) or dissociation kinetics (Kdissoc or Kd). The term “Kassoc” or “Ka,” as usedherein, refers to the association rate of a particular antibody-antigen interaction, whereas the term“Kdissoc” or “Kd,” as used herein, refers to the dissociation rate of a particular antibody-antigeninteraction. The term “KD,” as used herein, refers to the dissociation constant, which is obtainedfrom the ratio of Kd to Ka (i.e., Kd / Ka) and is expressed as a molar concentration (M). KD valuesfor TCRs can be determined using methods well established in the art. A method for determiningthe KD of a TCR is by using surface plasmon resonance, such as the biosensor system of Biacore®,or Solution Equilibrium Titration (SET) (see Friguet B et al. (1985) J. Immunol Methods; 77(2):305-319, and Hanel C et al. (2005) Anal Biochem; 339(1): 182-184).In some embodiments, the TCR with high avidity has a pMHC-TCR half-life of less thanabout 60 seconds (e.g., 5, 10, 15, 20, 25, 30, 40, 45, 55, 60, 65, 70, 75, or 80 seconds).As used herein, “T cell receptor (TCR)” is also called a T cell antigen receptor. A T cellreceptor refers to a receptor recognizing an antigen expressed on a cell membrane of a T cell that plays a central role in the immune system. TCRs have an α chain, β chain, γ chain, and δ chain, with which an αβ or γδ dimer is constituted. TCRs consisting of the combination of the former are called αβ TCRs, and TCRs consisting of the combination of the latter are called γδ TCRs. T cells having such TCRs are respectively called αβ T cells and γδ T cells. The TCRs are structurallyvery similar to a Fab fragment of an antibody produced by B cells and recognize antigen moleculesbound to an MHC molecule. Since a TCR gene of a mature T cell has undergone gene rearrangement, an individual has highly diverse TCRs that enable recognition of various antigens. Docket No.: FR 084276.00403 TCRs also form a complex by binding to a non-variable CD3 molecule at the cell membrane. CD3 has an amino acid sequence called ITAM (immunoreceptor tyrosine-based activation motif) in the intracellular region. This motif is considered to be involved in intracellular signaling. Each TCR chain is comprised of a variable domain (V) and a constant domain (C). A constant domain has a short cytoplasm section penetrating the cell membrane. A variable domain is present outside the cell and binds to an antigen-MHC complex. A variable domain has three hypervariable domains or regions called complementarity-determining regions (CDRs), which bind to an antigen-MHC complex. The three CDRs are called CDR1, CDR2, and CDR3. TCR gene rearrangement is similar to the process of B cell receptors known as immunoglobulins. For gene rearrangement of αβ TCRs, VDJ recombination of β chain is performed, followed by VJ recombination of an α chain. When the α chain is rearranged, the gene of the δ chain is deleted from the chromosome. Thus, a T cell having an αβ TCR would never have a γδ TCR simultaneously. In contrast, a signal via a γδ TCR in a T cell having the TCR suppresses the expression of β chain, so that a T cell having a γδ TCR would never have an αβ TCR simultaneously. In some embodiments, the immune cell comprises a lymphocyte. In some embodiments, the lymphocyte comprises a T cell or a natural killer (NK) cell. In some embodiments, the T cell comprises a CD8+ T cell or a CD4+ T cell. In some embodiments, the T cell comprises a human T cell. In some embodiments, the antigen may be a neoantigen or a tumor-associated antigen(TAA). As used herein, the term “antigen” is a molecule and / or substance that can generate peptidefragments that are recognized by a TCR and / or induces an immune response. An antigen maycontain one or more “epitopes.” In some embodiments, the antigen has several epitopes. Anepitope is recognized by a TCR, an antibody or a lymphocyte in the context of an MHC molecule. As used herein, the term “neoantigen” is an antigen that has at least one alteration thatmakes it distinct from the corresponding wild-type, parental antigen, e.g., via mutation in a tumor cell or post-translational modification specific to a tumor cell. A neoantigen can include a polypeptide sequence or a nucleotide sequence. A mutation can include a frameshift or non- frameshift indel, missense or nonsense substitution, splice site alteration, genomic rearrangement, or gene fusion, or any genomic or expression alteration giving rise to a neoORF. A mutation can also include a splice variant. Post-translational modifications specific to a tumor cell can include Docket No.: FR 084276.00403 aberrant phosphorylation. Post-translational modifications specific to a tumor cell can also include a proteasome-generated spliced antigen (Liepe et al., Science. 2016 Oct 21;354(6310):354-358).As used herein the term “tumor neoantigen” is a neoantigen present in a subject’s tumor cell ortissue but not in the subject’s corresponding normal cell or tissue. As used herein, the terms “tumor-associated antigen,” “TAA,” and “cancer antigen” referto any molecule (e.g., protein, peptide, lipid, carbohydrate, etc.) solely or predominantly expressed or over-expressed by a tumor cell and / or a cancer cell, such that the antigen is associated with tumor and / or cancer. The TAA / cancer antigen can also be expressed by normal, non-tumor, or non-cancerous cells. However, in such a situation, the expression of the TAA / cancer antigen bynormal, non-tumor, or non-cancerous cells is, in some embodiments, not as robust as theexpression of the TAA / cancer antigen by tumor and / or cancer cells. Thus, in some embodiments,the tumor and / or cancer cells overexpress the TAA and / or express the TAA at a significantlyhigher level as compared to the expression of the TAA by normal, non-tumor, and / or non- cancerous cells. In some embodiments, the phosphopeptides are fragments of TAAs or TAAsthemselves. The TAA can be an antigen expressed by any cell of any cancer or tumor, includingthe cancers and tumors described herein. The TAA can be a TAA of only one type of cancer or tumor, such that the TAA is associated with or characteristic of only one type of cancer or tumor. Alternatively, the TAA can be characteristic of more than one type of cancer or tumor. For example, the TAA can be expressed by both breast and prostate cancer cells and not expressed at all by normal, non-tumor, or non-cancer cells. In some embodiments, a set of amino acids or a subset of amino acids of a TCR comprises3 to 20 (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) or more amino acids. Insome embodiments, the set of amino acids or the subset of amino acids of a TCR is contained inthe CDR3β of the TCR. In some embodiments, the set of amino acids comprises one or more of Arg, Asn, Asp, Gly, Ile, Lue, and Phe, or a conservative substitution thereof. Suitable conservative substitutions of amino acids are known to those of skill in this art and may be made generally without altering the biological activity of the resulting molecule (See, e.g., Watson, et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. Co., p.224). In some embodiments, the method comprises determining the probability value by: Docket No.: FR 084276.00403 wherein p is the probability value, b0 represents a bias term, and R, N, D, G, I, L, and F areamino acids Arg, Asn, Asp, Gly, Ile, Leu, and Phe, respectively. The bias term b0 and the weightsWn were determined using a set of TCRs (with, e.g., 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, ormore TCRs). In some embodiments, the threshold value is about 0.4, 0.5, 0.6, 0.7, 0.8, or 0.9. In some embodiments, the threshold value is about 0.5 (e.g., 0.5). In some embodiments, the method comprises selecting a subset of amino acids from theset of amino acids that have solvent accessibility higher than about 15%, 20%, 25%, 30%, 35%,40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80%. In some embodiments, the methodcomprises selecting a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 30% (e.g., 30%). In some embodiments, the method comprises inputting the solvent accessibility of the subset of amino acids into the supervised machine learning model (e.g., logistic regression model). In some embodiments, the method comprises determining the solvent accessibility of each amino acid of the set of amino acids as a relative solvent excluded surface area (SESA). In some embodiments, the SESA is determined by normalizing surface area of an amino acid in the TCR against surface area of the amino acid in a reference state. In some embodiments, the methodcomprises determining solvent accessibility of an amino acid based on a three-dimensional (3D)model of the TCR. The TCR 3D model may be generated by any suitable methods (e.g., X-raycrystallography, NMR, cryo-EM, homology modeling, or de novo folding).In some embodiments, SESA can be computed with the MSMS package of the UCSFChimera software, as described by Goddard, T. D. et al. (Goddard, T. D. et al. Protein Sci. Publ.Protein Soc. 27, 14–25 (2018)). SESA can also be calculated by normalizing the surface area ofthe residue in the TCR of interest by its surface area in a reference state (Bendell, C. et al. BMCBioinformatics 15, 82 (2014)). In some embodiments, the method may further include inputting into the supervisedmachine learning model (e.g., logistic regression model) one or more additional characteristics of each amino acid of the set of amino acids. Docket No.: FR 084276.00403 In some embodiments, the additional characteristics comprise hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of β-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of β-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in β-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, α-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in α-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed α-helix, heat capacity, free energy in α-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of β-sheet, unweighted, normalized frequency of β-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof. In some embodiments, the additional characteristics comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof. In some embodiments, the method comprises determining a nucleotide sequence of theTCR by sequencing, e.g., deep sequencing or ultra-deep sequencing. Deep sequencing yields aunique genetic fingerprint that can be used to identify a person, and a trove of predictors of genetic medical diseases. Deep sequencing to identify epigenetic events including changes in DNA methylation and RNA expression can reveal the history and impact of environmental exposures. Ultra-deep sequencing is the sequencing of amplicons at a high depth of coverage with the goal of identifying the common and rare sequence variations. With sufficient depth of coverage, ultra- deep sequencing has the ability to fully characterize rare sequence variants down to less than 1%.Ultra-deep sequencing has been used to detect low- frequency HIV drug-resistant mutations oridentify rare somatic mutations in complex cancer samples. For tests such as non-invasive blood tests, the frequency of biomarker mutation could be lower than 1%. Docket No.: FR 084276.00403 High Avidity TCRs and T Cells In another aspect, this disclosure provides a TCR or antigen-binding fragment, or animmune cell, which is identified according to the methods described herein. In some embodiments,the immune cell comprises a lymphocyte. In some embodiments, the lymphocyte comprises a T cell or a natural killer (NK) cell. In some embodiments, the T cell comprises a CD8+ T cell or a CD4+ T cell. In some embodiments, the T cell comprises a human T cell. Lymphocytes are one subtype of white blood cells in the immune system. In some embodiments, lymphocytes may include tumor-infiltrating immune cells. Tumor-infiltrating immune cells consist of both mononuclear and polymorphonuclear immune cells (i.e., T cells, B cells, natural killer cells, macrophages, neutrophils, dendritic cells, mast cells, eosinophils, basophils, etc.) in variable proportions. In some embodiments, lymphocytes may include tumor- infiltrating lymphocytes (TILs). TILs are white blood cells that have left the bloodstream and migrated towards a tumor. TILs can often be found in the tumor stroma and within the tumor itself.In some embodiments, TILs are “young” T cells or minimally cultured T cells. In someembodiments, the young cells have a reduced culturing time (e.g., between about 22 to about 32 days in total). In some embodiments, the lymphocytes express CD27. In some embodiments, the lymphocytes may be autologous, allogeneic, syngeneic, or xenogeneic with respect to the subject. In some embodiments, the lymphocytes are autologous in order to reduce an immunoreactive response against the lymphocyte when reintroduced into the subject for immunotherapy treatment. In some embodiments, the T cells are CD8+ T cells. In some embodiments, the T cells are CD4+ cells. In some embodiments, the NK cells are CD 16+ CD56+ and / or CD57+ NK cells. NKs are characterized by their ability to bind to and kill cells that fail to express “self MHC / HLA antigens by the activation of specific cytolytic enzymes, the ability to kill tumor cells or other diseased cells that express a ligand for NK activating receptors, and the ability to release protein molecules called cytokines that stimulate or inhibit the immune response. Also within the scope of this disclosure is a variant of the TCR identified according to themethods disclosed herein and an immune cell comprising the variant of the TCR. Docket No.: FR 084276.00403 As used herein, the term “variant” refers to a first molecule that is related to a secondmolecule (also termed a “parent” molecule). The variant molecule can be derived from, isolatedfrom, based on or homologous to the parent molecule. A “functional variant” of a protein, as usedherein, refers to a variant of such protein that retains at least partially the activity of that protein.Functional variants may include mutants (which may be insertion, deletion, or replacement mutants), including polymorphs, etc. Also included within functional variants are fusion products of such protein with another, usually unrelated, nucleic acid, protein, polypeptide, or peptide. Functional variants may be naturally occurring or may be man-made. In some embodiments, a variant of a TCR may include one or more conservativemodifications. The TCR variant with one or more conservative modifications may retain the desired functional properties, which can be tested using the functional assays known in the art. As used herein, the term “conservative sequence modifications” refers to amino acidmodifications that do not significantly affect or alter the binding characteristics of the protein containing the amino acid sequence. Such conservative modifications include amino acid substitutions, additions, and deletions. Modifications can be introduced by standard techniques known in the art, such as site-directed mutagenesis and PCR-mediated mutagenesis. Conservativeamino acid substitutions are ones in which the amino acid residue is replaced with an amino acidresidue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art. These families include: amino acids with basic side chains (e.g., lysine, arginine, histidine); acidic side chains (e.g., aspartic acid, glutamic acid); uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine, tryptophan); nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine); beta-branched side chains (e.g., threonine, valine, isoleucine); and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine) includes one or more conservative modifications. The Cas protein with one or more conservative modifications may retain the desired functional properties, which can be tested using the functional assays known in the art. As used herein, the percent homology between two amino acid sequences is equivalent to the percent identity between the two sequences. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., % homology = # of identical positions / total # of positions x 100), taking into account the number of gaps, and the Docket No.: FR 084276.00403 length of each gap, which need to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm, as described in the non-limiting examples below. The percent identity between two amino acid sequences can be determined using thealgorithm of E. Meyers and W. Miller (Comput. Appl. Biosci., 4:11-17 (1988)), which has beenincorporated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. In addition, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch (J. Mol. Biol. 48:444-453(1970)) algorithm, which has been incorporated into the GAP program in the GCG softwarepackage (available at www.gcg.com), using either a Blossum62 matrix or a PAM250 matrix, and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6. The term “homolog” or “homologous,” when used in reference to a polypeptide, refers toa high degree of sequence identity between two polypeptides, or to a high degree of similarity between the three-dimensional structure or to a high degree of similarity between the active site and the mechanism of action. In some embodiments, a homolog has a greater than 60% sequence identity, and more preferably greater than 75% sequence identity, and still more preferably greaterthan 90% sequence identity, with a reference sequence. The term “substantial identity,” as appliedto polypeptides, means that two peptide sequences, when optimally aligned, such as by the programs GAP or BESTFIT using default gap weights, share at least 75% sequence identity. Apeptide or polypeptide “fragment” as used herein refers to a less than full-length peptide,polypeptide or protein. For example, a peptide or polypeptide fragment can have at least about 3, at least about 4, at least about 5, at least about 10, at least about 20, at least about 30, at least about 40 amino acids in length, or single unit lengths thereof. For example, fragment may be 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or more amino acids in length. There is no upper limit to the size of a peptide fragment. However, in some embodiments, peptide fragments can be less than about 500 amino acids, less than about 400 amino acids, less than about 300 amino acids or less than about 250 amino acids in length. Also within the scope of this disclosure are the variants, mutants, and homologs with significant identity to the TCR. For example, such variants and homologs may have sequences with at least about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, Docket No.: FR 084276.00403 about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% sequenceidentity with the sequences of TCRs described herein.In another aspect, this disclosure provides a method of producing an engineered T cell with high avidity against an antigen. In some embodiments, the method comprises transfecting or transducing a T cell with a nucleic acid molecule encoding a TCR identified according to themethod described herein. In some embodiments, the method comprises: (a) isolating a plurality ofimmune cells from a subject; (b) transfecting or transducing the plurality of immune cells with the vector described above; and (c) optionally expanding the transfected cells. Method of Treatments In another aspect, this disclosure provides a method of treating cancer in a subject. In someembodiments, the method comprises administering to the subject an immune cell identifiedaccording to the method described herein or an immune cell produced according to the method ofclaim described herein or a composition comprising the immune cell. In some embodiments, themethod comprises administering to the subject a therapeutically effective amount of immune cells identified according to the method described herein or a therapeutically effective amount of immune cells produced according to the method of claim described herein. In some embodiments, the method comprises administering to the subject an additional agent, such as an anti-cancer or anti-tumor agent. In some embodiments, the immune cell comprises a lymphocyte. In some embodiments, the lymphocyte comprises a T cell or a natural killer (NK) cell. In some embodiments, the T cell comprises a CD8+ T cell or a CD4+ T cell. In some embodiments, the T cell comprises a human T cell. In some embodiments, the immune cells can used in adoptive T cell therapy (ADT).Generally, adoptive T cell therapy relies on the in vitro expansion of endogenous, cancer-reactiveT cells. These T cells can be harvested from cancer patients, manipulated, and then reintroduced into the same or a different patient as a mechanism for generating productive tumor immunity. Insome embodiments, CD8+ cytotoxic T lymphocytes can be used in adoptive T cell therapy. CD4+ Docket No.: FR 084276.00403 T cells can also play an important role in maintaining CD8+ cytotoxic function, and transplantationof tumor reactive CD4+ T cells has been associated with some efficacy in metastatic melanoma.T cells used in adoptive therapy can be harvested from a variety of sites, including peripheral blood, malignant effusions, resected lymph nodes, and tumor biopsies. Although T cells harvested from the peripheral blood are easier to obtain technically, TILs obtained from biopsies may contain a higher frequency of tumor-reactive cells. Once harvested, T cells can be transfected with a vector as described above. In some embodiments, a TCR or antigen-binding fragment as disclosed has antigen specificity for an antigen that is characteristic of a disease or disorder. The disease or disorder can be any disease or disorder involving an antigen, such as but not limited to a tumor and / or a cancer, an infectious disease, or an autoimmune disease. In some embodiments, the subject is a human. In some embodiments, the subject has a cancer. In some embodiments, the subject is immune-depleted. As used herein, “cancer,” “tumor,” and “malignancy” all relate equivalently to hyperplasiaof a tissue or organ. If the tissue is a part of the lymphatic or immune system, malignant cells may include non-solid tumors of circulating cells. Malignancies of other tissues or organs may produce solid tumors. The methods described herein can be used in the treatment of lymphatic cells, circulating immune cells, and solid tumors. Cancers that can be treated include tumors that are not vascularized or are not substantially vascularized, as well as vascularized tumors. Cancers may comprise non-solid tumors (such as hematologic tumors, e.g., leukemias and lymphomas) or may comprise solid tumors. The types of cancers to be treated with the disclosed compositions include, but are not limited to, carcinoma,blastoma, and sarcoma, and certain leukemias or malignant lymphoid tumors, benign andmalignant tumors, and malignancies, e.g., sarcomas, carcinomas, and melanomas. Also includedare adult tumors / cancers and pediatric tumors / cancers. Hematologic cancers are cancers of the blood or bone marrow. Examples of hematologic (or hematogenous) cancers include leukemias, including acute leukemias (such as acute lymphocytic leukemia, acute myelocytic leukemia, acute myelogenous leukemia, promyelocytic, myelomonocytic, monocytic, and erythroleukemia), chronic leukemias (such as chronic myelocytic (granulocytic) leukemia, chronic myelogenous leukemia, and chronic lymphocytic Docket No.: FR 084276.00403 leukemia), polycythemia vera, lymphoma, Hodgkin’s disease, non-Hodgkin’s lymphoma (indolent and high-grade forms), myeloma Multiple, Waldenstrom’s macroglobulinemia, heavy chain disease, myelodysplastic syndrome, hairy cell leukemia, and myelodysplasia. Solid tumors are abnormal masses of tissue that usually do not contain cysts or liquid areas. Solid tumors can be benign or malignant. The different types of solid tumors are named for the type of cells that form them (such as sarcomas, carcinomas, and lymphomas). Examples of solid tumors, such as sarcomas and carcinomas, include fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma and other sarcomas, synovium, mesothelioma, Ewing tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, lymphoid malignancy, pancreatic cancer, breast cancer, lung cancer, ovarian cancer, prostate cancer, hepatocellular carcinoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, carcinoma of the sweat gland, medullary thyroid carcinoma, papillary thyroid carcinoma, sebaceous gland carcinoma of pheochromocytomas, carcinoma papillary, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, Wilms tumor, cervical cancer, testicular tumor, seminoma, bladder carcinoma, melanoma, and CNS tumors (such as glioma) (such as brainstem glioma and mixed gliomas), glioblastoma (also astrocytoma, CNS lymphoma, germinoma, medulloblastoma, Schwannoma craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, and brain metastasis). Non-limiting examples of tumors that can be treated by the methods described herein include, for example, carcinomas, lymphomas, sarcomas, blastomas, and leukemias. Non-limiting specific examples, include, for example, breast cancer, pancreatic cancer, liver cancer, lung cancer, prostate cancer, colon cancer, renal cancer, bladder cancer, head and neck carcinoma, thyroid carcinoma, soft tissue sarcoma, ovarian cancer, primary or metastatic melanoma, squamous cell carcinoma, basal cell carcinoma, brain cancers of all histopathologic types, angiosarcoma, hemangiosarcoma, bone sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endothelio sarcoma, lymphangiosarcoma, lymphangioendo-theliosarcoma, synovioma, testicular cancer, uterine cancer, cervical cancer, gastrointestinal cancer, mesothelioma, cancers associated with viral infection (such as but not limited to human papilloma virus (HPV) associated tumors (e.g., cancer cervix, vagina, vulva, head and neck, anal, and penile carcinomas)), Ewing’s tumor, leiomyosarcoma, Ewing’s sarcoma, Docket No.: FR 084276.00403 rhabdomyosarcoma, carcinoma of unknown primary (CUP), squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, Waldenstroom’s macroglobulinemia, papillary adenocarcinomas, cystadenocarcinoma, bronchogenic carcinoma, bile duct carcinoma, choriocarcinoma, seminoma,embryonal carcinoma, Wilms’ tumor, lung carcinoma, epithelial carcinoma, cervical cancer,testicular tumor, glioma, glioblastoma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, retinoblastoma, leukemia, neuroblastoma, small cell lung carcinoma, bladder carcinoma, lymphoma, multiple myeloma, medullary carcinoma, B cell lymphoma, T cell lymphoma, NK cell lymphoma, large granular lymphocytic lymphoma or leukemia, gamma-delta T cell lymphoma or gamma-delta T cell leukemia, mantle cell lymphoma, myeloma, leukemia,chronic myeloid leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, acutelymphocytic leukemia, hairy cell leukemia, hematopoietic neoplasias, thymoma, sarcoma, non- Hodgkin’s lymphoma, Hodgkin’s lymphoma, Epstein-Barr virus (EBV) induced malignancies of all types including but not limited to EBV-associated Hodgkin’s and non-Hodgkin’s lymphoma, all forms of post-transplant lymphomas including post-transplant lymphoproliferative disorder (PTLD), uterine cancer, renal cell carcinoma, hepatoma, and hepatoblastoma. Cancers that may be treated by methods described herein include, but are not limited to, cancer cells from the bladder, blood, bone, bone marrow, brain, breast, colon, esophagus, gastrointestine, gum, head, kidney, liver, lung, nasopharynx, neck, ovary, prostate, skin, stomach, testis, tongue, or uterus. The immune cells, as described, can be administered in a manner appropriate to the disease to be treated or prevented. The amount and frequency of administration will be determined by factors such as the condition of the patient, and the type and severity of the patient’s disease, although appropriate dosages can be determined by clinical trials. When “a therapeutically effective amount,” “an immunologically effective amount,” “aneffective antitumor quantity,” or “an effective tumor-inhibiting amount” is indicated, the preciseamount of the compositions of the present disclosure to be administered can be determined by a physician having account for individual differences in age, weight, tumor size, extent of infection or metastasis, and patient’s condition. It can generally be stated that a pharmaceutical composition Docket No.: FR 084276.00403 comprising the lymphocytes described herein can be administered at a dose of 104to 109cells / kg body weight, e.g., 105to 106cells / kg body weight, including all values integers within these intervals. The lymphocyte compositions can also be administered several times at these dosages. The cells can be administered using infusion techniques that are commonly known in immunotherapy (see, for example, Rosenberg et al., New Eng. J. of Med. 319: 1676, 1988). The optimal dose and treatment regimen for a particular patient can be readily determined by one skilled in the art of medicine by monitoring the patient for signs of the disease and adjusting the treatment accordingly. In some embodiments, the immune cells can be administered to the subject in a manner compatible with the dosage formulation and in such amount as is therapeutically effective. Dose ranges and frequency of administration can vary depending on, e.g., the nature of the population of cells (e.g., antigen-specific lymphocytes) produced by the methods described herein and the medical condition as well as parameters of a specific patient and the route of administration used. The administration of the immune cells or compositions thereof can be carried out in anyconvenient way, including infusion or injection (i.e., intravenous, intrathecal, intramuscular, intraluminal, intratracheal, intraperitoneal, or subcutaneous), transdermal administration, or othermethods known in the art. Administration can be once every two weeks, once a week, or moreoften, but the frequency may be decreased during a maintenance phase of the disease or disorder.In some embodiments, the immune cells or compositions thereof are administered by intravenousinfusion. Additional Definitions To aid in understanding the detailed description of the compositions and methods according to the disclosure, a few express definitions are provided to facilitate an unambiguousdisclosure of the various aspects of the disclosure. Unless otherwise defined, all technical andscientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Unless defined otherwise, all technical and scientific terms used herein have the meaningcommonly understood by a person skilled in the art to which this invention belongs. The followingreferences provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Docket No.: FR 084276.00403 Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics,5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper CollinsDictionary of Biology (1991). As used herein, the following terms have the meanings ascribed tothem below, unless specified otherwise. The term “machine learning,” as used herein, refers to a computer algorithm used to extractuseful information from a database by building probabilistic models in an automated way. The term “regression tree,” as used herein, refers to a decision tree that predicts values ofcontinuous variables. The term “supervised learning,” as used herein, refers to a data analysis using a well-defined (known) dependent variable. All regression and classification algorithms are supervised.In contrast, “unsupervised learning” refers to the collection of algorithms where groupings of thedata are defined without the use of a dependent variable. The term “test data” refers to a data set independent of the training data set, used to evaluatethe estimates of the model parameters (i.e., weights). As used herein, the term “clustering tree” refers to a hierarchical tree structure in whichobservations, such as organisms, genes, and polynucleotides, are separated into one or moreclusters. The root node of a clustering tree consists of a single cluster containing all observations,and the leaf nodes correspond to individual observations. A clustering tree can be constructed onthe basis of a variety of characteristics of the observations. Many techniques known in the art, e.g.,hierarchical clustering analysis, can be used to construct a clustering tree. A non-limiting exampleof a clustering tree is a phylogenetic, taxonomic or evolutionary tree.The terms “T cell” and “T lymphocyte” are interchangeable and used synonymously herein.As used herein, T cell includes thymocytes, naive T lymphocytes, immature T lymphocytes,mature T lymphocytes, resting T lymphocytes, or activated T lymphocytes. A T cell can be a Thelper (Th) cell, for example, a T helper 1 (Thl) or a T helper 2 (Th2) cell. The T cell can be ahelper T cell (HTL; CD4+ T cell) CD4+ T cell, a cytotoxic T cell (CTL; CD8+ T cell), a tumor-infiltrating cytotoxic T cell (TIL; CD8+ T cell), CD4+CD8+ T cell, or any other subset of T cells.Other illustrative populations of T cells suitable for use in particular embodiments include naiveT cells and memory T cells. Also included are “NKT cells,” which refer to a specialized population Docket No.: FR 084276.00403of T cells that express a semi-invariant ab T cell receptor, but also express a variety of molecularmarkers that are typically associated with NK cells, such as NK1.1. NKT cells include NK1.1+and NK1. G, as well as CD4+, CD4, CD8+, and CD8 cells. The TCR on NKT cells is unique in that it recognizes glycolipid antigens presented by the MHC I-like molecule CD Id. NKT cells can haveeither protective or deleterious effects due to their ability to produce cytokines that promote eitherinflammation or immune tolerance. Also included are”gamma-delta T cells (γδ T cells),” whichrefer to a specialized population that to a small subset of T cells possessing a distinct TCR on their surface, and unlike the majority of T cells in which the TCR is composed of two glycoproteinchains designated a- and b-TCR chains, the TCR in γδ T cells is made up of a g- chain and a d-chain. γδ T cells can play a role in immunosurveillance and immunoregulation and were found tobe an important source of IL-17 and to induce robust CD8+ cytotoxic T cell response. Alsoincluded are “regulatory T cells” or “Tregs,” which refer to T cells that suppress an abnormal orexcessive immune response and play a role in immune tolerance. Treg cells are typically transcription factor Foxp3-positive CD4+T cells and can also include transcription factor Foxp3 - negative regulatory T cells that are IL-10-producing CD4+T cells. The terms “natural killer cell” and “NK cell” are used interchangeably and usedsynonymously herein. As used herein, NK cell refers to a differentiated lymphocyte with a CD16+ CD56+ and / or CD57+ TCR- phenotype. NKs are characterized by their ability to bind to andkill cells that fail to express ‘self’ MHC / HLA antigens by the activation of specific cytolyticenzymes, the ability to kill tumor cells or other diseased cells that express a ligand for NK activating receptors, and the ability to release protein molecules called cytokines that stimulate or inhibit the immune response. The terms “treat” or “treatment” of a state, disorder or condition include: (1) preventing,delaying, or reducing the incidence and / or likelihood of the appearance of at least one clinical or sub-clinical symptom of the state, disorder or condition developing in a subject that may beafflicted with or predisposed to the state, disorder or condition, but does not yet experience ordisplay clinical or subclinical symptoms of the state, disorder or condition; or (2) inhibiting the state, disorder or condition, i.e., arresting, reducing or delaying the development of the disease or a relapse thereof or at least one clinical or sub-clinical symptom thereof; or (3) relieving the disease, i.e., causing regression of the state, disorder or condition or at least one of its clinical or sub-clinical symptoms. The benefit to a subject to be treated is either statistically significant or at least Docket No.: FR 084276.00403perceptible to the patient or to the physician. Thus, the term “treatment” includes preventing acondition from occurring in a patient, particularly when the patient is predisposed to acquiring the condition; reducing and / or inhibiting the condition and / or its development and / or progression; and / or ameliorating and / or reversing the condition. Insofar as some embodiments of the methods of the presently disclosed subject matter are directed to preventing conditions, it is understood thatthe term “prevent” does not require that the condition be completely thwarted. Rather, as usedherein, the term “preventing” refers to the ability of one of ordinary skill in the art to identify apopulation that is susceptible to the condition, such that administration of the compositions of the presently disclosed subject matter might occur prior to the onset of the condition. The term does not imply that the condition must be completely avoided. “Combination” therapy, as used herein, unless otherwise clear from the context, is meantto encompass administration of two or more therapeutic agents in a coordinated fashion and includes, but is not limited to, concurrent dosing. Specifically, combination therapy encompassesboth co-administration (e.g., administration of a co-formulation or simultaneous administration ofseparate therapeutic compositions) and serial or sequential administration, provided that administration of one therapeutic agent is conditioned in some way on administration of another therapeutic agent. For example, one therapeutic agent may be administered only after a different therapeutic agent has been administered and allowed to act for a prescribed period of time. See,e.g., Kohrt et al. (2011) Blood 117:2423.As used herein, the term “in vitro” refers to events that occur in an artificial environment,e.g., in a test tube or reaction vessel, in cell culture, etc., rather than within a multi-cellular organism. As used herein, the term “in vivo” refers to events that occur within a multi-cellularorganism, such as a non-human animal. It is noted here that, as used in this specification and the appended claims, the singularforms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise.The terms “including,” “comprising,” “containing,” or “having” and variations thereof aremeant to encompass the items listed thereafter and equivalents thereof as well as additional subject matter unless otherwise noted. Docket No.: FR 084276.00403 The phrases “in one embodiment,” “in various embodiments,” “in some embodiments,”and the like are used repeatedly. Such phrases do not necessarily refer to the same embodiment,but they may unless the context dictates otherwise. The terms “and / or” or means any one of the items, any combination of the items, or allof the items with which this term is associated. The word “substantially” does not exclude “completely,” e.g., a composition which is“substantially free” from Y may be completely free from Y. Where necessary, the word“substantially” may be omitted from the definition of this disclosure.As used herein, the term “approximately” or “about,” as applied to one or more values ofinterest, refers to a value that is similar to a stated reference value. In some embodiments, the term“approximately” or “about” refers to a range of values that fall within 25%, 20%, 19%, 18%, 17%,16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in eitherdirection (greater than or less than) of the stated reference value unless otherwise stated orotherwise evident from the context (except where such number would exceed 100% of a possiblevalue). Unless indicated otherwise herein, the term “about” is intended to include values, e.g.,weight percents, proximate to the recited range that are equivalent in terms of the functionality of the individual ingredient, the composition, or the embodiment. It is to be understood that wherever values and ranges are provided herein, all values and ranges encompassed by these values and ranges are meant to be encompassed within the scope ofthe present disclosure. Moreover, all values that fall within these ranges, as well as the upper orlower limits of a range of values, are also contemplated by the present application. As used herein, the term “each,” when used in reference to a collection of items, is intendedto identify an individual item in the collection but does not necessarily refer to every item in thecollection. Exceptions can occur if explicit disclosure or context clearly dictates otherwise.The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope ofthis disclosure unless otherwise claimed. No language in the specification should be construed asindicating any non-claimed element as essential to the practice of this disclosure. Docket No.: FR 084276.00403 All methods described herein are performed in any suitable order unless otherwiseindicated herein or otherwise clearly contradicted by context. In regard to any of the methodsprovided, the steps of the method may occur simultaneously or sequentially. When the steps ofthe method occur sequentially, the steps may occur in any order, unless noted otherwise. In cases in which a method comprises a combination of steps, each and every combination or sub-combination of the steps is encompassed within the scope of the disclosure, unless otherwise noted herein. Each publication, patent application, patent, and other reference cited herein is incorporated by reference in its entirety to the extent that it is not inconsistent with the presentdisclosure. Publications disclosed herein are provided solely for their disclosure prior to the filingdate of the present disclosure. Nothing herein is to be construed as an admission that the presentinvention is not entitled to antedate such publication by virtue of prior invention. Further, the datesof publication provided may be different from the actual publication dates, which may need to be independently confirmed. It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. Examples EXAMPLE 1 This example describes the materials and methods employed by the subsequent examples. Patients and regulatory issues Patients included stage III / IV metastatic melanoma, ovarian, non-small cell lung cancerand colorectal cancer patients (Table 2) and had received several lines of chemotherapy andimmunotherapy. Samples were collected and biobanked from patients enrolled under protocolsapproved by the respective institutional regulatory committees at the University of Pennsylvania, USA, and Lausanne University Hospital (CHUV), Switzerland. Patients and healthy donors’recruitment, study procedures, and blood withdrawal were approved by regulatory authorities andall patients signed written informed consents. Docket No.: FR 084276.00403 Identification of non-synonymous tumor mutations Genomic DNA from cryopreserved tumor tissue and matched PBMC was isolated using DNeasy kit (Qiagen) and subjected to whole exome capture and paired-end sequencing using theHiSeq2500 Illumina platform as described (Bobisse, S., et al. Nat Commun 9, 1092 (2018)). RNAwas extracted for RNA sequencing using the Total RNA Isolation RNeasy Mini Kit (Qiagen) according to the manufacturer’s protocol and sequenced on the same platform for paired end sequencing. Non-synonymous tumor mutations were identified from tumor tissues and matched blood cells. Samples from patients CRC1 and CRC2 and OvCa1-4 were analyzed as previously described(Bobisse, S., et al. Nat Commun 9, 1092 (2018)). Samples from patients Mel7-10 were analyzedwith NeoDisc V1.2 pipeline (Bassani-Sternberg, M. et al. Front. Immunol. 10, 1832 (2019)) thatincludes the GATK variant calling algorithm Mutect2, Mutect1, HaplotypeCaller and VarScan 2. NeoDisc v1.2 also determines the presence of each mutation and quantifies the expression of each mutant gene and mutation from RNAseq data. Predictions for binding to HLA class-I of allcandidate peptides of samples from patients CRC1 and CRC2, and OvCa1-4 were performed usingthe NetMHC v3.4 and netMHCpan-3.0 algorithms. Predictions for binding and immunogenicityon candidate peptides of samples from patients Mel7-10 were performed using the PRIMEalgorithm (Schmidt, J. et al. Cell Rep. Med. 2, 100194 (2021)). Candidate neoantigen-antigenpeptides (i.e., mutant 9mer and 10mer peptide sequences containing the somatically alteredresidue) with a %rank < 0.5 were synthesized. Antigen validation CD8 T cells (106mL-1) isolated (Dynabeads, Invitrogen) from cryopreserved PBMC were co-incubated with autologous irradiated CD8 and CD4-depleted PBMCs and peptides (1 µg mL-1, single peptide or pools of ≤ 50 peptides) in RPMI supplemented with 8 % human serum and IL-2 (20 IU mL-1) for 48 h and then 100 IU mL−1). IFN-γ Enzyme-Linked ImmunoSpot (ELISpot) andpeptide-MHC multimer staining assays were performed on day 12. T cell reactivity for everyneoantigens was validated by ≥ 2 independent experiments. ELISpot assays were performed using pre-coated 96-well ELISpot plates (Mabtech) and counted with Bioreader-6000-E (BioSys). Those with an average number of spots higher than the counts of the negative control (No Ag) plus 3times the standard deviation of the negative was considered as positive conditions. TILs were Docket No.: FR 084276.00403 generated from tumor enzymatic digestion by plating total dissociated tumor in p24-well plates at a density of 1x106cells / well in RPMI supplemented with 8% human serum and IL-2 (6000 IU mL-1). After 2-4 weeks, TILs were collected, and a fraction of the cultures underwent a rapid expansion(REP) for 14 days. T cell reactivity against predicted neoantigens was tested by IFN-γ ELISpot onpre-REP TILs, when available, and post-REP TILs as described above. Positivity was confirmed in ≥2 independent experiments. Isolation and expansion of antigen-specific CD8 T cells Circulating and tumor-infiltrating antigen-specific CD8 T cells were FACS sorted using reversible pMHC multimers (NTAmers), and were either used for TCR sequencing or cloned by limiting dilution. To this end, cells were plated in Terasaki plates and stimulated with irradiated feeder cells (PBMC from two donors) in RPMI supplemented with 8% human serum,phytohemagglutinin (1 μg mL−1), and IL-2 (150 IU mL−1). At the end of the expansion, pMHCmultimer-positive cells were ≥ 95% pure. Peptide synthesis Peptides produced by the Peptides and Tetramers Core Facility (PTCF) of the University of Lausanne were HPLC purified (≥90% pure), verified by mass spectrometry and kept lyophilized at -80°C. Production of NTAmers and peptide binding assay NTAmers (reversible pMHC multimers) were synthesized at the Peptide and TetramerCore Facility of the University of Lausanne as described (Schmidt, J. et al. J. Biol. Chem. 286,41723–41735 (2011)). NTAmers are composed of streptavidin-phycoerythrin (SA-PE; Invitrogen) complexed with biotinylated peptides carrying four Ni2+-nitrilotriacetic acid (NTA4) moieties and non-covalently bound to His-tagged pMHC monomers. For pMHC-TCR dissociation kinetics experiments, pMHC monomers were refolded with Cy5-labeled β2m. Briefly, β2m containing theS88C mutation was alkylated using Cy5-maleimide (Pierce), purified, and used for furtherrefolding assay. Peptide-MHC monomers were produced by refolding the different HLA heavy chains in the presence of labeled β2m and peptide of interest, purified on a Superdex S75 quantifiedby Bradford, aliquoted, and kept at -80°C until further use. Validation and quantification of peptidebinding was done by micro-scale refolding. Refolding with HLA heavy chains carrying a C- Docket No.: FR 084276.00403terminal BirA substrate peptide (BSP), Cy5-labeled β2m, and a test peptide were performedessentially as described (Schmidt, J. et al. J. Biol. Chem. 286, 41723–41735 (2011)). Human β2mwas mutated S88 to C and, after refolding, alkylated with maleimide-PEG2-Cy5 (Pierce, ThermoFisher Scientific) in PBS at pH 7.4. Refolding reactions were performed in 96 well plates at 4°C for 72h in the presence of 10 µM peptide. Incubation without peptide control. After centrifugation(4,000 rpm, 5min), the reaction mixtures were transferred into 96 well plates, and Cy5 fluorescenceread on a fluorescence plate reader (Modulus, Promega). All measurements were performed intriplicates, and data was processed using Excel (Microsoft).Structural avidity assay KIF1BS918F-specific T cells were obtained by co-transfecting Jurkat cells (Promega) with500 ng each of TCRα and TCRβ chain RNA together with 300 ng each of CD8α and CD8β RNA,using a Neon electroporation system (Thermo Fisher Scientific) as previously described (Arnaud,M. et al. Nat. Biotechnol. 1–5 (2021)).Antigen-specific CD8 T cell clones (i.e., obtained from isolation and expansion of primarycells) or transfected Jurkat cells (2x105cells) were incubated for 40min at 4°C with cognate NTAmers containing streptavidin-phycoerythrin and Cy5-labeled pMHC monomers in 50µLFACS buffer (PBS supplemented with 0.5% BSA and 2mM EDTA), as described (Hebeisen, M.et al. Cancer Res. 75, 1983–1991 (2015)). Irrelevant T cells were used to measure backgroundsignal, and values were systematically subtracted. Specific gMFI values were plotted and analyzedusing the GraphPad Prism software (v.7, GraphPad) fitting a one-phase exponential decay model. Functional avidity assay Functional avidity of antigen-specific CD8 T cell responses was assessed by performing invitro IFN-γ Enzyme-Linked ImmunoSpot (ELISpot, Mabtech) assay with limiting peptidedilutions (ranging from 10 μg mL-1 to 0.1 pg mL-1) as described (Viganò, S. et al. Clin. Dev.Immunol. 2012, 153863 (2012)). EC50 values were derived by dose-response curve analysis (log(peptide concentration) versus response) using GraphPad Prism software (v.7, GraphPad). The peptide concentration required to achieve a half-maximal cytokine response (EC50) wasdetermined and referred to as the functional avidity.CD8 T cell tropism assay Docket No.: FR 084276.00403 PBMCs, primary CD8 T cell clones, or primary CD8 T cells transduced with engineered TCR specific for NY-ESO-1 restricted by HLA-A*0201 were distributed in 48-well plates (6x105 / well) in RPMI supplemented with 8% human serum and IL-2 (150 U mL-1). Cells were stimulated at 37°C under 5% CO2 either with culture medium alone, phytohemagglutinin (PHA; Oxoid, 1 mg mL-1), OKT3 antibody (plate precoated with 30 ng mL-1, 5 ng mL-1, or 1 mg mL-1 in PBS), or 2x105T2 cells pulsed with cognate peptide (at 1 mM or 1nM). After 48h, cells were washed and replaced in culture for 48h at 37°C under 5% CO2 in RPMI supplemented with 8% human serum and IL-2 (150 IU mL-1). Half of the cells were analyzed by flow cytometry (BD LSRII flow cytometer) using the following panel of antibodies: Zombie Aqua™ dye (Biolegend), Pacific Blue™ anti-CD8 (SK1, Biolegend), PE-Texas Red anti-CD3d (7D6, Invitrogen), Brilliant Violet 650™ anti-CX3CR1 (2A9-1, Biolegend), Brilliant Violet 605™ anti-CD194 (CCR4) (L291H4, Biolegend), Brilliant Violet 711™ anti-CD197 (CCR7) (G043H7, Biolegend), FITC anti-CD49b (P1E6-C5, Biolegend), PerCP / Cy5.5 anti-CD195 (CCR5) (HEK / 1 / 85a, Biolegend), Brilliant Violet 650™ anti-CD196 (CCR6) (G034E3, Biolegend), PE anti-CD49a (TS2 / 7, Biolegend), PE / Cy7 anti-CD103 (Integrin αE) (Ber-ACT8, Biolegend), Brilliant Violet 510™ anti-CD183 (CXCR3) (G025H7, Biolegend). After 5 days of resting, the remaining cells were profiled with the same panel. TCRα and TCRβ repertoire sequencing mRNA was extracted using the Dynabeads mRNA DIRECT purification kit according to the manufacturer instructions (ThermoFisher) and was then amplified using the MessageAmp IIaRNA Amplification Kit (Ambion) with the following modifications: in vitro transcription wasperformed at 37°C for 16h. First strand cDNA was synthesized using the Superscript III (Thermofisher) and a collection of TRAV / TRBV specific primers. TCRs were then amplified byPCR (20 cycles with the Phusion from NEB) with a single primer pair binding to the constantregion and the adapter linked to the TRAV / TRBV primers added during the reverse transcription. A second round of PCR (25 cycles with the Phusion from NEB) was performed to add the Illumina adapters containing the different indexes. The TCR products were purified with AMPure XP beads(Beckman Coulter), quantified, and loaded on the MiniSeq instrument (Illumina) for deepsequencing of the TCRα / TCRβ chain. The TCR sequences were further processed using ad hoc Perl scripts to: (i) pool all TCR sequences coding for the same protein sequence; (ii) filter out allout-frame sequences; and (iii) determine the abundance of each distinct TCR sequence. TCR with Docket No.: FR 084276.00403a single read were not considered for the analysis. This methodology was previously described inArnaud et al. and Bobisse et al. ((Bobisse, S., et al. Nat Commun 9, 1092 (2018)); Arnaud, M. etal. Nat. Biotechnol. 1-5 (2021)).Clone TCRα and TCRβ sequencing mRNA was extracted using the Dynabeads mRNA DIRECT purification kit according tothe manufacturer's instructions (ThermoFisher). First strand cDNA was synthesized using oligodT and the Superscript III (Thermofisher). Second strand was performed using a collection of TRAV / TRBV specific primer (1 cycle with the Phusion from NEB). TCRs were then amplified by PCR (20 cycles with the Phusion from NEB) with a single primer pair binding to the constant region and the adapter linked to the TRAV / TRBV primers added during the reverse transcription. A second round of PCR (25 cycles with the Phusion from NEB) was performed to add the Illumina adapters containing the different indexes. The TCR products were purified with AMPure XP beads(Beckman Coulter), quantified, and loaded on the MiniSeq instrument (Illumina) for deepsequencing of the TCRα / TCRβ chain. The TCR sequences were further processed using ad hocPerl scripts to: (i) pool all TCR sequences coding for the same protein sequence; (ii) filter out allout-frame sequences; and (iii) determine the abundance of each distinct TCR sequence. TCRs witha single read were not considered for the analysis. This methodology was previously reported inSchmidt et al. (Schmidt, J. et al. Cell Rep. Med. 2, 100194 (2021)).TCR transduction TCRα / TCRβ chains were cloned in a pMSGV retroviral vector downstream of the blasticidin resistance gene followed by a P2A element. Viral particles were produced by mixing in 250 μl of Optimem medium (Life Technologies), pMSGV (1.25 μg), packaging plasmidspMD.gagpol (1.25 μg), and pMD.G (1.25 μg, VSV-G envelope protein) with 7.5 μl of MIRUSreagent (MIRUS Bio LLC, USA). After 20 min at RT, the mix was added slowly to 106293 T cells(Storck, A., et al. BioTechniques 63, 136–138 (2017)). After 48h, 50 μl of virus-containingsupernatant was collected and added to 106 primary CD8 T cells previously stimulated for 24hwith anti-CD3 / anti-CD28 beads. After 24h, the medium was changed, and blasticidin (Sigma-Aldrich) was added at 500 μg mL-1. TCR expression was checked by pMHC multimer staining after 4 days. Once TCR expressing cells reached >90% purity, they were used in functional and structural assays. Docket No.: FR 084276.00403 For experiments with NY-ESO-I-specific TCRs, the following procedure was followed. Full-length codon-optimized TRAV23.1 and TRBV13.1 chain sequences of a dominant HLA-A0201 / NY-ESO-I157-165 specific T cell clone of patient LAU155 (Hebeisen, M. et al. Cancer Res.75, 1983–1991 (2015)) were cloned in the pRRL third generation lentiviral vectors as an hPGK-AV23.1-IRES-BV13.1 construct and structure-based amino acid substitutions were introduced into the WT TCR sequence by point mutations. Lentiviral production was performed using thecalcium-phosphate method, and concentrated supernatant of lentiviral-transfected 293T cells wasused to infect primary CD8 T cells overnight. Levels of TCR transduction efficacy were monitoredby pMHC multimer staining (Hebeisen, M. et al. Cancer Res. 75, 1983-1991 (2015)).For experiments with KIF1BS918F-specific TCRs, the following procedure was followed. TCRα and TCRβ chains, divided by a Furin / GS linker / T2A, were cloned into a pCRRL-pGK lentiviral plasmid to produce high-titer replication-defective lentiviral particles, as previouslydescribed (Giordano-Attianese, G. et al. Nat. Biotechnol.38, 426–432 (2020)). For primary humanT cell transduction, CD8 T cells were negatively selected with beads (Miltenyi Biotec) fromPBMCs of a healthy donor (apheresis filter from anonymous healthy donors following the legal Swiss guidelines under project P_123 with informed consent of the donors and with ethics approval from the Canton of Vaud (Lausanne)), activated and transduced as previously reported(Giordano-Attianese, G. et al. Nat. Biotechnol. 38, 426–432 (2020)), with minor modifications.Briefly, CD8 T cells were incubated with lentiviral particles after 24 hours of activation with anti- CD3 / CD28 beads (Thermo Fisher Scientific) in R8 medium supplemented with 50 IU ml−1IL-2.After incubation at 37°C for 72h, beads were removed, and transduced cells were sorted with aBD FACSAria III or BD melody Cell Sorter for viability and expression of CD8 and mouse TCRβ- constant region. Sorted KIF1BS918F TCR-transduced CD8 T cells were then expanded for 10 days in R8 medium and 50 IU ml−1of IL-2 before mouse injection.Adoptive T cell transfer in immunodeficient IL-2 NOG mice and multispectralimmunofluorescence staining IL-2 NOG mice (Taconic) were maintained in a conventional animal facility at theUniversity of Lausanne under a specific pathogen-free status. Six- to nine-week-old female micewere anesthetized with isoflurane and subcutaneously injected with 106autologous human melanoma tumor cells (grown in DMEM medium supplemented with 10% FCS). Once the tumors Docket No.: FR 084276.00403became palpable (around day 7), 2-5x106 human tumor-specific CD8 T cell clones (Hebeisen, M.et al. Cancer Res. 75, 1983–1991 (2015)) were injected intravenously in the tail vein. Tumorvolumes were measured by caliper twice a week and calculated as follows: volume = length * width * width / 2. Mice were sacrificed by CO2 inhalation before the tumor volume exceeded 103mm3or when necrotic skin lesions were observed at the tumor site. The same experiment withNY-ESO-I-specific TCRs was repeated, but anti-CXCR3 monoclonal antibody (Biolegend) wasinjected i.p. (100 μg per mouse) at day 5 (simultaneously of ACT) and day 10. Additionally, for the CXCR3-blockade experiment with KIF1B-specific TCRs, anti-CXCR3 or isotype monoclonal antibodies (Biolegend) were injected ip. (100 μg per mouse) twice a week for two weeks starting the day of ACT. This study was approved by the Veterinary Authority of the Canton de Vaud (under license 3387) and performed in accordance with Swiss ethical guidelines. Tumors were then harvested,processed, and analyzed by in situ immunofluorescence labeling. Briefly, Multiplexed stainingwas performed on 4-micrometer formalin-fixed paraffin-embedded (FFPE) tissue sections on an automated Ventana Discovery Ultra staining module (Ventana, Roche). Slides were placed on thestaining module for deparaffinization, epitope retrieval (64min at 95°C), and endogenousperoxidase quenching (Discovery Inhibitor, 8min, Ventana). Multiplex staining consists of multiple rounds of staining. Each round includes non-specific sites blocking (Discovery Goat IgG and Discovery Inhibitor, Ventana), primary antibody incubation, secondary HRP-labeled antibody incubation for 16min (Discovery OmniMap anti-rabbit HRP (Ventana, # 760-4311) or anti-mouse HRP (Ventana, #760-4310)), OPALTMreactive fluorophore detection (Akoya Biosciences, Marlborough, MS, USA) that covalently label the primary epitope (incubation: 12min) and then antibodies heat denaturation. The sequence of antibodies used in the multiplex with the associatedOPAL is the following: 1st, rabbit anti-CD8 antibody (4 µg / ml, Clone SP16, Cellmarque, 1h, 37°C),OPAL520; and 2nd, rabbit anti-SOX10 antibody (1 µg / ml, Clone EP268, CellMarque, 1h, RT), OPAL690; Nuclei were visualized by a final incubation with Spectral DAPI (1 / 10, FP1490, Akoya Biosciences) for 12min. Multiplex IF images were acquired on Vectra 3.0 automated quantitative pathology imaging system (Akoya Biosciences). Tissue and panel-specific spectral library of each panel individual fluorophore and tumor tissue autofluorescence were acquired for an optimal IF signal un-mixing (individual spectral peaks) and multiplex analysis. IF stained slides were pre- scanned at 10x magnification. Using the Phenochart™ whole-slide viewer (Akoya Biosciences). Docket No.: FR 084276.00403 Whole tumor was selected and annotated for the high-resolution multispectral acquisition of images at 20x magnification. IF signal extractions were performed using inForm 2.3.0 imageanalysis software (Akoya Biosciences) (Kramer, A. S. et al. Sci. Rep. 8, 3418 (2018)) enabling aper-cell analysis of IF markers of multiplex stained tissue sections. The images were firstsegmented into tumor, stroma, and necrosis regions based on the cytokeratin staining using theinForm Tissue Finder™ algorithms. Individual cells were then segmented using the counterstained-based cell segmentation algorithm, based on DAPI staining. Quantification of theimmune cells was performed using the informed active learning phenotyping algorithm byassigning the different T cell phenotypes across several images. IF-stained cohorts are then batchprocessed, and data were exported and processed via an in-house developed R-script algorithm toretrieve every cell population. TCR reactivity validation Paired α and β chains were annotated based on single-cell TCR sequencing data. For TCR cloning, DNA sequences coding the full-length TCR chains were codon optimized and synthesized by GeneArt (Thermo Fisher Scientific) or with a BioXP System (Telesis Bio). Each DNA sequence included a T7 promoter upstream of the ATG codon, whereas human constant regions of α and β chains were replaced by corresponding homologous murine constant regions. DNA served as a template for in vitro transcription (IVT) and polyadenylation of RNAmolecules per the manufacturer’s instructions (Thermo Fisher Scientific). To confirm antitumorreactivity, TCRα and TCRβ RNA were transfected into recipient activated T cells. AutologousPBMCs were resuspended at 106cells mL-1in 48-well plates in R8 medium supplemented with 50 IU mL-1IL-2 (Proleukin). T cells were activated with Dynabeads Human T Activator CD3 / CD28 beads (Thermo Fisher Scientific) at a ratio of 0.75 beads: 1 total PBMCs. After 3 days of incubationat 37°C and 5% CO2, beads were removed, and activated T cells were cultured for two extra daysbefore electroporation or freezing. For the transfection of TCRαβ pairs into T cells, the Neon electroporation system (Thermo Fisher Scientific) was used. Cells were resuspended at 15-20x106cells mL-1in buffer R (buffer from the Neon kit) and mixed with 500 µg of TCRα chain RNA together with 500 µg of TCRβ chain RNA and electroporated with the following parameters: 1600V, 10ms, 3 pulses and 1325V, 10msec, 3 pulses, respectively. Electroporated cells were 6 hours and used in co-culture experiments for tumor recognition and functional avidity assays. To Docket No.: FR 084276.00403assess antitumor reactivity TCR RNA-electroporated T cells were incubated with IFNγ-treatedautologous tumor cells at a ratio of 5:1. After over-night culture, cells were collected and the upregulation of 4-1BB (CD137) was evaluated by staining with anti-4-1BB PE (4B4-1, Miltenyi), anti-CD3 APC Fire 50 (SK7, Biolegend) or anti-CD3 APC-H7 (SK7, BD Biosciences), anti-CD4 PE-CF594 (RPA-T4, BD Biosciences), anti-CD8 Pacific Blue™ (RPA-T8, BD Biosciences) and anti-mouse TCRβ-constant APC (H57-597, Thermo Fisher Scientific) and with viability dye Aqua (Thermo Fisher Scientific). The following experimental controls were included: mock (transfection with water), an irrelevant TCR (random crossmatch of a TCRα and β chain). Flow cytometry was performed using LSR Fortessa (BD Biosciences) or IntelliCyt iQue® Screener PLUS (Bucher Biotec) and analyzed with FlowJo v10 (TreeStar). Single-cell RNA and TCR sequencing Expanded TILs from patients Mel7-10 were resuspended in PBS + 0.04% BSA, and DAPI(Invitrogen) staining was performed. Live cells were sorted with a BD FACS Melody sorter and manually counted to assess viability with Trypan blue. Cells were then resuspended at 103cells µL-1with viability of >90% and subjected to a 10X Chromium instrument for the single-cellanalysis. The standard protocol of 10X Genomics was followed, and the reagents for theChromium Single Cell 5’ Library and V(D)J library (v1.0 Chemistry) were used.12,200 cells wereloaded per sample, with the targeted cell recovery of 7,000 cells according to the protocol. Singlecells were captured and lysed using microfluidic technology, and mRNA was reverse transcribedto barcoded cDNA using the provided reagents (10X Genomics). 14 PCR cycles were used toamplify cDNA, and the final material was divided into two fractions: the first fraction was target-enriched for TCRs and V(D)J library was obtained according to manufacturer protocol (10X Genomics). Barcoded VDJ libraries were pooled and sequenced by an Illumina HiSeq 2500Sequencer. The second fraction was processed for 5’ gene expression library following themanufacturer’s instruction (10X Genomics). Barcoded samples were pooled and sequenced by an Illumina HiSeq 4000 sequencer. The scRNA-seq reads were aligned to the GRCh38 reference genome and quantified usingCellranger count (10x Genomics, version 3.0.2). Filtered gene-barcode matrices that containedonly barcodes with unique molecular identifier (UMI) counts that passed the threshold for cell detection were used for further analysis. The number of genes per cell averaged 1,862 (median: Docket No.: FR 084276.004031,729), and the number of unique transcripts per cell averaged 4,886 (median: 4,169).18,378 cells(7,056 for Mel7, 3,656 for Mel8, 3,137 for Mel9 and 4,529 for Mel10) were obtained. Low qualitycells exhibiting more than 10% of mitochondrial reads were discarded from the analysis, resultingin a final set of 17,937 cells (6,916 for Mel7, 3,545 for Mel8, 3,059 for Mel9, and 4,417 for Mel10).The data was processed using the Seurat R package (version 3.2.2) as follows briefly: counts werelog-normalized using the NormalizeData function and then scaled using the ScaleData functionby regressing the mitochondrial, ribosomal contents and S phase and G2 / M phase scores. Dimensionality reduction was performed using the standard Seurat workflow by principalcomponent analysis followed by tSNE and UMAP projection (using the first 75 PCs). The k-nearest neighbors of each cell were found using the FindNeighbors function run on the first 75PCs, and followed by clustering at several resolutions using the FindClusters function. Cells wereannotated by looking at expression of the canonical PTPRC and CD3E markers, where all clusterswere found to be T cells. The cells were then classified as CD8-positive, CD4-positive, double-negative (DN), double-positive (DP), and Tγδ as follows: cells with non-null expression of CD8Aand null expression of CD4 were defined as CD8-positive (and vice-versa for CD4-positive). Cellsshowing non-null expression of both genes were classified as DP. Due to notorious dropout events in single-cell data, cells lacking the expression of both markers were classified as follows: if a cell belongs to a cluster (taking a fine resolution of 10) in which the 75thpercentile expression of CD8 was higher than its 75thpercentile expression of CD4, it was classified as CD8-positive (and vice-versa for CD4-positive cells). If the 75th percentile expressions of both markers equal 0, the cellswere classified as DN. Finally, cells with an average expression score of all TRG and TRD-relatedgenes higher than 0.3 were assigned to be Tγδ cells. This resulted in the final set of 10,947 CD8 Tcells, 5’922 CD4 T cells, 852 DP, 1 DN, and 132 Tγδ cells.VDJ sequencing data were aligned to the same human genome using the Cellranger VDJ(10x Genomics, version 3.1.0). Cells from the VDJ sequencing were mapped to the scRNAseqdata, and 90.7% of the T cells had a mapped TCR β-chain (84.6% for TCR α-chain).TCR-pMHC structure modeling and correlation with structural avidity The Rosetta “TCRmodel” protocol (Giordano-Attianese, G. et al. Nat. Biotechnol. 38,426–432 (2020)) was adapted to the approach and applied to find the respective templates andmodel TCR. The orientation of the TCR relative to the pMHC was performed based on TCR- Docket No.: FR 084276.00403pMHC templates retrieved from Protein Data Bank (Rose, P. W. et al. Nucleic Acids Res. 45,D271–D281 (2017)) and identified using sequence similarity. Side chains and backbones of theTCR-pMHC models were refined using the fast “relax” protocol in Rosetta (Nivón, L. G., et al.PloS One 8, e59004 (2013).). A total of 500 models were produced for each TCR-pMHC. These models were subsequently ranked based on a consensus approach that combines the Rosetta energyfunction as implemented in Rosetta (Gowthaman, R. & Pierce, B. G. Nucleic Acids Res. 46,W396–W401 (2018)) and the Discrete Optimized Potential Energy as implemented in Modeller(Webb, B. & Sali, A. Curr. Protoc. Protein Sci. 86, 2.9.1-2.9.37 (2016)). This consensus score corresponded to the sum of the normalized (Z-score) Rosetta and DOPE energies calculated over the peptide residues, as well as the CDRs and MHC residues within 6 Å from the peptide. For each TCR-pMHC, the best model according to the consensus score was selected for CDR loop refinement. The latter was performed by creating 100 alternative loop conformations using thekinematic closure loop modeling of Rosetta (Mandell, D. J., et al. Nat. Methods 6, 551–552 (2009))and subsequent refinement using the fast “relax” protocol. The final TCR:pMHC structural modelis the one with the highest number of favorable interactions within the top 5 high-score modelsover 600. Molecular graphics and analyses were performed with the UCSF Chimera package(Pettersen, E. F. et al. J. Comput. Chem. 25, 1605–1612 (2004)). Correlation between the meanstructural avidity of each pMHC-TCR pair with the number of non-polar, ^^^^^^^, and the numberof polar, ^^^^^^ , contacts between modeled TCR and pMHC was obtained via the equation: This equation represents a simplification of the binding free energy estimation (Zoete, V.,et al. J. Comput. Chem. 32, 2359–2368 (2011)), where ^ ^^^ ^ are weighting terms applied onthe number of apolar and polar contacts, respectively, and K is added to account for contributionsthat are not a function of the number of polar and non-polar contacts. The ^, ^ ^^^ ^ parametersare fitted by multiple linear regression against the experimental pMHC-TCR T1 / 2. The parameterswere optimized using 10 complexes (Table 3), and values of -62.89s, 2.647s, and 8.747s wereobtained for K, ^ ^^^ ^, respectively. The correlation coefficient R, the leave-one-out correlationcoefficient, the standard deviation, and the p-value are 0.8679, 0.6928, 24 s, and 0.005,respectively. The relevance of this correlation was assessed by a randomization test. The latterconsisted in attributing randomly, for each TCR, the T1 / 2 of another TCR (paying attention that Docket No.: FR 084276.00403 each of the ten T1 / 2obtained experimentally was re-attributed only once), before applying the same multiple linear regression. This randomization test was performed 10,000 times. A better correlation with the randomized T1 / 2than with the true T1 / 2was obtained for only 0.5% of the tests, in agreement with a p-value of 0.005. Hierarchical clustering of TCR sequences Acomputational pipeline was implemented based on a biophysicochemical approach(Atchley, W. R., et al. Proc. Natl. Acad. Sci. U. S. A. 102, 6395–6400 (2005)) that allows TCRcomparisons by analyzing the biophysicochemical properties of the 4-mer subunits that arepossible to construct from CDR3β and comparing them across all the TCRs under study. To provide insights into the clusters, structural models were created for the TCRs as described in the previous section. The clustering pipeline consists of 4 main steps. First, all possible slidingwindows of 4 residues that constitute the so-called 4-mer subunits were identified. The first 4 andthe last 3 residues of the CDR3β are excluded from this process because these residues usually do not contact the HLA peptide. Second, each 4-mer subunit is converted into a biophysicochemical representation using 5 Atchley factors that describe i) hydrophobicity, ii) secondary structure, iii)size / mass, iv) codon degeneracy, and v) electric charge. Third, for a pair of TCRs, all the n 4-mersubunits that are possible to construct from the first TCR with all the m possible 4-mer subunits ofthe second TCR were compared. This results in n*m matrices to compare for each pair of TCRs.The matrices’ comparison is performed via a Manhattan distance score normalized over themaximum possible distance. This score ranges from 0, for 4-mers sharing exactly the same biophysicochemical properties, and to 1, for 4-mers that have totally different biophysiochemical properties. Fourth, a distance tree is constructed using the smallest distance for each TCR pair. The generic hierarchical clustering algorithm UPGMA (unweighted pair group method witharithmetic mean) is used (Wheeler, T. & Kececioglu, et al. Oxf. Engl. 23, i559-68 (2007)). Theclustering analysis was finally applied to a total of 58 TCRs with known pMHC, 52 of which withknown avidity (Table 4).Logistic Regression to discriminate between low and high avidity TCRs Alogistic regression was implemented, based on the CDR3β amino acids that are enoughsolvent exposed and therefore able to interact with the peptide, to determine whether a TCR binds the cognate pMHC with high or low koffvalue. TCR avidity was correlated with CDR3β sequence Docket No.: FR 084276.00403as it is generally accepted that CDR3β is the most determining CDR for antigen specificity (Dash,P. et al. Nature 547, 89–93 (2017); Glanville, J. et al. Nature 547, 94–98 (2017)). The amino acidsused in the logistic regression resulted from an exhaustive exploration (details below) and correspond to the optimal solution found. The probability, p, of a TCR being high avidity is given by: R, N, D, G, I, L, and F take the value of 1 when the corresponding amino acid is presentwith a solvent accessibility higher than 30% in the CDR3β of the TCR 3D model, and 0 otherwise. The solvent accessibility of each CDR3β residue is determined as the relative solvent excluded surface area (SESA) computed with the MSMS package of the UCSF Chimera software, asdescribed by Goddard, T. D. et al. (Goddard, T. D. et al. Protein Sci. Publ. Protein Soc.27, 14–25(2018)). SESA is calculated by normalizing the surface area of the residue in the TCR of interestby its surface area in a reference state (Bendell, C. et al. BMC Bioinformatics 15, 82 (2014)). Thebias term b0 and the weights Wn were determined using a set of 48 TCRs (TCRs with undeterminedavidity and TCRs 5, 25, 26, 27, and 38 without a good 3D model were discarded and thereforewithout calculated solvent accessibility from Table 4, maximizing the likelihood that each avidityprediction for these TCRs is correct. The TCRs were divided into two sets: high avidity set with 11 TCRs (T1 / 2 > 60s) and low avidity set, with 37 TCRs (T1 / 2 < 60s). Before converging to this model, TCR avidity was correlated with CDR3β sequence as itis generally accepted that CDR3β is the most determining CDR for antigen specificity (Dash, P.et al. Nature 547, 89–93 (2017); Glanville, J. et al. Nature 547, 94–98 (2017)). Thevariables / parameters previously used resulted from an exhaustive exploration and corresponded tothe optimal solution found. The relationship between the outcome, i.e., the avidity, and differentpredictors, either binomial (presence or absence of the amino acid in CDR3, presence or absence of the amino acid that is sufficiently exposed in CDR3 when within the TCR structure) orcontinuous (frequency of the amino acid in CDR3) was explored. Combinations of 5-8 amino acidswere explored to alleviate overfitting thanks to 5-9 TCRs per explanatory variable (Vittinghoff, E.& McCulloch, C. E. Am. J. Epidemiol. 165, 710–718 (2007)). 277,746 multilinear regressions (MLR) were performed, and the combinations that gave the highest correlation coefficient R2, Docket No.: FR 084276.00403 were selected to be used in logistic regressions. The accuracy of the logistics regressions wasdetermined by the area under the ROC curve (AUC) and by the % of correct predictions, and thebest solution found in the MLR was confirmed to be the optimal solution to be used in the logistic regression. The best model obtained is the one described in the upper equations and has AUC=0.96, very close to 1, emphasizing the ability of the model to discriminate between high and low avidity.The threshold of the classifier was set to 0.5, and a high avidity structure if P>0.5 was predicted.Cross-validations were carried out, illustrating the robustness of the approach. The regressiontrained on the full set was then applied to a library of TCRs determined by single-cell sequencingfor four melanoma patients (Mel7-10), and high and low avidity CD8 TCRs were predicted. Thelibrary of TCRs was then tracked in bulk repertoires of blood and tumors (only β chain TCRinformation was considered). Statistical analyses Statistical analyses were performed with the GraphPad Prism software. Correlation analyses were performed using Pearson coefficient, nonparametric Spearman correlation,nonlinear regression, Mann-Whitney, Wilcoxon-paired, and log-rank tests, which are indicatedthroughout. For cumulative analyses of four melanoma patients (Mel7-10) in Fig. 4B, after pulling allTCRs for a given patient, i.e., the n infiltrating TCRs and m non-infiltrating TCRs, the frequency,F, of HA TCRs in this entire set of n+m TCRs was calculated. Then, for each of the n blood TCRsand m tumor TCRs, a probability of F / (n+m) was randomly chosen as high affinity. Subsequently,the fraction of ‘random’ HA or LA TCRs in the blood and in the tumor was determined. Thisprocess is repeated 1000 times. Finally, the P-value is calculated as the probability of gettinga %HAInf-%HANo value in the random sets equal or higher to the value in the real set.EXAMPLE 2 Neoantigen-specific CD8 T cells are structurally and functionally heterogeneous Neoepitopes are generally considered prototypical tumor rejection antigens. Yet, it remains unclear whether their clinical relevance stems from their tumor specificity alone or whether they truly drive better effector T cells relative to TAAs. To learn more, a library of 371 CD8 T cellclones recognizing 19 neoantigens was generated, TAAs and virus epitopes (Table 1) from 16 Docket No.: FR 084276.00403patients with melanoma, ovarian, lung or colorectal cancer and 6 healthy donors (Table 2), andinvestigated the functional and structural profiles of their TCRs in 190 and 338 clones (Fig. 1B),respectively, in accordance with an example process depicted in Fig. 1A.. Antigen-specific cellswere sorted using double-fluorescent reversible pMHC multimers (i.e., NTAmers), which avoidthe selective loss of high avidity cells. First, T cell structural avidity, intended as the strength of TCR binding to cognate pMHC,was assessed. This was determined through the dissociation kinetic (pMHC-TCR half-life, T1 / 2) of monomeric pMHCs and TCRs. Briefly, rapid decay of reversible pMHC multimers to pMHC monomers allows dissociation rate measurements of fluorescent monomeric pMHC off CD8 Tcells. Polyclonal responses against individual epitopes of any class in most patients or donors weredetected, with marked variance of T1 / 2among clones recognizing the same epitope in each classof antigen specificity. Overall, the structural avidities of neoantigen-specific CD8 T cells werehigher than that of TAA-specific CD8 T cells (Fig. 1C). Similar conclusions were drawn whenexclusively HLA-A*0201-restricted CD8 T cells or unique CDR3 sequences were examined(Table 1). This is the first proof of the long-proposed hypothesis that neoepitope-specific TCRsare of higher structural avidity than T cells directed against “self” tumor antigens (Schumacher, T.N., Scheper, W. & Kvistborg, P. Annu. Rev. Immunol. 37, 173–200 (2019); Balachandran, V. etal. Nature 551, (2017)).Next, the antigen sensitivity of each clone was assessed by IFNγ ELISpot, measuring thepeptide concentration required for half-maximal T cell activation (effect concentration 50%, EC50).Similar to pMHC-TCR T1 / 2, there was an important variance in EC50 among different clonesrecognizing the same peptide for each class of antigens, including for HLA-A*0201 restricted Tcell responses or genetically unique clonotypes. As expected, a positive correlation was observedbetween T1 / 2 and EC50, despite some variability. The latter was attributed to more reproducible and reliable measurements obtained by pMHC-TCR dissociation kinetics when compared to functionalassays, which also depend on T cell intrinsic regulatory mechanisms. To further test thereproducibility of the structural avidity parameter, six TCRαβ chain pairs were cloned into healthyperipheral blood T cells. Measurements of structural avidity remained more consistent betweenoriginal and recipient T cells, maintaining similar ranking between clones, as opposed to antigensensitivity. This supports the robustness of structural avidity as a biophysical parameter to profile T cells. Docket No.: FR 084276.00403 An in vitro pMHC refolding assay (Schmidt, J. et al. Biol. Chem. 292, 11840–11849(2017)) was used to validate the predicted affinity of each peptide for the cognate HLA allele (notshown). The overall ranges of pMHC affinity ruled out any important bias in measurements of antigen sensitivity due to low peptide-MHC interactions. Revealing the limitation of commonlyused algorithms for predicting epitope immunogenicity, there were poor correlations betweenmeasured structural avidity (or antigen sensitivity) with in silico predictors of pMHC affinity,stability or processing, mainly relying on the determination of antigen presentation (Fig. 1D). However, structural avidity was significantly correlated with immunogenicity predicted byPRIME (Schmidt, J. et al. Cell Rep. Med.2, 100194 (2021)) and with pMHC Dissimilarity-to-Self(DisToSelf) (Bjerregaard, A.-M. et al. Front. Immunol.8, 1566 (2017)) (Fig. 1D), also significantwhen viral epitopes were excluded and when only genetically unique clonotypes were considered.PRIME not only considers the binding capacity of a peptide to a given MHC but also integrates its propensity to be recognized by TCRs. DisToSelf determines the similarity (or dissimilarity) of a given peptide with the human proteome. Peptides with high DisToSelf scores are recognized by higher avidity T cells (Fig.1D). High structural avidity neoantigen-specific CD8 T cells reside in tumors Given the unveiled heterogeneity of TCR avidities for given tumor epitopes, whether TCRstrength discriminates cells with a propensity for tumor infiltration was investigated. Indeed, ifhigher avidity cells were to carry an antitumor response, they would be expected to be ratherenriched in the tumor microenvironment. Strikingly, TILs recognizing neoantigen- or TAA-epitope exhibited significantly superior antigen sensitivity relative to cognate peripheral bloodlymphocytes (PBLs) recognizing the same epitope across melanoma, ovarian, colorectal, and lungcancer patients. To assess whether differences in antigen sensitivity could be attributed to structural avidity attributes of TIL vs. PBL clones (Fig.2A), seven pairs of tumor-specific T cells originating fromTILs or PBLs. It was found that the structural avidity of TILs was significantly higher than that ofcognate PBLs across all studied cancers analyzed. Thus, antigen-specific T cells infiltrating tumors, particularly neoantigen-specific clones, display stronger structural avidity than their bloodcounterparts, including when genetically unique clonotypes are considered (Fig. 2B). Docket No.: FR 084276.00403 To better understand the relative enrichment of TILs in high avidity cells, the TCRs ofsorted primary CD8 PBLs and TILs recognizing the same neoepitope from the UTP20 proteinfrom patient Lung1 were sequenced. Neoantigen-specific T cells were oligoclonal, but only threeTCRs were shared between PBLs and TILs. Remarkably, clonotype 5, which was dominant in TILs (58.8% of neoepitope-specific cells), was only contributing to 1.4% of the PBL repertoire, while clonotype 1 was less frequent in TILs (9.6%) but dominant in PBLs (26.1%), and clonotype 3 showed similar frequency in TILs (17.2%) and in PBLs (16.2%). Notably, the structural avidity of UTP20-specific TCRs correlated with their frequency in the tumor compartment, and was the highest for clonotype 5, indicating that tumor-resident clones have higher structural avidity (Fig.2C). It was previously shown that molecular modeling of TCR and pMHC can accurately infer thestrength of their interaction (Bobisse, S. et al. Nat. Commun. 9, (2018)). Here, clone 5 TCRestablished significantly more favorable interactions with UTP20 pMHC than clone 1 TCR. Similar results were obtained in a second example in patient CRC1 for PHLPP2-specific TCRs confirming the association between structural avidity and tumor residence. These observations indicate that the preferential accumulation of high and low avidity clones in tumors and blood, respectively, is also true among clonotypes from the same antigen-specific repertoires. To experimentally validate the preferential tumor infiltration by high structural avidity Tcells (Fig. 3A), a well-characterized panel of NY-ESO-1157-165-specific TCRs with high (DMβ),intermediate (WT) and low (V49I) structural avidity were used. Their avidity covers the range ofviral-, neoantigen- and TAA-specific T cells. CD8 T cells of an HLA-A*0201 donor with DMβ,WT or V49I TCRs were stably transduced, and their structural and functional avidities wereprofiled (Fig.3B). Unlike V49I-transduced T cells, both WT and DMβ variants showed equivalentin vitro responsiveness to HLA-matched Me275 melanoma tumor expressing NY-ESO-1 (Fig.3B).ACT of 5x106T cells in interleukin-2 (IL-2) NOG mice bearing Me275 tumors indicated acorrelation between the in vivo efficacy and the structural but not the functional avidity of TCR-transduced T cells (Fig. 3C). Following ACT, DMβ-transduced CD8 T cells significantly betterinfiltrated tumors as compared to V49I- and WT-transduced cells (Fig. 3D), confirming higherengraftment propensity of high avidity clones. CXCR3-mediated tumor infiltration and control by high avidity T cells Docket No.: FR 084276.00403 Having established a relationship between T cell avidity and tumor homing, it washypothesized that high avidity cells may be endowed with a superior ability for tumor infiltration and retention (Fig.3A). Several studies reported that key chemokine receptors, especially CXCR3,may be required for tumor homing (Lanitis, E., et al. Ann. Oncol. Off. J. Eur. Soc. Med. Oncol.28, xii18–xii32 (2017)). The expression of a panel of chemokine receptors on seven pairs of lowand high avidity antigen-specific CD8 T cells was analyzed. CXCR3 was more strongly expressedand upregulated after short-term stimulation by high as compared to low avidity T cell clones. Thisobservation was specific to tumor homing-related molecules since no significant difference wasfound for chemokine receptors that are not specifically involved in tumor infiltration (e.g., CCR7,not shown). In addition to CXCR3, CD103, and CD49a (VLA-1), two major integrins associatedwith a tissue-residency phenotype, were both upregulated in high avidity clones. Therefore, T cellstructural avidity is associated with CXCR3 expression and, to a lower extend, with CD103 andCD49a expression. Higher CXCR3 expression was also observed on DMβ- relative to V49I- or WT-transducedT cells upon pMHC stimulation in vitro. Of interest, addition of an anti-CXCR3 antibody after ACT in IL-2 NOG mice bearing the Me275 melanoma tumor, known to express CXCR3 ligands,i.e., CXCL9 / 10 / 11 (Neubert, N. J. et al. Cancer Res.77, 1623–1636 (2017)), with DMβ-transducedT cells (Fig. 3A) significantly impaired tumor control (Fig. 3E). Consistently, lower densities of CD8 T cells were observed in animals treated with an anti-CXCR3 blocking antibody post-ACT (Fig. 3F). The inhibition of tumor control after ACT by blocking CXCR3 was further demonstrated in two additional models using neoantigen-specific TCRs. This confirms the critical contribution of CXCR3 in tumor homing and mechanistically links CXCR3 expression with high avidity clones and tumor infiltration. Biophysicochemical inference of tumor-specific T cells that engraft in tumors The above findings collectively indicate that tumor-infiltrating lymphocytes are enrichedin tumor-specific T cell clones endowed with high-avidity TCRs. It was also sought to developfurther methods to infer the avidity of clones for a given epitope (Fig. 4A). Homology modelingwas used to compare TCRs recognizing the same pMHC with high or low structural avidity, applied to five distinct antigens. The number of favorable interactions (bonds) of each TCR with Docket No.: FR 084276.00403 its cognate pMHC, inferred based on the modeled structures of its α and β chains and the cognatepMHC, was consistently higher for high structural avidity TCRs (Table 3), and significantlycorrelated with pMHC-TCR T1 / 2.A major limitation in identifying clinically relevant T cells is the lack of knowledge ofpossible cognate antigens. To solve this, it was hypothesized that high avidity TCRs may sharecommon sequence features (Fig.4A). Furthermore, it has been reported that highly frequent clones among TILs may not be tumor-specific. To overcome these limitations, and driven by the above molecular modeling results, whether specifically high structural avidity TCRs based on theirsequence analysis and without prior knowledge of their specificity can be inferred was investigated.58 individual TCRs recognizing 12 distinct pMHC were selected, for which TCR α and βsequences as well as structural avidities (Table 4) were known, and looked for structural patterns.Biophysical features of k-mers encoded based on the Atchley factors and a generic hierarchicalclustering algorithm (Ostmeyer, J., et al. Cancer Res.79, 1671–1680 (2019); Atchley, W. R., et al.Proc. Natl. Acad. Sci. U. S. A. 102, 6395–6400 (2005)) were used. It was found that CDR3βsequences in high avidity TCRs (T1 / 2 >60s) were significantly enriched in specific amino acidresidues (i.e., N, E, I, K, T, Y, V; all P<0.0001 compared to low avidity TCRs). Conversely, A, R,D, L, M, and P were more frequent in low-avidity TCRs (all P<0.0001) (FIG. 5 and Table 5).Next, hierarchical clustering was developed based on CDR3β motifs. Notably, a hotspotenriched in TCRs with high structural avidity, irrespective of their target, was identified. Thiscomprised 62% of all TCRs of intermediate or high avidity (T1 / 2 >10s), while outside of this cluster, 65% of TCRs had structural avidity <10s. Such enrichment was not observed in a control analysis with 1,000 random clustering, illustrating the significance of this observation (P<0.001) and indicating that some shared common CDR3β features were preferentially associated with higher structural avidity. Guided by this observation, a structure-based logistic regression model was derived topredict the structural avidity of TCRs of unknown specificity (Fig. 4A). It was applied to a panelof 58 TCRs, and it was able to accurately discriminate between high and low-avidity TCRs with an AUC of 0.96, with only one false-positive and one false-negative high avidity among the full set of TCRs (sensitivity of 0.91 and specificity of 0.97). Three cross-validations schemes weresuccessfully performed on viral peptides, TAAs, and neo-antigens and on different HLA alleles to Docket No.: FR 084276.00403 assess the robustness of the predictor, following the standard leave-20%-out protocol (cross- validations 1 and 2) as well as more challenging leave-one-epitope-out cross-validation (cross-validation 3). A success higher than 70% in all cases was achieved, which is significantly betterthan random, providing confidence in the algorithm for the purpose and data described herein. The structure-based logistic regression was then applied to identify high and low avidityTCRs in blood and tumors of four additional melanoma patients (Table 2). When analyzing totaltumor and blood TCR repertoires, it was consistently found enrichment in TCRs predicted to beof high avidity in tumors relative to blood (Fig.4B, P=0.05), therefore validating the preferentialtumor tropism for high avidity TCRs. Furthermore, two TCRs directed against neoantigenKIF1BS918F in enriched TILs from patient Mel8 were also identified (Table 2). Among the twoKIF1BS918FTCRs, and despite no differences in functional avidity, TCR#1 was predicted by the logistic regression to be of high avidity, while TCR#2 was predicted to be of low avidity and thesepredictions were validated experimentally (Fig. 4C). Autologous tumor cells were used to assessthe relative clinical efficacy of the two KIF1BS918F TCRs. Despite the fact that both TCRs weretumor-reactive in vitro, tumor control in mice bearing the autologous tumor was exclusivelyachieved after ACT with T cells transduced with the predicted and validated high avidity TCR but not with the low avidity TCR (Fig. 4D). This prototypical example illustrates the superior tumor infiltration and control by high avidity cells, even when they target the same neoepitope but also highlights the clinical relevance of the logistic regression to predict clinically relevant TCRs regardless of their specificity. Discussion The success of TIL-based immunotherapies for solid tumors relies on strong antitumoralactivity of adoptively transferred T cells. Major efforts were thus made to develop methodsallowing to estimate a priori the functional potential of tumor antigen-specific T cells. The strengthof T cell recognition is a key parameter, as it affects T cell activation, proliferation, infiltration,and effector functions, as well as longevity of T cell responses. Besides the structural avidity ofthe TCR, multiple coreceptors are implicated in determining T cell functional avidity. Cellularassays (cytotoxicity or cytokine production) have been traditionally used to determine antigensensitivity, for which EC50 represents a widely accepted parameter. However, cellular assay resultsdepend on the state of cellular activation or exhaustion, limiting their performance. The Docket No.: FR 084276.00403 dissociation kinetic measurement of pMHC from the TCR, a structural avidity parameter reflectingthe binding strength of a TCR, can be readily applied to viable T cells. Reversible pMHCmultimers can be used to reliably determine such dissociation kinetics, revealing T cell functionality independently of the cellular activation state. Here, tumor antigen-specific T cells in patients with solid tumors were comprehensivelyprofiled and compared with virus-specific cells. To do so, 371 T cell clones upon FACS-sortingwith reversible pMHC multimers were generated, precluding the loss of high avidity T cells proneto TCR-induced cell death induced by conventional pMHC multimers. It was found that virus-specific T cells show higher avidity and function than TAA-specific cells. Neoantigen-specific T cells displayed higher structural avidity than TAA-specific T cells. The superiority of neoantigen- specific T cells over TAA-specific ones was not recapitulated by IFNγ release. Functional avidityassays largely rely on T cell activation states and are more prone to intra- and inter-experimentalfluctuations. T cell responses of high avidity clonotypes can also be inhibited by exhaustionmechanisms. This legitimates structural avidity as a robust and reliable biomarker of T cell responsiveness. However, a broad heterogeneity was observed in the avidity of neoantigen-specific T cells, ranging from low avidities, comparable to TAAs-, to high avidities, comparable to virus-specific cells. A correlation between TCR avidity and T cell potency for the same antigen andacross antigen specificities was observed, providing the rationale for developing prioritizationalgorithms in TCR discovery. Of note, the strength of interaction between effector cells andcognate antigen is becoming an attractive parameter to measure T cell activation and predictefficacy. Notably, the importance of the binding strength between effector cells and tumors was also recently demonstrated with CAR T cells. The broad heterogeneity of TCR avidities supports the notion that neoantigens may strongly differ in their potential to mediate anti-tumor effects in vivo. Identification of neoantigensrelies on in silico prediction of antigen binding avidity to MHC molecules, with a discovery rate<5%, arguing that only a minor fraction of presented peptides are immunogenic. Indeed, no correlation was found between functional or structural parameters and prediction of peptide binding affinities, but a correlation was found with immunogenicity prediction through PRIME. This presumably reflects the importance of the mutation occurring at MHC anchor residues or directly in those in contact with the TCR, which is ultimately captured by molecular modeling. Docket No.: FR 084276.00403 It has been reported that clonally expanded T cells can reside in tumor tissue and adjacent normal tissue or blood. Here, by analyzing T cells from blood and tumor targeting the same tumorantigen, it was found that the latter consistently showed higher antigen sensitivity and structuralavidity. However, while TILs are enriched for high avidity T cell populations, common TCRclonotypes were also identified in PBL (albeit at lower frequencies), consistently with the presenceof tumor-reactive TIL clonotypes in the circulation. The association between structural avidity andtumor infiltration was seen across multiple epitopes and multiple patients, but also within antigen classes (TAAs and neoantigens) and within distinct clonotypic repertoires of neoantigen-specific T cells. Furthermore, despite the fact that T cells from both blood and tumors were systematically interrogated for each patient and each antigen, neoepitopes and TAAs were preferentially detectedin TIL and PBLs, respectively, consistently with the superior structural avidity of neoepitope-specific T cells. The high frequency of neoepitope-specific TCRs in tumors was recently shown in patients with lung cancer. It was also found that structural avidity is associated with CXCR3 expression, known topromote tumor infiltration, as well as CD103 (αEβ7) and CD49a (VLA-1) expression, both associated with tumor residency. CXCR3 blocking after ACT prevented tumor infiltration of high avidity tumor-specific T cells. Tumor infiltration and eradication probably rely on several othercomplementary parameters, like T cell expansion and persistence. A major issue in ACT therapyis the downregulation of antigen presentation, which could indeed favor high avidity T cells that are less dependent on antigen concentration. This was recently suggested in melanoma patients where TAA-related antigen expression was higher than that of neoantigen and was inverselycorrelated with the functional avidity of the respective antigen-specific T cells (Oliveira, G. et al.Nature 1-7 (2021)). The unique capacity to comprehensively profile CD8 T cells by measuring structuralavidity, linked to TCR biophysicochemical and sequence features, as well as structure modeling,allowing for building a novel structure-based logistic regression model of TCR avidity levelprediction of TCRs with unknown antigen specificity. To prevent overfitting of the model, thenumber of parameters entering the equation was limited, and several cross-validation tests wereperformed successfully to assess the robustness of the approach, following the standard leave- 20%-out protocol as well as a more challenging leave-one-epitope-out cross-validation. Thispredictor was able to identify TCRs with high avidity features in four melanoma patients. These Docket No.: FR 084276.00403 were found more frequently within tumors than blood at steady state, supporting the notion that high avidity TCRs preferentially home and reside in tumors. The data link neoantigen recognition, T cell functionality, and ability to infiltrate and residein tumors, indicating that the clinical relevance of neoantigen-specific T cells is not only relatedto their tumor-specificity but also to their higher functionality and their preferential ability toinfiltrate tumors. The data also indicate that tumor-specific CD8 T cells are highly heterogeneousand that measurements of structural avidity can be used for better selection of clinically relevantT cells, avoiding the use of poorly functional clonotypes, both for TAA- and neoantigen-specificT cells. High avidity T cells (i.e., preferentially TILs) should therefore be prioritized forpersonalized therapies, including TCR-based immunotherapy. EXAMPLE 3 Structural avidity is a robust and stable physical parameter. To assess any potential bias of cell culture condition in the measurements of antigen sensitivity and / or structural avidity, six pairs of TCRαβ chains (one high and one low functionality TCR for three distinct pMHCs: two neoantigens (SLC25A48 and PHLPP2) and one TAA (Melan-A) were sequenced and cloned, and naïve primary CD8 T cells from healthy donors were transduced with these TCRs, as described(Irving, M. et al. J. Biol. Chem. 287, 23068–23078 (2012).). The comparison of transduced cellsrelative to native T cells for each pair of TCRs shows some degree of variability in measurements of antigen sensitivity, yet the ranking of high vs low functionality cells for each antigen was maintained for all three pairs except PHLPP2 while the measurement of structural avidity remains more quantitatively consistent, confirming the robustness of this physical parameter to profile T cells. Features displaying consecutive G or several T are seen in the high avidity cluster region(GGGT (SEQ ID NO: 308) / GGGR (SEQ ID NO: 309) in TCR models #45, #48, #55, #58 orTDTQ (SEQ ID NO: 457) in TCR models #3, #43, #50, #53, Table 4), in agreement with the factthat G and T are more frequent in high avidity structures (17.7% and 10.4%, respectively) when compared to low avidity (17.0% and 8.9%, respectively). G and T are more frequent inside thecluster, with frequencies dropping by 5% and 4%, respectively, when outside the cluster (Table 5,all P<0.0001). Outside the cluster, high avidity structures with consecutive G, such as the virus-specific TCR (#4) with highest structural avidity (CDR3β is CASMGGAYNEQFFG (SEQ ID Docket No.: FR 084276.00403 NO: 310)) and the neoantigen-specific TCR (#17) with highest structural avidity (CDR3β isCASSITTSGGYEQYFG (SEQ ID NO: 311)) were observed. These two TCRs do not clustertogether and are not in the clustered box, because the cluster analysis is based on 4-mer featuresand not in 2-mer features. Three other TCRs (i.e., #36, #37 and #69) were of interest: with #36 ofhigh avidity, #37 of intermediate avidity and #69 of low avidity for the same pMHC. For these three TCRs, #36 contains 4 G in the CDR3β including 2 consecutive ones; #37 contains 4 non- consecutive G and #69 only contains 3 non-consecutive G. These observations indicate that high and intermediate avidity structures do exhibit a preference for G and more specifically for consecutive G. Nevertheless, this pattern is not exclusive from high and intermediate avidity TCRs as the G does not have any side chain and therefore establishes limited interactions with the peptide, justifying the presence of some low energy structures in the clustered region. High frequency of G in high avidity structures could be explained by the fact that it allows for loop rearrangements, sidechain flexibility of the residues nearby, and, consequently, more specific and strong interactionsof the nearby residues with the peptide. Amino acids N, E, I, T and Y show the largest variationsin frequencies between high and low-avidity CDR3β, all P<0.0001) (FIG. 5). Globally in the setof 58 complexes studied here, higher avidity between TCR and pMHC seems to be acquired by anenrichment in G that increases TCRs flexibility and adaptability together with enrichment in N, E, I, T and Y that can favor hydrophilic (H-bonds), hydrophobic as well as aromatic interactions with aromatic residues in the peptide (Π-interactions), depending on the peptide nature. Residues N and I contribute with positive weights (P<0.04) to the avidity prediction, inagreement with their higher frequencies in high avidity structures (FIG. 5 and Table 5), whileremaining amino acids in the model have negative weights as they are more frequent in low avidity structures, with the exception of G which is more frequent in the CDR3β of high avidity TCRs(FIG. 5 and Table 5), but more accessible to the solvent in low avidity TCRs.FIG. 5 details amino acid frequencies in the CDR3β independently of their 3D exposure.However, in the structure-based logistic regression, significantly better models were obtained byrepresenting the presence (1) or the absence (0) of the amino acid with a normalized solvent accessibility higher than 30% in the CDR3β of the TCR 3D model (the amino acids that have more chances to contact the peptide). Considering solvent accessibility of the residues, among others, G is now more frequent in low avidity structures. The decrease in G is expected since the expected role of these residues is to allow loop rearrangements that help strong-interacting residues nearby Docket No.: FR 084276.00403(such as N, E, I, T, and Y) to be displayed in front of the pMHC. As a result, G residues arehindered, and exhibit decreased solvent accessibility, and consequently, they are themselvesexposed with lower frequency. Three independent cross-validations were carried out: First cross-validation First cross-validation was carried out by randomly splitting the full set into 10 sets of 38 training TCRs (3 high avidity and 7 low avidity) and 10 testing TCRs (8 high avidity and 30 low avidity). For each of the 10 combinations, the model (logistic regression) was trained from the training set and tested the predictions on the test set. Among the 10 sets, the average accuracy was 0.76±0.01. Despite the fact that the training and test sets were significantly smaller, there was a significant correlation, illustrating the robustness of the approach. Notably, the different trainingand test sets include different compositions of neo-antigens, viral, TAA, and different alleles,highlighting the broad applicability of the model. Second cross-validation For the same 10 sets, the cross-validation was re-performed, this time including the featureselection step in the process. The amino acids (features) that were selected for the 10 training setsare shown in Table 6.As can be seen, RNDGIL (SEQ ID NO: 312) and F are the features selected in 6 / 10 setsand RNDL (SEQ ID NO: 458) and F were automatically selected as being the most relevantfeatures in all the 10 sets. In addition, the sign of the coefficient of these features remained the same in all 10 cross-validations (data not shown). These results support the robustness of the model and that the CDR3β amino acid representation used is the most relevant for the data under study. Among the 10 test sets, the average accuracy was 0.76±0.01. Third cross-validation A leave-one-epitope-out cross validation was also done, where the model was trained on all TCRs except those recognizing one given epitope, before applying the new model on TCRs recognizing this epitope. For example, the data were trained on a set excluding the viral peptideGLCTLVAML / HLA-A*02:01 (SEQ ID NO: 313), and then it was applied on a set that onlyincludes the TCRs recognizing the GLCTLVAML / HLA-A*02:01 (SEQ ID NO: 314). Then, the Docket No.: FR 084276.00403data were trained on the set excluding the neo antigen GRKLFGTHF / HLA-B*27:05 (SEQ ID NO:315), and it was applied on a set that only includes TCRs recognizing GRKLFGTHF / HLA-B*27:05 (SEQ ID NO: 316). It was done successively for the 12 epitopes under discussion.Ten leave-one-epitope-out sets, among the 12, converged to the same features (with thesame negative and positive signs) as in the full set: R, N, D, G, I, L, and F. The two sets thatconverge to different variables are the ones leaving the EAAGIGILTV (SEQ ID NO: 317) and theNILDAIAEI (SEQ ID NO: 318) epitopes out. However, they still keep 4 / 7 and 6 / 7 common features, respectively, with the full set. N and I, highlighted in yellow, are the unique amino acids contributing positively to the prediction in all 12 sets. When EAAGIGILTV (SEQ ID NO: 319) (the most frequent epitope) was removed fromthe training set, the smallest and most different training set was created. Here, 32% of the TCRsused in the full set were removed, all of them with low avidity, making the situation the mostdifferent compared to the full set, and, therefore, the most challenging. Despite this, the featuresR, N, G and I in the predictor were still kept, with the same positive and negative signs. The overallaccuracy leaving this epitope out is 88%, the smallest among all the sets but still satisfactory. The accuracy of the avidity predictions for TCRs recognizing this epitope, even if they were not trained on this epitope and they are all of low avidity, is 67%. The present cross-validation shows thatthere is predictive power for new epitopes, even in extreme cases, as the one excludingEAAGIGILTV (SEQ ID NO: 456).For these epitope-out cross validations, the overall average prediction is 72% for the validation sets (predictions for the validation epitope, the epitope not used in training). Thepredictions range from 25% to 100%. Of note, the number of TCRs per epitope ranges from 1 to12, with an average of 4 and a median of 3. So, in some cases, one single wrong prediction in thevalidation sets, drastically decreases the accuracy. All these three cross-validation schemes supportthe reliability of the predictor for the data from this study. The predictor is relevant for TAA,neoantigen, and viral peptides and different HLA alleles.The present disclosure is not to be limited in scope by the specific embodiments describedherein. Indeed, various modifications of the invention in addition to those described herein willbecome apparent to those skilled in the art from the foregoing description and the accompanyingfigures. Such modifications are intended to fall within the scope of the appended claims. Docket No.: FR 084276.00403Table 1. Specificity, antigen sensitivity, and structural profile of CD8 T cell clones.Patient / Protein Epitope SE SNV IDOrigi EC50 pMHClonotype (TCR - CDR3β chain) SEQHD code Q Clo n (M) C- ID ID ne TCR NO NO T1 / 2 (s) 1PBL 3.66 2.9 hTRBV20_CSARDVGLGIYEQYFG_hTR1 E-07 BJ02-7 2PBL - 2.9 hTRBV13_CASSLDPSGSPNEQFFG_hTR2 BJ02-1 3PBL 8.47 3.2 hTRBV20_CSARDVGLGIYEQYFG_hTR3 E-08 BJ02-7 4PBL 7.65 3.1 hTRBV03-4 E-08 1_CASSQGDLAWIPTEAFFG_hTRBJ01- 1 5PBL 2.54 4 hTRBV13_CASSLDPSGSPNEQFFG_hTR5 E-08 BJ02-1 6PBL 1.02 4.6 NDE-07 7PBL 1.10 3.4 hTRBV20_CSARDVGLGIYEQYFG_hTR6 E-08 BJ02-7 8TIL 4.03 4 hTRBV20_CSASPGLAEQFFG_hTRBJ02-7 E-08 1 9TIL 3.15 3.7 hTRBV14_CASSQDTGLSSYNEQFFG_h8 E-08 TRBJ02-1 10 TIL 8.20 4.4 hTRBV06-9 E-10 1_CASSELGLAGNEQFFG_hTRBJ02-1 11 TIL 1.24 4 hTRBV20_CSAERGLGQPQHFG_hTRBJ10 E-07 01-5 Mel1 Melan-A EAAGIGILT 152 12 TIL 3.59 4.1 hTRBV19_CASTSGELGQPQHFG_hTRB11 V E-09 J01-5 13 TIL 1.23 3.6 hTRBV27_CASSLSGLAGVEQYFG_hTR12 E-07 BJ02-7 14 TIL 1.28 36.3 hTRBV19_CASKWGALMNTEAFFG_hT13 E-07 RBJ01-1
[0002] Docket No.: FR 084276.00403 15 TIL 7.68 5 hTRBV27_CASSSLGATYEQYFG_hTRB14 E-08 J02-7 16 TIL 1.66 3.7 hTRBV27_CASSWTSGSPSEQFFG_hTR15 E-09 BJ02-1 17 TIL 1.17 3.9 hTRBV27_CASSLFSGSSGELFFG_hTRB16 E-09 J02-2 18 TIL 5.33 16.6 NDE-08 19 TIL 2.49 3.1 hTRBV03-17 E-07 1_CASSQGSLAGSEQYFG_hTRBJ02-7 20 TIL 6.81 7.9 hTRBV28_CASRVQGLGQPQHFG_hTR18 E-09 BJ01-5 21 TIL 1.00 3.6 hTRBV06-19 E-06 1_CASSELGLAGNEQFFG_hTRBJ02-1 22 TIL 6.47 5.9 hTRBV06-20 E-09 1_CASSELGLAGNEQFFG_hTRBJ02-1 1PBL - 37.4 hTRBV12-21 3_CASSLDRATNEKLFFG_hTRBJ01-4Mel2 MAGE EVDPIGHLY 153 2 PBL - 47.8 hTRBV12-22 A3 3_CASSLDRATNEKLFFG_hTRBJ01-4 3PBL - 37.1 hTRBV12-23 3_CASSLDRATNEKLFFG_hTRBJ01-4 4PBL - 46.8 hTRBV12-24 3_CASSLDRATNEKLFFG_hTRBJ01-4Mel3 NBEA LPQARRILL 154 S2272 1 PBL 1.03 83.8 NDL E-09 2PBL 7.06 - NDE-09 1PBL 9.55 13.5 NDE-09 2PBL 1.49 18 NDE-09 3PBL 3.59 4.8 hTRBV19_CASSMGQLILGYEQYFG_hT25 E-08 RBJ02-7
[0003] Docket No.: FR 084276.00403 4PBL 4.15 8.8 hTRBV02_CASSELERLKVYNSPLHFG_26 E-09 hTRBJ01-6 5PBL 3.83 3.5 hTRBV19_CASSMGQLILGYEQYFG_hT27 E-09 RBJ02-7 6PBL - 5.6 hTRBV19_CASSMGQLILGYEQYFG_hT28 RBJ02-7 7PBL - 5.2 hTRBV19_CASSMGQLILGYEQYFG_hT29 RBJ02-7Mel4 GP100 ITDQVPFSV 155 8 PBL - 9.4 hTRBV02_CASSELERLKVYNSPLHFG_30 hTRBJ01-6 9PBL - 13.4 hTRBV19_CASSMGQLILGYEQYFG_hT31 RBJ02-7 10 TIL 7.70 82.1 NDE-11 11 TIL 8.89 63.3 NDE-10 12 TIL - 68.1 ND13 TIL 1.56 83.6 NDE-10 14 TIL - 70.8 hTRBV19_CASSITTSGGYEQYFG_hTR32 BJ02-7 1PBL 2.07 6.4 hTRBV10-33 E-07 2_CASSYRGNSPLHFG_hTRBJ01-6Mel5 GP100 ITDQVPFSV 459 2 PBL 5.35 3.8 hTRBV19_CASSLRLAATIYNEQFFG_hT34 E-06 RBJ02-1 3PBL 7.33 9.4 hTRBV19_CASSARGYASPLHFG_hTRB35 E-07 J01-6 1PBL - 2.8 hTRBV02_CASIGLAKNIQYFG_hTRBJ036 2-4 2PBL - 2.2 hTRBV11-37 1_CASSFGGSSYEQYFG_hTRBJ02-7 3PBL - 3.3 hTRBV09_CASSPVWAGAYNEQFFG_h38 TRBJ02-1 4PBL - 2.7 hTRBV09_CASSPVWAGAYNEQFFG_h39 TRBJ02-1
[0004] Docket No.: FR 084276.00403 5PBL - 2.8 hTRBV09_CASSPVWAGAYNEQFFG_h40 TRBJ02-1 6PBL - 2.1 hTRBV02_CASIGLAKNIQYFG_hTRBJ041 2-4 7PBL - 2.3 hTRBV02_CASIGLAKNIQYFG_hTRBJ042 2-4 8PBL - 2.7 hTRBV11-43 1_CASSFGGSSYEQYFG_hTRBJ02-7Mel6 MAGEGLYDGMEH156 9 PBL 9.92 2.8 NDA10 L E-08 10 PBL 1.34 2.2 hTRBV02_CASIGLAKNIQYFG_hTRBJ044 E-07 2-4 11 PBL 8.56 2.5 NDE-10 12 PBL 6.92 2.3 hTRBV11-45 E-10 1_CASSFGGSSYEQYFG_hTRBJ02-7 13 PBL 1.67 2.7 hTRBV09_CASSPVWAGAYNEQFFG_h46 E-09 TRBJ02-1 14 PBL 1.03 2.8 hTRBV09_CASSPVWAGAYNEQFFG_h47 E-08 TRBJ02-1 15 PBL 6.81 2.4 hTRBV11-48 E-08 1_CASSFGGSSYEQYFG_hTRBJ02-7 16 PBL 2.69 2.3 hTRBV11-49 E-09 1_CASSFGGSSYEQYFG_hTRBJ02-7 17 PBL 1.12 1.6 NDE-07 1PBL 3.14 43.3 hTRBV05-50 E-08 4_CASTLSTGQGIYGYTFG_hTRBJ01-2 2PBL 6.02 53.2 hTRBV05-51 E-08 4_CASTLSTGQGIYGYTFG_hTRBJ01-2 3PBL 1.25 43.9 hTRBV05-52 E-07 4_CASTLSTGQGIYGYTFG_hTRBJ01-2 4PBL 9.67 42.6 NDE-08
[0005] Docket No.: FR 084276.00403 5PBL 1.06 46.3 hTRBV05-53 E-06 4_CASTLSTGQGIYGYTFG_hTRBJ01-2 6PBL 6.76 48.5 hTRBV05-54 E-07 4_CASTLSTGQGIYGYTFG_hTRBJ01-2 7PBL 1.53 48.9 NDE-07 8PBL - 51.8 hTRBV05-55 4_CASTLSTGQGIYGYTFG_hTRBJ01-2 9PBL - 55.7 NDCRC1 PHLPP2 QSDNGLDS 157 D1186 10 TIL 5.12 76.4 hTRBV10-56 DY2Y E-09 3_CAISGGSVGEQYFG_hTRBJ02-7 11 TIL 1.09 75.2 hTRBV10-57 E-08 3_CAISGGSVGEQYFG_hTRBJ02-7 12 TIL 4.44 80.2 NDE-09 13 TIL 8.75 75.8 NDE-09 14 TIL - 48.5 ND15 TIL - 65.4 hTRBV10-58 3_CAISGGSVGEQYFG_hTRBJ02-7 16 TIL - 64.4 hTRBV10-59 3_CAISGGSVGEQYFG_hTRBJ02-7 17 TIL - 77.4 ND18 TIL - 56.2 ND19 TIL 4.45 69 hTRBV10-60 E-09 3_CAISGGSVGEQYFG_hTRBJ02-7 20 TIL - 62.2 hTRBV10-61 3_CAISGGSVGEQYFG_hTRBJ02-7 21 TIL - 51.6 hTRBV10-62 3_CAISGGSVGEQYFG_hTRBJ02-7 22 TIL - 63.1 hTRBV10-63 3_CAISGGSVGEQYFG_hTRBJ02-7 23 TIL - 83.7 hTRBV10-64 3_CAISGGSVGEQYFG_hTRBJ02-7
[0006] Docket No.: FR 084276.00403 24 TIL - 63.9 hTRBV10-65 3_CAISGGSVGEQYFG_hTRBJ02-7 25 TIL - 84.6 ND26 TIL - 66.4 hTRBV10-66 3_CAISGGSVGEQYFG_hTRBJ02-7 27 TIL - 59.3 hTRBV10-67 3_CAISGGSVGEQYFG_hTRBJ02-7 28 TIL - 75.1 hTRBV10-68 3_CAISGGSVGEQYFG_hTRBJ02-7 29 TIL - 60.3 hTRBV10-69 3_CAISGGSVGEQYFG_hTRBJ02-7CRC1 PHLPP2 QSDNGLDS 158 D1186 30 TIL 8.30 90.6 hTRBV10-70 DY2Y E-09 3_CAISGGSVGEQYFG_hTRBJ02-7 31 TIL - 58.2 hTRBV10-71 3_CAISGGSVGEQYFG_hTRBJ02-7 32 TIL - 84.2 ND33 TIL - 80.6 ND34 TIL 9.25 52.7 hTRBV10-72 E-09 3_CAISGGSVGEQYFG_hTRBJ02-7 35 TIL - 91.2 ND36 TIL 4.95 84.8 hTRBV10-73 E-09 3_CAISGGSVGEQYFG_hTRBJ02-7 37 TIL 1.39 75.5 hTRBV10-74 E-09 3_CAISGGSVGEQYFG_hTRBJ02-7 38 TIL 5.35 86.4 hTRBV10-75 E-09 3_CAISGGSVGEQYFG_hTRBJ02-7 39 TIL 3.58 90.8 NDE-09 1PBL 1.00 34.5 NDE-09 2PBL - 29.2 ND3 PBL 3.49 51.8 NDE-10 4PBL 9.94 30.4 NDE-11
[0007] Docket No.: FR 084276.00403 5PBL 3.27 35.6 NDE-09 6PBL - 29.8 NDCRC2 NUP210 GLQAILVH 159 E849 7 PBL - 40.8 NDV V 8PBL - 23.1 ND9 PBL - 34.4 ND10 PBL - 29.3 ND11 PBL - 46.8 ND12 PBL - 33.2 ND13 PBL - 45.4 ND1 PBL 5.80 7.5 hTRBV04-76 E-07 2_CASSQDAETQYFG_hTRBJ02-5 2PBL 1.00 3 hTRBV04-77 E-07 2_CASSQDAETQYFG_hTRBJ02-5 3PBL 1.40 9.5 hTRBV04-78 E-07 2_CASSQDAETQYFG_hTRBJ02-5 4TIL 1.44 20.5 hTRBV12-79 E-08 3_CASSRTSPTDTQYFG_hTRBJ02-3 5TIL 8.18 31.8 hTRBV12-80 E-09 3_CASSRTSPTDTQYFG_hTRBJ02-3 6TIL 1.45 - NDE-08 7TIL 1.51 - NDE-08 8TIL 3.50 - NDE-09 9TIL 9.29 - NDE-09OvCa1 HHAT 10 TIL 7.03 - NDE-09 KQWLVWL160 L75F 11 TIL 9.87 - NDFL1E-09 12 TIL 1.72 - NDE-09
[0008] Docket No.: FR 084276.00403 13 TIL 3.52 - NDE-09 14 TIL 2.68 - NDE-09 15 TIL 2.34 - NDE-09 16 TIL 5.46 - NDE-09 17 TIL 5.02 - NDE-09 18 TIL 3.52 - NDE-09 19 TIL 4.26 - NDE-09 1PBL 8.73 - NDE-09 2PBL 2.01 4 hTRBV07-81 E-08 8_CASSWDSGYEQYFG_hTRBJ02-7 3PBL 3.66 - NDE-08 4PBL 1.52 3.7 hTRBV07-82 E-08 8_CASSWDSGYEQYFG_hTRBJ02-7 5PBL 1.79 3.9 hTRBV07-83 E-08 8_CASSWDSGYEQYFG_hTRBJ02-7OvCa2 ZCCHCGRKLFGTH161 P1265 6 PBL 1.37 3.6 hTRBV07-84 11 F1H E-08 8_CASSWDSGYEQYFG_hTRBJ02-7 7PBL 1.03 3.8 hTRBV07-85 E-08 8_CASSWDSGYEQYFG_hTRBJ02-7 8PBL 2.08 - NDE-09 9TIL 2.94 153.6 hTRBV07-86 E-08 8_CASSLDIGTYEQFFG_hTRBJ02-1 10 TIL 7.74 - hTRBV28_CASSLAGLNTEAFFG_hTRBJ87 E-08 01-1 11 TIL 1.15 - NDE-07
[0009] Docket No.: FR 084276.00403OvCa3 SCL25APYMFLSEW162 V200 1 PBL 7.90 2.6 hTRBV19_CASSIARVTEAFFG_hTRBJ088 48 I M E-07 1-1 2PBL 1.02 61.4 hTRBV19_CASSIGKTGKLFFG_hTRBJ089 E-09 1-4OvCa4 MAGEKVLEYVIK163 1 PBL - 9 NDA1 V 1PBL 3.20 43.1 hTRBV07-90 E-08 3_CASSVGSYNEQFFG_hTRBJ02-1 2PBL 1.20 39.3 hTRBV07-91 E-08 3_CASSVGSYNEQFFG_hTRBJ02-1OvCa5 MUC1 VLVCVLVA 164 3 PBL 3.90 62.2 NDL E-09 4PBL 8.90 40.1 hTRBV07-92 E-09 3_CASSVGSYNEQFFG_hTRBJ02-1 5PBL 5.40 39.2 hTRBV07-93 E-08 3_CASSVGSYNEQFFG_hTRBJ02-1 1PBL 4.24 12 NDE-07 2PBL - 12.7 ND3 PBL - 7.9 ND4 PBL 4.81 6.4 hTRBV05-94 E-07 4_CASFFSGGGTDTQYFG_hTRBJ02-3 5PBL - 5.9 hTRBV20_CSASRADTYEQYFG_hTRBJ95 02-7Lung1 MMP9 NILDAIAEI 165 F621L 6 PBL - 29.6 hTRBV07-96 2_CASTDTDTQYFG_hTRBJ02-3 7PBL 3.41 9.9 hTRBV05-97 E-07 4_CASFFSGGGTDTQYFG_hTRBJ02-3 8PBL - 12.7 hTRBV10-98 2_CASSLGQETQYFG_hTRBJ02-5 9PBL 4.22 10.7 NDE-07 10 PBL - 6.4 ND11 TIL 1.79 50.8 hTRBV07-99 E-09 9_CASSPIAGGTDTQYFG_hTRBJ02-3
[0010] Docket No.: FR 084276.00403 12 TIL 9.97 17.7 NDE-09 13 TIL 4.94 50.9 hTRBV07-100 E-10 9_CASSPIAGGTDTQYFG_hTRBJ02-3 14 TIL 1.04 16.3 NDE-09 15 TIL 1.18 25.1 NDE-09 16 TIL 2.71 41 hTRBV07-101 E-09 9_CASSPIAGGTDTQYFG_hTRBJ02-3Lung1 MMP9 NILDAIAEI 166 F621L 17 TIL 1.80 - NDE-08 18 TIL 4.80 - NDE-09 19 TIL 1.28 - NDE-08 20 TIL 9.26 - NDE-09 21 TIL 7.79 - hTRBV05-102 E-09 6_CASSLGGGRDEQYFG_hTRBJ02-7 22 TIL 1.24 - NDE-08 1PBL 7.21 10 hTRBV10-390 E-10 1_CASSDSTAKETQYFG_hTRBJ02-5 2PBL 2.26 61.2 hTRBV04-103 E-08 3_CASSQEESYEQYFG_hTRBJ02-7Lung1 UTP20 AMDLGIHK 167 D2661 3 PBL 2.01 54.7 hTRBV05-104 V H E-09 6_CASSLGGGRDTQYFG_hTRBJ02-3 4PBL 4.55 - NDE-09 5TIL 7.20 77.7 hTRBV11-105 E-10 1_CASSFQTGWNEQFFG_hTRBJ02-1 6TIL - 88.8 ND1 PBL - 293 ND
[0011] Docket No.: FR 084276.00403Lung2 Influenz GILGFVFTL 168 2 PBL - 138.5 NDa A MP 3PBL 2.39 - NDE-10 1PBL - 27.6 ND2 PBL - 44.6 ND3 PBL - 27.3 ND4 PBL - 18.2 ND5 PBL 4.96 18 NDE-09 6PBL 8.48 21.3 NDE-09 7PBL 1.54 25.4 NDE-09 8PBL 2.30 21 NDE-09 9PBL 6.84 33.7 NDE-09 10 PBL 9.12 20.5 NDE-09 11 PBL 6.41 25.8 NDE-09 12 PBL 8.87 17.3 NDE-09 13 PBL 4.22 22.8 NDE-09 14 PBL 2.00 16.1 NDE-09 15 PBL 1.24 31.5 NDE-09Lung3 EBVGLCTLVAM169 16 PBL 8.56 71.8 NDBMLF1 L E-11 17 PBL 2.25 42.6 NDE-09
[0012] Docket No.: FR 084276.00403 18 PBL 4.93 66.2 NDE-10 19 PBL - 33.8 ND20 PBL - 21.5 ND21 PBL - 47.9 ND22 TIL 6.21 85.8 NDE-10 23 TIL - 71.2 ND24 TIL - 60.3 ND25 TIL - 64.5 ND26 TIL 9.65 87.4 NDE-10 27 TIL 7.41 99.8 NDE-10 28 TIL - 77.8 ND29 TIL - 69.7 ND30 TIL 3.16 56.3 NDE-09 1PBL 1.15 9 hTRBV19_CATSGRSGDTQYFG_hTRBJ106 E-07 02-3HD1 InfluenzVSDGGPNL170 2 PBL 9.78 20.9 hTRBV11-107 a A PB1 Y E-08 2_CASSLDGQGPLYGYTFG_hTRBJ01-2 3PBL - 20.4 hTRBV19_CASSTRSSYEQYFG_hTRBJ0108 2-7 1PBL 9.10 83.7 NDE-13 2PBL - 72.6 ND3 PBL - 86.5 ND4 PBL - 54.1 ND5 PBL - 19.9 ND6 PBL - 47.2 ND7 PBL - 101.3 ND8 PBL 8.90 82 NDE-13 9PBL - 85 ND
[0013] Docket No.: FR 084276.00403 10 PBL - 83.3 ND11 PBL 4.40 69.3 NDE-12 12 PBL 2.80 89.1 NDE-13 13 PBL 8.00 72.2 NDE-12 14 PBL - 87 NDHD2 Influenz GILGFVFTL 171 15 PBL - 66.6 NDa A MP 16 PBL - 38.1 ND17 PBL - 119.7 ND18 PBL - 66.7 ND19 PBL - 53.9 ND20 PBL - 70.1 ND21 PBL - 62.1 ND22 PBL - 61.8 ND23 PBL - 50.2 ND24 PBL - 83.7 ND25 PBL - 64.1 ND26 PBL - 191.2 ND27 PBL - 74.3 ND28 PBL - 71.4 NDHD3 EBVGLCTLVAM172 1 PBL 7.40 102.7 hTRBV07-109 BMLF1 L E-09 3_CASSPGGQSTDTQYFG_hTRBJ02-3 2PBL - 81.3 hTRBV06-110 8_CASSENVGIGANVLTFG_hTRBJ02-6 3PBL - 68.5 hTRBV06-111 8_CASSENVGIGANVLTFG_hTRBJ02-6 4PBL - 108.4 ND5 PBL 2.78 98.6 hTRBV07-112 E-09 3_CASSPGGQSTDTQYFG_hTRBJ02-3 6PBL - 95.4 hTRBV06-113 8_CASSENVGIGANVLTFG_hTRBJ02-6
[0014] Docket No.: FR 084276.00403 7PBL - 62.5 hTRBV06-114 8_CASSENVGIGANVLTFG_hTRBJ02-6 8PBL - 67.1 hTRBV06-115 8_CASSENVGIGANVLTFG_hTRBJ02-6 9PBL 3.10 79.2 hTRBV06-116 E-10 8_CASSENVGIGANVLTFG_hTRBJ02-6 10 PBL 3.52 90.3 hTRBV06-117 E-10 8_CASSENVGIGANVLTFG_hTRBJ02-6 11 PBL - 80.1 hTRBV06-118 8_CASSENVGIGANVLTFG_hTRBJ02-6 12 PBL - 65 hTRBV06-119 8_CASSENVGIGANVLTFG_hTRBJ02-6 13 PBL 1.66 81.2 hTRBV07-120 E-10 3_CASSPGGQSTDTQYFG_hTRBJ02-3 14 PBL - 80.2 ND15 PBL 1.05 15.5 hTRBV20_CSARDRGLGNTIYFG_hTRB121 E-08 J01-3 16 PBL 5.25 88.5 hTRBV06-122 E-09 8_CASSENVGIGANVLTFG_hTRBJ02-6 17 PBL 9.58 71 hTRBV07-123 E-10 3_CASSPGGQSTDTQYFG_hTRBJ02-3 18 PBL 6.37 15.7 hTRBV20_CSARDRGLGNTIYFG_hTRB124 E-09 J01-3 19 PBL 2.10 109.4 NDE-09 20 PBL 1.19 17.1 hTRBV20_CSARDRGLGNTIYFG_hTRB125 E-08 J01-3 21 PBL 1.05 61.7 hTRBV07-126 E-10 3_CASSPGGQSTDTQYFG_hTRBJ02-3 22 PBL - 185.3 NDHD3 EBVGLCTLVAM173 23 PBL 4.53 68.4 NDBMLF1 L E-11 24 PBL 1.37 68.3 hTRBV07-127 E-10 3_CASSPGGQSTDTQYFG_hTRBJ02-3
[0015] Docket No.: FR 084276.0040325 PBL 3.18 108.6 NDE-1026 PBL - 74.3 ND27 PBL - 76.2 hTRBV06-128 8_CASSENVGIGANVLTFG_hTRBJ02-628 PBL - 88.4 hTRBV06-129 8_CASSENVGIGANVLTFG_hTRBJ02-629 PBL 1.77 101.5 hTRBV07-130 E-10 3_CASSPGGQSTDTQYFG_hTRBJ02-330 PBL - 91.5 ND31 PBL - 89.8 ND32 PBL - 62.2 hTRBV06-131 8_CASSENVGIGANVLTFG_hTRBJ02-633 PBL - 69.7 ND34 PBL - 76.8 hTRBV06-132 8_CASSENVGIGANVLTFG_hTRBJ02-635 PBL - 53.9 ND36 PBL - 52.4 ND37 PBL - 57.1 ND38 PBL - 62.9 hTRBV06-133 8_CASSENVGIGANVLTFG_hTRBJ02-639 PBL 7.40 84.5 hTRBV06-134 E-09 8_CASSENVGIGANVLTFG_hTRBJ02-640 PBL - 56.9 ND41 PBL - 89.3 hTRBV06-135 8_CASSENVGIGANVLTFG_hTRBJ02-61 PBL 1.15 102.3 hTRBV7-136 E-11 6_CASSLAPGATNEKLFFG_hTRBJ1-42 PBL 1.10 111.3 hTRBV27_CASSLNGGLPETQYFG_hTR137 E-10 BJ2-53 PBL 1.12 106.2 hTRBV27_CASSLNGGLPETQYFG_hTR138 E-11 BJ2-54 PBL 4.39 92.9 NDE-11
[0016] Docket No.: FR 084276.00403 5PBL 6.06 93.1 NDE-11HD4 hCMVNLVPMVAT174 6 PBL 9.26 109.1 hTRBV27_CASSLNGGLPETQYFG_hTR139 pp65 V E-11 BJ2-5 7PBL - 85.8 ND8 PBL 4.90 99.4 hTRBV20_CSARDNTVANYGYTFG_hT140 E-11 RBJ1-2 9PBL - 100.6 hTRBV20_CSARDNTVANYGYTFG_hT141 RBJ1-2 10 PBL 1.31 93.8 NDE-11 11 PBL 5.89 95.5 hTRBV20_CSARDNTVANYGYTFG_hT142 E-11 RBJ1-2 1PBL 6.73 227.8 hTRBV05-143 E-10 4_CASSSLATSTDTQYFG_hTRBJ02-3 2PBL 6.12 241.4 hTRBV05-144 E-10 4_CASSSLATSTDTQYFG_hTRBJ02-3 3PBL 9.09 142.5 hTRBV05-145 E-09 4_CASSSLATSTDTQYFG_hTRBJ02-3 4PBL 4.80 300.2 hTRBV05-146 E-10 4_CASSSLATSTDTQYFG_hTRBJ02-3 5PBL 5.14 294.5 hTRBV04-147 E-10 1_CASSQDGTNYGYTFG_hTRBJ01-2 6PBL 2.19 226.2 hTRBV05-148 E-10 4_CASSSLATSTDTQYFG_hTRBJ02-3 7PBL 2.69 171 hTRBV02_CASSEEETGGSPLHFG_hTR149 E-09 BJ01-6HD5 hCMV IPSINVHHY 175 8 PBL 7.76 211.9 hTRBV02_CASMGGAYNEQFFG_hTRBJ150 pp65 E-10 02-1 9PBL 3.17 201.1 hTRBV05-151 E-10 4_CASSSLATSTDTQYFG_hTRBJ02-3 10 PBL 5.31 - NDE-09 11 PBL 1.69 - NDE-10
[0017] Docket No.: FR 084276.00403 12 PBL 2.02 - NDE-09 13 PBL 3.41 - NDE-10 14 PBL 1.55 - NDE-10 1PBL - 111 ND2 PBL - 20.6 ND3 PBL - 16.7 ND4 PBL - 17.6 ND5 PBL - 19.4 ND6 PBL - 27.4 ND7 PBL - 18.5 ND8 PBL - 29.1 ND9 PBL - 25.8 NDHD6 EBVGLCTLVAM176 10 PBL - 17.3 NDBMLF1 L 11 PBL - 33.3 ND12 PBL - 21.5 ND13 PBL - 35 ND14 PBL - 46.8 ND15 PBL - 26.9 ND16 PBL - 85.5 ND17 PBL - 24.1 ND18 PBL - 55.5 ND19 PBL - 17.3 ND20 PBL - 19.9 ND21 PBL - 20.2 ND22 PBL - 22.8 ND23 PBL - 19.3 ND24 PBL - 74.7 ND25 PBL - 26.2 ND26 PBL - 21.4 ND27 PBL - 21.6 ND
[0018] Docket No.: FR 084276.00403
[0019] Docket No.: FR 084276.00403Table 2. Patients and donors description and clinical information.Code Tumor type Tumor stage Gender TreatmentMel1 Melanoma Mel2 Melanoma Advanced melanoma Advanced melan M F M Mel3 Melanoma oma Advanced melanoma naive lanoma Advance M Mel4 Me d melanoma Mel5 Melanoma Advanced melanoma Advanced Mel6 Melanoma melanomaF M MAGE-A3 vaccinationTargeted therapy-Checkpoint blockade- Targeted therapy naive Lymphodepleting chemotherapy, MART-1 vaccination and ACT of PBMCs Melan A and MART-1 peptide vaccines Mel7 Melanoma Advanced melanoma Mel8 Melanoma Advanced melanomaM M Checkpoint blockadeMelan-A (ELA), MAGE-A10, NY- Mel9 Melanoma Advanced melanoma MESO1 peptide Vaccines-Checkpoint blockade Targeted therapy-Checkpoint blockade- Targeted therapy Melanoma Mel10 Colorectal CRC1 Colorectal Advanced melanoma M M F Chemotherapy-Checkpoint blockade CRC2 Ovarian Ovarian F F F naive Ovarian OvCa1 Early-stage microsatellite instable colon OvCa2 Ovarian adenocarcinoma Early-stageF naiveOvCa3 microsatellite instable colon adenocarcinoma Advanced HGSOC
[0020] Docket No.: FR 084276.00403 Chemotherapy, targeted therapy and dendritic cell vaccines OvCa4 Advanced BRCA1 mutated HGSOC Chemotherapy and dendritic cell Advanced HGSOC vaccines Chemotherapy and dendritic cell vaccines Chemotherapy and dendritic cell vaccines Advanced HGSOC OvCa5 OvarianHigh grade serous ovarian carcinoma, BRCA1 mutatedF Chemotherapy and targeted therapyLung1 NSCLC Early-stage NSCLC squamous M naiveLung2 NSCLC Early-stage NSCLC adenocarcinoma M naiveLung3 NSCLC Early-stage NSCLC squamous F naiveCode Indication Gender TreatmentHD1 healthy donor NAF (21 yo) NA HD2 healthy donor NAM (60 yo) NA HD3 healthy donor NAF (24 yo) NA HD4 healthy donor NAF (35 yo) NA HD5 healthy donor NA NA NAHD6 healthy donor NA NA NA
[0021] Docket No.: FR 084276.00403Table 3. pMHC-TCR structural avidity and TCR pMHC interactions in the modeled complexes.Proteipeptide-MHC ID TCRα / β SEQT1 / 2 Number of Data n Clon ID (s)* predicted adjuste (patient) e NO interactions d via pola non- tota equatio r pola l n r GLCTLVAML / HL ClhTRAV05_CAEDSNARLMFG_hTRAJ31 178 16, 8 24 32 70,6A-A*02:01 (SEQ14 1 ID NO: 460) hTRBV20_CSARDRGLGNTIYFG_hTRBJ01-3 179EBV GLCTLVAML / HL Cl hTRAV30_CGTEGQMNTGFQKLVFG_hTRAJ180 &77.8 32 40 91,8(Lung3)A-A*02:01 (SEQ28 08 hTRBV06- 391 8 ID NO: 461) 8_CASSENVGIGANVLTFG_hTRBJ02-6 GRKLFGTHF / HLACl 9 hTRAV04_CLVGGPPTGNQFYFG_hTRAJ4181153,6 11 37 48 131,3-B*27:05 (SEQ ID9 hTRBV07- & 392 NO: 462) 8_CASSLDIGTYEQFFG_hTRBJ02-1 hTRAV8-6_CAANNNNDMRFG_hTRAJ43 182 hTRBV7-8_CASSWDSGYEQYFG_hTRBJ02-7 & 393 ZCCHC1 GRKLFGTHF / HLACl 7 3,8 4 13 17 6,51-B*27:05 (SEQ ID(OvCa2) NO: 463) AMDLGIHKV / HLCl 5 hTRAV25_CAGMDSSYKLIFG_hTRAJ12 183 77, 10 14 24 61,6A-A*02:01 (SEQ7 ID NO: 464) hTRBV11- 184 1_CASSFQTGWNEQFFG_hTRBJ02-1 UTP20 AMDLGIHKV / HLCl 1 hTRAV24_CAFINSGNTPLVFG_hTRAJ29 185 10, 5 11 16 10,0(Lung1)A-A*02:01 (SEQ0 ID NO: 465) hTRBV10- 186 1_CASSDSTAKETQYFG_hTRBJ02-5 NILDAIAEI / HLA-Cl 11 hTRAV12-312- 187 & 48,1 7 16 23 40,7A*02:01 (SEQ ID4_CAMRSIGGSNYKLTFG_hTRAJ53 394 NO: 466)
[0022] Docket No.: FR 084276.00403 hTRBV07- 9_CASSPIAGGTDTQYFG_hTRBJ02-3 hTRAV35_CAGHGNTGKLIFG_hTRAJ37 188MMP9 NILDAIAEI / HLA-Cl 7 hTRBV05- 189 9,9 4 11 15 1,2(Lung1)A*02:01 (SEQ ID4_CASFFSGGGTDTQYFG_hTRBJ02-3 NO: 467) ITDQVPFSV / HLA-Cl 14 hTRAV24_CAFAELWGGSQGNLIFG_hTR 190 & 70,8 4 26 30 40,9A*02:01 (SEQ IDAJ42 395 NO: 468) hTRBV19_CASSITTSGGYEQYFG_hTRBJ02- 7 hTRAV41_CASTNVGGSGNTPLVFG_hTR191 &AJ29 396 hTRBV19_CASSARGYASPLHFG_hTRBJ01-6 GP100 ITDQVPFSV / HLA-Cl 8 9,4 4 19 23 22,4(Mel4)A*02:01 (SEQ IDNO: 469)
[0023] Docket No.: FR 084276.00403 Table 4. List of pMHC-TCR and their structural avidity used for clustering analysis. CloneTCRα / β SEQ peptide-MHC SEQT1 / 2 ID specificity ID ID (s) TCR NO NO Model hTRAV30_CGTEGQMNTGFQKLVFG_hTRAJ08 192 GLCTLVAML / HLA-250 77.8 1hTRBV06-8_CASSENVGIGANVLTFG_hTRBJ02-6 & A*02:01 397 EBV hTRAV05_CAEDSNARLMFG_hTRAJ31 193 GLCTLVAML / HLA-251 16.1 2BMLF1 hTRBV20_CSARDRGLGNTIYFG_hTRBJ01-3 & A*02:01 PBLs 398 hTRAV24_CAFLTGTYKYIFG_hTRAJ40 194 IPSINVHHY / HLA-252 171 3hTRBV05-4_CASSSLATSTDTQYFG_hTRBJ02-3 & B*35:01 399 hTRAV12-2_CAGYSGTYKYIFG_hTRAJ40 195 IPSINVHHY / HLA-253 294.5 4hTRBV02_CASMGGAYNEQFFG_hTRBJ02-1 & B*35:01 400 Viral hCMV hTRAV16_FNKFYFG_hTRAJ21 196 IPSINVHHY / HLA-254 211.9 5Antigens pp65 PBLs hTRBV02_CASSEEETGGSPLHFG_hTRBJ01-6 & B*35:01 401 hTRAV22_CAGREVTGGGNKLTFG_hTRAJ10 197 IPSINVHHY / HLA-255 214.5 6hTRBV04-1_CASSQDGTNYGYTFG_hTRBJ01-2 & B*35:01 402 Influenza hTRAV21_CAVINAGNNRKLIWG_hTRAJ38 198 VSDGGPNLY / HLA-256 16.8 8A PB1 hTRBV11-2_CASSLDGQGPLYGYTFG_hTRBJ01-2 & A*01:01 PBLs 403 hTRAV27_CAGAGSQGNLIFG_hTRAJ42 199 VSDGGPNLY / HLA-257 293 9hTRBV19_CASSIRSSYEQYFG_hTRBJ02-7 & A*01:01 404 Influenza hTRAV27_CAGAGGGSQGNLIFG_hTRAJ42 200 VSDGGPNLY / HLA-258 ND 10A PB1 hTRBV19_CASSIRSSNEQFFG_hTRBJ02-1 & A*01:01 TILs 405 hTRAV04_CLNAGNNRKLIWG_hTRAJ38 201 VSDGGPNLY / HLA-259 ND 11hTRBV19_CASGLGLEQFFG_hTRBJ02-1 & A*01:01 406
[0024] Docket No.: FR 084276.00403 hTRAV08-1_CAGGGDRDDKIIFG_hTRAJ30 202 ITDQVPFSV / HLA-260 6.4 12hTRBV10-2_CASSYRGNSPLHFG_hTRBJ01-6 & A*02:01 407 GP100 hTRAV01-2_CAVPSYGQNFVFG_hTRAJ26 203 ITDQVPFSV / HLA-261 3.8 13PBLs hTRBV19_CASSLRLAATIYNEQFFG_hTRBJ02-1 & A*02:01 408 hTRAV41_CASTNVGGSGNTPLVFG_hTRAJ29 204 ITDQVPFSV / HLA-262 9.4 14hTRBV19_CASSARGYASPLHFG_hTRBJ01-6 & A*02:01 409 hTRAV19_CALSEGGGGADGLTFG_hTRAJ45 205 ITDQVPFSV / HLA-263 9 15hTRBV02_CASSELERLKVYNSPLHFG_hTRBJ01-6 & A*02:01 410 hTRAV10_CVVSARSGGSYIPTFG_hTRAJ06 206 ITDQVPFSV / HLA-264 6.2 16hTRBV19_CASSMGQLILGYEQYFG_hTRBJ02-7 & A*02:01 411 GP100 hTRAV24_CAFAELWGGSQGNLIFG_hTRAJ42 207 ITDQVPFSV / HLA-265 70.8 17TILs hTRBV19_CASSITTSGGYEQYFG_hTRBJ02-7 & A*02:01 412 hTRAV12-2_CAGGGSNYQLIWGAG_hTRAJ33 208 EAAGIGILTV / HLA-266 4 18hTRBV20_CSASPGLAEQFFG_hTRBJ02-1 & A*02:01 413 hTRAV12-2_CAVDVGARLMFG_hTRAJ31 209 EAAGIGILTV / HLA-267 3.7 19hTRBV14_CASSQDTGLSSYNEQFFG_hTRBJ02-1 & A*02:01 414TAAs hTRAV12-2_CAYQAGTALIFG_hTRAJ15210 EAAGIGILTV / HLA-268 4.8 20hTRBV06-1_CASSELGLAGNEQFFG_hTRBJ02-1 & A*02:01 415 hTRAV12-2_CAVNTGNQFYFG_hTRAJ49 211 EAAGIGILTV / HLA-269 4 21hTRBV20_CSAERGLGQPQHFG_hTRBJ01-5 & A*02:01 416 hTRAV12-2_CAPGGGYQKVTFG_hTRAJ13 212 EAAGIGILTV / HLA-270 4.1 22hTRBV19_CASTSGELGQPQHFG_hTRBJ01-5 & A*02:01 417 hTRAV12-2_CAVIHAGKSTFG_hTRAJ27 213 EAAGIGILTV / HLA-271 3.6 23hTRBV27_CASSLSGLAGVEQYFG_hTRBJ02-7 & A*02:01 418
[0025] Docket No.: FR 084276.00403 Melan-A hTRAV35_CAGVLGSARQLTFG_hTRAJ22 214 EAAGIGILTV / HLA-272 36.3 24TILs hTRBV19_CASKWGALMNTEAFFG_hTRBJ01-1 & A*02:01 419 hTRAV12-2_CAASIGFGNVLHCGSG_hTRAJ35 215 EAAGIGILTV / HLA-273 5 25hTRBV27_CASSSLGATYEQYFG_hTRBJ02-7 & A*02:01 420 hTRAV12-2_CAASIGFGNVLHCGSG_hTRAJ35 216 EAAGIGILTV / HLA-274 3.7 26hTRBV27_CASSWTSGSPSEQFFG_hTRBJ02-1 & A*02:01 421 hTRAV12-2_CAVTIGFGNVLHCGSG_hTRAJ35 217 EAAGIGILTV / HLA-275 3.9 27hTRBV27_CASSLFSGSSGELFFG_hTRBJ02-2 & A*02:01 422 hTRAV12-2_CAVGGGADGLTFG_hTRAJ45 218 EAAGIGILTV / HLA-276 3.1 28hTRBV03-1_CASSQGSLAGSEQYFG_hTRBJ02-7 & A*02:01 423 hTRAV12-2_CAVGGAAGNKLTFG_hTRAJ17 219 EAAGIGILTV / HLA-277 7.9 29hTRBV28_CASRVQGLGQPQHFG_hTRBJ01-5 & A*02:01 424 hTRAV12-2_CAVSSGFQKLVFG_hTRAJ08 220 EAAGIGILTV / HLA-278 3.4 30hTRBV13_CASSLDPSGSPNEQFFG_hTRBJ02-1 & A*02:01 425 Melan-A hTRAV12-2_CAVNDAGKSTFG_hTRAJ27 221 EAAGIGILTV / HLA-279 3.1 31PBLs hTRBV03-1_CASSQGDLAWIPTEAFFG_hTRBJ01- & A*02:01 1 426 hTRAV14_CAMRGPYNTDKLIFG_hTRAJ34 222 EAAGIGILTV / HLA-280 3.2 32hTRBV20_CSARDVGLGIYEQYFG_hTRBJ02-7 & A*02:01 427 PHLPP2 hTRAV23_CAAPMPMDTGRRALTFG_hTRAJ05 223 QSDNGLDSDY / HLA-281 69.3 36TILs hTRBV10-3_CAISGGSVGEQYFG_hTRBJ02-7 & A*01:01 428 hTRAV21_CAVSSGSARQLTFG_hTRAJ22 224 QSDNGLDSDY / HLA-282 47.8 37hTRBV05-4_CASTLSTGQGIYGYTFG_hTRBJ01-2 & A*01:01 429NeoAntigens PHLPP2hTRAV21_CAVGGSGSARQLTFG_hTRAJ22 225 QSDNGLDSDY / HLA-283 <3s 69PBLs hTRBV05-4_CASSPTTSGRIGELFFG_hTRBJ02-2 & A*01:01 430
[0026] Docket No.: FR 084276.00403 hTRAV04_CLVGGPPTGNQFYFG_hTRAJ49 226 GRKLFGTHF / HLA-284 153.6 38hTRBV07-8_CASSLDIGTYEQFFG_hTRBJ02-1 & B*27:05 431 ZCCHC11 hTRAV21_CAVRLTGQGAQKLVFG_hTRAJ54 227 GRKLFGTHF / HLA-285 ND 39TILs hTRBV28_CASSLAGLNTEAFFG_hTRBJ01-1 & B*27:05 432 ZCCHC11 hTRAV08-6_CAANNNNDMRFG_hTRAJ43 228 GRKLFGTHF / HLA-286 3.8 40PBLs hTRBV07-8_CASSWDSGYEQYFG_hTRBJ02-7 & B*27:05 433 HHAT hTRAV12-2_CAVNYNNARLMFG_hTRAJ31 229 KQWLVWLFL / HLA-287 6.7 41PBLs hTRBV04-2_CASSQDAETQYFG_hTRBJ02-5 & A*02:01 434 HHAT hTRAV38-2_CAFMDSNYQLIWGAG_hTRAJ33 230 KQWLVWLFL / HLA-288 26.2 43TILs hTRBV12-3_CASSRTSPTDTQYFG_hTRBJ02-3 & A*02:01 435 UTP20 hTRAV25_CAGMDSSYKLIFG_hTRAJ12 231 AMDLGIHKV / HLA-289 77.7 44TILs hTRBV11-1_CASSFQTGWNEQFFG_hTRBJ02-1 & A*02:01 436 hTRAV12-2_CAGGVDSNYQLIWGAG_hTRAJ33 232 AMDLGIHKV / HLA-290 ND 45hTRBV05-6_CASSLGGGRDEQYFG_hTRBJ02-7 & A*02:01 437 hTRAV24_CAFINSGNTPLVFG_hTRAJ29 233 AMDLGIHKV / HLA-291 10 46hTRBV10-1_CASSDSTAKETQYFG_hTRBJ02-5 & A*02:01 438 UTP20 hTRAV20_CAVSGGSYIPTFG_hTRAJ06 hTRBV04- 234 AMDLGIHKV / HLA-292 61.2 47PBLs 3_CASSQEESYEQYFG_hTRBJ02-7 & A*02:01 439 hTRAV19_CALIFNQAGTALIFG_hTRAJ15 235 AMDLGIHKV / HLA-293 52.7 48hTRBV05-6_CASSLGGGRDTQYFG_hTRBJ02-3 & A*02:01 440 hTRAV24_CAPNRDDKIIFG_hTRAJ30 hTRBV12- 236 NILDAIAEI / HLA-294 ND 493_CASATGVKLAKNIQYFG_hTRBJ02-4 & A*02:01 441 MMP9 hTRAV12-312- 237 NILDAIAEI / HLA-295 48.1 50TILs 4_CAMRSIGGSNYKLTFG_hTRAJ53 hTRBV07- & A*02:01 9_CASSPIAGGTDTQYFG_hTRBJ02-3 442
[0027] Docket No.: FR 084276.00403NeoAntigens hTRAV35_CAGHGNTGKLIFG_hTRAJ37238 NILDAIAEI / HLA-296 9.9 51hTRBV05-4_CASFFSGGGTDTQYFG_hTRBJ02-3 & A*02:01 443 hTRAV12-2_CAVRGNEKLTFG_hTRAJ48 239 NILDAIAEI / HLA-297 5.8 52hTRBV20_CSASRADTYEQYFG_hTRBJ02-7 & A*02:01 444 hTRAV13-1_CAASSMNRDDKIIFG_hTRAJ30 240 NILDAIAEI / HLA-298 29.6 53hTRBV07-2_CASTDTDTQYFG_hTRBJ02-3 & A*02:01 445 MMP9 hTRAV13-1_CAASINTDKLIFG_hTRAJ34 241 NILDAIAEI / HLA-299 6.4 55PBLs hTRBV05-4_CASFFSGGGTDTQYFG_hTRBJ02-3 & A*02:01 446 hTRAV12-2_CAVGGTSYGKLTFG_hTRAJ52 242 NILDAIAEI / HLA-300 12.6 57hTRBV10-2_CASSLGQETQYFG_hTRBJ02-5 & A*02:01 447 hTRAV27_CAGGNSGGYQKVTFG_hTRAJ13 243 NILDAIAEI / HLA-301 ND 58hTRBV05-4_CASFFSGGGTDTQYFG_hTRBJ02-3 & A*02:01 448 hTRAV21_CAVPSTSGTYKYIFG_hTRAJ40 244 PYMFLSEWI / HLA-302 62 59hTRBV19_CASSIGKTGKLFFG_hTRBJ01-4 & A*24:02 449 SLC25A48 hTRAV17_CATGGALGYGGSQGNLIFG_hTRAJ42 245 PYMFLSEWI / HLA-303 3.6 60PBLs hTRBV19_CASSIARVTEAFFG_hTRBJ01-1 & A*24:02 450 hTRAV12-1_CVVRANNARLMFG_hTRAJ31 246 PYMFLSEWI / HLA-304 4.1 65hTRBV07-2_CASSTGSSGELFFG_hTRBJ02-2 & A*24:02 451 hTRAV12-2_CGGSGTASKLTFG_hTRAJ44 247 VFSFTATLPF / HLA-305 <3s 66hTRBV05-1_CASSFSGSEQFFG_hTRBJ02-1 & A*24:02 452 FPR2 hTRAV19_CALSEWELNTNAGKSTFG_hTRAJ27 248 VFSFTATLPF / HLA-306 <3s 67PBLs hTRBV20_CSARKRGYREEAFFG_hTRBJ01-1 & A*24:02 453 FPR2 hTRAV20_CAVLSGNTGKLIFG_hTRAJ37 249 VFSFTATLPF / HLA-307 <3s 68TILs hTRBV20_CSARGQGNTEAFFG_hTRBJ01-1 & A*24:02 454
[0028] Docket No.: FR 084276.00403
[0029] Docket No.: FR 084276.00403Table 5. Amino acids frequencies in CDR3β for high (>60s) and low avidity TCRs used forclustering analyses. High- High- Insi Affinity Amino de Outs High- Affinity Low- Acid box ide box cluster Affinity (> outside within Affinity cluster 60 s) box box (< 60 s) cluster cluster A0,029 0,072 0,031 0,048 0,016 0,063R 0,052 0,023 0,01 0 0,032 0,039N 0,017 0,046 0,052 0,095 0,016 0,031D 0,069 0,033 0,031 0,024 0,048 0,05C 0 0 0 0 0 0Q 0,092 0,128 0,115 0,119 0,113 0,115E 0,075 0,112 0,135 0,119 0,129 0,089G 0,201 0,155 0,177 0,19 0,194 0,17H 0 0 0 0 0 0I 0,023 0,039 0,052 0,071 0,032 0,029L 0,075 0,105 0,052 0,048 0,065 0,105K 0,023 0,01 0,021 0 0,032 0,013M 0 0,007 0 0 0 0,005F 0,017 0,01 0,01 0,024 0 0,013P 0,04 0,03 0,01 0 0,016 0,039S 0,086 0,082 0,083 0,024 0,113 0,084T 0,121 0,076 0,104 0,095 0,113 0,089W 0 0,016 0,01 0,024 0 0,01Y 0,069 0,033 0,073 0,071 0,065 0,039V 0,011 0,023 0,031 0,048 0,016 0,016Amino acid frequencies in CDR3β (first 4 and last 3 residues removed - in agreement with thebiophysical-based clustering analysis).
[0030] Docket No.: FR 084276.00403 Table 6. The amino acids (features) selected for the 10 training sets. Set Features SEQ IDCorrelation %correct %correct NO Coefficient, LR High Avidity Low Avidity Full set RNDGILF 455 0.75 91% 97%Training-set1 RNDGILF 320 0.75 88% 100%Training-set2 RNDGILF 321 0.75 88% 97%Training-set3 RNDGILF 322 0.72 88% 97%Training-set4 RNDGILF 323 0.81 100% 97%Training-set5 RNDGLFV 324 0.78 75% 97%Training-set6 RNDILKF 325 0.81 100% 100%Training-set7 DGILFPY 326 0.76 75% 97%Training-set8 RNDGILF 327 0.75 88% 100%Training-set9 RNDGILF 328 0.78 100% 100%Training-set10 RNDILFW 329 0.80 100% 100%
[0031] Docket No.: FR 084276.00403 Table 7. Example amino acid sequences described in Fig. 1A. SEQ SEQ ID ID NO NO TRAV23 / D CAARAAGNKLT330 TRA TRBVS-1 CASSPLTSEYNE 360 TRBJV6 F J17 QFF 2-1 TRAV13-1 CAAPQGGKLIF 331 TRA TRBV7-2 CASNPGLHNEQF 361 TRBJJ23 F 2-1 TRAV13-1 CAASPLFDKLIF 332 TRA TRBV7-9 CASSASGRADNE 362 TRBJJ34 QFF 2-1 TRAV12-1 CWTSGTYKYIF 333 TRA TRBV6-1 CASSDPSPTDTQ 363 TRBJJ40 YF 2-3 TRAV9-2 CALSPMDSSYKL 334 TRA TRBVIO-3 CAISDWTSGKNN 364 TRBJIF J12 EQFF 2-1 TRAV24 CAFSNGNKLVF 335 TRA TRBV27 CASSLGGHGYN 365 TRBJJ47 EQFF 2-1 TRAV12-3 CAMSAHTNAGK 336 TRA TRBV13 CASSSQESYEQY 366 TRBJSTF J27 F 2-7 TRAV8-6 CAVRSDYKLSF 337 TRA TRBV6-2 CASSYPTAGANV 367 TRBJJ20 LTF 2-6 TRAV17 CATVPTGANSKL 338 TRA TRBV9 CASSVAGFYGYT 368 TRBJITF J56 F -2 TRAVI-2 CAVKRNNNARL 339 TRA TRBV14 CASSQDAATFLG 369 TRBJIMF J31 YTF -2 TRAVIO CWSLNTGGFKTI 340 TRA TRBV7-9 CASSASGRADNE 370 TRBJF J9 QFF 2-1 TRAV4 CLVGDKEGGKLI 341 TRA TRBV9 CASRERDPTDTQ 371 TRBJF J23 YF 2-3 TRAV22 CAVARFSGGYN 342 TRA TRBV7-2 CASSSGRGVTGA 372 TRBJKLIF J4 NVLTF 2-6 TRAV4 CLVGGTDNNDM 343 TRA TRBV12-3 CASALGTLYNEQ 373 TRBJRF J43 FF 2-1 TRAV12-1 CVVNPYNNNDM 344 TRA TRBVS-4 CASSSVSPFFATF 374 TRBJIRF J43 -2 TRAV3 CAVRDENARLM 345 TRA TRBV7-9 CASSIETQYF 375 TRBJF J31 2-5 TRAV13-1 CAAPQGGKLIF 346 TRA TRBV7-2 CAISSGIHNEQFF 376 TRBJJ23 2-1 TRAV4 CLVGPRGHNTD 347 TRA TRBV27 CASSTSRTGNTD 377 TRBJKLIF J34 TQYF 2-3 TRAVB-2 CAVSGKAGTALI 348 TRA TRBV7-9 CASSIETQYF 378 TRBJF JIS 2-5 TRAV38- CAYFSPQGGSEK349 TRA TRBV7-2 CASSIPPGQETQ 379 TRBJ2 / DV LVF J57 YF 2-5 TRAV13-1 CAASGRDDKIIF 350 TRA TRBV7-3 CASSLTGPAGGG 380 TRBJJ30 TQYF 2-5 TRAV27 CAGYPGSGNTG 351 TRA TRBVIO-3 CAIRGIDTEAFF 381 TRBJIKLIF J37 -1 Docket No.: FR 084276.00403TRAVIO CVVSAGGYIPTF 352 TRA TRBV7-9 CASSASGRADNE 382 TRBJJ6 QFF 2-1TRAV2 CAVDRDDKIIF 353 TRA TRBV3-1 CASSQDRTVGA 383 TRBJIJ30 AFF -1TRAV12-1 CWNGAGSYQLT 354 TRA TRBVS-5 CASSSLDRGGTD 384 TRBJF J28 TQYF 2-3TRAV3 CAVRDRGAALIF 355 TRA TRBV7-9 CASSLETGGTTY 385 TRBJJIS EQYF 2-7TRAV17 CATDPGGYNKLI 356 TRA TRBV7-3 CASSPAGGHTGE 386 TRBJF J4 LFF 2-2TRAV24 CAFSNGNKLVF 357 TRA TRBV6-1 CASSPGQGYGYT 387 TRBJIJ47 F -2TRAV30 CGMSLRGADGL 358 TRA TRBV7-2 CASSIPPGQETQ 388 TRBJTF J45 YF 2-5TRAV17 CATDERGAGGT 359 TRA TRBV28 CATRDRGWETQ 389 TRBJSYGKLTF J52 YF 2-SI
Claims
Docket No.: FR 084276.00403 CLAIMS What is claimed is:
1. A method of identifying a T cell receptor (TCR) with high avidity against an antigen or a Tcell comprising the TCR, comprising: selecting a set of amino acids in an amino acid sequence of a CDR3β of the TCR; determining solvent accessibility of each amino acid of the set of amino acids; inputting the solvent accessibility of each amino acid of the set of amino acids into a logistic regression model, wherein the logistic regression model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids; determining a probability value of a high avidity status of the TCR as an output of the logistic regression model; and identifying the TCR or the T cell as having high avidity against the antigen if the probability value is greater than or equal to a threshold value.
2. The method of claim 1, wherein the set of amino acids comprises one or more of Arg, Asn,Asp, Gly, Ile, Lue, and Phe.
3. The method of any one of the preceding claims, comprising determining the probabilityvalue by:wherein p is the probability value, b0 represents a bias term, and R, N, D, G, I, L, and F areamino acids Arg, Asn, Asp, Gly, Ile, Lue, and Phe, respectively.
4. The method of any one of the preceding claims, wherein the threshold value is about 0.5.
5. The method of any one of the preceding claims, comprising selecting a subset of aminoacids from the set of amino acids that have solvent accessibility higher than about 30%.Docket No.: FR 084276.004036. The method of claim 5, comprising inputting the solvent accessibility of the subset ofamino acids into the logistic regression model.
7. The method of any one of the preceding claims, comprising determining the solventaccessibility of each amino acid of the set of amino acids as a relative solvent excluded surface area (SESA).
8. The method of claim 7, wherein the SESA is determined by normalizing surface area of anamino acid in the TCR against surface area of the amino acid in a reference state.
9. The method of any one of the preceding claims, comprising determining solventaccessibility of an amino acid based on a three-dimensional model of the TCR.
10. The method of any one of the preceding claims, wherein the TCR with high avidity has apMHC-TCR half-life of less than about 60 seconds.
11. The method of any one of the preceding claims, comprising inputting into the logisticregression model one or more additional characteristics of each amino acid of the set of amino acids.
12. The method of claim 11, wherein the additional characteristics comprise hydrophilicityvalue, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of β-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of β-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in β-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, α-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in α-helix, average nonbonded energy per residue,Docket No.: FR 084276.00403bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed α-helix, heat capacity, free energy in α-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of β-sheet, unweighted, normalized frequency of β-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof.
13. The method of claim 12, wherein the additional characteristics comprise hydrophobicity,secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof.
14. The method of any one of the preceding claims, comprising determining a nucleotidesequence of the TCR by deep sequencing.
15. The method of any one of the preceding claims, wherein the antigen comprises aneoantigen or a tumor-associated antigen.
16. The method of any one of the preceding claims, wherein the T cell comprises a CD8+ Tcell or a CD4+ T cell.
17. A TCR or a T cell, which is identified according to the method of any one of the precedingclaims.
18. A method of producing an engineered T cell with high avidity against an antigen,comprising transfecting or transducing a T cell with a nucleic acid molecule encoding a TCR identified according to the method of any one of claims 1-16.
19. A method of treating cancer in a subject comprising administering to the subject a T cellidentified according to the method of any one of claims 1-16 or a T cell produced according tothe method of claim 17.
20. The method of claim 19, wherein the cancer is selected from adrenal gland tumors, biliarycancer, bladder cancer, brain cancer, breast cancer, carcinoma, central or peripheral nervousDocket No.: FR 084276.00403 system tissue cancer, cervical cancer, colon cancer, endocrine or neuroendocrine cancer or hematopoietic cancer, esophageal cancer, fibroma, gastrointestinal cancer, glioma, head and neck cancer, Li-Fraumeni tumors, liver cancer, lung cancer, lymphoma, melanoma, meningioma, multiple neuroendocrine type I and type II tumors, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteogenic sarcoma tumors, ovarian cancer, pancreatic cancer, pancreatic islet cell cancer, parathyroid cancer, pheochromocytoma, pituitary tumors, prostate cancer, rectal cancer, renal cancer, respiratory cancer, sarcoma, skin cancer, stomach cancer, testicular cancer, thyroid cancer, tracheal cancer, urogenital cancer, and uterine cancer.
21. A system for identifying a T cell receptor (TCR) with high avidity against an antigen or a Tcell comprising the TCR, wherein the system comprises one or more processors configured to: select a set of amino acids in an amino acid sequence of a CDR3β of the TCR; determine solvent accessibility of each amino acid of the set of amino acids; input the solvent accessibility of each amino acid of the set of amino acids into a logistic regression model, wherein the logistic regression model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids; determine a probability value of a high avidity status of the TCR as an output of thelogistic regression model; and identify the TCR or the T cell as having high avidity against the antigen if the probability value is greater than or equal to a threshold value.
22. The system of claim 21, wherein the set of amino acids comprises one or more of Arg, Asn,Asp, Gly, Ile, Lue, and Phe.
23. The system of any one of claims 21-22, wherein the one or more processors are furtherconfigured to determine the probability value by:wherein p is the probability value, b0 represents a bias term, and R, N, D, G, I, L, and F areamino acids Arg, Asn, Asp, Gly, Ile, Lue, and Phe, respectively.Docket No.: FR 084276.0040324. The system of any one of claims 21-23, wherein the threshold value is about 0.5.
25. The system of any one of claims 21-24, wherein the one or more processors are furtherconfigured to select a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 30%.
26. The system of claim 25, wherein the one or more processors are further configured to inputthe solvent accessibility of the subset of amino acids into the logistic regression model.
27. The system of any one of claims 21-26, wherein the one or more processors are furtherconfigured to determine the solvent accessibility of each amino acid of the set of amino acids as a relative solvent excluded surface area (SESA).
28. The system of claim 27, wherein the SESA is determined by normalizing surface area of anamino acid in the TCR against surface area of the amino acid in a reference state.
29. The system of any one of claims 21-28, wherein the one or more processors are furtherconfigured to determine solvent accessibility of an amino acid based on a three-dimensional model of the TCR.
30. The system of any one of claims 21-29, wherein the TCR with high avidity has a pMHC-TCR half-life of less than about 60 seconds.
31. The system of any one of claims 21-30, wherein the one or more processors are furtherconfigured to input into the logistic regression model one or more additional characteristics of each amino acid of the set of amino acids.
32. The system of claim 31, wherein the additional characteristics comprise hydrophilicityvalue, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of β-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue,Docket No.: FR 084276.00403 conformational parameter of β-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in β-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, α-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in α-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left- handed α-helix, heat capacity, free energy in α-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of β-sheet, unweighted, normalized frequency of β-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof.
33. The system of claim 31, wherein the additional characteristics comprise hydrophobicity,secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof.
Citation Information
Patent Citations
Prediction of t cell response to antigens
WO2023028595A1
Methods for predicting epitope specificity of t cell receptors
WO2023031207A1