Methods and systems for identifying clinically relevant t-cell receptors
Patent Information
- Application Number
- PCT/EP2024/087761
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-24
AI Technical Summary
Current methods for identifying clinically relevant T-cell receptors (TCRs) in cancer immunotherapy are hindered by the lack of a rigorous evaluation framework and the inclusion of TCRs with untested tumor specificity or insufficient affinity.
A method and system for identifying clinically relevant TCRs by determining a transcriptomic profile of T cells through single-cell RNA and TCR sequencing, using a classifier to predict tumor reactivity, and selecting candidate TCRs with high structural avidity or clustering into TCR clusters targeting distinct antigens.
The approach enables the precise identification of tumor-reactive TCRs, improving the efficacy of personalized T-cell therapies by selecting TCRs with optimal tumor specificity and affinity.
Abstract
Description
[0001] METHODS AND SYSTEMS FOR IDENTIFYING CLINICALLY RELEVANT T-CELL RECEPTORS
[0002] FIELD OF THE INVENTION
[0003] This invention relates to methods and systems for identifying clinically relevant T-cell receptors (TCRs).
[0004] BACKGROUND OF THE INVENTION
[0005] Adoptive cell transfer (ACT) of tumor-infiltrating lymphocytes (TILs) is a personalized immunotherapy approach with demonstrated superiority over second-line checkpoint blockade immunotherapy in melanoma patients. Recent studies indicate that the frequency of tumor-reactive T cells (TRTs) in melanoma tumors used to expand TILs, and their number in cognate TIL products infused to patients are important determinants of TIL- ACT efficacy. Moreover, it has been hypothesized that the paucity of TRTs in other solid tumors can explain the lower success rate of HL therapy in these patients. Alternatively, the use of T-cell receptor (TCR)-engineered T cells has shown promise for treating solid tumors. However, when personalized, this approach faces the challenge of including TCRs with untested tumor specificity or insufficient affinity, stressing the need to identify clinically relevant TCRs.
[0006] Recent progress in single-cell RNA and TCR sequencing technologies enabled the exploration of the unique transcriptomic profile of intratumoral neoantigen-specific and tumor- reactive T cells. This progress provides an opportunity to derive predictors of private TRTs capable of reliably identifying tumor-reactive over bystander TILs, thus opening the door for the development of personalized TCR-based therapies. However, any practical application has been hampered by the lack of a rigorous evaluation framework and robust benchmarking against external datasets. Moreover, predicting all TRTs indiscriminately may result in the selection of clinically suboptimal low-avidity TCRs (also occurring for some neoantigen-specific clonotypes).
[0007] Accordingly, there remains a need for methods and systems for identifying clinically relevant TCRs. SUMMARY OF THE INVENTION
[0008] This disclosure addresses the need mentioned above in a number of aspects. In one aspect, this disclosure provides a method of identifying a classifier for predicting tumor reactivity of T cells. In some embodiments, the method comprises: (a) determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells, wherein the transcriptomic profile comprises transcription level of each of a set of genes; (b) inputting the transcriptomic profile of the set of genes into a plurality of test classifiers; (c) determining, by each of the test classifiers, a result comprising tumor reactivity of each of the T cells; (d) determining a classification score of each of the test classifiers by a nested cross- validation framework based on the result determined by each of the test classifiers; and (e) ranking classification scores of the test classifiers and selecting, from the test classifiers, a candidate classifier having the classification score higher than a reference classification performance value.
[0009] In another aspect, this disclosure also provides a method of predicting tumor reactivity of T cells. In some embodiments, the method comprises: (i) determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells; wherein the transcriptomic profile comprises transcription level of each of a set of genes; (ii) inputting the transcriptomic profile of the set of genes into the classifier identified according to the method as described herein, (iii) determining tumor reactivity of each of the T cells by the classifier; and (iv) ranking tumor reactivity of the T cells and selecting, from the T cells, candidate T cells that are reactive to a tumor antigen (e.g., a tumor-associated antigen (TAA) or an intratumoral neoantigen).
[0010] In some embodiments, the nested cross-validation framework is implemented using a conventional or a leave-one-patient-out (LOPO) approach.
[0011] In some embodiments, the method further comprises finetuning hyperparameters of the candidate classifier by a cross-validation framework. In some embodiments, the cross-validation framework is implemented using a conventional or a leave-one-patient-out (LOPO) approach.
[0012] In some embodiments, the test classifiers comprise a logistic regression model. In some embodiments, the test classifiers comprise a signature scoring model. In some embodiments, the candidate classifier comprises a signature scoring model implemented using an edgeR-QFL.
[0013] In some embodiments, the hyperparameters comprise one or more of: (i) the choice of the log-Fold-Change or P-value for selecting differential expressed genes; (ii) number of genes in the set of the genes; (iii) whether to include only up-regulated genes or both up-regulated and down- regulated genes in the set of the genes; and (iv) a scoring method to rank tumor reactivity of the T cells.
[0014] In some embodiments, the scoring method comprises an Average scoring method, an AUCell scoring method, a Singscore scoring method, or a Ucell scoring method. In some embodiments, the hyperparameters comprise one or more of: using P-value for differential expression analysis, selecting 5 to 500 up-regulated and down-regulated genes, and scoring the tumor reactivity of the T cells with the Average scoring method.
[0015] In some embodiments, the result comprises a tumor reactivity score of each of the T cells. In some embodiments, tumor reactivity scores are scaled based on a mean and standard deviation obtained from training data.
[0016] In some embodiments, the classification scores are characterized by area-under-the-curve (AUC) scores. In some embodiments, the classification scores are characterized by Matthew’s Correlation Coefficient (MCC). In some embodiments, the candidate classifier has the highest classification score among the test classifiers.
[0017] In some embodiments, the tumor reactivity of the T cells comprises reactivity to a tumor antigen (e.g., a tumor-associated antigen (TAA) or an intratumoral neoantigen). In some embodiments, the set of genes comprises one or more genes set forth in Table 4. In some embodiments, the set of genes comprises CXCL13, LAG3, TOX, PDCD1, TNFRSF9, or a combination thereof.
[0018] In some embodiments, the method further comprises validating the candidate classifier against spurious learning.
[0019] In some embodiments, the T cells are from one or more patients. In some embodiments, the candidate classifier can exclude virus-specific T cells. In yet another aspect, this disclosure also provides a method of identifying a tumor-reactive T-cell receptor (TCR). In some embodiments, the method comprises: (a) identifying tumor- reactive T cells by: determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data and single cell TCR sequencing data of T cells, wherein the transcriptomic profile comprises transcription level of each of a set of genes, inputting the transcriptomic profile of the set of genes into the classifier identified according to the method as described herein, determining tumor reactivity of each of the T cells by the classifier, and ranking tumor reactivity of the T cells and selecting, from the T cells, candidate T cells that are reactive to a tumor antigen (e.g., an antigen comprising a tumor-associated antigen (TAA) or an intratumoral neoantigen); and identifying a tumor-reactive TCR from T-cell receptors (TCRs) contained the candidate T cells by performing at least one of: (i) identifying candidate TCRs having high structural avidity from the TCRs of the candidate T cells by: selecting a set of amino acids in an amino acid sequence of a CDR30 of each of the TCRs, determining solvent accessibility of each amino acid of the set of amino acids, inputting the solvent accessibility of each amino acid of the set of amino acids into a logistic regression model, wherein the logistic regression model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids, determining a probability value of a high avidity status of each of the TCRs as an output of the logistic regression model, and identifying candidate TCRs as having high avidity against the antigen if probability values of the candidate TCRs are greater than or equal to a threshold value; and (ii) clustering the candidate TCRs based on physicochemical properties thereof into one or more TCR clusters that target distinct antigens, and selecting a top tumor-reactive TCR from each of the one or more TCR clusters.
[0020] In some embodiments, the top tumor-reactive TCR is the TCR most reactive to the tumor antigen (e.g., an antigen comprising a tumor-associated antigen (TAA) or an intratumoral neoantigen) in each of the one or more TCR clusters.
[0021] In some embodiments, the set of amino acids comprises one or more of Arg, Asn, Asp, Gly, He, Lue, and Phe.
[0022] In some embodiments, the step of identifying candidate TCRs having high structural avidity comprises determining the probability value by: wherein p is the probability value, bo represents a bias term, and R, N, D, G, I, L, and F are amino acids Arg, Asn, Asp, Gly, He, Lue, and Phe, respectively.
[0023] In some embodiments, the threshold value is about 0.5.
[0024] In some embodiments, the step of identifying candidate TCRs having high structural avidity comprises selecting a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 30%.
[0025] In some embodiments, the step of identifying candidate TCRs having high structural avidity comprises inputting the solvent accessibility of the subset of amino acids into the logistic regression model.
[0026] In some embodiments, the step of identifying candidate TCRs having high structural avidity comprises determining the solvent accessibility of each amino acid of the set of amino acids as a relative solvent excluded surface area (SESA). In some embodiments, the SESA is determined by normalizing surface area of an amino acid in the TCR against surface area of the amino acid in a reference state.
[0027] In some embodiments, the step of identifying candidate TCRs having high structural avidity comprises determining solvent accessibility of an amino acid based on a three-dimensional model of the TCR.
[0028] In some embodiments, the TCR with high avidity has a pMHC-TCR half-life of less than about 60 seconds.
[0029] In some embodiments, the step of identifying candidate TCRs having high structural avidity comprises inputting into the logistic regression model one or more additional characteristics of each amino acid of the set of amino acids. In some embodiments, the additional characteristics comprise hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of 0-tum, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of P-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in 0-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, a-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in a-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed a-helix, heat capacity, free energy in a-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of 0- sheet, unweighted, normalized frequency of 0-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof.
[0030] In some embodiments, the additional characteristics comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof.
[0031] In some embodiments, the step of clustering the candidate TCRs comprises: (a) selecting a subset of amino acids in a first TCR and a corresponding subset of amino acids in a second TCR, wherein the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR have an identical number of amino acids; (b) determining an amino acid sum of differences in each of a plurality of physicochemical properties by performing a pairwise comparison between an amino acid in the subset of amino acids in the first TCR and a corresponding amino acid in the corresponding subset of amino acids in the second TCR; (c) repeating steps (a) to (b) for remaining amino acids in the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR; (d) determining a subset sum of differences between the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR; (e) repeating steps (a) to (d) for one or more subsets of amino acids in the first TCR and the second TCR; (f) determining an aggregate value of all subset sums of differences between the first TCR and the second TCR by assigning a weight value to each of the subset sums; and (g) identifying the first TCR and the second TCR as having similar specificity to an epitope if the aggregate value is smaller than a threshold value. In some embodiments, the step of clustering the candidate TCRs comprises: (1) generating a similarity matrix for a subset of amino acids of a TCR of the TCRs, wherein the similarity matrix comprises a plurality of physicochemical properties of each amino acid in the subset of amino acids; (2) repeating step (1) for one or more subsets of amino acids of the TCR; (3) repeating steps (1) to (2) for remaining TCRs of the plurality of immunological entities; and (4) performing a clustering analysis based on a distance between two corresponding similarity matrices of a pair of immunological entities to identify a subset of TCRs having similar specificity to the epitope.
[0032] In some embodiments, the distance is a Manhattan distance. In some embodiments, the clustering analysis comprises hierarchical clustering. In some embodiments, the hierarchical clustering comprises an unweighted pair group method with arithmetic mean (UPGMA).
[0033] In some embodiments, the epitope is located on peptide-MHC (pMHC). In some embodiments, the subset of amino acids comprises 3 to 8 amino acids. In some embodiments, the subset of amino acids comprises 3 to 8 consecutive amino acids. In some embodiments, the subset of amino acids consists of 4 amino acids. In some embodiments, one or more subsets of amino acids are selected from amino acids in CDRla, CDR2a, CDR3a, CDR10, CDR20, and CDR30. In some embodiments, the step of determining the aggregate value comprises assigning a weight value of about 30% to the subset of amino acids in CDR3a or CDR30. In some embodiments, the step of determining the aggregate value comprises assigning a weight value of about 10% to the subset of amino acids in CDRla, CDR2a, CDR10, or CDR20.
[0034] In some embodiments, the subset of amino acids does not include amino acids that are not solvent-exposed. In some embodiments, the subset of amino acids does not include amino acids in CDRla, CDR2a, CDR10, or CDR20 that have a relative solvent excluded surface area (SESA) of less than about 5%. In some embodiments, the subset of amino acids does not include amino acids in CDR3a or CDR30 that have a SESA of less than about 20%.
[0035] In some embodiments, the physicochemical properties comprise amino acid attributes selected from hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of 0-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of 0-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in 0-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, a-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in a-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed a-helix, heat capacity, free energy in a-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of 0- sheet, unweighted, normalized frequency of 0-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, and percentage of buried residues.
[0036] In some embodiments, the physicochemical properties comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, and electrostatic charge.
[0037] The foregoing summary is not intended to define every aspect of the disclosure, and additional aspects are described in other sections, such as the following detailed description. The entire document is intended to be related as a unified disclosure, and it should be understood that all combinations of features described herein are contemplated, even if the combination of features are not found together in the same sentence, paragraph, or section of this document. Other features and advantages of the invention will become apparent from the following detailed description. It should be understood, however, that the detailed description and the specific examples, while indicating specific embodiments of the disclosure, are given by way of illustration only, because various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description.
[0038] BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIGs. 1A, IB, 1C, ID, IE, IF, 1G, 1H, and II show TRTpred as a sensitive in silico predictor of tumor-reactive clonotypes. FIG. 1A shows an Alluvial plot depicting the fractions of cells and clones annotated as tumor-reactive or non-tumor-reactive (orphan or antigen (Ag)- specific) within the input data (n =10 melanoma patients). FIG. IB (top) shows a design of the 12 Logistic Regression (LR) and 9 signature scoring models with their hyperparameters. FIG. IB (bottom) shows a model selection framework estimating the generalization performance of the model through a leave-one-patient-out nested-cross-validation. FIG. 1 C shows evaluation of the 12 LR (circles) and 9 signature scoring (triangles) binary classifiers in function of Matthew's Correlation Coefficient (MCC) and the area under the curve (AUC) (Table 2). The panel shows the distribution of the best model score for tumor antigen-specific and viral-specific clones. FIG. ID shows a volcano plot displaying the differential gene expression analysis comparing tumor- reactive and non-tumor-reactive cells. The 90 upregulated and down-regulated genes obtained by edgeR-QFL are shown (Table 4). FIGs. IE and IF show Alluvial plots showing the fractions of cells and clones annotated as tumor-reactive or non-tumor-reactive (orphan or Ag-specific) within internal (FIG. IE) and external (FIG. IF) benchmarking data. FIG. 1 G shows a Receiver Operating Characteristic (ROC) curve of TRTpred applied to the input data, the internal and external benchmarking data (Oliveira, G. et al. Nat. 2021 1-7 (2021)). FIG. 1H shows a heatmap of AUC performance of existing tumor-reactive predictive signatures (X-axis) applied to input data and internal and external benchmarking data (Y-axis). FIG. II shows fraction of neoantigen-specific clones in three external datasets (Zheng, C. et al. Cancer Cell 40, 410-423. e7 (2022); Lowery, F. J. et al. Science 375, eabl5447 (2022); Hanada, K.-I. et al. Cancer Cell. 2022 May 9;40(5):479- 493. e6.) predicted as tumor-reactive using TRTpred.
[0040] FIGs. 2A, 2B, 2C, 2D, 2E, 2F, 2G, 2H, and 21 show applications of TRTpred for discovery of immune repertoires and validation of MixTRTpred. FIG. 2A shows richness (top) and clonality (bottom) of inferred tumor-reactive (+) and non-tumor-reactive (-) clones using TRTpred in 77=42 patients from 5 cohorts: internal data (melanoma 17=14); (melanoma, 17=4); (GI, 77=5); (77=1 melanoma, 77=2 breast and 77=12 GI); (lung, 77=4). Patients are grouped according to the cancer type. Metrics are displayed on logarithmic scale, and statistics were performed using a one-tailed t-test. FIG. 2B shows proportion of inferred TRT CD8+T cells in melanoma (n=19) and other solid tumors (n=23, as described in FIG. 2A). Statistics were performed using a one-tailed Wilcoxon test. FIG. 2C shows sequential multiplexed immunohistochemistry of patient 1 with hematoxylin and CD8 staining, illustrating the spatial distribution of tumor and stromal compartments. Upper panel: whole-tissue section; lower panel (white rectangle): zoom-in. Scale bars, 500 and 50 pm, for the whole-tissue section and zoom-in images, respectively. FIGs. 2D and 2E show cumulative frequency of inferred high avidity (FIG. 2D) or tumor-reactive (FIG. 2E) CD8+clones identified in microdissected areas of stroma and tumor in five melanoma patients (patients 1-5). Statistics were performed using a one-tailed Wilcoxon test. FIG. 2F shows TRTpred results depicting the distance matrix of the top 20 ranked tumor-reactive among high structural avidity clones. The five TCR clusters are described by the dendrogram obtained through hierarchical clustering. The five individual clones selected have the highest tumor-reactive score in each cluster (TCR#l-5, Table 6). FIG. 2G shows in vitro validation of the tumor reactivity of the five TCRs (TCR#1 -5) predicted through MixTRTpred. TCR-transfected (or mock) T cells were exposed to autologous tumor cells, and CD137 upregulation on CD8+T cells (mean±SD) was analyzed by flow cytometry. The color code of the different clones corresponds to that of FIG. 2F. FIGs. 2H and 21 show that IL-2 NOG mice were subcutaneously engrafted with tumor cells from patient 14, followed by adoptive transfer of TCR-transduced primary CD8+ T cells. FIG. 2H shows tumor-bearing mice received 5x106CD8+T cells transduced (day 11) with TCR1, TCR3 or TCR5 (Table 6). FIG. 21 shows that mice were adoptively transferred with infusion products containing 1x106total CD8+T cells transduced (day 11) with either TCR1 or TCR3 or TCR5 (Single TCRs) or with a pool of 1x106total CD8+T cells transduced with the three TCRs (TCR cocktail, 0.33xl06CD8+T cells transduced with each TCR). In FIGs. 2H and 21, 5x106CD8+untransduced T cells were transferred as control. Data shows mean±SEM.
[0041] FIG. 3 shows an example process of MixTRTpred for identifying clinically relevant TCRs.
[0042] DETAILED DESCRIPTION OF THE INVENTION
[0043] The adoptive transfer of genetically engineered cells expressing specific T-cell receptors (TCRs) represents a promising alternative to cancer cellular therapy. By exploiting the distinct transcriptomic profile of tumor-reactive T-cells (TRT) relative to bystander cells, this disclosure provides TRTpred, an in-silico predictor of tumor-reactive TCRs that enables discovery and translational applications. This disclosure further provides MixTRTpred, a combinatorial algorithm that integrates TRTpred with an avidity predictor and / or a TCR-clustering tool to identify clinically relevant TCRs for personalized T-cell therapy.
[0044] Model Selection and TRTpred
[0045] This disclosure provides a novel, robust model selection approach using, for example, nested cross-validation to evaluate models while fine-tuning hyperparameters.
[0046] Accordingly, in one aspect, this disclosure provides a method of identifying a classifier for predicting tumor reactivity of T cells. In some embodiments, the method comprises: (a) determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells, wherein the transcriptomic profile comprises transcription level of each of a set of genes; (b) inputting the transcriptomic profile of the set of genes into a plurality of test classifiers; (c) determining, by each of the test classifiers, a result comprising tumor reactivity of each of the T cells; (d) determining a classification score of each of the test classifiers by a nested cross-validation framework based on the result determined by each of the test classifiers; and (e) ranking classification scores of the test classifiers and selecting, from the test classifiers, a candidate classifier having the classification score higher than a reference classification performance value.
[0047] In another aspect, this disclosure also provides a method of predicting tumor reactivity of T cells. In some embodiments, the method comprises: (i) determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells; wherein the transcriptomic profile comprises transcription level of each of a set of genes; (ii) inputting the transcriptomic profile of the set of genes into the classifier identified according to the method as described herein, (iii) determining tumor reactivity of each of the T cells by the classifier; and (iv) ranking tumor reactivity of the T cells and selecting, from the T cells, candidate T cells that are reactive to a tumor antigen (e.g., a tumor-associated antigen (TAA) or an intratumoral neoantigen).
[0048] Also provided is a system for identifying a classifier for predicting tumor reactivity of T cells. In some embodiments, the system comprises one or more processors configured to: (a) determine a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells, wherein the transcriptomic profile comprises transcription level of each of a set of genes; (b) input the transcriptomic profile of the set of genes into a plurality of test classifiers; (c) determine, by each of the test classifiers, a result comprising tumor reactivity of each of the T cells; (d) determine a classification score of each of the test classifiers by a nested cross-validation framework based on the result determined by each of the test classifiers; and (e) rank classification scores of the test classifiers and select, from the test classifiers, a candidate classifier having the classification score higher than a reference classification value. In another aspect, this disclosure also provides a system of predicting tumor reactivity of T cells. In some embodiments, the system comprises one or more processors configured to: (i) determine a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells; wherein the transcriptomic profile comprises transcription level of each of a set of genes; (ii) input the transcriptomic profile of the set of genes into the classifier identified according to the method as described herein, (iii) determine tumor reactivity of each of the T cells by the classifier; and (iv) rank tumor reactivity of the T cells and select, from the T cells, candidate T cells that are reactive to a tumor antigen (e.g., a tumor- associated antigen (TAA) or an intratumoral neoantigen).
[0049] In some embodiments, the methods as disclosed herein can be used to predict clinical responses to a tumor-infiltrating lymphocytes (TILs) therapy or a peripheral blood lymphocyte (PBL) therapy.
[0050] In some embodiments, the sequence data also comprises single cell TCR sequencing data.
[0051] In some embodiments, the nested cross-validation framework is implemented using a conventional or a leave-one-patient-out (LOPO) approach. In some embodiments, the method further comprises finetuning hyperparameters of the candidate classifier by a cross-validation framework. In some embodiments, the cross-validation framework is implemented using a conventional or a leave-one-patient-out (LOPO) approach.
[0052] Cross-validation (CV) is a resampling technique used to test and train a model on different iterations. It can be used to estimate how accurately a predictive model will perform in practice and also to finetune a model’s hyperparameters. Nested cross-validation (NCV) is a machine learning method that estimates the generalization performance of a model while finetuning its hyperparameters .
[0053] Leave-one-patient-out cross-validation (LOPOCV) is a method used to assess the performance of machine learning algorithms. It involves leaving out a single biopsy sample as the test case.
[0054] In some embodiments, test classifiers may include a classifier based on support vector machines (SVM), neural networks, naive Bayes classifiers, decision trees, Adaptive Boosting (AdaBoost), Extreme Gradient Boosting (XGBoost), discriminant analysis, or nearest neighbors (kNN). Examples of regression-type supervised machine learning techniques include linear regression, lasso regression, ridge regression, elasticnet regression, partial least squares regression, polynomial regression, random forests, SVM, XGBoost, Adaboost, nonlinear regression, generalized linear models, decision trees, and neural networks.
[0055] In some embodiments, the test classifiers comprise a logistic regression model. In some embodiments, the test classifiers comprise a signature scoring model. A signature scoring model such as gene signature scoring can be used for evaluating the strength of biological signals in a transcriptome in, for example, bulk and single-cell RNA sequencing (scRNA-seq) data.
[0056] In some embodiments, the candidate classifier comprises an edgeR-LRT or an edgeR-QFL. In some embodiments, the candidate classifier comprises an edgeR-QFL model. edgeR analyzes differential expression of replicated count data. It can be used for analyzing digital gene expression data, such as count data from DNA sequencing technologies.
[0057] In some embodiments, the classification scores are scaled based on a mean and standard deviation obtained from training data.
[0058] In some embodiments, the classification scores are characterized by area-under-the-curve (AUC) scores. In some embodiments, the classification scores are characterized by Matthew’s Correlation Coefficient (MCC). The Matthew’s correlation coefficient (MCC) is a statistical tool that measures the difference between predicted and actual values. It can be used in machine learning to evaluate the quality of binary and multiclass classifications.
[0059] In some embodiments, the candidate classifier has the highest classification score among the test classifiers.
[0060] As used herein, “a reference classification value” refers to a classification score that can be a mean or average classification score of test classifiers. In some embodiments, a reference classification value can be arbitrarily set.
[0061] In some embodiments, the tumor reactivity of the T cells comprises reactivity to a tumor antigen (e.g., a tumor-associated antigen (TAA) or an intratumoral neoantigen). In some embodiments, the set of genes comprise one or more genes set forth in Table 4. In some embodiments, the set of genes comprise CXCL13, LAG3, TOX, PDCD1, TNFRSF9, or a combination thereof. In some embodiments, the method further comprises validating the candidate classifier against spurious learning. Spurious learning occurs when a machine learning model learns to rely on features that are not causally related to the target variable. These features are often irrelevant or unnatural. Spurious correlation, or spuriousness, occurs when two factors appear causally related to one another but are not. This can happen when: similar movement on a chart turns out to be coincidental; or a third “confounding” factor is not being accounted for. In some embodiments, spurious learning was checked through Y-random permutation tests wherein the training and testing procedure was applied to the data composed of randomly permuted tumor- reactive annotated clones, repeating the process 100 times.
[0062] In some embodiments, the T cells are from one or more patients. In some embodiments, the candidate classifier can exclude virus-specific T cells (VST).
[0063] In some embodiments, the step of determining the result by each of the test classifiers comprises determining the result by applying hyperparameters to each of the test classifiers. In some embodiments, the hyperparameters comprise one or more of: (i) the choice of the log-Fold- Change or P-value for selecting differentially expressed genes from differential expression analyses; (ii) number of genes in the set of the genes; (iii) whether to include only up-regulated genes or both up-regulated and down-regulated genes in the set of the genes; and (iv) a scoring method to rank tumor reactivity of the T cells. In some embodiments, the scoring method comprises an Average scoring method, an AUCell scoring method, a Singscore scoring method, or a Ucell scoring method. In some embodiments, the hyperparameters comprise one or more of: using P-value for differential expression analysis, selecting 5 to 500 (e.g., 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500) up-regulated and down-regulated genes, and scoring the tumor reactivity of the T cells with the Average scoring method.
[0064] Log fold change is the log ratio of expression values of a gene or transcript in two different conditions. It can be reported in logarithmic scale (base 2). A P-value, or probability value, is a measure of how likely it is that a result could have occurred under a certain hypothesis. It can be used in hypothesis testing to help researchers avoid publication errors. AUCell is a method that identifies cells with active gene regulatory networks in single-cell RNA-seq data. It uses AUC to calculate whether a critical subset of the input gene set is enriched within the expressed genes for each cell. AUCell is a rank-based method and is calculated for each cell individually. AUCell values can range from 0 to 1.
[0065] Singscore is a rank-based metric that measures gene set enrichment in single samples. It scores individual samples independently. Singscore can be used for small data sets or data integration because it produces stable scores. The scores are directly interpretable as a normalized mean percentile rank.
[0066] UCell is an R package that scores gene signatures in single-cell datasets. UCell scores are normalized U statistics that range from 0 to 1. They are based on the Mann-Whitney U statistic and are robust to dataset size and heterogeneity. UCell scores are mathematically related to the area under the ROC curve. UCell can be applied to any single-cell data matrix and includes functions to interact with Seurat objects.
[0067] As used herein, a “classifier,” a “machine learning model,” or a “model” refers to a set of algorithmic routines and parameters that can predict an output(s) for a process input based on a set of input features, with or without being explicitly programmed. A structure of the software routines (e.g., number of subroutines and relation between them) and / or the values of the parameters can be determined in a training process, which can use actual results of the process that is being modeled. Such systems or models are understood to be necessarily rooted in computer technology, and, in fact, cannot be implemented or even exist in the absence of computing technology. While machine learning systems utilize various types of statistical analyses, machine learning systems are distinguished from statistical analyses by virtue of the ability to learn without explicit programming and being rooted in computer technology. A neural network or an artificial neural network is one set of algorithms used in machine learning for modeling the data using graphs of neurons. Any network structure may be used. Any number of layers, nodes within layers, types of nodes (activations), types of layers, interconnections, learnable parameters, and / or other network architectures may be used. Machine training uses the defined architecture, training data, and optimization to learn values of the learnable parameters of the architecture based on the samples and ground truth of training data. A typical machine learning pipeline may include building a machine learning model from a sample dataset (referred to as a “training set”), evaluating the model against one or more additional sample datasets (referred to as a “validation set” and / or a “test set”) to decide whether to keep the model and to benchmark how good the model is, and using the model in “production” to make predictions or decisions against live input data captured by an application service. For training the model to be applied as a machine-learned model, training data is acquired and stored in a database or memory. The training data is acquired by aggregation, mining, loading from a publicly or privately formed collection, transfer, and / or access. Ten, hundreds, or thousands of samples of training data are acquired. The samples are from scans of different patients and / or phantoms. Simulation may be used to form the training data. The training data includes the desired output (ground truth), such as segmentation, and the input, such as protocol data and imaging data.
[0068] In some embodiments, the training set is used to create a single classifier using any now or hereafter-known methods. In other embodiments, a plurality of training sets is created to generate a plurality of corresponding classifiers. Each of the plurality of classifiers can be generated based on the same or different learning algorithm that utilizes the same or different features in the corresponding one of the pluralities of training sets.
[0069] Once trained, the machine-learned or trained classifier is stored for later application. The training determines the values of the learnable parameters of the network. The network architecture, values of non-learnable parameters, and values of the learnable parameters are stored as the machine-learned network. Once stored, the machine-learned network may be fixed. The same machine-learned network may be applied to different patients, different scanners, and / or with different imaging protocols for the scanning. The machine-learned network may be updated. As additional training data is acquired, such as through application of the network for patients and corrections by experts to that output, the additional training data may be used to re-train or update the training.
[0070] For the machine learning model, input data structures of subreads can be used for the training. The input data structure may correspond to a window of in a sequence. Each training sample can include one of the first plurality of first data structures and a pocket prediction score likelihood of a residue being a ligand binding residue. The training is performed by optimizing parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels and optionally the second labels when the first plurality of first data structures and optionally the second plurality of second data structures are input to the model. An output of the model specifies whether a residue in sequence is likely a ligand-binding residue. In some embodiments, the output of the model may include a probability of being in each of a plurality of states. The state with the highest probability can be taken as the state.
[0071] MixTRTpred for Identifying Clinically Relevant TCRs
[0072] In another aspect, this disclosure also provides a method of identifying a tumor-reactive TCR. In some embodiments, the method comprises: (a) identifying tumor-reactive T cells by: determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data and single cell TCR sequencing data of T cells, wherein the transcriptomic profile comprises transcription level of each of a set of genes, inputting the transcriptomic profile of the set of genes into the classifier identified according to the method as described herein, determining tumor reactivity of each of the T cells by the classifier, and ranking tumor reactivity of the T cells and selecting, from the T cells, candidate T cells that are reactive to a tumor antigen (e.g., a tumor-associated antigen (TAA) or an intratumoral neoantigen); and identifying a tumor-reactive TCR from TCRs contained the candidate T cells by performing at least one of: (i) identifying candidate TCRs having high structural avidity from the TCRs of the candidate T cells by: selecting a set of amino acids in an amino acid sequence of a CDR30 of each of the TCRs, determining solvent accessibility of each amino acid of the set of amino acids, inputting the solvent accessibility of each amino acid of the set of amino acids into a supervised machine learning model (e.g., a logistic regression model), wherein the supervised machine learning model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids, determining a probability value of a high avidity status of each of the TCRs as an output of the supervised machine learning model, and identifying candidate TCRs as having high avidity against the antigen if probability values of the candidate TCRs are greater than or equal to a threshold value; and (ii) clustering the candidate TCRs based on physicochemical properties thereof into one or more TCR clusters that target distinct antigens, and selecting a top tumor- reactive TCR from each of the one or more TCR clusters. In yet another aspect, this disclosure also provides a system for identifying a tumor-reactive TCR. In some embodiments, the system comprises one or more processors configured to: (a) identify tumor-reactive T cells by the one or more processors configured to: determine a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data and single cell TCR sequencing data of T cells, wherein the transcriptomic profile comprises transcription level of each of a set of genes, input the transcriptomic profile of the set of genes into the classifier identified according to the method as described herein, determine tumor reactivity of each of the T cells by the classifier, and rank tumor reactivity of the T cells and select, from the T cells, candidate T cells that are reactive to a tumor antigen (e.g., a tumor- associated antigen (TAA) or an intratumoral neoantigen); and identify a tumor-reactive TCR from TCRs contained the candidate T cells by the one or more processors configured to at least one of: (i) identify candidate TCRs having high structural avidity from the TCRs of the candidate T cells by: select a set of amino acids in an amino acid sequence of a CDR30 of each of the TCRs, determine solvent accessibility of each amino acid of the set of amino acids, input the solvent accessibility of each amino acid of the set of amino acids into a supervised machine learning model (e.g., a logistic regression model), wherein the supervised machine learning model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids, determine a probability value of a high avidity status of each of the TCRs as an output of the supervised machine learning model, and identify candidate TCRs as having high avidity against the antigen if probability values of the candidate TCRs are greater than or equal to a threshold value; and (ii) cluster the candidate TCRs based on physicochemical properties thereof into one or more TCR clusters that target distinct antigens, and select a top tumor-reactive TCR from each of the one or more TCR cluster.
[0073] In some embodiments, the top tumor-reactive TCR is the TCR that is most reactive to the tumor antigen in each of the one or more TCR clusters.
[0074] As used herein, the supervised machine learning model may include, without limitation, analytical learning, artificial neural network, backpropagation, boosting (meta-algorithm), bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, gaussian process regression, group method of data handling, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), multilinear subspace learning, naive Bayes classifiers, maximum entropy classifiers, conditional random fields, nearest neighbor algorithms, probably approximately correct learning (PAC) learning, ripple down rules, support vector machines, minimum complexity machines (MCM), random forests, ensembles of classifiers, ordinal classification, data pre-processing, and statistical relational learning.
[0075] In some embodiments, the supervised machine learning model may include classificationtype supervised machine learning techniques or regression-type supervised machine learning techniques. Examples of classification-type supervised machine learning techniques include support vector machines (SVM), neural networks, naive Bayes classifiers, decision trees, Adaptive Boosting (AdaBoost), Extreme Gradient Boosting (XGBoost), discriminant analysis, and nearest neighbors (kNN). Examples of regression-type supervised machine learning techniques include linear regression, lasso regression, ridge regression, elasticnet regression, partial least squares regression, polynomial regression, random forests, SVM, XGBoost, Adaboost, nonlinear regression, generalized linear models, decision trees, and neural networks.
[0076] In some embodiments, the supervised machine learning model comprises a logistic regression model. As used herein, the term “logistic regression model” refers to the widely used statistical model that, in its basic form, may use a logistic function to model a binary dependent variable; many more complex extensions exist. In regression analysis, logistic regression (or logit regression) may estimate the parameters of a logistic model; it is a form of binomial regression. Mathematically, a binary logistic model has a dependent variable with two possible values, such as pass / fail, win / lose, alive / dead or healthy / sick; these are represented by an indicator variable, where the two values are labeled “0” and “1.” In the logistic model, the log-odds (the logarithm of the odds) for the value labeled “1” is a linear combination of one or more independent variables (“predictors”); the independent variables can each be a binary variable (two classes, coded by an indicator variable) or a continuous variable (any real value). The corresponding probability of the value labeled “1” can vary between 0 (certainly the value “0”) and 1 (certainly the value “1”), hence the labeling; the function that converts log-odds to probability is the logistic function, hence the name.
[0077] Identification of TCRs with High Avidity
[0078] As used herein, the term “avidity” refers to an informative measure of the overall stability or strength of the pMHC-TCR complex. It is controlled by three major factors, including TCR epitope affinity, the valency of both the antigen and TCR, and the structural arrangement of the interacting parts. These factors also define the specificity of TCR, that is, the likelihood that the TCR is binding to a precise antigen epitope.
[0079] As used herein, the term “affinity” refers to the strength of interaction between TCR and antigen at single antigenic sites. Within each antigenic site, the variable region of the TCR “arm” interacts through weak non-covalent forces with antigens at numerous sites; the more interactions, the stronger the affinity.
[0080] In some embodiments, the avidity of a TCR may be defined by either pMHC-TCR association (Kassoo or Ka) or dissociation kinetics (Kdissoo or Ka). The term “Kassoo” or “Ka,” as used herein, refers to the association rate of a particular TCR-antigen interaction, whereas the term “Kdissoo” or “Ka,” as used herein, refers to the dissociation rate of a particular TCR-antigen interaction. The term “KD,” as used herein, refers to the dissociation constant, which is obtained from the ratio of Ka to Ka(i.e., Ka / Ka) and is expressed as a molar concentration (M). KD values for TCRs can be determined using methods well established in the art. A method for determining the KD of a TCR is by using surface plasmon resonance, such as the biosensor system of Biacore®, or Solution Equilibrium Titration (SET) (see Friguet B et al. (1985) J. Immunol Methods; 77(2): 305-319, and Hanel C et al. (2005) Anal Biochem; 339(1): 182-184).
[0081] In some embodiments, the TCR with high avidity has a pMHC-TCR half-life of less than about 60 seconds (e.g., 5, 10, 15, 20, 25, 30, 40, 45, 55, 60, 65, 70, 75, or 80 seconds).
[0082] In some embodiments, a set of amino acids or a subset of amino acids of a TCR comprises 3 to 20 (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) or more ammo acids. In some embodiments, the set of amino acids or the subset of amino acids of a TCR is contained in the CDR30 of the TCR.
[0083] In some embodiments, the set of amino acids comprises one or more of Arg, Asn, Asp, Gly, He, Lue, and Phe, or a conservative substitution thereof. Suitable conservative substitutions of amino acids are known to those of skill in this art and may be made generally without altering the biological activity of the resulting molecule (see, e.g., Watson, et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. Co., p. 224).
[0084] In some embodiments, the method comprises determining the probability value by: wherein p is the probability value, bo represents a bias term, and R, N, D, G, I, L, and F are amino acids Arg, Asn, Asp, Gly, He, Leu, and Phe, respectively. The bias term / v>and the weights Wn were determined using a set of TCRs (with, e.g., 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more TCRs).
[0085] In some embodiments, the threshold value is about 0.4, 0.5, 0.6, 0.7, 0.8, or 0.9. In some embodiments, the threshold value is about 0.5 (e.g., 0.5).
[0086] In some embodiments, the method comprises selecting a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80%. In some embodiments, the method comprises selecting a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 30% (e.g., 30%).
[0087] In some embodiments, the method comprises inputting the solvent accessibility of the subset of amino acids into the supervised machine learning model (e.g., logistic regression model). In some embodiments, the method comprises determining the solvent accessibility of each amino acid of the set of amino acids as a relative solvent excluded surface area (SESA). In some embodiments, the SESA is determined by normalizing surface area of an amino acid in the TCR against surface area of the amino acid in a reference state. In some embodiments, the method comprises determining solvent accessibility of an amino acid based on a three-dimensional (3D) model of the TCR. The TCR 3D model may be generated by any suitable methods (e.g., X-ray crystallography, NMR, cryo-EM, homology modeling, or de novo folding).
[0088] In some embodiments, SESA can be computed with the MSMS package of the UCSF Chimera software, as described by Goddard, T. D. et al. (Goddard, T. D. et al. Protein Sci. Publ. Protein Soc. 27, 14-25 (2018)). SESA can also be calculated by normalizing the surface area of the residue in the TCR of interest by its surface area in a reference state (Bendell, C. et al. BMC Bioinformatics 15, 82 (2014)).
[0089] In some embodiments, the method may further include inputting into the supervised machine learning model (e.g., logistic regression model) one or more additional characteristics of each amino acid of the set of amino acids. In some embodiments, the additional characteristics comprise hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of 0-tum, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of P-tum, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in P-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, a-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in a-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed a-helix, heat capacity, free energy in a-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of P-sheet, unweighted, normalized frequency of P-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof.
[0090] In some embodiments, the additional characteristics comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof.
[0091] In some embodiments, the method comprises determining a nucleotide sequence of the TCR by sequencing, e.g., deep sequencing or ultra-deep sequencing. Deep sequencing yields a unique genetic fingerprint that can be used to identify a person, and a trove of predictors of genetic medical diseases. Deep sequencing to identify epigenetic events including changes in DNA methylation and RNA expression can reveal the history and impact of environmental exposures. Ultra-deep sequencing is the sequencing of amplicons at a high depth of coverage with the goal of identifying the common and rare sequence variations. With sufficient depth of coverage, ultradeep sequencing can fully characterize rare sequence variants down to less than 1%. Ultra-deep sequencing has been used to detect low-frequency HIV drug-resistant mutations or identify rare somatic mutations in complex cancer samples. For tests such as non-invasive blood tests, the frequency of biomarker mutation could be lower than 1%. TCR Clusterins
[0092] In some embodiments, the step of clustering the candidate TCRs comprises: (a) selecting a subset of amino acids in a first TCR and a corresponding subset of amino acids in a second TCR, wherein the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR have an identical number of amino acids; (b) determining an amino acid sum of differences in each of a plurality of physicochemical properties by performing a pairwise comparison between an amino acid in the subset of amino acids in the first TCR and a corresponding amino acid in the corresponding subset of amino acids in the second TCR; (c) repeating steps (a) to (b) for remaining amino acids in the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR; (d) determining a subset sum of differences between the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR; (e) repeating steps (a) to (d) for one or more subsets of amino acids in the first TCR and the second TCR; (f) determining an aggregate value of all subset sums of differences between the first TCR and the second TCR by assigning a weight value to each of the subset sums; and (g) identifying the first TCR and the second TCR as having similar specificity to an epitope if the aggregate value is smaller than a threshold value.
[0093] In some embodiments, the step of clustering the candidate TCRs comprises: (1) generating a similarity matrix for a subset of amino acids of a TCR of the TCRs, wherein the similarity matrix comprises a plurality of physicochemical properties of each amino acid in the subset of amino acids; (2) repeating step (1) for one or more subsets of amino acids of the TCR; (3) repeating steps (1) to (2) for remaining TCRs of the plurality of immunological entities; and (4) performing a clustering analysis based on a distance between two corresponding similarity matrices of a pair of immunological entities to identify a subset of TCRs having similar specificity to the epitope.
[0094] As used herein, “similarity” refers to the degree molecules are similar for molecules such as an immunological entity (e.g. , TCR) binder (e.g. , antigen) or epitope or a part thereof. Similarity can be determined based on a difference in physicochemical properties. Generally, the concept encompasses a broadly defined “structural similarity.” Although not wishing to be bound by any theory, it is understood that antibodies, TCRs, BCRs, or the like binding to an epitope belonging to an identical cluster can be assigned to a disease, disorder, symptom, physiological phenomenon, or the like in the same category when immunological entities or epitopes are classified based on such similarity in some of the embodiments of the present invention. Therefore, a variety of diagnoses (incidence of cancer, compatibility of administered drug, and the like) are made possible by studying whether there are antibodies, TCRs, BCRs, or the like that react to the same epitope cluster by using the disclosed methods.
[0095] As used herein, a “corresponding” amino acid refers to an amino acid which has, or is expected to have, in a certain polypeptide molecule, similar action as a predetermined amino acid in a benchmark polypeptide, and for enzyme molecules, refers to an amino acid which is present at a similar position in an active site and makes a similar contribution to catalytic activity. It is preferable to define identical residues when investigating a corresponding amino acid. A corresponding amino acid can be a specific amino acid subjected to, for example, cysteination, glutathionylation, S — S bond formation, oxidation (e.g., oxidation of methionine side chain), formylation, acetylation, phosphorylation, glycosylation, myristylation, or the like. Alternatively, a corresponding amino acid can be an amino acid responsible for dimerization. Such a “corresponding” amino acid or nucleic acid may be a region or a domain (e.g., V region, D region, or the like) over a certain range. Thus, it is referred to herein as a “corresponding” region, subset, or domain in such a case.
[0096] As used herein, the term “clustering analysis” refers to using a clustering algorithm to identify groups of members that are similar. It divides a data set so that records with similar content are in the same group, and groups are as different as possible from each other. When the categories are unspecified, this is sometimes referred to as unsupervised clustering. When the categories are specified a priori, this is sometimes referred to as supervised clustering.
[0097] In some embodiments, the clustering algorithms comprise partitional clustering, hierarchical clustering, k-nearest neighbor (KNN), K-means and fuzzy clustering, and Kohonen self-organizing maps clustering. In some embodiments, clustering includes agglomerative “bottom-up” or divisive “top-down” hierarchical clustering, distance “partition” clustering, and alignment clustering.
[0098] In some embodiments, the clustering analysis comprises a hierarchical clustering. In some embodiments, the hierarchical clustering comprises an unweighted pair group method with arithmetic mean (UPGMA). In some embodiments, the distance is a Manhattan distance. In some embodiments, the distance is a normalized distance obtained, e.g., by dividing each of them by the maximum calculated value for all possible pairs.
[0099] In some embodiments, the distance between two TCRs ranges from 0 (e.g., for two unrelated TCRs) to 1 (e.g., for two TCRs bearing an identical set of 4 residues on each considered loop, e.g., only CDR30 or all CDRs). In some embodiments, a threshold value can be used as an indicator to classify the similarity between TCRs. For example, a pair with a maximum distance found by clustering analysis using a hierarchical clustering methodology (e.g., group average method (average linkage clustering), nearest neighbor method (NN method), K-NN method, Ward method, furthest neighbor method, or centroid method) of less than a specific value can be deemed to be in identical cluster. Examples of such a value include, but are not limited to, less than 1 , less than 0.95, less than 0.9, less than 0.85, less than 0.8, less than 0.75, less than 0.7, less than 0.65, less than 0.6, less than 0.55, less than 0.5, less than 0.45, less than 0.4, less than 0.35, less than 0.3, less than 0.25, less than 0.2, less than 0.15, less than 0.1, less than 0.05, and the like. In some embodiments, a specific threshold value can be set for evaluation. For example, about 0.9 can be used to distinguish whether an entity belongs to an identical group or another group. To increase the degree of separation, the threshold value can be appropriately raised. When, for example, about 0.9 is used, the threshold value can be set higher to about 0.95 or the like.
[0100] In yet another aspect, this disclosure also provides a method of identifying one or more TCRs that bind specifically to an epitope. The method comprises: selecting a candidate TCR that binds specifically to the epitope; identifying at least one TCR that has similar specificity to an epitope with the candidate TCR according to the method as described above; and identifying at least one TCR as the one or more TCRs that bind specifically to the epitope.
[0101] In some embodiments, the subset of amino acids comprises 3 to 10 (e.g., 3, 4, 5, 6, 7, 8, 9, 10) amino acids. In some embodiments, the subset of amino acids comprises 3 to 10 (e.g., 3, 4, 5, 6, 7, 8, 9, 10) consecutive amino acids. In some embodiments, the subset of amino acids consists of 4 amino acids (e.g., 4 consecutive residues).
[0102] In some embodiments, it is important for predicting epitope specificity by considering the contribution of residues in all CDRs (e.g., CDRla, CDR2a, CDR3a, CDRip, CDR2P, and CDR3P of a TCR). Accordingly, in some embodiments, one or more subsets of amino acids are selected from ammo acids in CDRla, CDR2a, CDR3a, CDR10, CDR20, and CDR30.
[0103] In some embodiments, the step of determining an aggregate value in the method described herein comprises assigning a weight value of about 10% to about 35% (e.g., about 20%, about 25%, about 30%, about 35%) to the subset of amino acids in CDR3a or CDR30. In some embodiments, step (f) determining an aggregate value comprises assigning a weight value of about 5% to about 15% (e.g., about 5%, about 8%, about 10%, about 12%, about 14%, about 15%) to the subset of amino acids in CDRla, CDR2a, CDR10, or CDR20.
[0104] In some embodiments, the subset of amino acids does not include amino acids that are not solvent-exposed. In some embodiments, the subset of amino acids does not include amino acids in CDRla, CDR2a, CDR10, or CDR20 that have a relative solvent excluded surface area (SESA) of less than about 3% to about 8% (e.g., about 3%, about 4% about, 5%, about 6%, about 7%, about 8%). In some embodiments, the subset of amino acids does not include amino acids in CDR3a or CDR30 that have a SESA of less than about 15% to about 25% (e.g., about 15%, about 18%, about 20%, about 22%, about 25%).
[0105] In some embodiments, the physicochemical properties comprise amino acid attributes selected from hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of 0-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of 0-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in 0-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, a-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in a-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed a-helix, heat capacity, free energy in a-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of 0- sheet, unweighted, normalized frequency of 0-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, and percentage of buried residues.
[0106] In some embodiments, the physicochemical properties comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, and electrostatic charge.
[0107] In another aspect, this disclosure provides a TCR or antigen-binding fragment, or an immune cell, which is identified according to the methods described herein. In some embodiments, the immune cell comprises a lymphocyte. In some embodiments, the lymphocyte comprises a T cell or a natural killer (NK) cell. In some embodiments, the T cell comprises a CD8+ T cell or a CD4+ T cell. In some embodiments, the T cell comprises a human T cell. In some embodiments, T cell may be obtained from blood or from a tumor.
[0108] Lymphocytes are one subtype of white blood cells in the immune system. In some embodiments, lymphocytes may include tumor-infiltrating immune cells. Tumor-infiltrating immune cells consist of both mononuclear and polymorphonuclear immune cells (i.e., T cells, B cells, natural killer cells, macrophages, neutrophils, dendritic cells, mast cells, eosinophils, basophils, etc.) in variable proportions. In some embodiments, lymphocytes may include tumorinfiltrating lymphocytes (TILs). TILs are white blood cells that have left the bloodstream and migrated towards a tumor. TILs can often be found in the tumor stroma and within the tumor itself. In some embodiments, TILs are “young” T cells or minimally cultured T cells. In some embodiments, the young cells have a reduced culturing time (e.g., between about 22 to about 32 days in total). In some embodiments, the lymphocytes express CD27.
[0109] In some embodiments, the lymphocytes may be autologous, allogeneic, syngeneic, or xenogeneic with respect to the subject. In some embodiments, the lymphocytes are autologous to reduce an immunoreactive response against the lymphocyte when reintroduced into the subject for immunotherapy treatment.
[0110] In some embodiments, the T cells are CD8+ T cells. In some embodiments, the T cells are CD4+ cells. In some embodiments, the NK cells are CD 16+ CD56+ and / or CD57+ NK cells. NKs are characterized by their ability to bind to and kill cells that fail to express “self MHC / HLA antigens by the activation of specific cytolytic enzymes, the ability to kill tumor cells or other diseased cells that express a ligand for NK activating receptors, and the ability to release protein molecules called cytokines that stimulate or inhibit the immune response.
[0111] Also within the scope of this disclosure is a variant of the TCR identified according to the methods disclosed herein and an immune cell comprising the variant of the TCR.
[0112] As used herein, the term “variant” refers to a first molecule that is related to a second molecule (also termed a “parent” molecule). The variant molecule can be derived from, isolated from, based on or homologous to the parent molecule. A “functional variant” of a protein, as used herein, refers to a variant of such protein that retains at least partially the activity of that protein. Functional variants may include mutants (which may be insertion, deletion, or replacement mutants), including polymorphs, etc. Also included within functional variants are fusion products of such protein with another, usually unrelated, nucleic acid, protein, polypeptide, or peptide. Functional variants may be naturally occurring or may be synthetic.
[0113] In some embodiments, a variant of a TCR may include one or more conservative modifications. The TCR variant with one or more conservative modifications may retain the desired functional properties, which can be tested using the functional assays known in the art.
[0114] As used herein, the term “conservative sequence modifications” refers to amino acid modifications that do not significantly affect or alter the binding characteristics of the protein containing the amino acid sequence. Such conservative modifications include amino acid substitutions, additions, and deletions. Modifications can be introduced by standard techniques known in the art, such as site-directed mutagenesis and PCR-mediated mutagenesis. Conservative amino acid substitutions are ones in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art. These families include: amino acids with basic side chains (e.g., lysine, arginine, histidine); acidic side chains (e.g., aspartic acid, glutamic acid); uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine, tryptophan); nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine); beta-branched side chains (e.g., threonine, valine, isoleucine); and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine) includes one or more conservative modifications.
[0115] As used herein, the percent homology between two amino acid sequences is equivalent to the percent identity between the two sequences. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., % homology = # of identical positions / total # of positions x 100), taking into account the number of gaps, and the length of each gap, which need to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm, as described in the non-limiting examples below.
[0116] The percent identity between two amino acid sequences can be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl. Biosci., 4: 11-17 (1988)), which has been incorporated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. In addition, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch (J. Mol. Biol. 48:444-453 (1970)) algorithm, which has been incorporated into the GAP program in the GCG software package (available at www.gcg.com), using either a Blossum62 matrix or a PAM250 matrix, and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6.
[0117] The term “homolog” or “homologous,” when used in reference to a polypeptide, refers to a high degree of sequence identity between two polypeptides, or to a high degree of similarity between the three-dimensional structure or to a high degree of similarity between the active site and the mechanism of action. In some embodiments, a homolog has a greater than 60% sequence identity, and more preferably greater than 75% sequence identity, and still more preferably greater than 90% sequence identity, with a reference sequence. The term “substantial identity,” as applied to polypeptides, means that two peptide sequences, when optimally aligned, such as by the programs GAP or BESTFIT using default gap weights, share at least 75% sequence identity.
[0118] A peptide or polypeptide “fragment” as used herein refers to a less than full-length peptide, polypeptide or protein. For example, a peptide or polypeptide fragment can have at least about 3, at least about 4, at least about 5, at least about 10, at least about 20, at least about 30, at least about 40 amino acids in length, or single unit lengths thereof. For example, fragment may be 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or more amino acids in length. There is no upper limit to the size of a peptide fragment. However, in some embodiments, peptide fragments can be less than about 500 amino acids, less than about 400 amino acids, less than about 300 amino acids or less than about 250 amino acids in length. Also within the scope of this disclosure are the variants, mutants, and homologs with significant identity to the TCR. For example, such variants and homologs may have sequences with at least about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% sequence identity with the sequences of TCRs described herein.
[0119] In another aspect, this disclosure provides a method of treating cancer in a subject. In some embodiments, the method comprises administering to the subject an immune cell identified according to the method described herein or an immune cell produced according to the method of claim described herein or a composition comprising the immune cell. In some embodiments, the method comprises administering to the subject a therapeutically effective amount of immune cells identified according to the method described herein or a therapeutically effective amount of immune cells produced according to the method of claim described herein.
[0120] In some embodiments, the method comprises administering to the subject an additional agent, such as an anti-cancer or anti-tumor agent.
[0121] In some embodiments, the immune cell comprises a lymphocyte. In some embodiments, the lymphocyte comprises a T cell or a natural killer (NK) cell. In some embodiments, the T cell comprises a CD8+ T cell or a CD4+ T cell. In some embodiments, the T cell comprises a human T cell.
[0122] In some embodiments, the immune cells can be used in adoptive T cell therapy (ADT). Generally, adoptive T cell therapy relies on the in vitro expansion of endogenous, cancer-reactive T cells. These T cells can be harvested from cancer patients, manipulated, and then reintroduced into the same or a different patient as a mechanism for generating productive tumor immunity. In some embodiments, CD8+ cytotoxic T lymphocytes can be used in adoptive T cell therapy. CD4+ T cells can also play an important role in maintaining CD8+ cytotoxic function, and transplantation of tumor-reactive CD4+ T cells has been associated with some efficacy in metastatic melanoma. T cells used in adoptive therapy can be harvested from a variety of sites, including peripheral blood, malignant effusions, resected lymph nodes, and tumor biopsies. Although T cells harvested from the peripheral blood are easier to obtain technically, TILs obtained from biopsies may contain a higher frequency of tumor-reactive cells. Once harvested, T cells can be transfected with a vector comprising a nucleotide sequence encoding the TCR.
[0123] In some embodiments, a TCR or antigen-binding fragment as disclosed has antigen specificity for an antigen that is characteristic of a disease or disorder. The disease or disorder can be any disease or disorder involving an antigen, such as but not limited to a tumor and / or a cancer, an infectious disease, or an autoimmune disease.
[0124] In some embodiments, the subject is a human. In some embodiments, the subject has a cancer. In some embodiments, the subject is immune-depleted.
[0125] As used herein, “cancer,” “tumor,” and “malignancy” all relate equivalently to hyperplasia of a tissue or organ. If the tissue is a part of the lymphatic or immune system, malignant cells may include non-solid tumors of circulating cells. Malignancies of other tissues or organs may produce solid tumors. The methods described herein can be used in the treatment of lymphatic cells, circulating immune cells, and solid tumors.
[0126] Cancers that can be treated include tumors that are not vascularized or are not substantially vascularized, as well as vascularized tumors. Cancers may comprise non-solid tumors (such as hematologic tumors, e.g., leukemias and lymphomas) or may comprise solid tumors. The types of cancers to be treated with the disclosed compositions include, but are not limited to, carcinoma, blastoma, and sarcoma, and certain leukemias or malignant lymphoid tumors, benign and malignant tumors, and malignancies, e.g., sarcomas, carcinomas, and melanomas. Also included are adult tumors / cancers and pediatric tumors / cancers.
[0127] Cancers that may be treated by methods described herein include, but are not limited to, cancer cells from the bladder, blood, bone, bone marrow, brain, breast, colon, esophagus, gastrointestine, gum, head, kidney, liver, lung, nasopharynx, neck, ovary, prostate, skin, stomach, testis, tongue, or uterus.
[0128] The immune cells, as described, can be administered in a manner appropriate to the disease to be treated or prevented. The amount and frequency of administration will be determined by factors such as the condition of the patient, and the type and severity of the patient’s disease, although appropriate dosages can be determined by clinical trials. When “a therapeutically effective amount,” “an immunologically effective amount,” “an effective antitumor quantity,” or “an effective tumor-inhibiting amount” is indicated, the precise amount of the compositions of the present disclosure to be administered can be determined by a physician having account for individual differences in age, weight, tumor size, extent of infection or metastasis, and patient’s condition. It can generally be stated that a pharmaceutical composition comprising the lymphocytes described herein can be administered at a dose of 104to 109cells / kg body weight, e.g., 105to 106cells / kg body weight, including all values integers within these intervals. The lymphocyte compositions can also be administered several times at these dosages. The cells can be administered using infusion techniques that are commonly known in immunotherapy (see, for example, Rosenberg et al., New Eng. J. of Med. 319: 1676, 1988). The optimal dose and treatment regimen for a particular patient can be readily determined by one skilled in the art of medicine by monitoring the patient for signs of the disease and adjusting the treatment accordingly.
[0129] In some embodiments, the immune cells are administered to the subject in a manner compatible with the dosage formulation and in such amount as is therapeutically effective. Dose ranges and frequency of administration can vary depending on, e.g., the nature of the population of cells (e.g., antigen-specific lymphocytes) produced by the methods described herein and the medical condition as well as parameters of a specific patient and the route of administration used.
[0130] The administration of the immune cells or compositions thereof can be carried out in any convenient way, including infusion or injection (i.e., intravenous, intrathecal, intramuscular, intraluminal, intratracheal, intraperitoneal, or subcutaneous), transdermal administration, or other methods known in the art. Administration can be once every two weeks, once a week, or more often, but the frequency may be decreased during a maintenance phase of the disease or disorder. In some embodiments, the immune cells or compositions thereof are administered by intravenous infusion.
[0131] Additional Definitions
[0132] To aid in understanding the detailed description of the compositions and methods according to the disclosure, a few express definitions are provided to facilitate an unambiguous disclosure of the various aspects of the disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0133] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless specified otherwise.
[0134] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. In some embodiments, the flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, a segment, or a portion of instructions, which may include one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0135] These computer readable program instructions may be provided to a processor of a general- purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0136] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0137] It will be understood that, although the terms “first,” “second,” etc., may be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Thus, a first element, component, region, layer or section discussed below could be termed a second element, component, region, layer or section without departing from the teachings of example embodiments.
[0138] Unless specifically stated otherwise, as apparent from the above discussion, it is appreciated that throughout the description, discussions utilizing terms such as “processing,” “performing,” “receiving,” “computing,” “calculating,” “determining,” “identifying,” “displaying,” “providing,” “merging,” “combining,” “running,” “transmitting,” “obtaining,” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (or electronic) quantities within the computer system memories or registers or other such information storage, transmission or display devices.
[0139] As used herein, the term “if may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event],” depending on the context.
[0140] As used herein, “T cell receptor (TCR)” is also called a T cell antigen receptor. A T cell receptor refers to a receptor recognizing an antigen expressed on a cell membrane of a T cell that plays a central role in the immune system. TCRs have an a chain, 0 chain, y chain, and 8 chain, with which an a or y8 dimer is constituted. TCRs consisting of the combination of the former are called a0 TCRs, and TCRs consisting of the combination of the latter are called y8 TCRs. T cells having such TCRs are respectively called a0 T cells and y8 T cells. The TCRs are structurally very similar to a Fab fragment of an antibody produced by B cells and recognize antigen molecules bound to an MHC molecule. Since a TCR gene of a mature T cell has undergone gene rearrangement, an individual has highly diverse TCRs that enable recognition of various antigens. TCRs also form a complex by binding to a non-variable CD3 molecule at the cell membrane. CD3 has an amino acid sequence called IT AM (immunoreceptor tyrosine-based activation motif) in the intracellular region. This motif is involved in intracellular signaling. Each TCR chain is comprised of a variable domain (V) and a constant domain (C). A constant domain has a short cytoplasm section penetrating the cell membrane. A variable domain is present outside the cell and binds to an antigen-MHC complex. A variable domain has three hypervariable domains or regions called complementarity-determining regions (CDRs), which bind to an antigen-MHC complex. The three CDRs are called CDR1, CDR2, and CDR3. TCR gene rearrangement is similar to the process of B cell receptors known as immunoglobulins. For gene rearrangement of a0 TCRs, VD J recombination of 0 chain is performed, followed by VJ recombination of an a chain. When the a chain is rearranged, the gene of the 8 chain is deleted from the chromosome. Thus, a T cell having an a0 TCR would never have a y8 TCR simultaneously. In contrast, a signal via a y8 TCR in a T cell having the TCR suppresses the expression of 0 chain, so that a T cell having a y8 TCR would never have an a0 TCR simultaneously. In some embodiments, the immune cell comprises a lymphocyte. In some embodiments, the lymphocyte comprises a T cell or a natural killer (NK) cell. In some embodiments, the T cell comprises a CD8+ T cell or a CD4+ T cell. In some embodiments, the T cell comprises a human T cell.
[0141] As used herein, the term “antigen” is a molecule and / or substance that can generate peptide fragments that are recognized by a TCR and / or induces an immune response. An antigen may contain one or more “epitopes.” In some embodiments, the antigen has several epitopes. An epitope is recognized by a TCR, an antibody or a lymphocyte in the context of an MHC molecule. In some embodiments, an “antigen” may be a neoantigen or a tumor-associated antigen (TAA).
[0142] As used herein, the term “neoantigen” is an antigen that has at least one alteration that makes it distinct from the corresponding wild-type, parental antigen, e.g., via mutation in a tumor cell or post-translational modification specific to a tumor cell. A neoantigen can include a polypeptide sequence or a nucleotide sequence. A mutation can include a frameshift or non- frameshift indel, missense or nonsense substitution, splice site alteration, genomic rearrangement, or gene fusion, or any genomic or expression alteration giving rise to a neoORF. A mutation can also include a splice variant. Post-translational modifications specific to a tumor cell can include aberrant phosphorylation. Post-translational modifications specific to a tumor cell can also include a proteasome-generated spliced antigen (Liepe et al., Science. 2016 Oct 21;354(6310):354-358). As used herein the term “tumor neoantigen” is a neoantigen present in a subject’s tumor cell or tissue but not in the subject’s corresponding normal cell or tissue.
[0143] As used herein, the terms “tumor-associated antigen,” “TAA,” and “cancer antigen” refer to any molecule (e.g., protein, peptide, lipid, carbohydrate, etc.) solely or predominantly expressed or over-expressed by a tumor cell and / or a cancer cell, such that the antigen is associated with tumor and / or cancer. The TAA / cancer antigen can also be expressed by normal, non-tumor, or non-cancerous cells. However, in such a situation, the expression of the TAA / cancer antigen by normal, non-tumor, or non-cancerous cells is, in some embodiments, not as robust as the expression of the TAA / cancer antigen by tumor and / or cancer cells. Thus, in some embodiments, the tumor and / or cancer cells overexpress the TAA and / or express the TAA at a significantly higher level as compared to the expression of the TAA by normal, non-tumor, and / or non- cancerous cells. In some embodiments, the phosphopeptides are fragments of TAAs or TAAs themselves. The TAA can be an antigen expressed by any cell of any cancer or tumor, including the cancers and tumors described herein. The TAA can be a TAA of only one type of cancer or tumor, such that the TAA is associated with or characteristic of only one type of cancer or tumor. Alternatively, the TAA can be characteristic of more than one type of cancer or tumor. For example, the TAA can be expressed by both breast and prostate cancer cells and not expressed at all by normal, non-tumor, or non-cancer cells. As used herein, the term “classification” refers to any number or other characters that are associated with a particular property of a sample. The classification can be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale from 1 to 10 or 0 to 1). The term “cutoff’ or “threshold” refers to a predetermined number used in an operation. For example, a cutoff value can refer to a classification score as used above. A threshold value may be a value above or below which a particular classification applies. Either of these terms can be used in either of these contexts.
[0144] The term “machine learning,” as used herein, refers to a computer algorithm used to extract useful information from a database by building probabilistic models in an automated way.
[0145] The term “regression tree,” as used herein, refers to a decision tree that predicts values of continuous variables.
[0146] The term “supervised learning,” as used herein, refers to a data analysis using a well- defined (known) dependent variable. All regression and classification algorithms are supervised. In contrast, “unsupervised learning” refers to the collection of algorithms where groupings of the data are defined without the use of a dependent variable.
[0147] The term “test data” refers to a data set independent of the training data set, used to evaluate the estimates of the model parameters (i.e., weights).
[0148] As used herein, the term “clustering tree” refers to a hierarchical tree structure in which observations, such as organisms, genes, and polynucleotides, are separated into one or more clusters. The root node of a clustering tree consists of a single cluster containing all observations, and the leaf nodes correspond to individual observations. A clustering tree can be constructed based on a variety of characteristics of the observations. Many techniques known in the art, e.g., hierarchical clustering analysis, can be used to construct a clustering tree. A non-limiting example of a clustering tree is a phylogenetic, taxonomic or evolutionary tree.
[0149] As used herein, the term “in vitro” refers to events that occur in an artificial environment, e.g., in a test tube or reaction vessel, in cell culture, etc., rather than within a multi-cellular organism.
[0150] As used herein, the term “in vivo” refers to events that occur within a multi-cellular organism, such as a non-human animal. It is noted here that, as used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise.
[0151] The terms “including,” “comprising,” “containing,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional subject matter unless otherwise noted.
[0152] The phrases “in one embodiment,” “in various embodiments,” “in some embodiments,” and the like are used repeatedly. Such phrases do not necessarily refer to the same embodiment, but they may unless the context dictates otherwise.
[0153] The terms “and / or” or means any one of the items, any combination of the items, or all the items with which this term is associated.
[0154] The word “substantially” does not exclude “completely,” e.g., a composition which is “substantially free” from Y may be completely free from Y. Where necessary, the word “substantially” may be omitted from the definition of this disclosure.
[0155] As used herein, the term “approximately” or “about,” as applied to one or more values of interest, refers to a value that is similar to a stated reference value. In some embodiments, the term “approximately” or “about” refers to a range of values that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the stated reference value unless otherwise stated or otherwise evident from the context (except where such number would exceed 100% of a possible value). Unless indicated otherwise herein, the term “about” is intended to include values, e.g., weight percents, proximate to the recited range that are equivalent in terms of the functionality of the individual ingredient, the composition, or the embodiment.
[0156] It is to be understood that wherever values and ranges are provided herein, all values and ranges encompassed by these values and ranges are meant to be encompassed within the scope of the present disclosure. Moreover, all values that fall within these ranges, as well as the upper or lower limits of a range of values, are also contemplated by the present application.
[0157] As used herein, the term “each,” when used in reference to a collection of items, is intended to identify an individual item in the collection but does not necessarily refer to every item in the collection. Exceptions can occur if explicit disclosure or context clearly dictates otherwise. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of this disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of this disclosure.
[0158] All methods described herein are performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. Regarding any of the methods provided, the steps of the method may occur simultaneously or sequentially. When the steps of the method occur sequentially, the steps may occur in any order, unless noted otherwise.
[0159] In cases in which a method comprises a combination of steps, each and every combination or sub-combination of the steps is encompassed within the scope of the disclosure, unless otherwise noted herein.
[0160] Each publication, patent application, patent, and other reference cited herein is incorporated by reference in its entirety to the extent that it is not inconsistent with the present disclosure. Publications disclosed herein are provided solely for their disclosure prior to the filing date of the present disclosure. Nothing herein is to be constmed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0161] It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims.
[0162] Examples
[0163] EXAMPLE 1
[0164] This example describes the materials and methods employed by the subsequent examples.
[0165] Cancer patient data collection
[0166] Single-cell RNA / TCR-Seq data from baseline tumors were obtained from internal patients with locally advanced (stage III) or metastatic (stage IV) cutaneous melanoma who have progressed on at least one standard first line therapy. Eleven patients were enrolled in a single- center phase I clinical study of TIL therapy (ClinicalTrials.gov NCT03475134) and three additional melanoma patients with available tumor samples were used as an internal benchmarking of the predictor of tumor-reactivity. Tumor samples were obtained following surgery and processed as previously described for single-cell analysis (Chiffelle, J. et al. bioRxiv 2023.07.21.544585 (2023); Barras, D. et al. bioRxiv 2022.12.23.519261 (2022)). TRTpred was applied on external data from Oliveira et al. (Oliveira, G. et al. Nat. 2021 1-7 (2021) (melanoma, 77=4); Zheng et al. (Zheng, C. et al. Cancer Cell 40, 410-423. e7 (2022)) (GI, 77=5); Lowery et al. (Lowery, F. J. et al. Science 375, eabl5447 (2022)) (77 =1 melanoma, 77=2 breast and / 7=12 GI) and Hanada etal. (Hanada KI, etal. Cancer Cell. 2022 May 9;40(5):479-493.e6) (lung, 77=4). All data was processed following the authors guidelines. All patients are referenced in Table 1.
[0167] In vivo study
[0168] Interleukin-2 (IL-2) NOD / Shi-scid IL2rgamma(null) (NOG) mice, constitutively expressing human IL-2, were obtained from Taconic Biosciences and maintained in a conventional animal facility at the University of Lausanne under specific pathogen-free status, with dark / light cycles of 12 h, humidity 55% (±10%) and a temperature of 22 °C (±1 °C). Twenty to twenty-five- week-old male and female mice (randomly selected) were used in the experiment. This study was approved by the Veterinary Authority of the Canton de Vaud (under the license 3746) and performed following Swiss ethical guidelines. Optional: mice were monitored 3 times a week and given a score based on their weight, behavior, physical condition, dehydration, breathing and tumor burden. As per the protocol, animals reaching a defined score were sacrificed.
[0169] TCR cloning and tumor-reactivity validation
[0170] TCRs from the ten internal melanoma patients were selected annotated following the rationale described in Chiffelle et al. (Chiffelle, J. et al. bioRxiv 2023.07.21.544585 (2023)). Briefly, TCR antitumor reactivity was interrogated by transferring RNA coding for TCRo.p pairs into recipient activated T cells and Jurkat cell line (TCR / CD3 Jurkat-luc cells (NF AT), Promega, stably transduced with human CD8a / p and TCRa / p CRISPR-KO). Electroporated cells were cocultured with IFNy-treated autologous tumor cells and tumor-reactivity was assessed through CD137-upregulation or bioluminescence assay for T cells and Jurkat cells, respectively. 102 tumor-reactive and 123 non-tumor-reactive CD8+TCRs were acquired from 10 patients that were used for building the final model of tumor-reactivity prediction and 46 tumor-reactive and 44 non- tumor-reactive CD8+TCRs from four patients used for the model benchmarking (Table 1). For the dose-response curve, autologous activated T cells electroporated with TCR 1, 3 or 5 were cocultured with tumor cells in fFNy assay, using pre-coated 96-well ELISpot plates (Mabtech). Transfected activated T cells were plated at 5x103cells per well and challenged with limiting dilutions of IFNy-treated autologous tumor cells (range 1x105to 250 cells). After 18-20 hours of incubation, cells were removed, the plate developed according to the manufacturer’s instructions and counted using a Bioreader 6000-E (BioSys).
[0171] Statistical models to predict tumor reactivity
[0172] Two different approaches were used to predict cell- wise tumor-specificity: the signature score and the logistic regression (LR) approach. The models first predict cell-wise tumorspecificity from scRNA-seq data which is then inferred on the TCR repertoire. The clone-wise score corresponds to the largest tumor reactivity score obtained by any cell from a given clone.
[0173] The signature score approach uses differential expression analysis methods to derive a signature of tumor-specificity which is then used to score cells. To allow comparison between the training and testing scores, the score was scaled based on the mean and standard deviation obtained from the training data. A threshold was then identified to stratify cells into tumor-specific and nonspecific cells maximizing accuracy. The score thresholds were identified in the training data and applied on the testing data. For this approach, 9 different models were constructed upon the selection of 9 differential expression analysis methods (DESeq2-Wald / -LRT, edgeR-QFL / -LRT, limma-trend / -voom, Seurat-LR / -negbinomial / -wilcox). These methods were carefully selected based on differential expression analysis benchmarking articles (Mou, T., et al. Front. Genet. 10, 1331 (2020); Soneson, C. & Robinson, M. D. Bias. Nat. Methods 2018 154 15, 255-261 (2018)) and applied following best practices (see the section “Differential Expression analysis methods”). Other parameters such as how to select the genes from the differential expression analysis method (z. e. , by considering only the log-Fold-Change or the P-value), the number of genes in the signature (i.e., signature length), whether to take only the up-regulated genes or both up and down regulated genes (i.e., signature side) and the signature score method (Average, AUCell, Singscore, UCell) were defined as hyperparameters to fine-tune.
[0174] The LR approach uses a standard logistic regression coupled with an elastic-net regularization. The type and strength of regularization was defined, respectively, by the a and A hyperparameters, where a = 0 corresponds to Ridge regression and a = 1 to Lasso regression. From this definition, LR models were derived based on two different feature spaces: First the scaled expression data (RNA) and secondly the Principal Components (PCs). For the RNA feature space, the genes correlating more than 80% were removed. Because of the large dimensionality, and in addition to the elastic net, filtering non-essential features were also tested. For both feature spaces, two different dimensionality reduction methods were used. The first consists in keeping only features associated with tumor-specificity (Wilcoxon - P-value < 1%). For the RNA feature space, the differential expression analysis method mentioned above was also applied only to keep significant genes (Bonferroni adjusted p-value < 0.05). Finally, another model was constructed solely on PCs explaining more than 10'4of the explained variances. The combination led to 10 RNA and 2 PCs based models.
[0175] Training and evaluation framework
[0176] The evaluation of the 21 models and their associated hyperparameters combinations were performed using a Nested Cross Validation (NCV). This robust framework allows to iteratively train and test the models on different data partitions called folds. For this application, a leave-one- patient-out (LOPO) NCV designed to partition the data into training folds composed of data from all patients but one which constitutes the testing fold was used. Iteratively, this approach allows to simulate the evaluation of the model on new unobserved data from other patients. For the sake of robustness, the performance metric chosen to evaluate the model is the reliable Matthews Correlation Coefficient (MCC) (Chicco, D. & Jurman, G. BMC Genomics 21, 1-13 (2020)). The LOPO evaluation serves as the final model evaluation and the best model’s hyperparameters were fined tuned using a similar LOPO Cross Validation (CV). A final tumor-specificity model was obtained by training the model using its best configuration on the whole input data.
[0177] Model benchmarking
[0178] To establish the robustness of the model's association with tumor reactivity, the training and evaluation framework was subjected to y-randomization tests by applying the same methods with data composed of randomly permuted tumor-reactive annotated clones, repeating the process 100 times. This extensive analysis yielded an average MCC of 0 and an average accuracy of 50%, indicative of the model's immunity to spurious learning and providing strong validation for its credibility. The efficacy of the model was further validated on 100 and 205 CD8+tumor-reactive annotated clonotypes, sourced from 4 internal and 4 external melanoma tumor biopsies (Oliveira, G. et al. Nat. 2021 1-7 (2021). To further explore the generalization potential of TRTpred, it was applied on 3 external TILs dataset (Zheng, C. et al. Cancer Cell 40, 410-423. e7 (2022) (GI, n=5), (Lowery, F. J. et al. Science 375, eabl5447 (2022) (melanoma, 77=l;breast, n=2„ GI, 77=12) and (Hanada KI, et al. Cancer Cell. 2022 May 9;40(5):479-493.e6) (Lung, 77=4). Finally, the discriminant power of the different models was compared to others by computing the Area Under the ROC Curve (AUC) of the models and external signature on the different data sources. The external signatures were applied on each data by using the signature score method described in the respective manuscripts. If not mentioned a simple average signature score was computed.
[0179] Differential expression analysis methods
[0180] 9 different differential expression analysis methods found to work best in scRNA-seq data were applied following best practices guidelines (Soneson, C. & Robinson, M. D. Nat. Methods 2018 154 15, 255-261 (2018); Squair, J. W. et al. Nat. Commun. 2021 121 12, 1-15 (2021)). These methods were grouped into two classes: the methods developed for single-cell RNA- sequencing data (LR, Negative-Binomial & Wilcoxon) and the methods originally developed for bulk RNA-sequencing data (edgeR, DESeq2, and limma), named pseudo-bulk method for the sake of clarity. The single-cell methods were implemented using the Seurat FindMarkers function on the loglO-normalized UMI counts with all filters (min.pct, only.pos, logfc. threshold, min. cells. roup) disabled. To ensure the performance of the pseudo-bulk methods, they were applied on the clone-average loglO-normalized UMI counts. This was done to obtain data resembling bulk-RNA-sequencing data distribution (which reduces inconsistencies in pseudo-bulk methods, while retaining the behavior of the clone transcription. EdgeR was applied both with the likelihood ratio test (edgeR-LRT - with default dispersion estimate) and with the quasi-likelihood approach (edge-QFL). DESeq2 was applied both with the Wald test of the negative binomial model coefficients (DESeq2-Wald) and with a likelihood ratio test compared to a reduced model (DESeq2-LRT). Limma was applied using two approaches: one incorporating the mean-variance trend into the empirical Bayes procedure at the gene level (limma-trend) and the other incorporating the mean-variance trend by assigning a weight to each individual observation (limma-voom). Log-transformed counts per million values computed by edgeR were provided as input to limma-trend. Tumor microdissection and RNA extraction
[0181] Consecutive sections from fresh-frozen blocks were cut at 8 m in a cryostat and mounted on pre-cooled PET slides (Leica, Mannheim, Germany) at -20°C for Ih. Tissue sections were fixed in ethanol, stained with freshly prepared cresyl violet, and cleared in ethanol. Each fresh-frozen section was microdissected within 20 min after staining, using the Laser Microdissection Systems Leica LMD7000. Laser parameters were set as follows: laser power of 39 mW, a wavelength of 349 nm, pulse frequency of 664 Hz, pulse energy of 58 j. Microdissected tissues were collected in 0.5 mL tubes in / M4 later® solution (Thermo Lisher Scientific, MA, USA) and kept at -20°C until RNA extraction. sRNA was extracted using RNeasy Plus Micro Kit (Qiagen, Hilden, Germany) according to the supplier’s protocol. Total RNA was quantified using Qubit™ RNA HS Assay Kit and Qubit® Lluorometer (Thermo Lisher Scientific, MA, USA). RNA fragment size was analyzed with the 2100 Bioanalyzer (Agilent, CA, USA) using Eukaryote total RNA Pico assay (Agilent, CA, USA).
[0182] Sequential Multiplexed immunohistochemistry
[0183] Fresh-frozen tissue sections were cut at 4 pm, fixed in paraformaldehyde (PF A) (4%) overnight, and permeabilized in 0.5% Triton X-100 in PBS for 30 min. Slides were subjected to heat-mediated antigen retrieval in a citrate buffer of pH 6.0 for 10 min. Endogenous peroxidases, non-specific proteins, endogenous biotins and avidins were blocked with corresponding blocking solutions from Dako®. The first primary antibody was applied to tissue sections, followed by a biotinylated-secondary antibody and a streptavidin-HRP complex, and revealed by AEC chromogen. Tissues were counter-stained by Harris hematoxylin for 1 min and coated by a glass coverslip with an aqueous mounting solution. Slides were scanned into MRXS images using Pannoramic 250 Flash III scanner (3D Histech, Budapest, Hungary). Glass coverslips were removed by immersion in hot water and AEC staining was eliminated by immersion in ethanol of increasing concentrations. Antibodies were stripped by boiling tissue sections in a citrate buffer of pH 6.0 for 10 to 20 min, and putative residual antibodies were blocked by Fab fragments directed against the host of the previous primary antibody used. Multiplexed IHC consisted of sequential cycles of staining with primary antibodies revealed by AEC chromogen, tissue section scanning, removal of AEC chromogen with ethanol, antibody stripping and blocking with Fab fragments. Primary antibodies used for multiplexed IHC were F0XP3 (clone ab99963, Abeam®, dilution 1: 50) and CD8 (clone, C8 / 144B, Dako®, dilution 1: 20).
[0184] Bulk TCR a and P sequencing
[0185] Bulk TCR-Sequencing was performed as described (Genolet, R. etal. Cell reports methods 3, (2023)). Briefly, mRNA was isolated using the Dynabeads mRNA DIRECT purification kit (Life Technologies) and was amplified using the MessageAmp II aRNA Amplification Kit (Ambion) with the following modifications: in vitro transcription (IVT) was performed at 37°C for 16 hours. First, strand cDN A was synthesized using Superscript III (Thermo Fisher Scientific) and a collection of TRAV / TRBV-specific primers. TCRs were then amplified by PCR (20 cycles with the Phusion from New England Biolabs, NEB) with a single primer pair binding to the constant region and the adapter linked to the TRAV / TRBV primers added during the reverse transcription. A second round of PCR cycle (25 cycles with the Phusion from NEB) was performed to add the Illumina adapters containing the different indexes. The TCR products were purified with AMPure XP beads (Beckman Coulter), quantified, and loaded on the MiSeq or NextSeqlOOO instrument (Illumina) for deep sequencing of the TCRa / TCRP chain. The TCR sequences were further processed using ad hoc Perl scripts to (i) pool all TCR sequences coding for the same protein sequence; (ii) filter out all out-frame sequences; and (iii) determine the abundance of each distinct TCR sequence. TCR sequences with a single read were not considered for analysis.
[0186] TCR repertoire metrics
[0187] To analyze TCR repertoires, two metrics were used: (i) the richness corresponding to the total number of unique clones present in the repertoire and (ii) the clonality described by the metric 1-Pielou’s evenness, as previously disclosed (Chiffelle, J. et al. Curr. Opin. Biotechnol. 65, 284- 295 (2020)).
[0188] MixTRTpred - integration of TCR structural avidity and TCR clustering
[0189] Prediction of TCR structural avidity were performed as described (Schmidt, J. et al. Cell Reports Med. 2, 100194 (2021)). In brief, a binary logistic regression, based on the CDR30 amino acids that are sufficiently solvent exposed to interact with the cognate peptide, was used to determine whether a TCR is likely to bind the corresponding pMHC with high or low structural avidity (koff). Avidity levels were also computed with this model on assessable patients, i.e., patients provided with scTCR-sequencing data composed of both alpha and beta CDR3 chain information. TCR clustering using TCRpcDist was performed as described (Perez, M. A. S. et al. TCRpcDist: bioRxiv 2023.06.15.545077 (2023)). TCRpcDist is a novel and fast structure-based approach that calculates similarities between TCRs using a metric related to the physicochemical properties of solvent-exposed amino-acids of the most important residues of this receptor.
[0190] The TIL repertoire of patient 14 was analyzed and filtered through the three predictors, TRTpred, high structural avidity predictor (Schmidt, J. etal. Cell Reports Med. 2, 100194 (2021))) and TCRpcDist. To combine the three axes, the low structural avidity predicted clones were first filtered out, and the resulting clones were ranked according to their tumor reactivity score. Finally, TCR clustering was applied on the subset of top 20 tumor-reactive clones. The distance matrix obtained through TCRpcDist went through hierarchical clustering (agglomerative method: unweighted pair group method with arithmetic mean) leading to a dendrogram. The latter was split in five distinct TCR clusters given the downstream in vitro and in vivo validation. The selection of five clusters was arbitrary and can be adapted depending on the clinical context.
[0191] TCR transduction for in vivo experiment
[0192] TCRs n° 1, 3, and 5 were selected among the five in vitro-v Xidate tumor-reactive TCRs from patient 14 (Table 6) to be tested in vivo. TCRa and TCR0 chains (both including a mouse constant region), divided by a Furin / GS linker / T2A, were cloned into a pCRRL-pGK lentiviral plasmid to produce high-titer replication-defective lentiviral particles, as described (Giordano- Attianese, G. etal. Nat. Biotechnol. 38, 426-432 (2020)), with minor modifications. Briefly, 293T cells were plated at 10xl06in T-150 tissue culture flask and transfected the following day with 7 pg pVSVG (VSV glycoprotein expression plasmid), 18 pg ofR874 (Rev and Gag / Pol expression plasmid) and 15 pg of pCRRL transgene plasmid using a mix of 180 pl of Turbofect (Thermo Fisher) and 3 ml of Optimem media (Life Technologies). After overnight culture at 37 °C, 1 mM Sodium Butyrate (Sigma) and 4 mM caffeine (Sigma) were added to the media and cells incubated at 33 °C until viral supernatant collection (48h post-transfection). Viral particles were concentrated with Lenti-X-concentrator (Takara) following the manufacturer’s instructions. For primary human T cell transduction, CD8+T cells were negatively selected with beads (Miltenyi Biotec) from PBMCs of a healthy donor, activated and transduced as previously reported (Schmidt, J. etal. Nat. Commun. 2023 141 14, 1-15 (2023); Giordano-Attianese, G. et al. Nat. Biotechnol. 38, 426-432 (2020)), with minor modifications. Briefly, CD8+T cells were incubated with lentiviral particles after 24h activation with anti-CD3 / CD28 beads (Thermo Fisher Scientific) in R8 medium supplemented with 50 lU ml-1IL-2. After incubation at 37 °C for 72h, beads were removed, and transduced cells stained with an APC-conjugated anti-mouse constant beta antibody (eBioscience) followed by sorting with anti APC microbeads (Mitenyi Biotec). Sorted TCR-transduced CD8+T cells were then expanded for 10 days in R8 medium and 50 lU ml-1of IL-2 before mouse injection. Tumor-reactivity of transduced cells was assessed in IFNy ELISpot assays using pre-coated 96- well ELISpot plates (Mabtech). Transduced cells were plated at 5x103cells per well and challenged with IFNy-treated autologous tumor cells at a 1 : 1 ratio. After 18-20 hours of incubation, cells were removed, the plate developed according to the manufacturer’s instructions and counted using a Bioreader 6000-E (BioSys).
[0193] Adoptive T cell transfer in immunodeficient IL-2 NOG mice
[0194] IL2-NOG mice were shaved, anesthetized with isoflurane and subcutaneously injected with 106autologous human melanoma tumor cells from patient 14 (grown in RPMI 1640 GlutaMAX (Gibco) medium supplemented with 10% fetal bovine serum (FBS, Gibco), 100 lU / pl of penicillin, 100 pg / ml of streptomycin (Bio-Concept)). Once tumors became palpable (day 11), l *106or 5*106human tumor specific CD8+T cell clones, diluted in 100 pl of PBS, were injected intravenously in the tail vein, according to the treatment arms, with 4-6 mice per condition except some cases where 3 mice were considered. Tumor volumes were measured by caliper twice a week and calculated as follows: volume = length x width x width / 2. Mice were sacrificed by pentobarbital injection (150 mg / kg) before the tumor volume exceeded 103mm3or when necrotic skin lesions were observed at the tumor site.
[0195] Data analyses and computation
[0196] Data analyses were performed using R Statistical Software (V.4.0.3). All data processing and analysis was performed using the R dplyr and base libraries. The nested and simple cross- validation were performed using an in-house R library developed to control the models and hyperparameters throughout the folds. The R library glmnet was used to build the LR models and their specifications. Parallelization of the computation was allowed using the foreach library. The differential expression analysis methods were computed using the appropriate R libraries (Seurat, Umma, edgeR, DESeq2) as well as the signature score methods (AUCell, UCell, singscore, GSEABase). Statistical analyses were performed using the standard stats library. The statistical tests used, and their specifications are described in the figure legends. Parametric tests, for comparing two or more groups, were applied only on normally distributed variables validated with Anderson-Darling, D’Agostino-Pearson omnibus, Shapiro-Wilk and Kolmogorov-Smirnov tests (GraphPad 9.1.0), otherwise, non-parametric tests were used.
[0197] Plotting description
[0198] The figures were generated in R Statistical Software (V.4.0.3) with the ggplot2 R package. Alluvial plots were generated using the ggalluvial R package. The distance heatmaps were performed using the pheatmap function from the pheatmap R package. Plotting of scRNA-seq derived UMAP was achieved using Seurat R package functions. Schematic figures were created with BioRender.com. All figures were reprocessed using Adobe Illustrator (Al) 2020 solely for esthetical purposes.
[0199] EXAMPLE 2
[0200] This example describes TRTpred, an antigen -agnostic in silica TRT predictor developed and extensively benchmarked within a machine learning framework. Its superiority was demonstrated compared to existing predictive TRT signatures. TRTpred was also applied to successfully explore the immune repertoire of tumor-reactive TILs across various tumor indications and microenvironments. Finally, by integrating a high-avidity TCR predictor and a TCR clustering algorithm (TCRpcDist), a comprehensive algorithm (referred to as MixTRTpred) was engineered for the selection of clinically relevant TCRs from the pool of TRTs, which was subsequently validated in vitro and in vivo (Fig. 3).
[0201] To build TRTpred, 235 CD8+clonotypes, annotated as tumor-reactive (n=\ 12) or non- tumor-reactive («=123), curated from 10 metastatic melanoma patients were used (Fig. 1A and Table 1). Leveraging both their scTCR- and scRNA-seq profiles, a suite of 21 binary classifiers including logistic regression (LR) and signature scoring methods were trained, which underwent thorough fine-tuning (Fig. IB). The models were evaluated using the leave-one-patient-out (LOPO) nested cross-validation framework, providing insights into their generalization performance when faced with unseen data from new patients. While LR models exhibited commendable area-under-the-curve (AUC) scores, one of the signature scoring models (edgeR- QFL) emerged as the most generalizable, leading to TRTpred after training on the whole dataset (Fig. 1C and Table 2). Y-randomization tests further validated TRTpred’s credibility against spurious learning. As an illustration of TRTpred, for clonotypes from the dataset with known specificity, a clear discrimination between virus -specific bystander TILs and TILs specific for tumor-associated antigens (TAAs) or tumor-restricted neoantigens were observed (Fig. 1C and Table 3). TRTpred’s signature is composed of several genes associated with exhaustion (e.g., CXCL13, LAG3, TOX, PDCD1, or TNFRSF9, Fig. ID and Table 4).
[0202] To benchmark TRTpred, its performance was evaluated using two independent datasets, sourced internally or externally by Oliveira et al. (Oliveira, G. et al. Phenotype, specificity and avidity of anti-tumor CD8+ T cells in melanoma. Nat. 2021 1-7 (2021)), containing 90 and 205 CD8+TRT annotated clonotypes, respectively (Fig. 1E-F and Table 1). Also, TRTpred was compared with predictive signatures from recent studies (Lowery, F. J. etal. Science 375, eabl5447 (2022); Zheng, C. et al. Cancer Cell 40, 410-423. e7 (2022); Hanada, K.-I. etal. Cancer Cell. 2022 May 9;40(5):479-493.e6; Oliveira, G. etal. Nat. 2021 1-7 (2021); van derLeun, A. M., et al. Nat. Rev. Cancer 2020 204 20, 218-232 (2020)). TRTpred exhibited consistent performance across both sets of benchmarking data (with an accuracy of 83% and 87.8% for the internal and external datasets, respectively), and outperformed previous models (Fig. 1G-H). Moreover, TRTpred successfully identified 52 from the 63 neoantigen-specific clonotypes with undisclosed / unknown tumor reactivity status in publicly available datasets from 24 patients with diverse cancer types (Fig. II and Table 1).
[0203] Next, TRTpred was applied to interrogate the tumor repertoires from 42 patients with melanoma (w=19) as well as gastrointestinal (GI, n=17), lung ( / ? 4) or breast ( / ? 2) cancer (Table 1). Consistently across the different tumor types, inferred TRT repertoires were richer and more clonal than cognate bystander counterparts (Fig. 2A and Table 5). Also, in line with the expectations based on the clinical efficacy of TIL therapy reported in melanoma versus other solid tumors, a higher proportion of CD8+TRTs was identified in melanoma relative to other solid tumors (Fig. 2B and Table 5).
[0204] Tumor-specific TIL clonotypes, especially those with high avidity TCRs, accumulate preferentially within the tumor islets, i.e., the tumor cell clusters circumscribed by surrounding stroma, while these clonotypes are largely diluted in the surrounding stroma. To further challenge the prediction tools, the spatial distribution of predicted TRTs was examined. Taking advantage of TCR repertoires obtained from microdissected tumor core and stroma areas from five melanoma patients with scRNA / TCRseq data (Fig. 2C and Table 1), a predictor of TCR structural avidity was first applied, and the higher frequency of clones inferred to have high TCR avidity within the tumor islet compartment was validated (Fig.2D). Furthermore, by applying TRTpred, it was found that TRTs were more frequent in tumor islets, while bystander TIL clones accumulated preferentially in the stroma (Fig. 2E). This analysis was corroborated by analyzing neoantigen / TAA- or virus-specific T cells in a representative patient and in cumulative data (Table 3), validating the preferential infiltration of TRTs within tumor cores.
[0205] Besides exploratory applications, the ability to accurately distinguish TRTs represents a unique opportunity to identify private clinically relevant tumor-reactive TCRs for personalized TCR T-cell therapy. For this purpose, two key qualitative features of TRTs must be considered. First, all TRTs may not be equally relevant as they can be equipped with low avidity TCRs, even if they target neoantigens. Furthermore, in the perspective of personalized TCR-engineered T-cell therapy using multiple TCRs, it is key to generate cell products targeting multiple distinct antigens to limit tumor escape.
[0206] In summary, this example highlights the significance of a robust model selection approach using nested-cross-validation to evaluate the model while fine-tuning the hyperparameters. This has never been done in the single-cell RNA-sequencing (scRNA-seq) field and allowed a finding that a more complex model (LR) is not necessarily better for generalization purposes. Also, TRTpred allows a fair comparison between the training and testing scores, and the score was scaled based on the mean and standard deviation obtained from the training data. A threshold was then identified to stratify cells into tumor-specific and non-specific cells maximizing accuracy. The score thresholds were identified in the training data and applied on the testing data.
[0207] One of the best performing differential expression analysis methods is edgeR applied with the quasi-likelihood approach (edge-QFL). Since this method was originally developed for bulk RNA-sequencing data, edge-R QFL is considered a “pseudo-bulk” method. To ensure the performance of these methods, they were applied on the clone-average loglO-normalized UMI counts.
[0208] Other parameters such as how to select the genes from the differential expression analysis method (e.g., by considering only the log-Fold-Change or the P-value), the number of genes in the signature (e.g., signature length), whether to take only the up-regulated genes or both up and down regulated genes (e.g., signature side) and the signature score method (e.g., Average, AUCell, Singscore, and UCell) were defined as hyperparameters to fine-tune. It was found that the best hyperparameters were selecting the genes with the P-value, selecting 90 up and down regulated genes and scoring the cells with a simple average scoring method.
[0209] EXAMPLE 3
[0210] The proposed methodology, named MixTRTpred, hinges on (i) applying TRTpred to generate a ranked list of tumor-reactive clones, and (ii) filtering out clones with inferred low structural avidity TCRs (Fig. 1A). A TCR clustering tool, i.e., TCRpcDist, was subsequently applied to group TCRs with similar physicochemical properties, selecting the top tumor-reactive scoring TCR in each TCR-cluster to maximize the diversity of targeted antigens. As a result, an optimized list of inferred tumor-reactive TCRs with high structural avidity targeting distinct antigens was obtained.
[0211] To validate MixTRTpred’s efficacy, it was used for a patient where autologous patient- derived xenograft (PDX) tumors were available. From a total of 188 inferred tumor-reactive and high -avidity clones, the top-ranking clonotypes (17=20) were selected, and TCRpcDist was used to generate five distinct TCR clusters (Fig. 2F). By further selecting the top tumor-reactive TCRs from each cluster, five highly selected TCR candidates were obtained (Table 6). Consistently with TRTpred’ s overall accuracy, all TCRs (5 / 5) demonstrated tumor reactivity in vitro (Fig. 2G). Encouraged by these results, three of these TCRs (spanning, agnostically, the entire range of in situ frequency) were further selected to investigate their potential in controlling the in vivo growth of autologous PDX tumors. Transfer of 5x106T cells engineered with each individual TCR controlled tumor growth (Fig. 2H), and the two TCRs with the highest tumor reactivity scores and highest functional avidity eradicated tumors (Fig. 2H and Table 6). To test the advantage of transferring multiple TCRs, a suboptimal T-cell product (1x106cells per ACT dose) was used, and the efficacy of cells transduced with only one of the individual TCRs (1-3) was compared to a cocktail of all three TCRs. Multivalent ACT products yielded better tumor control than products containing only one type of TCR-transduced cells (Fig. 21). Finally, applying MixTRTpred to all patients with assessable data («=37, Table 1), more than 10 clinically relevant clones per patient (median=102) were consistently identified, except for one patient with only few sequenced clones, thus demonstrating the applicability of MixTRTpred for TCR T-cell therapy.
[0212] This example describes TRTpred, a rigorously benchmarked predictor of tumor-reactive clonotypes outcompeting other existing tools. TRTpred enabled a granular interrogation of TIL repertoire and spatial distribution and revealed a large range of richness and clonality of TRT repertoires in tumors. These metrics, reflecting the abundance of tumor-reactive cells in tumors, are useful in predicting responses to checkpoint immunotherapies. Furthermore, the results indicate that the abundance of TRT in tumors predicts clinical responses to TIL ACT in melanoma, offering further opportunities for better patient selection.
[0213] A minimum of 10 distinct TCRs per patient was found using MixTRTpred, a novel and stringent approach for the selection of clinically relevant TCRs, integrating TRTpred with a high- avidity TCR predictor and a TCR clustering algorithm (TCRpcDist). The results described herein support the clinical relevance of inferred TCRs for adoptive immunotherapy, and indicate that TRTpred is instrumental to either select patients who may benefit from TIL ACT or to select clinically relevant TCRs (using MixTRTpred) for TCR T-cell therapy in remaining patients who would not be eligible for TIL therapy. The accuracy of MixTRTpred indicates that personalized TCR-based therapy can be achieved for many patients with solid tumors.
[0214] The present disclosure is not to be limited in scope by the specific embodiments described herein. Indeed, various modifications of the invention, in addition to those described herein, will become apparent to those skilled in the art from the foregoing description and the accompanying figures. Such modifications are intended to fall within the scope of the appended claims.
[0215] Table 1: Cohorts’
[0216]
[0217]
[0218] Table 2: Model selection results
[0219]
[0220] *Description of the models can be found in Example 1
[0221] Table 3: Neoantigen, TAA, and Viral specificities
[0222] Table 4: TRTpred’ signature
[0223] Table 5: T-cell repertoire’s metrics for the 42 patients
[0224]
[0225]
Claims
CLAIMSWhat is claimed is:
1. A method of identifying a classifier for predicting tumor reactivity of T cells, comprising: determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells, wherein the transcriptomic profile comprises transcription level of each of a set of genes; inputting the transcriptomic profile of the set of genes into a plurality of test classifiers; determining, by each of the test classifiers, a result comprising tumor reactivity of each of the T cells; determining a classification score of each of the test classifiers by a nested cross-validation framework based on the result determined by each of the test classifiers; and ranking classification scores of the test classifiers and selecting, from the test classifiers, a candidate classifier having the classification score higher than a reference classification performance value.
2. A method of predicting tumor reactivity of T cells, comprising: determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data of T cells; wherein the transcriptomic profile comprises transcription level of each of a set of genes; inputting the transcriptomic profile of the set of genes into the classifier identified according to the method of claim 1 ; determining tumor reactivity of each of the T cells by the classifier; and ranking tumor reactivity of the T cells and selecting, from the T cells, candidate T cells that are reactive to a tumor antigen.
3. The method of any one of claims 1-2, wherein the nested cross-validation framework is implemented using a conventional or a leave-one-patient-out (LOPO) approach.
4. The method of claim 3, further comprising finetuning hyperparameters of the candidate classifier implemented using a conventional or a leave-one-patient-out (LOPO) approach.
5. The method of any one of the preceding claims, wherein the test classifiers comprise a logistic regression model.
6. The method of any one of claims 1-4, wherein the test classifiers comprise a signature scoring model.
7. The method of claim 6, wherein the candidate classifier comprises a signature scoring model implemented using edgeR-QFL.
8. The method of any one of the preceding claims, wherein the result comprises a tumor reactivity score of each of the T cells.
9. The method of claim 8, wherein tumor reactivity scores are scaled based on a mean and standard deviation obtained from training data.
10. The method of claim 4, wherein the hyperparameters comprise one or more of(i) the choice of the log-Fold-Change or P-value for selecting differential expressed genes; (ii) number of genes in the set of the genes; (iii) whether to include only up-regulated genes or both up- regulated and down-regulated genes in the set of the genes; and (iv) a scoring method to rank tumor reactivity of the T cells.
11. The method of claim 9, wherein the scoring method comprises an Average scoring method, an AUCell scoring method, a Singscore scoring method, or a Ucell scoring method.
12. The method of claim 10, wherein the hyperparameters comprise one or more of: using P- value for differential expression analysis, selecting 5 to 500 up-regulated and down-regulated genes, and scoring the tumor reactivity of the T cells with the Average scoring method.
13. The method of any one of the preceding claims, wherein the classification scores are characterized by area-under-the-curve (AUC) scores.
14. The method of any one of claims 1-12, wherein the classification scores are characterized by Matthew’s Correlation Coefficient (MCC).
15. The method of any one of the preceding claims, wherein the candidate classifier has the highest classification score among the test classifiers.
16. The method of any one of the preceding claims, wherein the tumor reactivity of the T cells comprises reactivity to a tumor antigen.
17. The method of any one of the preceding claims, wherein the set of genes comprise one or more genes set forth in Table 4.
18. The method of any one of the preceding claims, wherein the set of genes comprise CXCL13, LAG3, TOX, PDCD1, TNFRSF9, or a combination thereof.
19. The method of any one of the preceding claims, further comprising validating the candidate classifier against spurious learning.
20. The method of any one of the preceding claims, wherein the sequence data further comprises single cell TCR sequencing data.
21. The method of any one of the preceding claims, wherein the T cells are from one or more patients.
22. The method of any one of the preceding claims, wherein the candidate classifier is capable of excluding virus-specific T cells.
23. A method of identifying a tumor-reactive T-cell receptor (TCR), comprising:(a) identifying tumor-reactive T cells by: determining a transcriptomic profile of each of the T cells based on the sequence data comprising single cell RNA sequencing data and single cell TCR sequencing data of Tcells, wherein the transcriptomic profile comprises transcription level of each of a set of genes, inputting the transcriptomic profile of the set of genes into the classifier identified according to the method of any one of claims 1 -22, determining tumor reactivity of each of the T cells by the classifier, and ranking tumor reactivity of the T cells and selecting, from the T cells, candidate T cells that are reactive to a tumor antigen; and(b) identifying a tumor-reactive TCR from T-cell receptors (TCRs) contained the candidate T cells by performing at least one of:(i) identifying candidate TCRs having high structural avidity from the TCRs of the candidate T cells by: selecting a set of amino acids in an amino acid sequence of a CDR30 of each of the TCRs, determining solvent accessibility of each amino acid of the set of amino acids, inputting the solvent accessibility of each amino acid of the set of amino acids into a logistic regression model, wherein the logistic regression model performs an aggregated analysis based on the solvent accessibility of each amino acid of the set of amino acids and a weight value assigned to each amino acid of the set of amino acids, determining a probability value of a high avidity status of each of the TCRs as an output of the logistic regression model, and identifying candidate TCRs as having high avidity against the antigen if probability values of the candidate TCRs are greater than or equal to a threshold value; and(ii) clustering the candidate TCRs based on physicochemical properties thereof into one or more TCR clusters that target distinct antigens, and selecting a top tumor-reactive TCR from each of the one or more TCR clusters.
24. The method of claim 23, wherein the top tumor-reactive TCR is the TCR that is most reactive to the tumor antigen in each of the one or more TCR clusters.
25. The method of any one of claim 23-24, wherein the set of amino acids comprises one or more of Arg, Asn, Asp, Gly, He, Lue, and Phe.
26. The method of any one of claims 23-25, comprising determining the probability value by:wherein p is the probability value, bo represents a bias term, and R, N, D, G, I, L, and F are amino acids Arg, Asn, Asp, Gly, He, Lue, and Phe, respectively.
27. The method of any one of claims 23-26, wherein the threshold value is about 0.5.
28. The method of any one of claims 23-27, comprising selecting a subset of amino acids from the set of amino acids that have solvent accessibility higher than about 30%.
29. The method of claim 28, comprising inputting the solvent accessibility of the subset of amino acids into the logistic regression model.
30. The method of any one of claims 23-29, comprising determining the solvent accessibility of each amino acid of the set of amino acids as a relative solvent excluded surface area (SESA).
31. The method of claim 30, wherein the SESA is determined by normalizing surface area of an amino acid in the TCR against surface area of the amino acid in a reference state.
32. The method of any one of claims 23-31, comprising determining solvent accessibility of an amino acid based on a three-dimensional model of the TCR.
33. The method of any one of claims 23-32, wherein the TCR with high avidity has a pMHC- TCR half-life of less than about 60 seconds.
34. The method of any one of claims 23-33, comprising inputting into the logistic regression model one or more additional characteristics of each amino acid of the set of amino acids.
35. The method of claim 34, wherein the additional characteristics comprise hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of P-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of P-turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in P-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, a-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in a-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of lefthanded a-helix, heat capacity, free energy in a-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of P-sheet, unweighted, normalized frequency of P-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, percentage of buried residues, or a combination thereof.
36. The method of claim 35, wherein the additional characteristics comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, electrostatic charge, or a combination thereof.
37. The method of any one of claims 23-36, wherein the step of clustering the candidate TCRs comprises:(a) selecting a subset of amino acids in a first TCR and a corresponding subset of amino acids in a second TCR, wherein the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR have an identical number of amino acids;(b) determining an amino acid sum of differences in each of a plurality of physicochemical properties by performing a pairwise comparison between an amino acid in the subset of amino acids in the first TCR and a corresponding amino acid in the corresponding subset of amino acids in the second TCR;(c) repeating steps (a) to (b) for remaining amino acids in the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR;(d) determining a subset sum of differences between the subset of amino acids in the first TCR and the corresponding subset of amino acids in the second TCR;(e) repeating steps (a) to (d) for one or more subsets of amino acids in the first TCR and the second TCR;(f) determining an aggregate value of all subset sums of differences between the first TCR and the second TCR by assigning a weight value to each of the subset sums; and(g) identifying the first TCR and the second TCR as having similar specificity to an epitope if the aggregate value is smaller than a threshold value.
38. The method of any one of claims 23-36, wherein the step of clustering the candidate TCRs comprises:(1) generating a similarity matrix for a subset of amino acids of a TCR of the TCRs, wherein the similarity matrix comprises a plurality of physicochemical properties of each amino acid in the subset of amino acids;(2) repeating step (1) for one or more subsets of amino acids of the TCR;(3) repeating steps (1) to (2) for remaining TCRs of the plurality of immunological entities; and(4) performing a clustering analysis based on a distance between two corresponding similarity matrices of a pair of immunological entities to identify a subset of TCRs having similar specificity to the epitope.
39. The method of claim 38, wherein the distance is a Manhattan distance.
40. The method of claim 38, wherein the clustering analysis comprises a hierarchical clustering.
41. The method of claim 40, wherein the hierarchical clustering comprises an unweighted pair group method with arithmetic mean (UPGMA).
142. The method of any one of claims 37-41, wherein the epitope is located on peptide-MHC (pMHC).
43. The method of any one of claims 37-42, wherein the subset of amino acids comprises 3 to 8 amino acids.
44. The method of claim 43, wherein the subset of amino acids comprises 3 to 8 consecutive amino acids.
45. The method of any one of claims 43-44, wherein the subset of amino acids consists of 4 amino acids.
46. The method of any one of claims 37-45, wherein the one or more subsets of amino acids are selected from amino acids in CDRla, CDR2a, CDR3a, CDRip, CDR2P, and CDR3p.
47. The method of claim 46, wherein the step of determining the aggregate value comprises an aggregate value comprises assigning a weight value of about 30% to the subset of amino acids in CDR3a or CDR3p.
48. The method of claim 46, wherein the step of determining the aggregate value comprises determining an aggregate value comprises assigning a weight value of about 10% to the subset of amino acids in CDRla, CDR2a, CDRip, or CDR2p.
49. The method of any one of claims 37-48, wherein the subset of amino acids does not include amino acids that are not solvent-exposed.
50. The method of claim 49, wherein the subset of amino acids does not include amino acids in CDRla, CDR2a, CDRip, or CDR2P that have a relative solvent excluded surface area (SESA) of less than about 5%.
51. The method of claim 49, wherein the subset of amino acids does not include amino acids in CDR3a or CDR3P that have a SESA of less than about 20%.
52. The method of any one of claims 37-51, wherein the physicochemical properties comprise amino acid attributes selected from hydrophilicity value, polar requirement, long range nonbonded energy per atom, negative charge, positive charge, size, normalized relative frequency of bend, normalized frequency of P-turn, molecular weight, relative mutability, normalized frequency of coil, average volume of buried residue, conformational parameter of P- turn, residue volume, isoelectric point, optimized propensity to form reverse turn, chou-fasman parameter of coil conformation, information measure for loop, free energy in P-strand region, side chain volume, amino acid composition of total proteins, average relative probability of helix, a-helix indices, relative frequency of occurrence, helix-coil equilibrium constant, amino acid composition, number of codon(s), net charge, normalized frequency of turn, relative frequency in a-helix, average nonbonded energy per residue, bulkiness, normalized relative frequency of coil, refractivity, normalized frequency of left-handed a-helix, heat capacity, free energy in a-helical region, hydrophobicity factor, normalized frequency of extended structure, normalized frequency of P-sheet, unweighted, normalized frequency of P-sheet, information measure for pleated-sheet, hydropathy index, eisenberg hydrophobic index, average side chain orientation angle, average interactions per side chain atom, transfer free energy, and percentage of buried residues.
53. The method of claim 52, wherein the physicochemical properties comprise hydrophobicity, secondary structure propensity, size / mass, amino acid composition, codon degeneracy, and electrostatic charge.