Methods and systems for analysis of receptor interactions

The computational framework addresses the low signal-to-noise issue in TCR-pMHC binding data by preprocessing and normalization, enabling accurate prediction of TCR-pMHC binding through deep learning, thus enhancing the reliability of TCR-pMHC binding predictions.

JP7827757B2Active Publication Date: 2026-03-10REGENERON PHARMACEUTICALS INC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing high-throughput TCR-pMHC binding data suffer from low signal-to-noise ratios and suboptimal prediction accuracy due to limitations in computational methods, particularly in identifying complex sequence patterns from full-length TCR sequences, leading to challenges in reliably distinguishing true binding events from nonspecific background.

Method used

A computational framework that includes data preprocessing and normalization steps to filter low-quality cells, adjust for background noise, and perform cell-wise and pMHC-wise normalization, followed by a predictive model training using deep learning algorithms to identify reliable TCR-pMHC binding events.

Benefits of technology

Enhances the accuracy of predicting TCR-pMHC binding specificity by effectively distinguishing true binding events from noise, improving the reliability and precision of TCR-pMHC binding predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827757000023
    Figure 0007827757000023
  • Figure 0007827757000024
    Figure 0007827757000024
  • Figure 0007827757000025
    Figure 0007827757000025
Patent Text Reader

Abstract

To provide a computational framework for high-throughput mapping, validating, and predicting receptor sequence interactions.SOLUTION: A method performed by a computer comprises: receiving single cell sequencing data comprising single cell sequence data, dextramer sequence data, and single cell T-Cell Receptor (TCR) sequence data; filtering, from the dextramer sequence data, based on the single cell sequence data, data associated with low-quality cells; adjusting, based on a measure of background noise, the dextramer sequence data; filtering, from the dextramer sequence data, based on the single cell TCR- data, data according to a presence or an absence of an α-chain or a β-chain; and identifying data remaining in the normalized filtered dextramer sequence data as associated with reliable TCR-pMHC binding events.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 013,480, filed April 21, 2020, U.S. Provisional Patent Application No. 63 / 090,498, filed October 12, 2020, and U.S. Provisional Patent Application No. 63 / 111,395, filed November 9, 2020. The contents of these earlier applications are incorporated herein by reference in their entireties. [Background technology]

[0002] T cell antigen specificity, mediated through the T cell receptor (TCR), is a hallmark of cellular immunity. TCRs are heterodimeric proteins present on the surface of T cells and generally consist of an α chain and a β chain. The TCR α and β chain genes are composed of separate V, D (β chain only), and J segments that are combined by somatic recombination during T cell development. This gene rearrangement generates a highly diverse TCR repertoire (estimated at 1015–1061 potential in humans) to ensure efficient control of viral infections and other pathogen-induced diseases. TCR diversity is primarily expressed in the complementarity-determining region (CDR) loops (CDR1, CDR2, and CDR3), which bind peptides presented by major histocompatibility complex (MHC) proteins and thus directly determine the specificity of T cell pMHC binding.

[0003] Although the factors underlying TCR-pMHC recognition are not fully understood, recent studies have shown that T cells that bind to specific pMHC share common TCR sequence characteristics. When selected, it is possible to predict the specific binding probability of unseen TCR sequences based on the learned TCR sequence characteristics. However, these studies were limited by the volume and diversity of training data generated by traditional single multimer sorting assays or antigen rechallenge assays. Further understanding of TCR-pMHC specific binding requires innovations in both computational and experimental methods. 10xGenomics recently published a dataset from a highly multiplexed pooled dextramer binding immune profiling platform that combines feature-barcoded dextramers with single-cell TCR sequencing. While this approach allows for the generation of high-dimensional pMHC-specific binding data at the single-cell level using paired T cell α and β chain sequences, other large-scale pooled multimer approaches can only estimate the composition of pMHC-specific binding T cells.

[0004] Like other high-throughput technologies, highly multiplexed Dexter binding data are often associated with low signal-to-noise ratios. This makes it bioinformatically challenging to reliably identify TCR-pMHC binding events using such large binding datasets. We observed unexpectedly high inter-HLA and inter-pMHC associations from the binding events provided by 10x Genomics (Figure 11A). This low signal-to-noise dataset requires more sophisticated computational normalization methods to distinguish true TCR-pMHC binding events from nonspecific background.

[0005] Next-generation screening technologies have increased the amount of available TCR-pMHC binding data, making state-of-the-art functional classifiers more feasible for computationally validating and subsequently predicting TCR-pMHC-specific recognition. While early TCR-pMHC binding classifier results have been encouraging, they were limited to CDR loop sequences and therefore unable to learn global complex sequence patterns from full-length TCR sequences, resulting in suboptimal prediction accuracy for highly diverse pMHC-binding TCRs. Taking advantage of the ability of deep learning algorithms to learn complex patterns, several deep learning frameworks have recently been proposed to uncover binding patterns in large, highly complex TCR sequence datasets.

[0006] In this study, we describe a computational framework to map, computationally validate, and predict TCR-pMHC-specific recognition using highly multiplexed dextramer binding data. Summary of the Invention

[0007] Disclosed is a method that includes receiving single cell sequencing data including single cell sequence data, dextramer sequence data, and single cell T cell receptor (TCR) sequence data; filtering from the dextramer sequence data data associated with low quality cells based on the single cell sequence data; adjusting the dextramer sequence data based on a measurement of background noise; filtering from the dextramer sequence data data according to the presence or absence of an alpha chain or a beta chain based on the single cell TCR data; and identifying data remaining in the normalized filtered dextramer sequence data that is associated with reliable TCR-pMHC binding events.

[0008] receiving single cell sequence data, dextramer sequence data, and T cell receptor (TCR) sequence data of a single cell; determining a number of genes based on the single cell sequence data for each cell represented in the dextramer sequence data; removing data associated with cells whose number of genes is outside a gene threshold range from the dextramer sequence data; determining a fraction of mitochondrial gene expression based on the single cell sequence data for each cell represented in the dextramer sequence data; removing data associated with cells whose fraction of mitochondrial gene expression exceeds a gene expression threshold from the dextramer sequence data; determining selected dextramer sequence data based on the dextramer sequence data, wherein the selected dextramer sequence data includes selected test dextramer sequence data, negative control dextramer sequence data, and unselected dextramer sequence data, and the unselected dextramer sequence data includes unselected test dextramer sequence data; determining, for each cell, a maximum negative control dextramer signal based on the negative control dextramer sequence data; determining, for each cell represented in the dextramer sequence data, a maximum selected dextramer signal based on the selected test dextramer sequence data; determining, for each cell represented in the dextramer sequence data, a maximum unselected dextramer signal based on the unselected test dextramer sequence data; estimating a dextramer binding background noise based on the maximum negative control dextramer signal; estimating a dextramer sorting gate efficiency based on the maximum selected dextramer signal and the maximum unselected dextramer signal; determining a background noise measurement based on the dextramer binding background noise and the dextramer sorting gate efficiency; subtracting, for each cell represented in the dextramer sequence data, the background noise measurement from the dextramer signal associated with each cell;Disclosed is a method that includes, for each cell represented in the dextramer sequence data, performing cell-wise normalization on the dextramer signal associated with each cell; for each cell represented in the dextramer sequence data, performing pMHC-wise normalization; for each cell represented in the dextramer sequence data, determining the presence or absence of at least one alpha chain and at least one beta chain based on the TCR sequence data of a single cell; removing from the normalized dextramer sequence data data data associated with cells having only an alpha chain, only a beta chain, or multiple alpha or beta chains based on the presence or absence of at least one alpha chain and at least one beta chain; and identifying data remaining in the normalized dextramer sequence data as associated with reliable TCR-pMHC binding events.

[0009] Disclosed is a method including: performing TCR-pMHC binding specificity data normalization in dextramer sequence data to identify a plurality of TCR-pMHC binding events; determining a training dataset comprising a plurality of TCR sequences based on the normalized dextramer sequence data, each TCR sequence associated with a binding affinity; determining a plurality of characteristics for a predictive model based on the plurality of TCR sequences; training a predictive model with the plurality of characteristics based on a first portion of the training dataset; testing the predictive model based on a second portion of the training dataset; and outputting the predictive model based on the testing.

[0010] Disclosed is a method that includes presenting an unknown TCR sequence to a trained predictive model, where the trained predictive model is trained based on a training dataset generated by the disclosed method; and predicting binding affinity using the trained predictive model.

[0011] Receiving single cell sequence data, dextramer sequence data, and T cell receptor (TCR) sequence data of a single cell; for each cell represented in the dextramer sequence data, determining a number of genes based on the single cell sequence data; removing data from the dextramer sequence data associated with cells whose number of genes is outside a gene threshold range; for each cell represented in the dextramer sequence data, determining a fraction of mitochondrial gene expression based on the single cell sequence data; removing data from the dextramer sequence data associated with cells whose fraction of mitochondrial gene expression exceeds a gene expression threshold; determining selected dextramer sequence data based on the dextramer sequence data; the selected dextramer sequence data includes selected test dextramer sequence data and negative control dextramer sequence data; for each cell represented in the dextramer sequence data, determining a maximum negative control dextramer signal based on the negative control dextramer sequence data; determining the maximum selected dextramer signal for each cell based on the selected test dextramer sequence data; estimating the dextramer binding background noise based on the maximum negative control dextramer signal and the maximum selected dextramer signal; determining the presence or absence of at least one α chain and at least one β chain for each cell represented in the dextramer sequence data based on the TCR sequence data of a single cell; removing data associated with cells having only an α chain, only a β chain, or multiple α or β chains from the dextramer sequence data based on the presence or absence of at least one α chain and at least one β chain; determining the ratio of the dextramer signal within the cell to the sum of all dextramers to the cell (a measure of dextramer binding specificity to the cell) for each dextramer that binds to a given TCR clonotype of each cell represented in the dextramer sequence data;Disclosed is a method that includes determining the fraction of T cells within a clone that bind to a particular dextramer (a measure of dextramer binding specificity for the clonal type to which the cell belongs); for each dextramer that binds to a given cell represented in the dextramer sequence data, determining a corrected dextramer signal associated with each dextramer that binds to the cell based on the measure of dextramer binding specificity for the cell and the measure of dextramer binding specificity for the clonal type to which the cell belongs; for each cell represented in the dextramer sequence data, performing cell-wise normalization on the dextramer signal associated with each cell; for each cell represented in the dextramer sequence data, performing pMHC-wise normalization; and identifying the data remaining in the normalized dextramer sequence data as associated with reliable TCR-pMHC binding events based on a threshold value.

[0012] Disclosed is an apparatus configured to perform any of the disclosed methods.

[0013] Disclosed is a computer readable medium having processor executable instruction embodiments configured to cause an apparatus to perform any of the disclosed methods.

[0014] Additional advantages of the disclosed method and compositions will be set forth in part in the description which follows, and in part will be understood from the description, or may be learned by practice of the disclosed method and compositions. The advantages of the disclosed method and compositions will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention as claimed. [Brief explanation of the drawings]

[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments of the disclosed methods and compositions and, together with the description, serve to explain the principles of the disclosed methods and compositions.

[0016] [Figure 1] FIG. 1 illustrates an exemplary operating environment.

[0017] [Figure 2] Figure 2 illustrates the experimental approach for generating multi-omic high-throughput TCR-pMHC binding data. PBMC T cells from healthy human donors were labeled for sorting on CD8+ T cells. Sorted CD8+ T cells were stained with a pool of 50 dCODE Dextramer antibodies. Dextramer-positive CD8+ T cells were sorted by flow cytometry and individually captured as input for 10x Genomics single-cell sequencing library preparation. Three libraries were generated for gene expression, cell surface protein / dCODE expression, and paired TCR sequences for each CD8+ T cell.

[0018] [Figure 3] FIG. 3 illustrates an exemplary method.

[0019] [Figure 4] FIG. 4 illustrates an exemplary method.

[0020] [Figure 5] FIG. 5 illustrates an exemplary method.

[0021] [Figure 6A]Figures 6A and 6B show an example of the ICON (Integrative Context-specific Normalization) workflow scheme. a. From top left to bottom left: I. Distribution of dCODE Dextramers raw expression in UMIs (unique molecular identifiers). Maximum dCODE Dextramers expression in UMIs for CD8+ cells from Dex_sorted (maximum UMI representing a test of Dextramers from Dextramers sorted CD8+ T cells), NC_dex (maximum UMI representing a test of negative control Dextramers from Dextramers sorted CD8+ T cells), and Dex_unsorted (maximum UMI representing a test of stained Dextramers rather than sorted control CD8+ cells). II. Filtering of low-quality cells based on single-cell RNA-seq. Each dot represents a T cell. Red dots represent unhealthy cells. III. Estimation of dextramer binding background noise (P99.9) and dextramer sorting gate efficiency (argmaxDs,u) based on dCODE dextramer expression data. III. Adjustment for background noise by subtracting Max(P99.9,argmaxDs,u). V. Cell- and pMHC-wise normalization of background-subtracted dextramer expression. VI. Selection of cells with a single paired TCR αβ chain. VII. Distribution of normalized dextramer expression. UMI*: Normalized UMI. For details, see Methods. b. TCR-pMHC binding specificity of expanded TCR clonotypes. Up to 50 TCR clones from donor 1 are plotted along with their binding specificity and match. Circles indicate that at least one member of a clonotype was classified as specific for a particular pMHC. Circle size indicates the total clonotype size within a donor. Circle color indicates the percentage of cells within a clonotype that bind dextramer ("binding match"). Left panel: 10x Genomics identified ~50 clonotypes using an exhaustive cutoff. Right panel: ~50 clonotypes from the pMHC repertoire containing ~50 clonotypes from 10x Genomics for donor 1. [Figure 6B] Same as above.

[0022] [Figure 7A] Figures 7A-7E show the pMHC binding landscape of the 10x Genomics Dextramers binding data. a. Network of identified pMHC-specific binding T cell repertoires. Each node represents a pMHC repertoire and a pie chart of the number of unique paired TCRs from each donor that bind to that pMHC. Donor 1 is gray, Donor 2 is red, and Donor 4 is yellow. Node size indicates the total number of T cells that bind to that pMHC. Each edge represents a unique TCR shared by two pMHCs. Edge thickness indicates the number of shared unique TCRs. b. The majority of identified binders interact with seven pMHCs. c. Venn diagram of unique paired binding TCRs identified from Donor 1, Donor 2, and Donor 3. d. Composition of unique paired TCR αβ chains.

[0001] According to TCRB, 1:1 means one unique TCR β chain paired with one unique TCR α chain; 1:2 and binding to the same pMHC means a unique pair of TCRs with a shared β chain, but different α chains recognize the same pMHC; 1:2 and binding to >=2 pMHC means a unique pair of TCRs with a shared β chain, but different α chains recognize different pMHC. According to TCRA, 1:1 means one unique TCR α chain paired with one unique TCR β chain; 1:2 and binding to the same pMHC means a unique pair of TCRs with a shared α chain, but different β chains recognize the same pMHC; 1:2 and binding to >=2 pMHC means a unique pair of TCRs with a shared α chain, but different β chains recognize different pMHC. e. TCR-pMHC binding specificity and TCR cross-HLA recognition. Left, pie chart of T cell binding to one pMHC or at least two pMHC. Right, pie chart of T cells: HLA type-matched binding, supertype-matched binding or cross-typed binding. [Figure 7B] Same as above. [Figure 7C] Same as above. [Figure 7D] Same as above. [Figure 7E] Same as above.

[0023] [Figure 8A] Figures 8A-8D show a convolutional neural network (CNN)-based classification of TCR-pMHC-bound TCRs. a. CNN-based TCR sequence classification framework. Left panel: V and J segments (from alpha and beta) were transformed into embedding vectors. A trainable embedding was used for the amino acids comprising the CDR3 alpha or beta sequences, and a one-dimensional CNN was applied to the embedding. All embeddings were then concatenated together and fed through a concatenated layer. A SoftMax layer was then used to output the sequence class probability. In the right panel, a toy example illustrates the input and output of the deep learning sequence classifier. For details, see the Methods section. b. ROC curve of the CNN-based classifier with binomial mode using 11 curated paired TCR pMHC-binding repertoires. Binders are unique TCRs that bound a specific pMHC, and non-binders are unique TCRs that bound the other 10 pMHCs. Paired α and β TCR sequences were used as input data. c. Comparison of classification power between CNN-based and distance-based binary classifiers with the same definitions of binders and non-binders as described in b. Paired α and β TCR sequences were used as input data (Methods). d. Correlation of pMHC repertoire diversity, measured by Shannon entropy, with predictive performance between CNN-based and distance-based classifiers. ΔAUC = CNN-based AUC - distance-based AUC. [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 8D] Same as above.

[0024] [Figure 9A]Figures 9A-4E show CNN-based classification of the top seven pMHC-binding repertoires identified from the 10x Genomics dataset. a. ROC curves of the CNN-based classifier in binomial mode using the seven pMHC-binding repertoires identified from the 10x Genomics high-throughput dataset. Binders are unique TCRs that bound to a particular pMHC, and non-binders are unique TCRs that bound to the other six pMHCs. Paired α and β TCR sequences were used as input data. b. ROC curves for prediction results of the CNN-based classifier using independent test datasets from VDJdb: A*02:01_GILGFVFTL_Flu-MP_Influenza, A*02:01_ELAGIGILTV_MART-1_Cancer, A*02:01_GLCTLVAML_BMLF1_EBV, and A*11:01_AVFDRKSDAK_EBNA-3B_EBV binding T cells and another set of MART-1 (REGN_A*02:01_ELAGIGILTV_MART-1_Cancer) binders from an independent in-house experiment (Methods). The module was trained with pMHC repertoires identified from 10x Genomics data for prediction. c. Classification performance comparison using TCRα only, TCRβ only, or paired TCRα and β chains as sequence input. d. Use of T cell V and J gene segments for T cells binding these seven pMHCs. Less than 5% of gene segments were combined and are shown in grey. e. CDR3 motifs of the 10 most predictable paired TCRs from the seven pMHC repertoires. [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 9D] Same as above. [Figure 9E] Same as above.

[0025] [Figure 10A]Figures 10A-10E show the immunophenotype of pMHC-binding CD8+ T cells. a. Classification of pMHC-binding cells. Clusters were visualized by UMAP, with cell types represented by different colors. b. Heatmap of gene or protein expression of cell-type marker genes to annotate CD8+ T cell subpopulations. pMHC-binding landscape by CT cell immune subtype. Bars indicate the number of pMHC-binding T cells on a log2 scale. d. Expanded clonotypes are enriched in the non-naive compartment. Each dot represents a unique TCR clone. e. Percentage of HLA-matched and mismatched binding among naive and non-naive binding T cells. Tpm: peripheral memory cells; Tcm: central memory cells; Tem: effector memory cells; Temra: highly differentiated effector memory cells; Other: other memory cells with marker expression CD43loKLRG1hiCD127. [Figure 10B] Same as above. [Figure 10C] Same as above. [Figure 10D] Same as above. [Figure 10E] Same as above.

[0026] [Figure 11A-1] Figures 11A-11B show the TCR-pMHC binding specificities of clonotypes expanded from binding events identified by 10xGenomics from each donor. Up to 50 clonotypes are plotted along with their binding specificities and concordances. a. Circles indicate that at least one member of a clonotype was classified as specific for a particular pMHC. Circle size indicates the total intradonor clonotype size. Circle color indicates the percentage of cells within a clonotype that bind dextramer ("binding concordance"). b. Scatter plot of cell sorting results from 10xGenomics Donors 3 and 4 (Methods) CD8+ T cells reassessing dextramer binding. [Figure 11A-2] Same as above. [Figure 11A-3] Same as above. [Figure 11A-4] Same as above. [Figure 11B] Same as above.

[0027] [Figure 12A] Figures 12A-12F show examples of background estimation and modulation of dextramer binding signals for 10x Genomics high-throughput data. Dex_sorted (maximum UMI testing dextramer from dextramer-sorted CD8+ T cells), NC_dex (maximum UMI testing negative control dextramer from dextramer-sorted CD8+ T cells), and Dex_not sorted (maximum UMI testing dextramer stained, not sorted control CD8+ cells). a. Scatter plot of the number of detected genes versus the percentage of mitochondrial gene expression using single-cell RNA data. Each dot represents a cell. Red dots are dead cells or doublets. b. Distribution of dextramer expression data before and after the ICON process. c and d. Estimation of dextramer sorting efficiency. c. Accumulated distribution of dextramer UMI. Each dot is a data point for a unique dextramer UMI. d. p-value distribution of the KS test (Dex_selected vs. not Dex_selected) using one dextramer UMI data point as a sliding window. The gray dashed line is the threshold for dextramer selection efficiency. e. Scatter plot of Dex_selected before (x-axis) and after (y-axis) background subtraction for each donor. f. E'e density distribution. E'e: logarithmic rank of each dextramer signal within the cell (Methods). The blue dashed line is for the threshold for pMHC-specific binding. [Figure 12B] Same as above. [Figure 12C] Same as above. [Figure 12D] Same as above. [Figure 12E] Same as above. [Figure 12F] Same as above.

[0028] [Figure 13A]Figures 13A-13C show the binding specificities of the expanded clonotypes identified by this study for three donors. Up to 50 T cell clones are plotted along with their binding specificities and matches. Circle size indicates T cell clone size. Circle color indicates the percentage of cells within the clone that bind dextramer, which is a binding match. [Figure 13B] Same as above. [Figure 13C] Same as above.

[0029] [Figure 14A] Figures 14A and 14B show ROC curves for distance-based classifiers using curated pMHC-binding repertoires. b. Shannon entropy scores for curated pMHC-binding repertoires. [Figure 14B] Same as above.

[0030] [Figure 15A]Figures 15A-15C show the characteristics of the top seven pMHC-binding T cell repertoires. a. Pie chart of the proportion of T cell binding matched, matched supertypes, and mismatched HLA types. b. Power law of unique T cell clone sizes for the top seven pMHC-binding repertoires. Regression smoothing was used for fitting. c. Simpsons diversity index and TCRB generation probability of TCR-pMHC repertoires. The R package vegan was used to calculate the Simpsons diversity index. The TCRB CDR3 amino acid sequence generation probability of each pMHC-specific binder was calculated using OLGA. The fraction of the repertoire (represented by red triangles) specific for each pMHC was then obtained as the sum of the generation probabilities for each of the corresponding CDR3 sequences, as described by Sethna et al. The results show that the net fraction of TCRs specific for these pMHCs is large (ranging from 10 to 10), in a sense defined by the inverse of the number of independent TCR recombination events (10), meaning that any individual is likely to have these binding T cells in their T repertoire. Each point on the TCR generation probability diagram represents a unique T cell clone, and the colored bars indicate T cell clone size. [Figure 15B] Same as above. [Figure 15C] Same as above.

[0031] [Figure 16A] Figures 16A-16C show classification of TCR-pMHC binding TCRs. a. Distance and distance distribution of pMHC binders and non-binders using α chain only, β chain only, and paired αβ chains. b. ROC curve for distance-based classifier using the top 7 pMHC binding repertoires identified from the 10x Genomics high-throughput dataset. Paired α and β TCR sequences were used as input data. c. Comparison of classification power of CNN-based and distance-based classifiers. [Figure 16B] Same as above. [Figure 16C] Same as above.

[0032] [Figure 17A] Figures 17A and 17B show the CDR3 motifs of four pMHC-binding repertoires from the VDJdb duplication and the top seven pMHC repertoires identified from the 10x Genomics high-throughput data. b. ROC curve for a multinomial CNN-based classifier using seven pMHC-binding repertoires identified from the 10x Genomics high-throughput dataset. Paired α and β TCR sequences were used as input data. [Figure 17B] Same as above.

[0033] [Figure 18A] Figures 18A and 18B show examples of clusters of pMHC-binding CD8+ cells using single cell RNA-seq data. a. By cluster number. b. Overlaid with donor information. [Figure 18B] Same as above.

[0034] [Figure 19] FIG. 19 is a table containing information about the T cell donors used in the disclosed studies.

[0035] [Figure 20] FIG. 20 lists the dCODE dextramer reagents used in the disclosed studies and the NetMHC peptide-HLA allele binding predictions.

[0036] [Figure 21] FIG. 21 is a table summarizing the pMHC-TCR binding phenomenon.

[0037] [Figure 22] Figure 22 shows TCR-pMHC repertoire diversity and peptide characteristics.

[0038] [Figure 23] Figure 23 shows an overview of 11 pMHC repertoires collated from the VDJdb and McPAS.

[0039] [Figure 24-1] Figure 24 shows expanded TCR clonotype pMHC specificity in binders identified by 10x Genomics. Up to 50 TCR cell clones from donors 1-4 are plotted along with their binding specificity and match. Circles indicate that at least one member of a clonotype was classified as specific for a particular pMHC. Circle size indicates the total clonotype size within a donor. Circle color indicates the percentage of cells within a clonotype that bind dextramer ("binding match"). [Figure 24-2] Same as above. [Figure 24-3] Same as above. [Figure 24-4] Same as above.

[0040] [Figure 25A]Figures 25A-G show the identification and characterization of pMHC-binding T cells from high-throughput pMHC binding data. (A) ICON (Integrated Context-Specific Normalization) workflow scheme. RT: fraction of T cells within a clone that bind a specific dextramer; RC: ratio of intracellular dextramer signal to the sum of all dextramers bound to the cell. (B) pMHC-binding landscape network of dextramer binders identified by ICON. Each node represents the pMHC repertoire and is presented as a pie chart of the number of unique paired TCRs from each donor that bind pMHC. Node size indicates the total number of unique TCRs that bind to a given pMHC. Each edge represents a unique TCR shared by two pMHCs. Edge thickness represents the number of shared unique TCRs. (C) Correlation of flow sorting results in ICON with estimated single dextramer binding compared to the abundance of pMHC-binding T cells. The number of dextramers used for validation was 21. (D) Uniqueness and overlap of identified pMHC-binding TCRs among donors 1, 2, 3, 4, and V. (E) The majority of identified binders interact with nine pMHCs. (F) V and J gene segment utilization for T cell binding to these nine pMHCs. Gene segments less than 5% combined are shown in gray. (G) HLA-restricted and unrestricted binding. [Figure 25B] Same as above. [Figure 25C] Same as above. [Figure 25D] Same as above. [Figure 25E] Same as above. [Figure 25F] Same as above. [Figure 25G] Same as above.

[0041] [Figure 26A]Figures 26A-D show processing of high-throughput data using ICON. (A) Scatter plot of the number of detected genes versus the percentage of mitochondrial gene expression using single-cell RNA data. Each dot represents a cell. Red dots are dead cells or doublets. (B) Distribution of dextramer signal in UMIs from negative control and test dextramers. Sorted_nc: negative control dextramer; sorted_dex: test dextramer. (C) Scatter plot of RT vs. RC. RC is the ratio of intracellular dextramer signal to the sum of all dextramers that bind to T cells. RT is the fraction of T cells within a clone that binds a specific dextramer. (D) Hierarchical clustering of pMHC-binding T cells identified by ICON. Each row represents a dextramer and each column represents a T cell. [Figure 26B] Same as above. [Figure 26C] Same as above. [Figure 26D] Same as above.

[0042] [Figure 27] FIG. 27 shows pooled Dextramer FACS gating for fluorescence activated sorting (FACS) of Dextramer+ T cells from donor V.

[0043] [Figure 28A] Figures 28A-B show single oligo-dextramer sorting. (A) Representative gating for fluorescence-activated sorting (FACS) of dextramer-positive T cells. T cells were previously enriched from donor V peripheral blood mononuclear cells (PBMCs) and then stained with single oligo-dextramers. The following sequential gating strategy was utilized to isolate the desired dextramer+ population for sorting. (B) Scatter plot of single oligo-dextramer cell sorting results for each of the 21 test dextramers and two negative control dextramers. [Figure 28B] Same as above.

[0044] [Figure 29] FIG. 29 is a table outlining pMHC-TCR binding events ICON identified from high-throughput pMHC binding data.

[0045] [Figure 30A] Figures 30A-B show characteristics of pMHC-binding T cells identified by ICON from high-throughput datasets. (A) Power law of unique T cell clone size for the top nine most abundant pMHC-binding T cell repertoires. (B) Shannon diversity scores for the top nine pMHC repertoires. [Figure 30B] Same as above.

[0046] [Figure 31A] Figures 31A-C show the performance of the TCRAI model and the gold standard dataset. (A) Schematic of the TCRAI framework for the model, which receives input for CDR3, and V, J genes of both α and β chains. The trained TCRAI model generates a numerical fingerprint and prediction for a given TCR. (B) ROC curves for TCRAI classification performance using eight curated public TCR-pMHC binding repertoires. Binders are unique TCRs that bind to a specific pMHC, and non-binders are unique TCRs that bind to other pMHC. Paired α and β TCR sequences were used as input data. FPR: false positive rate; TPR: true positive rate. (C) Classification performance comparison. TCRAI was compared to the predictive classifiers NetTCR, TCRdist, and DeepTCR. Area under the ROC curve (AUC) scores for NetTCR and TCRdist were generated using the original classifier with default parameters. AUC scores for DeepTCR (a multinomial classifier) ​​were derived from a slightly modified and hyperparameter-optimized version of DeepTCR (Methods) to compare it with these binary classifiers, NetTCR and TCRdist. For comparison, the binomial mode of TCRAI was used. [Figure 31B] Same as above. [Figure 31C] Same as above.

[0047] [Figure 32A-1] Figures 32A-C show the ROC performance of the TCR antigen-specificity classifier (a and b). (c) shows the ROC curve of the TCRAI in multinomial format using nine pMHC binding repertoires identified from a high-throughput dataset. Paired α and β TCR sequences were used as input data. FPR: false positive rate; TPR: true positive rate. [Figure 32A-2] Same as above. [Figure 32B] Same as above. [Figure 32C] Same as above.

[0048] [Figure 33] FIG. 33 is a table showing a comparison of TCR antigen-specific classifiers.

[0049] [Figure 34A]Figures 34A-D show TCRAI performance in the high-throughput dataset. (A) ROC curves for TCRAI on the top nine most abundant pMHC-binding repertoires. Binders are unique TCRs that bind to a specific pMHC, and non-binders are unique TCRs that bind to other pMHC. Paired α and β TCR sequences were used as input data. FPR: false positive rate; TPR: true positive rate. (B) Classification performance comparison using TCRα only, TCRβ only, or paired TCRα and β chains as sequence input. (C) ROC curves from independent testing of four overlapping pMHC repertoires between the curated public dataset and the high-throughput dataset. TCRAI was trained with pMHC repertoires identified from the high-throughput dataset and tested in the curated public dataset. (D) UMAPs of both training (high-throughput data) and testing ("gold standard" data) TCRAI fingerprints extracted from the high-throughput trained model. The right panel shows strong overlap between the A*02:01_ELAGIGILTV_MART-1_cancer training and test sets, while the right panel shows poor overlap between the A*02:01_NLVPMVATV_pp65_CMV training and test datasets. The black circles highlight areas with little overlapping fingerprints of bound TCRs. [Figure 34B] Same as above. [Figure 34C] Same as above. [Figure 34D] Same as above.

[0050] [Figure 35] Figure 35. ROC curve for TCRAI in multinomial format using nine pMHC-binding repertoires identified from a high-throughput dataset. Paired α and β TCR sequences were used as input data. FPR: false positive rate; TPR: true positive rate.

[0051] [Figure 36A]Figures 36A-B show a comparison of TCRAI fingerprints between models trained on different datasets. (A) Comparison of the "gold standard" TCR fingerprints generated by models trained on high-throughput and high-throughput data for two cases not shown in Figure 3d shows good overlapping binders in both cases. (B) The inference problem was reversed: training a model using the "gold standard" data and calculating fingerprints for the "gold standard" and high-throughput TCRs. For the A*02:01_NLVPMVATV_pp65 / CMV case, where cross-dataset performance is poor, the model trained on the "gold standard" data, which contains TCRs from many donors, separates a large group of binding TCRs. However, the high-throughput binding TCRs primarily come from a single donor, and this donor only has binding TCRs from a small cluster of TCR space that does not fully represent the range of binding TCRs occurring in the broader population. Black circles highlight TCRs unique to the high-throughput data. [Figure 36B] Same as above.

[0052] [Figure 37A]Figures 37A-G show TCR cluster characteristics. (A) Clustering of TCRAI fingerprints of high-confidence TCRs identified from a high-throughput dataset by a model trained to predict A*02:01_GILGFVFTL_Flu-MP_Influenza binders reveals two TCR clusters: Cluster 0 (orange) and Cluster 1 (green). (B) Dextramer signal (UMI) distribution for Clusters 0 and 1. (C) Conserved CDR3 motifs and gene usage in these two clusters of Flu peptide-binding TCRs. For Cluster 0, gene usage is shown for the 30 most common unique quartets so that significant variations can be seen in a single plot. (D) 3D structures of Flu peptide-binding TCR-pMHC binding complexes for a Cluster 0 TCR (PDB 2VLJ) and a Cluster 1 TCR (PDB 5JHD). In the top panel, only non-peptide residues within 0.4 nm (4 Å) of the Phe-5 ring (pink -strand, blue -strand, green MHC) are shown. In the bottom panel, comparison of peptide structures of TCR-pMHC binding complexes in cluster 0 and cluster 1. (E) Clustering of TCRAI fingerprints of TCRs with high-confidence binding to A*02-01_GLCTLVAML_BMLF1_EBV from the high-throughput dataset. (F) Distribution of dextramer signals (UMI) in EBV peptide-binding clusters 0–2. (G) Conserved CDR3 motifs and gene usage in these three clusters of EBV peptide-binding TCRs. [Figure 37B] Same as above. [Figure 37C] Same as above. [Figure 37D] Same as above. [Figure 37E] Same as above. [Figure 37F] Same as above. [Figure 37G] Same as above.

[0053] [Figure 38A]Figures 38A-F show the immunophenotype of pMHC-binding CD8+ T cells. (A) Classification of pMHC-binding cells. Clusters were visualized by UMAP, with cell types represented by different colors. (B) Heatmap of CD8+ T cell type marker gene and protein expression. *: Protein expression measured by CITE-seq. (C) pMHC-binding landscape by T cell immune subtype. Bars indicate the number of pMHC-binding T cells on a log2 scale. (D) Expanded clonotypes are enriched in the non-naive compartment. Each dot represents a unique TCR clone. (E) Pie chart describes the subpopulations of pMHC-binding CD8+ T cells. (F) Percentage of HLA-matched and mismatched binding in naive and non-naive binding T cells. Tpm: peripheral memory cells; Tcm: central memory cells; Tem: effector memory cells; Temra: highly differentiated effector memory cells; Others: other memory cells with marker expression CD43loKLRG1hiCD127. [Figure 38B] Same as above. [Figure 38C] Same as above. [Figure 38D] Same as above. [Figure 38E] Same as above. [Figure 38F] Same as above.

[0054] [Figure 39] Figure 39 shows the importance of VJ gene information. The error in AUC when comparing models trained using full inputs or only gene inputs is calculated by propagating the error in AUC for each model (full or gene), without assuming covariance between the outcomes. The error in AUC for each model was either the difference between the average AUC for the best hyperparameters in MCCV and the final model trained with those hyperparameters, or the standard deviation of the AUC in MCCV, whichever was larger. ΔAUC=AUCfull-AUCgene.

[0055] [Figure 40A]Figures 40A-B show TCR cluster characteristics. (A) Dextramar signal distribution of all five TCR clusters identified for A*02-01_GLCTLVAML_BMLF1_EBV as shown in the fingerprint space of Figure 4e. (B) Motif and gene usage of EBV peptide-binding TCR clusters 3 and 4. [Figure 40B] Same as above.

[0056]

[0057] [Figure 41] FIG. 41 illustrates an exemplary operating environment.

[0058] [Figure 42-1] FIG. 42 shows an exemplary method. [Figure 42-2] Same as above. [Figure 42-3] Same as above.

[0059] [Figure 43] FIG. 43 shows an exemplary method.

[0060] [Figure 44] FIG. 44 shows an exemplary method.

[0061] [Figure 45] FIG. 45 shows an exemplary method.

[0062] [Figure 46-1] FIG. 46 shows an exemplary method. [Figure 46-2] Same as above. [Figure 46-3] Same as above. DETAILED DESCRIPTION OF THE INVENTION

[0063] Understanding of the disclosed methods and compositions can be facilitated by reference to the following detailed description of specific embodiments and examples included therein, as well as the drawings and accompanying description.

[0064] A. Definitions It is to be understood that the methods and compositions of the present disclosure are not limited to the particular methodology, protocols, and reagents described, as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is defined solely by the appended claims.

[0065] It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, a reference to a "TCR" includes a plurality of such TCRs, a reference to a "dextramers" is a reference to one or more dextramers and equivalents thereof known to those skilled in the art, and so forth.

[0066] The term "subject" or "donor" may refer to an animal, such as a mammalian species (preferably a human) or an avian (e.g., avian) species. More specifically, the subject or donor may be a vertebrate, e.g., a mammal, such as a mouse, a primate, a monkey, or a human. Animals include livestock, sport animals, and pets. The subject or donor may be a healthy individual, an individual with symptoms or signs, or an individual suspected of having a disease or a predisposition to a disease, or an individual in need of treatment or suspected of needing treatment. In some embodiments, the subject donor is a human, such as a human with or suspected of having cancer.

[0067] As used herein, the term "barcode" generally refers to a label that can be attached to a molecule (e.g., a dextramers, cells) and convey information about the molecule. For example, a DNA barcode can be a polynucleotide sequence attached to each dextramers, and a common sequencing barcode can be a polynucleotide sequence attached during sequencing. This barcode can then be sequenced. The presence of the same barcode on multiple sequences can provide information about the origin of the sequence. For example, a barcode may indicate that the sequence came from a specific dextramers. A barcode can also indicate that the sequence came from a specific cell / dextramers combination.

[0068] As used herein, the term "sequencing" or "sequencer" refers to any of a number of techniques used to determine the sequence of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Exemplary sequencing methods include targeted sequencing, single molecule real-time sequencing, exon sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy end sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, double-stranded sequencing, cycle sequencing, single base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification with lower denaturing temperature PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, short-range sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome These include, but are not limited to, Genetic Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as a commercially available genetic analyzer from Illumina or Applied Biosystems.

[0069] A "polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) linked by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides usually range in size from a few monomeric units, e.g., 3-4, to several hundred monomeric units. When a polynucleotide is represented by a sequence of letters, such as "ATGCCTG," it will be understood that the nucleotides are in 5'→3' order from left to right, and that "A" represents adenosine, "C" represents cytosine, "G" represents guanosine, and "T" represents thymidine, unless otherwise indicated. The letters A, C, G, and T may be used as standard in the art to refer to the base itself, a nucleoside, or a nucleotide that includes the base.

[0070] The term "DNA (deoxyribonucleic acid)" refers to a chain of nucleotides containing deoxyribonucleosides, each containing one of four nucleobases: adenine (A), thymine (T), cytosine (C), and guanine (G). The term "RNA (ribonucleic acid)" refers to a chain of nucleotides containing four types of ribonucleosides, each containing one of four nucleobases: A, uracil (U), G, and C. Particular pairs of nucleotides specifically bind to each other in a complementary manner (called complementary base pairs). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand consisting of nucleotides complementary to those of the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence" or "nucleic acid sequencing read" refers to any information or data indicating the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule of nucleic acid such as DNA or RNA (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, or a fragment). It should be understood that the present teachings contemplate sequence information obtained using all available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide discrimination systems, pyrosequencing, ion-based or pH-based detection systems, and electronic signature-based systems.

[0071] "Optional" or "optionally" means that a subsequently described event, circumstance, or material may or may not occur, may be present, or may not be present, and that the description encompasses instances when said event, circumstance, or material occurs and does not occur, or is present and is not present.

[0072] Throughout this specification and the claims, the word "comprise" and variations of this word, such as "comprising" and "comprises," mean "including, but not limited to," and are not intended to exclude, for example, other additional things, components, integers, or steps. In particular, for methods described as including one or more steps or actions, each step is specifically intended to include what is recited (unless the step includes a limiting term such as "consisting of"), meaning that each step is not intended to exclude, for example, other additional things, components, or steps not recited in the step.

[0073] "Exemplary" means "one example of" and is not intended to convey an indication of a preferred or ideal configuration. "Etc." is not used in a limiting sense, but rather for illustrative purposes.

[0074] Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such ranges are expressed, the ranges that are specifically contemplated and considered disclosed are from the one particular value and / or the other particular value, unless the context specifically dictates otherwise. Similarly, when values ​​are expressed as approximations, by using the antecedent "about," it will be understood that the particular value forms another embodiment, and specifically contemplates embodiments that should be considered disclosed, unless the context specifically dictates otherwise. It will be further understood that the endpoints of each of these ranges are significant both in relation to the other endpoint, and independently of the other endpoint, unless the context specifically dictates otherwise. Finally, it should be understood that all individual values ​​and subranges of values ​​falling within explicitly disclosed ranges are also specifically contemplated and should be considered disclosed, unless the context specifically dictates otherwise. The foregoing applies regardless of whether, in a particular instance, some or all of these embodiments are explicitly disclosed.

[0075] B. Methods for identifying reliable receptor-pMHC binding and methods of use thereof In some embodiments, the described methods and systems can identify reliable TCR-pMHC binding by analyzing multi-omics high-throughput binding data. The methods and systems may be referred to herein as ICON (Integrated Context-Specific Normalization).

[0076] Disclosed is a method that includes receiving single cell sequence data, dextramer sequence data, and single cell receptor sequence data; filtering from the dextramer sequence data data associated with low quality cells based on the single cell sequence data; adjusting the dextramer sequence data based on a measurement of background noise; filtering from the dextramer sequence data data based on the presence or absence of specific receptor sequences based on the single cell receptor data; and identifying data remaining in the normalized filtered dextramer sequence data that is associated with reliable receptor-pMHC binding events.

[0077] The single cell sequence data and corresponding receptor sequence data can be from several cell types, including T cells (αβ or γδ) and B cells. Thus, by way of example, a method is disclosed that includes receiving single cell sequence data, dextramer sequence data, and single cell TCR sequence data; filtering data associated with low-quality cells from the dextramer sequence data based on the single cell sequence data; adjusting the dextramer sequence data based on a measurement of background noise; filtering data from the dextramer sequence data based on the presence or absence of α or β chains based on the single cell TCR data; and identifying data remaining in the normalized filtered dextramer sequence data associated with reliable TCR-pMHC binding.

[0078] 1. Data Acquisition A method for acquiring, receiving, and / or determining multi-omic high-throughput binding data is disclosed. As shown in FIG. 1 , the system 100 can include a single-cell immune profiling platform 102. The single-cell immune profiling platform 102 can be configured to generate multi-omic high-throughput binding data (e.g., sequence data 104). In one embodiment, the multi-omic high-throughput binding data can include one or more of single cell sequence data, dextramer sequence data, and / or single cell receptor sequence data. The single cell sequence data can include, for example, RNA-seq data. The dextramer sequence data can include, for example, dCODE-dextramer-seq and / or cell surface protein expression sequencing, also referred to as CITE-seq (Cellular Index of Transcriptomes and Epitopes by Sequencing). The single cell receptor sequence data can include, for example, TCR-seq data, such as TCR-seq data of a single cell versus an αβ chain (or γδ chain).

[0079] In some aspects, multi-omic high-throughput binding data can be previously generated and incorporated into the disclosed methods. In some aspects, multi-omic high-throughput binding data can be generated as part of the disclosed methods.

[0080] In some embodiments, peripheral blood mononuclear cells (PBMCs) from a healthy human donor may be labeled for sorting on cells, such as T cells or B cells, forming a single cell immune profiling platform 102, as shown in FIG. 2. In some embodiments, the cells may be T cells (e.g., CD4+ or CD8+ cells). In some embodiments, the T cells may be αβ T cells or γδ T cells. In some embodiments, the cells may be B cells. Thus, when labeled for sorting, the label may be a CD4, CD8, or B cell-specific label.

[0081] In some embodiments, once a cell type of interest has been sorted, the sorted cells can then be sorted for cells that bind to a specific peptide-major histocompatibility complex (MHC) (pMHC). In some embodiments, the cells can be combined with a set of dextramers, such as dCODE™ dextramers. In some embodiments, dCODE™ Dextramer® technology can be used. A dextramer can contain two or more MHCs, a peptide presented by each MHC, and a DNA barcode. In some embodiments, a pool of dextramers is used. In some embodiments, the pool of dextramers can include, but is not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 single dextramers, each containing a different pMHC. In some embodiments, the dextramer pool includes two or more single dextramers each containing a different pMHC. In some embodiments, the two or more MHCs on a single dextramer are identical and therefore present the same peptide. In some embodiments, the MHC can be MHC class I (MHC I) or MHC class II (MHC II). In some embodiments, the DNA barcode includes one or more primer sequences, a peptide-MHC (pMHC)-specific barcode, and a unique molecular identifier. In some embodiments, the dextramer can further include a label. For example, the label may be a fluorescent label. In some embodiments, cells that bind to a specific pMHC are sorted based on the label on the dextramer. In some embodiments, cells that bind to a specific pMHC are sorted based on a labeled antibody specific to the dextramer.

[0082] In some embodiments, cell sorting for a particular cell type and cell sorting for cells that recognize dextramers can be performed simultaneously or sequentially.

[0083] In some embodiments, after sorting the cells that bind to the pMHC-containing dextramers, each cell and the corresponding dextramer can be sequenced. In some embodiments, the cell sequence and the dextramer sequence (e.g., the DNA barcode sequence derived from the dextramer) all have a common sequencing barcode, which allows determining which cell sequence was associated with which dextramer sequence. In some embodiments, NextGEM technology can be used for sequencing. The common sequencing barcode is different from the DNA barcode in the dextramer.

[0084] In some embodiments, sequencing of cells bound to pMHC-containing dextramers provides sequence data 104, which may include single-cell sequence data, dextramer sequence data, and single-cell receptor sequence data. In some embodiments, the single-cell sequence data includes sequences from the entire cell genome or transcriptome. Thus, in some embodiments, the single-cell sequence data includes gene expression data. In some embodiments, the dextramer sequence data includes DNA barcode sequences. In some embodiments, the single-cell receptor sequence data includes the sequence of a specific receptor. For example, the single-cell receptor sequence data includes single-cell TCR or B-cell receptor (BCR) sequence data. In some embodiments, the single-cell TCR sequence data includes paired TCR sequence data. In some embodiments, the paired TCR sequence data includes sequence data for an α chain and a β chain, if present, for each cell. In some embodiments, the paired TCR sequence data includes sequence data for a γ chain and a δ chain, if present, for each cell. Thus, for each of the methods and examples described herein, sequencing of the alpha and beta chains can be interchanged with sequencing of the gamma and delta chains.

[0085] Returning to the system 100 shown in FIG. 1 , in one embodiment, the sequence data 104 may be provided to a computing device 106. The computing device 106 may be, for example, a smartphone, a tablet, a laptop computer, a desktop computer, a server computer, etc. The computing device 106 may include one or more servers. The computing device 106 may be configured to generate, store, maintain, and / or update various data structures, including databases for storage of one or more of the sequence data 102. The computing device 106 may be configured to operate one or more application programs, such as an integrated context-specific normalization (ICON) module 108 and / or a prediction module 110. The ICON module 108 and the prediction module 110 may be stored and / or configured to operate separately on the same computing device or on separate computing devices.

[0086] In some embodiments, the ICON module 108 can be configured to analyze received sequence data 104 (e.g., multi-omics high-throughput binding data, single-cell sequence data, dextramers sequence data, single-cell receptor sequence data, etc.). The sequence data 104 may include sequence information as well as meta-information. The sequence data 104 can be saved in any suitable file format, including, for example, a VCF file, a FASTA file, or a FASTQ file, as known to those skilled in the art. FASTA and FASTQ are common file formats used to save raw sequence reads from high-throughput sequencing. A FASTQ file stores an identifier for each sequence read, a sequence, and a quality score string for each read. A FASTA file stores only the identifier and the sequence. Other file formats are also contemplated.

[0087] 3, the ICON module 108 can be configured to perform a method 300 that includes filtering low-quality cells from the sequence data 104 (e.g., dextramer sequence data) in step 310, adjusting the sequence data 104 for background noise in step 320, selecting T cells with paired αβ chains in the sequence data 104 in step 330, applying dextramer signal correction to the sequence data 104 in step 340, performing cell- and / or pMHC-wise dextramer signal normalization and binder identification on the sequence data 104 in step 350, and identifying the data remaining in the normalized dextramer sequence data as associated with reliable TCR-pMHC binding events in step 360. In one embodiment, the ICON data processing may be performed in a donor-, cell-, and / or dextramer-specific manner.

[0088] Filtering low-quality cells from the sequence data 104 in step 310 may include single-cell RNA-seq-based filtering of low-quality cells. The ICON module 108 can be configured to filter low-quality cells, such as doublets and dead cells. Cells with an unexpectedly high number of genes for detected T cells (e.g., >2500 genes per cell) may be classified as doublets, and cells with a high fraction of mitochondrial gene expression (e.g., a ratio of mitochondrial gene expression UMI to total gene expression UMI >0.4) or too few detected genes (<200 genes per cell) may be classified as dead cells. Data associated with low-quality cells may be removed from the sequence data 104 (e.g., dextramar sequence data).

[0089] In one embodiment, filtering low-quality cells from the sequence data 104 in step 310 may include: for each cell represented in the dextramer sequence data, determining a number of genes based on the sequence data of a single cell; removing from the dextramer sequence data data data associated with cells whose number of genes is outside a gene threshold range (the gene threshold range may be, for example, about 200 to about 2,500 genes); for each cell represented in the dextramer sequence data, determining a fraction of mitochondrial gene expression based on the sequence data of a single cell; and removing from the dextramer sequence data data data associated with cells whose fraction of mitochondrial gene expression is above a gene expression threshold. The gene expression threshold may be about 40 percent of the total unique molecular identifier count.

[0090] Adjusting the sequence data 104 for background noise in step 320 may include a single-cell dCODE-dextramer sequence-based background adjustment. In one embodiment, two types of background noise controls designed for the dextramer binding assay include a negative control dextramer derived from dextramer-stained and sorted CD8+ T cells (denoted as nc, NC_dex), and a negative control dextramer derived from dextramer-stained CD8+ T cells without sorting on dextramer (denoted as Dex_unsorted, du). To examine the signal and noise distribution, the maximum dextramer signal in each cell's UMI (unique molecular identifier), representing the best binding of each cell, may be selected. Specifically, the non-specific dextramer binding signal of a cell is expressed as Max(nc1, ..., nc n ), where the maximum dextramer signal of n negative control dextramers included the dextramer pool. The dextramer binding signal of cells from the dextramer-stained and sorted sample (denoted as ds, Dex_sorted) is the maximum dextramer signal in the UMI of m test dextramers, Max(ds1, ..., dsm ) Similarly, the dextramer binding signal of cells from a non-Dex-sorted sample may be expressed as Max(du1, ..., du m ), Max(du, ..., du 44 ) P of nonspecific dextramer binding signals in UM 99.9 may be selected as the non-specific dextramer binding cutoff (absolute outliers of the negative dextramer control may be excluded).

[0091] To estimate the noise that may be introduced by the cell sorting process, cumulative analysis of dextramer binding signals between Dex_sorted and non-Dex_sorted samples may be compared to determine a cutoff for dextramer sorting efficiency. Kolmogorov-Smirnov test (KS test) p-values ​​may be calculated by comparing cumulative curves of dextramer-sorted and non-dextramer-sorted samples using each data point (dextramer UMI) as a sliding window. Dex_sorted and Dex_unsorted (argmaxD s,u The dextramer UMI, which defines the maximum difference in dextramer binding signal between dextramers, may be used as a threshold to estimate dextramer sorting efficiency. A measure of the estimated background noise (d) of a dextramer-sorted sample may be defined as: d=maximum(P 99.9 , argmaxD s,u ) The dextramer signal (UMI) for each test dextramer in the sorted cells may be corrected by subtracting a measurement of the estimated background noise (d). E c =E s -d

[0092] In one embodiment, adjusting the data for background noise in step 320 may include determining selected dextramer sequence data and unselected dextramer sequence data based on the dextramer sequence data. The selected dextramer sequence data may include selected test dextramer sequence data (dex_selected) and negative control dextramer sequence data (nc_dex). The unselected dextramer sequence data may include unselected test dextramer sequence data (dex_unselected). In step 320, method 300 determines, for each cell represented in the dextramer sequence data, a maximum negative control dextramer signal (Max(nc1,...,nc n In step 320, the method 300 may determine the maximum selected dextramer signal (Max(ds1,...,ds)) based on the selected test dextramer sequence data (dex_selected) for each cell represented in the dextramer sequence data. m In step 320, method 300 may determine, for each cell represented in the dextramer sequence data, a maximum unsorted dextramer signal Max(du,...,du) based on the unsorted test dextramer sequence data (dex_unsorted). m ) may be determined.

[0093] In step 320, the method 300 calculates the dextramer binding background noise (P) based on the maximum negative control dextramer signal. 99.9 ) and the dextramer sorting gate efficiency (argmaxD) was calculated based on the maximum sorted dextramer signal and the maximum unsorted dextramer signal. s,u ) may be estimated. The dextramer selection gate efficiency can be calculated, for example, by the Max(ds1,...,ds m) and Max(du,...,du) of the unsorted dextramers sequence data m ) may be determined by the largest difference between

[0094] The method 300, in step 320, comprises determining the dextramer-bound background noise (P 99.9 ) and dextramer sorting gate efficiency (argmaxD s,u ), a measure of background noise (d) was determined, and for each cell represented in the dextramer sequence data, the measure of background noise (d) was calculated as the dextramer signal (E) associated with each cell. c =E s -d) may be subtracted.

[0095] In one embodiment, selecting T cells with paired αβ chains in sequence data 104 in step 330 may include determining, for each cell represented in the dextramer sequence data, the presence or absence of at least one α chain and at least one β chain based on the TCR sequence data of a single cell, and removing from the dextramer sequence data data data associated with cells with only an α chain, only a β chain, or multiple α or β chains based on the presence or absence of at least one α chain and at least one β chain. Step 330 may also include removing any data from the dextramer sequence data that is not associated with cells with a single paired γδ chain. Thus, the same steps for adjusting for background noise in step 320 can be performed with respect to the presence or absence of a γ chain and / or a δ chain.

[0096] Selecting T cells with paired αβ chains in sequence data 104 in step 330 may include removing any data from the dextramer sequence data that is not associated with cells with a single paired αβ chain. Single cell receptor sequence data (e.g., single cell TCR-seq data) may be used to determine data associated with T cells with only α chains, only β chains, and multiple α or β chains, and such data may be removed from sequence data 104 (e.g., dextramer sequence data). For T cells with multiple detected α or β chains, the α or β chain with the highest UMI count may be assigned to the respective T cell. For example, if a T cell has four detected α chains and four β chains, the β chain with the highest UMI may be selected from the list of all β chains. Similarly for the α chain. The α or β chain selected from this process may be assigned to the cell.

[0097] Method 300 may include applying a dextramer signal correction to sequence data 104 in step 340. In step 340, the dextramer signal in sequence data 104 may be corrected to obtain corrected dextramer sequence data. While each dextramer has optimal binding conditions, it is not possible to arrange experimental conditions in a multiplexed dextramer binding assay so that they are optimal for each dextramer. This results in multiple dextramers binding to the same T cell / clone. To correct for this effect, dextramer signals may be penalized when simultaneously binding to the same T cell / clone using the following technique:

[0098] j th Binds to dextramer th The background noise-subtracted dextramer signal for T cells was measured using E ij Defining it as i th j about T cells th The fraction of the dextramer signal that was due to dextramer binding is further shown as follows:

[0099]

number

[0100] i th TCR clonotype of T cells i and T kij Clonotype k binds to dextramer j as i The number of T cells belonging to j th Dextramer-binding clonotype k i The fractions of T cells belonging to the genotype are as follows:

[0101]

number

[0102] Using these quantities, j th Binds to dextramer th The corrected dextramer signal for T cells is calculated as follows:

[0103] S ij =E ij (RC ij ) 2 RT kj

[0104] In step 350, method 300 may normalize the corrected dextramer sequence data by performing cell-wise normalization on the dextramer signal associated with each cell for each cell represented in the dextramer sequence data and / or pMHC-wise normalization for each cell represented in the dextramer sequence data. Such normalization may result in normalized dextramer sequence data. Step 350 may further include binder identification. To equalize all dextramer binding signals, the corrected dextramer binding signal may be a normalized log-ratio across the 44 test dextramers in the cell. Subsequently, pMHC-wise normalization may be performed based on the log-rank distribution. A normalized dextramer UMI > 0 was empirically selected as the cutoff for pMHC-specific binders.

[0105] In one embodiment, the corrected dextramer sequence data may be normalized in step 350. For example, cell-wise normalization may be performed based on the log-rank distribution for each cell, and / or pMHC-wise normalization may be performed to make the dextramer binding signals comparable to each other. c The adjusted dextramer binding signal of may be normalized across the test dextramers and then across all cells according to the following equation:

number

number

[0106] The method 300 may further identify remaining data in the normalized dextramer sequence data that are associated with reliable TCR-pMHC binding events in step 360. Such data may be considered part of a training dataset for use in the machine learning process. The resulting processed sequence data 104 (e.g., training dataset) may be provided to the prediction module 110.

[0107] C. Using Reliable Receptor-pMHC Binding for Machine Learning 4, the prediction module 110 is described. The prediction module 110 may be configured to use machine learning (“ML”) techniques to train, based on analysis of one or more training datasets 410, by a training module 420, at least one ML module 430 configured to predict binding affinities for a given receptor sequence.

[0108] The training dataset 410 may include one or more receptor sequences, one or more gene identifiers, a binding status, and an identifier of the peptide to which the receptor sequence bound (if any). The binding status may indicate "yes" for receptor sequences that bound the peptide or "no" for receptor sequences that did not bind the peptide. For receptor sequences that bound the peptide, the identifier of the peptide can be used to identify the antigen associated with the peptide. Such data may be derived, in whole or in part, from sequence data 104 processed by the ICON module 108. In one embodiment, a TCR-CDR3 amino acid sequence may be determined from sequence data 104, including associated V, D, and J gene identifiers, an indicator indicating the binding status (yes, no), and an identifier of the peptide to which the TCR-CDR3 amino acid sequence bound. The TCR-CDR3 amino acid sequence may be coded with numbers representing 20 possible amino acids. Padding may be applied to the sequence as needed. The V and J gene identifiers may be one-hot coded to provide a categorical and separate representation of the gene identifiers in the computational space. The encoded TCR-CDR3 amino acids and V and J gene identifiers may be concatenated together to represent one TCR that is recorded and associated with a label indicating the binding status (yes, no). The label may further indicate the particular peptide that the TCR bound. One or more TCR records may be combined to provide a training dataset 410.

[0109] A subset of TCR records may be randomly assigned to a training dataset 410 or a test dataset. In some implementations, the assignment of data to a training dataset or a test dataset may not be completely random. In this case, one or more criteria may be used during the assignment. In general, any suitable method may be used to assign data to a training dataset or a test dataset, while ensuring that the distribution of yes and no labels is somewhat similar in the training dataset and the test dataset.

[0110] The training module 420 may train the ML module 430 by extracting a feature set from a plurality of TCR records (e.g., labeled as yes) in the training dataset 410 via one or more feature selection techniques. The training module 420 may train the ML module 430 by extracting a feature set from the training dataset 410 that includes statistically significant features of the positive examples (e.g., labeled as yes) and statistically significant features of the negative examples (e.g., labeled as no).

[0111] The training module 420 may extract feature sets from the training dataset 410 in a variety of ways. The training module 420 may perform feature extraction multiple times, each time using a different feature extraction technique. In one example, feature sets generated using different techniques may each be used to generate a different machine learning-based classification model 440. For example, the feature set with the highest quality metric may be selected for use in training. The training module 420 may use the feature sets to build one or more machine learning-based classification models 440A-440N configured to indicate whether a novel receptor sequence (e.g., with an unknown binding state) likely or likely does not bind to a peptide or pMHC.

[0112] The training dataset 410 may be analyzed to determine any dependencies, associations, and / or correlations between features and yes / no labels in the training dataset 410. The identified correlations may have the form of a list of features associated with different yes / no labels. As used herein, the term "feature" may refer to any characteristic of an item of data that can be used to determine whether an item of data falls within one or more particular categories. By way of example, the features described herein may include one or more sequence patterns, amino acid sequences of one or both alpha and beta chains, names of the v and j gene segments of one or both alpha and beta chains.

[0113] The feature selection technique may include one or more feature selection rules, which may include feature occurrence rules, which may include determining which features occur a threshold number of times in the training dataset 410 and identifying those features that meet the threshold as candidate features.

[0114] A single feature selection rule may be applied to select features, or multiple feature selection rules may be applied to select features. Feature selection rules may be applied in a cascading manner, where feature selection rules are applied in a particular order and apply to the results of previous rules. For example, feature generation rules may be applied to the training dataset 410 to generate a first list of features. The final list of candidate features may be analyzed by further feature selection techniques to determine one or more candidate feature sets (e.g., feature sets that can be used to predict binding). Any suitable computational technique may be used to identify the candidate feature sets using any feature selection technique, such as a filter method, a wrapper method, and / or an embedding method. One or more candidate feature sets may be selected according to a filter method. Filter methods include, for example, Pearson's correlation, linear discriminant analysis, analysis of variance (ANOVA), chi-square, combinations thereof, etc. Feature selection according to a filter method is independent of any machine learning algorithm. Instead, features may be selected based on scores in various statistical tests for correlation with an outcome variable (e.g., yes / no).

[0115] As another example, one or more candidate feature sets may be selected using a wrapper method. The wrapper method may be configured to use a subset of features and train a machine learning model using the subset of features. Features may be added and / or removed from the subset based on inferences drawn from previous models. Wrapper methods include, for example, forward feature selection, backward feature reduction, recursive feature reduction, combinations thereof, and the like. As an example, one or more candidate feature sets may be identified using forward feature selection. Forward feature selection is an iterative method that starts with no features in the machine learning model. In each iteration, the feature that best improves the model is added until adding a new variable no longer improves the performance of the machine learning model. As an example, one or more candidate feature sets may be identified using backward elimination. Backward reduction is an iterative method that starts with all features in the machine learning model. In each iteration, the lowest-ranked feature is removed until no improvement is observed upon feature removal. Recursive feature elimination may be used to identify one or more candidate feature sets. Recursive feature reduction is a greedy optimization algorithm that aims to find the feature subset with the best performance. Recursive feature reduction builds a model iteratively, setting aside the best or worst performing features at each iteration. Recursive feature reduction builds the next model with the remaining features until all features are exhausted. Recursive feature reduction then ranks the features based on their order of reduction.

[0116] As a further example, one or more candidate feature sets may be selected using an embedding method, which combines the qualities of filter and wrapper methods. Embedding methods include, for example, least absolute shrinkage and selection operator (LASSO) and ridge regression, which implement a penalty function to reduce overfitting. For example, LASSO regression implements L1 regularization, which applies a penalty equal to the absolute value of the coefficient magnitude, while ridge regression implements L2 regularization, which applies a penalty equal to the square of the coefficient magnitude.

[0117] After the feature set is generated by the training module 420, the training module 420 may generate a machine learning-based classification model 440 based on the feature set. A machine learning-based classification model may refer to a complex mathematical model for data classification generated using machine learning techniques. In one example, the machine learning-based classification model 440 may include a map of support vectors representing boundary features. In this example, the boundary features may be selected from and / or represent the highest-ranked features in a feature set.

[0118] The training module 420 may use the feature sets extracted from the training dataset 410 to build machine learning-based classification models 440A-440N for each classification category (e.g., yes, no). In some examples, the machine learning-based classification models 440A-440N may be combined into a single machine learning-based classification model 440. Similarly, the ML module 430 may represent a single classifier containing single or multiple machine learning-based classification models 440 and / or multiple classifiers containing single or multiple machine learning-based classification models 440.

[0119] The extracted features (e.g., one or more candidate features) may be combined in a classification model trained using machine learning approaches, such as discriminant analysis; decision trees; nearest neighbor (NN) algorithms (e.g., k-NN models, replicator NN models, etc.); statistical algorithms (e.g., Bayesian networks, etc.); clustering algorithms (e.g., k-means, mean shift, etc.); neural networks (e.g., reservoir networks, artificial neural networks, etc.); support vector machines (SVMs); logistic regression algorithms; linear regression algorithms; Markov models or chains; principal component analysis (PCA) (e.g., for linear models); multilayer perceptron (MLP) ANNs (e.g., for nonlinear models); reservoir network replication (e.g., for nonlinear models, typically for time series); random forest classification; combinations thereof, and / or the like. The resulting ML module 430 may include a decision rule or mapping for each candidate feature for assigning binding states to novel receptor sequences.

[0120] In one embodiment, the training module 420 may train the machine learning-based classification model 440 as a convolutional neural network (CNN). The CNN may include at least one convolutional feature layer and three fully connected layers leading to a final classification layer (softmax). The final classification layer may be finally applied to combine the outputs of the fully connected layers using a softmax function known in the art.

[0121] The candidate feature and ML module 430 may be used to predict the binding states (and associated peptides) of multiple TCR records in a test dataset. In one example, the results of each TCR record include a confidence level corresponding to the likelihood or probability that the receptor sequence binds to the peptide. The confidence level may be a value between zero and one, which may represent the likelihood that the receptor sequence belongs to a yes / no binding state with respect to one or more peptides. In one example, when there are two states (e.g., yes and no), the confidence level may correspond to a value p, which refers to the likelihood that a particular receptor sequence belongs to the first state (e.g., yes). In this case, the value 1-p may refer to the likelihood that a particular receptor sequence belongs to the second state (e.g., no). Generally, when there are more than two states, multiple confidence levels may be provided for each test receptor sequence and for each candidate feature. The best-performing candidate feature may be determined by comparing the results obtained for each test receptor sequence with the known yes / no binding state for each test receptor sequence. Generally, the best-performing candidate feature will have results that closely match the known yes / no binding state.

[0122] The best performing candidate features may be used to predict a yes / no binding status of the receptor sequence with respect to one or more peptides. For example, a novel TCR sequence may be determined / received. The novel TCR sequence may be applied to an ML module 430, which may classify the novel TCR sequence as either binding (yes) or not binding (no) and an indication of a binding peptide based on the best performing candidate feature.

[0123] 5 is a flowchart illustrating an example training method 500 for generating an ML module 530 using a training module 420. The training module 420 can implement supervised, unsupervised, and / or semi-supervised (e.g., reinforcement-based) machine learning-based classification models 440. The method 500 illustrated in FIG. 5 is an example of a supervised learning method; variations of this example training method are discussed below, however, other training methods can be similarly implemented to train unsupervised and / or semi-supervised machine learning models.

[0124] The training method 500 may determine (e.g., access, receive, retrieve, etc.) first sequence data processed by the ICON module 108 in step 510. The sequence data may include a labeled set of receptor sequences. The labels may correspond to a binding status (e.g., yes or no) and the identity of the peptide to which the receptor sequence bound.

[0125] The training method 500 may generate a training data set and a test data set in step 520. The training data set and the test data set may be generated by randomly assigning labeled receptor sequences to either the training data set or the test data set. In some implementations, the assignment of labeled receptor sequences as training or test samples may not be completely random. As an example, a majority of the labeled receptor sequences may be used to generate the training data set. For example, 75% of the labeled receptor sequences may be used to generate the training data set, and 25% may be used to generate the test data set.

[0126] The training method 500 may, in step 530, determine (e.g., extract, select, etc.) one or more features that can be used by a classifier to distinguish among different classes of binding states (e.g., yes vs. no) for one or more peptides, for example. As an example, the training method 500 may determine the set of features from labeled receptor sequences. In a further example, the set of features may be determined from labeled receptor sequences other than the receptor sequences labeled in either the training data set or the test data set. In other words, the labeled receptor sequences may be used for feature determination rather than for training a machine learning model. Such labeled receptor sequences may be used to determine an initial set of features, which may be further reduced using the training data set.

[0127] According to the training method 500, one or more machine learning models may be trained at 540 using one or more features. In one example, the machine learning models may be trained using supervised learning. In another example, other machine learning techniques may be used, including unsupervised learning and semi-supervised learning. The machine learning models trained at 540 may be selected based on different criteria depending on the problem being solved and / or the data available in the training dataset. For example, machine learning classifiers may be subject to different degrees of bias. Thus, more than one machine learning model may be trained at 540 and optimized, improved, and cross-validated at 550.

[0128] The training method 500 may select one or more machine learning models to build a predictive model at 560. The predictive model may be evaluated using a test dataset. The predictive model may analyze the test dataset and generate predicted binding states at step 570. The predicted binding states may be evaluated at step 580 to determine whether such values ​​achieve a desired level of accuracy. The performance of a predictive model may be evaluated in a number of ways based on classification of multiple true positives, false positives, true negatives, and / or false negatives of multiple data points represented by the predictive model.

[0129] For example, a false positive of a predictive model may refer to the number of times the predictive model incorrectly classifies a receptor sequence as bound when in fact it is not. Conversely, a false negative of a predictive model may refer to the number of times the machine learning model classifies a receptor sequence as unbound when in fact it is bound. True negatives and true positives may refer to the number of times the predictive model correctly classifies one or more receptor sequences as bound or unbound. Related to these measurements are the concepts of recall and precision. Generally, recall refers to the ratio of true positives to the sum of true positives and false negatives, thereby quantifying the sensitivity of the predictive model. Similarly, precision refers to the ratio of true positives to the sum of true positives and false positives. When such a desired level of precision is reached, the training phase ends and the predictive model (e.g., ML module 430) may be output in step 590. However, when the desired level of precision is not reached, subsequent iterations of the training method 500 may be performed, beginning in step 510, with variations, such as considering a larger collection of sequence data.

[0130] In one embodiment, a flexible framework for the study of TCR-pMHC specificity, referred to herein as TCRAI, is provided. In one embodiment, TCRAI may utilize Tensorflow 2. TCRAI is highly modular, allowing for tailoring to model building. Any number of V(D)J genes and CDR regions of a TCR may be defined as inputs to the model in textual format. A method for processing these inputs into numerical form in an unlearnable manner can be selected via a "processor" object, which converts text to numeric representations. These numeric inputs can then be further processed in a learnable manner via an "extractor" object, which forms the building blocks of the neural network and provides as their output vector representation of the input data, referred to herein as a TCRAI fingerprint. TCRAI fingerprints may be concatenated via a single numeric vector into a single TCRAI fingerprint describing the input TCR. The TCRAI fingerprint may then be passed through a "closer" object, which forms the final building block of the neural network, to generate a prediction on the input TCR. TCRAI provides several such pre-built processors, extractors, and closers. TCRAI can be configured to perform binomial, polynomial, regression, and / or other tasks by choosing to build different closer objects. In one embodiment, TCRAI may be used to build a model to make a prediction of whether a given TCR can bind to a particular pMHC complex.

[0131] In one embodiment, TCRAI may utilize 1D convolution and batch normalization for CDR3 sequences and a low-dimensional representation for genes, which results in model normalization and forces the model to learn stronger gene associations.

[0132] In one embodiment, the input information of the TCR may be processed in a numeric format. For each CDR3 sequence, the amino acids may be converted to integers, and the integer vector may be coded in one-hot representation. For V and J genes, a dictionary of gene types to integers may be constructed for each V and J gene and used to convert each gene to an integer.

[0133] The neural network architecture applied to the processed input information may include an embedding layer and a convolutional network. Specifically, the processed CDR3 residues may be embedded in a 16-dimensional space through a learned embedding, and the resulting numerical CDR3 may be fed through one or more (e.g., three) 1D convolutional layers. In one embodiment, a filter with dimensions [64, 128, 256], kernel width [5, 4, 4], and stride [1, 3, 3] may be used. Each convolution may be activated by exponential linear unit activation, followed by dropout and batch normalization. After these three convolutional blocks, global max pooling may be applied to the final feature, and this process encodes each CDR3 by a 256-length vector, a "CDR3 fingerprint." The processed gene input for each gene may be one-hot coded and embedded into a reduced-dimensional space (e.g., 16 for V genes and 8 for J genes) via a learned embedding, thereby providing a "gene fingerprint" for each gene as a vector. All selected CDR3 and gene fingerprints may then be concatenated into a single vector, the "TCRAI fingerprint." The TCRAI fingerprint may be passed through one final fully connected layer to provide a binomial prediction (single output value, sigmoid activation), a regression prediction (single output, no activation), or a multinomial prediction (multiple output values, softmax activation).

[0134] In one embodiment, TCR sequencing files may be collected as raw multi-omics high-throughput binding data in a CSV format. The sequencing files may be analyzed to obtain the amino acid sequences of the CDR3s after removing non-productive sequences. Clones with different nucleotide sequences but identical amino acid sequences from the CDR3s and V, D, and J genes may be aggregated together under a single TCR. Thus, each TCR record may include a single pair of α and β TCR chains, each with its own CDR3 amino acid sequence and V and J genes.

[0135] The data may be divided into a training set (e.g., 76.5%), a validation set (e.g., 13.5%), and a left-pruned test set (e.g., 10%) for each model, followed by 5-fold Monte Carlo cross-validation (MCCV) on the training set. The model may be trained by minimizing the cross-entropy loss via the Adam optimizer, where the cross-entropy loss may be weighted for each class by 1 / (number of classes * fraction of samples in that class). To prevent overfitting, early stopping may be coupled via the left-pruned validation data set; in this case, the model stops training when the validation loss increases more than five times and the weights of the model with the smallest validation loss are restored. When training a large number of models, only the learning rate and batch size need to be adjusted during cross-validation. After cross-validation, the optimal implementation of the hyperparameters may be selected, and the model may be re-trained on the full training set, using the validation set to control early stopping. The re-trained model may then be evaluated on the left-pruned test set.

[0136] The TCRAI model can generate both a prediction about TCRs that bind to a particular pMHC (one of many pMHCs, in the polynomial case) and a numerical vector (TCRAI fingerprint) that describes that TCR within the context of the question of whether it binds to that pMHC (e.g., by encoding the paired αβ chain CDR3 amino acid sequences and V and J genes of each TCR into a one-dimensional input vector).

[0137] In one embodiment, the distribution of fingerprints may be analyzed to identify groups of TCRs with different binding modes. Fingerprints can be reduced to a two-dimensional space, for example, using UMAP: uniform manifold approximation and projection for dimensionality reduction. When using a model trained on one dataset to estimate fingerprints on another unseen dataset, a UMAP projector can be fitted using the TCRs from the training dataset and used to transform the TCRs from the unseen set.

[0138] When clustering TCR fingerprints, the fingerprints of all TCRs in a dataset can be projected into a two-dimensional space as described above, and then those TCRs that are strong true positives (STPs, binomial prediction >0.95) can be selected. These STPs can then be clustered in the two-dimensional space using, for example, a k-means classifier. Other clustering algorithms may also be used. TCRs from within each cluster can then be collected and used to construct CDR3 motif logos (using weblogos), gene usage, and / or UMI distributions by pairing unique TCR clonotypes within the cluster with all repeated clonotypes in the high-throughput data.

[0139] D.How to use In one embodiment, a trained predictive model (e.g., a machine learning classifier) ​​may be used to predict the binding status of a TCR sequence with respect to one or more peptides. A TCR sequence may be submitted to the machine learning classifier. The machine learning classifier may predict the likelihood that the TCR sequence will bind to one or more particular peptides. Similarly, multiple TCR sequences may be submitted to the machine learning classifier. The machine learning classifier may predict, for each TCR sequence in the multiple TCR sequences, the likelihood that each TCR sequence will bind to one or more particular peptides. In one embodiment, the machine learning classifier can generate a TCR-peptide map shown in the example output below. [Table 1]

[0140] Thus, the generated TCR-peptide map can be used to quickly identify peptides that likely bind to the subject's TCR sequence. A biological sample (e.g., blood) can be obtained from the subject, and cells can be isolated and sequenced. The subject's TCR sequence can be identified and compared to the TCR-peptide map to identify peptides that are most likely to bind to the subject's TCR sequence.

[0141] In some embodiments, identifying and evaluating antigen-specific T cells can be used to better understand drug activity in monotherapy and combination therapy settings, distinguish between potent anti-tumor T cell signatures, screen for immunogenic epitopes in a haplotype-associated manner, develop novel vaccines and TCR therapies, and develop peptide binding algorithms based on TCR sequence characteristics.

[0142] In some embodiments, a method for identifying a subject using the binding pattern of the subject's TCR is disclosed. For example, blood may be drawn (first blood draw), cells from the blood may be processed through a single cell-based immune profiling platform, and the resulting data may be processed according to the ICON method described herein. In some embodiments, the cells are exposed to various dextramers containing pMHC from a wide range of immunogens. After performing the ICON method as described herein, a reliable TCR binding pattern can be determined. In some embodiments, the TCR binding pattern represents the specificity of the TCR to the immunogen on the dextramer. Blood can then be drawn at a different time point (days, weeks, months, or years later) from the first blood draw (second blood draw). In some embodiments, the second blood draw is about 10 15 Although there are many possible TCR sequences, it is unlikely that the TCR binding pattern will change, so it is expected that the second blood draw will likely contain T cells with TCRs with sequences different from those present in the first blood draw. Cells from the second blood draw may be exposed to the same dextramers used in the first blood draw, and the resulting data analyzed according to the ICON method. Regardless of the different TCR sequences, the binding data from the first and second blood draws can be compared to determine whether they both come from the same subject.

[0143] In some embodiments, a method for identifying a subject using machine learning to predict the subject's TCR binding pattern is disclosed. Reliable TCR binding data can be identified according to the ICON method described herein. In some embodiments, reliable TCR binding data can be used to train a machine learning classifier described herein. The trained machine learning classifier can be used to predict the subject's specific TCR binding pattern. In some embodiments, blood can be drawn (first blood draw), and the TCR binding pattern can be predicted using the trained machine learning classifier. The blood can then be drawn at a different time point (days, weeks, months, years later) from the first blood draw (second blood draw). In some embodiments, the second blood draw is about 10 15 Although there are many possible TCR sequences, it is expected that the TCR binding pattern is unlikely to change, and therefore the first blood draw is likely to contain T cells with TCRs having sequences different from those present in the first blood draw. Regardless of the different TCR sequences, a trained machine learning classifier may be used to predict the second TCR binding pattern using data derived from the second blood draw. The second blood draw can be predicted to be from the same subject as the first blood draw based on the TCR signature.

[0144] In some embodiments, TCR or BCR binding patterns can be established using the methods described herein.In some embodiments, having reliable TCR data identified using the methods described herein allows someone, such as a medical professional, to estimate the antigenic history or vaccine history of a subject.In some embodiments, reliable TCR data identified using the ICON method described herein allows someone, such as a medical professional, to infer which pathogen a subject has been exposed to or which country a subject has visited.For example, the existence of TCR binding data for pathogens that only exist in Africa can indicate that a subject has been in Africa and has been exposed to these pathogens.

[0145] In some embodiments, reliable TCR data identified using the ICON method described herein can assess a subject's current immune status. For example, blood may be drawn (first blood draw), blood-derived cells may be processed through a single cell-based immune profiling platform, and the resulting data may be processed according to the ICON method described herein to obtain TCR binding data. In some embodiments, the dextramers used to establish TCR binding data contain tumor-specific pMHC. Thus, once the TCR binding data is normalized using the ICON method and reliable TCR binding data is established, the presence of a predicted tumor-specific TCR can be determined. For example, reliable TCR data can be used in the disclosed machine learning (CNN) method, and thus blood from a subject can be analyzed for the presence of a predicted tumor-specific TCR. Thus, the presence of a tumor-specific TCR can result in early detection of cancer before any tumor or cancer symptoms are detected.

[0146] In some embodiments, a method for selecting T cells for T cell-based therapy is disclosed. In some embodiments, training data can be accumulated using the disclosed methods of machine learning classification. In some embodiments, the classifier can assign a probability of pMHC binding to each TCR sequence tested. In some embodiments, the tested TCR sequences are associated with T cells, which may be derived from primary or secondary cell cultures. This avoids the need to perform binding assays on all T cells tested to determine whether each T cell has a TCR specific for a different pMHC. Instead, the classifier is trusted to determine the probability of TCR-pMHC binding. Thus, those TCRs classified as highly selective for a particular pMHC, and T cells containing such TCRs, can be used in T cell therapy. In some embodiments, because only the most reliable binding data was used to generate the training data used to classify the TCR associated with the selected T cells, T cells identified via the machine learning classifier can provide safer cell therapy than those T cells identified via binding assays.

[0147] In some embodiments, an immune monitoring method is disclosed. In some embodiments, blood can be collected from a subject undergoing immunotherapy (e.g., vaccine treatment, immune checkpoint treatment), and cells, particularly T cells, can be classified as having or not having specificity for a target epitope based on training data established by the disclosed machine learning approach. In some embodiments, if T cells are determined to have specificity for a target epitope, then it can be predicted whether the subject will respond to immunotherapy or will respond to immunotherapy. For example, if the immunotherapy is a vaccine that induces an immune response to a cancer-specific antigen, T cells obtained from the subject are classified based on their probability of binding to the cancer-specific antigen. If T cells with a high probability of binding to the cancer-specific antigen are selected based on training data obtained using single cell immune profiling technology and ICON, then the subject will be considered a responder to immunotherapy (e.g., vaccine).

[0148] In some embodiments, methods of TCR epitope mapping using the disclosed methods are disclosed. In some embodiments, TCR epitope mapping is a term that refers to the process of identifying the specific (possibly shortest) amino acid sequence of an epitope of a particular antigen that is recognized by a T cell (CD4+ and / or CD8+) receptor, while potentially stimulating a long-term and cytotoxic immune response. During the disclosed single cell immune profiling platform technology, dextramers can be used, and all different epitopes from one or more antigens of interest can be displayed on the dextramers. In other words, a single dextramers can contain pMHC, and the peptide of the pMHC is a single epitope from one or more antigens of interest, and sufficient dextramers are used so that all epitopes of the one or more antigens are present on the pMHC on the dextramers. T cells can be exposed to dextramers in the disclosed single cell immune profiling platform, with dextramers containing a single epitope from one or more antigens of interest, and sufficient dextramers are used so that all epitopes of the one or more antigens of interest are present in the pMHC of the dextramer. The sequence data of the single cell, the dextramer sequence data, and the TCR sequence data of the single cell obtained from single cell immune profiling can provide data about T cells that bind to different dextramers (e.g., epitopes). The single cell immune profiling data is then processed using ICON as described herein, thus providing binding data for those cells that have the most reliable binding to one or more epitopes of the one or more antigens of interest. In some embodiments, machine learning classification of TCRs that bind to one or more epitopes of the one or more antigens of interest can be used to predict which T cells from a subject will be reactive to a particular antigen (e.g., a tumor antigen). E. Kit

[0149] The above materials, as well as other materials, can be packaged together in any suitable combination as a kit useful for practicing or assisting in practicing the disclosed methods. A given kit is useful if the kit components are designed and adapted for use together in the disclosed methods. For example, a kit for generating single cell sequencing data is disclosed, where the kit includes reagents for single cell immune profiling. In some embodiments, the kit can include one or more of the disclosed dextramers containing pMHC. In some embodiments, the kit can include NextGEM sequencing materials. In some embodiments, the kit can include multi-omics high-throughput binding data, including one or more of single cell sequence data, dextramer sequence data, and / or single cell receptor sequence data.

[0150] Example The following examples illustrate the present methods and systems as they relate to the detection of colorectal cancer and are not intended to be limiting.

[0151] A. Example 1 1.Results i. Multi-omics high-throughput TCR-pMHC binding data. 10x Genomics recently generated a scalable, publicly available TCR-pMHC binding dataset. In their initial report, they assessed the binding properties of over 150,000 CD8+ T cells from four HLA-haplotyped healthy donors (Figure 19) across 44 pMHC dextramers using a single-cell-based immune profiling platform to directly detect antigen binding to T cells, while simultaneously sequencing T cell αβ chain pairs and transcriptomes (Figure 2). The dextramar pool spanned eight HLA alleles and consisted of epitopes with known common viral and cancer responses (Figure 20).

[0152] This paper describes a highly multiplexed dextramer binding data set generated at the single cell level. 10xGenomics used a simple approach to determine pMHC-binding TCRs by applying comprehensive cutoffs for background noise and non-specific dextramer binding to all donors. However, unexpectedly, a large number of promiscuous cross-HLA and cross-peptide associations were found from the TCR-pMHC binding events identified by this approach, especially in donors 3 and 4 (Figure 11A). During further investigation, data from donor 3 was excluded from this study due to data quality issues (Figure 11B).

[0153] To robustly identify reliable binding events from such high-throughput TCR-pMHC binding data, we developed the ICON integrated context-specific normalization method (Figure 6A, Figure 12, and Methods). The ICON data normalization process was performed in a donor-specific context by acquiring multi-omics high-throughput binding data from each donor separately as input data. Briefly, single-cell transcriptome data was used to select good-quality cells (raw and singleton). Then, both negative control dextramers (n = 6) and dextramer-unsorted samples were used for each donor as background controls to empirically estimate the background binding noise for each donor. Subsequently, the raw dextramer binding signals were corrected by subtracting the estimated background noise for each donor separately. The corrected dextramer signals were then normalized across cells and pMHC to directly generate equivalent dextramer binding signals. The distribution of ICON-normalized dextramer binding signals and binding specificities of expanded T cell clones demonstrates that ICON significantly increased the signal-to-noise ratio of high-throughput TCR-pMHC binding data (Figures 6A and 6B and 12B and 13).

[0154] ii. TCR-pMHC binding events identified from 10x Genomics high-throughput data. Applying ICON, a total of 20,843 CD8+ T cells were identified from 1,514 unique T cell clones binding to 29 pMHCs from three donors (Figure 7A, Figure 21 and Methods). The number of unique TCR-pMHC interactions identified from this high-throughput dataset is comparable in size to the total number of paired αβ TCRs in the VDJdb. Of the pMHC-binding TCRs, 98.9% of the total TCRs (94.7% of the unique TCRs) bound to seven pMHC:B TCRs. * 08:01_RAKFKQLL_BZLF1_EBV, A * 02:01_GILGFVFTL_Flu-MP_Influenza, A * 11:01_IVTDFSVIK_EBNA-3B_EBV,A * 03:01_KLGGALQAK_IE-1_CMV, A * 11:01_AVFDRKSDAK_EBNA-3B_EBV, A * 02:01_GLCTLVAML_BMLF1_EBV and A * 02:01_ELAGIGILTV_MART-1_Binds to cancer (Figure 7B and Figures 16 and 17).

[0155] The most common HLA haplotypes in the dextramer pool (A * 02:01) (Figures 14 and 15), donors 1 and 2 share a significant fraction of unique TCR-pMHC reactants (n=38) (Figure 7C). Donor 4 has A * 02:01 negative and has a different HLA haplotype from donors 1 and 2 (Figure 19). No shared pMHC-binding TCR sequences were observed between donor 4 and donors 1 and 2 binding (Figure 7C), indicating that the TCR-pMHC binding pattern is most likely HLA-restricted.

[0156] Interestingly, 37% of TCRs with a shared β chain paired with a different α chain. This percentage was slightly lower for shared TCR α chains (30.9%). Although the majority of TCRs with a shared α or β chain (approximately 92%) bound to the sample pMHC, approximately 8% of them recognized a different pMHC (Figure 7D), indicating that αβ pairing information is essential for accurate estimation of TCR functionality.

[0157] We suggest that TCR dual specificity (specificity vs. degeneracy) is a key feature of immune response mechanisms that significantly distinguish self from foreign peptides to avoid autoimmune reactions while maintaining broad antigen coverage. Indeed, we observed highly specific yet promiscuous TCR-pMHC interactions. 98.7% of unique TCRs bind to one specific pMHC, while the remaining TCRs interact with two or three pMHCs (Figure 7E and A). While we observed TCRs capable of interacting with multiple epitopes, these TCR-pMHC interactions generally follow an HLA-type-specific pattern. Over 99.3% of binding events were HLA-matched, and of these, 11.6% were HLA-matched to HLA A antigens that share similar primary anchor positions of the presented peptide. * 03-Supertype family member HLA A * 03:01 and A * However, 0.7% of binding events are cross-HLA type interactions.

[0158] iii. Convolutional neural network (CNN)-based classification of T cell antigen specificity. With this large and diverse TCR-pMHC binding dataset, a more robust functional classifier for computationally validating or prioritizing these binding events is desirable. Recent studies have demonstrated that convolutional neural networks (CNNs) can learn high-dimensional information from TCR sequences and thus robustly predict TCR-pMHC binding. A CNN-based framework was adapted for validating and / or predicting TCR-pMHC binding. Briefly, paired αβ chain CDR3 amino acid sequences and the V and J genes of each TCR were encoded into one-dimensional input vectors. Specifically, trainable embeddings were used to encode the CDR3 amino acid sequences and transform the V and J gene segments into vectors. The CNN structure may include one convolutional feature layer and three fully connected layers leading to a final classification layer (Figure 8A and Methods). To address potential bias that may be introduced by having an unbalanced number of binding and non-binding TCRs for a given pMHC, a class-weighted cost function was used for training (Methods).

[0159] To evaluate the performance of this CNN-based model, 11 pMHC-specific binding T cell repertoires were generated by conventional single multimer binding assays and antigen rechallenge assays as the gold standard dataset (Figure 23). Each curated pMHC-binding repertoire was divided into training, validation, and test sets. The CNN-based model was able to classify the antigen-binding specificities of curated TCRs with an average area under the curve (AUC) of 0.90 ((AUC) = 0.90) (Figure 8B). The CNN-based classifier was compared with TCR sequence similarity, a distance-based classifier. The CNN-based classifier outperformed distance-based predictive models, particularly for highly diverse pMHC repertoires (Figure 14) (Figure 8C). The classification performance difference (ΔAUC) between the CNN-based and distance-based classifiers positively correlated with the diversity of the pMHC-binding T cell repertoire, as measured by Shannon entropy (Figure 8D).

[0160] iv. Classification of pMHC-binding repertoires identified from 10× Genomics high-throughput data. Next, the CNN-based classifier was applied to the top seven pMHC-binding repertoires identified from the 10x Genomics binding data (Figure 7B and Figure 15). The seven pMHC repertoires were classified with an average (AUC) of 0.89 (Figure 9A). In these data, as in the curated dataset, the CNN-based classifier outperformed the distance-based model (Figure 16). To further computationally validate these binding TCRs, we compared four pMHC repertoires (A) that also had binding TCRs in the curated dataset. * 02:01_ELAGIGILTV_MART-1, A * 02:01_GILGFVFTL_Flu-MP, A * 02:01_GLCTLVAML_BMLF1_EBV, and A * CNN-based classifiers were used with additional A / B data from four curated repertoires and an independent in-hospital antigen rechallenge experiment (Methods). * We trained using four repertoires identified from the 10x Genomics dataset to predict the 02:01_ELAGIGILTV_MART-1 binding repertoire. Figure 9B shows the prediction results, which were comparable to the high performance in the training set.

[0161] Historically, TCR β chain sequencing has often been used to infer T cell antigen-binding specificity due to its higher combining capacity compared to the α chain. To quantitatively assess the contribution of TCR α and β chains in predicting TCR-pMHC interactions, either the α chain or the β chain was used as input to a CNN-based classifier instead of the paired αβ chain. Performance using the paired αβ chain was better than the α or β chain alone, with an average increase in AUC of 16% (Figure 9C). We observed unbalanced α and β chain contributions to the prediction of TCR-pMHC-specific recognition. For example, the β chain contribution was dominant in the A*02:01_GILGFVFTL_Flu-MP_Influenza repertoire, while the α chain was dominant in the A*02:01_GILGFVFTL_Flu-MP_Influenza repertoire.* 11:01_AVFDRKSDAK_EBNA-3B_EBV and A * 02:01_ELAGIGILTV_MART-1_ was more important in predicting cancer-specific binders (Figure 9C). Similarly, different levels of conservation of TCR VJ gene usage were observed between the α and β chains of these seven pMHC repertoires (Figure 9D). Furthermore, V gene usage was significantly different from A * With the exception of the dominant TRBV19 usage in the 02:01_GILGFVFTL_Flu-MP_Influenza repertoire, the α chain is generally more conserved than the β chain, which may partially explain the disproportionate classification performance between the α and β chains. Again, these results collectively demonstrate the importance of αβ pairing for accurate inference of TCR-pMHC interactions.

[0162] To further understand the conserved TCR sequence characteristics underlying classification, motif conservation of CDR3 amino acid sequences was explored among the 10 most predictive TCR sequences for each of these seven pMHC repertoires (Figure 9E). Consistent with VJ gene usage, motif conservation is generally more evident in α-chain CDR3s than in β-chain CDR3s (Figures 9E and 9D). For the four pMHC repertoires for which the VDJdb also possesses CDR3 amino acid motifs, the motifs identified from the 10× Genomics data are similar to those derived from the VDJdb (Figure 9E and Figure 17A). Collectively, the results indicate that pMHC-specific TCRs identified from high-throughput datasets are reliable binding partners and that CNN-based models can capture important conserved TCR sequence characteristics.

[0163] Immunophenotype of v.pMHC-binding CD8+ T cells. Combined information on antigen specificity and T cell phenotype has been reported to be important for the clinical success of immunotherapies such as vaccination. Multi-omics data generated by the 10x Genomics immune profiling platform allow us to link T cell antigen specificity with various T cell phenotypes. Gene (single cell RNA-seq) and surface protein (CITE-seq) expression levels from this multi-omics dataset were used to separate pMHC-binding CD8+ T cells into subpopulations (Methods and Figure 18). The identified subpopulations were then annotated according to previously described CD8+ T cell subtype marker genes: naive cells (CD45RA+CD45RO-CD62LhiCD127hi), central memory cells (Tcm, CD45RA-CD45RO+CD62L+), T effector memory cells (Tem, CD45RA-CD45RO+CD62L-), peripheral memory cells (Tpm, CD62L+CD127hi), highly differentiated effector cells (Temra, CD45RA+CD45RO-CD127loGZMBhi), and other memory cells (CD43loKLRG1hiCD127-) (Figures 10A and 10B).

[0164] 98.6% of pMHC-binding T cells were memory cells enriched in expanded T cell clones (Figure 10D), indicating that these T cells were selected by a specific immune response and therefore likely to be responsive and reliable binders. The majority of these memory T cells bound common viral epitopes (e.g., influenza, EBV, CMV), and CD8+ pMHC-binding T cells from each donor showed different distributions of memory cell subsets. For example, donor 1 had primarily Tpm and Tcm cells, while donor 2 had primarily Tem and Tpm cells, and donor 4 had primarily Temra cells (Figures 10C and 10D).

[0165] The majority of pMHC-binding T cells expressed a memory phenotype, but 1.3% of them were naive cells. These naive cells had more diverse pMHC interactions than non-naive cells and frequently bound endogenous antigens, tumor-associated antigens (e.g., MART-1), or antigens derived from viruses (e.g., HIV) that the donor had encountered (Figure 10C and Figure 20). Interestingly, the percentage of naive T cells with cross-HLA binding was significantly higher than that of non-naive cells (Figure 10E). These results suggest that healthy donor T cell repertoires, particularly naive cells, may respond to unencountered or rare antigens and retain cross-reactivity. Further assays are needed to assess whether these cells can support functional T cell responses.

[0166] 2. Essay We developed a method (Icon) capable of reliably identifying TCR-pMHC interactions by significantly increasing the signal-to-background ratio in highly multiplexed 10x Genomics TCR-pMHC binding data. Having appropriate controls (negative control dextramers and non-dextramer-sorted T cell samples) is essential for accurately estimating background noise, a factor that proved crucial for reliably identifying TCR-pMHC binding events. While ICON was developed on a single dataset consisting of a single pool of multiplexed dextramers, this method can be generalized to query pMHC-TCR binding data from a broader pool of pMHC dextramers as more multiplexed datasets are generated.

[0167] This study demonstrates the robustness of this CNN-based classifier in predicting TCR-pMHC-specific binding and demonstrates the potential for computational prediction to practically (vs. experimentally) study T cell antigen-specific recognition. Immunomonitoring of T cell antigen-specific recognition was applied to determine immune responses to specific antigens (e.g., tumor-specific antigens and peptide vaccines) and their potential correlation with clinical outcomes in patients undergoing immunotherapy. However, experimental mapping of TCR sequences to antigen specificities is costly and labor-intensive. With appropriate training data for specific pMHCs, the classifier presented here can assign a probability of pMHC binding to each TCR sequence of interest without performing binding assays. This study validates the multinomial prediction mode of this classifier (Figure 17B), which may be used to select highly specific TCRs for safe T cell-related therapies.

[0168] The results show that the majority (>30%) of TCRs that bind to a particular pMHC share one chain and differ in a second chain, indicating that T cell clonality must be determined by data using paired αβ chains. Furthermore, 8% of these TCRs that share a single chain can bind to different pMHCs. This is consistent with the ability to predict TCR antigen specificity using paired TCR chains, which is 16% higher than using either chain alone. Therefore, sequencing paired αβ chains from single cells is likely more powerful for accurately examining T cell repertoire clonality and TCR-pMHC binding specificity.

[0169] The ability to assess biologically relevant T cell reactivity is important for investigating and monitoring immune responses to pathogens and other disease states. We observed that the majority of recovered T cell reactivity (98.6%) matched the appropriate HLA type / supertype and, further, that the phenotype of multimer-positive cells was largely restricted to the memory T cell compartment, indicating that relevant memory reactivity from prior functional T cell responses can be resolved with this technology. Paired αβ TCR sequencing revealed multiple TCR sequences specific to individual multimers, enhancing broad antigen immune responses against common viral loads.

[0170] Although low-level HLA-mismatch reactivity was recovered, it was significantly enriched in unexpanded naive T cells compared with memory subsets, potentially revealing antigen-specific interactions against previously unexposed targets or those that did not culminate in functional T cell responses. Furthermore, a range of TCR avidity was recovered in these experiments, which may contribute to the detection of unexpected binding patterns. Dextramers are highly multimerizable, potentially detecting a broader range of TCR binding avidity than traditional tetramer reagents. Furthermore, the broad range of fluorescent dextramer intensities was screened for multimer-positive gating, capturing even low-frequency, lower-avidity TCR interactions in this highly sensitive single-cell assay.

[0171] 3. Method i.10× Genomics Single Cell Immune Profiling Dataset The 10x Genomics data used for this study was downloaded from support.10xgenomics.com / single-cell-vdj / datasets.

[0172] ii. Single cell RNA-seq data QC CD8+ cells from each donor were selected for downstream analysis by the following criteria: RNA signature count <=2500 and >200 genes detected per cell, and a mitochondrial percentage that was less than 40 percent of the total UMI (unique molecular identifier) ​​count.

[0173] iii.Classification of pMHC-binding T cells The Seurat V3 single cell sequencing analysis R package33, 34 was used for classification analysis based on single cell R NA-seq data. Because significant enrichment of TCR VJ gene usage was observed in identified pMHC-binding T cells, TCR genes were removed from the classification. Thus, cell clusters are not dominated by their shared VJ gene usage. All other gene expression in identified binding T cells was then normalized and quantified using Seurat V3 default parameters. PCA was normalized, transformed, and UMI counts were performed on variably expressed genes. The top 10 PCs were used for cell classification. UMAP was used for classification visualization (Figure 17).

[0174] iv. Generation of CDR3 motifs from the most predictable pMHC-binding TCR pairs The CDR3 amino acid sequences of the α and β chains from the 10 most predictive TCRs were aligned using COBALT (www.ncbi.nlm.nih.gov / tools / cobalt / cobalt.cgi). The aligned CDR3 amino acid sequences were input into WebLogo35 using default parameters to generate motifs.

[0175] v. Selection of reported pMHC specific binding pairs to TCR The raw files are available at VDJdb28 (vdjdb.cdr3.net / ) and The Data were downloaded from the Pathology-associated TCR database (friedmanlab.weizmann.ac.il / McPAS-TCR / ). Data were processed according to the following criteria to obtain pMHC TCR binding: for the VDJdb, paired α or β chain CDR3 amino acid sequences were required for each "complex.id," TCRs annotated with "source" were removed from 10x genomics, and data were filtered for "species" = "human." For McPAS-TCR, a known "epitope.ID" was required in the complete data, with "CDR3.alpha.aa" and "CDR3.beta.aa." Similarly, for the VDJdb, human TCRs were filtered.

[0176] vi. Normalization of TCR-pMHC binding data We developed the Integrated Context Specific Normalization (ICON) method. It takes multi-omics single-cell sequencing data generated from the 10x Genomics Immune Map platform as input data and performs TCR-pMHC binding specificity data normalization to identify reliable binding events. The multi-omics datasets include single-cell RNA-seq, paired αβ chain single-cell TCR-seq, dCODE-dextramar-seq, and cell surface protein expression sequencing, also referred to as CITE-seq (Cellular Index of Transcriptomes and Epitopes by Sequencing). ICON involves the following major steps (Figure 6A and Figure 12).

[0177] Single-cell RNA-seq-based filtering of low-quality cells. It filters low-quality cells such as doublets and dead cells. Cells with unexpectedly high numbers of genes detected (e.g., >2500 genes per cell) were classified as doublets, while cells with a high fraction of mitochondrial gene expression (e.g., ratio of mitochondrial gene expression UMI to total gene expression UMI >0.4) or too few detected genes (<200 genes per cell) were classified as dead cells (Figure 12A).

[0178] Single-cell dCODE-Dextramer-seq-based background control. Two types of background noise controls were designed for the dextramer binding assay and used in the analysis: one was a negative control dextramer (n = 6) derived from dextramer-stained and sorted CD8+ T cells (denoted as nc, NC_dex), and the other was dextramer-stained CD8+ T cells without dextramer sorting. To examine the signal-to-noise distribution, the maximum dextramer signal in each cell's UMI (unique molecular identifier) ​​was selected, representing the best binding of each cell. Specifically, the nonspecific dextramer binding signal of a cell was represented as Max(nc1,...,nc6), and the maximum dextramer signal of the six negative control dextramers included in the dextramer pool. The dextramer binding signals of cells from the dextramer stained and sorted samples (denoted as ds, Dex_sorted) were calculated as Max(ds1,...,ds), which is the maximum dextramer signal at the UMI of the 44 tested dextramers. 44 Similarly, the dextramer binding signal of cells from non-Dex-sorted samples is expressed as Max(du,...,du 44 The distribution of these three dextramer signals before ICON processing is shown in the upper panel of Figure 12B. The P of nonspecific dextramer binding signals in UMI 99.9 (absolute outliers in the negative dextramer control were excluded) was selected as the non-specific dextramer binding cutoff for each donor.

[0179] To estimate the noise potentially introduced by the cell sorting process, we compared cumulative analysis of dextramer binding signals between Dex_sorted and non-Dex_sorted samples to determine a cutoff for dextramer sorting efficiency (Figure 12C). For each donor, the Kolmogorov-Smirnov test (KS test) p-value was calculated by comparing cumulative curves of dextramer-sorted and non-dextramer-sorted samples using each data point (dextramer UMI) as a sliding window. A sigmoidal decreasing p-value curve indicates enrichment of dextramer binding signals in dextramer-sorted samples compared to non-dextramer-sorted samples, whereas a V-shaped curve suggests a loose cell sorting gate (Figure 12D). The dextramer UMI, which defines the maximum difference in dextramer binding signal between Dex_sorted and Dex_unsorted (argmax D_(s,u)), was used as a threshold to estimate the dextramer sorting efficiency for V-shaped samples. Finally, the background noise of dextramer-sorted samples was defined as: d=maximum(P 99.9 ,argmaxDs,u)

[0180] The dextramer signal (UMI) for each of the 44 tested dextramers in sorted cells was corrected by subtracting the estimated background (Figure 12E): E c =E s -d

[0181] Cell-wise normalization was then performed based on the log-rank distribution for each cell. pMHC-wise normalization was performed to make the dextramer binding signals comparable to each other. The adjusted dextramer binding signals of the sorted cells, Ec, were normalized across the 44 tested dextramers and then normalized across all cells according to the following equation: Ec^'>=0.9 was empirically selected as the cutoff for pMHC-specific binders (Figure 12F).

number

[0182] Selection of T cells with a single pair of αβ chains based on single-cell TCR-seq. T cells with only α chains, only β chains, and multiple α or β chains were removed. Only T cells with a single pair of αβ chains were used in this study.

[0183] The ICON normalization process was performed separately for each donor.

[0184] vii. Antigen-specific T cell expansion and antigen re-exposure to identify MART-1-binding T cells HLA A * Peripheral blood mononuclear cells (PBMCs) from the 02:01 individual were isolated by Ficoll-Paque Plus gradient isolation. PBMCs were seeded onto culture plates in T cell medium (CellGenix Dendritic Cell Medium, Cat. No. 20801-0500 + 5% human serum AB (Sigma, Cat. No. H3667)) + 1% penicillin / streptomycin / L-glutamine (ThermoFisher, Cat. No. 10378-016), 5 ng / ml of the T cell accessory cytokines IL-7 and IL-15 (CellGenix, Cat. Nos. 1410-050 and 1413-050, respectively), and 10 U / ml of IL-2 (Peprotech, Cat. No. 200-0), and 10 μg / ml of the A*02:01-restricted MART-1 epitope ELAGIGILTV (Genscript). Cultures were fed fresh medium and cytokines every 2 days for 1 week. On day 7 of culture, cells were transfected with fluorescently labeled dextramer HLA-A. *Antigen-specific CD8+ T cell expansion was assessed by flow cytometry following staining with 02:01 MART-1 ELAGIGILT (Immudex, Cat. No. WB2162-PE). For antigen rechallenge assays, peptides were added to T cell expansion cultures after 7 days of expansion. 24 hours after restimulation, cells were harvested and stained with fluorescently labeled antibodies for CD3 (BD ​​Biosciences, Cat. No. 612750), CD8 (BD Biosciences, Cat. No. 612889), CD69 (BD Biosciences, Cat. No. 564364), CCR7 (Biolegend, Cat. No. 353218), CD45RO (Biolegend, Cat. No. 304238), CD137 (Biolegend, Cat. No. 309828), and CD25 (Biolegend, Cat. No. 356104). Fluorescence-activated cell sorting (FACS) was performed using an Astrios cell sorter (Beckman Coulter) with gating on the forward scatter plot, side scatter plot, and fluorescence channel to select live cells while excluding debris and doublets. Single CD3+CD8+CD45RO+CD137+ cells were selected for further processing using a 100 μm nozzle.

[0185] The sorted cells were then loaded onto Chromium Single Cell 5' chips (10x Genomics, catalog no. 10x Genomics) and processed through the Chromium Controller to generate GEMs (gel beads in emulsion). RNA-Seq libraries were prepared using the Chromium Single Cell 5' Library & Gel Bead Kit (10x Genomics, catalog no. 10x Genomics) according to the manufacturer's protocol.

[0186] viii. Regeneron oligo-tagged dextramers staining and sorting for 10x Genomics donor 3 and donor 4 10xGenomics kindly provided cryopreserved PBMCs from donors 3 and 4 for use in reassessing CD8+ T cell dextramer binding capacity. CD8+ T cells were enriched using Miltenyi CD8+ T cell negative enrichment (Mitenyi). Cells were then incubated with benzonase (Millipore) and dasatinib (Axon) for 45 minutes, followed by staining with oligo-tagged dextramer pool (Immudex, Figure 21) for 30 minutes at room temperature. Cells were then transfected with CD3 (BD Cells were stained for 30 minutes on ice using fluorescently labeled CD4 (BD Biosciences, catalog number 612750), CD4 (BD Biosciences, catalog number 563919), CD8 (BD Biosciences, catalog number 612889), CCR7 (Biolegend, catalog number 353218), and CD45RO (Biolegend, catalog number 304238) and CITE-seq antibodies. Viable cells were selected using an Astrios cell sorter (Beckman Coulter) by fluorescence-activated cell sorting (FACS) gating on the forward scatter plot, side scatter plot, and fluorescence channel, while excluding debris and doublets. Single CD3+CD8+Dextramer+ cells were sorted using a 100 μm nozzle for further processing (Figure 11).

[0187] Recently, distance-based classification of TCR sequence similarity was reported. This method, TCRdist, is a weighted distance-based method for predicting TCR-pMHC binding specificity based on the sequence space of TCR CDR regions guided by structural information about pMHC binding. The nearest neighbor (NN) distance (the average TCRdist between a receptor and its nearest neighbors in the repertoire) was further calculated to measure receptor density within the repertoire. For each pMHC repertoire, a binder was defined as a TCR that binds to a given pMHC. NN distances were calculated between each binding TCR and each set of pMHC binders that removed the given TCR. NN distances were separated based on the known specificity of each TCR. For each binary classifier of pMHC, receiver operating characteristic (ROC) curves and areas under the ROC curve (AUC) were calculated using the plotROC R package. Briefly, ROC curves were generated by calculating the sensitivity and specificity at several NN distance thresholds for each classifier, where TCRs would be classified as binding to a given pMHC if their NN distance was below the given threshold.

[0188] ix. CNN-based classification A weighted binary classifier is adapted based on a deep learning framework, which includes three major steps with adjustments to meet specific needs.

[0189] x. Input data formatting TCR sequencing files were collected as raw formatted files in 10x Genomics. The sequencing files were analyzed to obtain the amino acid sequences of the CDR3 after removing non-productive sequences. Clones with different nucleotide sequences but the same matching amino acid sequences from the CDR3 and V, D, and J genes were aggregated together under one TCR. Therefore, each TCR record used here contains the amino acid sequences of a single pair of α and β TCRs, CDR3, V, and J genes. For model runs using only the α chain TCRB-CDR3 amino acid sequence, the β chain gene was removed from the input. Similar removal was performed for the β chain-only model.

[0190] xi. Data conversion Each TCR-CDR3 amino acid sequence was coded with numbers representing 20 possible amino acids. Only sequences that matched the IUPAC (International Union of Pure and Applied Chemistry) amino acid code were retained. For TCRs of different lengths, zero padding was applied to a maximum length of 40. Features were further extracted from the amino acid sequences using a trainable embedding layer. V and J genes were one-hot coded to provide a taxonomic and separate representation of gene names in the computational space. The coded sequence and gene name were concatenated together to represent a single TCR record. This data conversion process was applied before training all networks.

[0191] xii. Single TCR sequence classifier We adapted this method to provide a generalized, conventional neural network architecture for training TCRs, focusing on sample- or repertoire-level prediction. We focused on optimizing single TCR sequence prediction. To achieve this, we removed T cell clone size from the input data. Furthermore, we applied a single translation-invariant layer to the sequence, followed by three fully connected convolutional layers for the final output. The network was modeled using the Adam The network was trained using Optimizer (learning rate = 0.001) to minimize the cross-entropy loss between the soft-maximum logarithm and one-hot coded representations of the network's separate categorical outputs. This approach was modified by using a biologically meaningful kernel size 439 to capture potential motifs. To account for imbalanced class representation in the training data, a weighted cross-entropy loss function was applied using the following formula:

number

number

[0192] Monte Carlo cross-validation (MCCV) training was performed by retaining a fixed number of TCRs for validation and testing, respectively. An early stopping algorithm was implemented using the validation set of sequences, where Monte Carlo sampling was performed with 20 iterations. Receiver operating characteristic (ROC) curves for the sequence classifiers were calculated based on the test set after averaging all MCCV predictions.

[0193] B. Example 2 1.Results i. Identification of pMHC-specific binding TCRs from high-throughput binding data 10x Genomics recently generated a scalable, publicly available TCR-pMHC binding dataset. In their initial report, they assessed the binding properties of over 150,000 CD8+ T cells from four HLA-haplotyped healthy donors (Table 1, donors 1–4) across 44 pMHC dextramers using the single-cell-based immune profiling platform ImmuneMap to directly detect antigen binding to T cells, while simultaneously sequencing T cell αβ chain pairs and transcriptomes (Figure 2). The dextramer pool spanned eight HLA alleles and consisted of epitopes with known common viral and cancer responses (Table 2). [Table 2] [Table 3-1] [Table 3-2] [Table 3-3]

[0194] Here, we describe a highly multiplexed dextramer binding dataset generated at the single-cell level using paired T cell α and β chain sequences. 10x Genomics applied exhaustive cutoffs for background noise and nonspecific dextramer binding to all donors and dextramers to identify pMHC-binding TCRs (18). Not surprisingly, we found an unexpectedly large number of promiscuous TCR-pMHC binding events in the data provided by 10x Genomics (Figure 24). To robustly identify reliable binding events from such high-throughput TCR-pMHC binding data, we developed ICON (Figure 25A, Figure 26A-D and Materials and Methods). ICON data processing was performed in a donor-, cell-, and dextramer-specific context. Briefly, single-cell transcriptome data was used to select good-quality cells (live and singleton). Negative control dextramers (n = 6) were then used to empirically estimate background binding noise for each donor. The raw dextramer binding signals were then corrected by subtracting the estimated background noise for each donor separately. As previous studies have shown that paired αβ chains synergistically mediate TCR-pMHC recognition, T cells with paired αβ chains were selected as candidates for pMHC-binding T cells. The T cell dextramer binding signals were further corrected by penalizing dextramers that simultaneously bound to the same T cell / clone. Finally, the dextramer binding signals were normalized across cells and MHCs to directly equate them (Figure 25A, Figure 26A-D, and Methods). To evaluate the performance of ICON, the pMHC-binding specificity of CD8+ T cells was assessed from another healthy donor (donor V) using the same dextramer panel (Figure 27 and Materials and Methods). ICON was able to engage 91% of sequenced T cells with paired αβ chains to their antigen targets. To estimate the specificity of ICON, 21 individual dextramer binding essays were performed using T cells from the same donor, donor V (ee and Materials and Methods).Flow cytometry results show the relative abundance of T cells binding to these 21 Dextramers identified from ICON (Figure 25C).

[0195] Applying ICON, we identified a total of 53,062 CD8+ T cells belonging to 5,721 unique T cell clones that bound 37 pMHCs from five donors (Figure 25B, Figure 29). This suggests that TCR dual specificity (specificity vs. degeneracy) is a key feature of the immune response mechanism, significantly distinguishing self from foreign peptides to avoid autoimmune reactions while maintaining broad antigen coverage. Indeed, 99.6% of unique TCRs bind to one specific pMHC, while the remaining TCRs interact with two pMHCs (Figure 25B). Furthermore, these TCR-pMHC interactions generally follow an HLA-type-specific pattern. 94% of binding events were HLA-matched, of which 6% were HLA-matched with HLA A antigens that share similar primary anchor positions of the presented peptides. * 03-Supertype family member HLA A * 03:01 and A * The most common HLA haplotypes (A) in the dextramer pool (Tables 1 and 2) * Donors 1 and 2, with their respective TCR-pMHC binding patterns (02:01), shared a significant fraction (n=44) of unique TCR-pMHC interactions (Figure 25D, G), supporting the notion that TCR-pMHC binding patterns are most HLA-restricted. However, 6% of binding events are cross-HLA type interactions. HLA-type mismatched binding T cells tend to have smaller clones or be singletons (antigen naive).

[0196] Of all pMHC-binding TCRs, 99% of total TCRs (96% of unique TCRs) are pMHC:B * 08:01_RAKFKQLL_BZLF1_EBV (T cell count: 18,468 / unique TCR count: 479), A *02:01_GILGFVFTL_Flu-MP_Influenza (T cell count: 8,365 / unique TCR count: 1,095), A * 11:01_IVTDFSVIK_EBNA-3B_EBV (T cell count: 5,438 / unique TCR count: 149), A * 03:01_KLGGALQAK_IE-1_CMV (T cell count: 3,899 / unique TCR count: 2,865), A * 11:01_AVFDRKSDAK_EBNA-3B_EBV (T cell count: 1,579 / unique TCR count: 95), A * 02:01_GLCTLVAML_BMLF1_EBV (T cell count: 1,886 / unique TCR count: 117), A * 02:01_ELAGIGILTV_MART-1_cancer (T cell count: 297 / unique TCR count: 293), B * 35:01_IPSINVHHY_pp65_CMV (T cell count: 6,986 / unique TCR count: 280) and A *02:01_NLVPMVATV_pp65_CMV (T cell count: 5,612 / unique TCR count: 164) (Figure 25E). To further understand the characteristics of the conserved TCR sequences underlying the classification, TCR VJ gene usage was examined for these nine pMHC repertoires. In addition to the enrichment reported in previous studies, such as TRBV19 and TRAV27 in the influenza repertoire, TRAV5 and TRBV20-1 in the BMLF1_EBV repertoire, and TRBV6-5 in NLVPMVATV_pp65_CMV, we also found TRAV12-2 in the MART-1_cancer repertoire, TRAV21, TRAV35, TRBV11-2, and TRBV6-6 in the IVTDFSVIK_EBNA-3B_EBV repertoire, and AV We found extensive usage of TRAV8-3, TRAV13-1, and TRBV28 in FDRKSDAK_EBNA-3B_EBV, TRAV13-1, TRAV13-2, and TRBV12-3 in the BZLF1_EBV repertoire, TRAV12-1, TRAV41, TRBV2, and TRBV20-1 in IPSINVHHY_pp65_CMV, and TRAV23 / D6 and TRBV12-4 in NLVPMVATV_pp65_CMV (Figure 25F). Consistent with conserved VJ gene usage, Shannon diversity indices and TCR clone size distributions suggested that each pMHC-binding T cell repertoire underwent different degrees of expansion in response to their target peptides (Figures 30A and B).

[0197] ii.TCRAI: A neural network classifier index of T cell antigen specificity With the large and diverse TCR-pMHC binding events identified, robust functional classifiers for rapid validation of these binding events are desirable. Recent studies have demonstrated that neural networks (CNNs) can learn high-dimensional information from TCR sequences and thus robustly predict TCR-pMHC binding.

[0198] The Python package, TCRAI, was developed using TensorFlow 2 and provides a flexible framework for the study of TCR-pMHC specificity (Figure 31A). The highly modular nature of the TCRAI package allows for easy adjustment of model construction. Briefly, the TCRAI framework works as follows: Any number of V(D)J genes and CDR regions of TCRs can be defined as inputs to the model in textual format. A method for processing these inputs into numeric form in an unlearnable manner can be selected via a "processor" object, which converts text to numeric representations. These numeric inputs can then be further processed in a learnable manner via an "extractor" object, which forms the building blocks of the neural network and provides their output vector representation of the input data, called fingerprints. These fingerprints are concatenated via a single numeric vector into a single TCRAI fingerprint that describes the input TCR. This TCRAI fingerprint is then passed through a "closer" object, which forms the final building block of the neural network, to generate a prediction on the input TCR. The TCRAI package provides several such pre-built processors, extractors, and closers, and is easily extensible to new variants: it allows performing binomial, polynomial, regression, or other tasks simply by choosing to build a different closer object.

[0199] To evaluate the performance of TCRAI, we conducted a literature search of currently available methods (Table 3) and compared its classification index with four leading methods in this field: GLIPH2, DeepTCR, NetTCR, and TCRdist. For comparison, eight pMHC-specific binding T cell repertoires were collated with at least 50 unique paired αβ chain TCRs generated by traditional single multimer binding assays or antigen rechallenge assays as the gold standard dataset (Table 4 and Materials and Methods). DeepTCR, NetTCR, and TCRdist are predictive models like TCRAI. The area under the receiver operator characteristic (ROC) curve (AUROC / AUC), a standard measure of classification success for these predictive models, indicates that TCRAI and DeepTCR, which have similar neural network frameworks, perform better than TCRdist and NetTCR. Overall, TCRAI performs more consistently and better than DeepTCR (Figures 31e and 32B). Because GLIPH2 was designed to cluster TCR sequences into distinct groups of specificities, the sensitivity and specificity (calculated at a model threshold that maximized the two geometric means) of these four predictive models were measured for comparison with GLIPH2. The comparison results showed that TCRAI had the best balanced sensitivity and specificity (Figure 33). Some methods with objectives different from those of TCRAI were not included in the comparison. For example, ALICE is designed to detect groups of homologous / expanded TCRs. TcellMatch uses cell-specific covariates (e.g., gene expression) rather than just TCR sequences as input, and its performance was tested on 10x Genomics immune map data at a high signal-to-noise ratio without further refinement.

[0200] [Table 4] [Table 5]

[0201] iii. Classification of pMHC-binding TCRs identified from high-throughput data Next, we applied TCRAI to the nine most abundant pMHC-binding repertoires, ICONs, identified from the high-throughput data (Figure 25E). TCRs in these nine pMHC repertoires were classified with TCRAI in binomial mode, with an average AUC of 0.88. Similar predictive performance was also observed using TCRAI in multinomial mode (Figures 34A and 35; hereafter, TCRAI results are from predictive performance unless otherwise specified). Historically, TCR β chain sequencing has often been used to infer T cell antigen-binding specificity due to its higher complex potential compared to the α chain. To quantitatively assess the contribution of TCR α and β chains in predicting TCR-pMHC interactions, either the α or β chain was used as input to TCRAI instead of the paired αβ chain. Performance using paired αβ chains was better than either the α or β chains alone, with an average increase in AUC of 0.2 (Figure 34B). Consistent with previous studies, these results collectively demonstrate the importance of αβ pairing for accurate inference of TCR-pMHC interactions. The predictive performance of the β chain was not necessarily better than that of the α chain, indicating the importance of the α chain in TCR-pMHC-specific recognition, which has often been overlooked previously.

[0202] To further validate the performance of TCRAI, we analyzed four pMHC repertoires (A * 02:01_ELAGIGILTV_MART-1, A * 02:01_GILGFVFTL_Flu-MP, A * 02:01_GLCTLVAML_BMLF1_EBV and A * 02:01_NLVPMVATV_pp65_CMV) was used. TCRAI was trained using four repertoires identified from the high-throughput dataset and predicted four curated repertoires. Figure 34C shows that the prediction results were generally comparable to the performance on the training set. However, A *The performance of TCRAI when inferring on 02:01_NLVMVATV_pp65_CMV was significantly worse than the other three pMHCs. To understand the difference in performance, we examined the TCRAI fingerprint space of the model (Materials and Methods). * In the case of 02:01_ELAGIGILTV_MART-1_cancer and two other pMHC (Figure 36A), the binding TCRs from the high-throughput and curated datasets overlap spatially in fingerprint space, while the overlap is significantly worse for pp65_CMV (Figures 34D and 36B). This poor overlap is due to 98.2% of the pp65_CMV-binding TCRs in the high-throughput dataset coming from a single donor (Figure 29), thereby representing a small subspace of possible binding TCRs, while the public data contains TCRs from a range of donors representing a larger range of TCR space. This result also highlights the importance of diverse datasets for training robust TCR antigen prediction models.

[0203] iv. Characterization of pMHC-specific TCRs To investigate the properties of TCRs that bind to a given pMHC, we analyzed how the TCRAI classifier model places TCRs within its fingerprint space (Materials and Methods). TCR fingerprints derived from the classifier model allow the discovery of specific groups of TCRs with conserved gene usage and CDR3 motifs. These groups often exhibit distinct binding capacities and distinct structural binding modes.

[0204] TCR A *Clustering 02:01_GILGFVTL_Flu-MP_Influenza leads to two well-separated clusters in TCRAI fingerprint space (Figure 37A). The constructed α and β-CDR3 motifs and gene usage indicate that cluster 0 has a strongly conserved xRSx motif in the β chain and TRB19 and TRAJ42 gene usage, while the smaller cluster 1 has the highly conserved gene usage TRBV19 / TRBJ1-2 / TRAV38-1 / TRAJ52 (Figure 37C). The dextramer signal (unique molecular identifier in UMI) distribution showed that TCRs in cluster 0 have stronger binding to Flu dextramer than those in cluster 1 (Figure 37B). The results suggest that the A, which is thought to be linked to the "uncharacterized" pMHC complex, * This is consistent with the well-known strong conservation of CDR3 motifs and TCRBV19 gene usage in 02:01_GILGFVLTL_Flu-responsive T cells. Further comparison with the recently identified classes of A*02:01_GILGFVL_Flu-binding TCRs linked clusters 0 and 1 to their groups I (canonical) and II (novel), respectively. Previous studies have also found that group I TCRs have stronger binding than group II TCRs. The 3D structures of TCR-pMHC binding complexes proposed in the prior art suggest that these two TCR groups have different binding modes due to the highly conserved motifs / residues, thereby causing different Phe-5 ring rotations of the Flu peptide in these two complexes (Figure 37D).

[0205] The TCRs binding to eight other pMHCs were also characterized. *The results for the 02:01_GLCTLVAML_BMLF1_EBV-binding TCR are particularly interesting. Previous studies have observed a dominant open TCR constructed from TRBV20-1 / TRBJ1-2 / TRAV5 / TRAJ31. However, previous analyses of the TCR population binding to this pMHC focused on the highly populated TRAV5 TCR. The current experiments unbiasedly identified five clusters of TCRs within the TCRAI fingerprint space (Figure 37E). Clusters 1 and 2 represent classical HLA*02:01_GLCTLVAML-published TCRs, but the two clusters are divided based on their β-chain gene usage (Figure 37G). Cluster 0 contains TCRs following gene usage (TRBV2 / TRBJ2-2) and a β-chain CDR3 motif not present elsewhere. TCRs belonging to this novel group exhibited different binding capacities relative to the standard TCR clusters (clusters 1 and 2), as evidenced by reduced dextramer UMI numbers (Figure 37F), indicating lower affinity and partly explaining why this TCR group has not yet been recognized.

[0206] Immunophenotype of v.pMHC-binding CD8+ T cells. Combined information on antigen specificity and T cell phenotype has been reported to be important for the clinical success of immunotherapies such as vaccination. Multi-omics data generated by the ImmuneMap platform allow us to link T cell antigen specificity with T cell phenotype. Using gene (single-cell RNA-seq) and surface protein (CITE-seq, Cellular Index of Transcriptome and Epitopes by Sequencing) expression data from this multi-omics dataset, we grouped pMHC-binding CD8+ T cells into subpopulations (Figure 38A and Materials and Methods). The identified subpopulations were then annotated according to previously described CD8+ T cell subtype marker genes: naive cells (CD45RA+CD62L+CD127hi), central memory cells (Tcm, CD45RA-CD62L+CD127+EOMEShighTBETlow), T effector memory cells (Tem, CD45RA-CD62LlowCD127+GZMB+), peripheral memory cells (Tpm, CD62L+CD127hiGZMB+), highly differentiated effector cells (Temra, CD45RA+CD127loGZMBhi), and other memory cells (CD43loKLRG1hiCD127-) (Figure 38A and B).

[0207] Ninety-six percent of pMHC-binding T cells were memory cells enriched in expanded T cell clones (Figure 38E and D), indicating that these T cells were selected by a specific immune response and therefore likely to be responsive and reliable binders. The majority of these memory T cells bound common viral epitopes (e.g., influenza, EBV, CMV), and pMHC-binding T cells from each donor showed different distributions of memory cell subsets. For example, donors 1 and 2 contained primarily Tpm, while donor V contained primarily Tem, and donors 3 and 4 contained primarily Temra cells (Figure 38C and D).

[0208] The majority of pMHC-binding T cells expressed a memory phenotype, but 4% of them were naive cells. These naive cells had more diverse pMHC interactions than non-naive cells and frequently bound tumor-associated antigens (e.g., MART-1), endogenous antigens, or antigens derived from viruses (e.g., HIV) that the donor had encountered (Figure 38C). Interestingly, the percentage of naive T cells with cross-HLA binding was significantly higher than that of non-naive cells (Figure 38F). These results suggest that healthy donor T cell repertoires, particularly naive cells, may respond to unencountered or rare antigens and retain cross-reactivity. Further assays are needed to assess whether these cells can support functional T cell responses.

[0209] 2. Essay High-throughput TCR-pMHC binding data offers an attractive route to furthering our understanding of TCR antigen recognition. However, this type of data is often associated with a high signal-to-noise ratio. Here, we present a framework of computational tools, including a novel method, ICON, that can identify reliable TCR-pMHC interactions by significantly increasing the signal-to-noise ratio in highly multiplexed TCR-pMHC binding data with excellent sensitivity and specificity. ICON calculates noise-corrected dextramer signals in a parameter-free manner, making it easily generalizable to pMHC-TCR binding data from broader pMHC dextramer pools and potentially extensible to normalizing protein binding signals in single-cell space, such as with CITE-seq.

[0210] In this study, we developed a Python package, TCRAI, that demonstrates the robustness of deep learning classifiers in predicting TCR-pMHC specific binding. Due to the importance of the CDR3 region in determining TCR specificity for a given antigen, it is tempting to build predictive models utilizing only this information, as others have. However, due to highly conserved gene usage for many pMHCs, we find that VJ gene usage is a key predictor of TCRAI, especially for the small number of unique pMHC-binding TCRs in our dataset. We observed that the predictive performance of models receiving CDR3 information outperformed gene-level-only models when the number of pMHC-binding TCRs was on the order of at least 100 (Figure 39), indicating that these models require this volume of data to extract useful sequence motifs from CDR3s.

[0211] We demonstrated that TCRAI can not only perform state-of-the-art classification of TCR-pMHC specific binding, but also identify groups of TCRs with distinct binding properties. Combining dextramer UMIs with TCR sequence information enabled investigation of the differential binding capacities between these groups. This finding demonstrates that as the amount of high-throughput TCR pMHC binding data increases, so too will our ability to discover new TCR motifs and combine them not only with UMIs but also with broader multi-omics data. For example, the ability to investigate differential transcriptional regulation of T cell receptor signaling between groups of TCRs with distinct binding mechanisms is highly exciting not only for broad scientific questions but also for the development of T cell therapeutics.

[0212] T cell antigen-specific recognition can potentially be studied virtually (rather than experimentally) using the TCRAI. Immune monitoring of T cell antigen-specific recognition has been applied to determine immune responses to specific antigens (e.g., SARS-COV2, tumor-specific antigens, and peptide vaccines) and their potential correlation with clinical outcome disease severity in patients undergoing immunotherapy. However, experimental mapping of TCR sequences to antigen specificity is costly and labor-intensive. With appropriate training data for specific pMHCs, the TCRAI classifier presented herein can assign pMHC binding probabilities to each TCR sequence of interest without performing binding assays. This study validated the multinomial prediction mode of this classifier (Figure 35), implying that it can be used to select highly specific TCRs for safe T cell-related therapies.

[0213] The ability to assess biologically relevant T cell reactivity is important for investigating and monitoring immune responses to pathogens and other disease states. The majority (94%) of recovered T cell reactivity matched the appropriate HLA type / supertype, and furthermore, the phenotype of multimer-positive cells was largely restricted to the memory T cell compartment, indicating that relevant memory reactivity from prior functional T cell responses can be resolved with this technology. Paired αβ TCR sequencing revealed multiple TCR sequences specific to individual multimers, enhancing broad antigen immune responses against common viral loads.

[0214] Although low-level HLA-mismatch reactivity was recovered, it was significantly enriched in unexpanded naive T cells compared with memory subsets, potentially revealing antigen-specific interactions against previously unexposed targets or those that did not culminate in functional T cell responses. Furthermore, a range of TCR avidity could be recovered in these experiments, potentially contributing to the detection of unexpected binding patterns. Dextramers are highly multimerizable, potentially detecting a broader range of TCR binding avidity than traditional tetramer reagents. Furthermore, because a wide range of fluorescent dextramer intensities was sorted using multimer-positive gating, even low-frequency, low-activity TCR interactions were captured in this highly sensitive single-cell assay.

[0215] 3. Materials and Methods i.10× Genomics Single Cell Immune Profiling Dataset The 10x Genomics data used for this study was downloaded from support.10xgenomics.com / single-cell-vdj / datasets.

[0216] ii. Identification of pMHC-binding T cell phenotypes The Seurat V3 single cell sequencing analysis R package was used for classification analysis based on single cell R NA-seq data. Because significant enrichment of TCR VJ gene usage was observed in identified pMHC-binding T cells, TCR genes were removed from the classification. Thus, cell clusters are not dominated by their shared VJ gene usage. All other gene expression in identified binding T cells was then normalized and quantified using Seurat V3 default parameters. PCA was normalized, transformed, and UMI counts were performed on variably expressed genes. The top 10 PCs were used for cell classification. UMAP was used for classification visualization.

[0217] iii. Selection of reported pMHC specific binding pairs to TCR Raw files were downloaded from the VDJdb (42) (vdjdb.cdr3.net / ) and The Pathology-associated TCR database (friedmanlab.weizmann.ac.il / McPAS-TCR / ). Data were processed to obtain pMHC TCR binding according to the following criteria: for VDJdb, paired α or β chain CDR3 amino acid sequences were required for each "complex.id." TCRs annotated with "source" were removed from 10x Genomics and filtered for "species" = "human." For McPAS-TCR, a known "epitope.ID" was required in the complete data, with "CDR3.alpha.aa" and "CDR3.beta.aa." Similarly, for VDJdb, human TCRs were filtered.

[0218] iv. Normalization of high-throughput TCR-pMHC binding data To identify reliable TCR-pMHC interactions, we developed ICON, an integrated context-specific normalization method. It takes multi-omics single-cell sequencing data generated from multiplexed multimer binding platforms, such as the 10x Genomics immune map, as input data, including single-cell RNA-seq, single-cell TCR-seq of paired αβ chains, dCODE-dextramer-seq, and cell surface protein expression sequencing, also known as CITE-seq. ICON involves the following major steps (Figure 25A and Figure 26):

[0219] Step 1: Single-cell RNA-seq-based filtering of low-quality cells.

[0220] It filters out low-quality cells such as doublets and dead cells. T cells with an unexpectedly large number of genes (e.g., >2500 genes per cell) were classified as doublets, and cells with a high fraction of mitochondrial gene expression (e.g., ratio of mitochondrial gene expression to total gene expression >0.2) or too few detected genes (<200 genes per cell) were classified as doublets (Figure 26A).

[0221] Step 2: Single-cell dCODE-Dextramer-seq-based background estimation

[0222] Six negative control dextramers were designed to estimate background noise from multiplexed dextramer binding assays. To examine signal and noise distribution, the maximum dextramer signal in the UMI (unique molecular identifier) ​​of the negative control dextramer and test dextramer for each cell was used to represent the worst noise and best dextramer for each T cell. The density distribution of these two types of dextramer signals is shown in Figure 26B. The background cutoff (gray dashed line in Figure 26B) was empirically selected for each donor.

[0223] Step 3: Selection of T cells with paired αβ chains based on single cell TCR-seq.

[0224] T cells with only a single chain were removed. For T cells with multiple α or β chains detected, the one with the highest UMI count was assigned to each T cell.

[0225] Step 4: Dextramar signal correction

[0226] Although each dextramer has its own optimal binding conditions, it is not possible to arrange experimental conditions such that multiplexed dextramer binding assays are optimal for each dextramer. This results in multiple dextramers binding to the same T cell / clone, as observed in this high-throughput dataset (Figure 26C). To correct for this effect, the following technique was used to penalize dextramer signals when simultaneously binding to the same T cell / clone.

[0227] j th Binds to dextramer th The background noise-subtracted dextramer signal for T cells was measured using E ij Defining it as i th j about T cells th The fraction of the dextramer signal that was due to dextramer binding is further shown as follows:

number

[0228] i th TCR clonotype of T cells i and T_(k ij ) as clonotype k that binds to dextramer j i The number of T cells belonging to j th Dextramer-binding clonotype k i The fractions of T cells belonging to the genotype are as follows:

number

[0229] Using these quantities, the corrected dextramer signal is calculated as j th Binds to dextramer th The T cells are calculated as follows: S ij =E ij (RC ij ) 2 RTkj

[0230] Step 5: Cell- and pMHC-wise dextramer signal normalization and binder identification

[0231] To equalize all dextramer binding signals, the corrected dextramer binding signals were log-ratio normalized across the 44 tested dextramers in the cells. Subsequently, pMHC-wise normalization was performed based on the log-rank distribution. A normalized dextramer UMI > 0 was empirically selected as the cutoff for pMHC-specific binders.

[0232] v. Regeneron Oligo-tagged Dextramer Staining and Sorting CD8+ T cells were enriched from healthy donor PBMCs using Miltenyi CD8+ T cell negative enrichment (Mitenyi). Cells were then incubated with benzonase (Millipore) and dasatinib (Axon) for 45 minutes, followed by staining with oligo-tagged dextramer pool (Immudex, see Table 2) for 30 minutes at room temperature. Cells were then stained for 30 minutes on ice with fluorescently labeled CD3 (BD ​​Biosciences, catalog no. 612750), CD4 (BD Biosciences, catalog no. 563919), CD8 (BD Biosciences, catalog no. 612889), CCR7 (Biolegend, catalog no. 353218), and CD45RA (Biolegend, catalog no. 304238) and CITE-seq antibodies. Viable cells were selected using an Astrios cell sorter (Beckman Coulter) by fluorescence-activated cell sorting (FACS) gating on the forward scatter plot, side scatter plot, and fluorescence channel, while excluding debris and doublets. A 100 μm nozzle was used to select single CD3+CD8+Dextramer+ cells for further processing.

[0233] vi. Construction of neural network-based classification index TCRAI TCRAI provides a flexible framework for the design of TCR classifiers, but uses a specific and consistent architecture throughout this work, which is described in detail below. Aside from its flexible architecture, some key differences from the DeepTCR architecture are the use of 1D convolution and batch normalization for CDR3 sequences, and a lower-dimensional representation for genes. These changes result in improved model normalization, allowing the model to learn stronger gene associations.

[0234] To process the TCR input information in numeric form, the following method was applied: For each CDR3 sequence, the amino acids are first converted to integers, and then these integer vectors are coded into one-hot representation. For the V and J genes, a dictionary of gene types to integers is constructed separately for each V and J gene, and these are used to convert each gene to an integer.

[0235] The neural network architecture applied to the processed input information includes an embedding layer and a convolutional network. Specifically, the processed CDR3 residues are embedded into a 16-dimensional space via a learned embedding, and the resulting numerical CDR3 is fed through three 1D convolutional layers using filters for dimension, kernel width, and stride length. Each convolution is activated by exponential linear unit activation, followed by dropout and batch normalization. After these three convolutional blocks, global max pooling is applied to the final feature, and this process encodes each CDR3 with a 256-length vector, the "CDR3 fingerprint." The processed gene input for each gene is one-hot coded via a learned embedding and embedded into a reduced-dimensional space (16 for V genes and 8 for J genes), thereby providing a "gene fingerprint" for each gene as a vector. All selected CDR3 and gene fingerprints are then concatenated into a single vector, the "TCRAI fingerprint." The TCRAI fingerprint is passed through one final fully connected layer to give binomial predictions (single output value, sigmoid activation), regression predictions (single output, no activation), or multinomial predictions (multiple output values, softmax activation). In this work, we focus on binomial and multinomial predictions.

[0236] TCR sequencing files were collected as raw formatted files in 10x Genomics. The sequencing files were analyzed to obtain the amino acid sequences of the CDR3 after removing non-productive sequences. Clones with different nucleotide sequences but the same matching amino acid sequences from the CDR3 and V, D, and J genes were grouped together under one TCR. Therefore, each TCR record used here contains a single pair of α and β TCR chains, with the CDR3 amino acid sequence and V and J genes for each chain.

[0237] The data was split into training (76.5%), validation (13.5%), and left-pruned test sets (10%) for each model, followed by 5-fold Monte Carlo cross-validation (MCCV) on the training set. Models were trained by minimizing the cross-entropy loss via the Adam optimizer, where the cross-entropy loss was calculated for each class as a weighted sum of 1 / (number of classes). * The weights are weighted by the fraction of samples in that class. To prevent overfitting, early stopping is tied to the left-pruned validation data set; in this case, the validation loss is increased more than five times, and the model stops training when the weights of the model with the smallest validation loss are restored. Due to the large number of models being trained here, only the learning rate and batch size are adjusted during cross-validation. After cross-validation, the optimal implementation of the hyperparameters is selected, and the model is retrained on the full training set, using the validation set to control early stopping. The retrained model is then evaluated on the left-pruned test set.

[0238] vii. TCRAI Fingerprint Analysis The TCRAI model generates both a prediction about TCRs that bind to a particular pMHC (or one of many pMHCs, in the multinomial case) and a "fingerprint," a numeric vector that describes the TCR within the context of the question of whether it can bind to that pMHC. The distribution of these fingerprints is analyzed to understand how the model works and to identify groups of TCRs with different binding modes. UMAP is used to reduce the fingerprints to a two-dimensional space. When using a model trained on one dataset to estimate fingerprints on another, unseen dataset, a UMAP projector is fitted using TCRs from the training dataset and the projector is used to transform TCRs from the unseen set.

[0239] When clustering TCR fingerprints, the fingerprints of all TCRs in the dataset are projected into two-dimensional space as described above, and then those TCRs that are strong true positives (STPs, binomial prediction >0.95) are selected. These STPs are then clustered in two-dimensional space using a k-means classifier. TCRs from within each cluster are then collected and used to construct CDR3 motif logos (using weblogo), gene usage, and UMI distributions by pairing unique TCR clonotypes within the cluster with all repeated clonotypes in the high-throughput data.

[0240] viii.DeepTCR modification The DeepTCR method was adapted to construct a binary classifier with the adjustments described below.

[0241] For each TCR recording, a single pair of α and β TCR chains was used, along with the CDR3 amino acid sequence and V and J genes for each chain only, along with the input provided to the TCRAI package. That is, clonality, MHC, or D gene usage were not included in the DeepTCR model. The final output layer was adjusted to give a single binomial output, and the model's hyperparameters were optimized for the problem at hand in the context of the DeepTCR framework.

[0242] FIG. 41 is a block diagram depicting an environment 4100 including non-limiting examples of a computing device 4101 (e.g., computing device 106) and a server 4102 connected through a network 4104. In one aspect, some or all steps of any described method can be performed on a computing device described herein. The computing device 4101 can include one or more computers configured to store one or more of sequence data 104 (e.g., single cell sequence data, dextramers sequence data, and single cell receptor sequence data), training data 410 (e.g., labeled receptor sequence data), ICON module 108, prediction module 110, etc. The server 1402 can include one or more computers configured to store the sequence data 104. Multiple servers 4102 can communicate with the computing device 4101 through the network 4104. In one embodiment, the server 1402 can comprise a repository for data generated by the single cell immune profiling platform 102.

[0243] In terms of hardware architecture, the computing device 4101 and the server 4102 may be digital computers that generally include a processor 4108, a memory system 4110, an input / output (I / O) interface 4112, and a network interface 4114. These components (4108, 4110, 4112, and 4114) are communicatively coupled via a local interface 4116. The local interface 4116 may be, for example, but not limited to, one or more buses or other wired or wireless connections known in the art. The local interface 4116 may have additional elements (omitted for simplicity) to enable communication, such as controllers, buffers (caches), drivers, repeaters, and receivers. Additionally, the local interface may include address, control, and / or data connections to enable appropriate communication between the aforementioned components.

[0244] The processor 4108 may be a hardware device for executing software, particularly that stored in the memory system 4110. The processor 4108 may be any custom-made or commercially available processor, a central processing unit (CPU), a coprocessor among several processors associated with the computing device 4101 and the server 4102, a semiconductor-based microprocessor (in the form of a microchip or chipset), or generally any device for executing software instructions. When the computing device 4101 and / or the server 4102 are in operation, the processor 4108 may be configured to execute software stored in the memory system 4110 to communicate data to and from the memory system 4110 and generally control the operation of the computing device 4101 and the server 4102 in accordance with the software.

[0245] The I / O interface 4112 can be used to receive user input from and / or provide system output to one or more devices or components. User input may be provided, for example, via a keyboard and / or mouse. System output may be provided via a display device and printer (not shown). The I / O interface 41412 may include, for example, a serial port, a parallel port, a small computer system interface (SCSI), an infrared (IR) interface, a radio frequency (RF) interface, and / or a universal serial bus (USB) interface.

[0246] The network interface 4114 can be used to transmit and receive from the computing device 4101 and / or the server 4102 over the network 4104. The network interface 4114 may include, for example, a 10BaseT Ethernet adapter, a 100BaseT Ethernet adapter, a LAN PHY Ethernet adapter, a Token Ring adapter, a wireless network adapter (e.g., WiFi, cellular, satellite), or any other suitable network interface device. The network interface 4114 may include address, control, and / or data connections to enable appropriate communication over the network 4104.

[0247] The memory system 4110 may include any one or combination of volatile memory elements (e.g., random access memory (RAM, such as DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, hard drives, tape, CD-ROM, DVD-ROM, etc.). Furthermore, the memory system 4110 may incorporate electronic, magnetic, optical, and / or other types of storage media. It should be noted that the memory system 4110 may have a distributed architecture in which various components are located remotely from each other but can be accessed by the processor 4108.

[0248] The software in memory system 4110 may include one or more software programs, each of which includes an ordered list of executable instructions for implementing logical functions. In the example of Figure 41, the software in memory system 4110 of computing device 4101 may include sequence data 104, training data 410, ICON module 108, prediction module 110, and a suitable operating system (O / S) 4118. In the example of Figure 41, the software in memory system 4110 of server 4102 may include sequence data 104 and a suitable operating system (O / S) 4118. Operating system 4118 essentially controls the execution of other computer programs and provides scheduling, input-output control, file and data management, memory management, and communication control, and related services.

[0249] For purposes of illustration, application programs and other executable program components, such as the operating system 4118, are illustrated herein as separate blocks, with the understanding that such programs and components may reside at various times in different storage components of the computing device 4101 and / or the server 4102. An implementation of the training module 220 may be stored on or transmitted through some form of computer-readable media. Any of the methods of the present disclosure may be performed by computer-readable instructions embodied on a computer-readable medium. A computer-readable medium may be any available medium that can be accessed by a computer. By way of example, and not intended to be limiting, computer-readable media may include “computer storage media” and “communications media.” “Computer storage media” may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Exemplary computer storage media may include RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disks (DVDs) or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer.

[0250] In one embodiment, the ICON module 108 and / or the prediction module 110 may be configured to perform method 4200, shown in FIG. 42 . Method 4200 may be implemented in whole or in part by a single computing device, multiple electronic devices, and the like. Method 4200 may include, in step 4201, receiving single cell sequence data, dextramer sequence data, and single cell T cell receptor (TCR) sequence data. The single cell sequence data may include RNA-seq data, the dextramer sequence data may include dCODE-dextramer-seq data, and the single cell T cell receptor (TCR) sequence data may include TCR-seq data.

[0251] The method 4200 may include, in step 4202, determining the number of genes for each cell represented in the dextramer sequence data based on the sequence data of a single cell.

[0252] Method 4200 may include removing data associated with cells with a number of genes outside a gene threshold range from the dextramer sequence data, in step 4203. By way of example, the gene threshold range may be from about 200 genes to about 2,500 genes.

[0253] The method 4200 may include, at step 4204, determining, for each cell represented in the dextramer sequence data, a fraction of mitochondrial gene expression based on the single cell sequence data.

[0254] Method 4200 may include, in step 4205, removing from the dextramer sequence data data associated with cells in which the fraction of mitochondrial gene expression exceeds a gene expression threshold. The gene expression threshold may be approximately 40 percent of the total unique molecular identifier count.

[0255] Method 4200 may include, in step 4206, determining based on the dextramer sequence data and the unselected dextramer sequence data. The selected dextramer sequence data may include selected test dextramer sequence data and negative control dextramer sequence data. The unselected dextramer sequence data may include unselected test dextramer sequence data.

[0256] Method 4200 may include, in step 4207, determining a maximum negative control dextramer signal for each cell represented in the dextramer sequence data based on the negative control dextramer sequence data. The maximum negative control dextramer signal is defined as (Max(nc1,...,nc n )) where n is the number of negative control dextramers.

[0257] Method 4200 may include, in step 4208, determining a maximum sorted dextramer signal based on the sorted test dextramer sequence data for each cell represented in the dextramer sequence data. The maximum sorted dextramer signal is defined as (Max(ds1,...,ds m )) where m is the number of dextramers tested.

[0258] Method 4200 may include, in step 4209, determining a maximum sorted dextramer signal based on the unsorted test dextramer sequence data for each cell represented in the dextramer sequence data. The maximum unsorted dextramer signal is defined as (Max(du,...,du m )) where m is the number of dextramers tested.

[0259] The method 4200 may include, in step 4210, estimating the dextramer binding background noise based on the maximum negative control dextramer signal. The dextramer binding background noise is calculated as (P 99.9 )

[0260] The method 4200 may include, in step 4211, estimating a dextramer sorting gate efficiency based on the maximum selected dextramer signal and the maximum unselected dextramer signal. The dextramer sorting gate efficiency is calculated as (argmaxD s,u ) The dextramer sorting gate efficiency can be expressed as (Max(ds1,...,ds m )) and (Max(du,...,du m )) may be determined as the maximum difference between

[0261] The method 4200 may include determining a measure of background noise based on the dextramer binding background noise and the dextramer sorting gate efficiency at step 4212. The measure of background noise may be represented as (d).

[0262] Method 4200 may include, in step 4213, for each cell represented in the dextramer sequence data, subtracting a measure of background noise from the dextramer signal associated with each cell. Subtracting the measure of background noise from the dextramer signal associated with each cell may include subtracting the measure of background noise from the dextramer signal associated with each cell by: (E c =E s -d).

[0263] Method 4200 may include, at step 4214, for each cell represented in the dextramer sequence data, performing cell-wise normalization on the dextramer signal associated with each cell. Performing cell-wise normalization may include:

number

[0264] Method 4200 may include, at step 4215, performing pMHC-wise normalization for each cell represented in the dextramer sequence data. Performing pMHC-wise normalization may include:

number

[0265] Method 4200 may include, in step 4216, determining the presence or absence of at least one alpha chain and at least one beta chain for each cell represented in the dextramer sequence data based on the TCR sequence data of the single cell.

[0266] Method 4200 may include, in step 4217, removing data from the normalized dextramer sequence data that are associated with cells having only an alpha chain, only a beta chain, or multiple alpha or beta chains based on the presence or absence of at least one alpha chain and at least one beta chain.

[0267] The method 4200 may include, in step 4218, identifying remaining data in the normalized dextramer sequence data that is associated with reliable TCR-pMHC binding events.

[0268] Method 4200 may further include training a predictive model based on data associated with reliable TCR-pMHC binding events, and may further include predicting the binding state of the newly presented receptor sequence with the trained predictive model.

[0269] In one embodiment, the ICON module 108 and / or the prediction module 110 may be configured to perform method 4300, shown in FIG. 43. Method 4300 may be performed in whole or in part by a single computing device, multiple electronic devices, and the like. Method 4300 may include, at step 4310, receiving single cell sequencing data including single cell sequence data, dextramer sequence data, and single cell T cell receptor (TCR) sequence data. The single cell sequence data may include RNA-seq data, the dextramer sequence data may include dCODE-dextramer-seq data, and the single cell T cell receptor (TCR) sequence data may include TCR-seq data.

[0270] Method 4300 may include, in step 4320, filtering data associated with low-quality cells from the Dextramar sequence data based on the sequence data of a single cell. Filtering data associated with low-quality cells from the Dextramar sequence data based on the sequence data of a single cell may include: for each cell represented in the Dextramar sequence data, determining a number of genes based on the sequence data of the single cell; removing data from the Dextramar sequence data associated with cells whose number of genes is outside a gene threshold range; for each cell represented in the Dextramar sequence data, determining a fraction of mitochondrial gene expression based on the sequence data of the single cell; and removing data from the Dextramar sequence data associated with cells whose fraction of mitochondrial gene expression exceeds a gene expression threshold. The gene threshold range may be between about 200 genes and about 2,500 genes. The gene expression threshold may be about 40 percent of the total unique molecular identifier count.

[0271] Method 4300 may include adjusting the dextramer sequence data based on the background noise measurement in step 4330. Method 4300 may further include determining selected dextramer sequence data based on the dextramer sequence data, wherein the selected dextramer sequence data includes selected test dextramer sequence data, negative control dextramer sequence data, and unselected dextramer sequence data, and wherein the unselected dextramer sequence data includes unselected test dextramer sequence data. Method 4300 may further include determining a maximum negative control dextramer signal for each cell represented in the dextramer sequence data based on the negative control dextramer sequence data, determining a maximum sorted dextramer signal for each cell represented in the dextramer sequence data based on the sorted test dextramer sequence data, and determining a maximum unsorted dextramer signal for each cell represented in the dextramer sequence data based on the unsorted test dextramer sequence data. The maximum negative control dextramer signal is defined as (Max(nc1,...,nc n )), where n is the number of negative control dextramers. The maximum selected dextramer signal may be expressed as (Max(ds1,...,ds m )), where m is the number of dextramers tested. The maximum unsorted dextramer signal may be expressed as (Max(du,...,du m )) where m is the number of dextramers tested.

[0272] Adjusting the dextramer sequence data based on the measured background noise can include estimating the dextramer binding background noise based on the maximum negative control dextramer signal, estimating the dextramer sorting gate efficiency based on the maximum selected dextramer signal and the maximum unselected dextramer signal, determining a measured background noise (d) based on the dextramer binding background noise and the dextramer sorting gate efficiency, and subtracting, for each cell represented in the dextramer sequence data, the measured background noise from the dextramer signal associated with each cell. The measured background noise may be expressed as (d). Subtracting the measured background noise from the dextramer signal associated with each cell can be expressed as (E c =E s -d). Method 4300 may further include normalizing the dextramer sequence data. Normalizing the dextramer sequence data may include, for each cell represented in the dextramer sequence data, performing cell-wise normalization on the dextramer signal associated with each cell, and / or performing pMHC-wise normalization for each cell represented in the dextramer sequence data. Performing cell-wise normalization may include:

number

number

[0273] Method 4300 may include filtering data based on the presence or absence of an α chain or a β chain from the dextramer sequence data based on the TCR data of a single cell at step 4340. Filtering data based on the presence or absence of an α chain or a β chain from the dextramer sequence data based on the TCR data of a single cell may include determining, for each cell represented in the dextramer sequence data, the presence or absence of at least one α chain and at least one β chain based on the TCR sequence data of the single cell, and removing data associated with cells having only an α chain, only a β chain, or multiple α or β chains from the normalized dextramer sequence data based on the presence or absence of at least one α chain and at least one β chain.

[0274] The method 4300 may include, at step 4350, identifying data remaining in the normalized filtered dextramer sequence data that is associated with reliable TCR-pMHC binding events.

[0275] Method 4300 may further include training a predictive model based on the data remaining in the normalized filtered dextramer sequence data. Method 4300 may further include predicting the binding state of the newly presented receptor sequence with the trained predictive model.

[0276] In one embodiment, the ICON module 108 and / or the prediction module 110 may be configured to perform method 4400, shown in FIG. 44. Method 4400 may be performed in whole or in part by a single computing device, multiple electronic devices, and the like. Method 4400 may include, at step 4410, performing TCR-pMHC binding specificity data normalization on the dextramer sequence data to identify a plurality of TCR-pMHC binding events. Performing TCR-pMHC binding specificity data normalization on the dextramer sequence data to identify a plurality of TCR-pMHC binding events may include some or all of method 4200 and / or method 4300.

[0277] Method 4400 may include, in step 4420, determining a training dataset including a plurality of TCR sequences based on the normalized dextramer sequence data, each TCR sequence associated with a binding affinity. Determining a training dataset including a plurality of TCR sequences based on the normalized dextramer sequence data, each TCR sequence associated with a binding affinity, may include determining, for each TCR sequence of the plurality of TCR sequences, a paired αβ chain CDR3 amino acid sequence, a V gene identifier, and a J gene identifier, and encoding, for each TCR sequence of the plurality of TCR sequences, the paired αβ chain CDR3 amino acid sequence, the V gene segment sequence, and the J gene segment sequence into a one-dimensional input vector. Encoding the paired αβ chain CDR3 amino acid sequence for each TCR sequence of the plurality of TCR sequences includes converting each alphabetical representation of the amino acids into a numeric representation of the amino acids. Encoding the V gene identifier and the J gene identifier for each TCR sequence of the plurality of TCR sequences includes one-hot encoding to generate taxonomic and discrete representations of gene names in a computational space.

[0278] The method 4400 may further include clustering the one-dimensional input vector into one or more clusters. Clustering the one-dimensional input vector into one or more clusters includes applying a KNN clustering algorithm to the one-dimensional input vector. The one or more clusters are indicative of connection strength.

[0279] The method 4400 may include determining a plurality of features for a predictive model based on the plurality of TCR sequences at step 4430. The predictive model may include a weighted binary classifier or a convolutional neural network (CNN).

[0280] Method 4400 may include, at step 4440, training a multi-feature predictive model based on the first portion of the training dataset. Training the multi-feature predictive model based on the first portion of the training dataset may include training a convolutional neural network (CNN). Training the multi-feature predictive model based on the first portion of the training dataset may include applying a class-weighted cost function.

[0281] The method 4400 may include, at step 4450, testing the predictive model based on a second portion of the training dataset.

[0282] The method 4400 may include, at step 4460, outputting a predictive model based on the testing.

[0283] Method 4400 may further include submitting the unknown TCR sequence to the trained prediction model and predicting the binding affinity with the trained prediction model.

[0284] In one embodiment, the ICON module 108 and / or the prediction module 110 may be configured to perform method 4500, shown in FIG. 45. Method 4500 may be implemented, in whole or in part, by a single computing device, multiple electronic devices, and the like. Method 4500 may include, in step 4510, submitting unknown TCR sequences to a trained prediction model, and training the trained prediction model based on a training data set resulting from TCR-pMHC binding specificity data normalization. Method 4500 may also include, in step 4510, performing TCR-pMHC binding specificity data normalization on the dextramer sequence data to identify a plurality of TCR-pMHC binding events. Performing TCR-pMHC binding specificity data normalization on the dextramer sequence data to identify a plurality of TCR-pMHC binding events may include some or all of method 4200 and / or method 4300.

[0285] Method 4500 may include predicting the binding affinity with the trained predictive model at step 4520. The predictive model may include a weighted binary classifier or a convolutional neural network (CNN).

[0286] Method 4500 may include determining a training dataset based on the normalized dextramer sequence data, the training dataset including a plurality of TCR sequences, each TCR sequence associated with a binding affinity. The training dataset can include a plurality of TCR sequences, each TCR sequence associated with a binding affinity. The training dataset can include paired αβ chain CDR3 amino acid sequences, V gene identifiers, J gene identifiers, and binding affinities (e.g., yes / no).

[0287] Method 4500 may include training a multi-feature predictive model based on a first portion of the training dataset. Training the multi-feature predictive model based on the first portion of the training dataset includes training a convolutional neural network (CNN). Training the multi-feature predictive model based on the first portion of the training dataset includes training a convolutional neural network (CNN) having a single translation-invariant layer applied to each TCR sequence, followed by three fully connected convolutional layers in a final output layer. Training the multi-feature predictive model based on the first portion of the training dataset includes applying a class-weighted cost function. Training a multi-feature predictive model based on a first portion of the training dataset includes training a neural network by embedding one-hot coded V and J genes of each chain of the TCR sequence via the learned embeddings, and concatenating these embeddings together with the output of a convolutional neural network for each CDR3 to provide the embedded CDR3 and form a 1D numeric vector representing the TCR, followed by passing each numeric TCR sequence through a final fully concatenated layer.

[0288] In one embodiment, the ICON module 108 and / or the prediction module 110 may be configured to perform method 4400, shown in Figure 44. Method 4400 may be performed in whole or in part by a single computing device, multiple electronic devices, and the like. Method 4400 may include, at 4601, receiving single cell sequence data, dextramer sequence data, and T cell receptor (TCR) sequence data of the single cell.

[0289] The method 4400 may include, in step 4602, determining the number of genes for each cell represented in the dextramer sequence data based on the sequence data of the single cell.

[0290] The method 4400 may include, in step 4603, removing data from the dextramer sequence data that are associated with cells where the number of genes is outside a gene threshold range.

[0291] The method 4400 may include, in step 4604, determining, for each cell represented in the dextramer sequence data, a fraction of mitochondrial gene expression based on the single cell sequence data.

[0292] The method 4400 may include, at 4605, removing from the dextramer sequence data data associated with cells in which the fraction of mitochondrial gene expression exceeds a gene expression threshold.

[0293] The method 4400 may include, at 4606, determining selected dextramer sequence data based on the dextramer sequence data, the selected dextramer sequence data including selected test dextramer sequence data and negative control dextramer sequence data.

[0294] The method 4400 may include, at 4607, determining a maximum negative control dextramer signal for each cell represented in the dextramer sequence data based on the negative control dextramer sequence data.

[0295] The method 4400 may include, at 4608, determining a maximum sorted dextramer signal based on the sorted test dextramer sequence data for each cell represented in the dextramer sequence data.

[0296] The method 4400 may include, at 4609, estimating the dextramer binding background noise based on the maximum negative control dextramer signal and the maximum selected dextramer signal.

[0297] The method 4400 may include, at 4610, determining, for each cell represented in the dextramer sequence data, the presence or absence of at least one alpha chain and at least one beta chain based on the TCR sequence data of the single cell.

[0298] Method 4400 may include, at 4611, removing data from the dextramer sequence data that are associated with cells having only an alpha chain, only a beta chain, or multiple alpha or beta chains based on the presence or absence of at least one alpha chain and at least one beta chain.

[0299] Method 4400 may include, at 4612, determining, for each dextramer that binds to a given cell represented in the dextramer sequence data, a ratio of the dextramer signal within the cell to the sum of all dextramers that bind to the cell (a measure of dextramer binding specificity for the cell). Determining, for each dextramer that binds to a given cell represented in the dextramer sequence data, a ratio of the dextramer signal within the cell to the sum of all dextramers that bind to the cell may include: th T cell binding th For dextramer, the dextramer signal E with background noise subtracted ij and determining

number

[0300] Method 4400 may include, at 4613, for each dextramer that binds to a given TCR clonotype of each cell represented in the dextramer sequence data, determining the fraction of T cells within the clone that binds to the particular dextramer (a measure of dextramer binding specificity for the clonotype to which the cell belongs). Determining the fraction of T cells within the clone that binds to the particular dextramer for each dextramer that binds to a given TCR clonotype of each cell represented in the dextramer sequence data may include: th T cell TCR clonotypek i Determining the clonotype k that binds to dextramers i The number of T cells belonging to T kij and determining

number

[0301] Method 4400 may include, at 4641, determining, for each dextramer that binds to a given cell represented in the dextramer sequence data, a corrected dextramer signal associated with each dextramer that binds to the cell based on the measured dextramer binding specificity for the cell and the measured dextramer binding specificity for the clonal type to which the cell belongs. Determining, for each dextramer that binds to a given cell represented in the dextramer sequence data, a corrected dextramer signal associated with each dextramer that binds to the cell based on the measured dextramer binding specificity for the cell and the measured dextramer binding specificity for the clonal type to which the cell belongs may comprise S ij =E ij (RC ij ) 2 RT kjBy evaluating i th T cell binding th This may include determining a corrected dextramer signal for the dextramer.

[0302] Method 4400 may include, for each cell represented in the dextramer sequence data, performing cell-wise normalization on the dextramer signal associated with each cell.

[0303] The method 4400 may include, at 4615, performing pMHC-wise normalization for each cell represented in the dextramer sequence data.

[0304] The method 4400 may include, at 4616, identifying remaining data in the normalized dextramer sequence data as associated with reliable TCR-pMHC binding events based on a threshold value.

[0305] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the method and compositions described herein which equivalents are intended to be encompassed by the following claims.

Claims

1. performing computational normalization of the TCR-pMHC binding specificity data in the dextramer sequence data to distinguish between multiple TCR-pMHC binding events; identifying, by the computer, a plurality of TCR sequences based on the normalized dextramer sequence data; generating, by the computer, a one-dimensional input vector comprising an encoded pair of αβ chain CDR3 amino acid sequences, a V gene segment sequence, and a J gene segment sequence for each TCR sequence of the plurality of TCR sequences; determining, by the computer, a training data set including at least the one-dimensional input vector; determining, by the computer, a plurality of characteristics for a machine learning model based on the plurality of TCR sequences; generating the machine learning model by the computer; training the machine learning model by the computer based on the training dataset and according to the plurality of characteristics; A computer-implemented method comprising:

2. performing, by the computer, normalization of TCR-pMHC binding specificity data in the dextramer sequence data to identify the plurality of TCR-pMHC binding events; receiving, by the computer, single cell sequencing data, the single cell sequencing data including single cell sequence data, dextramer sequence data, and single cell T cell receptor (TCR) sequence data; filtering, by the computer, data associated with low-quality cells from the dextramar sequence data by removing data associated with cells whose number of genes is outside a gene threshold range or whose fraction of mitochondrial gene expression exceeds a gene expression threshold based on the single cell sequence data; For each cell represented in the dextramer sequence data, subtracting by the computer a measure of background noise from the dextramer signal associated with each cell; filtering data from the dextramer sequence data by the presence or absence of alpha or beta chains by removing data associated with cells having only alpha chains, only beta chains, or multiple alpha or beta chains based on the TCR data of a single cell; identifying by said computer the data remaining in the filtered dextramer sequence data as said plurality of TCR-pMHC binding events; The method of claim 1 , comprising:

3. Generating by the computer a one-dimensional input vector comprising encoded paired αβ chain CDR3 amino acid sequences, V gene segment sequences, and J gene segment sequences for each TCR sequence of the plurality of TCR sequences, comprises: determining, by the computer, the paired αβ chain CDR3 amino acid sequences, the V gene segment sequence, and the J gene segment sequence for each TCR sequence of the plurality of TCR sequences; encoding the paired αβ chain CDR3 amino acid sequences, the V gene segment sequence, and the J gene segment sequence for each TCR sequence of the plurality of TCR sequences into a one-dimensional input vector by the computer; The method of claim 1 , comprising:

4. The method described in claim 3, wherein for each TCR sequence of the plurality of TCR sequences, encoding the paired αβ chain CDR3 amino acid sequences by the computer includes converting the alphabetical representation of each amino acid into a numerical representation of the amino acid by the computer.

5. The method of claim 3, wherein for each TCR sequence of the plurality of TCR sequences, the computer encoding the V gene segment sequence and the J gene segment sequence includes one hot encoding to generate a categorical and separate representation of gene names in a computational space.

6. The method of claim 3, further comprising clustering the one-dimensional input vector into one or more clusters by the computer.

7. The method described in claim 6, wherein clustering the one-dimensional input vector into one or more clusters by the computer includes applying a KNN clustering algorithm to the one-dimensional input vector by the computer.

8. The method described in claim 6, wherein the one or more clusters are indicators of binding strength.

9. The method of claim 1, wherein the trained machine learning model includes a weighted binary classification index or a convolutional neural network (CNN).

10. The method of claim 1, wherein training the machine learning model according to the plurality of characteristics based on the training dataset comprises training a neural network by embedding one-hot coded V and J genes of each chain of the TCR sequence via learned embeddings, and concatenating these embeddings with the output of a convolutional neural network for each CDR3 fed with the embedded CDR3s to form a 1D numeric vector representing the TCR, and subsequently passing each numeric TCR sequence through a final fully connected layer.

11. 10. The method of claim 1, wherein the computer-generated training of the machine learning model according to the plurality of features based on the training dataset comprises the computer-generated applying a class-weighted cost function.

12. The method of claim 1, further comprising testing the machine learning model by the computer based on a test dataset.

13. submitting an unknown TCR sequence to the trained prediction model by said computation; predicting binding affinity using the trained predictive model; The method of claim 1 further comprising:

14. submitting subject TCR sequence data to the machine learning model by the computer; determining a subject TCR binding pattern based on the subject TCR sequence data using the machine learning model; determining by the computer a likelihood that a subject associated with the TCR sequence data has migrated to one or more locations based on a repository of antigen locations and the subject's TCR binding patterns; The method of claim 1 further comprising:

15. generating a TCR binding pattern for the subject based on the remaining data in the normalized dextramer sequence data that correlates with reliable TCR-pMHC binding events; At a subsequent time, receiving by the computer second single cell sequence data, second dextramer sequence data, and second single cell T cell receptor (TCR) sequence data for the subject; determining by the computer a second TCR binding pattern based on the second single cell sequence data, the second dextramer sequence data, and the second single cell TCR sequence data for the subject; identifying, by the computer, the subject based on a comparison of the TCR binding pattern with the second TCR binding pattern for the subject; The method of claim 2 further comprising:

16. Receiving single cell sequence data, dextramers sequence data, and T cell receptor (TCR) sequence data of the single cell by a computer; determining by the computer the number of genes for each cell represented in the dextramers sequence data based on the single cell sequence data; removing, by the computer, data associated with cells whose number of genes is outside a gene threshold range from the dextramers sequence data; determining by the computer a fraction of mitochondrial gene expression for each cell represented in the dextramers sequence data based on the single cell sequence data; removing, by said computer, from said dextramersequence data, data associated with cells whose fraction of mitochondrial gene expression exceeds a gene expression threshold; Based on the dextramer sequence data, Selected dextramers sequence data, including selected test dextramers sequence data and negative control dextramers sequence data; determining unselected dextramer sequence data, including unselected test dextramer sequence data, by said computer; determining by the computer a maximum negative control dextramer signal for each cell represented in the dextramer sequence data based on the negative control dextramer sequence data; determining by the computer the maximum selected dextramer signal for each cell represented in the dextramer sequence data based on the selected test dextramer sequence data; determining by the computer a maximum unsorted dextramer signal for each cell represented in the dextramer sequence data based on the unsorted test dextramer sequence data; said computer estimating a dextramer binding background noise based on said maximum negative control dextramer signal; estimating by the computer a dextramer sorting gate efficiency based on the maximum selected dextramer signal and the maximum unselected dextramer signal; determining by the computer a measure of background noise based on the dextramer binding background noise and the dextramer sorting gate efficiency; for each cell represented in the dextramer sequence data, subtracting by the computer the measure of the background noise from the dextramer signal associated with each cell; for each cell represented in the dextramer sequence data, performing by the computer a cell-wise normalization of the dextramer signal associated with each cell; performing, by said computer, pMHC-wise normalization for each cell represented in said dextramer sequence data; determining, by the computer, for each cell represented in the dextramer sequence data, the presence or absence of at least one alpha chain and at least one beta chain based on the TCR sequence data of the single cell; removing, by the computer, from the normalized dextramer sequence data, data associated with cells having only an alpha chain, only a beta chain, or multiple alpha chains or beta chains based on the presence or absence of the at least one alpha chain and the at least one beta chain; identifying by said computer the data remaining in said normalized dextramer sequence data as being associated with a reliable TCR-pMHC binding event; A computer-implemented method comprising:

17. The method described in claim 16, wherein the gene threshold range is from about 200 genes to about 2,500 genes.

18. The method described in claim 16, wherein the gene expression threshold is approximately 40 percent of the total number of unique molecular identifiers.

19. The method described in claim 16, wherein estimating the dextramer sorting gate efficiency by the computer based on the maximum selected dextramer signal and the maximum unselected dextramer signal includes determining by the computer the maximum difference between the maximum selected dextramer signal and the maximum unselected dextramer signal.

20. The method of claim 16, further comprising training a machine learning model by the computer based on the data associated with reliable TCR-pMHC binding events.

21. The method described in claim 20, further comprising predicting the binding state of a newly presented receptor sequence by the computer according to a trained machine learning model.

22. The method of claim 21, further comprising: presenting target TCR sequence data to the machine learning model by the computer; determining a subject TCR binding pattern based on the subject TCR sequence data using the machine learning model; determining by the computer a likelihood that a subject associated with TCR sequence data has migrated to one or more locations based on a repository of antigen locations and the subject TCR binding patterns; 21. The method of claim 20, further comprising:

23. The method of claim 16, further comprising generating a TCR binding pattern for the subject by the computer based on the data remaining in the normalized dextramer sequence data that is associated with reliable TCR-pMHC binding events.

24. At a subsequent time, receiving by said computer second single cell sequence data, second dextramer sequence data, and second single cell T cell receptor (TCR) sequence data for said subject; determining by the computer a second TCR binding pattern based on the second single cell sequence data, the second dextramer sequence data, and the second single cell TCR sequence data for the subject; identifying, by the computer, the subject based on a comparison of the TCR binding pattern with the second TCR binding pattern for the subject; 24. The method of claim 23, further comprising:

Citation Information

Patent Citations

  • Signs of proliferation and prognosis in gastrointestinal cancer

    JP2010539973A

  • Determination of antigen recognition by barcode labeling of mhc multimers

    JP2017518375A

  • Efficient clustering of immunological entities

    JP2019160261A

  • Systems and methods for sequencing t cell receptors and uses thereof

    WO2018107178A1

  • Ranking system for immunogenic cancer-specific epitopes

    WO2018183980A2